All articles
Cost29 April 2026 · 8 min read · Nevastack Engineering

Cutting inference cost per token without cutting quality

Inference cost is usually the largest recurring line item in a production LLM deployment, and it is also the most controllable one once a system is live. This post covers the levers that reduce cost per token meaningfully without degrading output quality, along with the ones that sound promising but rarely deliver in practice.

Quantization is the highest-leverage change

Running a model at 8-bit or 4-bit precision instead of full precision typically reduces GPU memory footprint substantially and increases the number of concurrent requests a single GPU can serve, which lowers cost per token directly. Quality impact varies by model and task, so measure it against your own evaluation set rather than assuming it is negligible.

For most business applications, well-implemented 8-bit quantization is close to lossless, and many models tolerate 4-bit quantization with only minor quality impact on tasks that do not require precise numerical reasoning. Test both before committing to production.

Batching and continuous batching

Serving frameworks that support continuous batching, where new requests join an in-flight batch rather than waiting for the current batch to complete, substantially improve GPU utilization for variable-length generation workloads. This is one of the largest efficiency gains available without touching the model itself.

The tradeoff is a small increase in latency variability, since a request's processing can be interleaved with others. For most applications this tradeoff favors continuous batching, but latency-critical endpoints should measure the actual impact under their traffic pattern.

Caching repeated or similar requests

Exact-match caching for identical prompts is straightforward and effective for applications with repeated queries, such as FAQ-style retrieval. Semantic caching, which matches queries that are similar but not identical, can extend this benefit but requires careful tuning of the similarity threshold to avoid returning a cached answer for a meaningfully different question.

Cache invalidation matters as much as caching itself. A cache that serves stale answers after underlying documents change creates a correctness problem that can be harder to debug than the cost problem caching was meant to solve.

  • Use exact-match caching for high-repetition query patterns first, since it carries no correctness risk
  • Add semantic caching cautiously, with a conservative similarity threshold
  • Tie cache invalidation to the same event that triggers document re-indexing

Model routing by task complexity

Not every request needs the largest available model. Routing simple classification, extraction, or short-answer tasks to a smaller, cheaper model while reserving a larger model for open-ended generation can reduce average cost per request significantly without a measurable quality loss on the simpler tasks.

Building a reliable router requires an evaluation step to confirm the smaller model performs acceptably on its assigned task category. A router that misclassifies task difficulty and sends complex requests to a small model will degrade quality in ways that are easy to miss in aggregate metrics.

What does not reliably help

Aggressive prompt shortening to save input tokens often removes context the model needs, increasing the odds of a wrong or incomplete answer that requires a retry, which negates the token savings. Optimize prompts for clarity and necessary context first, and treat token count reduction as a secondary concern.

Switching to a cheaper model purely on advertised price per token, without evaluating quality on your specific task, frequently increases total cost once retries, follow-up clarifications, and downstream error correction are accounted for. Evaluate cost per successfully completed task, not cost per token in isolation.

Want this applied to your data?

We scope private AI projects in one call and start with a pilot you can evaluate.

Make room for your next idea.

Bring your data, models, and compute into one workspace. Start with a project and build from there.