Syntax Station

Insights / AI & Agents

How to Cut LLM Costs Without Hurting Quality

Practical techniques that reduce AI running costs in production: model routing, prompt caching, shorter context, batching and knowing when a smaller model is good enough.

By Syntax Station Engineering · · 3 min read

Key takeaways

  • Measure cost per completed task, not per token. It is the number that maps to business value.
  • Route easy requests to small models and reserve large models for hard ones.
  • Most cost hides in context: long system prompts, full chat histories and too many retrieved passages.
  • Use provider prompt caching and batch APIs for anything that is not time-critical.

AI features have a habit of looking affordable in the prototype and expensive in production. The prototype handles a few hundred requests. Production handles hundreds of thousands, each carrying more context than anyone realized.

The good news is that most AI bills can be cut substantially without users noticing any change in quality. Here is where to look, in the order we usually check.

Start by measuring the right thing

Track cost per completed task: per resolved ticket, per processed invoice, per generated report. Token prices change often and differ by provider, while cost per task tells you whether the feature makes economic sense. Log tokens in and out for every call, tagged by feature and step, so you can see where the money goes.

1. Trim the context

Input tokens are usually the largest part of the bill.

  • System prompts. Prompts grow as teams patch edge cases. Review them quarterly. Remove duplicated rules and examples that tests show are no longer needed.
  • Conversation history. Do not resend the whole chat every turn. Summarize older turns or keep only what is relevant.
  • Retrieved passages. Sending twenty passages "to be safe" multiplies costs. A good reranking step lets you send five that matter.

2. Route requests by difficulty

Not every request needs the best model. A common pattern:

Request typeModel tier
Classification, routing, extractionSmall, fast model
Standard answers and draftsMid-sized model
Complex reasoning, long documents, codeLarge model

A lightweight classifier, or simple rules, decide which tier handles each request. Many teams find most traffic can be served by the cheaper tiers.

3. Use caching

  • Prompt caching. Major providers discount repeated prompt prefixes. Put stable content (instructions, tool definitions, reference documents) at the start of the prompt and variable content at the end to benefit.
  • Response caching. Identical or near-identical questions, such as "what are your opening hours?", can be answered from a cache without calling a model at all.
  • Embedding caching. Never re-embed a document that has not changed.

4. Batch what is not urgent

Overnight report generation, bulk classification and data enrichment do not need instant answers. Batch APIs typically cost significantly less than real-time calls, and queues let you smooth load and stay inside rate limits.

5. Control output length

Ask for structured outputs (JSON with defined fields) instead of free prose when the result feeds into software. Set maximum output lengths. Shorter, structured output is cheaper, faster and easier to validate.

6. Consider self-hosting at scale

Open-weight models have become strong for many tasks. Self-hosting can make sense when volume is high and steady, when data must stay in your infrastructure, or when you need a fine-tuned model. Factor in GPU costs, autoscaling, monitoring and the engineers who keep it running.

7. Make cost part of your test suite

Every change to prompts, models or retrieval should run against your evaluation set and report both quality and cost. That way a "small prompt improvement" that doubles cost is caught before it ships. See how to evaluate LLM applications.

A realistic outcome

On a typical support assistant, combining context trimming, routing and caching often cuts running costs by more than half while answer quality, measured on the same test set, stays flat. The exact number depends on where you start, which is why measurement comes first.

Frequently asked questions

Why are my LLM API costs so high?

Usually because every request sends far more context than it needs: long instructions, entire conversation histories and many retrieved documents. Input tokens add up quickly at volume.

Are open-source models cheaper than API models?

Sometimes. Self-hosting removes per-token fees but adds GPU, engineering and operations costs. It tends to pay off at high, steady volume or when data must stay in your own infrastructure.

Does using a smaller model reduce quality?

For classification, extraction and short structured outputs, small models often perform as well as large ones. Test them against your own examples before deciding.

Related reading