Most LLM bills are not high because the model is expensive. They are high because the application sends the same tokens over and over, uses a large model for jobs a small one could do, and lets the model write far more than anyone reads. If you want to save token cost, the fix is usually engineering discipline, not a cheaper vendor.
This guide walks through the techniques we apply when a client's AI feature works well but the invoice is growing faster than usage. None of them require a rewrite. Most can be shipped in a sprint, and together they often cut spend substantially without users noticing any change in quality.
Start by measuring cost per task, not cost per month
A monthly invoice tells you almost nothing. You need to know what one unit of useful work costs: one resolved support ticket, one booked appointment, one summarized contract. Without that number you cannot tell whether a change helped or whether traffic simply dropped.
Log these fields on every model call and tag them with a task or feature name:
- Model name and version
- Input tokens, cached input tokens and output tokens
- Latency and number of retries
- The feature, tenant and user that triggered the call
- Whether the task succeeded (resolved, booked, accepted by the user)
Then divide total spend for a feature by successful outcomes. A support assistant that costs very little per call but needs six back-and-forth turns and two tool loops per ticket can easily cost more per resolution than a smarter, pricier model that gets there in two turns. Cost per successful task is the only number that captures that.
Once this is in place, sort features by total spend. In most products, one or two flows account for the bulk of the bill. Optimize those first and ignore the rest until they matter.
Use prompt caching for the parts that never change
Most production prompts have a large static prefix: system instructions, tool definitions, policy text, few-shot examples, sometimes a product catalog. For example, a 4,000-token system prompt sent on every call to an assistant that handles 50,000 messages a day is 200 million input tokens daily before the user has said anything.
The major providers now offer prompt caching. You mark a stable prefix, the provider stores the processed version, and later calls that reuse the exact same prefix are billed at a steep discount and usually return faster. The details differ by vendor (some cache automatically, some need explicit cache markers, and cache lifetimes are often measured in minutes unless you pay for longer), so check the current pricing page for the model you use.
To get real cache hits, structure prompts deliberately:
- Put everything static at the top: system prompt, tool schemas, reference documents.
- Put anything dynamic at the bottom: the user's message, retrieved chunks, timestamps.
- Never inject a timestamp, request ID or user name into the system prompt. One changed character near the start breaks the cache for everything after it.
- Keep tool definitions in a stable order. Sorting them differently per request silently kills caching.
Track the ratio of cached input tokens to total input tokens per feature. If it is low on a flow with a long static prompt, something in your prefix is changing.
Route each request to the smallest model that can do the job
Using the flagship model for everything is the most common source of waste we see. Classification, extraction, language detection, intent routing and short rewrites are well within the ability of small, fast models that cost a fraction as much per token.
A simple routing layer looks like this:
def pick_model(task):
if task.kind in ("classify", "extract", "detect_language"):
return "small-model"
if task.kind == "draft_reply" and task.context_tokens < 3000:
return "mid-model"
return "large-model" # multi-step reasoning, tool use, long context
Start with rules like these, based on task type. Only add a learned router or a "try small, escalate on low confidence" pattern after you have evaluation data showing where the small model fails. Escalation works, but it needs a reliable signal (a schema validation failure, a confidence score, a checker model) or it becomes a second bill on top of the first.
Routing decisions should be tested, not assumed. Run the same evaluation set through each candidate model and compare quality and cost per successful task side by side. Sometimes the bigger model is cheaper overall because it needs fewer retries.
Trim the context and cap the output
Input tokens usually dominate volume, while output tokens are typically priced several times higher per token. Both can be cut.
Conversation history
Chat apps often resend the full history on every turn, so cost grows with every message. Keep the last few turns verbatim and replace older turns with a running summary. For task-oriented assistants (booking, order status), store structured state such as the selected date or order number instead of the transcript that produced it.
Retrieved documents and tool output
RAG pipelines frequently pull ten chunks when two would do. Rerank, keep only the top few, and cap total retrieved tokens per request. Every tool schema you attach is also input on every call, so expose only the tools relevant to the current step, and strip raw API responses before passing them back. A CRM lookup that returns 200 fields when the model needs five is pure waste.
Output length
Set an explicit maximum output length per feature, and tell the model what format you want: "Reply in at most three sentences" or "Return JSON matching this schema, no commentary." A JSON object with four fields is shorter, cheaper and easier to validate than a paragraph that a second step then has to parse.
Reasoning models deserve special attention. Thinking tokens are billed as output, and a model can spend thousands of them on a question that did not need deep reasoning. Where the provider allows it, set a reasoning budget or effort level per task, and keep high-effort reasoning for the flows that actually benefit.
Batch the slow work and cache the repeated answers
Batch APIs
Not every model call is user-facing. Nightly report generation, document tagging, embedding backfills, evaluation runs and bulk translation can all wait minutes or hours. Most major providers offer a batch API that processes requests asynchronously, typically within 24 hours, at a significant discount compared to real-time calls. Move anything that is not on a user's critical path into a queue and submit it in batches. Combined with prompt caching on a shared prefix, large document jobs become much cheaper than the same work done synchronously.
Response caching
Prompt caching reduces the cost of a call. Response caching removes the call entirely. If many users ask the same thing ("What are your opening hours?", "How do I reset my password?"), there is no reason to generate the answer fresh each time.
| Approach | How it works | Good for | Watch out for |
|---|---|---|---|
| Exact-match cache | Hash the normalized input, return the stored answer | Repeated FAQ-style queries, deterministic transforms | Low hit rate on free-text input |
| Semantic cache | Embed the query, return a stored answer if similarity is above a threshold | Paraphrased versions of common questions | Wrong answers if the threshold is loose |
| Precomputed answers | Generate answers offline for known intents | High-volume support flows | Keeping answers current when policies change |
Two rules keep response caches safe. First, scope the cache by tenant and by permissions, so one customer never receives an answer generated from another customer's data. Second, attach an expiry and invalidate on source changes, so a price update does not keep serving last month's figure.
Put guardrails on the bill itself
Agents add a new risk: a loop that keeps calling tools and the model until it hits a limit you forgot to set. Protect yourself with hard caps on steps per task, tokens per task and spend per tenant per day, and alert when a single task exceeds a threshold. If you sell an AI feature on a flat subscription, per-tenant budgets are what stop one heavy user from wiping out the margin on the whole plan.
Review cost per task weekly alongside quality metrics. Teams that treat token spend as a product metric, owned by the same people who own the feature, tend to keep it under control. Teams that leave it to finance discover the problem at the end of the quarter. If your bill is already a concern, a focused cost optimization review across both cloud and model spend is usually the quickest way to find the big wins.
How Softzee can help
We build and tune production AI systems, from voice agents to support assistants, and cost per task is something we instrument from day one. If your AI features are working but the spend is not, talk to our team and we will help you find where the tokens are going.