Imagine you’re asked to automate loan approvals. The policy fits in one paragraph:
A loan is approved if credit score is 700 or higher, age is between 21 and 65, and income is at least €10,000. Scores from 650 to 699 go to manual review.
With an LLM API at hand, the first design almost writes itself: send the policy and the applicant’s data to the model and ask for a verdict. It works in the demo, and it takes an afternoon to build.
Then real traffic arrives. Every application becomes a model call, every answer takes a second or two, and the bill grows with each new applicant. Once in a while the same applicant gets two different answers on two different days, which is hard to explain to a customer and even harder to explain to a regulator.
To be clear, most software doesn’t work like this. Nobody routes a checkout or a login form through a language model, and a developer who knows the loan policy upfront would simply write it as code. The trap appears in the harder cases, where the logic isn’t known when the software is written. The same feeling shows up with messy inputs, long multi-step workflows and anything sold as “agentic”: the problem looks too complex for ordinary code, so calling a model on every request seems like the only option. That impression is wrong, and it’s worth questioning every time.
We ran into exactly this while building SmartDecision, a rules engine we’re working on at Wingravity. Its users write their own policies in plain language, change them whenever the business changes, and can have dozens of them. No developer can code those in advance, which makes a model call per request look unavoidable. We ended up somewhere else. The model reads each policy once, when it’s written, and turns it into a decision table: credit score, age and income go in, a decision comes out, one row per rule. A person checks the table and pins a few test cases against it. After that, every loan check is a plain table lookup that answers in milliseconds, gives the same result every time and costs next to nothing.
Writing the table also turned out to be easy work for a model. A 4-billion-parameter model on an ordinary server handles it well, because it only has to get the table right once, with a human checking the result.
That design choice is the core idea of this post: most teams spend AI at the wrong moment. They call a model every time a user does something, when they could have spent that intelligence once, while building the product, and shipped plain, deterministic software that runs cheaply forever.
Runtime inference is the biggest cost in AI
To set the scope: everything below is about the AI stack itself, so salaries and general hosting stay out of the comparison.
An AI-powered product today usually includes much more than an LLM behind an API. There are vector databases, embedding pipelines, orchestration frameworks, guardrails, observability and eval tooling. Look at a typical LangChain app or any of the elaborate agent harnesses out there and you’ll count a dozen moving parts.
Most of those parts are fixed or near-fixed infrastructure costs. You pay roughly the same for them whether you have a hundred users or a hundred thousand. Model calls grow with every user, every request and every month, and they keep growing for as long as the product exists. In an AI-powered product with real traffic, inference is by far the largest line on the bill.
That makes timing the important question: at what point in the product’s life do you pay for the AI you use?
Two places to spend AI
There are two moments in a product’s life where you can use a model:
Build time. AI writes code, generates test cases, produces data, drafts content, derives rules or labels examples. A human reviews the output, and it gets committed, stored or shipped as ordinary software.
Runtime. The shipped product calls a model on every interaction. The user clicks, the model thinks, the user waits.
The costs of these two behave very differently.
Build-time inference is bounded. Even heavy agent-assisted development has an end. You spend it while making the product, and then you stop. If you generate ten thousand product descriptions, you pay for ten thousand product descriptions, once.
Runtime inference is unbounded. It scales with usage for as long as the product exists. If you generate a product description on every page view, you pay for it every time someone looks at the page. Tony Perez puts it well: “If you let AI run your operations live, you are renting intelligence on every single heartbeat.”
Pay attention to who benefits from that rent. As Trey Reeme of Trabian points out, when a vendor’s business model is tied to token usage, more agent activity means more revenue. Every time an agent runs a workflow, the work goes back through the model and the meter keeps ticking. The pitch you hear most often, that agents will do the work for you in production indefinitely, happens to be the most profitable one for the people selling tokens.
Runtime inference also brings costs beyond the invoice:
- Latency. A lookup takes milliseconds, while a model call often takes seconds.
- Non-determinism. The same input can produce different outputs, which makes bugs hard to reproduce. In regulated industries like banking, a model improvising in production is a compliance risk on top of a technical one.
- Harder testing. A function can be unit-tested with exact expectations, whereas model output can only be sampled and scored.
- Vendor dependency. Every request now depends on a third-party API being up, fast and priced the way it was last month.
Once build-time output is reviewed and committed, it behaves like any other code or data, and none of these problems apply.
The common mistakes
These are the patterns we see most often, where runtime inference is doing work that could have happened once, at build time.
1. Classifying and routing with an LLM call
Ticket routing, intent detection, spam filtering, tagging and sentiment analysis are classic classification problems with a known, finite set of outputs. Asking a large model to handle them on every request is like hiring a lawyer to sort your mail.
A better approach is to use AI at build time to write a rules-based classifier, or to help label a dataset you can train a small, cheap model on. You keep the benefit of the model’s understanding and stop paying for it on every request.
Tony Perez describes doing exactly this for customer support. He used AI to work through 1,000 manually reviewed tickets and built a support engine out of them. Today it handles about 80% of tickets on its own, answers only when its confidence is above 90%, and leaves the harder 10 to 20% to a person. In his words: “I did not build a dependency on AI. I used AI to build infrastructure that outlives it.”
2. Generating content on the fly
Product descriptions, onboarding emails, page copy, help text and notification messages often depend on a known set of inputs. In that case they can be pre-generated, reviewed by a human and stored.
The review step may matter even more than the savings. Pre-generated content can be checked before anyone sees it, while content generated live goes straight to the user.
3. Agents deciding steps that were never in question
A lot of “agentic” workflows ask a model to decide what to do next, at every step, for processes that are completely predictable. Fetch the order, check the status, send the email, update the record. The model picks the same sequence every single time, and you pay for each of those decisions.
If the workflow is predictable, write the workflow. Let AI help you write it as normal code, with normal branches and normal error handling, and keep the model for the one step that actually needs judgment, if there is one.
4. Harnesses that multiply every request
This one is easy to miss because it hides inside the tooling. In a typical agent harness, one user request can turn into:
- a planning call,
- several tool-selection calls,
- retries on malformed output,
- a reflection or validation step,
- a summarization pass,
and each of those re-sends a large context. A single answer on the user’s screen can easily cost ten calls and tens of thousands of tokens.
Frameworks meant to make AI apps easier to build often make them far more expensive to run. Much of that orchestration logic could have been written once, at build time, as regular code.
5. Using a large model for a small job
Extracting a date from a sentence, checking whether a string is an address or translating a fixed set of UI labels are tasks that frontier models are wildly overqualified for. A small model, a local model or a regular expression an AI wrote for you will usually do just as well, at a fraction of the cost and latency.
Ramp’s engineering team put numbers on this: for the same $100K, you can buy 5 billion tokens of the smartest model on the market, or 210 billion tokens of an open-weight model, 42 times more. Their rule of thumb is simple. Well-understood, well-scoped work gets the cheapest model that passes your quality bar, and the frontier model is reserved for novel, ambiguous or high-stakes problems. They also explain why that rarely happens by default: “Nobody has ever been fired for buying the frontier.”
What this looks like in numbers
Let’s go back to SmartDecision and do a back-of-the-napkin calculation: the way it works today, against a version that asks an LLM for every verdict. I’ll count tokens instead of money, because prices differ a lot between providers and change every few months.
The assumptions:
- A manager owns 2 to 3 projects with 10 to 15 decisions each, so roughly 30 decisions.
- Drafting one decision table takes about 3,000 tokens: the instructions with the table schema and examples, the policy itself and the generated table. Allowing three attempts per decision for rewording and redrafting brings it to about 10,000 tokens per decision.
- Resolving a decision with an LLM means sending the instructions, the policy and the inputs, and getting a short answer back. Kept lean, that’s about 1,000 tokens per call.
- Each decision is called 1,000 times a day by the applications that use it. For pricing or eligibility checks, that’s a quiet day.
Build time, per manager: 30 decisions × 10,000 tokens = 300,000 tokens, paid once. The number only moves when someone writes or rewrites a policy.
Runtime with an LLM, per manager: 30 decisions × 1,000 calls × 1,000 tokens = 30 million tokens a day, or around 900 million tokens a month. It keeps growing with every new decision, every new application that calls one, and every month the product is in use.
In other words, the runtime version burns through the entire build-time budget after 300 resolutions. At 30,000 calls a day, that happens in about fifteen minutes. A company with ten people owning policies would be looking at 300 million tokens a day, while the table-based version answers the same calls in milliseconds on a server it already runs.
Your own numbers will be different, but the gap between a one-off cost and a recurring one will look much the same.
Build heavy, run lean
AI makes build time incredibly cheap. Work that used to sit on the “nice to have, if we ever find the time” list now takes an afternoon:
- more tests and better coverage,
- handling for edge cases you would have ignored,
- internal tooling and admin panels,
- seed data, fixtures and realistic test datasets,
- migration scripts, validation rules and documentation.
So use it aggressively, upfront. Spend the tokens while you’re building, and let the AI produce more code, tests, data and rules than your team could ever write by hand. Then ship a product that runs on plain software, with fewer API calls, smaller infrastructure, a lighter maintenance load and far fewer surprises at 3 a.m.
Trabian’s COO describes the same idea with a thesis from his colleague Matt: “Agents should build as many systems as possible and run as few systems as possible.” In their platform, a subject-matter expert describes a workflow in plain language, an agent drafts it, a human refines and approves the exact version, and from then on it runs on deterministic infrastructure with retries, human checkpoints and a full audit trail. If a step hasn’t been built yet, the system halts instead of improvising.
This is how we work at Wingravity. We lean heavily on AI during development, and we’re deliberate about what ends up calling a model in production. Our clients get the speed of AI-assisted delivery without signing up for an AI bill that grows every month.
When runtime inference is worth it
Some problems genuinely need a model at runtime:
- Open-ended user input, where you can’t predict what people will type or ask.
- Real conversation, where the model’s response is the product itself.
- Unbounded input spaces, such as summarizing arbitrary documents or answering questions over content that changes constantly.
For these, calling a model at runtime is the right tool, and the cost is worth paying.
A simple test you can apply to any model call in your product:
Could I write down the answer ahead of time, or the rules that produce it?
If yes, do that. Use AI to help you write it, then ship the result.
If the honest answer is no, keep the call, and make it as small, cheap and infrequent as you can. Pick the cheapest model that does the job, and if the user doesn’t need the answer instantly, use a provider’s flex or batch tier. Ramp notes these are often around half the price of standard speed.
Audit your model calls
Open your codebase and find every place your product calls a model. For each one, ask:
- Does this run on every request, or once?
- Is the set of possible outputs actually finite?
- Could this have happened at build time, with the result stored or turned into code?
- If it must happen at runtime, does it need this particular model, or would a smaller one do?
You’ll likely find that a large share of your runtime inference is doing work that could have been done once. Move that work to build time, and you’ll end up with a product that’s faster, cheaper, easier to test and far more predictable.
Further reading
These articles shaped my thinking on this topic and are well worth your time:
- AI Builds the Tooling. AI Doesn’t Run the Tooling. by Tony Perez
- A Future Enabled by AI, Not Run by It by Trey Reeme, Trabian
- You’re Spending Too Much on AI. You’re Also Using Too Little. by Anand Kuchibotla, Kedar Thakkar and Rahul Sengottuvelu, Ramp
Keep reading
CTO @ Wingravity and Co-Founder of Slashscore. Passionate about building tools that empower developers and teams. Have a question or want to collaborate? Reach out via our contact form.



