Home / Guides
LLM cost management and AI FinOps guides
Class1 estimates the monthly cost impact of an AI code change inside the pull request, before it becomes a bill. These guides explain the method — probabilistic forecasting, fit-for-purpose model selection, and calibration against real actuals — and define the vocabulary.
Why estimate before merge
The expensive decision is the architecture, and it is made at review time.
Most AI cost tools are observability dashboards: they chart money that has already been spent. By the time a spike shows up, the callsite, the model choice and the retry policy are already in production and expensive to reverse. Class1 moves the question earlier, to the moment a change is proposed, and answers the only one that matters before merge: what does this do to next month's bill?
Because the answer is a forecast under uncertainty, it is a distribution, not a point. Class1 reports P50 and P90 (and P95 when a budget owner wants a stricter ceiling), decomposes the drivers, and attaches an estimate class that says how much the number should be trusted given the evidence behind it. A gate can fail a pull request whose P90 monthly delta exceeds a declared budget, and it never blocks a change that does not raise cost.
Moving left also aligns incentives. Developers get immediate feedback on how their architectural choices (like context limits or model selection) affect unit economics before they are committed to trunk. Finance gets a predictable, pre-approved risk envelope instead of a surprise invoice at the end of the month. This shifts the dynamic from reactive cost-cutting to proactive engineering.
How the estimate stays honest
Confidence is reported separately from the dollar figure, and it has to be earned.
A P90 number can be mathematically sharp and still rest on immature inputs, so Class1 reports the conditional risk band and the estimate class as two different facts. A brand-new install ships at Class 5 and only moves toward Class 1 as real estimate-to-actual pairs accumulate through reconciliation against actual bills.
The basis is frozen and effective-dated, so the same diff yields the same estimate, and as that basis ages the accuracy band widens by design. If you want the evidence behind the numbers, the trust page lists the tests, assumptions and open gaps, and the ledger shows the frozen price, spec, capability and actuals data.
Deeper context
The Class1 engine
At the heart of our platform lies a strict adherence to the principles of cost engineering, adapted for the unpredictable nature of Large Language Models. Unlike traditional software where compute costs are deterministic and tied directly to traffic, AI feature costs are highly variable. They depend on model selection, input token length, output token variability, the frequency of retries due to hallucinations or malformed JSON, and dynamic fallback chains.
To capture this complexity, Class1 employs Monte Carlo simulations using Common Random Numbers (CRN). By modeling the system before and after a pull request across thousands of synthetic scenarios, we isolate the true cost delta of your architectural change from background noise. This approach allows us to assign an AACE (Association for the Advancement of Cost Engineering) estimate class to the forecast, communicating both the expected cost and the confidence interval.
Furthermore, this estimation does not happen in a vacuum. The Blue Book ledger records historical execution data, creating a closed-loop calibration system. As your team merges changes and actuals are observed in production, the engine updates its actuarial tables. This continuous feedback loop ensures that our pre-merge budget gates remain accurate and actionable, preventing catastrophic cost overruns before they reach production while maintaining developer velocity.
Search-demand library