Seeded stochastic tests
Tests assert invariants on seeded Monte Carlo runs: percentile ordering, monotonicity, exact-zero self-delta, and medallion round-trips.
Trust
The product declares what is measured, inferred, assumed, excluded, or unpriced. That honesty is a sales asset.
Trust posture
Class1 does not sell certainty. It sells a disciplined estimate with a declared basis. The report shows the conditional risk band and the maturity class separately, because those are different uncertainties. A P90 number can be mathematically sharp and still depend on immature inputs; the buyer deserves to see both facts. This separation of epistemic (knowledge-based) and aleatory (random) uncertainty is standard practice in heavy civil engineering, and we believe it belongs in software engineering.
The product also refuses to confuse events with validation. Usage events help measure retry tails, fallback rates, output distributions, and demand spikes. Estimate-actual pairs validate forecasts. That distinction is what prevents a customer with lots of telemetry but few closed loops from claiming Class 1 maturity too early. Without actuals, a model is just an educated guess.
The trust page exists because the product will be used in merge decisions. If a gate can block a PR, the team needs to know how it behaves, when it does nothing, which assumptions are in scope, which data is frozen, and which open items remain outside the current product. Transparency is the only way to earn the right to block a deployment.
Tests assert invariants on seeded Monte Carlo runs: percentile ordering, monotonicity, exact-zero self-delta, and medallion round-trips.
Forecasts use frozen, effective-dated snapshots and past estimates are never silently rewritten. As the price or capability basis ages, estimate decay widens the accuracy band and steps the class back toward Class 5 (one step per staleness half-life) - Class1's extension of AACE classification, so confidence cannot outlive its basis.
The P90 is contingency: known-unknowns the engine can model from the diff and workload (retry storms, output drift, context growth, fallback, demand spikes). It is not management reserve for unknown-unknowns - future scope growth or new features re-enter only through the actuals loop. You approve against a conditional risk band, not a total-project allowance.
Ingest preserves provider events in raw_payload and keeps cached cost components disjoint.
Events measure factor shapes. Estimate class improves only from scarce estimate-actual pairs.
The policy gate cannot block a PR whose P90 delta is zero or negative.
GOAL.md and NEEDS_HUMAN.md keep L2/L3 items out of marketing fantasy.
Evidence checklist
Current open items
Proof command
GOAL checkboxes are not enough. A task counts only when the acceptance test exists and the suite is green.
make test
make demo
make capability
make footprint
Deeper context
At the heart of our platform lies a strict adherence to the principles of cost engineering, adapted for the unpredictable nature of Large Language Models. Unlike traditional software where compute costs are deterministic and tied directly to traffic, AI feature costs are highly variable. They depend on model selection, input token length, output token variability, the frequency of retries due to hallucinations or malformed JSON, and dynamic fallback chains.
To capture this complexity, Class1 employs Monte Carlo simulations using Common Random Numbers (CRN). By modeling the system before and after a pull request across thousands of synthetic scenarios, we isolate the true cost delta of your architectural change from background noise. This approach allows us to assign an AACE (Association for the Advancement of Cost Engineering) estimate class to the forecast, communicating both the expected cost and the confidence interval.
Furthermore, this estimation does not happen in a vacuum. The Blue Book ledger records historical execution data, creating a closed-loop calibration system. As your team merges changes and actuals are observed in production, the engine updates its actuarial tables. This continuous feedback loop ensures that our pre-merge budget gates remain accurate and actionable, preventing catastrophic cost overruns before they reach production while maintaining developer velocity.