Blue Book ledger

Private LLM pricing intelligence and rate-basis methodology

Class1 keeps a frozen, effective-dated, provenance-aware data layer so every PR estimate can be reproduced and audited.

10,980 pricing rows effective 2026-09-29
3,626 structured price rows cache, batch, tiers
11,906 spec rows litellm plus models.dev
6,637 price trend models historical token index
1,258 cloud instances 122 dated versions

Why the ledger matters

The hard part is not displaying a price. The hard part is knowing which price was used, when, and why.

LLM vendors change prices, aggregators rename models, context tiers appear, cache discounts differ by provider, and frontier model releases can reset capability assumptions overnight. A PR estimate that uses a live feed without freezing the basis is difficult to reproduce later. Without a frozen basis, two identical pull requests evaluated on different days could yield completely different cost forecasts, leading to confusion and mistrust between engineering and finance.

Class1 treats the pricing basis like an estimator treats a cost book. Raw artifacts are preserved with provenance, normalized into silver rows, and curated into a gold snapshot. The estimate can say which basis it used, how old that basis is, and which items were measured, inferred, assumed, excluded, or unpriced. This rigorous approach ensures that every budget decision is defensible and traceable.

That makes the Blue Book commercially important. A public basis proves the open-core product. A private basis becomes the paid pilot: customer-specific rates, private actuals, internal model allowances, team allocation rules, and variance reports that cannot be copied from a public dashboard. It transitions AI cost management from an observational exercise to an active governance framework.

Medallion discipline

Bronze raw source. Silver normalizer. Gold estimate basis.

Prices, specs, capability, cloud cost, actuals, carbon, water, and materials all live as frozen snapshots.

The product can refresh the basis, but an estimate never silently depends on a live feed. Staleness is modeled as estimate decay.

What a buyer gets from the Blue Book

A credible estimate needs more than a token price table.

Structured ratesCache reads, cache writes, batch discounts, reasoning tokens, and context tiers can change the cost-per-call without any code changing the model name.
Capability joinsSWE-Bench Verified coding grades let the product explain why the cheapest per-token model may be expensive per completed task.
Actuals bridgeFOCUS and provider usage rows give the calibration loop a place to land after the PR has shipped.
Cloud-side basisVector databases, managed endpoints, storage, egress, and GPU commitments belong in the fully loaded AI cost discussion.
Footprint basisCarbon intensity, water scarcity, WUE, and materials coefficients give the environmental report traceable assumptions.
pricing.json10,980 model rows, USD per 1M tokens
pricing_structure.json3,626 structured rates for cache, batch, tiers and related pricing shapes
price_index.json6,637 models with historical token-price trend data
capability.json35 distinct models with SWE-Bench Verified coding grades (39 raw rows incl. dated variants)
spec_sheet.json11,906 model metadata records, with raw source retention
actuals_index.json8 public actuals sources feeding the cloud-cost basis

Capability sample

Real SWE-Bench Verified coding grades, one row per model.

ModelProviderSWE-Bench scoreDate
claude-opus-4-5anthropic79.2%2025-12-05
Doubao-Seed-Codeunknown78.8%2025-09-28
gemini-3-pro-previewgoogle77.4%2025-11-20
claude-sonnet-4-20250514anthropic76.8%2025-08-04
gpt-5openai75.6%2025-09-01
claude-sonnet-4-5anthropic74.8%2025-11-03
claude-sonnet-4.5anthropic73.8%2025-11-03
claude-4-opus-20250514anthropic73.2%2025-05-22
gemini-3-5-flashgoogle71.8%2026-09-01
gpt-5-2025-08-07openai71.8%2025-08-07
kimi-k2-0905-previewmoonshot71.2%2025-10-14
claude-3-7-sonnet-20250219anthropic70.4%2025-05-15

Deeper context

The Class1 engine

At the heart of our platform lies a strict adherence to the principles of cost engineering, adapted for the unpredictable nature of Large Language Models. Unlike traditional software where compute costs are deterministic and tied directly to traffic, AI feature costs are highly variable. They depend on model selection, input token length, output token variability, the frequency of retries due to hallucinations or malformed JSON, and dynamic fallback chains.

To capture this complexity, Class1 employs Monte Carlo simulations using Common Random Numbers (CRN). By modeling the system before and after a pull request across thousands of synthetic scenarios, we isolate the true cost delta of your architectural change from background noise. This approach allows us to assign an AACE (Association for the Advancement of Cost Engineering) estimate class to the forecast, communicating both the expected cost and the confidence interval.

Furthermore, this estimation does not happen in a vacuum. The Blue Book ledger records historical execution data, creating a closed-loop calibration system. As your team merges changes and actuals are observed in production, the engine updates its actuarial tables. This continuous feedback loop ensures that our pre-merge budget gates remain accurate and actionable, preventing catastrophic cost overruns before they reach production while maintaining developer velocity.

Proprietary basis

The Blue Book is private

The public Ledger explains the coverage, provenance discipline and role of the Class1 Blue Book without publishing the underlying pricing records, source archive, capability mappings or historical snapshots.

Pilot outputs expose only the evidence needed to review a specific estimate. The proprietary rate basis remains part of the Class1 service and is not offered as a public download.