COST REPORT
Northwind Q3 cost review
Run 2026-08-18 · completed with warnings · pinned to snap-2026-09-01-a7f3, snap-2026-09-01-c118. Built to be cropped: every figure carries its own grade.
Document Summariser is 110% of the attributed bill and is already on the right model — it needs its reasoning level turned down, not a new one. The obvious move comes out $2,155/month more expensive than doing nothing.
This rests on 1 of 7 workloads with MEASURED evidence, from $5.80 of replay. The rest is modelled and says so, per figure.
Tokens and cost in scope — read from your logs
Your provider's own records, 2026-07-24 to 2026-08-28 · 36 days · 2 accounts. Nothing here was replayed.
Requests
14.8M
model calls in window
Input tokens
36.18B
27.06B fresh · 9.12B cached
Cached input
25%
billed at a tenth of the input rate
Output tokens
2.76B
539.7M of it reasoning
Total tokens
38.94B
input + output
Token cost
$52,450
$43,709/month at this rate · $1.35 per 1M
Counted over the 36-day scope window, so the total is larger than a month. The monthly equivalent above still runs ahead of the $39,000 baseline, and should: this counts every readable call in the scope's accounts, while the baseline prices only the spend discovery could attribute to a workload — the difference is the unattributed traffic declared further down. Cost here is tokens against the published price list, which is why the per-1M rate reconciles with the total beside it. gemini-2.5-flash is excluded from scope and from these counts.
Cost across the scope window
36 days · the dashed line is the daily mean
A monthly figure only means something if the month was ordinary. Days well above the mean are what a rate hides — $1456.95 a day here, and the run is priced from the whole window rather than any one of them.
Where the spend sits
By model, across the scope window
- gemini-2.5-pro$37,241
21.06B tokens · $1.77/1M · 23% cached
- gpt-4.1$12,587
5.65B tokens · $2.23/1M · 8% cached
- gpt-4.1-mini$1,817
4.34B tokens · $0.42/1M · 19% cached
- gemini-2.5-flash-lite$423
5.44B tokens · $0.08/1M · 53% cached
- gpt-4o-mini$383
2.45B tokens · $0.16/1M · 33% cached
Rate, not just total: a model can be a small share of the bill and still be the dearest thing you run. Colours match the dashboard, so a model is the same hue wherever it appears.
First run of this schedule — there is nothing to compare it against yet.
Coverage by provider
Rolled up across connectors, not averaged into one number
- GVGoogle Vertex AI$34,80094%LOGGED
- AFAzure AI Foundry$13,20071%LOGGED
A provider without a billing export can be read, but its cost is modelled from token counts rather than taken from an invoice — so it can never carry LOGGED. Coverage is weighted by spend, so a small account cannot flatter a large one.
These are whole-account figures — $48,000 across both providers. The baseline below prices only what this assessment selected, $39,000; the $9,000 difference is spend in these accounts that the scope left out.
$12,480 could not be attributed to any workload
32% of the spend in scope. The amount is known; what it belongs to is not. Declared as its own figure rather than folded into a grade band.
$46,950 sits in accounts nothing is reading
9 accounts have live traffic and no connector, so none of it appears anywhere above. This figure exists because a coverage percentage over the connected estate would otherwise look complete.
- kestrel-ops-77 · $6,900
- orchid-intake-01 · $2,400
- nw-staging-02 · $180
- sub-5512-atlas · $8,400
The scope — 3 applications, 7 workloads
Every workload sits under the application it belongs to, and both name the connector they were read through
Knowledge Base
label · app=knowledge-baseread from a log label · 3 workloads
wf-8f3a21c9 · confidence high
$29,200MEASUREDwf-5c22e918 · confidence medium
$1,200MODELLEDwf-2ad88e71 · confidence high
$240MODELLED
Customer Support
declared · declared by customerthe customer told us this name · 2 workloads
wf-41bd07e2 · confidence high
$3,600MODELLEDwf-09e5b6f4 · confidence medium
$1,800MODELLED
Catalogue
hash · sha256:4f1c…9ab2derived from a fingerprint — no label existed · 2 workloads
wf-77c1d3aa · confidence high
$2,400MODELLEDwf-b61f4c0d · confidence low
$560ASSUMED
What the replay proved
The only stage that ran your traffic, and the only evidence graded MEASURED
Document Summariser · Knowledge Base
reasoning high · cap 8,192 → gemini-2.5-pro · reasoning low
Calls replayed
120
of 500 approved
Tokens spent
1.6M
1.5M in · 80.0k out
Replay cost
$6.40
cap $8
Latency
1.85s
median, on the tested config
Saving if adopted
$900/mo
$29,200 → $28,300
120 calls stand for $29,200/month of traffic — a sample, not a census. The verdict says the cheaper configuration produced output that held against the original on those calls; it does not promise every call behaves the same.
The other 6 workloads in scope were not replayed. Their recommendations are modelled from token counts and carry the weaker grade wherever they appear — no figure in this report claims otherwise.
Today, and three priced futures
Each independently graded — S0 is what you already pay
S0BaselineWhat the project costs today
$39,000
LOGGEDS1Lift-and-shiftEverything onto the newer model, no re-sizing
$41,155+$2,155
MODELLED$38,448–$44,691More expensive than doing nothing.S2Right-sizedEach workload onto its own recommended path
$37,105−$1,895
MODELLED$35,740–$38,890S3Right-sized + leversS2 plus batching, caching, reasoning level, response caps
$34,470−$4,530
MODELLED$32,591–$36,845
Bars are drawn to the same scale, and the dashed line is today's bill — anything crossing it costs more than doing nothing. The pale band behind each bar is the range, so a wide band is a figure you should not act on as though it were a point. Totals inherit the weakest grade among their inputs, and every scenario is scaled to the selected scope rather than the whole project.
What this run could not say
Was this run useful?
What you say here shows against the run in history