Run run-nw-3

COST REPORT

Northwind Q3 cost review

Run 2026-09-01 · completed · pinned to snap-2026-09-01-a7f3, snap-2026-09-01-c118. Built to be cropped: every figure carries its own grade.

Document Summariser is 92% of the attributed bill and is already on the right model — it needs its reasoning level turned down, not a new one. The obvious move comes out $2,567/month more expensive than doing nothing.

This rests on 1 of 7 workloads with MEASURED evidence, from $6.40 of replay. The rest is modelled and says so, per figure.

Tokens and cost in scope — read from your logs

Your provider's own records, 2026-07-24 to 2026-08-28 · 36 days · 2 accounts. Nothing here was replayed.

LOGGED

Requests

14.8M

model calls in window

Input tokens

36.18B

27.06B fresh · 9.12B cached

Cached input

25%

billed at a tenth of the input rate

Output tokens

2.76B

539.7M of it reasoning

Total tokens

38.94B

input + output

Token cost

$52,450

$43,709/month at this rate · $1.35 per 1M

Fresh input69%· full rateCached input23%· a tenth of the rateOutput6%· the dearest tokensReasoning1%· billed as output

Counted over the 36-day scope window, so the total is larger than a month. The monthly equivalent above still runs ahead of the $39,000 baseline, and should: this counts every readable call in the scope's accounts, while the baseline prices only the spend discovery could attribute to a workload — the difference is the unattributed traffic declared further down. Cost here is tokens against the published price list, which is why the per-1M rate reconciles with the total beside it. gemini-2.5-flash is excluded from scope and from these counts.

Cost across the scope window

36 days · the dashed line is the daily mean

MODELLED

A monthly figure only means something if the month was ordinary. Days well above the mean are what a rate hides — $1456.95 a day here, and the run is priced from the whole window rather than any one of them.

Where the spend sits

By model, across the scope window

  1. gemini-2.5-pro$37,241

    21.06B tokens · $1.77/1M · 23% cached

  2. gpt-4.1$12,587

    5.65B tokens · $2.23/1M · 8% cached

  3. gpt-4.1-mini$1,817

    4.34B tokens · $0.42/1M · 19% cached

  4. gemini-2.5-flash-lite$423

    5.44B tokens · $0.08/1M · 53% cached

  5. gpt-4o-mini$383

    2.45B tokens · $0.16/1M · 33% cached

Rate, not just total: a model can be a small share of the bill and still be the dearest thing you run. Colours match the dashboard, so a model is the same hue wherever it appears.

Against the previous run — 2026-08-25

Same schedule, same scope. Only what the run could see has moved.

Saving found

$2,257

+$55

was $2,202

Coverage

81%

+2 pts

was 79%

Right-sized (S2)

$36,743

−$55

was $36,798

Replay spend

$6.40

+$0.30

was $6.10

The baseline does not move between runs — the bill is the bill. What moves is how much of it this run could read, and every saving is scaled to that: a run that saw 81% of the estate cannot claim a saving on the part it could not.

Every figure states how it was obtained:MEASUREDLOGGEDMODELLEDASSUMED

Coverage by provider

Rolled up across connectors, not averaged into one number

  • GVGoogle Vertex AI$34,800
    94%LOGGED
  • AFAzure AI Foundry$13,200
    71%LOGGED

A provider without a billing export can be read, but its cost is modelled from token counts rather than taken from an invoice — so it can never carry LOGGED. Coverage is weighted by spend, so a small account cannot flatter a large one.

These are whole-account figures — $48,000 across both providers. The baseline below prices only what this assessment selected, $39,000; the $9,000 difference is spend in these accounts that the scope left out.

$7,410 could not be attributed to any workload

19% of the spend in scope. The amount is known; what it belongs to is not. Declared as its own figure rather than folded into a grade band.

$46,950 sits in accounts nothing is reading

9 accounts have live traffic and no connector, so none of it appears anywhere above. This figure exists because a coverage percentage over the connected estate would otherwise look complete.

  • kestrel-ops-77 · $6,900
  • orchid-intake-01 · $2,400
  • nw-staging-02 · $180
  • sub-5512-atlas · $8,400

The scope — 3 applications, 7 workloads

Every workload sits under the application it belongs to, and both name the connector they were read through

In scope: $39,000 LOGGEDAttributed: $31,590 LOGGED

Knowledge Base

label · app=knowledge-base

read from a log label · 3 workloads

  • Document Summarisercomplex-agenticinferred from logs

    wf-8f3a21c9 · confidence high

    $29,200
    MEASURED
  • Internal Searchmoderateinferred from logs

    wf-5c22e918 · confidence medium

    $1,200
    MODELLED
  • Meeting Notessimpleinferred from logs

    wf-2ad88e71 · confidence high

    $240
    MODELLED

Customer Support

declared · declared by customer

the customer told us this name · 2 workloads

  • Support Triagesimpledeclared

    wf-41bd07e2 · confidence high

    $3,600
    MODELLED
  • Chat Assistantmoderatedeclared

    wf-09e5b6f4 · confidence medium

    $1,800
    MODELLED

Catalogue

hash · sha256:4f1c…9ab2

derived from a fingerprint — no label existed · 2 workloads

  • Product Classifiersimpleinferred from logs

    wf-77c1d3aa · confidence high

    $2,400
    MODELLED
  • Contract Extractortier unknown — we couldn't tellinferred from logs

    wf-b61f4c0d · confidence low

    $560
    ASSUMED

What the replay proved

The only stage that ran your traffic, and the only evidence graded MEASURED

MEASURED

Document Summariser · Knowledge Base

reasoning high · cap 8,192gemini-2.5-pro · reasoning low

Output holds

Calls replayed

120

of 500 approved

Tokens spent

1.6M

1.5M in · 80.0k out

Replay cost

$6.40

cap $8

Latency

1.85s

median, on the tested config

Saving if adopted

$900/mo

$29,200 → $28,300

120 calls stand for $29,200/month of traffic — a sample, not a census. The verdict says the cheaper configuration produced output that held against the original on those calls; it does not promise every call behaves the same.

The other 6 workloads in scope were not replayed. Their recommendations are modelled from token counts and carry the weaker grade wherever they appear — no figure in this report claims otherwise.

Today, and three priced futures

Each independently graded — S0 is what you already pay

  1. S0BaselineWhat the project costs today

    $39,000

    LOGGED
  2. S1Lift-and-shiftEverything onto the newer model, no re-sizing

    $41,567+$2,567

    MODELLED$38,342–$45,779More expensive than doing nothing.
  3. S2Right-sizedEach workload onto its own recommended path

    $36,743$2,257

    MODELLED$35,117–$38,868
  4. S3Right-sized + leversS2 plus batching, caching, reasoning level, response caps

    $33,603$5,397

    MODELLED$31,366–$36,433

Bars are drawn to the same scale, and the dashed line is today's bill — anything crossing it costs more than doing nothing. The pale band behind each bar is the range, so a wide band is a figure you should not act on as though it were a point. Totals inherit the weakest grade among their inputs, and every scenario is scaled to the selected scope rather than the whole project.

What this run could not say

cannot saytime-to-first-token is not present in cloud inference logs, from any source, on any runcannot saygemini-2.5-flash was excluded from scope, so nothing here speaks to it

Was this run useful?

What you say here shows against the run in history