Run run-vg-2

COST REPORT

Vega research copilot

Run 2026-08-27 · completed with warnings · pinned to snap-2026-08-27-9d04. Built to be cropped: every figure carries its own grade.

Document Summariser is 129% of the attributed bill and is already on the right model — it needs its reasoning level turned down, not a new one. The obvious move comes out $1,208/month more expensive than doing nothing.

This rests on 1 of 7 workloads with MEASURED evidence, from $14.90 of replay. The rest is modelled and says so, per figure.

Tokens and cost in scope — read from your logs

Your provider's own records, 2026-06-20 to 2026-07-25 · 36 days · 1 account. Nothing here was replayed.

LOGGED

Requests

3.6M

model calls in window

Input tokens

8.78B

8.28B fresh · 502.0M cached

Cached input

6%

billed at a tenth of the input rate

Output tokens

527.6M

no reasoning tokens in scope

Total tokens

9.30B

input + output

Token cost

$14,989

$12,490/month at this rate · $1.61 per 1M

Fresh input89%· full rateCached input5%· a tenth of the rateOutput6%· the dearest tokensReasoning0%· billed as output

Counted over the 36-day scope window, so the total is larger than a month. The monthly equivalent above still runs ahead of the $25,740 baseline, and should: this counts every readable call in the scope's accounts, while the baseline prices only the spend discovery could attribute to a workload — the difference is the unattributed traffic declared further down. Cost here is tokens against the published price list, which is why the per-1M rate reconciles with the total beside it. gpt-4.1-ft-vega01 is excluded from scope and from these counts.

Cost across the scope window

36 days · the dashed line is the daily mean

MODELLED

A monthly figure only means something if the month was ordinary. Days well above the mean are what a rate hides — $416.35 a day here, and the run is priced from the whole window rather than any one of them.

Where the spend sits

By model, across the scope window

  1. gpt-4.1$13,499

    5.87B tokens · $2.30/1M · 4% cached

  2. gpt-4.1-mini$1,489

    3.43B tokens · $0.43/1M · 9% cached

Rate, not just total: a model can be a small share of the bill and still be the dearest thing you run. Colours match the dashboard, so a model is the same hue wherever it appears.

Against the previous run — 2026-07-28

Same schedule, same scope. Only what the run could see has moved.

Saving found

$1,063

−$55

was $1,118

Coverage

58%

−3 pts

was 61%

Right-sized (S2)

$24,580

+$55

was $24,525

Replay spend

$14.90

+$2.70

was $12.20

The baseline does not move between runs — the bill is the bill. What moves is how much of it this run could read, and every saving is scaled to that: a run that saw 58% of the estate cannot claim a saving on the part it could not.

Every figure states how it was obtained:MEASUREDLOGGEDMODELLEDASSUMED

Coverage by provider

Rolled up across connectors, not averaged into one number

  • AFAzure AI Foundry$31,800
    58%LOGGED

A provider without a billing export can be read, but its cost is modelled from token counts rather than taken from an invoice — so it can never carry LOGGED. Coverage is weighted by spend, so a small account cannot flatter a large one.

These are whole-account figures — $31,800 across both providers. The baseline below prices only what this assessment selected, $25,740; the $6,060 difference is spend in these accounts that the scope left out.

$10,811 could not be attributed to any workload

42% of the spend in scope. The amount is known; what it belongs to is not. Declared as its own figure rather than folded into a grade band.

$46,950 sits in accounts nothing is reading

9 accounts have live traffic and no connector, so none of it appears anywhere above. This figure exists because a coverage percentage over the connected estate would otherwise look complete.

  • kestrel-ops-77 · $6,900
  • orchid-intake-01 · $2,400
  • nw-staging-02 · $180
  • sub-5512-atlas · $8,400

The scope — 3 applications, 7 workloads

Every workload sits under the application it belongs to, and both name the connector they were read through

In scope: $25,740 LOGGEDAttributed: $14,929 LOGGED

Knowledge Base

label · app=knowledge-base

read from a log label · 3 workloads

  • Document Summarisercomplex-agenticinferred from logs

    wf-vg021c9 · confidence high

    $19,272
    MEASURED
  • Internal Searchmoderateinferred from logs

    wf-vg4e918 · confidence medium

    $792
    MODELLED
  • Meeting Notessimpleinferred from logs

    wf-vg68e71 · confidence high

    $158
    MODELLED

Customer Support

declared · declared by customer

the customer told us this name · 2 workloads

  • Support Triagesimpledeclared

    wf-vg107e2 · confidence high

    $2,376
    MODELLED
  • Chat Assistantmoderatedeclared

    wf-vg3b6f4 · confidence medium

    $1,188
    MODELLED

Catalogue

hash · sha256:4f1c…9ab2

derived from a fingerprint — no label existed · 2 workloads

  • Product Classifiersimpleinferred from logs

    wf-vg2d3aa · confidence high

    $1,584
    MODELLED
  • Contract Extractortier unknown — we couldn't tellinferred from logs

    wf-vg54c0d · confidence low

    $370
    ASSUMED

What the replay proved

The only stage that ran your traffic, and the only evidence graded MEASURED

MEASURED

Document Summariser · Knowledge Base

reasoning high · cap 8,192gemini-2.5-pro · reasoning low

Output holds

Calls replayed

120

of 1,200 approved

Tokens spent

1.6M

1.5M in · 80.0k out

Replay cost

$6.40

cap $18

Latency

1.85s

median, on the tested config

Saving if adopted

$594/mo

$19,272 → $18,678

120 calls stand for $19,272/month of traffic — a sample, not a census. The verdict says the cheaper configuration produced output that held against the original on those calls; it does not promise every call behaves the same.

The other 6 workloads in scope were not replayed. Their recommendations are modelled from token counts and carry the weaker grade wherever they appear — no figure in this report claims otherwise.

Today, and three priced futures

Each independently graded — S0 is what you already pay

  1. S0BaselineWhat the project costs today

    $25,643

    LOGGED
  2. S1Lift-and-shiftEverything onto the newer model, no re-sizing

    $26,851+$1,208

    MODELLED$25,333–$28,834More expensive than doing nothing.
  3. S2Right-sizedEach workload onto its own recommended path

    $24,580$1,063

    MODELLED$23,815–$25,581
  4. S3Right-sized + leversS2 plus batching, caching, reasoning level, response caps

    $23,102$2,541

    MODELLED$22,049–$24,435

Bars are drawn to the same scale, and the dashed line is today's bill — anything crossing it costs more than doing nothing. The pale band behind each bar is the range, so a wide band is a figure you should not act on as though it were a point. Totals inherit the weakest grade among their inputs, and every scenario is scaled to the selected scope rather than the whole project.

What this run could not say

cannot saytime-to-first-token is not present in cloud inference logs, from any source, on any runcannot saygpt-4.1-ft-vega01 was excluded from scope, so nothing here speaks to it

Was this run useful?

What you say here shows against the run in history

  • S. Vahid 2026-08-28

    Coverage was too low to act on. We need the research subscription connected before this run means anything.