COST REPORT
Vega research copilot
Run 2026-07-28 · completed · pinned to snap-2026-08-27-9d04. Built to be cropped: every figure carries its own grade.
Document Summariser is 123% of the attributed bill and is already on the right model — it needs its reasoning level turned down, not a new one. The obvious move comes out $1,271/month more expensive than doing nothing.
This rests on 1 of 7 workloads with MEASURED evidence, from $12.20 of replay. The rest is modelled and says so, per figure.
Tokens and cost in scope — read from your logs
Your provider's own records, 2026-06-20 to 2026-07-25 · 36 days · 1 account. Nothing here was replayed.
Requests
3.6M
model calls in window
Input tokens
8.78B
8.28B fresh · 502.0M cached
Cached input
6%
billed at a tenth of the input rate
Output tokens
527.6M
no reasoning tokens in scope
Total tokens
9.30B
input + output
Token cost
$14,989
$12,490/month at this rate · $1.61 per 1M
Counted over the 36-day scope window, so the total is larger than a month. The monthly equivalent above still runs ahead of the $25,740 baseline, and should: this counts every readable call in the scope's accounts, while the baseline prices only the spend discovery could attribute to a workload — the difference is the unattributed traffic declared further down. Cost here is tokens against the published price list, which is why the per-1M rate reconciles with the total beside it. gpt-4.1-ft-vega01 is excluded from scope and from these counts.
Cost across the scope window
36 days · the dashed line is the daily mean
A monthly figure only means something if the month was ordinary. Days well above the mean are what a rate hides — $416.35 a day here, and the run is priced from the whole window rather than any one of them.
Where the spend sits
By model, across the scope window
- gpt-4.1$13,499
5.87B tokens · $2.30/1M · 4% cached
- gpt-4.1-mini$1,489
3.43B tokens · $0.43/1M · 9% cached
Rate, not just total: a model can be a small share of the bill and still be the dearest thing you run. Colours match the dashboard, so a model is the same hue wherever it appears.
First run of this schedule — there is nothing to compare it against yet.
Coverage by provider
Rolled up across connectors, not averaged into one number
- AFAzure AI Foundry$31,80058%LOGGED
A provider without a billing export can be read, but its cost is modelled from token counts rather than taken from an invoice — so it can never carry LOGGED. Coverage is weighted by spend, so a small account cannot flatter a large one.
These are whole-account figures — $31,800 across both providers. The baseline below prices only what this assessment selected, $25,740; the $6,060 difference is spend in these accounts that the scope left out.
$10,039 could not be attributed to any workload
39% of the spend in scope. The amount is known; what it belongs to is not. Declared as its own figure rather than folded into a grade band.
$46,950 sits in accounts nothing is reading
9 accounts have live traffic and no connector, so none of it appears anywhere above. This figure exists because a coverage percentage over the connected estate would otherwise look complete.
- kestrel-ops-77 · $6,900
- orchid-intake-01 · $2,400
- nw-staging-02 · $180
- sub-5512-atlas · $8,400
The scope — 3 applications, 7 workloads
Every workload sits under the application it belongs to, and both name the connector they were read through
Knowledge Base
label · app=knowledge-baseread from a log label · 3 workloads
wf-vg021c9 · confidence high
$19,272MEASUREDwf-vg4e918 · confidence medium
$792MODELLEDwf-vg68e71 · confidence high
$158MODELLED
Customer Support
declared · declared by customerthe customer told us this name · 2 workloads
wf-vg107e2 · confidence high
$2,376MODELLEDwf-vg3b6f4 · confidence medium
$1,188MODELLED
Catalogue
hash · sha256:4f1c…9ab2derived from a fingerprint — no label existed · 2 workloads
wf-vg2d3aa · confidence high
$1,584MODELLEDwf-vg54c0d · confidence low
$370ASSUMED
What the replay proved
The only stage that ran your traffic, and the only evidence graded MEASURED
Document Summariser · Knowledge Base
reasoning high · cap 8,192 → gemini-2.5-pro · reasoning low
Calls replayed
120
of 1,200 approved
Tokens spent
1.6M
1.5M in · 80.0k out
Replay cost
$6.40
cap $18
Latency
1.85s
median, on the tested config
Saving if adopted
$594/mo
$19,272 → $18,678
120 calls stand for $19,272/month of traffic — a sample, not a census. The verdict says the cheaper configuration produced output that held against the original on those calls; it does not promise every call behaves the same.
The other 6 workloads in scope were not replayed. Their recommendations are modelled from token counts and carry the weaker grade wherever they appear — no figure in this report claims otherwise.
Today, and three priced futures
Each independently graded — S0 is what you already pay
S0BaselineWhat the project costs today
$25,643
LOGGEDS1Lift-and-shiftEverything onto the newer model, no re-sizing
$26,914+$1,271
MODELLED$25,317–$29,000More expensive than doing nothing.S2Right-sizedEach workload onto its own recommended path
$24,525−$1,118
MODELLED$23,720–$25,578S3Right-sized + leversS2 plus batching, caching, reasoning level, response caps
$22,971−$2,672
MODELLED$21,863–$24,372
Bars are drawn to the same scale, and the dashed line is today's bill — anything crossing it costs more than doing nothing. The pale band behind each bar is the range, so a wide band is a figure you should not act on as though it were a point. Totals inherit the weakest grade among their inputs, and every scenario is scaled to the selected scope rather than the whole project.
What this run could not say
Was this run useful?
What you say here shows against the run in history