Run run-vg-1

COST REPORT

Vega research copilot

Run 2026-07-28 · completed · pinned to snap-2026-08-27-9d04. Built to be cropped: every figure carries its own grade.

Document Summariser is 123% of the attributed bill and is already on the right model — it needs its reasoning level turned down, not a new one. The obvious move comes out $1,271/month more expensive than doing nothing.

This rests on 1 of 7 workloads with MEASURED evidence, from $12.20 of replay. The rest is modelled and says so, per figure.

Tokens and cost in scope — read from your logs

Your provider's own records, 2026-06-20 to 2026-07-25 · 36 days · 1 account. Nothing here was replayed.

LOGGED

Requests

3.6M

model calls in window

Input tokens

8.78B

8.28B fresh · 502.0M cached

Cached input

6%

billed at a tenth of the input rate

Output tokens

527.6M

no reasoning tokens in scope

Total tokens

9.30B

input + output

Token cost

$14,989

$12,490/month at this rate · $1.61 per 1M

Fresh input89%· full rateCached input5%· a tenth of the rateOutput6%· the dearest tokensReasoning0%· billed as output

Counted over the 36-day scope window, so the total is larger than a month. The monthly equivalent above still runs ahead of the $25,740 baseline, and should: this counts every readable call in the scope's accounts, while the baseline prices only the spend discovery could attribute to a workload — the difference is the unattributed traffic declared further down. Cost here is tokens against the published price list, which is why the per-1M rate reconciles with the total beside it. gpt-4.1-ft-vega01 is excluded from scope and from these counts.

Cost across the scope window

36 days · the dashed line is the daily mean

MODELLED

A monthly figure only means something if the month was ordinary. Days well above the mean are what a rate hides — $416.35 a day here, and the run is priced from the whole window rather than any one of them.

Where the spend sits

By model, across the scope window

  1. gpt-4.1$13,499

    5.87B tokens · $2.30/1M · 4% cached

  2. gpt-4.1-mini$1,489

    3.43B tokens · $0.43/1M · 9% cached

Rate, not just total: a model can be a small share of the bill and still be the dearest thing you run. Colours match the dashboard, so a model is the same hue wherever it appears.

First run of this schedule — there is nothing to compare it against yet.

Every figure states how it was obtained:MEASUREDLOGGEDMODELLEDASSUMED

Coverage by provider

Rolled up across connectors, not averaged into one number

  • AFAzure AI Foundry$31,800
    58%LOGGED

A provider without a billing export can be read, but its cost is modelled from token counts rather than taken from an invoice — so it can never carry LOGGED. Coverage is weighted by spend, so a small account cannot flatter a large one.

These are whole-account figures — $31,800 across both providers. The baseline below prices only what this assessment selected, $25,740; the $6,060 difference is spend in these accounts that the scope left out.

$10,039 could not be attributed to any workload

39% of the spend in scope. The amount is known; what it belongs to is not. Declared as its own figure rather than folded into a grade band.

$46,950 sits in accounts nothing is reading

9 accounts have live traffic and no connector, so none of it appears anywhere above. This figure exists because a coverage percentage over the connected estate would otherwise look complete.

  • kestrel-ops-77 · $6,900
  • orchid-intake-01 · $2,400
  • nw-staging-02 · $180
  • sub-5512-atlas · $8,400

The scope — 3 applications, 7 workloads

Every workload sits under the application it belongs to, and both name the connector they were read through

In scope: $25,740 LOGGEDAttributed: $15,701 LOGGED

Knowledge Base

label · app=knowledge-base

read from a log label · 3 workloads

  • Document Summarisercomplex-agenticinferred from logs

    wf-vg021c9 · confidence high

    $19,272
    MEASURED
  • Internal Searchmoderateinferred from logs

    wf-vg4e918 · confidence medium

    $792
    MODELLED
  • Meeting Notessimpleinferred from logs

    wf-vg68e71 · confidence high

    $158
    MODELLED

Customer Support

declared · declared by customer

the customer told us this name · 2 workloads

  • Support Triagesimpledeclared

    wf-vg107e2 · confidence high

    $2,376
    MODELLED
  • Chat Assistantmoderatedeclared

    wf-vg3b6f4 · confidence medium

    $1,188
    MODELLED

Catalogue

hash · sha256:4f1c…9ab2

derived from a fingerprint — no label existed · 2 workloads

  • Product Classifiersimpleinferred from logs

    wf-vg2d3aa · confidence high

    $1,584
    MODELLED
  • Contract Extractortier unknown — we couldn't tellinferred from logs

    wf-vg54c0d · confidence low

    $370
    ASSUMED

What the replay proved

The only stage that ran your traffic, and the only evidence graded MEASURED

MEASURED

Document Summariser · Knowledge Base

reasoning high · cap 8,192gemini-2.5-pro · reasoning low

Output holds

Calls replayed

120

of 1,200 approved

Tokens spent

1.6M

1.5M in · 80.0k out

Replay cost

$6.40

cap $18

Latency

1.85s

median, on the tested config

Saving if adopted

$594/mo

$19,272 → $18,678

120 calls stand for $19,272/month of traffic — a sample, not a census. The verdict says the cheaper configuration produced output that held against the original on those calls; it does not promise every call behaves the same.

The other 6 workloads in scope were not replayed. Their recommendations are modelled from token counts and carry the weaker grade wherever they appear — no figure in this report claims otherwise.

Today, and three priced futures

Each independently graded — S0 is what you already pay

  1. S0BaselineWhat the project costs today

    $25,643

    LOGGED
  2. S1Lift-and-shiftEverything onto the newer model, no re-sizing

    $26,914+$1,271

    MODELLED$25,317–$29,000More expensive than doing nothing.
  3. S2Right-sizedEach workload onto its own recommended path

    $24,525$1,118

    MODELLED$23,720–$25,578
  4. S3Right-sized + leversS2 plus batching, caching, reasoning level, response caps

    $22,971$2,672

    MODELLED$21,863–$24,372

Bars are drawn to the same scale, and the dashed line is today's bill — anything crossing it costs more than doing nothing. The pale band behind each bar is the range, so a wide band is a figure you should not act on as though it were a point. Totals inherit the weakest grade among their inputs, and every scenario is scaled to the selected scope rather than the whole project.

What this run could not say

cannot saytime-to-first-token is not present in cloud inference logs, from any source, on any runcannot saygpt-4.1-ft-vega01 was excluded from scope, so nothing here speaks to it

Was this run useful?

What you say here shows against the run in history