Example chat provider
gpt-oss-120b on H100, interactivity floor 100 tokens per second per user, vllm, fp4. Provider: unnamed.
For this workload, 37 percent of your cost variance over the next six months comes from GPU rental prices, 44 percent from your own volume, 15 percent from changes in how efficiently your stack serves tokens, and 6 percent from the gap between your provider's rate and the index. A price hedge can address the first of those. Reservations address the first and part of the second. Nothing standard addresses the third.
Page 1 of 7
Your workload as the model sees it
| intake field | value | status |
|---|---|---|
| customer | Example chat provider | observed |
| model | gpt-oss-120b | observed |
| hardware | h100 | observed |
| precision | fp4 | observed |
| framework | vllm | observed |
| mean prompt tokens per request | 1,633 | observed |
| mean generated tokens per request | 106 | observed |
| requests per month | 575,000,000 | observed |
| monthly demand growth, percent | 3 | observed |
| demand forecast error, percent a month | 10 (shape default) | assumed |
| interactivity floor, tokens per second per user | 100 | observed |
| demand shape | conversation | observed |
| custom shape | none | observed |
| contract type | on_demand | observed |
| reserved GPUs today | 0 | observed |
| reserved rate, USD per GPU-hour | 0 | observed |
| reserved months remaining | 0 | observed |
| on-demand rate, USD per GPU-hour | 3.07 | observed |
| reservation quote in hand, USD per GPU-hour | none, reserved rate assumed at 2.61 USD per GPU-hour | assumed |
| provider | unnamed | observed |
| horizon, months | 12 | observed |
| budget overrun threshold, multiple of plan | 1.25 | observed |
| index volatility scenario, percent a year | 35 | scenario |
| index drift scenario, percent a year | 0 | scenario |
| basis volatility scenario, percent | 10 | scenario |
| risk premium scenario, percent a year | 0 | scenario |
| settlement index correlation scenario | 1 | scenario |
| custom benchmark | none | observed |
Derived from the intake and the benchmarks
| input share (prompt tokens over all tokens) | 94 percent | derived |
| effective yield, tokens per second per GPU | 5,023 | derived |
| single 1k/1k benchmark yield, tokens per second per GPU | 3,154 | derived |
| ratio, effective over single benchmark | 1.59 | derived |
A single 1k/1k benchmark would misprice this mix by 59 percent, the gap between its cost per million tokens and the two-phase figure.
Prompt tokens are prefill work: the GPU reads the whole prompt in one pass, which is bound by arithmetic throughput. Generated tokens are decode work: each token needs a full pass over the model weights and the growing context, which is bound by memory bandwidth. A GPU therefore spends a different amount of time on a prompt token than on a generated token, and a workload's cost depends on its mix, not only on its total token count.
Page 2 of 7
Your yield and how it moves
| yield | tokens per second per GPU | expected progress per month | shock standard deviation per month | status |
|---|---|---|---|---|
| prefill, kappa_p | 5,750 | -2 percent | 8 percent | derived derived |
| decode, kappa_d | 1,703 | +8 percent | 15 percent | derived derived |
| configuration (prompt/generated tokens) | 1k/1k/ | 8k/1k/ | 1k/8k/ |
|---|---|---|---|
| two-phase fit error, share of measured GPU-seconds per request | +20 percent | -1 percent | +0 percent |
2 distinct runs between 2026-03-27 and 2026-05-17; framework tags vllm:v0.18.0, vllm:v0.21.0. The latest published run is older than sixty days, so yields may have moved since.
Expected progress is the mean monthly log change of each yield across the calibration window (8 transitions); shocks are the residuals, resampled jointly for prefill and decode at the snapshot level. Any reservation or hedge sized on today's yield is sized for a quantity that expected progress alone will change by the figures above every month.
Page 3 of 7
Where your cost risk comes from
- GPU price, 37 percent of six-month cost variance: futures or reservations address it.
- Provider basis, 6 percent: nothing standard addresses it; only a contract that settles on your provider's own rate would.
- Your demand, 44 percent: reservations address part of it; forecasting is the lever.
- Serving yield, 15 percent: nothing addresses it; progress-aware sizing avoids compounding it.
- Utilization, 0 percent: your autoscaling policy is the lever.
Diagnostic: the five Shapley shares at six months sum to 103 percent (the exact value is one hundred; the difference is Monte Carlo estimation error of the closed Sobol indices). Shares are simulated under the base scenario with 20,000 paths.
Page 4 of 7
Procurement policies compared
| policy | expected cost, percent of on demand, six months | HE of cumulative cost, six months, percent | P(overrun), percent of paths | ES95, percent of plan | reserved GPUs | idle cash, month six | overflow, percent of compute, month six |
|---|---|---|---|---|---|---|---|
| on demand, unhedged | 100 | 0 | 20 | +75 | 0 | 0 USD | 100 |
| reserve for peak (p95 sizing) assumed | 107 | 80 | 12 | +51 | 132 | 22,700 USD (8 percent of month-six on-demand cost) | 10 |
| reserve for shape-implied peak assumed | 120 | 80 | 30 | +65 | 149 | 25,600 USD (9 percent of month-six on-demand cost) | 10 |
| reserve baseline assumed | 85 | 35 | 6 | +47 | 48 | 0 USD (0 percent of month-six on-demand cost) | 50 |
| reserve baseline plus progress-aware strip assumed | 85 | 44 | 5 | +41 | 48 | 0 USD (0 percent of month-six on-demand cost) | 50 |
| on demand plus minimum-variance futures strip | 100 | 40 | 15 | +59 | 0 | 0 USD | 100 |
All policy figures are simulated under the base scenario. P(overrun) is the share of simulated paths whose cumulative six-month cost exceeds 1.25 times plan; ES95 is the mean overrun of plan in the worst five percent of paths. Reserved rows marked assumed use a reserved rate based on an assumed 15 percent discount from your on-demand rate; the reserve-for-peak rows still overflow on demand when simulated demand exceeds the reserved capacity, so neither is free of on-demand exposure.
This comparison assumes a reserved rate of 85 percent of spot (assumed, 2.61 USD per GPU-hour), a demand shape whose always-on trough is 49 percent of the mean, expected efficiency progress of -2 percent per month in the prefill yield that dominates this mix, and a flat index. Change any of these and the ranking can change; page 7 shows how far.
Under these inputs, reserve baseline plus progress-aware strip has the lowest expected shortfall, at +41 percent of plan in the worst five percent of simulated six-month paths.
Page 5 of 7
Proxy-index hedge simulation
| sizing rule | contracts, month one | contracts, month six | contracts, month twelve | HE of cumulative cost, three months | six months | twelve months |
|---|---|---|---|---|---|---|
| fixed conversion, today's yields | 98 | 113 | 136 | 37 | 39 | 38 |
| progress-aware | 98 | 115 | 145 | 37 | 39 | 38 |
| minimum variance | 99 | 121 | 160 | 38 | 40 | 39 |
HE is one minus the variance of hedged cumulative cost divided by the variance of unhedged on-demand cost; a negative value means the strip added variance. At twelve months the minimum-variance strip is 118 percent of the fixed-conversion strip and the progress-aware strip is 106 percent of it. Each contract count already reflects your basis level of 1.000 index units per GPU-hour. Settlement correlation sensitivity, minimum-variance HE of six-month cost: correlation 1.0 gives 38 percent; correlation 0.9 gives 31 percent; correlation 0.7 gives 18 percent.
One-index theoretical bound
Under the assumptions below the largest share of cost variance any linear hedge on the index can remove is 37 percent at six months (37 at three months, 35 at twelve), from a coefficient of variation of 25 percent for the index and 31 percent for your quantity exposure. simulated
- positive index level S and positive quantity exposure X
- S and X independent
- a hedge linear in S
- settlement period matching the exposure month
- no correlation between basis and the index
- frictionless sizing (no margin, fees or liquidity limits)
- procurement and settlement indices coincide
Applies when the procurement and settlement indices coincide. Proxy-index mismatch can only reduce achievable hedge effectiveness relative to this idealised case.
Range at six months
Minimum-variance HE of six-month cost across the sensitivity grid: 18 / 38 / 64 percent (minimum, median, maximum). The grid is one-at-a-time changes to index volatility (25, 35, 60 percent), demand forecast error (half, base, one and a half times), basis volatility (0, 10, 20 percent), efficiency shock scale (0.5, 1, 1.5), adoption lag (0, 1, 3 months) and settlement correlation (1.0, 0.9, 0.7), each at N = 12000 paths.
Page 6 of 7
Sensitivity, provenance, scope
| quantity | status | source | retrieved (UTC) |
|---|---|---|---|
| index price today | observed | Ornn Compute Price Index, H100 SXM, 2026-09-06 | 2026-09-07T17:43:42Z |
| customer on-demand rate, workload mix, volume, contract | observed | intake file | 2026-09-24T21:41:08Z |
| prefill and decode yields | derived | two-phase fit to InferenceX snapshots 2025-10-20 to 2026-06-15 | 2026-09-07T17:45:53Z |
| expected efficiency progress and shock pool | derived | two-phase fit to InferenceX snapshots 2025-10-20 to 2026-06-15 | 2026-09-07T17:45:53Z |
| demand shape factors (utilization, reserved utilization, trough share, peak ratio) | derived | Azure LLM inference traces 2024 via TokenYield v0.2 | 2026-09-07T17:48:32Z |
| demand forecast error | assumed | shape default | not applicable |
| reserved rate (theta) | assumed | based on an assumed 15 percent discount from your on-demand rate | not applicable |
| index volatility, drift, basis volatility, risk premium, settlement correlation | scenario | intake scenarios | not applicable |
| utilization noise, basis persistence, contract size 730 GPU-hours | assumed | TokenYield v0.2 base parameters | not applicable |
| hedge effectiveness, policy costs, Shapley shares, sensitivity | simulated | Monte Carlo, base scenario, N = 20000 paths, seed 20260908 | 2026-09-24T21:41:08Z |
Ornn's Compute Price Index is a volume-weighted, winsorized mean of executed on-demand rental transactions over a rolling one-hour window, from verified providers. It is not the CME settlement index, which is published by Silicon Data.
Data files used
| file | sha256 (first twelve characters) | bytes | retrieved (UTC) |
|---|---|---|---|
| seed/inferencex/gpt-oss-120b/2025-10-20.json | 89d6cb8ddb65 | 549,775 | 2026-09-07T17:45:53Z |
| seed/inferencex/gpt-oss-120b/2025-11-15.json | f34926a8a9e5 | 542,645 | 2026-09-07T17:45:55Z |
| seed/inferencex/gpt-oss-120b/2025-12-15.json | 93d9316d9673 | 564,049 | 2026-09-07T17:45:58Z |
| seed/inferencex/gpt-oss-120b/2026-01-15.json | 629f4626e6b9 | 635,981 | 2026-09-07T17:46:01Z |
| seed/inferencex/gpt-oss-120b/2026-02-15.json | d3d2dc4178e1 | 636,091 | 2026-09-07T17:46:04Z |
| seed/inferencex/gpt-oss-120b/2026-03-15.json | c8781fc84a1c | 638,833 | 2026-09-07T17:46:06Z |
| seed/inferencex/gpt-oss-120b/2026-04-15.json | 192e1dc64c9c | 680,047 | 2026-09-07T17:46:09Z |
| seed/inferencex/gpt-oss-120b/2026-05-15.json | b3ec642040e2 | 701,378 | 2026-09-07T17:46:12Z |
| seed/inferencex/gpt-oss-120b/2026-06-15.json | acabe31fb7fc | 752,142 | 2026-09-07T17:46:14Z |
| seed/ornn/h100.json | 07a517ba660f | 5,576 | 2026-09-07T17:43:42Z |
Scenario: index volatility 35 percent a year, index drift 0 percent a year, basis volatility 10 percent, risk premium 0 percent, settlement correlation 1.0, demand forecast error 10 percent a month, growth 3 percent a month, 20,000 paths, seed 20260908. Intake hash e84bdffdedd5.
This report is a simulation under stated scenarios, calibrated to public benchmark, index and trace data as of the retrieval dates above and to the inputs you provided. It is not a forecast and not financial advice. The settlement index for CME compute futures is published by Silicon Data; prices here are from Ornn's transaction-based index. Serving-yield dynamics are estimated from public benchmarks for the stated model and hardware over the stated period and may not transfer to your stack. Results are conditional on the model, hardware, workload construction, period and scenarios examined.
Page 7 of 7