Furrow report, 24 September 2026. Furrow 1.0.0, engine tokenyield-0.2.

Example chat provider

gpt-oss-120b on H100, interactivity floor 100 tokens per second per user, vllm, fp4. Provider: unnamed.

cost per million tokens, two-phase, at your rate, autoscaled utilization derived
0.212
USD per million tokens at an index of 3.07 and your rate of 3.07 USD per GPU-hour
GPU-hours needed in the next quarter derived
220,600
billable GPU-hours over months one to three under expected efficiency progress
share of six-month cost variance that is GPU price simulated
37
Shapley share of the variance of cumulative cost over months one to six
reserved fleet with the lowest expected shortfall simulated
48
GPUs under these inputs: reserve baseline plus progress-aware strip
based on an assumed 15 percent discount from your on-demand rate

For this workload, 37 percent of your cost variance over the next six months comes from GPU rental prices, 44 percent from your own volume, 15 percent from changes in how efficiently your stack serves tokens, and 6 percent from the gap between your provider's rate and the index. A price hedge can address the first of those. Reservations address the first and part of the second. Nothing standard addresses the third.

Page 1 of 7

Your workload as the model sees it

intake fieldvaluestatus
customerExample chat providerobserved
modelgpt-oss-120bobserved
hardwareh100observed
precisionfp4observed
frameworkvllmobserved
mean prompt tokens per request1,633observed
mean generated tokens per request106observed
requests per month575,000,000observed
monthly demand growth, percent3observed
demand forecast error, percent a month10 (shape default)assumed
interactivity floor, tokens per second per user100observed
demand shapeconversationobserved
custom shapenoneobserved
contract typeon_demandobserved
reserved GPUs today0observed
reserved rate, USD per GPU-hour0observed
reserved months remaining0observed
on-demand rate, USD per GPU-hour3.07observed
reservation quote in hand, USD per GPU-hournone, reserved rate assumed at 2.61 USD per GPU-hourassumed
providerunnamedobserved
horizon, months12observed
budget overrun threshold, multiple of plan1.25observed
index volatility scenario, percent a year35scenario
index drift scenario, percent a year0scenario
basis volatility scenario, percent10scenario
risk premium scenario, percent a year0scenario
settlement index correlation scenario1scenario
custom benchmarknoneobserved

Derived from the intake and the benchmarks

input share (prompt tokens over all tokens)94 percentderived
effective yield, tokens per second per GPU5,023derived
single 1k/1k benchmark yield, tokens per second per GPU3,154derived
ratio, effective over single benchmark1.59derived

A single 1k/1k benchmark would misprice this mix by 59 percent, the gap between its cost per million tokens and the two-phase figure.

0 50% 100% prompt tokens (prefill) generated tokens (decode) 94% 6%
Derived from the intake: share of your tokens that are prompt tokens (prefill work) versus generated tokens (decode work).

Prompt tokens are prefill work: the GPU reads the whole prompt in one pass, which is bound by arithmetic throughput. Generated tokens are decode work: each token needs a full pass over the model weights and the growing context, which is bound by memory bandwidth. A GPU therefore spends a different amount of time on a prompt token than on a generated token, and a workload's cost depends on its mix, not only on its total token count.

Page 2 of 7

Your yield and how it moves

yieldtokens per second per GPUexpected progress per monthshock standard deviation per monthstatus
prefill, kappa_p5,750-2 percent8 percentderived derived
decode, kappa_d1,703+8 percent15 percentderived derived
configuration (prompt/generated tokens)1k/1k/8k/1k/1k/8k/
two-phase fit error, share of measured GPU-seconds per request+20 percent-1 percent+0 percent
25-10 25-11 25-12 26-01 26-02 26-03 26-04 26-05 26-06 1.0 1.2 1.4 1.6 1.8 yield, first snapshot = 1 prefill yield kappa_p decode yield kappa_d expected progress, prefill expected progress, decode
Derived from public benchmark snapshots: prefill and decode yields by snapshot from 2025-10-20 to 2026-06-15, rebased to the first snapshot; dashed lines are the expected-progress paths; small ticks mark published runs.

2 distinct runs between 2026-03-27 and 2026-05-17; framework tags vllm:v0.18.0, vllm:v0.21.0. The latest published run is older than sixty days, so yields may have moved since.

Expected progress is the mean monthly log change of each yield across the calibration window (8 transitions); shocks are the residuals, resampled jointly for prefill and decode at the snapshot level. Any reservation or hedge sized on today's yield is sized for a quantity that expected progress alone will change by the figures above every month.

Page 3 of 7

Where your cost risk comes from

months 1 to 3 months 1 to 6 months 1 to 12 0 50% 100% share of cost variance 36% 37% 39% 10% 6% 44% 44% 43% 14% 15% 17% GPU price provider basis your demand serving yield utilization
Simulated, base scenario: Shapley shares of the variance of cumulative cost by factor at three, six and twelve months; the raw Shapley estimates sum to slightly more than one hundred percent because of Monte Carlo error.

Diagnostic: the five Shapley shares at six months sum to 103 percent (the exact value is one hundred; the difference is Monte Carlo estimation error of the closed Sobol indices). Shares are simulated under the base scenario with 20,000 paths.

Page 4 of 7

Procurement policies compared

policyexpected cost, percent of on demand, six monthsHE of cumulative cost, six months, percentP(overrun), percent of pathsES95, percent of planreserved GPUsidle cash, month sixoverflow, percent of compute, month six
on demand, unhedged100020+7500 USD100
reserve for peak (p95 sizing) assumed1078012+5113222,700 USD (8 percent of month-six on-demand cost)10
reserve for shape-implied peak assumed1208030+6514925,600 USD (9 percent of month-six on-demand cost)10
reserve baseline assumed85356+47480 USD (0 percent of month-six on-demand cost)50
reserve baseline plus progress-aware strip assumed85445+41480 USD (0 percent of month-six on-demand cost)50
on demand plus minimum-variance futures strip1004015+5900 USD100

All policy figures are simulated under the base scenario. P(overrun) is the share of simulated paths whose cumulative six-month cost exceeds 1.25 times plan; ES95 is the mean overrun of plan in the worst five percent of paths. Reserved rows marked assumed use a reserved rate based on an assumed 15 percent discount from your on-demand rate; the reserve-for-peak rows still overflow on demand when simulated demand exceeds the reserved capacity, so neither is free of on-demand exposure.

0.50 0.75 1.00 1.25 1.50 1.75 2.00 2.25 2.50 six-month cost relative to plan 0 50% 100% share of simulated paths below plan threshold 1.25 on demand on demand plus futures strip reserve baseline reserve for peak (p95) plan (cost equals budget) overrun threshold, 1.25 times plan
Simulated, base scenario: distribution of cumulative six-month cost relative to plan under each policy; a steeper curve means a narrower range of outcomes.

This comparison assumes a reserved rate of 85 percent of spot (assumed, 2.61 USD per GPU-hour), a demand shape whose always-on trough is 49 percent of the mean, expected efficiency progress of -2 percent per month in the prefill yield that dominates this mix, and a flat index. Change any of these and the ranking can change; page 7 shows how far.

Under these inputs, reserve baseline plus progress-aware strip has the lowest expected shortfall, at +41 percent of plan in the worst five percent of simulated six-month paths.

Page 5 of 7

Proxy-index hedge simulation

1 2 3 4 5 6 7 8 9 10 11 12 delivery month 100 110 120 130 140 150 160 contracts fixed conversion (reference) progress-aware minimum variance
Simulated, base scenario: contracts per delivery month under the three sizing rules; one contract is one month of rent for one GPU, taken as 730 GPU-hours.
sizing rulecontracts, month onecontracts, month sixcontracts, month twelveHE of cumulative cost, three monthssix monthstwelve months
fixed conversion, today's yields98113136373938
progress-aware98115145373938
minimum variance99121160384039

HE is one minus the variance of hedged cumulative cost divided by the variance of unhedged on-demand cost; a negative value means the strip added variance. At twelve months the minimum-variance strip is 118 percent of the fixed-conversion strip and the progress-aware strip is 106 percent of it. Each contract count already reflects your basis level of 1.000 index units per GPU-hour. Settlement correlation sensitivity, minimum-variance HE of six-month cost: correlation 1.0 gives 38 percent; correlation 0.9 gives 31 percent; correlation 0.7 gives 18 percent.

One-index theoretical bound

Under the assumptions below the largest share of cost variance any linear hedge on the index can remove is 37 percent at six months (37 at three months, 35 at twelve), from a coefficient of variation of 25 percent for the index and 31 percent for your quantity exposure. simulated

Applies when the procurement and settlement indices coincide. Proxy-index mismatch can only reduce achievable hedge effectiveness relative to this idealised case.

Range at six months

Minimum-variance HE of six-month cost across the sensitivity grid: 18 / 38 / 64 percent (minimum, median, maximum). The grid is one-at-a-time changes to index volatility (25, 35, 60 percent), demand forecast error (half, base, one and a half times), basis volatility (0, 10, 20 percent), efficiency shock scale (0.5, 1, 1.5), adoption lag (0, 1, 3 months) and settlement correlation (1.0, 0.9, 0.7), each at N = 12000 paths.

Page 6 of 7

Sensitivity, provenance, scope

20% 40% 60% 80% minimum-variance HE of six-month cost adoption lag basis volatility efficiency shock scale demand forecast error index volatility 0 months 38% 3 months 44% 0% 31% 20% 41% 0.5x 32% 1.5x 43% half 25% 1.5x 56% 25% 24% 60% 64%
Simulated, one input at a time: minimum-variance HE of six-month cost when each input moves between its low and high grid value with the others at the base scenario; the dashed line is the base scenario.
quantitystatussourceretrieved (UTC)
index price todayobservedOrnn Compute Price Index, H100 SXM, 2026-09-062026-09-07T17:43:42Z
customer on-demand rate, workload mix, volume, contractobservedintake file2026-09-24T21:41:08Z
prefill and decode yieldsderivedtwo-phase fit to InferenceX snapshots 2025-10-20 to 2026-06-152026-09-07T17:45:53Z
expected efficiency progress and shock poolderivedtwo-phase fit to InferenceX snapshots 2025-10-20 to 2026-06-152026-09-07T17:45:53Z
demand shape factors (utilization, reserved utilization, trough share, peak ratio)derivedAzure LLM inference traces 2024 via TokenYield v0.22026-09-07T17:48:32Z
demand forecast errorassumedshape defaultnot applicable
reserved rate (theta)assumedbased on an assumed 15 percent discount from your on-demand ratenot applicable
index volatility, drift, basis volatility, risk premium, settlement correlationscenariointake scenariosnot applicable
utilization noise, basis persistence, contract size 730 GPU-hoursassumedTokenYield v0.2 base parametersnot applicable
hedge effectiveness, policy costs, Shapley shares, sensitivitysimulatedMonte Carlo, base scenario, N = 20000 paths, seed 202609082026-09-24T21:41:08Z

Ornn's Compute Price Index is a volume-weighted, winsorized mean of executed on-demand rental transactions over a rolling one-hour window, from verified providers. It is not the CME settlement index, which is published by Silicon Data.

Data files used

filesha256 (first twelve characters)bytesretrieved (UTC)
seed/inferencex/gpt-oss-120b/2025-10-20.json89d6cb8ddb65549,7752026-09-07T17:45:53Z
seed/inferencex/gpt-oss-120b/2025-11-15.jsonf34926a8a9e5542,6452026-09-07T17:45:55Z
seed/inferencex/gpt-oss-120b/2025-12-15.json93d9316d9673564,0492026-09-07T17:45:58Z
seed/inferencex/gpt-oss-120b/2026-01-15.json629f4626e6b9635,9812026-09-07T17:46:01Z
seed/inferencex/gpt-oss-120b/2026-02-15.jsond3d2dc4178e1636,0912026-09-07T17:46:04Z
seed/inferencex/gpt-oss-120b/2026-03-15.jsonc8781fc84a1c638,8332026-09-07T17:46:06Z
seed/inferencex/gpt-oss-120b/2026-04-15.json192e1dc64c9c680,0472026-09-07T17:46:09Z
seed/inferencex/gpt-oss-120b/2026-05-15.jsonb3ec642040e2701,3782026-09-07T17:46:12Z
seed/inferencex/gpt-oss-120b/2026-06-15.jsonacabe31fb7fc752,1422026-09-07T17:46:14Z
seed/ornn/h100.json07a517ba660f5,5762026-09-07T17:43:42Z

Scenario: index volatility 35 percent a year, index drift 0 percent a year, basis volatility 10 percent, risk premium 0 percent, settlement correlation 1.0, demand forecast error 10 percent a month, growth 3 percent a month, 20,000 paths, seed 20260908. Intake hash e84bdffdedd5.

This report is a simulation under stated scenarios, calibrated to public benchmark, index and trace data as of the retrieval dates above and to the inputs you provided. It is not a forecast and not financial advice. The settlement index for CME compute futures is published by Silicon Data; prices here are from Ornn's transaction-based index. Serving-yield dynamics are estimated from public benchmarks for the stated model and hardware over the stated period and may not transfer to your stack. Results are conditional on the model, hardware, workload construction, period and scenarios examined.

Page 7 of 7