Know what your AI will cost to run, before the bill arrives.

Describe the workload once. Furrow reads public benchmarks and prices, plays out twenty thousand futures, and writes a seven-page report.

Built on public data from

InferenceXOrnnAzure traces
OrnnAzure tracesCME Group
Azure tracesCME GroupSilicon DataTokenYield
CME GroupSilicon DataTokenYield
Silicon DataTokenYieldInferenceX
TokenYieldInferenceXOrnn

[

FEATURES

]

Four numbers. Every one of them tagged.

0.1

Two-phase cost

$ furrow run

chat_h100.json

> fitting yields...

> simulating 20,000 paths

step 4 of 6

80%

> rendering report_

Prefill and decode, priced separately

A chat mix is 94 percent prompt tokens. A single 1k/1k benchmark overstates its cost by 59 percent. Furrow fits both yields and prices your mix.

0.2

Variance shares

variance/

├── price 37%

├── demand 44%

└── yield 15%

├── basis 6%

└── utilization 0%

└── sum 103%

20,000 paths · 5 factors

Where the bill actually moves

Twenty thousand simulated futures, then a Shapley split of the cost variance across price, demand, yield and basis.

0.3

Procurement policies

reserve baseline

rank 1

48 GPUs · chat example

six-month cost · 85% of on demand

P(overrun) 5 percent

ES95 +41 percent

variance removed 44 percent

Reserve, buy on demand, or hedge

Six policies on the same simulated paths: expected cost, variance removed, overrun probability, downside and idle capacity.

[

USE CASES

]

Every question the compute line gets asked. Furrow answers it.

Renew a contract

Plan capacity

Plan capacity

Price per token

Price per token

Right-size the reservation.

Enter your fleet, rate and months remaining. Furrow puts your contract beside the alternatives with idle capacity and overflow made explicit.

Baseline, peak and p95 fleets compared

Idle cash and overflow on demand shown

Ranking under the discount you were quoted

Renew a contract

Renew a contract

Plan capacity

Price per token

Price per token

From traffic to GPU-hours, with progress built in.

A quarterly plan that carries expected efficiency progress and a forecast error you chose, instead of a spreadsheet conversion at today's throughput.

GPU-hours for months one to twelve

Expected yield progress applied month by month

PR with detailed summary of every change made

Renew a contract

Renew a contract

Plan capacity

Plan capacity

Price per token

The real cost of a million tokens.

Prefill and decode priced separately, then your rate and the utilization your demand shape allows, with the single-benchmark figure beside it.

Two-phase fit to public benchmarks

Your rate and your utilization, not the index

Runs the suite and fixes failures before the PR

[

HOW IT WORKS

]

Three steps. Then read seven pages.

1

Describe the workload

About ten questions your infrastructure lead already knows: tokens per request, requests a month, latency floor, GPU, contract.

2

Furrow simulates

Twenty thousand futures for price, demand, yield and basis, calibrated to public benchmarks, index history and traffic traces. A few seconds on a laptop.

3

Read and hand over

A seven-page report, every number tagged observed, derived, assumed, scenario or simulated. Print it to A4 or walk it on a call.

[

BENEFITS

]

Less guesswork. More numbers you can defend.

Nothing fabricated

Every number carries one of five tags. If a value was assumed, the page says so beside it.

Runs on your laptop

No login, no server, no upload. Customer files live in one folder on your machine.

Reproducible to the digit

One command regenerates the paper's 970 numbers from archived data and fixed seeds.

Seconds, not weeks

Intake to seven pages in a few seconds. A scenario re-run during a call takes a few more.

Honest ranges

Hedge effectiveness is shown as a range across scenarios, 18 to 64 percent for the chat example, not one point.

Public data, cited

InferenceX benchmarks, the Ornn index and Azure traces, each archived with its hash and retrieval time.

[

WORKLOADS

]

Five workloads. What the same engine says about each.

[

RESOURCES

]

Why this matters. In the words of the people building the market.

[

ACCESS

]

Three ways in. No fabricated prices.

One report

Quarterly

3 REPORTS

Sample

Free

Read the seven-page report for the chat example and the article behind it. No account, no form.

Features

Seven pages, every number tagged

gpt-oss-120b on H100, chat mix

Base scenario, 20,000 paths

Scope statement included

Your report

Operator-run

Quote

Bring your traffic numbers and on-demand rate. Furrow is run for you, one customer at a time, and walked through on a thirty-minute call.

Features

Your intake, your rate, your contract

Scenario re-runs during the call

PDF and HTML report delivered

A reservation quote grounds the reserved rows

Custom benchmark for uncovered models

Research

Open

The public package: the paper's engine, the yield dataset and the pre-registered observed phase.

Features

Engine reproduces the paper (MIT)

Yield dataset (CC BY 4.0)

Pre-registration with locked hashes

Fourteen anchor results checked

No customer data inside

One report

Quarterly

3 REPORTS

Sample

Free

Read the seven-page report for the chat example and the article behind it. No account, no form.

Features

Seven pages, every number tagged

gpt-oss-120b on H100, chat mix

Base scenario, 20,000 paths

Scope statement included

Your report

Operator-run

$23

Bring your traffic numbers and on-demand rate. Furrow is run for you, one customer at a time, and walked through on a thirty-minute call.

Features

Your intake, your rate, your contract

Scenario re-runs during the call

PDF and HTML report delivered

A reservation quote grounds the reserved rows

Custom benchmark for uncovered models

Research

Open

The public package: the paper's engine, the yield dataset and the pre-registered observed phase.

Features

Engine reproduces the paper (MIT)

Yield dataset (CC BY 4.0)

Pre-registration with locked hashes

Fourteen anchor results checked

No customer data inside

[

FAQ

]

Before you run it. Everything you need to know.

What does Furrow actually produce?

What does Furrow actually produce?

A seven-page HTML report from one intake file: GPU-hours, cost per million tokens, where the cost variance comes from, procurement policies compared, and what a GPU futures strip could and could not remove.

What does Furrow actually produce?

Is this a forecast?

Is this a forecast?

No. It is a simulation under stated scenarios, calibrated to public benchmark, index and trace data. Every number carries a tag saying whether it was observed, derived, assumed, a scenario or simulated.

Is this a forecast?

Which models and hardware are covered?

Which models and hardware are covered?

gpt-oss-120b and DeepSeek-R1-0528 on H100, H200 and B200 from public InferenceX snapshots. Any other model works with your own benchmark run, and the borrowed yield dynamics are labelled assumed.

Which models and hardware are covered?

Does my data leave my machine?

Does my data leave my machine?

No. Furrow runs offline from seeded public data. Customer files live in one folder on your laptop, and the only network calls fetch public benchmark and index data when you ask.

Does my data leave my machine?

Does it tell me to hedge?

Does it tell me to hedge?

It sizes a hypothetical futures position three ways and reports the variance each would remove, as a range across scenarios. It never recommends taking one.

Does it tell me to hedge?

Where do the numbers come from?

Where do the numbers come from?

TokenYield Working Paper 1 v0.2. The engine is the paper's code with two additive parameters, and one command regenerates every figure and macro of the paper from archived data.

Where do the numbers come from?

Your bill won't explain itself.

Give Furrow a workload. Get seven pages back.

Know what your AI will cost to run, before the bill arrives.

Describe the workload once. Furrow reads public benchmarks and prices, plays out twenty thousand futures, and writes a seven-page report.

Built on public data from

InferenceXOrnnAzure traces
OrnnAzure tracesCME Group
Azure tracesCME GroupSilicon DataTokenYield
CME GroupSilicon DataTokenYield
Silicon DataTokenYieldInferenceX
TokenYieldInferenceXOrnn

[

FEATURES

]

Four numbers. Every one of them tagged.

0.1

Two-phase cost

$ furrow run

chat_h100.json

> fitting yields...

> simulating 20,000 paths

step 4 of 6

80%

> rendering report_

Prefill and decode, priced separately

A chat mix is 94 percent prompt tokens. A single 1k/1k benchmark overstates its cost by 59 percent. Furrow fits both yields and prices your mix.

0.2

Variance shares

variance/

├── price 37%

├── demand 44%

└── yield 15%

├── basis 6%

└── utilization 0%

└── sum 103%

20,000 paths · 5 factors

Where the bill actually moves

Twenty thousand simulated futures, then a Shapley split of the cost variance across price, demand, yield and basis.

0.3

Procurement policies

reserve baseline

rank 1

48 GPUs · chat example

six-month cost · 85% of on demand

P(overrun) 5 percent

ES95 +41 percent

variance removed 44 percent

Reserve, buy on demand, or hedge

Six policies on the same simulated paths: expected cost, variance removed, overrun probability, downside and idle capacity.

[

USE CASES

]

Every question the compute line gets asked. Furrow answers it.

Renew a contract

Plan capacity

Price per token

Right-size the reservation.

Enter your fleet, rate and months remaining. Furrow puts your contract beside the alternatives with idle capacity and overflow made explicit.

Baseline, peak and p95 fleets compared

Idle cash and overflow on demand shown

Ranking under the discount you were quoted

Renew a contract

Plan capacity

Price per token

From traffic to GPU-hours, with progress built in.

A quarterly plan that carries expected efficiency progress and a forecast error you chose, instead of a spreadsheet conversion at today's throughput.

GPU-hours for months one to twelve

Expected yield progress applied month by month

PR with detailed summary of every change made

Renew a contract

Plan capacity

Price per token

The real cost of a million tokens.

Prefill and decode priced separately, then your rate and the utilization your demand shape allows, with the single-benchmark figure beside it.

Two-phase fit to public benchmarks

Your rate and your utilization, not the index

Runs the suite and fixes failures before the PR

[

HOW IT WORKS

]

Three steps. Then read seven pages.

1

Describe the workload

About ten questions your infrastructure lead already knows: tokens per request, requests a month, latency floor, GPU, contract.

2

Furrow simulates

Twenty thousand futures for price, demand, yield and basis, calibrated to public benchmarks, index history and traffic traces. A few seconds on a laptop.

3

Read and hand over

A seven-page report, every number tagged observed, derived, assumed, scenario or simulated. Print it to A4 or walk it on a call.

[

BENEFITS

]

Less guesswork. More numbers you can defend.

Nothing fabricated

Every number carries one of five tags. If a value was assumed, the page says so beside it.

Runs on your laptop

No login, no server, no upload. Customer files live in one folder on your machine.

Reproducible to the digit

One command regenerates the paper's 970 numbers from archived data and fixed seeds.

Seconds, not weeks

Intake to seven pages in a few seconds. A scenario re-run during a call takes a few more.

Honest ranges

Hedge effectiveness is shown as a range across scenarios, 18 to 64 percent for the chat example, not one point.

Public data, cited

InferenceX benchmarks, the Ornn index and Azure traces, each archived with its hash and retrieval time.

[

WORKLOADS

]

Five workloads. What the same engine says about each.

[

RESOURCES

]

Why this matters. In the words of the people building the market.

[

ACCESS

]

Three ways in. No fabricated prices.

One report

Quarterly

3 REPORTS

Sample

Free

Read the seven-page report for the chat example and the article behind it. No account, no form.

Features

Seven pages, every number tagged

gpt-oss-120b on H100, chat mix

Base scenario, 20,000 paths

Scope statement included

Your report

Operator-run

Quote

Bring your traffic numbers and on-demand rate. Furrow is run for you, one customer at a time, and walked through on a thirty-minute call.

Features

Your intake, your rate, your contract

Scenario re-runs during the call

PDF and HTML report delivered

A reservation quote grounds the reserved rows

Custom benchmark for uncovered models

Research

Open

The public package: the paper's engine, the yield dataset and the pre-registered observed phase.

Features

Engine reproduces the paper (MIT)

Yield dataset (CC BY 4.0)

Pre-registration with locked hashes

Fourteen anchor results checked

No customer data inside

One report

Quarterly

3 REPORTS

Sample

Free

Read the seven-page report for the chat example and the article behind it. No account, no form.

Features

Seven pages, every number tagged

gpt-oss-120b on H100, chat mix

Base scenario, 20,000 paths

Scope statement included

Your report

Operator-run

$23

Bring your traffic numbers and on-demand rate. Furrow is run for you, one customer at a time, and walked through on a thirty-minute call.

Features

Your intake, your rate, your contract

Scenario re-runs during the call

PDF and HTML report delivered

A reservation quote grounds the reserved rows

Custom benchmark for uncovered models

Research

Open

The public package: the paper's engine, the yield dataset and the pre-registered observed phase.

Features

Engine reproduces the paper (MIT)

Yield dataset (CC BY 4.0)

Pre-registration with locked hashes

Fourteen anchor results checked

No customer data inside

[

FAQ

]

Before you run it. Everything you need to know.

What does Furrow actually produce?

What does Furrow actually produce?

A seven-page HTML report from one intake file: GPU-hours, cost per million tokens, where the cost variance comes from, procurement policies compared, and what a GPU futures strip could and could not remove.

What does Furrow actually produce?

Is this a forecast?

Is this a forecast?

No. It is a simulation under stated scenarios, calibrated to public benchmark, index and trace data. Every number carries a tag saying whether it was observed, derived, assumed, a scenario or simulated.

Is this a forecast?

Which models and hardware are covered?

Which models and hardware are covered?

gpt-oss-120b and DeepSeek-R1-0528 on H100, H200 and B200 from public InferenceX snapshots. Any other model works with your own benchmark run, and the borrowed yield dynamics are labelled assumed.

Which models and hardware are covered?

Does my data leave my machine?

Does my data leave my machine?

No. Furrow runs offline from seeded public data. Customer files live in one folder on your laptop, and the only network calls fetch public benchmark and index data when you ask.

Does my data leave my machine?

Does it tell me to hedge?

Does it tell me to hedge?

It sizes a hypothetical futures position three ways and reports the variance each would remove, as a range across scenarios. It never recommends taking one.

Does it tell me to hedge?

Where do the numbers come from?

Where do the numbers come from?

TokenYield Working Paper 1 v0.2. The engine is the paper's code with two additive parameters, and one command regenerates every figure and macro of the paper from archived data.

Where do the numbers come from?

Your bill won't explain itself.

Give Furrow a workload. Get seven pages back.

Know what your AI will cost to run, before the bill arrives.

Describe the workload once. Furrow reads public benchmarks and prices, plays out twenty thousand futures, and writes a seven-page report.

Built on public data from

InferenceXOrnnAzure traces
OrnnAzure tracesCME Group
Azure tracesCME GroupSilicon DataTokenYield
CME GroupSilicon DataTokenYield
Silicon DataTokenYieldInferenceX
TokenYieldInferenceXOrnn

[

FEATURES

]

Four numbers. Every one of them tagged.

0.1

Two-phase cost

$ furrow run

chat_h100.json

> fitting yields...

> simulating 20,000 paths

step 4 of 6

80%

> rendering report_

Prefill and decode, priced separately

A chat mix is 94 percent prompt tokens. A single 1k/1k benchmark overstates its cost by 59 percent. Furrow fits both yields and prices your mix.

0.2

Variance shares

variance/

├── price 37%

├── demand 44%

└── yield 15%

├── basis 6%

└── utilization 0%

└── sum 103%

20,000 paths · 5 factors

Where the bill actually moves

Twenty thousand simulated futures, then a Shapley split of the cost variance across price, demand, yield and basis.

0.3

Procurement policies

reserve baseline

rank 1

48 GPUs · chat example

six-month cost · 85% of on demand

P(overrun) 5 percent

ES95 +41 percent

variance removed 44 percent

Reserve, buy on demand, or hedge

Six policies on the same simulated paths: expected cost, variance removed, overrun probability, downside and idle capacity.

[

USE CASES

]

Every question the compute line gets asked. Furrow answers it.

Renew a contract

Plan capacity

Price per token

Right-size the reservation.

Enter your fleet, rate and months remaining. Furrow puts your contract beside the alternatives with idle capacity and overflow made explicit.

Baseline, peak and p95 fleets compared

Idle cash and overflow on demand shown

Ranking under the discount you were quoted

Renew a contract

Plan capacity

Price per token

From traffic to GPU-hours, with progress built in.

A quarterly plan that carries expected efficiency progress and a forecast error you chose, instead of a spreadsheet conversion at today's throughput.

GPU-hours for months one to twelve

Expected yield progress applied month by month

PR with detailed summary of every change made

Renew a contract

Plan capacity

Price per token

The real cost of a million tokens.

Prefill and decode priced separately, then your rate and the utilization your demand shape allows, with the single-benchmark figure beside it.

Two-phase fit to public benchmarks

Your rate and your utilization, not the index

Runs the suite and fixes failures before the PR

[

HOW IT WORKS

]

Three steps. Then read seven pages.

1

Describe the workload

About ten questions your infrastructure lead already knows: tokens per request, requests a month, latency floor, GPU, contract.

2

Furrow simulates

Twenty thousand futures for price, demand, yield and basis, calibrated to public benchmarks, index history and traffic traces. A few seconds on a laptop.

3

Read and hand over

A seven-page report, every number tagged observed, derived, assumed, scenario or simulated. Print it to A4 or walk it on a call.

[

BENEFITS

]

Less guesswork. More numbers you can defend.

Nothing fabricated

Every number carries one of five tags. If a value was assumed, the page says so beside it.

Runs on your laptop

No login, no server, no upload. Customer files live in one folder on your machine.

Reproducible to the digit

One command regenerates the paper's 970 numbers from archived data and fixed seeds.

Seconds, not weeks

Intake to seven pages in a few seconds. A scenario re-run during a call takes a few more.

Honest ranges

Hedge effectiveness is shown as a range across scenarios, 18 to 64 percent for the chat example, not one point.

Public data, cited

InferenceX benchmarks, the Ornn index and Azure traces, each archived with its hash and retrieval time.

[

WORKLOADS

]

Five workloads. What the same engine says about each.

[

RESOURCES

]

Why this matters. In the words of the people building the market.

[

ACCESS

]

Three ways in. No fabricated prices.

One report

Quarterly

3 REPORTS

Sample

Free

Read the seven-page report for the chat example and the article behind it. No account, no form.

Features

Seven pages, every number tagged

gpt-oss-120b on H100, chat mix

Base scenario, 20,000 paths

Scope statement included

Your report

Operator-run

Quote

Bring your traffic numbers and on-demand rate. Furrow is run for you, one customer at a time, and walked through on a thirty-minute call.

Features

Your intake, your rate, your contract

Scenario re-runs during the call

PDF and HTML report delivered

A reservation quote grounds the reserved rows

Custom benchmark for uncovered models

Research

Open

The public package: the paper's engine, the yield dataset and the pre-registered observed phase.

Features

Engine reproduces the paper (MIT)

Yield dataset (CC BY 4.0)

Pre-registration with locked hashes

Fourteen anchor results checked

No customer data inside

One report

Quarterly

3 REPORTS

Sample

Free

Read the seven-page report for the chat example and the article behind it. No account, no form.

Features

Seven pages, every number tagged

gpt-oss-120b on H100, chat mix

Base scenario, 20,000 paths

Scope statement included

Your report

Operator-run

$23

Bring your traffic numbers and on-demand rate. Furrow is run for you, one customer at a time, and walked through on a thirty-minute call.

Features

Your intake, your rate, your contract

Scenario re-runs during the call

PDF and HTML report delivered

A reservation quote grounds the reserved rows

Custom benchmark for uncovered models

Research

Open

The public package: the paper's engine, the yield dataset and the pre-registered observed phase.

Features

Engine reproduces the paper (MIT)

Yield dataset (CC BY 4.0)

Pre-registration with locked hashes

Fourteen anchor results checked

No customer data inside

[

FAQ

]

Before you run it. Everything you need to know.

What does Furrow actually produce?

What does Furrow actually produce?

A seven-page HTML report from one intake file: GPU-hours, cost per million tokens, where the cost variance comes from, procurement policies compared, and what a GPU futures strip could and could not remove.

What does Furrow actually produce?

Is this a forecast?

Is this a forecast?

No. It is a simulation under stated scenarios, calibrated to public benchmark, index and trace data. Every number carries a tag saying whether it was observed, derived, assumed, a scenario or simulated.

Is this a forecast?

Which models and hardware are covered?

Which models and hardware are covered?

gpt-oss-120b and DeepSeek-R1-0528 on H100, H200 and B200 from public InferenceX snapshots. Any other model works with your own benchmark run, and the borrowed yield dynamics are labelled assumed.

Which models and hardware are covered?

Does my data leave my machine?

Does my data leave my machine?

No. Furrow runs offline from seeded public data. Customer files live in one folder on your laptop, and the only network calls fetch public benchmark and index data when you ask.

Does my data leave my machine?

Does it tell me to hedge?

Does it tell me to hedge?

It sizes a hypothetical futures position three ways and reports the variance each would remove, as a range across scenarios. It never recommends taking one.

Does it tell me to hedge?

Where do the numbers come from?

Where do the numbers come from?

TokenYield Working Paper 1 v0.2. The engine is the paper's code with two additive parameters, and one command regenerates every figure and macro of the paper from archived data.

Where do the numbers come from?

Your bill won't explain itself.

Give Furrow a workload. Get seven pages back.