For funders · one launch

What the $200k buys.

At a planning band of $3 to $4 per H200-hour, $200k buys approximately 50,000 to 67,000 H200-hours. Depending on the model size, 10 to 15 benchmarks across two or three checkpoints or model variants will be a meaningful jump in safety testing.

TaskComputeNotes
Initial benchmarks$20k–$30kTen to fifteen benchmarks across initial checkpoints and scaffolds
Replication and elicitation$30k–$50kRepeats, alternate scaffolds, capability elicitation and confirmation runs
Agentic and human-directed red-teaming$40k–$70kLong-horizon environments and adaptive branches from observed failures
Interpretability and diagnosis$25k–$50kMulti-week H200-node blocks, activation capture and targeted analysis
Failure analysis and retesting$20k–$40kReproducing findings and repeating affected evaluations after a voluntary model change
Storage, egress and variance reserve$10k–$20kLarge intermediate artifacts, failed runs, provider variance and workload variance
Total range$145k–$260kCampaign-dependent

This type of compute arrangement is identical to the existing workstream of many labs. As a partner we help reproduce and diagnose findings, retest, and run safety fine-tuning before launch where a lab wants help fixing what we found.

A clean model may finish well below $200k and a difficult model may run well above it. Unused compute rolls into the next launch, so this number is an approximation rather than a hard cap. The moment an evaluation discovers something important is the wrong moment to run out of compute; we need to stay flexible in the spend while making the cluster fair to use for all partners involved.

A year

We cannot predict how many open-weight models will launch next year. Last year is the best guide: 42 releases with 1,000 or more Hugging Face likes. At the $200k estimate per launch, a year like that costs $8.4M in compute. At the table's range, $6.1M to $10.9M.

Counts from the ledger, last 365 days. The per-launch estimate is the table above.

Back to funders