Costs
What it costs.
| Work | Price to the user | Terms |
|---|---|---|
| Launch evaluation of a frontier open model | Free to the lab | $200k all-in per evaluation, funded by philanthropy |
| Compute for third-party evaluations of frontier open models | $1.55 per H200 GPU-hour | Set below published H200 rates |
| Academic, safety hub and independent safety research | At or below cost | Allocation depends on the proposed work and available capacity |
| Incident investigations | By arrangement | Supported with GPUs when capacity allows |
Compute is billed per H200 GPU-hour. Access, scope and capacity are agreed before work begins, and pre-release launch evaluations take priority during release windows.
37% below the median on-demand H200 rate of $2.47 per GPU-hour on September 28, 2026 (GPU Finder, 31 single-GPU listings). The cheapest H200 in stock that day was $2.00 at PrimeIntellect; Runpod Secure Cloud listed $4.59.What AI safety compute should cost ↗
Launch evaluations
How we assess a model for $200k.
A substantial campaign includes benchmarks, replication, targeted red-teaming, diagnosis and retesting. We budget $200k per major model launch, with smaller iterations being considerably less expensive.
We target three weeks from agreed kickoff through delivering findings and scoped retesting.
We record GPU hours, queue time and cost per completed evaluation and release these metrics after anonymizing the model and lab we worked with. We are always thinking about where the constraints are greatest (hardware, lab endpoints, researcher time, eval complexity) and work to address these bottlenecks.
See the campaign sequenceOne scoped evaluation
- Evaluation compute and environmentsIncluded
- Research and delivery staffIncluded
- Investigation, reproduction and scoped retestingIncluded
- Storage and data transferIncluded
- Allocated operations and contingencyIncluded
- Private findings and resource statementIncluded
Risks and Cautions
Compute is fungible. Free evaluation can release a lab’s own budget for capability work. Usage restrictions and monitoring reduce particular risks; they cannot eliminate that indirect effect.
Evaluation can also help a lab improve its models. We accept that trade-off where we expect the safety benefit to outweigh the capability benefit. Our case is that these models are likely to be released anyway, and that additional scrutiny before release can expose risks while there is still time to respond.
Let’s scope a grant.
We can walk through the campaign plan, budget and remaining technical questions.