Fund evaluations before models ship.
Fund evaluations that help labs find and understand dangerous capabilities before a model is released. Support the benchmarks, replication and follow-up that make a finding useful.
Establish the team, evaluation tools and compute capacity, and run evals on a suite of open frontier model releases in Q4.
Sustain evaluation delivery and supporting open model safety research on both sides of the Pacific.
Programme funding keeps researchers available, evaluation tools maintained and compute ready for short release windows.
Safety research competes with capability research.
The same GPUs can make a model more capable or help researchers understand its risks. When those jobs share a budget, capability research can take priority. We give safety research compute of its own.
Why the GPUs matter
At the lab
Target modelWeights stay here
Pacific Compute
Evaluation compute- Attacker models
- Judge models
- Parallel experiments
- Method development
Endpoint-based evaluation · The first H200 node is rented and live.