Research program
The eight foundations
The research program is organized around eight mutually reinforcing foundations. They are not presented as immutable truths. They are the project's current best understanding of the infrastructure needed for frontier AI evaluation, and they are intended to evolve as evidence improves.
Foundation 1
Dynamic Evaluation Protocols
Static benchmarks inevitably lose relevance as models improve. Evaluation systems should evolve alongside capability growth through continuously updated methodologies, adaptive testing, and ongoing validation.
Foundation 2
Held-Out Evaluations
High-consequence evaluations require mechanisms that reduce benchmark leakage and gaming while preserving scientific validity. Secure, independently managed evaluation resources become increasingly important as capabilities grow.
Foundation 3
High-Stakes Capability Evaluation
Some capabilities warrant greater scrutiny because of their potential societal impact. Evaluation infrastructure should distinguish between ordinary capabilities and domains where evidence quality must be substantially stronger.
Foundation 4
Independent Expert Review
Trust increases when evaluation incorporates diverse, technically qualified, independent expertise. Independent review strengthens credibility while reducing institutional blind spots and conflicts of interest.
Foundation 5
Third-Party Auditor Ecosystem
Long-term evaluation capacity is unlikely to scale through developers alone. An ecosystem of qualified, accountable, independent evaluators can improve resilience, specialization, and public confidence.
Foundation 6
Progressive Standards and Requirements
Institutional expectations rarely emerge all at once. Voluntary practices may gradually mature into industry norms, procurement expectations, insurance requirements, certification mechanisms, or other structured forms over time.
Foundation 7
Incentives and Prestige
Organizations respond to incentives. Recognition, credibility, reputation, transparency, and demonstrated excellence should increasingly reward participation in rigorous evaluation rather than mere claims.
Foundation 8
Global Interoperability
Frontier AI development is international. Evaluation infrastructure should strive for compatibility across jurisdictions while respecting regional differences and encouraging collaboration instead of fragmentation.
Public editions of all eight foundation papers and the project's foundational sources are available in the Library, each with its version, status, downloads, and correction record.
Read the foundations overview →