Standards Body Standards Body

Independent research and institutional design

Foundations for Frontier AI

Standards Body develops research, frameworks, shared terminology, and institutional designs for more credible frontier AI evaluation, assurance, standards, and governance.

Standards Body is not currently a regulator, accreditation body, certification body, or governmental authority.

Identity and authority

What Standards Body is today

Standards Body is an independent research and institutional-design project developing foundations for frontier AI evaluation, assurance, standards, and governance. Its present work consists of research, framework development, shared terminology, institutional design, and public knowledge infrastructure. Its long-term purpose is to help make consequential frontier AI claims more credible, current, independently reviewable, and internationally interpretable.

The project holds authority only over its own research and publications. It does not regulate, certify, accredit, approve evaluators, or set requirements for anyone. That boundary is stated on every publication and is explained in full on a dedicated page.

What Standards Body can and cannot do →

Why this work exists

Frontier AI needs stronger institutional foundations

Frontier AI is developing faster than the institutions responsible for evaluating and governing it. Developers largely evaluate their own models. Public benchmarks become optimized and stale. External evaluators receive inconsistent access. Evaluation methods change without common version control, and terms like audit, certification, and accreditation are used imprecisely across the field.

History suggests durable institutions are rarely built by beginning with enforcement. Medicine required clinical methods before licensing systems. Engineering required measurement before building codes. Aviation required investigation before regulation. Frontier AI, by the same pattern, requires evaluation infrastructure before reliable governance.

Standards Body therefore prioritizes foundational infrastructure before institutional authority: shared language, shared measurements, shared evidence disciplines, and institutional designs that others can inspect, criticize, and improve.

Research program

The eight foundations

The research program is organized around eight mutually reinforcing foundations. They are not presented as immutable truths. They are the project's current best understanding of the infrastructure needed for frontier AI evaluation, and they are intended to evolve as evidence improves.

Foundation 1

Dynamic Evaluation Protocols

Static benchmarks inevitably lose relevance as models improve. Evaluation systems should evolve alongside capability growth through continuously updated methodologies, adaptive testing, and ongoing validation.

Foundation 2

Held-Out Evaluations

High-consequence evaluations require mechanisms that reduce benchmark leakage and gaming while preserving scientific validity. Secure, independently managed evaluation resources become increasingly important as capabilities grow.

Foundation 3

High-Stakes Capability Evaluation

Some capabilities warrant greater scrutiny because of their potential societal impact. Evaluation infrastructure should distinguish between ordinary capabilities and domains where evidence quality must be substantially stronger.

Foundation 4

Independent Expert Review

Trust increases when evaluation incorporates diverse, technically qualified, independent expertise. Independent review strengthens credibility while reducing institutional blind spots and conflicts of interest.

Foundation 5

Third-Party Auditor Ecosystem

Long-term evaluation capacity is unlikely to scale through developers alone. An ecosystem of qualified, accountable, independent evaluators can improve resilience, specialization, and public confidence.

Foundation 6

Progressive Standards and Requirements

Institutional expectations rarely emerge all at once. Voluntary practices may gradually mature into industry norms, procurement expectations, insurance requirements, certification mechanisms, or other structured forms over time.

Foundation 7

Incentives and Prestige

Organizations respond to incentives. Recognition, credibility, reputation, transparency, and demonstrated excellence should increasingly reward participation in rigorous evaluation rather than mere claims.

Foundation 8

Global Interoperability

Frontier AI development is international. Evaluation infrastructure should strive for compatibility across jurisdictions while respecting regional differences and encouraging collaboration instead of fragmentation.

Public editions of all eight foundation papers and the project's foundational sources are available in the Library, each with its version, status, downloads, and correction record.

Read the foundations overview →

Library

From the Library

Every Standards Body publication carries an explicit status, version, and authority notice, with immutable released versions and a visible correction route.

Public Foundational Essay SB-PUB-2026-0001 Version 1.0 Published July 17, 2026 Current

Static Benchmarks Are Not Enough: Why Frontier AI Evaluation Must Be Continuously Maintained

Static benchmarks remain useful, but they are insufficient as the sole basis for consequential frontier AI evaluation. This essay argues that such evaluation should be organized through versioned, governed, and continuously maintained protocols.

Correspondence

Review, criticism, and corrections are welcome

Standards Body invites factual error reports, citation problems, methodological criticism, and status or authority concerns about anything it publishes. Corrections, supersession, and withdrawal are recorded, and released versions are preserved without silent edits.

Standards Body does not operate a confidential or anonymous submission channel. Anyone seeking to raise confidential concerns about an AI developer should use established legal and whistleblower channels rather than this site.

Contact and correction routes →