Applications to attend are open · call for posters closes 2 October

Why we run this forum

The full reasoning behind the 2026 programme — the problem, its causes, and the three object-level areas where technical work most directly unblocks governance.

The evidence dilemma

Policymakers face a genuine bind. Acting on limited evidence risks policy that is ineffective or actively harmful. Waiting for stronger evidence leaves society exposed to misuse, loss of control and systemic failure.

That reads like a judgement call about risk appetite. We think it is not. The dilemma is produced by three specific, addressable failures in how evidence about frontier systems is generated, shared and checked — which is the whole premise of the forum: use technical solutions to guide the policies we adopt.

Three focus areas

As frontier AI systems become more capable, governance depends increasingly on technical facts and infrastructure: what systems can do, how much compute they use, and whether claims about them can be independently verified. Laws are hard to enforce if governments cannot measure capabilities or monitor deployed systems. Technical AI governance asks how that substrate can make governance more informed, precise and enforceable.

Each answers one of the three causes, and is briefed against it below.

Evidence exists — but isn’t shared

Developers hold the training data, the internal evaluations and the usage telemetry. What reaches policymakers and independent researchers is self-reported and unverifiable, and the cost of frontier-scale compute prices replication out of reach.

Trustworthy auditing

Third-party access & risk modelling

Third-party evaluators work through black-box APIs, on roughly a week’s notice, and rarely get a reliable upper bound on what a model can do. The map from a benchmark score to a real-world risk pathway stays largely undrawn — which is precisely the map policymakers need in order to write a guardrail that binds.

Read the brief

What blocks it

Developers hold the training data, internal evaluations and usage telemetry. What reaches the outside world is self-reported and unverifiable. Model scorecards restate the problem rather than solve it. Frontier-scale compute prices independent replication out of reach.

What we want in the room

Evaluators who can say what structured access would actually let them measure; policy teams who can say what finding would change a decision; and the people who have to translate one into the other.

Questions on the table

What does a defensible upper bound on capability look like? What evaluation window is long enough for long-horizon tasks? What would third-party access look like if it were designed for statistical power rather than reputational cover?

Background reading
Frontier AI auditing: toward rigorous third-party assessment ↗

Evidence exists — but can’t be verified

Even well-founded policy stalls at enforcement. States have commercial and strategic reasons to move fast, and no mechanism to confirm that a counterpart is doing what it says it is doing.

Verifiable AI

Hardware verification, scalable oversight & international coordination

An agreement is only as good as the ability to check it. Tamper-resistant hardware that keeps a trustworthy record of what a chip was used for — how much compute ran, whether an evaluation took place — turns unverifiable self-reporting into claims a counterparty can test.

Read the brief

What blocks it

States face commercial and strategic incentives to move quickly, and have no way to confirm that a counterpart is doing what it claims. Coordination fails not because the policy is wrong but because compliance is uncheckable.

Why hardware

A chip that can produce an unforgeable certificate about its own usage makes a claim auditable without exposing model weights or commercial secrets. That is the rare mechanism that survives adversarial incentives. ERA is building hardware security as a sub-field of AI safety, and this track is where that work meets the people who would have to enforce it.

Questions on the table

What does technical compliance actually look like on the ground? Where does scalable oversight substitute for verification, and where can it not? Which claims are worth verifying first?

Background reading
Governing through the cloud ↗
Faster AI diffusion through hardware-based verification ↗
Recommended directions for scalable oversight ↗

Evidence doesn’t exist yet

Safety teams get weeks to design policy they could plausibly implement. Evaluators get days to test. Long-horizon and multi-agent capability has no settled measurement — so some risks are simply unmapped.

AI R&D automation

Measuring its effects on AI progress and oversight

As models take on more of the research loop, the pace of capability gain and the difficulty of overseeing it move together. This track asks what can actually be measured about automated AI R&D — before the measurement problem outruns the measurers.

Read the brief

What blocks it

Models are beginning to distinguish evaluation settings from real deployment, and to behave differently in each. A dangerous capability that only shows up outside the test is a capability the governance pipeline never sees.

The under-examined case

It is not clear that multi-agent systems are being evaluated as systems at all — or whether composition introduces risk that no single-model evaluation would surface.

Questions on the table

How much does automated R&D actually accelerate progress, and by what measure? What does a valid evaluation look like when the subject can tell it is being evaluated? Who is responsible for evaluating a system of agents?

Background reading
Measuring AI R&D automation ↗

The working pathways

Framing justifies the programme; pathways are what we actually design against. Every session format on the programme exists to serve one of these.

Connecting the spheres

Technical AI governance is a real sub-field now, but a nascent one. There are still few people who understand both sides well. Putting evaluators, engineers and officials in sustained contact for two days is itself the intervention.

Fostering collaborations

A paper on arXiv is not the end of the job. We schedule introductions before the forum, run office hours and 1:1s throughout, and keep the poster hall and organisation fair open all day so that a conversation can become a working relationship.

Bringing talent in

The organisation fair, the founder pitch sessions and the dedicated ERA Fellowship poster session exist to move people and funding into the field — not just to discuss it.

Design