Proposed result architecture · design contract

One evaluation cannot responsibly
produce one kind of truth.

Facts, measurement, and credentialing require different evidence and different decision authority. The benchmark keeps them visibly separate.

Inspect the research position
01
Factual layerSelected layer

Verified production record

What happened, where it happened, and what this person actually contributed.

  • 01Attributable artifacts
  • 02Operating context
  • 03Contribution boundary
  • 04Reference corroboration
Decision authorityEvidence operations

Records preserve facts and provenance. They do not infer capability or confer standing.

Proposed shared work challenge

Same production problem.
Different operating conditions.

A common task substrate can examine a practitioner, an AI system, and the two working together. It does not presume their scores are interchangeable.

Common task substrateShared objective · declared condition adapters and resources
  1. 01
    FrameDiagnose the consequential problem
  2. 02
    BuildCreate and defend the intervention
  3. 03
    OperateRespond to a live failure injection
  4. 04
    ProveLeave tests, decisions, and outcome evidence
Condition-applicable rubrics, assistance rules, and evidence-capture plans are frozen before an attempt begins.
Active conditionHuman + AI
Subject
Operating pair
Observed
Delegation, verification, intervention, escalation, and accountability for the outcome.
Disclosure
Contribution is resolved across the practitioner, system, and operating environment.

Comparability is a research question. Shared tasks do not justify a shared human–model leaderboard without empirical equivalence evidence.

Research design only

No challenge has been administered, no attempt data are admitted, and no human, model, team, score, tier, or credential is represented here.

Release disclosure · evidence cutoff 1 Sep 2026

What this release can support.

Transparency belongs in the record—not on a victory scoreboard. Current absences remain inspectable below.

Public draftv0.1.0
ConstructPublished10 domains · 55 families
Human studyNot startedNo participant evidence admitted
Official scoringNot authorizedNo cut score, tier, or rank
External validationNot recordedIndependent review remains a gate
Open the current evidence boundary

No admitted human-participant record supports a capability, hiring, ranking, tier, or credential interpretation.

No independent validation, accreditation, external-majority commission approval, or reproduction is represented.

Publication readiness is tracked separately from measurement validity; neither can be advanced by interface polish.