How this benchmark is built

A benchmark is only as good as
the part you are allowed to inspect.

Most assessments ask you to trust a number. This one publishes its construct, its evidence rules, its calculation boundary, and the list of things it cannot yet do. This page is that list. It is written to be read by someone deciding whether to take the project seriously.

Outcome architecture

Three outputs, three authorities, one direction of travel.

The most common request is for a single score. Facts, measurement, and certification answer different questions, carry different error modes, and belong to different decision-makers. Collapsing them into one number is what makes most people-scores indefensible. Each layer below may cite the one before it. None may overwrite it.

  1. Layer 1

    Verified production record

    Decided by Fact verification
    May claim
    State corroborated facts: a deployment existed in a declared window, this person contributed under a declared attribution class, an outcome measure was defined and observed.
    May not claim
    Infer capability, proficiency, readiness, a score, a tier, or a rank.
    Buildable now
  2. Layer 2

    Capability profile

    Decided by Human review and measurement
    May claim
    Give a versioned, multidimensional interpretation for a declared role context, carrying uncertainty, coverage, maturity, and contradictions.
    May not claim
    Produce a global number, claim universal readiness, or transfer to a role or use it was never studied on.
    Needs job analysis, anchored rubrics, measured rater agreement
  3. Layer 3

    Governed credential tier

    Decided by A non-conflicted credential authority
    May claim
    Issue a bounded classification with an issuer, version, expiry, status, and a route to appeal.
    May not claim
    Make a claim true because the payload is signed, or imply capability outside the stated interpretation.
    Needs a seated commission

A single score is not a setting that can be switched on. It would be a new interpretation and a new use, and it would have to pass a staged authorization gate before any release could carry it. Read the gate.

Build ledger

What is built, what is not, and what unlocks the next step.

Eight parts of the system, each with the same three columns. The middle column is the one most projects leave out.

Build state by area, showing what exists, what does not, and the gate that unlocks the next step
AreaBuiltNot builtNext gate
01ConstructTen domains and 55 capability families, versioned in a released ontology with stable identifiers and provenance.No competencies, atomic signals, or rubric criteria are released. Two of five ontology levels are populated.A job analysis that shows the families describe work people are actually paid to do.
02Evidence modelSix maturity levels from claimed to field-validated, assessed on the link between a claim and a capability rather than on a whole artifact. Correlated evidence is grouped by lineage so copies cannot count as corroboration.No real evidence has entered the model. Every fixture is synthetic.Consent, intake, retention, and deletion controls that a privacy reviewer signs.
03Applied workA proposed assessment task for each of the 55 families, with observable behaviors, a minimum evidence set, a constraint injected mid-session, and explicit limits on what the exercise cannot show.No task has been administered to anyone. All task text is public, which makes every one of them practice material rather than an operational form.Sequestered task forms, rotation, and exposure control written and enforced.
04Human reviewA blinded two-reader design with structured observations, reason codes, and automatic routing to an independent adjudicator when readers disagree.No reader is qualified, calibrated, or appointed. Agreement has never been measured.Anchored rubric levels, a rater calibration set, and a reported agreement coefficient.
05CalculationA deterministic projection over an immutable release: canonical JSON, integer basis points, digest-bound manifests, and a frozen replay that reproduces the same output from the same inputs.It runs on synthetic fixtures only. No weights are calibrated and no thresholds exist.A standard-setting study, and the thirteen receipts of the scalar authorization gate if a single number is ever proposed.
06AI evaluationA provider-neutral evaluator that can propose typed observations and structurally cannot write a score: its release manifest denies score and credential permissions, and artifact text stays untrusted data.The shipped adapter uses no learned model. No agreement study against human raters exists, and no bias audit has been run.A pre-registered agreement and counterfactual-bias study before any model output influences a reported result.
07GovernanceA published firewall specifying a seven-to-nine seat commission with an external majority, an external chair, a twenty-four month affiliation lookback, and an appointments panel that LockedIn cannot sit on.No member is seated. Every accountability role reads Unassigned. The stage is I0, founding stewardship.Two named external reviewers, then an interim review panel, then the commission.
08PublicationA method paper, a claims register that fails closed, a dated source ledger with an explicit non-adoption column, and 28 primary and methodological sources.No empirical finding has been released. No study has been fielded. Authors are unassigned and no archival identifier is claimed.Named authors, an archived version, and a finding that survived outside review.
The conflict, stated plainly

LockedIn Labs trains, employs, and staffs the people this benchmark would measure.

That is a real conflict and disclosure does not resolve it. The published firewall answers it structurally: the roles that pay for the work are separated from the roles that would decide an outcome. None of it is operational yet, and the register says so.

  • Trains the engineerBlocked from the rating pathNo rating role for anyone who trained the subject in the prior 24 months
  • Employs the engineerBlocked from the rating pathNo access to candidate evidence
  • Staffs them to clientsBlocked from the rating pathNo influence on any result
  • Operates the platformPermitted, under auditRuns the service under audit
  • Funds the programPermitted, under auditAppears in a public funder register, holds no vote

The commission that would hold every protected decision is specified but not seated. Until members exist, the correct description is proposed. Read the firewall or inspect the accountability record.

Check the work

Everything above is checkable.

  • The construct

    Every domain and family, with its released definition, proposed task, and evidence gap.

    Open the explorer
  • The method

    Construct, task design, evidence model, calculation boundary, limitations, and sources.

    Read the method paper
  • The claims

    What may be said today, what may not, and which gate changes each answer.

    Open the status register
  • The naming decision

    Why there is no headline score, and what the system is called instead.

    Read the decision record