A benchmark is only as good as
the part you are allowed to inspect.
Most assessments ask you to trust a number. This one publishes its construct, its evidence rules, its calculation boundary, and the list of things it cannot yet do. This page is that list. It is written to be read by someone deciding whether to take the project seriously.
Three outputs, three authorities, one direction of travel.
The most common request is for a single score. Facts, measurement, and certification answer different questions, carry different error modes, and belong to different decision-makers. Collapsing them into one number is what makes most people-scores indefensible. Each layer below may cite the one before it. None may overwrite it.
Layer 1 Verified production record
Decided by Fact verification- May claim
- State corroborated facts: a deployment existed in a declared window, this person contributed under a declared attribution class, an outcome measure was defined and observed.
- May not claim
- Infer capability, proficiency, readiness, a score, a tier, or a rank.
Layer 2 Capability profile
Decided by Human review and measurement- May claim
- Give a versioned, multidimensional interpretation for a declared role context, carrying uncertainty, coverage, maturity, and contradictions.
- May not claim
- Produce a global number, claim universal readiness, or transfer to a role or use it was never studied on.
Layer 3 Governed credential tier
Decided by A non-conflicted credential authority- May claim
- Issue a bounded classification with an issuer, version, expiry, status, and a route to appeal.
- May not claim
- Make a claim true because the payload is signed, or imply capability outside the stated interpretation.
A single score is not a setting that can be switched on. It would be a new interpretation and a new use, and it would have to pass a staged authorization gate before any release could carry it. Read the gate.
What is built, what is not, and what unlocks the next step.
Eight parts of the system, each with the same three columns. The middle column is the one most projects leave out.
| Area | Built | Not built | Next gate |
|---|---|---|---|
| 01Construct | Ten domains and 55 capability families, versioned in a released ontology with stable identifiers and provenance. | No competencies, atomic signals, or rubric criteria are released. Two of five ontology levels are populated. | A job analysis that shows the families describe work people are actually paid to do. |
| 02Evidence model | Six maturity levels from claimed to field-validated, assessed on the link between a claim and a capability rather than on a whole artifact. Correlated evidence is grouped by lineage so copies cannot count as corroboration. | No real evidence has entered the model. Every fixture is synthetic. | Consent, intake, retention, and deletion controls that a privacy reviewer signs. |
| 03Applied work | A proposed assessment task for each of the 55 families, with observable behaviors, a minimum evidence set, a constraint injected mid-session, and explicit limits on what the exercise cannot show. | No task has been administered to anyone. All task text is public, which makes every one of them practice material rather than an operational form. | Sequestered task forms, rotation, and exposure control written and enforced. |
| 04Human review | A blinded two-reader design with structured observations, reason codes, and automatic routing to an independent adjudicator when readers disagree. | No reader is qualified, calibrated, or appointed. Agreement has never been measured. | Anchored rubric levels, a rater calibration set, and a reported agreement coefficient. |
| 05Calculation | A deterministic projection over an immutable release: canonical JSON, integer basis points, digest-bound manifests, and a frozen replay that reproduces the same output from the same inputs. | It runs on synthetic fixtures only. No weights are calibrated and no thresholds exist. | A standard-setting study, and the thirteen receipts of the scalar authorization gate if a single number is ever proposed. |
| 06AI evaluation | A provider-neutral evaluator that can propose typed observations and structurally cannot write a score: its release manifest denies score and credential permissions, and artifact text stays untrusted data. | The shipped adapter uses no learned model. No agreement study against human raters exists, and no bias audit has been run. | A pre-registered agreement and counterfactual-bias study before any model output influences a reported result. |
| 07Governance | A published firewall specifying a seven-to-nine seat commission with an external majority, an external chair, a twenty-four month affiliation lookback, and an appointments panel that LockedIn cannot sit on. | No member is seated. Every accountability role reads Unassigned. The stage is I0, founding stewardship. | Two named external reviewers, then an interim review panel, then the commission. |
| 08Publication | A method paper, a claims register that fails closed, a dated source ledger with an explicit non-adoption column, and 28 primary and methodological sources. | No empirical finding has been released. No study has been fielded. Authors are unassigned and no archival identifier is claimed. | Named authors, an archived version, and a finding that survived outside review. |
LockedIn Labs trains, employs, and staffs the people this benchmark would measure.
That is a real conflict and disclosure does not resolve it. The published firewall answers it structurally: the roles that pay for the work are separated from the roles that would decide an outcome. None of it is operational yet, and the register says so.
- Trains the engineerBlocked from the rating pathNo rating role for anyone who trained the subject in the prior 24 months
- Employs the engineerBlocked from the rating pathNo access to candidate evidence
- Staffs them to clientsBlocked from the rating pathNo influence on any result
- Operates the platformPermitted, under auditRuns the service under audit
- Funds the programPermitted, under auditAppears in a public funder register, holds no vote
The commission that would hold every protected decision is specified but not seated. Until members exist, the correct description is proposed. Read the firewall or inspect the accountability record.
Everything above is checkable.
The construct
Every domain and family, with its released definition, proposed task, and evidence gap.
Open the explorerThe method
Construct, task design, evidence model, calculation boundary, limitations, and sources.
Read the method paperThe claims
What may be said today, what may not, and which gate changes each answer.
Open the status registerThe naming decision
Why there is no headline score, and what the system is called instead.
Read the decision record
