Identity and eligibility
Confirm the participant and authorization context without treating identity as competence.
Realistic work. Changing constraints. Inspectable artifacts. Independent corroboration. Outcomes with attribution and uncertainty intact.
Each layer answers a different question. None is allowed to substitute for the others.
Confirm the participant and authorization context without treating identity as competence.
Observe decisions, artifacts, revisions, and recovery—not recall of benchmark language.
Ask references about specific work, scope, contribution, and outcomes they directly observed.
Separate what changed from why it changed, who contributed, and what remained outside the practitioner’s control.
The first pilot should prove depth before breadth. These are task hypotheses, not a released test form.
Given interview excerpts from a sponsor, operations lead, frontline user, security reviewer, and an absent downstream team, propose a stakeholder and decision map for the next discovery step.
Midway through the task, disclose that the executive sponsor cannot approve workflow changes and that a contractor performs the highest-volume step.
This exercise would not establish effectiveness in live customer interviews or politically sensitive environments. A polished map would not establish that the candidate earned trust from the represented stakeholders.
Every stage produces an inspectable receipt. A missing receipt stops the next state.
Critical tasks, conditions, artifacts, and non-inference limits are frozen before participant work.
Cognitive interviews test whether the exercise elicits the intended capability without irrelevant barriers.
Reviewers train on shared anchors, score overlapping units, and resolve material drift.
A consented, preregistered cohort completes tasks under declared conditions and accommodations.
Independent outcomes are linked only after leakage, attrition, access, and attribution controls pass.
A qualified, conflict-disclosed panel may propose criterion anchors after evidence exists. No cut score is authorized.
The pilot preserves capability, evidence, uncertainty, and readiness as different outputs. A future standard-setting study may propose criterion anchors, but v0.1 cannot emit an official threshold.
Inspect the scoring boundaryCriterion-level observations from work; no résumé or title multiplier.
Coverage, independence, maturity, and provenance remain visible.
Missingness, disagreement, task conditions, and limitations travel with every interpretation.
A governed decision state, separate from proficiency and evidence maturity.
No recruitment, real evidence intake, result publication, certification, ranking, or employment decision is authorized.