Skip to publication
LockedIn Labs ResearchAI-native role context · method + construct draft 001

The benchmark for work that has to work.

A proposed evidence-led framework for AI-native Forward Deployed Engineering across ambiguous discovery, AI-system delivery, production operation, measurable outcomes, and field judgment.

Document typeMethod + construct draft
Unit of observationCapability claim + evidence context
Empirical human record0 admitted observations
External reviewNot recorded
Official decisionsNot authorized
01
Abstract

A title is cheap. Field evidence is not.

Forward Deployed Engineering is an operating model used beyond AI alone. This release deliberately narrows the role context to AI-native FDE: practitioners joining software, AI systems, enterprise constraints, production operations, and customer judgment in one field role. Current first-party role records illustrate both AI-native and non-AI-exclusive uses of the title; they inform the boundary but do not validate it. F-01F-02

This public draft defines a vendor-neutral AI-native role context, a 55-family capability ontology, an ordinal evidence model, a proposed sequence of applied observations, and the boundary for a future deterministic measurement system. It treats the candidate’s work—not prestige, fluency, or artifact volume—as the primary object of review. Conventional résumé screens, title matching, and short multiple-choice tests cannot by themselves establish that a person can do this work. V-01V-02V-03V-05V-06

The portable FDE core and AI-native specialization remain distinguishable. Later security, financial-infrastructure, telecommunications, defense, industrial, or other context packs require their own practice analysis, blueprint, validation evidence, and release decision; this draft does not assume transportability.

The intended result is a reviewable multidimensional record: what was observed, what supports it, how much of the role was covered, where critical gates stand, and what remains unknown. It is deliberately not a universal human score.

Research thesis

Capability claims become decision-grade only when the benchmark can connect an observed behavior to a job-relevant task, an attributable artifact, an operating context, and a reviewable outcome. V-02V-05L-06B

RQ1
What work distinguishes AI-native Forward Deployed Engineering from adjacent applied-AI, delivery, consulting, solutions, and platform roles?
RQ2
Which observations are sufficiently direct, attributable, independent, recent, and job-relevant to support a capability claim?
RQ3
Can evaluators apply anchored criteria reliably without rewarding prestige, verbosity, tool familiarity, or access to privileged production contexts?
02
Formation protocol · definition 0.1

A benchmark is a chain of warranted decisions.

The ontology is one layer. Authority requires a declared use, representative work, standardized evidence, measurement quality, and a governed release.

Formal definitionLL-FDE / DEF-001

The LockedIn FDE Benchmark is a versioned, purpose-bound assessment system for collecting, reviewing, and interpreting evidence of demonstrated Forward Deployed Engineering capability in a defined role context.

v0.1 scope: FDE is the broader operating model; AI-native deployment is the specialization this release examines first. Reuse in another forward-deployed context requires a separately versioned domain pack, practice analysis, content blueprint, validation plan, and release decision. F-01F-02V-01V-02

A complete release binds its population and use, construct, content blueprint, administration, evidence protocol, measurement and decision rules, quality evidence, and governance record. Missing layers remain visible; software completion cannot substitute for them. V-01V-02V-06M-12

01Capability framework

What the work contains.

Ontology, definitions, role boundaries, and observable capability hypotheses.

Available as public draft
02Assessment benchmark

How the work is observed.

Blueprint, tasks, conditions, evidence, rubrics, measurement, uncertainty, and validity record.

Under formation
03Certification scheme

Who may be formally attested.

Identity, eligibility, impartial decision, certificate lifecycle, renewal, complaints, and appeals.

Not operational
Formation stack

Seven layers must resolve to one version.

Current evidence state
  1. 01

    Population + use

    Who, which role context, what interpretation, and which decisions are in scope.

    Defined for public review
  2. 02

    Construct

    The work, capability boundaries, excluded attributes, and adjacent-role distinctions.

    Public draft
  3. 03

    Content blueprint

    Job activity → behavior → task → artifact → rubric criterion → coverage.

    Not yet established
  4. 04

    Administration

    Conditions, tools, assistance, accommodations, security, identity, and missingness.

    Protocol proposed
  5. 05

    Measurement

    Scoring, uncertainty, critical gates, decision rules, and interpretation limits.

    Experimental only
  6. 06

    Quality evidence

    Content, response process, reliability, fairness, outcome, and consequence studies.

    Not yet established
  7. 07

    Governance + release

    Conflicts, independent review, corrections, appeals, reproducibility, and retirement.

    Source rules available
Formation inputs

Experience enters as a named evidence class—not as invisible authority.

Every input has an admission record, a limited purpose, and a boundary on what it can establish.

Standards + guidanceDated ledger

Professional standards, regulator guidance, and official technical sources shape the validation questions and use boundaries. They do not make this draft conformant. V-01V-02M-12L-06B

Reviewed inputs
Peer-reviewed researchSelected literature

Work samples, structured review, rater quality, validity arguments, and qualitative reporting inform hypotheses. Published coefficients are never imported as this benchmark’s performance. V-03V-04V-05V-06V-07V-08

Not a systematic review
Practitioner + executive evidence0 admitted

The sponsor-reported 500/50 target has no admissible records in this repository. It contributes no finding until the intake, coding, quality, and aggregate receipts reproduce it. M-18M-19M-20

Corpus unsubstantiated
Steward experienceHypothesis source

Founder, product, and delivery experience may propose scenarios and failure modes only when named, dated, conflicted, and traceable. It is not independent empirical evidence.

Register not published
Benchmark operationsSelected precedents

Versioning, audit, reproducibility, status classes, and correction patterns draw from operating benchmarks and NIST measurement practice—not from model-to-human score transfer. M-01M-04M-23M-24M-25M-28M-29M-30M-09

Adapted, not copied
Synthetic engineering evidenceAvailable in source

Fixtures and adversarial tests can establish schema, determinism, and failure behavior. They cannot establish human reliability, validity, fairness, or job performance.

Software evidence only
Claim-to-evidence map

What shaped each design decision—and what must still be earned.

Open job-analysis protocol

FDE construct + 55 familiesV-01V-02V-09V-06

Steward synthesis, published frameworks, market observation, and a released ontology.

Job/practice analysis, critical incidents, work products, importance/criticality ratings, adjacent-role study, and external content review.

Proposed construct

Hands-on task sequenceV-03V-05M-09L-06B

Work-sample research, evidence-centered assessment design, job-relevance guidance, and open-ended benchmark precedents.

Content blueprint, task solvability, response-process evidence, equivalent forms, accessibility, accommodations, and representative pilots.

Protocol proposed

Structured evidence defenseV-02V-07V-08

Structured-interview and rater-method literature plus a non-conflicted human-review requirement.

Rater selection/training, anchored exemplars, double review, agreement/reliability, drift, decision consistency, and automation-bias studies.

Protocol proposed

References + identityV-10M-17

Structured corroboration, provenance, and risk-based identity-proofing guidance.

Incremental utility, burden, false-match, privacy, accessibility, fraud, contribution-attribution, and redress studies.

Separate future inputs

Multidimensional resultV-01V-02V-04M-23

Intended-use analysis, evidence lineage, explicit missingness, and an experimental deterministic scorer.

Calibration, measurement error, generalizability, floors, compensation, sensitivity, standard setting, and independent replication.

No official score

Certification or hiring useM-12L-06L-06BV-02

Personnel-assessment, employment-use, impartiality, accommodation, and certification requirements are treated as design constraints.

Role-specific validity evidence, operational impartiality, fairness/adverse-impact evidence, accommodations, appeals, security, and an authorized certification scheme.

Prohibited in v0.1
03
Construct map · ontology 0.1.0

The role is a system of work, not a bag of skills.

Five reader arcs organize the released ten-domain, 55-capability-family ontology. The arcs aid navigation only; they are not score weights or subscales.

A
6 capability families

Frame the consequential problem

Discover the operating reality, contract the outcome, and reject unsafe or uneconomic work before implementation begins.

01Discovery6
B
21 capability families

Engineer the intervention

Design and implement the software, AI, data, and architectural boundaries needed to change a live system.

02Systems703AI systems804Architecture6
C
14 capability families

Cross the enterprise boundary

Translate controls into system behavior, move through production safely, and retain the evidence needed to operate and recover.

05Enterprise706Production7
D
10 capability families

Land durable change

Demonstrate outcomes, lead contested decisions, and sequence modernization without confusing activity with impact.

07Outcomes408Leadership409Modernization2
P
4 capability families

Apply a domain pack

Preserve the vendor-neutral core while adding separately versioned constraints for a specific industry or operating environment.

10Domain packs4
Included

Problem framing, system implementation, AI behavior, architecture, enterprise controls, production operation, outcomes, leadership, modernization, and contextual domain packs.

Excluded as direct inputs

Employer or school prestige, compensation, title, fame, social following, training purchase, sponsorship, friendship, and tool-brand allegiance.

04
Evidence model

Proof strengthens by directness—not by volume.

The same deployment described in a résumé, slide, repository, and reference is not four independent confirmations. Provenance and lineage remain attached.

E2
Evidence maturity · ordinal, not arithmetic

The practitioner can perform it.

Capability observed in a controlled lab, live simulation, architecture defense, paired exercise, or supervised challenge.

Required proof form
Controlled demonstration trajectory
Current release boundary
May describe evidence; does not authorize an official score.

E0–E5 describes the maturity of support for one capability claim. The levels are not evenly spaced points and are never averaged into a universal human score.

01SubjectWho did the work?
02ObservationWhat happened?
03ArtifactWhat remains?
04CorroborationWho independently saw it?
05OutcomeWhat changed?
05
Applied task design · hypothesis

Observe the engineer across the change curve.

The proposed assessment hypothesis combines ambiguity, implementation, changing constraints, an operating event, and an evidence defense. Job analysis and response-process research must still establish representativeness.

01
Applied assessment event

The problem arrives incomplete on purpose.

The candidate receives a consequential outcome, partial system context, conflicting stakeholder needs, and at least one hard operating constraint. The first observation is whether they clarify the right unknowns—not whether they begin coding quickly.

Problem decompositionConstraint discoveryAcceptance criteriaRefusal and escalation judgment
Retained evidenceQuestion log · scope contract · initial risk register

Job relevance first

Every scenario criterion and probe must trace to an approved content blueprint for a defined role context.

Equivalent paths

Accommodation and alternate demonstration are part of measurement quality, not exceptions added after launch.

Failures are findings

Broken tasks, contamination, burden, scorer disagreement, and protocol deviations belong in the published record.

06
Conceptual calculation boundary · unvalidated

Deterministic where software helps. Human where judgment matters.

The calculation boundary may organize eligible human-reviewed findings. It may not manufacture evidence, hide missingness, or make the final credentialing decision.

01
Version control

Bind the instrument

Resolve immutable benchmark, ontology, rubric, evaluator, assistance, recency, and experimental-policy identifiers before touching a result.

EmitsVersion manifest
Published recordCapability profile+Support confidence+Coverage+Critical gates+Limits

No official score exists. Numeric weights, thresholds, score bands, percentiles, pass rates, and a comparison population remain unvalidated. Software completion cannot activate them. V-01V-02M-23

07
Study register

The construct is drafted. Its empirical claims still have to be earned.

Protocol, recruitment, data, analysis, limitations, and release decisions are separate artifacts. Product maturity cannot substitute for study evidence.

Pending substantiationExcluded from benchmark findings
Commissioned interview corpus · sponsor-reported target

The sponsor reports a target of at least 500 AI practitioners and at least 50 executives. Zero records have been admitted.

The repository contains no recruitment frame, fieldwork dates, participant ledger, consent record, interview instrument, recording or transcript sample, vendor completion receipt, deduplication record, coding codebook, or quality-control report. Until those materials are received and checked, this target is not a completed sample and contributes zero admissible observations to the benchmark.

Admission requirements
  • Population definition, sampling and recruitment path
  • Participant-level pseudonymous ledger and deduplication
  • Consent, dates, interviewer assignment and fieldwork receipt
  • Instrument version, recordings/transcripts and QA sample
  • Coding codebook, analyst agreement and exclusion log
  • Conflicts, payment, missingness and limitations disclosure
StudyQuestion and designState
S01

Role and critical-incident study

Which observable decisions, artifacts, contexts, and consequences define AI-native Forward Deployed Engineering across declared settings?

Critical-incident interviews, work-product inventory, practitioner/manager panels, and importance–frequency–consequence ratings.
Protocol proposed
S02

Content and response-process study

Do tasks elicit the intended capability rather than test-taking strategy, tool familiarity, presentation polish, or inaccessible interaction patterns?

Blueprint review, cognitive walkthroughs, think-aloud analysis, accommodations, and scenario failure review.
Not started
S03

Reliability and evaluator study

Can trained evaluators reach sufficiently consistent, explainable judgments across domains, assistance modes, and repeated cases?

Double scoring, adjudication, generalizability analysis, drift checks, disagreement severity, and automation-bias probes.
Not started
S04

Field-outcome validation

For a defined role and intended use, what does a released profile predict beyond less costly, less burdensome alternatives?

Pre-registered, longitudinal, role-specific criterion study with missingness, subgroup, burden, and adverse-consequence analysis.
Blocked on earlier studies
Activation decision

What must be true before an official result can exist

01Defined construct, population, roles and intended uses02Externally reviewed ontology, blueprint and rubrics03At least two designed calibration cohorts04Measured rater reliability and adjudication quality05Calibrated uncertainty, coverage, floors and interpretation06Fairness, accessibility, accommodation and burden evaluation07Operational correction, appeal, privacy and security controls08Independent measurement review and governed release decision
08
Validity argument + evidence register

Credibility is a set of questions with receipts.

The draft names the interpretations it hopes to support and the evidence each requires. None of the four inference steps is authorized today.

InferenceProposed claimRequired warrantState

Scoring

Reviewers can translate observed work into defensible criterion findings.

Anchors, rater training, response-process evidence, agreement/reliability, and decision consistency.

Not established

Generalization

The sampled tasks represent the declared FDE role context.

Job analysis, content blueprint, task/form sampling, and generalizability evidence.

Not established

Extrapolation

Performance under benchmark conditions relates to independently measured field work.

Criterion definition, temporal separation, attribution controls, longitudinal evidence, and replication.

Not established

Decision

A profile improves a defined decision beyond less costly and less burdensome alternatives.

Incremental validity, standard setting, fairness, consequence, burden, and actual-use evidence.

Not authorized
VE-01Not established

Content

Does the benchmark represent important FDE work for the declared role context?

Minimum evidenceJob analysis · blueprint · panel method · ratings · dissent · coverage
VE-02Not established

Response processes

Do tasks elicit the intended capability rather than irrelevant strategy or interface friction?

Minimum evidenceCognitive walkthroughs · think-aloud evidence · accommodations · deviations
VE-03Not established

Internal structure + reliability

Are observations and decisions sufficiently consistent for the proposed interpretation?

Minimum evidenceRater agreement/reliability · task effects · drift · decision consistency · uncertainty
VE-04Not established

Relations to outcomes

Does the result relate to later, independently measured work outcomes as proposed?

Minimum evidencePre-registered criteria · temporal separation · attribution · cross-validation · replication
VE-05Not established

Fairness + accessibility

Are construct-irrelevant barriers identified and are alternate demonstrations comparable?

Minimum evidenceAccessibility · accommodations · subgroup/error analysis · adverse impact · remediation
VE-06Not established

Consequences + actual use

What happens when the benchmark is used, misused, appealed, or wrong?

Minimum evidenceBurden · false decisions · privacy · gaming · appeals · unintended-consequence monitoring
VE-07Synthetic foundation only

Reproducibility + security

Can a qualified reviewer replay a result without exposing controlled material?

Minimum evidenceFrozen inputs · code · environment · logs · hashes · contamination controls · replay

The evidence strands adapt the validity, reliability, fairness, intended-use, and documentation questions in the testing and personnel-selection standards. They are a research program, not a self-awarded quality rating. V-01V-02L-06B

09
Limitations

Read the boundary before the benchmark.

These limitations are release facts, not footnotes. A future study may narrow them; an attractive interface cannot.

01

No field study

No real-person job-analysis, reliability, fairness, criterion-validity, or longitudinal-outcomes dataset exists in the repository.

02

No calibrated scores

Weights, thresholds, bands, percentiles, readiness labels, and pass rates remain hypotheses. Evidence maturity is ordinal.

03

AI-native scope only

FDE is a broader operating model. v0.1 defines an AI-native role context and cannot be assumed to transfer to security, payments, telecom, defense, industrial, or other FDE contexts.

04

Access is not ability

Production access, public repositories, recognizable employers, fluent English, and expensive tools are unevenly distributed and cannot stand in for capability.

05

Verification can burden people

Identity proofing, references, artifact collection, monitoring, and long tasks create privacy, exclusion, coercion, and accessibility risks.

06

Employment use is high consequence

This draft is not authorized for automated or sole-source hiring, rejection, ranking, compensation, promotion, discipline, or termination.

Appropriate now

Public review, ontology critique, scenario prototyping, schema and workflow testing, and research-protocol development.

Not appropriate now

Claiming certification, issuing scores or pass/fail decisions, ranking people, or treating a prototype profile as validated hiring evidence.

10
References and source ledger

Primary sources, visible influence, explicit non-adoption.

This is a targeted design review—not a systematic review. A citation means reviewed and relevant; it does not mean implemented, conformant, endorsed, or validated.

F-01
Primary employer role record

OpenAI — Forward Deployed Engineer (FDE), San Francisco

Documents one current AI-native FDE instantiation spanning customer discovery, system design, hands-on build, frontier-model production rollout, measurable workflow impact, and field feedback to product and research. A job posting is market evidence, not a universal construct or validation study.

F-02
Primary employer role record

Stripe — Forward Deployed Engineer, Professional Services

Documents use of the FDE operating model for payment integrations and production software in customer environments, with AI familiarity listed as preferred rather than defining the role. It supports a non-AI-exclusive boundary but does not establish prevalence or transport validity.

V-01
Professional standard · AERA · APA · NCME

Standards for Educational and Psychological Testing (2014)

Validity, reliability, fairness, intended interpretation, documentation, and test-use responsibilities frame the measurement program. No conformance claim.

V-02
Professional practice principles

SIOP Principles for the Validation and Use of Personnel Selection Procedures, 5th edition (2018)

Job analysis, evidence for intended use, criterion quality, transportability, fairness, and documentation inform the proposed study sequence. No validation has been completed.

V-03
Peer-reviewed meta-analysis · DOI 10.1111/j.1744-6570.2005.00714.x

Roth, Bobko & McFarland (2005) — A meta-analysis of work sample test validity

Work-sample evidence is relevant but study design and estimates vary. The paper motivates study-specific criterion validation; no published coefficient is imported into this benchmark.

V-04
Peer-reviewed meta-analysis · DOI 10.1037/apl0000994

Sackett et al. (2022) — Revisiting meta-analytic estimates of validity in personnel selection

The critique of range-restriction corrections reinforces that borrowed validity estimates cannot establish this instrument’s performance. Every coefficient must come from its declared study and population.

V-05
Assessment-design foundation · DOI 10.1002/j.2333-8504.2003.tb01908.x

Mislevy, Almond & Lukas (2003) — A brief introduction to evidence-centered design

The explicit claim–evidence–task relationship informs the proposed content blueprint. It does not validate this occupational construct or require a specific psychometric model.

V-06
Argument-based validity framework · DOI 10.1111/jedm.12000

Kane (2013) — Validating the interpretations and uses of test scores

The benchmark must state its proposed interpretation-and-use argument, assumptions, warrants, rebuttals, and evidence gaps. A written argument is not itself validation.

V-07
Peer-reviewed review · DOI 10.1111/j.1744-6570.1997.tb00709.x

Campion, Palmer & Campion (1997) — A review of structure in the selection interview

Job-related fixed questions, standardized probing and administration, anchored scales, and evaluator structure inform the proposed evidence defense.

V-08
Peer-reviewed methods article · DOI 10.1177/1094428106296642

LeBreton & Senter (2008) — Interrater reliability and interrater agreement

Agreement and reliability answer different questions; the statistic and interpretation must match the scale, decision, and aggregation rule. No threshold is imported.

V-09
U.S. government assessment-practice guidance

U.S. Office of Personnel Management — Job Analysis

Task and competency analysis plus explicit linkage to assessment content inform the proposed practice-analysis and blueprint artifacts. No federal approval is implied.

V-10
U.S. government assessment-practice guidance

U.S. Office of Personnel Management — Reference Checking

Common job-related questions and standardized ratings support references as corroboration in a multi-method process—not proof of identity or competence.

M-18
Professional research-disclosure program

AAPOR Disclosure Standards and Transparency Initiative

Sponsor, conductor, population, recruitment, mode, dates, instrument, processing, quality controls, incentives, and limitations shape the interview-corpus gate. No membership claim.

M-19
Peer-reviewed 32-item reporting checklist · DOI 10.1093/intqhc/mzm042

COREQ — Consolidated criteria for reporting qualitative research

Research-team, study-design, context, analysis, and reporting disclosures inform admission of an interview corpus. Reporting completeness is not a quality finding.

M-20
Peer-reviewed reporting guidance · DOI 10.1097/ACM.0000000000000388

SRQR — Standards for Reporting Qualitative Research

Purpose, sampling, ethics, collection, analysis, trustworthiness, limitations, funding, and conflicts inform the proposed qualitative publication package.

M-01
First-party operating methodology

Artificial Analysis benchmarking methodology

Versioned methods, chart-level provenance, multidimensional reporting, and separation of methods from analysis inspired the publication architecture—not the human construct.

M-04
U.S. government voluntary framework

NIST Artificial Intelligence Risk Management Framework 1.0

Govern–Map–Measure–Manage, contextual evaluation, monitoring, human oversight, and contestability inform the research-control questions. No conformance claim.

M-23
U.S. government measurement-science report · DOI 10.6028/NIST.AI.800-3

NIST AI 800-3 — Expanding the AI Evaluation Toolbox with Statistical Models

Explicit estimands, modeling assumptions, dependence, and uncertainty inform benchmark reporting. Its AI-system models and results do not transfer to people.

M-24
U.S. government initial public draft · DOI 10.6028/NIST.AI.800-2.ipd

NIST AI 800-2 — Practices for Automated Benchmark Evaluations

Measurement-target definition, test conditions, contamination, reproducibility, uncertainty, and reporting inform operations. This is emerging AI guidance, not a final human-assessment standard.

M-25
Industry-consortium operating rules

MLCommons submission rules and independent-audit policies

Version-specific rules, structured submissions, environment metadata, compliance checks, reproducibility, audit, corrections, and result-status separation inform release operations.

M-28
Academic benchmark framework and public result record

Stanford CRFM HELM — versioned capabilities evaluation

Declarative configurations, prompt-level auditability, versioned leaderboards, reproduction commands, and explicit coverage limits inform release packaging—not human capability claims.

M-29
Academic benchmark, harness, dataset, and public artifact record

SWE-bench — evaluation and submission artifacts

Pinned revisions, containerized execution, per-instance logs, pass@1 disclosure, model/scaffold separation, re-grading, and verification reruns inform applied-task operations. Static-task contamination remains a warning.

M-30
U.S. government measurement guidance

NIST Technical Note 1297 — measurement uncertainty

Explicit uncertainty components, evaluation methods, coverage basis, and reporting conditions inform the future result-card contract. The guidance does not validate a human assessment model.

M-09
Primary benchmark implementation and paper

METR RE-Bench repository and paper

Long, open-ended environments, protected material, attempt-level conditions, and reproducible release artifacts inform task-design questions. It does not validate FDE measurement.

M-12
International personnel-certification standard

ISO/IEC 17024:2026

Impartiality, separation from training, consistent decisions, records, confidentiality, and human oversight are design inputs. This draft is not a certification scheme or accredited body.

M-17
U.S. government technical guidance

NIST SP 800-63A-4 — Identity Proofing and Enrollment

Risk assessment, data minimization, comparable proofing paths, fraud management, notice, and redress inform identity-gate design. No federal assurance claim.

L-06
U.S. federal enforcement guidance

EEOC — Employment Tests and Selection Procedures

Job relevance, intended-use validation, accommodation, and employer/vendor responsibility constrain any future employment use.

L-06B
U.S. federal guidance

EEOC — Uniform Guidelines questions and answers

Criterion definition, validation strategy, recordkeeping, and intended-use specificity inform the proposed validation program.

11
Scholarly and publication disclosures

The record includes who is accountable—and what is missing.

No review, affiliation, author, funding, ethics, license, or archival signal is implied by the visual design. Missing records remain publication blockers.

01

Organizational author

LockedIn Labs Research

02

Individual authors + CRediT roles

Not yet published; required before an authored external release

03

Funding

Final public funding statement not recorded; no outside sponsor is represented as endorsing this draft

04

Conflict of interest

LockedIn Labs is founder/operator and may also train, employ, or serve FDE practitioners and customers

05

Human-participant research

No admitted human dataset; no IRB or equivalent approval or exemption is claimed

06

Data + materials

Public specifications, schemas, code, and synthetic fixtures; no admitted real-person evidence

07

Review status

Not externally peer reviewed; no independent validation, accreditation, or independent reproduction recorded

08

AI assistance

AI-assisted tools supported source discovery, drafting, software development, and testing; they are not evidence or decision authorities

09

Persistent publication record

Canonical URL, DOI, archival deposit, and operative licenses remain publication blockers

Publication record

A citation that tells the truth about the release.

Use the title, version, status, and access date until a governed release assigns accountable authors, a canonical URL, archive, and persistent identifier.

Recommended citationAPA-style
LockedIn Labs Research. (2026). LockedIn FDE Benchmark v0.1 — Public method and construct draft (Version 0.1.0-draft). Source preview; not a validated assessment, completed study, certification, or authorized hiring instrument.
Paper ID
LL-FDE-METHOD-0.1
Version
0.1.0-draft
DOI
Not assigned
Release date
Not assigned
Canonical URL
Pending publication decision
Source-review cutoff
1 September 2026
Peer review
Not externally peer reviewed
License
Pending publication decision
Corrections
Append-only release history proposed

Citation identifies source and status only. This draft is not a validated assessment, completed study, certification, or authorized hiring instrument.

Run the instrument

Inspect the evidence workflow behind the paper.

Walk through the proposed brief, identity, applied work, field evidence, outcomes, references, and human-review stages.

Open assessment workspace
Inspect the output

See a verification record with its limits intact.

Review a synthetic employer-facing record that preserves capability shape, source lineage, lifecycle, uncertainty, and recourse.

Open verification record