A title is cheap. Field evidence is not.
Forward Deployed Engineering is an operating model used beyond AI alone. This release deliberately narrows the role context to AI-native FDE: practitioners joining software, AI systems, enterprise constraints, production operations, and customer judgment in one field role. Current first-party role records illustrate both AI-native and non-AI-exclusive uses of the title; they inform the boundary but do not validate it. F-01F-02
This public draft defines a vendor-neutral AI-native role context, a 55-family capability ontology, an ordinal evidence model, a proposed sequence of applied observations, and the boundary for a future deterministic measurement system. It treats the candidate’s work—not prestige, fluency, or artifact volume—as the primary object of review. Conventional résumé screens, title matching, and short multiple-choice tests cannot by themselves establish that a person can do this work. V-01V-02V-03V-05V-06
The portable FDE core and AI-native specialization remain distinguishable. Later security, financial-infrastructure, telecommunications, defense, industrial, or other context packs require their own practice analysis, blueprint, validation evidence, and release decision; this draft does not assume transportability.
The intended result is a reviewable multidimensional record: what was observed, what supports it, how much of the role was covered, where critical gates stand, and what remains unknown. It is deliberately not a universal human score.
Capability claims become decision-grade only when the benchmark can connect an observed behavior to a job-relevant task, an attributable artifact, an operating context, and a reviewable outcome. V-02V-05L-06B
- RQ1
- What work distinguishes AI-native Forward Deployed Engineering from adjacent applied-AI, delivery, consulting, solutions, and platform roles?
- RQ2
- Which observations are sufficiently direct, attributable, independent, recent, and job-relevant to support a capability claim?
- RQ3
- Can evaluators apply anchored criteria reliably without rewarding prestige, verbosity, tool familiarity, or access to privileged production contexts?
A benchmark is a chain of warranted decisions.
The ontology is one layer. Authority requires a declared use, representative work, standardized evidence, measurement quality, and a governed release.
The LockedIn FDE Benchmark is a versioned, purpose-bound assessment system for collecting, reviewing, and interpreting evidence of demonstrated Forward Deployed Engineering capability in a defined role context.
v0.1 scope: FDE is the broader operating model; AI-native deployment is the specialization this release examines first. Reuse in another forward-deployed context requires a separately versioned domain pack, practice analysis, content blueprint, validation plan, and release decision. F-01F-02V-01V-02
A complete release binds its population and use, construct, content blueprint, administration, evidence protocol, measurement and decision rules, quality evidence, and governance record. Missing layers remain visible; software completion cannot substitute for them. V-01V-02V-06M-12
What the work contains.
Ontology, definitions, role boundaries, and observable capability hypotheses.
Available as public draftHow the work is observed.
Blueprint, tasks, conditions, evidence, rubrics, measurement, uncertainty, and validity record.
Under formationWho may be formally attested.
Identity, eligibility, impartial decision, certificate lifecycle, renewal, complaints, and appeals.
Not operationalSeven layers must resolve to one version.
- 01Defined for public review
Population + use
Who, which role context, what interpretation, and which decisions are in scope.
- 02Public draft
Construct
The work, capability boundaries, excluded attributes, and adjacent-role distinctions.
- 03Not yet established
Content blueprint
Job activity → behavior → task → artifact → rubric criterion → coverage.
- 04Protocol proposed
Administration
Conditions, tools, assistance, accommodations, security, identity, and missingness.
- 05Experimental only
Measurement
Scoring, uncertainty, critical gates, decision rules, and interpretation limits.
- 06Not yet established
Quality evidence
Content, response process, reliability, fairness, outcome, and consequence studies.
- 07Source rules available
Governance + release
Conflicts, independent review, corrections, appeals, reproducibility, and retirement.
Experience enters as a named evidence class—not as invisible authority.
Every input has an admission record, a limited purpose, and a boundary on what it can establish.
Professional standards, regulator guidance, and official technical sources shape the validation questions and use boundaries. They do not make this draft conformant. V-01V-02M-12L-06B
Work samples, structured review, rater quality, validity arguments, and qualitative reporting inform hypotheses. Published coefficients are never imported as this benchmark’s performance. V-03V-04V-05V-06V-07V-08
The sponsor-reported 500/50 target has no admissible records in this repository. It contributes no finding until the intake, coding, quality, and aggregate receipts reproduce it. M-18M-19M-20
Founder, product, and delivery experience may propose scenarios and failure modes only when named, dated, conflicted, and traceable. It is not independent empirical evidence.
Versioning, audit, reproducibility, status classes, and correction patterns draw from operating benchmarks and NIST measurement practice—not from model-to-human score transfer. M-01M-04M-23M-24M-25M-28M-29M-30M-09
Fixtures and adversarial tests can establish schema, determinism, and failure behavior. They cannot establish human reliability, validity, fairness, or job performance.
What shaped each design decision—and what must still be earned.
FDE construct + 55 familiesV-01V-02V-09V-06
Steward synthesis, published frameworks, market observation, and a released ontology.
Job/practice analysis, critical incidents, work products, importance/criticality ratings, adjacent-role study, and external content review.
Proposed constructHands-on task sequenceV-03V-05M-09L-06B
Work-sample research, evidence-centered assessment design, job-relevance guidance, and open-ended benchmark precedents.
Content blueprint, task solvability, response-process evidence, equivalent forms, accessibility, accommodations, and representative pilots.
Protocol proposedStructured evidence defenseV-02V-07V-08
Structured-interview and rater-method literature plus a non-conflicted human-review requirement.
Rater selection/training, anchored exemplars, double review, agreement/reliability, drift, decision consistency, and automation-bias studies.
Protocol proposedReferences + identityV-10M-17
Structured corroboration, provenance, and risk-based identity-proofing guidance.
Incremental utility, burden, false-match, privacy, accessibility, fraud, contribution-attribution, and redress studies.
Separate future inputsMultidimensional resultV-01V-02V-04M-23
Intended-use analysis, evidence lineage, explicit missingness, and an experimental deterministic scorer.
Calibration, measurement error, generalizability, floors, compensation, sensitivity, standard setting, and independent replication.
No official scoreCertification or hiring useM-12L-06L-06BV-02
Personnel-assessment, employment-use, impartiality, accommodation, and certification requirements are treated as design constraints.
Role-specific validity evidence, operational impartiality, fairness/adverse-impact evidence, accommodations, appeals, security, and an authorized certification scheme.
Prohibited in v0.1The role is a system of work, not a bag of skills.
Five reader arcs organize the released ten-domain, 55-capability-family ontology. The arcs aid navigation only; they are not score weights or subscales.
Frame the consequential problem
Discover the operating reality, contract the outcome, and reject unsafe or uneconomic work before implementation begins.
Engineer the intervention
Design and implement the software, AI, data, and architectural boundaries needed to change a live system.
Cross the enterprise boundary
Translate controls into system behavior, move through production safely, and retain the evidence needed to operate and recover.
Land durable change
Demonstrate outcomes, lead contested decisions, and sequence modernization without confusing activity with impact.
Apply a domain pack
Preserve the vendor-neutral core while adding separately versioned constraints for a specific industry or operating environment.
Problem framing, system implementation, AI behavior, architecture, enterprise controls, production operation, outcomes, leadership, modernization, and contextual domain packs.
Employer or school prestige, compensation, title, fame, social following, training purchase, sponsorship, friendship, and tool-brand allegiance.
Proof strengthens by directness—not by volume.
The same deployment described in a résumé, slide, repository, and reference is not four independent confirmations. Provenance and lineage remain attached.
The practitioner can perform it.
Capability observed in a controlled lab, live simulation, architecture defense, paired exercise, or supervised challenge.
- Required proof form
- Controlled demonstration trajectory
- Current release boundary
- May describe evidence; does not authorize an official score.
E0–E5 describes the maturity of support for one capability claim. The levels are not evenly spaced points and are never averaged into a universal human score.
Observe the engineer across the change curve.
The proposed assessment hypothesis combines ambiguity, implementation, changing constraints, an operating event, and an evidence defense. Job analysis and response-process research must still establish representativeness.
The problem arrives incomplete on purpose.
The candidate receives a consequential outcome, partial system context, conflicting stakeholder needs, and at least one hard operating constraint. The first observation is whether they clarify the right unknowns—not whether they begin coding quickly.
Job relevance first
Every scenario criterion and probe must trace to an approved content blueprint for a defined role context.
Equivalent paths
Accommodation and alternate demonstration are part of measurement quality, not exceptions added after launch.
Failures are findings
Broken tasks, contamination, burden, scorer disagreement, and protocol deviations belong in the published record.
Deterministic where software helps. Human where judgment matters.
The calculation boundary may organize eligible human-reviewed findings. It may not manufacture evidence, hide missingness, or make the final credentialing decision.
Bind the instrument
Resolve immutable benchmark, ontology, rubric, evaluator, assistance, recency, and experimental-policy identifiers before touching a result.
The construct is drafted. Its empirical claims still have to be earned.
Protocol, recruitment, data, analysis, limitations, and release decisions are separate artifacts. Product maturity cannot substitute for study evidence.
The sponsor reports a target of at least 500 AI practitioners and at least 50 executives. Zero records have been admitted.
The repository contains no recruitment frame, fieldwork dates, participant ledger, consent record, interview instrument, recording or transcript sample, vendor completion receipt, deduplication record, coding codebook, or quality-control report. Until those materials are received and checked, this target is not a completed sample and contributes zero admissible observations to the benchmark.
- Population definition, sampling and recruitment path
- Participant-level pseudonymous ledger and deduplication
- Consent, dates, interviewer assignment and fieldwork receipt
- Instrument version, recordings/transcripts and QA sample
- Coding codebook, analyst agreement and exclusion log
- Conflicts, payment, missingness and limitations disclosure
Role and critical-incident study
Which observable decisions, artifacts, contexts, and consequences define AI-native Forward Deployed Engineering across declared settings?
Critical-incident interviews, work-product inventory, practitioner/manager panels, and importance–frequency–consequence ratings.Content and response-process study
Do tasks elicit the intended capability rather than test-taking strategy, tool familiarity, presentation polish, or inaccessible interaction patterns?
Blueprint review, cognitive walkthroughs, think-aloud analysis, accommodations, and scenario failure review.Reliability and evaluator study
Can trained evaluators reach sufficiently consistent, explainable judgments across domains, assistance modes, and repeated cases?
Double scoring, adjudication, generalizability analysis, drift checks, disagreement severity, and automation-bias probes.Field-outcome validation
For a defined role and intended use, what does a released profile predict beyond less costly, less burdensome alternatives?
Pre-registered, longitudinal, role-specific criterion study with missingness, subgroup, burden, and adverse-consequence analysis.What must be true before an official result can exist
Credibility is a set of questions with receipts.
The draft names the interpretations it hopes to support and the evidence each requires. None of the four inference steps is authorized today.
Scoring
Reviewers can translate observed work into defensible criterion findings.
Anchors, rater training, response-process evidence, agreement/reliability, and decision consistency.
Not establishedGeneralization
The sampled tasks represent the declared FDE role context.
Job analysis, content blueprint, task/form sampling, and generalizability evidence.
Not establishedExtrapolation
Performance under benchmark conditions relates to independently measured field work.
Criterion definition, temporal separation, attribution controls, longitudinal evidence, and replication.
Not establishedDecision
A profile improves a defined decision beyond less costly and less burdensome alternatives.
Incremental validity, standard setting, fairness, consequence, burden, and actual-use evidence.
Not authorizedContent
Does the benchmark represent important FDE work for the declared role context?
Response processes
Do tasks elicit the intended capability rather than irrelevant strategy or interface friction?
Internal structure + reliability
Are observations and decisions sufficiently consistent for the proposed interpretation?
Relations to outcomes
Does the result relate to later, independently measured work outcomes as proposed?
Fairness + accessibility
Are construct-irrelevant barriers identified and are alternate demonstrations comparable?
Consequences + actual use
What happens when the benchmark is used, misused, appealed, or wrong?
Reproducibility + security
Can a qualified reviewer replay a result without exposing controlled material?
The evidence strands adapt the validity, reliability, fairness, intended-use, and documentation questions in the testing and personnel-selection standards. They are a research program, not a self-awarded quality rating. V-01V-02L-06B
Read the boundary before the benchmark.
These limitations are release facts, not footnotes. A future study may narrow them; an attractive interface cannot.
No field study
No real-person job-analysis, reliability, fairness, criterion-validity, or longitudinal-outcomes dataset exists in the repository.
No calibrated scores
Weights, thresholds, bands, percentiles, readiness labels, and pass rates remain hypotheses. Evidence maturity is ordinal.
AI-native scope only
FDE is a broader operating model. v0.1 defines an AI-native role context and cannot be assumed to transfer to security, payments, telecom, defense, industrial, or other FDE contexts.
Access is not ability
Production access, public repositories, recognizable employers, fluent English, and expensive tools are unevenly distributed and cannot stand in for capability.
Verification can burden people
Identity proofing, references, artifact collection, monitoring, and long tasks create privacy, exclusion, coercion, and accessibility risks.
Employment use is high consequence
This draft is not authorized for automated or sole-source hiring, rejection, ranking, compensation, promotion, discipline, or termination.
Public review, ontology critique, scenario prototyping, schema and workflow testing, and research-protocol development.
Claiming certification, issuing scores or pass/fail decisions, ranking people, or treating a prototype profile as validated hiring evidence.
Primary sources, visible influence, explicit non-adoption.
This is a targeted design review—not a systematic review. A citation means reviewed and relevant; it does not mean implemented, conformant, endorsed, or validated.
OpenAI — Forward Deployed Engineer (FDE), San Francisco
Documents one current AI-native FDE instantiation spanning customer discovery, system design, hands-on build, frontier-model production rollout, measurable workflow impact, and field feedback to product and research. A job posting is market evidence, not a universal construct or validation study.
Stripe — Forward Deployed Engineer, Professional Services
Documents use of the FDE operating model for payment integrations and production software in customer environments, with AI familiarity listed as preferred rather than defining the role. It supports a non-AI-exclusive boundary but does not establish prevalence or transport validity.
Standards for Educational and Psychological Testing (2014)
Validity, reliability, fairness, intended interpretation, documentation, and test-use responsibilities frame the measurement program. No conformance claim.
SIOP Principles for the Validation and Use of Personnel Selection Procedures, 5th edition (2018)
Job analysis, evidence for intended use, criterion quality, transportability, fairness, and documentation inform the proposed study sequence. No validation has been completed.
Roth, Bobko & McFarland (2005) — A meta-analysis of work sample test validity
Work-sample evidence is relevant but study design and estimates vary. The paper motivates study-specific criterion validation; no published coefficient is imported into this benchmark.
Sackett et al. (2022) — Revisiting meta-analytic estimates of validity in personnel selection
The critique of range-restriction corrections reinforces that borrowed validity estimates cannot establish this instrument’s performance. Every coefficient must come from its declared study and population.
Mislevy, Almond & Lukas (2003) — A brief introduction to evidence-centered design
The explicit claim–evidence–task relationship informs the proposed content blueprint. It does not validate this occupational construct or require a specific psychometric model.
Kane (2013) — Validating the interpretations and uses of test scores
The benchmark must state its proposed interpretation-and-use argument, assumptions, warrants, rebuttals, and evidence gaps. A written argument is not itself validation.
Campion, Palmer & Campion (1997) — A review of structure in the selection interview
Job-related fixed questions, standardized probing and administration, anchored scales, and evaluator structure inform the proposed evidence defense.
LeBreton & Senter (2008) — Interrater reliability and interrater agreement
Agreement and reliability answer different questions; the statistic and interpretation must match the scale, decision, and aggregation rule. No threshold is imported.
U.S. Office of Personnel Management — Job Analysis
Task and competency analysis plus explicit linkage to assessment content inform the proposed practice-analysis and blueprint artifacts. No federal approval is implied.
U.S. Office of Personnel Management — Reference Checking
Common job-related questions and standardized ratings support references as corroboration in a multi-method process—not proof of identity or competence.
AAPOR Disclosure Standards and Transparency Initiative
Sponsor, conductor, population, recruitment, mode, dates, instrument, processing, quality controls, incentives, and limitations shape the interview-corpus gate. No membership claim.
COREQ — Consolidated criteria for reporting qualitative research
Research-team, study-design, context, analysis, and reporting disclosures inform admission of an interview corpus. Reporting completeness is not a quality finding.
SRQR — Standards for Reporting Qualitative Research
Purpose, sampling, ethics, collection, analysis, trustworthiness, limitations, funding, and conflicts inform the proposed qualitative publication package.
Artificial Analysis benchmarking methodology
Versioned methods, chart-level provenance, multidimensional reporting, and separation of methods from analysis inspired the publication architecture—not the human construct.
NIST Artificial Intelligence Risk Management Framework 1.0
Govern–Map–Measure–Manage, contextual evaluation, monitoring, human oversight, and contestability inform the research-control questions. No conformance claim.
NIST AI 800-3 — Expanding the AI Evaluation Toolbox with Statistical Models
Explicit estimands, modeling assumptions, dependence, and uncertainty inform benchmark reporting. Its AI-system models and results do not transfer to people.
NIST AI 800-2 — Practices for Automated Benchmark Evaluations
Measurement-target definition, test conditions, contamination, reproducibility, uncertainty, and reporting inform operations. This is emerging AI guidance, not a final human-assessment standard.
MLCommons submission rules and independent-audit policies
Version-specific rules, structured submissions, environment metadata, compliance checks, reproducibility, audit, corrections, and result-status separation inform release operations.
Stanford CRFM HELM — versioned capabilities evaluation
Declarative configurations, prompt-level auditability, versioned leaderboards, reproduction commands, and explicit coverage limits inform release packaging—not human capability claims.
SWE-bench — evaluation and submission artifacts
Pinned revisions, containerized execution, per-instance logs, pass@1 disclosure, model/scaffold separation, re-grading, and verification reruns inform applied-task operations. Static-task contamination remains a warning.
NIST Technical Note 1297 — measurement uncertainty
Explicit uncertainty components, evaluation methods, coverage basis, and reporting conditions inform the future result-card contract. The guidance does not validate a human assessment model.
METR RE-Bench repository and paper
Long, open-ended environments, protected material, attempt-level conditions, and reproducible release artifacts inform task-design questions. It does not validate FDE measurement.
ISO/IEC 17024:2026
Impartiality, separation from training, consistent decisions, records, confidentiality, and human oversight are design inputs. This draft is not a certification scheme or accredited body.
NIST SP 800-63A-4 — Identity Proofing and Enrollment
Risk assessment, data minimization, comparable proofing paths, fraud management, notice, and redress inform identity-gate design. No federal assurance claim.
EEOC — Employment Tests and Selection Procedures
Job relevance, intended-use validation, accommodation, and employer/vendor responsibility constrain any future employment use.
EEOC — Uniform Guidelines questions and answers
Criterion definition, validation strategy, recordkeeping, and intended-use specificity inform the proposed validation program.
The record includes who is accountable—and what is missing.
No review, affiliation, author, funding, ethics, license, or archival signal is implied by the visual design. Missing records remain publication blockers.
Organizational author
LockedIn Labs Research
Individual authors + CRediT roles
Not yet published; required before an authored external release
Funding
Final public funding statement not recorded; no outside sponsor is represented as endorsing this draft
Conflict of interest
LockedIn Labs is founder/operator and may also train, employ, or serve FDE practitioners and customers
Human-participant research
No admitted human dataset; no IRB or equivalent approval or exemption is claimed
Data + materials
Public specifications, schemas, code, and synthetic fixtures; no admitted real-person evidence
Review status
Not externally peer reviewed; no independent validation, accreditation, or independent reproduction recorded
AI assistance
AI-assisted tools supported source discovery, drafting, software development, and testing; they are not evidence or decision authorities
Persistent publication record
Canonical URL, DOI, archival deposit, and operative licenses remain publication blockers
A citation that tells the truth about the release.
Use the title, version, status, and access date until a governed release assigns accountable authors, a canonical URL, archive, and persistent identifier.
LockedIn Labs Research. (2026). LockedIn FDE Benchmark v0.1 — Public method and construct draft (Version 0.1.0-draft). Source preview; not a validated assessment, completed study, certification, or authorized hiring instrument.
- Paper ID
- LL-FDE-METHOD-0.1
- Version
- 0.1.0-draft
- DOI
- Not assigned
- Release date
- Not assigned
- Canonical URL
- Pending publication decision
- Source-review cutoff
- 1 September 2026
- Peer review
- Not externally peer reviewed
- License
- Pending publication decision
- Corrections
- Append-only release history proposed
Citation identifies source and status only. This draft is not a validated assessment, completed study, certification, or authorized hiring instrument.
