FDE BenchmarkPublic document reader
Available in sourcev0.1 · live public draft · official release pending

Canonical source view. This renders the bundled repository document as readable HTML. Source metadata remains in the repository, and relative document links remain inside this reader.

The FDE Benchmark v0.1: a technical report on measuring forward-deployed engineering work

Author: LockedIn Labs Research
Publication class: technical report accompanying the v0.1 public draft
Date: September 3, 2026
Canonical URL: https://fdebenchmark.org/report
Release posture: v0.1 Public Draft; official numeric scoring disabled; no participant evidence admitted

Abstract

The LockedIn FDE Benchmark is a public, versioned attempt to describe forward-deployed engineering work well enough that a claim about it can be checked. This report accompanies the opening of the benchmark's public draft to search on September 3, 2026, at https://fdebenchmark.org. It describes what the benchmark measures (ten domains and 55 capability families, released as ontology 0.1.0), how evidence is graded (an ordinal ladder from E0, a claim, to E5, independent field corroboration), what the benchmark would eventually produce (a verified production record, a capability profile, and a governed credential tier, kept deliberately separate), and how the same production problem could be attempted by a person, an AI system, and the two together without pretending the results are interchangeable. It describes the governance the steward has bound itself to, including a naming decision that retired the phrase "FDE Score" and the withdrawal of a staffing funnel from the benchmark's own origin. And it is explicit about what this release does not contain: no participant evidence, no scores, no certification, and no ranking of people. The argument is that a benchmark for consequential human work has to publish its rules before its results, and has to be built so that the organization running it cannot quietly profit from what it measures.

1. The problem

As of the benchmark's September 1, 2026 review of first-party records, four large organizations were using four different labels for people who sound like they do the same job. OpenAI distinguishes a Forward Deployed Engineer from a Forward Deployed Software Engineer, with the first owning discovery through adoption and the second building alongside. Accenture describes some of its practitioners as forward deployed engineers and, internally, as reinvention deployed engineers. Wipro describes a global FDE talent pool and a separate plan to certify 10,000 Front-Line Delivery Experts. Cognizant has a Frontier Certified workforce model. Those are the first-party records reviewed in the benchmark's first field note; none of them is doing anything wrong. Each label is a reasonable organizational choice.

The buyer's problem is that the labels have started to look equivalent while the populations behind them are not. "Available to" is not "trained." "Trained" is not "certified." "Certified" is not "deployed." And "deployed" does not mean a measurable outcome was attributable to the person in the room. The benchmark's second field note walks that ladder using the announcements that produced it, and finds the market already knows the difference: Anthropic's own partner tiers combine certified individuals with production deployments and public customer stories rather than counting certificates alone.

The title is spreading faster than the evidence of who holds it. That problem lands on the enterprise at the last mile, where a general model meets a particular claims workflow, a particular health record, a particular maintenance window. The fourth field note puts the hard skill there: not prompting, but owning the boundary between a powerful model and an organization's data, permissions, controls, and acceptance conditions. Nobody fails at the demo. Organizations fail at production, and the person who was supposed to carry the system across that line is the one whose capability is hardest to verify from a résumé.

The claim of this report is simple. The unit of proof for forward-deployed engineering is the production record, not the title and not a number derived from the title. And the institution that measures it has to be designed so that it cannot cheat, because the organization founding this benchmark also trains, employs, and staffs the population it would measure.

2. What the benchmark measures

The construct is demonstrated capability to move an ambiguous enterprise problem through discovery, architecture, hands-on engineering, evaluation of AI and software behavior, integration, security and governance, production operation, outcome substantiation, and capability transfer. The v0.1 release narrows the role context to AI-native forward-deployed engineering. Forward deployment is the older and broader operating model; AI-native deployment is the specialization measured first, and the benchmark card says plainly that transfer to non-AI contexts is not claimed.

The construct is released as a versioned ontology, version 0.1.0, with ten domains and 55 capability families. The released JSON is the same file the public explorer reads; the site cannot show a domain that the release does not carry. The domains, in released order:

  1. Forward-Deployed Discovery and Problem Framing (6 families)
  2. Software and Systems Engineering (7)
  3. AI-Native and Agentic Engineering (8)
  4. Architecture and Technical Judgment (6)
  5. Enterprise Readiness, Security, and Governance (7)
  6. Production Deployment and Operations (7)
  7. Business Outcomes and Value Realization (4)
  8. Forward-Deployed Communication and Leadership (4)
  9. Modernization and Transformation (2)
  10. Domain Specialization (4)

Two things about that list deserve attention. The largest domain is AI-native engineering, with eight families, because that is the specialization this release examines first. The smallest is modernization, with two, which says more about how much the steward currently knows than about how much the work matters. Family counts are not weights. The released domain file does carry weights in basis points that sum to 10,000, and the release manifest that binds it says synthetic_only: true and official_numeric_output_allowed: false. The weights exist so the reference software can be tested. They have not been validated against anything.

Each domain in the explorer carries one observable signal, the kind of thing a reviewer would look for. For discovery: rejects an attractive AI use case when the operating constraint makes it unsafe or uneconomic. For production: shows how a launch was observed, rolled back, remediated, and returned to service. For outcomes: separates contribution, correlation, and causal confidence in a claimed business outcome. These are hypotheses. The method paper groups the domains into five reader arcs and says in the same breath that the arcs are for navigation, not scoring.

What the construct excludes as direct inputs matters as much as what it includes: employer or school prestige, compensation, title, fame, social following, training purchase, sponsorship, friendship, and tool-brand allegiance. The scoring architecture turns several of those into commitments the public site repeats on every visit. LockedIn employment adds no points. Training completion adds no points by itself. Unknown evidence becomes uncertainty, not a low score.

The honest gap: no job analysis has been performed. The construct rests on steward synthesis, published frameworks, market observation, and the released ontology. The benchmark formation protocol lists what is missing before the construct can be called established: practice analysis, critical incidents, work-product inventory, importance and criticality ratings, adjacent-role study, and external content review. A job-analysis preregistration candidate exists. No participant has been recruited.

3. The evidence ladder

The benchmark grades evidence, not people, and it grades it on an ordinal ladder. The six levels are defined in the public site's source (apps/public-site/src/data.ts) and the specification:

  • E0, Claimed. A résumé, application, profile, or interview assertion. Useful for discovery; never sufficient for a verified capability claim.
  • E1, Knowledge validated. Knowledge demonstrated through a structured assessment, credible credential, or oral examination tied to a public rubric. The practitioner can explain it.
  • E2, Observed. Capability observed in a controlled lab, live simulation, architecture defense, paired exercise, or supervised challenge. The practitioner can perform it.
  • E3, Demonstrated. Real engineering artifacts substantiate the capability and the practitioner's contribution, subject to provenance and review.
  • E4, Production proven. The capability was exercised in a live system with operating evidence. Production access is context, not an automatic quality judgment.
  • E5, Field validated. Reserved for independent corroboration. A future E5 state would require production evidence, individual contribution, and material claims to pass an approved field-validation protocol. It is not operational in v0.1.

The ladder is ordinal. E4 is more direct than E2; it is not twice as good, and the benchmark does not add levels together. The method paper's phrasing is that proof strengthens by directness, not by volume, and the corollary is the most important rule in the evidence model: the same deployment described in a résumé, a slide deck, a repository, and a reference letter is one lineage group, not four confirmations. The reference scorer collapses lineage before any judgment of diversity or corroboration, so that five views of one deployment remain one deployment.

Two further rules keep the ladder from becoming a proxy for privilege. Production access is context, not credit; a person whose production evidence cannot leave a regulated building is not marked down for it. And missing evidence returns an explicit Insufficient Evidence state rather than a low estimate. The scorer cannot infer weakness from silence.

E5 is reserved on purpose. The steward could have called it "reference-checked" and moved on. Instead the level is held back until a field-validation protocol exists that says who may corroborate, how a corroborator's own conflicts are handled, and what a corroboration receipt looks like. Until then the top of the ladder is empty, and the site says so.

4. The outcome architecture

Ask what the benchmark would produce for a person and the natural answer is a score. The founder's original framing, as the naming decision of September 2, 2026 records it, was exactly that: a proud number on a résumé, raised by shipping production work. The decision record explains why that framing was retired. The single number contained three products with different owners, different error modes, and different legal footprints.

The three-layer outcome architecture separates them:

  • The Verified Production Record is the factual layer. What happened, where, and what this person actually contributed. Its authority is evidence operations. It preserves facts and provenance; it does not infer capability or confer standing.
  • The Capability Profile is the measurement layer. A versioned, multidimensional interpretation of observed work, with domain-level estimates, coverage and uncertainty, rater agreement, and an instrument version. Its authority is a measurement panel. It does not become a psychometric instrument until reliability, validity, fairness, and use claims are supported, and none are.
  • The Governed Credential Tier is the issuance layer. A shareable, verifiable status with an issuer, scope, version, expiry, and a recourse path. Its authority is a future non-conflicted credential body. Verifying the issuer would not establish that the underlying claim is true.

The rule that binds them, in the words the public site uses: evidence may advance a record, but it may not silently advance a credential. Each transition requires a named authority, a versioned rule, a conflict check, and an appeal path. The same person or team may not hold sole approval over all three layers for a case or a release.

Why not a scalar? Not because one is impossible, but because it is a new interpretation with its own validity argument, and nobody has made that argument yet. The future scalar-score authorization gate is a published research proposal, not a refusal. It sets out thirteen evidence-readiness receipts, then an independent method review, then a governance decision, in that order, so the final vote cannot become a prerequisite for the review it is supposed to follow. Even a scalar that passed every gate would be an optional, purpose-bound projection. It could not replace the record, the profile, or the Insufficient Evidence state.

The obvious objection is that a benchmark without a number is not a benchmark. Model benchmarks publish leaderboards; buyers want a figure for a slide. The objection is fair about model benchmarks and wrong about this one. A model can be re-run a thousand times under a pinned environment, and a wrong number costs a lab a footnote. A person cannot be re-run, the conditions of the attempt change the result, and a wrong number costs someone a job. The personnel-testing standards the firewall cites (the Standards for Educational and Psychological Testing, the SIOP Principles, the EEOC Uniform Guidelines) require validity evidence for each intended use before a score touches a consequential decision. Publishing a scalar before that evidence exists would not make the benchmark more rigorous. It would make it a leaderboard of people, which the specification prohibits.

5. The shared work challenge design

The most distinctive research proposal in the v0.1 draft is the Shared Work Challenge Battery. The idea is that a single versioned production problem can be attempted under three declared conditions: a human working without generative AI under an explicit tools policy, an AI system or agent scaffold working under an explicit model, tool, memory, authority, and resource envelope, and a human with a declared AI system working as one operating pair.

Every scenario binds one immutable manifest: the problem and enterprise context, the acceptance contract, inputs and hidden state, deliverable schemas, observable events, rubrics, safety and authority boundaries, a learnability review, and contamination controls. The three conditions reference the same manifest. Condition adapters may change interaction mechanics where predeclared; they may not change the problem, the acceptance criteria, or the available facts.

The public site illustrates the design with one synthetic scenario, SWC-01: recover a failing AI-assisted claims workflow whose production release is breaching its service objective, exposing sensitive data in traces, and generating inconsistent decisions. The shared task has four stages: frame the problem, build and defend the intervention, operate through a live failure injection, and prove the result by leaving tests, decisions, and outcome evidence behind. What the benchmark would observe differs by condition. For the human alone: problem framing, implementation judgment, technical defense, operational handoff. For the system alone: task completion, tool use, recovery behavior, evidence trace, failure handling. For the pair: delegation, verification, intervention, escalation, and accountability for the outcome.

The design's most important sentence is a refusal. Comparability is a research question. A shared scenario improves experimental control; it does not make the constructs identical. Interfaces, time, compute, information access, tool authority, and interaction costs differ materially between a CPU-minute and a human-minute. Results therefore remain condition-specific unless a later linking study supports a particular comparison, and the firewall reserves any human-versus-model comparability claim to the proposed Commission. This is not a spectacle about whether AI beats humans. It is a way of asking, for a particular task and operating boundary, how the work gets done and where it fails.

No challenge has been administered. No attempt data are admitted. No human, model, team, or score is represented anywhere in the design.

6. Governance

LockedIn Labs founded this benchmark, funds it, and stewards the repository. It also trains forward-deployed engineers, employs them, and staffs them to clients. The public build record states the conflict in a heading rather than a footnote. Disclosure does not resolve it. The design has to prevent a conflicted role from controlling the affected decision.

The independence firewall is the instrument. It defines a staged independence model, I0 through I4, and places the project at I0, founding stewardship, where no proposed body is seated. Before any official human assessment, result, credential, or tier, an external-majority FDE Benchmark Commission must be seated: seven to nine voting members, more than half externally independent, an external chair and an external measurement lead, at least one occupational-measurement expert, at least one member representing candidates or the public, and LockedIn Labs with its affiliates never forming an approval majority. Independence means no material compensation from LockedIn in the prior 24 months. Founding appointments are confirmed by a temporary panel of three to five external members that LockedIn cannot sit on, direct, or override, and once seated, LockedIn cannot unilaterally remove an external member.

Underneath the Commission sits a role and data firewall for operations. A trainer cannot rate anyone they trained in the prior 24 months. A recruiter or sales lead cannot review candidate evidence, set an outcome, or decide an appeal. Compensation cannot depend on pass rates or issuance. Assessment data cannot enter a staffing, employment, sales, or training system by default. The firewall calls itself a project constitutional safeguard, informed by ISO/IEC 17024 and the NCCA standards, with no conformance claimed, and treats a separate legal entity as a structural option the Commission must decide on, not a universal requirement.

The firewall was written before it was needed. Then it was needed.

Between August 31 and September 2, 2026, the benchmark's public origin served a talent-matching prototype at /find-an-fde. A visitor could describe an engagement; the page ranked a roster of invented, fictional candidates against it, labeled them by fit, and ended with two links to LockedIn Labs' sales contact form. A "Find an FDE" button sat in the primary navigation of every page. The project's own deployment evidence brackets the window: the route appears as a passing live route in the evidence record published at 21:30 UTC on September 1 and is absent from the record published at 00:29 UTC on September 2. It was certified live, by the project's own acceptance check, for at least two hours and fifty-nine minutes. The check passed because it tested the routes it was given, and the funnel was one of them.

The withdrawal decision record, dated September 3, 2026, records what was done. The route was unmounted, then the source was deleted rather than left unreferenced. The publication verifier now fails the build if the sales destination, the linked platform host, the call-to-action phrases, the fixture person's name, or the retired phrase "FDE Score" appears anywhere in the public site's source. The record is served at the withdrawn address, so a stale link resolves to the explanation. It also records what the evidence cannot bound: deployments before September 1 were published by hand with no commit reference, so the project cannot say when the route first appeared. That gap is why every deployment since runs through a commit-pinned pipeline and writes an evidence file.

The naming decision of the day before belongs in the same account. Beyond retiring "FDE Score," it fixed the method description as "AI-assisted observation, human-rated, deterministically computed," required the Commission to be described as proposed until its roster and receipts are public, reserved the words "official," "independent," "validated," "certified," and "accredited" until the corresponding receipt exists, and rewrote sponsor tiers so that no payment confers a vote, a seat, a logo, or endorsement language.

None of this makes the benchmark independent. The firewall says so itself: a policy, a committee name, or a separate company does not by itself establish impartiality. What the record shows is a conflict made visible and a breach corrected in the open, with the control that prevents recurrence named. Any structural answer the Commission eventually gives should be judged against whether it would have prevented this one.

7. What this release does not claim

The status register is the authoritative list, and it controls whenever marketing copy or a roadmap date disagrees with it. The short version:

  • No participant evidence has been admitted. The authority page's boundary ledger reads: human participant evidence, not admitted; official results or credentials, not authorized; independent validation, not recorded. The sponsor-reported target of 500 practitioner and 50 executive interviews has no admissible records and contributes no finding.
  • No score exists. Official numeric scoring is disabled in code and policy. The reference scorer runs on labeled synthetic fixtures only. No weights, thresholds, bands, percentiles, or comparison population have been validated.
  • No certification, verification, or credential is issued. The specification prohibits representing a person as certified, verified, or field validated by the v0.1 system.
  • No person is ranked. A named-person leaderboard is not authorized, and the specification bars views that reconstruct one through filtering or sorting.
  • No validity evidence has been established. Of the seven validity strands in the method paper, six are marked not established and the seventh, reproducibility, rests on a synthetic foundation only.
  • No governance body is seated. The Commission, the steering and integrity committees, the measurement and fairness advisory group, and the appeals panel are proposed. No member, chair, meeting, or decision is claimed.
  • No outside organization endorses this. No partner, sponsor, signatory, validator, or reviewer is represented. The four studies the method paper registers (S01 to S04) are, respectively, protocol proposed, not started, not started, and blocked on the earlier three, and the public study register records no admitted participant records.
  • No individual author is named. Individual authors, a guarantor, a methods reviewer, and a release authority are recorded as unassigned in the accountability record. No IRB approval or exemption is claimed because no human dataset exists. No DOI has been assigned. The draft is not externally peer reviewed.

The benchmark's credibility depends on those absences being stated as absences. The build record's rule is that every count on the page is derived from a released artifact and every gap is written as a gap. That rule applies to this report as well.

8. How to take part

The public draft opens to search with one inbound channel: the contact form, delivered by email to a monitored LockedIn Labs mailbox. Nothing submitted enters the benchmark or reaches a commercial team. It carries six intents.

  • Correct a published claim (/contact?intent=correction). Every publication page on the site links here as its correction channel, and corrections are recorded under the publication integrity workflow.
  • Offer to review (/contact?intent=reviewer). Measurement, evaluation-integrity, legal, accessibility, or field expertise for a named review role under the reviewer appointment protocol. No review badge is synthesized for an open role.
  • Take part in a study (/contact?intent=participant). For practitioners who would consider a future consented study. Nothing is open; the form records interest only, and the participant consent and rights protocol governs what must be true before anyone is asked for anything.
  • Enterprise or design-partner interest (/contact?intent=partner). Not a staffing request, and not routed to sales.
  • Report a security issue (/contact?intent=security), referenced from /.well-known/security.txt.
  • Raise a conduct or conflict concern (/contact?intent=conduct).

The source is public under operative licenses: Apache-2.0 for software, schemas, fixtures, scripts, and the public site; CC BY 4.0 for specification, policy, ontology, and research content. Every document cited in this report is served in the site's own document reader, and the released ontology is the same JSON the explorer renders. A reader who wants to check a claim here does not need to ask permission.

9. Roadmap

The public roadmap is gate-based. It describes sequence and exit criteria, and it says in its first paragraph that a calendar target cannot waive a security, privacy, measurement, accessibility, governance, legal, or integrity gate. The foundation phase, the public kernel, is what this report describes. What follows it, in order:

Controlled-pilot preparation seats the external-majority Commission at independence stage I1 before any real-person pilot is authorized, completes job analysis and external content review, designs the Shared Work Challenge Battery, obtains counsel review for each intended jurisdiction, and builds the private operating boundary that the public site deliberately does not contain.

Feasibility calibration recruits a permissioned cohort of roughly 20 to 30 practitioners, not to prove validity but to learn whether the construct, evidence model, and reviewer process are usable. A cohort of that size cannot support percentiles, population norms, subgroup fairness claims, or predictive employment claims, and the roadmap says so. Failed hypotheses and broken graders are to be published at a non-identifying level.

A verified-profile beta, at the earliest in months four to six and not tied to a promised date, would test a controlled verification service in the three-layer sequence: the factual record first, the profile second, and a credential-interoperability candidate third with issuance still disabled. A v1.0 consideration, at the earliest in months seven to twelve, requires at least two calibration cohorts, externally reviewed rubrics, met reliability thresholds, operating appeals, independence stage I3, an external-majority Commission vote for the exact release, and a published structural-entity decision. Predictive employment claims are not implied by a v1.0 label.

One direction in the steward's current planning is not yet in the normative roadmap, and this report records it as an intention rather than a plan the gates recognize. Since September 2, 2026, the steward's working position has been to measure the engagement before the engineer: to make a registry of forward-deployed engagements, with their operating context, acceptance conditions, and outcome evidence, the first measurand, and to derive any person-level record and credential from that registry later rather than the other way around. That ordering follows from the outcome architecture, where the factual record precedes the profile. But it has not been written into the roadmap, the status register, or the specification, and the governance question of when a person-level credential could be issued at all remains open. Until those documents change, the engagement registry is a stated direction, and the roadmap above is the plan.

The roadmap also names what it defers: a general-purpose FDE foundation model, public leaderboards, cost-per-person indices, automated employment recommendations, and any staffing funnel inside the neutral benchmark surface.

References

Site pages, by canonical URL:

  • Benchmark home: https://fdebenchmark.org/
  • Capability explorer (ten domains, 55 families): https://fdebenchmark.org/explorer
  • Method paper: https://fdebenchmark.org/method
  • Research program: https://fdebenchmark.org/research
  • Research authority record: https://fdebenchmark.org/authority
  • Publication desk: https://fdebenchmark.org/publications
  • Build record: https://fdebenchmark.org/build
  • Contact form: https://fdebenchmark.org/contact
  • Withdrawal record, served at the withdrawn address: https://fdebenchmark.org/find-an-fde

Bundled source documents, all served in the site's document reader:

Disclosures

LockedIn Labs is the founder and current steward of the benchmark and may also train, employ, staff, and provide services to the practitioners and organizations the benchmark would measure. The project is at independence stage I0. No finalized public funding statement is on record, and no outside sponsor is represented as having funded, reviewed, or endorsed this report. AI-assisted tools supported source discovery, drafting, software development, and testing; they are not evidence or decision authorities, and LockedIn Labs Research remains accountable for every claim. This report is not externally peer reviewed. No participant data were used because none exist.

Suggested citation: LockedIn Labs Research. (2026). The FDE Benchmark v0.1: a technical report on measuring forward-deployed engineering work. LockedIn FDE Benchmark, version 0.1.0-draft. https://fdebenchmark.org/report