Publication desk
Benchmark architectureVersion 1.06 min

Domain expertise is part of the engineering stack.

Healthcare, legal, semiconductor, and other consequential deployments change the constraints that define competent work. A vendor-neutral core therefore needs separately governed domain packs—not one generic interview loop.

LR
Corporate authorLockedIn Labs Research
Review statePublic-draft editorial review
LockedIn Labs · Field noteOpenAI Healthcare · Wipro Applied AI CoE
Claim boundary

Read the thesis at the strength of its evidence.

This supports

  • A source-grounded editorial interpretation of the current FDE market
  • Questions the benchmark should test through observable work and evidence
  • A documented rationale for specific construct and publication choices

This does not support

  • A population estimate, pass rate, or claim about how many people are qualified
  • Predictive validity, certification, ranking, or an employment decision
  • Independent endorsement of the benchmark or LockedIn Labs

An AI system does not enter “the enterprise.” It enters a claims workflow, a clinical operation, a manufacturing line, a contract-review process, a finance control, or a public service. The organization may purchase a general model. The deployment succeeds or fails in a particular domain.

That is why domain expertise should not be treated as a soft accessory to Forward Deployed Engineering. It is part of the engineering stack: the knowledge required to identify the real unit of work, interpret constraints, define failure, select evidence, and recognize when a technically valid answer is operationally wrong.

The implication is not that every FDE must be a lifelong industry specialist. It is that a person should not be represented as ready for a high-stakes domain merely because they can build a general agent.

The market is already specializing the role

OpenAI's current FDE hiring surface includes general roles alongside healthcare and legal positions. Its healthcare FDE specification calls for knowledge of payer and provider operations, electronic health records, interoperability standards such as HL7 and FHIR, privacy and authorization, human review, auditability, and customer-specific evaluation thresholds. Those are not decorative keywords. They change the architecture, evidence, rollout, and definition of a safe outcome.

Wipro's June 2026 Applied AI Center of Excellence announcement describes FDEs combining business-process knowledge, technology-landscape knowledge, and hands-on model expertise inside client environments. Its initial industry work spans mortgage, healthcare, airlines, manufacturing, and consumer sectors.

Cognizant's expanded work with Anthropic provides another signal. Its July 2026 announcement describes Claude systems across manufacturing, life sciences, and insurance, including a contract-intelligence deployment and a tool for underwriting research with deployment-specific reported outcomes. These are vendor-reported examples, not independent impact studies, but they demonstrate that evaluation and value are expressed in domain workflows rather than generic model capability.

Anthropic's work with Deloitte similarly emphasizes industry-specific solutions and a Center of Excellence intended to move implementations from pilot to production at scale. The partnership announcement does not establish individual competence, but it shows why services firms organize model expertise together with industry and implementation knowledge.

Domain knowledge changes what “good” means

A generic engineer can ask whether a response is accurate. A domain-ready FDE must ask whether the response uses the right source, reaches the right operator, preserves required approvals, fits the system of record, and can be acted on under the governing rules.

In healthcare, a seemingly useful output may create risk if it crosses the line between administrative support and clinical judgment, mishandles protected information, or cannot be reconciled with the patient record. In insurance, faster underwriting research is valuable only if sources, adverse-action obligations, authority, and review are handled correctly. In manufacturing, a recommendation can be analytically sound but unusable if it ignores equipment state, maintenance windows, safety controls, or the cost of stopping a line.

Domain expertise also improves discovery. Operators rarely describe their needs as a clean technical specification. They use local language, inherited process names, informal exceptions, and risk judgments developed through experience. The FDE must learn enough to distinguish a real constraint from a habit, and an attractive automation from a dangerous simplification.

This does not mean subject-matter experts and engineers are interchangeable. The strongest deployment model joins them. Domain readiness includes knowing when the practitioner can decide, when a subject-matter expert must decide, and how that authority is represented in the system.

Questions for hiring and evaluation

For a domain-specific FDE claim, ask:

  1. Which workflow do you understand end to end? Ask the candidate to name actors, systems of record, decisions, exceptions, and downstream consequences.
  2. Which domain constraint changed your architecture? Look for a concrete privacy, safety, regulatory, operational, or economic decision.
  3. What does a severe error look like here? Generic hallucination language is insufficient; require domain-specific harm and detection paths.
  4. Who held decision authority? Ask where engineering judgment ended and operator, clinical, legal, risk, or compliance authority began.
  5. What domain evidence supported launch? This could include expert review, representative cases, process metrics, audit evidence, or controlled rollout results.
  6. Which result did the domain team reject? A practitioner should be able to explain why a technically plausible result failed operational review.
  7. What would not transfer to another industry? Domain maturity includes recognizing the limit of one's own pattern.

An evaluation should use realistic artifacts—policies, process maps, data definitions, interface constraints, and contradictory stakeholder accounts—rather than trivia about industry terminology. The task is to apply domain knowledge to a deployment decision, not pass a vocabulary quiz.

LockedIn Benchmark implication

The LockedIn Benchmark should use a vendor-neutral core and separately versioned domain packs. The core can assess discovery, architecture, implementation, evaluation, production operation, outcomes, and field learning. A domain pack should add the workflows, constraints, evidence expectations, and failure modes that materially change performance in that context.

Domain claims should be explicit and bounded. “FDE-qualified” and “healthcare deployment-ready” should not be treated as identical. A candidate may demonstrate strong core capability with no observed evidence in a particular sector. The result should say exactly that, without converting unobserved domain experience into a negative judgment about general ability.

Domain packs also require their own governance. They should be built with qualified practitioners and affected operators, reviewed for accessibility and fairness, versioned as technology and rules change, and validated against relevant work. A universal composite score should not hide a critical domain gap.

Limitations

The cited sources are first-party role descriptions and partnership announcements. Their examples are selective and, where outcomes are reported, vendor-authored rather than independently reproduced. They do not establish which domains require separate assessment, which knowledge predicts performance, or how much experience is sufficient. LockedIn Labs has not completed the field research, job analysis, domain-panel review, or criterion validation needed to publish domain-readiness decisions. The core-plus-domain-pack design is a proposed research direction, not a validated classification system.