Running an Internal Forward-Deployed Engineering Practice

A practical system for governing FDE engagements, evidence, releases, incidents, capacity, ownership, value, and the practice itself.

[ fig.1 ] management_loops
WIDEST SCOPEEACH LOOP ENDS IN A DECISIONPRACTICEexpand / reorganise / closeis this still the right arrangement?PORTFOLIOfund / sequence / stopwhich commitments deserve capacity?OPERATIONcontinue / correct / transferdoes the live arrangement remain useful?RELEASErelease / narrow / pausedoes evidence cover this boundary?ENGAGEMENTbuild / change / stopdoes the proposed workflow address the problem?EVIDENCE RISESCOMMITMENT DESCENDSTHE NEXT DECISION STARTS THE FOLLOWING LOOPor returns to an earlier loop when evidence requires a change
evidence rises, commitment descendsAUTO
On this page

The previous article established how the practice gets authority, funding, evidence, and an operating owner. This article starts when the first engagement enters delivery. The practice must decide whether a proposed solution works, who can operate it, and what the enterprise should do next.

The main case is an illustrative investment-banking transaction-screening application. It prepares a source-linked brief for a potential acquisition target. A banking analyst checks the brief before a senior banker decides whether deeper work is justified. The trading and manufacturing cases are also illustrative. They appear only when they expose a different operating condition.

I would answer these questions through five management loops. Each loop asks one question, gathers evidence, records a decision, and funds the next commitment. The sections move from engagement evidence to release, live operation, portfolio value, and practice renewal. A finding can return the work to an earlier loop when its evidence changes the decision.

The practice lead governs the portfolio and approves practice-level commitments. The FDE lead owns delivery evidence for one engagement. The business workflow owner owns the process and the meaning of a useful result. The service owner owns deployment, support, dependencies, and technical change. Platform and control owners own shared capabilities and approval requirements. One person may hold more than one role, but each responsibility needs a named owner.

An engagement is a bounded piece of work that investigates a business problem, develops a proposed solution, and establishes who can operate it. A workflow is the business process that the solution supports. The practice should increase commitment only when evidence supports the next decision.

The table names each loop and the decision it governs. Each loop has a decision, evidence, owner, and next commitment.

Management loop Core question Main decision
Engagement Does the proposed workflow address a diagnosed problem? Investigate, build, change direction, or stop
Release Does evidence support operation within a defined boundary? Release, narrow, contain, or pause
Operation Does the live arrangement remain useful and supportable? Continue, correct, transfer, or retire
Portfolio Which commitments deserve limited enterprise capacity? Fund, sequence, pause, or stop
Practice Does the FDE model remain the right organisational response? Expand, reorganise, partner, or close

The sections follow these loops in order. First, the practice governs an engagement from requirements to release. It then governs live operation, ownership, portfolio value, and its own future. After each loop, the next decision starts the following loop or returns to an earlier loop when evidence requires a change.

Govern each engagement through evidence

Refine the engagement brief before committing to a solution

Requirements refinement is the first activity in the engagement loop. It converts a requested feature into a precise account of the work. It defines users, constraints, decisions, and the intended result. The practice should require this account before committing to a technical design.

The process starts with a case walkthrough. A case walkthrough follows representative work from its trigger to the final business decision. The team records each action, source, system, wait, exception, handoff, and decision. Formal procedures describe the intended process. Case walkthroughs reveal how staff perform it.

This method identifies application requirements and the actual operating problem. An operating problem is the work, delay, error, or exposure that the enterprise needs to change. Several causes can produce the same symptom. Therefore, the team must distinguish those causes before choosing a solution.

For example, a sponsor might request an agent that screens acquisition targets. The request names a technology and a broad task. It does not identify the operating problem. Analysts might struggle with source collection, inconsistent measures, peer selection, or senior-review delays. Each cause requires a different response.

Consequently, the team should use several completed target screens as walkthrough cases. Deal analysts can explain the request, sources, reconciliations, exceptions, assumptions, and review outcome. The team should separate active analyst effort from waiting for data or senior judgement. This evidence shows where work accumulates and why.

Each walkthrough must include every role affected by the proposed solution. The banking analyst prepares the screen, while the senior banker uses it for a decision. The engagement brief must describe the evidence, gaps, and assumptions that let both roles review the result without rebuilding the work.

The team should capture the result in an engagement brief. An engagement brief records the diagnosed problem, supported cases, proposed result, exclusions, principal uncertainty, and next decision. It also identifies credible alternatives. These might include better source access, a conventional data pipeline, a changed screening template, or an AI application. The team should choose among them only after comparing the evidence.

The requested agent remains one proposed solution until this work ends. The team should not assume that an agent or AI application is the answer. Inconsistent source data may make better data preparation the stronger investment. Conversely, analysts may have consistent data but spend time tracing each figure. A source-linked AI brief may then address the diagnosed problem. A changed recommendation is useful progress when evidence supports it.

Define how the business result will become visible

Next, the practice needs a measurement plan for the proposed solution. A measurement plan defines the business result, eligible population, comparison, observations, and interpretation limits. The proposed solution is the change being tested. The plan asks whether that change improves the business process around transaction screening. A model evaluation answers a narrower question about application behaviour.

The first step in writing the measurement plan is to define the intended result and its audience. Leaders may need evidence of lower effort, faster decisions, more complete source coverage, or increased screening capacity. Source coverage means whether the brief includes the required authoritative sources for an eligible target. These results support different decisions. Therefore, the plan must state which result supports the next funding or release decision.

The plan must then define the eligible population. For transaction screening, relevant characteristics include sector, transaction type, source availability, target complexity, and reviewer seniority. The plan should measure effort and quality across the complete workflow. That boundary includes collection, preparation, correction, review, clarification, and manual fallback.

Attribution means deciding how much of an observed change is associated with the proposed solution rather than with other changes. In this case, the question is whether the transaction-screening application caused a change in effort, time, quality, or capacity. Attribution requires a comparison with the process that would otherwise operate. One group contains eligible screens completed with the application. A second group contains comparable screens completed through the existing process. Sector, complexity, source availability, workload, and reviewer experience can affect the result.

Consequently, the practice should choose the strongest practical comparison method. Random allocation assigns eligible screens to the application or existing process by chance. It can support stronger causal interpretation when both routes are acceptable, the assignment does not breach controls, and the business can tolerate the scheduling constraint. Other formal options include phased introduction, matched concurrent cases, or a prospective baseline. A phased introduction starts with one defined group before another. Matched cases compare screens with similar characteristics. A prospective baseline records the existing process before the proposed solution starts. The plan should explain its method and remaining uncertainty.

Historical records of completed transaction screens can help when they contain the necessary detail. However, they may omit analyst effort, failed attempts, manual fallback, or case complexity. The practice may then need prospective measurement before release. This follows from the comparison design because missing variables weaken attribution.

Subsequently, the team must watch for behaviour changes. Analysts may reserve difficult targets for manual work. Reviewers may share the new screening template outside the pilot. Staffing may also change during the observation period. These differences can make comparison groups represent different work.

The measurement should include failed, abandoned, and unused application attempts. Measuring only successful briefs answers a narrower question about the application's best cases. The business process still includes recovery and manual work when the proposed solution fails or is not used. Those cases belong in the result when the practice claims an improvement to the complete screening process.

Sometimes the enterprise should act before it has a complete productivity baseline. An urgent control improvement may address an observed exposure. The practice should state that purpose directly. It should not invent a savings claim to justify a different benefit.

Therefore, the impact measurement plan should accompany the delivery plan. It compares the proposed solution with the existing screening process and defines success before behaviour changes. Business owners and finance partners should agree on its definitions. This agreement reduces later disputes about the result.

Use delivery stages to answer one question at a time

The practice must control how commitment grows. I would use delivery stages for this purpose. A delivery stage answers one decision through a bounded period of work. Each stage starts with a question and ends with a recorded decision.

The stage record should state the evidence, limitations, decision, owner, and reconsideration condition. This structure also defines the progress report. The report prepares the next decision and does not become a separate reporting exercise.

These delivery stages sit inside the engagement management loop. They are checkpoints within one engagement. They are not additional management loops. When a release or operation review identifies a new requirement, the practice returns to the relevant engagement stage.

Delivery stage Question Evidence Possible decision
Requirements refinement Which part of the workflow needs intervention? Case walkthroughs, alternatives, and an updated brief Investigate one cause, change the proposal, or stop
Feasibility experiment Can the proposed approach resolve the main technical uncertainty? Representative cases, important exceptions, and reviewer feedback Build a limited workflow, investigate another cause, or stop
Limited workflow Can the complete application operate with controls and recovery? End-to-end tests, state records, access checks, and manual fallback Begin controlled release or correct the design
Controlled release Does the application help eligible work under operating conditions? Workflow results, quality, use, support demand, and incidents Continue, narrow, pause, or prepare expansion
Expansion or transfer Do evidence and operating capability cover the changed scope? Population assessment, receiving-team exercises, and capacity commitments Expand, retain the boundary, change ownership, or stop

The practice lead should identify the decision-maker before each stage begins. A feasibility result may belong to the FDE and domain leads. A release decision may also require business, service, and control owners. An expansion decision usually needs portfolio and funding authority. Unclear authority can leave completed evidence without an actionable decision.

Consequently, a negative stage result can still represent progress. It can prevent investment based on an unsupported assumption. The next section shows how the first experiment produces that result efficiently.

Make the first experiment answer the most important uncertainty

An experiment is a small, time-bounded test of one proposed solution against real or representative transaction-screening cases. It should answer one question that could change the delivery decision. A broad prototype can demonstrate many capabilities while leaving that question unresolved. Therefore, the first experiment should address the uncertainty most likely to change the delivery decision. For transaction screening, that uncertainty may concern source availability, entity matching, reporting periods, or review effort. Access and information definitions may need investigation before model behaviour becomes the main concern.

Suppose the central question concerns whether the application can assemble a reviewable brief from permitted sources. A small exercise should follow that complete path. It retrieves information, prepares comparisons, exposes gaps, and asks a banking analyst to review the result. The exercise needs realistic targets and important exceptions. Its purpose is to expose a weak assumption before larger investment.

I would begin with the simplest design that can answer the question. A fixed sequence of retrieval, extraction, checking, and review may serve the workflow adequately. An autonomous agent adds decisions about which tools to use, when to retry, and when to stop. Those capabilities deserve their additional complexity only when the task requires them. Anthropic's engineering guidance makes a similar distinction between simpler workflows and more autonomous designs. Anthropic's guidance.

The first result may reveal that the original technical goal was incomplete. Accurate summaries, for example, can still require reviewers to open every source document. The team should identify the reason before changing the design. Reviewers may need visible citations, explicit gaps, a different structure, or source information that the application never retrieved. The next experiment can test one of those explanations. This is more useful than tuning a prompt against a vague quality goal.

A trading-surveillance application shows the same logic in a different financial process. Its alert summaries may read well but combine events from different accounts or time windows. A small test can use known alerts to check account identity and event sequencing. If the context is wrong, the practice should fix data selection before adding autonomous tool use.

Each experiment should end with a decision that follows from its result. Adequate evidence may justify the next delivery stage. If a business rule is unclear, the relevant domain owner must resolve the rule before further build work. An unsuitable source may require a narrower scope or another approach. A negative finding becomes valuable when it prevents a larger investment based on the same unsupported assumption.

In addition, the team should record what the experiment cannot establish. A small set of representative cases can reveal feasibility problems without estimating rare failure rates. Successful expert review in a quiet session has a clear limit. It does not establish ordinary review quality under production workload. Keeping these boundaries visible prevents early evidence from carrying more weight than it can support.

Build the complete limited workflow

Once the approach appears feasible, the next commitment should produce a complete limited workflow. This is the smallest end-to-end service that can generate operating evidence. It includes normal use, control, state, review, support, and recovery. A demonstration can omit these concerns because its builders can intervene. A service needs explicit behaviour when information is missing or processing fails.

For transaction screening, the proposed solution begins when a deal analyst submits a target for an initial screen through an approved internal channel. The application retrieves permitted sources, prepares peer and period comparisons, identifies gaps, and presents a source-linked brief. The analyst can accept the preparation, reject it, or request more information. The application records the result. Later users can then distinguish a generated proposal from a reviewed judgement.

State clarity means preserving the difference between a model proposal, a human review, and a completed business decision. That distinction should exist in application state as well as in the interface. A model may propose structured information. However, application code must validate required fields and permitted transitions before saving consequential changes. A response that says a task is complete does not establish that the system reached the correct state. The service must preserve what the model claims and what the application actually did.

The same principle applies to failure handling. If retrieval loses access to a source, the output should identify missing evidence rather than imply a complete review. If a retry occurs after partial processing, the system should avoid duplicate cases or repeated consequential actions. The service operator also needs a way to identify incomplete work and return it to an appropriate manual process.

A trading application requires similar state clarity around its permitted actions. Retrieving an order record, explaining an alert, and changing alert status are different events. Each event has different authority and consequences. The application must record which event occurred and who authorised it. A useful explanation cannot silently become permission to alter a supervisory record.

The practice should keep a version record for each evaluation and release. The record identifies the application build and the material model, source, policy, and data changes. This makes a later regression easier to investigate. It also gives the release authority a clear account of what it approved.

Resource limits belong in this complete design. The application needs suitable bounds for execution duration, retries, tool calls, concurrency, and spending. The values should follow from measured demand and operational requirements. The team must also observe the full workflow. Slow retrieval or a review backlog can dominate despite a fast model response.

The limited workflow now provides something concrete to evaluate. It includes the decisions, evidence, state changes, and recovery paths that the enterprise intends to operate. The evaluation can therefore assess a working arrangement rather than an isolated response.

Evaluate the failures that would change the release decision

Evaluation begins with error analysis. Error analysis is the practice of grouping observed failures by the requirement or decision they affect. It connects the evaluation back to the requirements brief. An incorrect date, omitted source, unauthorised retrieval, and unrecoverable interruption are different error classes. The plan must preserve these distinctions. The release authority can then understand what each result supports.

For transaction screening, the evaluation team can use the following table. “Responsibility” names the requirement that must hold. “Concrete failure” gives a case that violates it. “Evidence required” names the record needed to decide whether the requirement held.

Responsibility Concrete failure Evidence required
Select the right evidence The brief uses figures for the wrong legal entity Source identity, entity matching, and analyst assessment
Identify missing information The brief omits a filing required for the comparison Screening requirements and stated evidence gaps
Interpret reporting periods The brief compares trailing results with a different period Source dates, measure definition, and resulting comparison
Respect access A request retrieves another deal team's restricted material Authorisation checks and the actual retrieval record
Maintain workflow state A retry creates duplicate work State transitions and controlled recovery tests
Support human judgement Reviewers accept a conclusion that the evidence does not support Observed review behaviour and independent case assessment
Recover operation A dependency failure leaves cases without a route forward Failure exercises, operator actions, and recovery results

Different assessment methods fit different questions. Code can check schemas, calculations, permissions, and state transitions. Domain experts can assess whether the evidence supports a business conclusion. A model grader, sometimes called an LLM-as-a-judge, can compare qualitative outputs against a rubric when experts validate that rubric. Each method only supports the property and conditions it tested. A passing answer-quality check does not prove correct permissions or acceptable review effort.

Anthropic's agent evaluation guidance distinguishes code, model, and human graders, and emphasises the wider system and its resulting outcomes. Anthropic's evaluation guidance. For the transaction-screening application, that means inspecting the saved screen state and evidence trail alongside the generated brief. The final brief cannot reveal every action that preceded it.

The general evaluation principle is coverage of the conditions that matter to the release decision. A strong average can conceal weak performance for one sector, source, or transaction type. Therefore, the evaluation collection should include representative cases and deliberate challenge cases. Representative cases describe expected use. Challenge cases probe known risks and requirements. A deliberately difficult set cannot estimate ordinary failure frequency.

I would therefore separate assessments intended to represent normal usage from targeted challenge tests. Both inform the release decision, but they answer different questions. The normal sample helps assess expected behaviour for the proposed population. The challenge collection tests particular failures and restrictions that deserve attention even when they occur rarely.

Development and assessment material should also remain distinguishable. Engineers naturally adapt the application to cases they inspect repeatedly. Reserved cases provide a better check on whether a change generalises beyond those examples. Newly observed failures should join a collection of regression tests, so later changes cannot quietly recreate the same error. The team should also retain fresh assessment material where feasible. All material must respect data restrictions.

The graders themselves require evaluation. A model grader can accept a plausible but unsupported explanation or reject a valid answer with unfamiliar wording. Experts can disagree because the rubric is unclear or the business policy leaves room for judgement. Inspecting those disagreements helps the team improve the assessment instead of treating every recorded label as unquestionable truth.

User acceptance provides another useful observation with a narrower meaning. A reviewer may accept a correct output. They may also tolerate an error or lack enough time for further checking. Acceptance therefore cannot automatically serve as a correctness label. The practice needs independent assessment where the decision requires evidence about actual quality.

Agent evaluation adds a further distinction between acceptable outcomes and permitted routes to those outcomes. Several tool sequences may complete a task correctly, so requiring one exact sequence can reject valid behaviour. Conversely, an agent may reach a correct final answer after an unauthorised intermediate action. The evaluation should allow legitimate variation while checking constraints that must hold throughout execution.

Finally, the evaluation report needs to explain its uncertainty. A small sample with no observed failure cannot establish a strong claim about rare events. Required testing depends on consequence, variation, expected frequency, and the confidence needed for the release decision. Where release depends on a quantified reliability claim, the team should obtain suitable statistical support. A convenient universal test count would not resolve those questions.

[ fig.2 ] zero_failures
[ interactive fig.2 ]What a clean run does not proveno failure observed is not a rate

How to use: set the number of evaluation cases that ran with no failure, then set the failure rate the release decision relies on. The top bar shows every rate a clean run of that size still leaves possible. The lower bars compare the run you did with the run the claim needs.

FAILURE RATES STILL CONSISTENT WITH A CLEAN RUNNOT RULED OUT BY THIS RUNRULED OUT1 in 101 in 1001 in 1,0001 in 10,000RARER FAILURESMORE COMMON FAILURES1 in 34YOUR CLAIM: 1 in 100THE RUN YOU DID100THE RUN THE CLAIM NEEDS299one-sided 95 percent bound, independent cases, no failure observed
100 cases
fewer than 1 in 100
100 clean cases cannot support a claim of fewer than 1 in 100. A failure rate as common as 1 in 34 is still consistent with seeing no failure in 100 cases. Supporting the claim takes about 299 clean cases, 3 times the run.
100 clean cases cannot support a claim of fewer than 1 in 100.A failure rate as common as 1 in 34 is still consistent with seeing no failure in 100 cases. Supporting the claim takes about 299 clean cases, 3 times the run.
a clean run rules rates out. it does not prove oneLIVE

Release into conditions the evidence actually covers

A release decision authorises one version of the transaction-screening solution for one defined population and operating arrangement. The combination means the software version, the kinds of targets it may screen, the permitted sources, the human review process, and the people responsible for operation. Successful evaluation provides evidence for that combination. It does not authorise another coverage group, source class, or application action. Therefore, the team should define scope in terms that users and service owners can recognise.

Shadow operation is a controlled release mode in which the solution processes permitted transaction screens but does not make the business decision. It produces a brief alongside the existing screening process. This allows comparison with existing work before users rely on the brief. It still uses confidential data, infrastructure, and budget. The enterprise must assess these obligations before starting. Shadow operation cannot reveal every effect of live reliance because users do not yet depend on the output.

A limited live release allows a small, named group of banking analysts to use the brief for eligible targets. The initial group should represent the approved target population and have enough support for review. Analysts need the agreed evidence presentation and review procedure. The service owner needs a way to pause processing and identify affected screens. The group should also practise the return to the existing manual process. This makes recovery credible before a larger release.

I would capture the release decision in a short record. The record should link the approved transaction-screening scope to its supporting evaluation. It should identify the responsible service owner, unresolved limitations, control decisions, review procedure, and recovery procedure. The record preserves what the practice approved and why. It cannot replace a missing assessment or make unsupported use acceptable through documentation alone.

Expansion means extending the approved transaction-screening solution to a new population, higher volume, new capability, or new authority. It creates a new release decision because operating conditions change. Another coverage group may use different permissions or valuation conventions. Higher volume may exceed senior-review capacity. A new communication tool may allow consequential actions. The practice must assess each changed condition before approval.

This staged release handles mixed evidence. The solution may work well for public-company screens but poorly for private targets. The practice can keep the supported public-company scope while investigating the private-target failure. The next release decision then concerns that defined gap. A local success does not become an unsupported enterprise claim.

[ fig.3 ] release_boundary
WHAT ONE RELEASE DECISION APPROVESONE RELEASE DECISIONONE VERSIONbuild, model, sources, policyONE POPULATIONpublic-company targetsPERMITTED SOURCESthe filings it may retrieveONE REVIEW PROCESSanalyst checks, senior banker decidesNAMED OPERATORSthe service owner and support routeTHE EVIDENCE STOPS HEREPRIVATE TARGETSa population the evidence did not coverA SECOND COVERAGE GROUPdifferent permissions and conventionsA TOOL THAT WRITES TO A SYSTEMa new authorityVOLUME BEYOND SENIOR REVIEWthe review process no longer holdsEACH IS A NEW RELEASE DECISIONEVALUATION IS EVIDENCE FOR THE BOX AND NOTHING OUTSIDE ITthe public-company scope stays approved while the private-target gap is investigated
evaluation covers the box and nothing outside itFIG

Release therefore creates the operating evidence that the practice must review. The next loop asks whether the live arrangement remains useful, controlled, and supportable.

Operate released applications

Release changes the practice's responsibility. The application now creates business, control, service, and support evidence. The practice needs a regular process for acting on that evidence.

Use operational review to choose the next action

After release, the practice needs an operating review. An operating review is a decision meeting for one live engagement. It records what happened since the previous decision, checks whether the approved arrangement still works, and chooses the next commitment. This prevents the practice from expanding, transferring, or continuing a solution without current evidence.

The practice lead should establish the review before release. Participants should include the business owner, FDE lead, service owner, and relevant platform or control owners. The frequency should reflect change and consequence. A controlled release may need frequent review. A stable service can later use the receiving team's normal service cycle.

The review should record four evidence streams before interpreting them. Each stream answers a different question and supports different actions.

Evidence stream Question Example evidence Supported decisions
Business result Does the workflow improve eligible work? Total effort, completion time, capacity, quality, and cost Continue, redesign, narrow, or stop
Quality and control Does the application operate within its authority? Factual failures, access checks, state failures, and incidents Release, contain, correct, or pause
Use and review Can staff use the result effectively? Eligible use, abandonment, fallback, reviewer effort, and corrections Change output, training, capacity, or workflow
Operation and ownership Can the enterprise sustain the service? Health, support demand, recovery tests, and receiving-team capability Transfer, retain support, or defer expansion

Each stream leads to different decisions. A delayed integration may require a new delivery commitment. High abandonment may require an output or process change. An access failure means pausing the affected retrieval path and checking affected screens. Combining these findings into one score would hide important distinctions.

I would organise the review around the next commitment for the solution. A source dependency is an unavailable or unreliable filing, data feed, or document service needed for the approved screen. If it blocks release, the practice must name the owner who can resolve it and the alternative if it remains unavailable. If reviewer effort exceeds the expected benefit, investigate the cause before expanding use. If the service owner cannot diagnose failures, change the transfer plan or support commitment.

A useful review can also conclude that another investigation would not change the next commitment. The practice may already know enough to narrow the scope, choose a simpler design, or stop the proposed solution. Treating every uncertainty as a new task can prolong work without improving a decision. The practice lead should distinguish information that would change the next commitment from information that would merely make the report longer.

The decision record should name the responsible person and the condition for reconsideration. This creates continuity between reviews. Leaders do not need to reconstruct the engagement each time. The record also shows which evidence changed an earlier judgement.

The progress report now has a defined role. It summarises the stage, evidence streams, decisions, owners, and actions due. Consequently, leaders can assess the engagement without treating task counts as proof of value.

Use incidents to improve FDE ownership and shared platform controls

Incidents in a released solution belong first to its service owner. The FDE practice should remain involved when the incident exposes unclear ownership, an unmet transfer condition, or a shared platform weakness. The practice should preserve those lessons in the engagement and portfolio records. It should not replace the enterprise incident process.

Suppose a filing contains text that the application interprets as instructions to retrieve unrelated deal information. The first responsibility concerns the affected capability and possible exposure. Operators should contain the behaviour and follow the enterprise's incident process. Prompt changes can wait until the team understands what occurred.

The investigation should follow the complete sequence. Which document did the application read? Which identity performed the retrieval? What information did the tool expose, and which authorisation checks applied? Those details distinguish an incorrect model suggestion from a system that allowed the suggestion to become an unauthorised action.

The team should also establish the likely population of affected cases. The first report does not necessarily mark the first occurrence. Relevant application versions, data sources, tool configurations, and logs can help identify where the same conditions existed. The investigation can then guide both immediate remediation and the evidence required before restoring the capability.

Corrective work should address the mechanism that allowed the failure. The tool may offer excessive access, or the system may rely on model judgement to enforce a permission rule. Untrusted document content may influence instructions that control privileged actions. A prompt change can contribute to the response. However, access enforcement and application design must carry their assigned responsibilities.

After correction, the team should reproduce the failure under controlled conditions and test relevant variations. That assessment needs to examine the actual enforcement mechanism, not merely whether the model now produces a reassuring answer. Remaining uncertainty should influence the restored scope. One successful regression case does not prove that the team eliminated every related failure.

The resulting actions may belong to several teams. Application engineers can revise tool use, while platform or identity teams may need to change shared enforcement. Each correction needs an owner and a maintenance commitment. This process allows an incident to improve later deployments. It also preserves the lesson after the original team leaves.

Track adoption only when it informs an FDE decision

Adoption tracking asks whether eligible users use the released solution for eligible transaction screens, and whether that use replaces or adds to existing work. The practice should track eligible use, completion, abandonment, fallback, and reviewer effort. A login count cannot show business value.

The practice needs an explicit denominator. Exclude users who had no eligible screen during the observation period. Include eligible screens that users rejected or returned to the manual process. This keeps adoption connected to the approved population and exposes whether the solution fits the real process.

Adoption is useful to the FDE practice when it explains a decision. Low use may indicate missing permissions, weak evidence, poor timing, or an unsuitable handoff. High use may still add effort or bypass a control. Therefore, the operating review should connect adoption to quality, result, and support demand before changing the solution or expanding its scope.

Account for the review work that automation creates

The adoption review may reveal that the application moves work instead of removing it. Faster brief preparation can increase the rate reaching senior bankers. If review capacity stays fixed, the queue may grow. The preparation task becomes faster while the complete screening workflow gains little.

[ fig.4 ] review_queue
THE SAME SCREENING WORKFLOW, BEFORE AND AFTER THE APPLICATIONBEFOREPREPAREthe analyst preparesSENIORREVIEWcapacity unchangeddecided: the sameAFTERPREPAREthe application prepares, the analyst checksWAITING FOR SENIOR REVIEWSENIORREVIEWcapacity unchangeddecided: the sameFASTER PREPARATION MOVED THE WAIT. THE DECISION ARRIVES NO SOONER.illustrative. completion time and net effort are the measures that show it
faster preparation moved the wait, not the decisionFIG

Human review therefore needs an operating capacity model. The team should observe how much effort reviewers spend checking evidence, resolving uncertainty, correcting output, and recording a decision. Difficult cases and periods of higher demand need attention because they can expose weaknesses that a quiet demonstration conceals. Reviewers also need the skills and information required to recognise the application's likely failure modes.

The first response may concern the output itself. A useful screening brief links each important claim to its source. It distinguishes facts from interpretation and exposes missing information. This structure shows which judgement remains with the analyst or banker. More generated text can hide important omissions.

Review intensity may vary when the evidence and applicable policy support that variation. Some cases may require full assessment, while others may permit narrower checks or sampling. The relevant authority needs to approve that arrangement for the actual use. A model's stated confidence does not, by itself, establish the reliability required to reduce human review.

If review remains the limiting step, the enterprise has four operating choices. It can improve evidence, add qualified capacity, narrow eligible work, or retain manual processing. The cause and expected benefit determine the choice. Expansion without addressing the constraint sends more work into the same queue.

The measurement plan should capture that consequence. Net effort includes review and correction, while completion time includes relevant waiting. Keeping these measures together distinguishes a useful local improvement from relocated effort or delay. That distinction becomes essential when the practice later describes the application's value.

Reassess the application when its conditions change

The release evidence describes an application under particular conditions. Production creates continuing obligations because those conditions do not remain fixed. A model update can change behaviour. A source system can alter field meaning. A business policy can invalidate a correct interpretation. The operating team needs a practical route from those changes to an appropriate reassessment.

I would maintain an inventory of consequential dependencies and the people responsible for them. A proposed model change should connect to relevant evaluation cases, operating requirements, and release decisions. The comparison should examine quality, control behaviour, latency, cost, and reviewer effort where those properties matter. Improvement in one area can accompany regression in another.

Other changes alter the scope of the application itself. Adding a tool that can write to a business system changes authority. Supporting another region can change data and process requirements. A new language can expose gaps in evaluation coverage. The service owner should recognise these as changes to the operating arrangement and assess the differences that matter.

The manual route also needs maintenance. Staff can lose familiarity with an automated procedure. Dependencies can also make an old fallback unavailable. Recovery exercises should follow the service's importance and observed changes. An existing runbook provides a starting point, but only current evidence can show whether the intended recovery remains practical.

These continuing obligations explain why ownership and evidence received attention during setup. They are part of operating the application, rather than exceptional tasks that appear only after something goes wrong. Early funding and assignment help the receiving team sustain the result after the engagement ends.

Operation produces evidence that can support transfer and shared learning. The next loop separates those responsibilities so neither remains implicit.

Transfer responsibility and useful learning

An engagement should leave two durable results. One team must operate the application. Later teams should also receive any supported learning from the work.

Set and verify the operating handoff

The receiving team is the team that will run the released transaction-screening solution after the FDE engagement. It may be an investment-banking technology team, a product team, or a shared service team. The practice should name it during setup and test its readiness before ending the engagement.

Transfer should happen after the solution operates within its approved scope and before the FDE team closes its delivery commitment. The receiving team should deploy a routine change, run the evaluation, inspect an unsuccessful screen, and explain the solution's limitations. It should also exercise recovery and manual fallback. These tests show whether the team can operate the service.

The business workflow owner must understand the same arrangement. They should know which targets the solution supports, what the manual route involves, and who can approve an expansion of that scope. The service owner remains accountable for deployment, support, dependencies, and technical changes. Technical ownership cannot settle an unassigned business decision.

The transfer record is a short acceptance record. It names the service owner, business workflow owner, support route, access required, approved scope, known limitations, recovery procedure, and evidence from the readiness exercises. When an exercise reveals a gap, the record identifies the remaining work, its owner, and its consequence for transfer.

Some responsibilities may remain with the FDE practice or another specialist service. A shared evaluation service may support several receiving teams through central maintenance. Continuing specialist involvement needs a defined service, capacity, and funding commitment. Otherwise, informal requests to the original engineers can hide the fact that responsibility never moved.

The handoff is complete when the receiving team can perform the agreed exercises and accept the defined responsibilities. If it cannot, leaders should change the design, fund capability, extend specialist support, or delay transfer. A signed record cannot substitute for demonstrated operating capability.

Turn a local result into useful learning for later teams

An engagement produces many observations, but those observations do not all justify shared investment. Here, “later teams” means the shared platform team or another FDE team working on a related financial process. The practice needs to distinguish what worked locally from what those teams can safely reuse. That distinction allows successful delivery to improve the platform without turning every local decision into an enterprise standard.

The screening application might produce a retrieval component, a valuation definition, and an evaluation collection. The retrieval component may support several workflows. The valuation definition may apply only to one coverage group. The evaluation structure may generalise while its source material remains restricted. Each output needs an owner and stated boundaries.

I would bring a proposed shared improvement to the shared platform product owner with a short evidence note. The note should show the repeated problem, affected screens, current alternatives, likely users, operating cost, and known limits. It should include a failing example, the decision it blocked, and evidence from a second case. The product owner can then select a new component, a product change, or better documentation. Acceptance should include a maintenance and support owner.

A second adopter provides a practical test of the shared proposal. A manufacturing supplier-certification workflow may test shared source-ownership and effective-date metadata. Procurement owns the meaning of certificate validity. The platform team can observe which metadata transfers and which rules remain local. Adaptation effort and support demand provide better reuse evidence than an import count.

Negative findings also deserve a maintained record when they affect future decisions. An application may prove unsuitable because available evidence cannot support its intended judgement. Review effort may also exceed the benefit. The finding should explain those conditions and the evidence that would justify reconsideration. It should not turn one local result into an unsupported claim about every possible use of the technology.

This gives the practice a broader account of progress. Shared software can remove repeated implementation. Clear source definitions can remove repeated investigation. Evaluation methods can reduce repeated disagreement. The useful test concerns what later teams can do with those results. Calling an artefact reusable does not establish that anyone benefits from maintaining it.

Evaluation evidence deserves particular care because its meaning can change over time. A new policy can invalidate an old label. A new transaction type can expose missing coverage. Repeated development against the same examples can also weaken their independence. The evaluation owner must update the collection while preserving the history behind earlier results.

Shared controls have similar boundaries. A retrieval pattern may rely on a particular identity model, processing location, or data classification. Another application needs to check those conditions before treating the pattern as suitable. An earlier approval supplies relevant evidence for the new assessment; it does not automatically authorise a different use.

These obligations should appear in the shared investment decision. The assets need a maintenance owner. Otherwise, later teams may rely on evidence that no longer describes their applications. The practice should therefore make the continuing cost of shared learning visible alongside the delivery effort it might remove.

Transfer and shared learning create continuing commitments. The practice lead must compare them with new requests and other uses of enterprise capacity.

Manage the portfolio and its value

The practice lead must connect engagement decisions to finite enterprise capacity. This requires a portfolio view, specific measures, and consistent value definitions.

Manage the portfolio as a set of continuing commitments

The practice lead must connect each engagement to the capacity of the whole team. New requests compete with delivery, incident support, transfer work, shared improvements, and professional development. Accepting another solution therefore commits more than build time. The commitment can continue after launch when ownership or shared maintenance remains unresolved.

[ fig.5 ] portfolio_tails
ONE TEAM, THREE ENGAGEMENTS. ILLUSTRATIVE.TIMESCREENINGBUILDSUPPORT · REASSESSMENTTRANSFERREDthe tail moved tooTRADING SURVEILLANCEBUILDSUPPORT · REASSESSMENTSUPPLIER CERTIFICATIONBUILDSUPPORT · REASSESSMENTNOWA NEW REQUESTBUILD?arrives into capacity two open tails already holdA COMMITMENT IS LONGER THAN ITS BUILDonly transfer ends the tail. an unresolved owner keeps it on the practice
a commitment is longer than its buildFIG

I would maintain a portfolio view for every engagement. The view should identify the stage, next decision, accountable owner, unresolved dependency, committed capacity, and expected support after release. Active engineering and waiting should remain distinct. Waiting may release development time, but it still requires coordination and context recovery. Blocked projects are not costless. Treating them that way creates excess demand when several dependencies resolve together.

Portfolio field What the practice lead records
Engagement and owner The solution, business workflow owner, FDE lead, and service owner
Current stage Requirements, feasibility, limited workflow, controlled release, transfer, or operation
Next decision The choice needed before more capacity or scope is committed
Evidence status Established findings, unresolved uncertainty, and linked records
Dependencies Data, platform, control, domain, or receiving-team work that can block progress
Capacity Engineering, domain review, platform, control, and support commitments
Future obligation Operation, maintenance, evaluation, support, and reassessment after release

The appropriate limit on concurrent work should follow from observed capacity and the portfolio's demands. Engineers need time for review, documentation, support, and transfer as well as implementation. An allocation based only on visible development tasks will understate those commitments. The practice should measure where time goes before turning a convenient staffing ratio into a permanent rule.

A blocked engagement also needs an explicit decision. Progress may require unavailable capacity from a data or identity team. Leaders can fund that capacity, narrow the proposal, or pause. The record should identify the dependency and the condition for restarting work. Keeping the project nominally active does not make the unresolved commitment easier to manage.

Portfolio review should compare the next increment of spending with current alternatives. A nearly completed application can still deserve cancellation if its remaining costs and obligations exceed its likely benefit. A difficult investigation can deserve continuation if its result will resolve an important uncertainty. Past spending explains how the project reached its current state; the next decision depends on what additional work can achieve.

Shared engineering needs the same scrutiny. A proposed platform improvement should identify the future work it can remove and the teams likely to use it. An urgent local requirement should identify the consequence of delay. The practice lead and product owners can then compare concrete choices. A reusable component does not always deserve priority over a local result.

The portfolio may also reveal a different organisational need. Repeated work within one domain can justify a permanent product team. Repeated delays around one integration can justify platform investment. A mature workflow family may need enablement more than intensive delivery. Recognising these patterns helps the enterprise allocate responsibility. It also prevents the practice from pursuing project count as its goal.

Define measures as evidence for specific questions

Measures are recorded observations that help the practice answer a management question. Delivery speed, business value, control performance, ownership, and shared learning describe different responsibilities. A single score can hide a weak result in one area behind a strong result in another. I would therefore define each measure with its question and interpretation limit.

Question Evidence that helps Why interpretation needs care
Does work progress through delivery? Time by stage, waiting, rework, and active capacity Faster delivery can reflect easier case selection
Do applications remain useful? Continued use with quality, outcome, and support evidence Continued existence does not establish value
Can receiving teams sustain operation? Demonstrated capability and subsequent support demand A signed handover can conceal dependence
Does learning improve later work? Adaptation effort and maintained improvements used by later teams Reuse counts omit suitability and maintenance
Does specialist assistance change? Assistance required for comparable workflow families Aggregate demand changes as the portfolio changes
Does the practice justify further investment? Outcomes, costs, continuing obligations, and credible alternatives Attribution and observation periods limit conclusions

Comparable observation periods matter when reviewing application survival. A recent release has had less time to develop operating problems or produce benefits than an older service. Reporting both as equivalent successes can create a misleading picture. Planned retirement also differs from abandonment. An application can produce value and later become unnecessary after a business or technology change.

The share of deployments requiring assistance needs similar care. That share may rise because the enterprise begins harder workflows while established workflows become easier. It may fall because the FDE practice lacks capacity or because teams stop recording informal help. Neither movement independently establishes whether the platform or practice improved.

A more useful investigation compares similar workflow families and examines why assistance remains necessary. Repeated help with a resolved integration suggests one problem. Support for a new permission model or business requirement suggests another. Combining that information with effort, quality, and actual use makes the measure more informative than the aggregate percentage alone.

The strength of the conclusion should match the evidence. A component adoption count can demonstrate distribution. A claim that the component reduces delivery effort needs observations about integration and support against a suitable comparison. A reliability claim needs failures and observation periods. Maintaining those distinctions allows the practice to report progress without turning every activity measure into a causal claim.

Explain value for the complete screening process

The business value discussion should begin with what changed in the complete screening process. A solution can reduce analyst effort, improve quality, shorten waiting, or enable more target reviews. These outcomes can all matter, but they require different evidence. They do not all create a corresponding reduction in expenditure.

Published research illustrates why a general productivity assumption is insufficient. A study of AI assistance in customer support reported improvements that varied across workers. METR's 2025 developer experiment found slower completion in its studied setting. Its 2026 update described selection problems in a later experiment. Brynjolfsson, Li, and Raymond, METR's original study, METR's update. These studies concern different tasks and tools; none estimates the value of this proposed FDE practice.

For transaction screening, I would begin with the human effort required to complete comparable screens at the required quality. Preparation, review, correction, coordination, and fallback belong within the affected workflow. The relevant difference concerns total effort within that boundary. Faster preparation may produce little benefit if senior review adds equivalent work.

For a comparable case category, the relationship is straightforward:

Net human effort change per case = comparison effort per case minus observed effort per case.

The difficult part concerns measurement and comparability. The team should explain how it records effort, how cases vary, and which population the result represents. Category results can then combine using an explicit case mix. Otherwise, a shift towards simpler submissions may look like an application improvement even when performance within each category remains unchanged.

The calculation should also respect what the underlying measure already includes. If the comparison covers all eligible cases, it already reflects non-use and manual fallback. Applying another adoption discount would count that effect twice. Likewise, review and correction should not receive a second subtraction when the effort measure already contains them. Explicit boundaries make the calculation more defensible than a chain of arbitrary discounts.

A measured reduction in effort initially establishes released capacity. The enterprise can use that capacity to clear a backlog, improve service, or complete additional work with existing staff. Those uses can create value without changing payroll. Assigning a salary-based value to the time describes an economic estimate, rather than an observed cash saving.

Cash savings require evidence of reduced expenditure. A cancelled external service or a lower contractor commitment can provide that evidence, subject to timing and offsetting costs. Avoided cost makes a different claim: the enterprise would otherwise incur an expense. That claim needs a credible plan or demand forecast. An expense cannot become a saving when nobody intended to incur it.

Revenue and quality benefits require their own causal explanation. Faster screening may support earlier transaction work. However, many other conditions determine whether a transaction proceeds. Better evidence preparation may reduce review errors, but the team must define and observe those errors. Estimates of avoided losses should remain separate from observed outcomes.

Finance involvement helps establish these distinctions before delivery. Agreeing definitions and measurement boundaries early gives the team a clearer way to interpret both positive and disappointing results. Late finance approval creates avoidable disagreement. The team should establish the comparison before delivery.

The practice's contribution also needs a boundary. An application may benefit from the underlying model, platform capabilities, process changes, vendor expertise, and FDE work together. The full outcome cannot automatically belong to the FDE practice. Assessing the practice requires a credible account of how its delivery arrangement compares with alternatives for similar work.

Match costs and benefits to the same decision

A useful value account includes the resources required to produce and sustain the measured result. Discovery, engineering, domain review, integration, evaluation, infrastructure, model usage, change management, and support all contribute. Shared services also have costs even when another department pays them. The enterprise needs an explicit allocation method that remains consistent across comparisons.

The time boundary matters as much as the cost categories. Comparing a mature service's recurring benefit with a short pilot's partial cost can overstate the case. Assigning all platform costs to one engagement can also distort the result. This is especially true when benefits span the enterprise. Costs and benefits should concern the same relevant period, population, and responsibility.

Historical expenditure and future commitments serve different decisions. Historical costs describe what the enterprise spent to obtain the current result. Future costs help determine whether to continue, expand, replace, or retire the service. A large historical investment does not justify further spending. Conversely, an expensive investigation can still support an inexpensive operating service.

Forecasts should identify the assumptions that connect observed results to future value. Eligible volume, case mix, measured effect, adoption, operating cost, and maintenance needs can all change the outcome. The team should use measured inputs where available and show how consequential uncertainties affect the result. A forecast becomes more useful when leaders can see which assumption would change the decision.

This approach also helps explain an engagement with mixed results. The application may release capacity while incurring higher model costs, or improve review quality without shortening completion time. Reporting those effects separately allows the business to assess the actual tradeoff. Compressing them prematurely into a single return figure can obscure the reason the enterprise might still choose to continue.

Value evidence informs the practice decision. The final loop asks whether the current FDE arrangement still fits the recurring work and obligations.

How should the practice change when the work changes?

The practice itself needs periodic review. Its current shape should follow recurring work, observed value, dependencies, and credible alternatives.

Expand, reorganise, or close the practice

Successful delivery can create pressure to hire more FDEs, but additional engineers help only when engineering capacity limits worthwhile work. The current constraint may concern domain review, data access, release decisions, or receiving teams. Adding delivery capacity can increase demand on those already constrained groups. The expansion decision should therefore begin with evidence about where work waits and why.

[ fig.6 ] where_work_waits
WHERE EACH ENGAGEMENT IS WAITING. ILLUSTRATIVE.FDEENGINEERINGDATAACCESSDOMAINREVIEWRELEASEDECISIONRECEIVINGTEAMNOTHING WAITS HERESCREENING, SECOND GROUPTRADING SURVEILLANCESCREENING, PRIVATE TARGETSSUPPLIER CERTIFICATIONADD MORE ENGINEERSwidens the empty columnand sends more work into the four that are fullTHE REMEDY BELONGS WHERE THE WORK WAITSengagements waiting on the same integration justify that investment before another team
the remedy belongs where the work waitsFIG

If engagements repeatedly wait for the same identity integration, a shared investment may improve delivery more than another FDE team. If reviewers cannot assess the output volume, more prototypes will not resolve the operating problem. Examining waiting, rework, support demand, and dependency ownership helps leaders identify which capacity the enterprise actually needs.

Expansion must also preserve the conditions that made the initial arrangement work. Founding engineers may rely on relationships or informal access to experts that new staff will not automatically inherit. The practice should make those commitments explicit before assuming another team can repeat the result. Technical review, domain participation, and funded receiving capacity need to expand with delivery demand.

Different workflow families may eventually need different arrangements. A mature family can use standard components and occasional advice, while a new family requires intensive collaboration. A strategically important application may deserve a permanent product team. The practice can support those transitions while reserving specialist deployment capacity for work that still requires it.

Professional development belongs within that capacity decision. Engineers need time to compare findings, review implementations, maintain shared methods, and develop domain judgement. Continuous assignment to urgent delivery can leave little room to improve subsequent work. The practice should recognise this commitment and assess whether it produces useful changes in later engagements.

Periodic renewal should return to the original organisational problem. Does the enterprise still need a distinct arrangement for this class of work? Does the practice provide an acceptable combination of outcomes, cost, quality, and continuing ownership compared with available alternatives? Those questions can justify expansion, a revised mandate, a transfer of responsibilities, or closure.

No single metric can settle that decision. Reduced assistance may reflect stronger capability or hidden support. Greater reuse may accompany substantial maintenance. Continued service operation may coexist with limited value. Leaders need the underlying evidence and an account of the obligations that would move if the practice changed form.

Ending the practice also requires an operating plan. Supported applications, shared assets, access responsibilities, evaluation collections, and outstanding maintenance still need owners. Closing the team does not remove the work its engagements created. A responsible transition assigns that work and the capacity required to perform it.

Bring the decisions together in a practice review

The transaction-screening example can now show what the operating system produces. Suppose the brief reduces analyst preparation for public-company targets. However, private-target evidence remains unreliable and senior-review queues are growing. The banking technology team can deploy routine changes but still needs help with retrieval failures. Another coverage group also requests access under different permissions and valuation conventions.

First, the unsupported private-target population requires a scope decision. The team may narrow processing, correct the behaviour, or pause that capability. Next, leaders must address senior-review capacity before adding intake. The transfer plan must cover retrieval diagnosis and any continuing specialist support.

Subsequently, the second coverage group requires an assessment of its differences before rollout. This work can also test whether a retrieval improvement serves more than one application. Finance partners can evaluate the observed result within the agreed measurement boundary. The practice lead can then compare further investment with other portfolio commitments.

The review therefore connects technical behaviour, business use, operational capacity, ownership, and shared learning. Each finding has a consequence for the next action, and each action has someone responsible for it. The practice must understand the whole workflow and make the next decision well. It must then make that decision real.

The operating model is complete when every finding reaches a decision, an owner, and a funded next action. That is how the practice learns without turning activity into proof of value.

Part one: what forward-deployed engineering does. Part two: setting up an internal practice.