Setting Up an Internal Forward-Deployed Engineering Practice
How I would establish an internal AI FDE practice that can finish valuable workflows and leave them with an owner.
On this page
Part one showed why a shared AI platform still needs engineers who can complete business processes. The platform provides common capabilities such as model access, retrieval, and tool execution. FDEs combine those capabilities with local data, rules, decisions, and operating procedures. A proposed solution can then serve a real process while the platform receives evidence about recurring technical needs.
Transaction screening makes the organisational problem concrete. Deal analysts, data owners, control owners, and service teams each hold part of the work. An FDE can build an integration, but the FDE cannot grant access to a confidential deal workspace, approve a screening method, or create operating capacity in another department.
The setup must connect engineering responsibility to the authority and time held by those other roles. It must also make the business result, its supporting evidence, and its future owner explicit. This article develops that setup through the transaction-screening case, with manufacturing used only when it reveals a different constraint. The cases are illustrative, not measured delivery results.
The setup has five connected decisions. First, diagnose the delivery problem and define the proposed solution. Next, assign decision rights and choose an organisational home. Then fund and staff the work. After that, establish data, evaluation, and operation. Finally, authorise a bounded engagement and review what the practice learned.
| Setup decision | Question it answers |
|---|---|
| Diagnose and define | What problem and proposed solution deserve attention? |
| Assign and place | Who can decide, and where should the practice sit? |
| Fund and staff | What capacity and skills can support the commitment? |
| Enable and operate | Which data, controls, evaluation, and ownership make use possible? |
| Authorise and learn | What bounded engagement tests the practice and informs the next commitment? |
The sections follow these decisions in order. Later decisions can return to an earlier one when evidence shows that the proposed arrangement cannot work.
This article uses four role names. The practice lead manages the FDE portfolio and capacity. The FDE lead owns delivery evidence for one engagement. The business workflow owner owns the process and the meaning of a useful result. The service owner accepts deployment, support, dependencies, and technical change. One person may hold more than one role, but each responsibility needs a named owner.
An engagement is a bounded commitment to investigate, develop, and assess one proposed solution. It has a defined population, budget, decision owner, and next commitment. The FDE practice closes or renews the commitment when evidence supports that choice.
Diagnose and define the delivery problem
Start with a delivery problem, not a team design
An internal AI FDE practice should begin with a repeated delivery problem. A large enterprise may already have a central AI platform, several product teams, data specialists, and business functions with valuable ideas. The existence of those resources does not show that a new practice is needed. The question is whether important workflows repeatedly fail to move from a promising capability to an operating result.
Consider an investment banking team that wants to reduce the work involved in first-pass transaction screening. Deal analysts collect public filings, approved internal research, market data, and early diligence material. They compare a target with selected peers, investigate gaps, and prepare a screening brief for a senior banker deciding whether deeper work is justified. An AI platform can extract figures and draft explanations. The proposed application still needs the right sources, the correct reporting period, consistent entity and currency definitions, evidence for each conclusion, and a controlled route to review.
The visible symptom might be that the project has spent months in pilot status. That symptom does not identify the cause. The banking team may lack an engineer who can build the connection to research, market-data, and deal systems. The data team may not have approved the proposed use of confidential transaction information. The valuation or risk owner may not have time to define how exceptions should be handled. The existing product team may have built a useful extraction service but no one has responsibility for the review process around it.
The same symptom can appear in manufacturing for different reasons. A supplier-certification application may extract fields accurately, but procurement reviewers may still need to establish the supplier's legal entity, certificate validity, and applicable product category. The delay could come from document access, unclear procurement rules, or a review queue that cannot absorb more prepared cases. Treating every delay as a shortage of engineers would direct investment at the wrong constraint.
I would define the existing delivery arrangement before proposing a new team. By delivery arrangement, I mean the path a request follows from a business problem through technical design, data and control decisions, release approval, and continuing service ownership. For the banking example, I would trace who requested the screening application, which team built the pilot, who supplied research and market-data access, who defined a valid comparison, who could approve production use, and who would respond when a source or review rule changed.
Tracing several requests produces a useful pattern. One request may stop because the platform lacks permission-aware retrieval. Another may stop because no team can spend time with deal analysts to model the exceptions. A third may reach production but create so much review work that analysts return to the old process. Those are three different problems. The first suggests platform investment, the second suggests a delivery-capacity gap, and the third suggests that the proposed workflow does not remove enough effort.
The review should include applications that succeeded. An investment-banking technology team may already own the process, data access, and service operation for a reporting workflow. Extending that team may be better than introducing an FDE handoff. A standard document application may need only configuration. An internal FDE practice is justified when the same unresolved connection appears across valuable workflows and existing teams cannot own the whole result.
This diagnosis also protects the practice from becoming a general AI help desk. Executive demonstrations, training sessions, routine platform support, and long-term maintenance can all create legitimate demand. They should have clear routes and owners. The FDE mandate should remain focused on workflows that require engineers to work across business and technology boundaries, complete a useful result, and return durable learning to the platform or receiving teams.
That diagnosis gives the practice a bounded reason to exist. The next question is who can make the decisions that delivery needs.
Define the business process and proposed solution
Once the enterprise understands the recurring gap, it needs a concrete unit of work. I would define that unit as a business process and a proposed solution, rather than as an abstract capability. The business process has a starting event, actions, decisions, exceptions, and users. The proposed solution may be an agent, retrieval system, conventional application, process change, or another design. This distinction turns “improve AI adoption in investment banking” into something that engineers and business owners can inspect together.
The transaction-screening process can be described precisely. A deal team submits a potential acquisition target for an initial review. The proposed solution retrieves permitted public filings, internal research, market data, and early diligence documents. It extracts financial measures, compares the target with selected peers, identifies gaps or conflicts, and presents source evidence for each finding. A banking analyst reviews the brief and either confirms the facts, requests clarification, or routes a question to a valuation or risk specialist. A senior banker decides whether deeper work is justified.
Each step gives the FDE team a question to answer. Which source is authoritative for a financial measure? Which identity may retrieve a document from the deal workspace? How does the application distinguish the relevant reporting period from an older figure? Which conflicts can an analyst resolve, and which require escalation? What evidence must remain attached to the screening decision? The process is now specific enough to reveal missing data, methodology decisions, integration work, and evaluation needs.
The initial engagement should cover a manageable population without removing the difficulty that makes the problem worth solving. I might select potential acquisition targets in one sector at the initial screening stage. The application would prepare a source-linked screening brief for those cases, but it would not issue a trading instruction, approve a valuation, or communicate with a client. The banking owner would retain the decision about further work, while the FDE team would own the application and the engineering required to make the evidence usable.
The boundary should include the next handoff because that is where the business result becomes visible. If analysts must reopen every filing and reconstruct each comparison, the application has produced another document rather than reducing screening effort. Including analyst review allows the team to examine evidence quality, review time, clarification requests, and manual fallback. A technical demonstration can stop at a generated explanation. A workflow engagement must continue until the next person can use the result.
The manufacturing comparison follows the same logic. A first engagement might cover supplier certificates for one product category in two plants. It would identify the required certificate types, retrieve permitted documents, show missing or expired evidence, and route exceptions to procurement. It would not decide whether a supplier is approved. The scope is useful because it contains real variation while keeping the population small enough to understand.
I would write the initial workflow brief in terms that each participant can recognise. For the transaction-screening engagement, a brief might look like this:
- Problem: Deal analysts spend substantial time collecting filings, research, and market data before deciding whether a potential transaction merits deeper work.
- Supported cases: Potential acquisition targets in one sector at the initial screening stage, using approved public filings, internal research, market data, and early diligence documents.
- Proposed result: A source-linked screening brief that extracts financial measures, compares selected peers, flags conflicts, and lists missing evidence.
- Out of scope: Investment recommendation, valuation approval, trading instruction, client communication, and use of non-public data outside the approved deal team.
- Main uncertainty: Whether the available sources contain consistent entity and period definitions for a brief that bankers can review faster without hiding exceptions.
- Evidence for the next decision: Comparable manual effort, factual correction rate, source-access behaviour, reviewer acceptance, and the receiving banking technology team's ability to operate the application.
This brief does several jobs at once. It defines the population, states the result, protects the decision that must remain with the deal team, and identifies the uncertainty that the first investigation should resolve. It also prevents the application from quietly expanding into valuation approval or client communication because those features seem technically possible.
Make dependencies visible so somebody can resolve them
The engagement brief makes dependencies visible, but visibility has a practical purpose. A dependency is a condition held by another person or team that must be satisfied before the proposed solution can operate. In the transaction-screening example, dependencies include access to approved sources, the definition of a valid comparison, analyst review time, release approval, and a team that can maintain the service. Naming them tells the practice where progress can stop and who can change the condition.
I would record each dependency with its decision owner, required contribution, and consequence if it remains unresolved. If the data owner cannot permit retrieval from a confidential deal workspace, the engineering team should not promise to build a production application that requires it. If the valuation or risk owner cannot clarify how a comparison should be treated, the team should test whether the application can expose the ambiguity and route it for judgement. If the receiving team cannot provide support capacity, the scope or operating model must change before release.
This is why a workflow brief cannot remain a document that only the FDE team reads. The banking owner, source owner, security or control owner, platform product owner, release authority, and service owner need to agree that the named conditions are real. Their agreement does not make every dependency available. It makes the unresolved condition explicit, so leaders can fund it, narrow the workflow, choose another design, or stop.
Assign and place the practice
Assign each decision to the authority that can make it
The next step is to turn visible dependencies into explicit decisions. Responsibility for delivering the screening application does not give the FDE team authority over every dependency. The team needs a route to the people who hold that authority, and those people need a clear account of the decision they are being asked to make.
| Decision | Accountable role | Required contribution |
|---|---|---|
| Define the screening result and supported cases | Banking workflow owner | Process boundary, acceptance criteria, and authority to change the screening procedure |
| Decide whether source data may be used | Data and control owner | Purpose, identities, fields, retention, and approved processing conditions |
| Define what counts as a valid comparison | Valuation or risk owner | Rules for entity, period, currency, source authority, and exception handling |
| Build and maintain the application | FDE engineering lead and later service owner | Technical design, integration, testing, deployment, support, and change management |
| Provide shared AI primitives | AI platform product owner | Model access, retrieval, tool controls, evaluation support, and maintenance commitments |
| Approve production release | Designated release authority | Review of technical, business, control, and operational evidence |
| Operate the service after engagement | Technical service owner and banking workflow owner | Staff, access, incident response, source updates, and continuing funding |
These are decision responsibilities, not a required list of separate hires. One person may hold several roles in a small business unit. A large enterprise may require independent control assessment for certain data or release decisions. The FDE practice should fit those existing authorities rather than create an informal approval process that competes with them.
The banking example shows why this distinction matters. Suppose analysts want the application to read detailed deal-room documents, but the data owner permits only selected files for the approved purpose. The FDE team can explain which screening gaps the restriction creates and propose alternatives, such as retrieving a narrow document set under the requesting analyst's identity. The data owner must decide whether the use is permitted. Escalation can resolve the tradeoff, but it does not transfer data authority to the engineers.
Some decisions need an immediate route rather than a scheduled meeting. If an operator suspects that the application retrieved data for the wrong target or deal team, someone must be able to pause processing and identify affected cases. The incident path should name that person before production release. The FDE practice can implement containment, but the enterprise must decide who has authority to stop the workflow and who assesses the business consequence.
With decision rights mapped, the enterprise can choose where the practice should sit. Its organisational home must protect capacity and preserve routes to business and platform decisions.
Put the practice where it can get decisions made
Once the decision owners are known, the enterprise can choose where the internal AI FDE practice should sit. This is a reporting and operating choice, not a matter of finding the most fashionable organisation chart. The practice needs a home that can protect engineering capacity, maintain a route to business decisions, and return recurring technical requirements to the shared platform.
A platform-aligned practice reports into the group that owns the AI platform. That home supports common engineering standards, access to platform specialists, and a direct route for turning deployment evidence into new primitives. It works well when the recurring problem is that several business units need similar integrations and no single product team can absorb the work. Its risk is distance from the banking or manufacturing process. A platform leader must then protect time for domain participation and give the FDE practice authority to work with business owners.
A business-technology practice reports closer to the functions whose workflows it changes. That home can make process ownership, analyst time, and business value easier to secure. It works well when the main difficulty lies in coordinating a business process across several systems. Its risk is that local delivery becomes detached from shared platform improvement. The practice needs a formal route to the platform product owner, with enough capacity to maintain common primitives.
A federated practice combines both arrangements. Engineers may report through the platform organisation while each engagement has a named business sponsor and service owner. The arrangement can work when the enterprise needs common technical standards and strong local process ownership. It requires one leader to resolve conflicts when platform work, banking delivery, and manufacturing delivery compete for the same engineers. Calling the practice federated does not resolve that conflict by itself.
The specific choice should follow the recurring dependency pattern. If investment banking, manufacturing, and other functions repeatedly need the same retrieval, identity, or evaluation improvements, place the practice close to the platform and establish protected business participation. If the shared primitives are mature but workflows repeatedly stall on business process change, place the practice closer to business technology and establish a funded route to platform engineering. If both conditions are present, use a federated model with one accountable practice lead and explicit capacity commitments on both sides.
Team Topologies offers useful language for describing how those relationships change over time. A close collaboration may be necessary while the FDE team learns the screening process. The team may then consume a stable retrieval service while focusing on transaction logic. Before transfer, the FDE team may facilitate the receiving service team's ability to operate the application. These are different modes of interaction, and naming the change helps leaders reduce intensive collaboration when the work becomes established. Team Topologies.
Decide what belongs in the platform and what remains local
The transaction-screening workflow will quickly reveal requests that could be implemented locally or proposed as shared platform improvements. The practice needs a way to distinguish them. A primitive should move into the platform when it has a stable meaning across applications, a clear interface, several plausible consumers, and an owner willing to maintain it. A rule should remain local when its meaning depends on one business policy, one target, or one workflow decision.
For example, a deal team may need the banking definition of a comparable peer or a permitted screening source. That definition belongs with the valuation or risk owner. Several applications may also need retrieval results to include source ownership, effective dates, and provenance. Those are candidates for a shared retrieval capability because they describe information about the source rather than the meaning of a particular transaction decision.
Manufacturing can provide a second test. Procurement may need a rule that a certificate must cover the legal entity supplying a particular product category. The rule is local to that procurement process. The ability to retrieve a certificate's issuing entity, issue date, and expiry date may serve other workflows. A proposed shared primitive should therefore expose those facts and leave the application to interpret them according to its own policy.
The distinction should be tested with a second adopter, rather than decided from the first application's code alone. If a retrieval change helps both transaction screening and supplier certification without adding contradictory assumptions, the platform team has evidence for shared investment. If adapting the change requires embedding banking-specific methodology or manufacturing-specific document rules, the local applications should retain those parts. A component can be technically reusable while still imposing too much semantic coupling to be a good platform feature.
The platform product owner must accept more than the code. Acceptance means agreeing to own defects, compatibility changes, documentation, security review, usage support, and future requests. Without that commitment, the FDE practice may transfer a repository while retaining the real responsibility. A shared component becomes a platform capability only when its maintenance and support have an owner and funding.
Useful learning can also remain outside the platform. A banking-specific interpretation of a valuation field, a maintained evaluation collection, or a documented negative result may help a later team without becoming a central service. The practice should record the scope, evidence, and owner of that learning. Later teams can then decide whether it applies to their workflow instead of treating an earlier application as an automatic precedent.
The platform-versus-local distinction now has a practical test. The next decision concerns how the enterprise funds the work and staffs the judgement it requires.
Fund and staff the practice
Fund the business result, shared learning, and continuing service
Funding determines which obligations the practice must recognise. Evaluation maintenance means keeping the test cases, labels, policies, and comparison data current after the initial application is released. If an engagement budget covers only implementation, those obligations appear later as unplanned support work.
The investment-banking business unit should have a direct economic stake in the transaction-screening engagement. That does not mean banking must pay for every shared AI platform capability. It means banking funds the work that changes its screening process and can therefore see the resulting capacity, quality, or control benefit quickly. A business unit that funds a bounded workflow has a reason to define the result carefully, supply expert time, and decide whether the observed benefit justifies continued use.
I would separate three allocations even when the accounts sit in one budget. The engagement allocation covers discovery, banking participation, integration, application engineering, evaluation, and the initial release. The shared platform allocation covers primitives and improvements that serve several applications. The service allocation covers operation, model and data usage, support, incident response, source updates, and continuing evaluation. Naming the allocations prevents the enterprise from calling an application complete while leaving its future costs unassigned.
For the banking example, the engagement budget might pay for connecting approved research, market-data, and deal systems, modelling the screening cases, building the evidence package, running domain review, and testing the controlled release. The banking service owner would fund the application's normal operation and changes to the screening process. The platform team would fund a retrieval improvement only if it accepted responsibility for maintaining that improvement for other users. Existing platform services can reduce the incremental cost, but they do not remove the need to account for it.
Those allocations also explain why business-unit funding can be valuable. Central platform funding can make shared capability available, but it does not prove that a particular workflow deserves to exist. Banking funding creates a direct test of value: does the application reduce total screening effort, improve evidence quality, shorten a material delay, or enable work that the team could not otherwise complete? The answer should use an agreed comparison and include review, correction, coordination, and manual fallback. A faster extraction step alone is not the business result.
The same principle applies to chargeback. Charging investment banking for every platform request can discourage useful adoption. Funding everything centrally can attract requests whose sponsors have little reason to prioritise carefully. I would first observe demand and behaviour across an initial period. Record which business units request work, which workflows pass the intake conditions, how much engineering and expert time they require, which dependencies block them, and what continuing support they create. Compare those observations with the value and learning each engagement produces.
That evidence allows the practice lead and business leaders to choose an allocation method. They might fund discovery centrally, require a business unit to fund workflow delivery, and retain shared platform work in the platform budget. They might allocate service costs to the team that operates each application. They might use a common fund for high-value control improvements. The permanent allocation mechanism should follow observed demand and obligations, rather than a generic rule applied before the enterprise understands its workload.
The practice lead also needs a portfolio view. A portfolio decision is a decision about which of several possible engagements receives limited engineering, domain, platform, and operating capacity. Each proposal should show its expected business result, evidence already available, unresolved dependencies, continuing obligations, and credible alternatives. A transaction-screening application may compete with a manufacturing supplier workflow or a platform investment that would unblock both. Comparing those choices requires their different purposes to remain visible.
Funding and portfolio choices define what the practice can promise. The next section turns that promise into data, evaluation, and an operating destination.
Select an engagement that can test the practice
The first engagement should test both the application and the internal AI FDE arrangement. It needs a worthwhile workflow, a plausible technical path, access to suitable evidence, authority to change the relevant process, and a receiving team that can operate the result. These conditions are gates, not ingredients in a single score. A proposal with prohibited data access cannot compensate by claiming a large potential benefit. A proposal with no credible technical path cannot compensate with a senior sponsor.
For each proposal, I would write down the condition that remains uncertain and the evidence that would change the intake decision. In investment banking, uncertainty may concern whether entity and period definitions are consistent enough for peer comparison. In manufacturing, it may concern whether supplier documents contain the information required for reliable entity matching. A small discovery engagement can investigate uncertain extraction or data quality. The absence of permitted access can block delivery until the data owner resolves it.
Among feasible proposals, compare expected value, urgency, learning potential, risk, and opportunity cost in concrete terms. A transaction-screening workflow may affect a recurring origination process and provide a clear manual comparison. A manufacturing workflow may expose a reusable document-retrieval requirement. Neither should win because it receives a higher average score. Leaders should be able to say what decision the comparison supports and why the selected engagement uses the practice's scarce capacity well.
I would initially favour a related workflow family when the enterprise wants the practice to learn quickly. For example, transaction screening across two coverage groups can reveal which source and comparison requirements generalise. A second group may expose differences in entity identifiers, currencies, or review rules. Those differences provide useful evidence about what belongs in shared primitives and what remains local. Selecting unrelated demonstrations would make that learning harder to interpret.
Similarity should not become the only criterion. A related workflow still needs a banking owner with authority, protected analyst time, permitted data, a receiving team, and a credible outcome comparison. A workflow that offers excellent learning but no route into operation is a research project, not an FDE engagement. The first deployment should contain enough real difficulty to test the practice while keeping the consequences of failure manageable.
The initial decision is the investment committee or accountable technology and business leadership authorising that bounded investigation. It should name the workflow owner, FDE practice lead, participating platform and control teams, discovery scope, budget, uncertainty, and evidence required for the next commitment. It should not promise production delivery while a decisive dependency remains unknown. The first technical exercise should expose that dependency as efficiently as practical.
Hire for judgment, then change the assessment for AI-assisted coding
The transaction-screening workflow gives the practice a basis for staffing. An FDE needs to frame an unfamiliar business problem, build maintainable integrations, reason about data and permissions, evaluate system behaviour, and make progress with people who hold different authorities. The practice may draw on existing platform, data, design, product, and control teams. Dedicated hiring should fill the capability that the diagnosed delivery problem repeatedly lacks.
AI-assisted coding changes the implementation part of that job. Anthropic describes Claude Code being used across its organisation to navigate unfamiliar codebases, write tests, debug production issues, build automation, and create documentation. Its examples include legal, marketing, and data teams building useful tools with an agentic coding environment. How Anthropic teams use Claude Code.
OpenAI's account of its Codex experiment describes a related shift. Engineers spend less time writing individual lines and more time specifying intent, shaping repository knowledge, designing tool environments, and building feedback loops that let agents produce verifiable work. The human responsibility for architecture, ambiguous decisions, and ownership remains. Harness engineering, Building an AI-native engineering team.
This changes what the FDE practice should look for. The strongest candidate may not be the person who can type code fastest without assistance. The practice needs people who can decide what should be built, give an agent enough context to act, inspect the result, design tests that expose failure, and recognise when the agent has misunderstood the business requirement. An FDE who accepts a plausible screening brief without checking entity, period, permission, and evidence assumptions creates risk regardless of who wrote the code.
The interview should therefore permit approved AI coding tools while making the reasoning observable. A practical exercise could provide a small repository, an incomplete transaction-screening workflow, conflicting requirements, and a data-access limitation. Candidates could use Claude Code, Codex, or Copilot in the permitted environment. The assessment would examine how they inspect the repository, clarify the problem, decompose the work, instruct the agent, review generated changes, run tests, handle an access restriction, and explain what remains uncertain.
The exercise should include a requirement that the tool cannot solve through code generation alone. For example, the candidate might discover that the field called “enterprise value” has different meanings in two source systems. A strong response would pause implementation, ask which source is authoritative, and record the unresolved interpretation. A weak response would normalise both fields into one schema and rely on a test that checks only syntax. This distinguishes engineering judgement from successful prompting.
I would use a second exercise that tests failure handling and maintainability. The candidate could receive an agent-generated pull request containing a plausible integration, incomplete tests, and an unsafe permission assumption. The task would be to review the change, identify the risk, ask the agent for a correction, and decide which checks must remain human-owned. This resembles the practice's real work more closely than a timed algorithm question.
The interview should still examine fundamentals. A candidate needs to reason about data models, APIs, identity, state transitions, testing, observability, and system failure. The practice can ask for an architecture discussion without an agent, then compare it with the candidate's tool-assisted implementation. It can also ask the candidate to explain every important design choice in plain language to a banking owner. FDE work fails when technical and business explanations cannot meet.
Internally, the practice should provide a safe development environment for agentic coding. Repository instructions, architecture notes, test commands, data handling rules, and deployment constraints should be legible to both humans and agents. CI, static checks, evaluation cases, and review workflows should produce feedback that an agent can use. These controls do not make generated code correct. They make errors easier to find and reduce the amount of human attention spent on routine checks.
Career incentives must recognise the resulting work. If advancement rewards only launches or lines of code, engineers have little reason to improve repository structure, evaluation, transfer, or the receiving team's capability. I would recognise sound problem framing, reliable systems, useful shared primitives, clear negative findings, effective use of agents, and the ability to leave a workflow with an owner. The practice should develop engineers who can move between deployment, platform, and product work without making permanent dependence on one exceptional person part of the design.
Enable and operate the proposed solution
Establish data access and data meaning together
After the first engagement is authorised, data access becomes an engineering dependency. The team should specify the information required for the transaction-screening workflow, the systems that hold it, the identities that may retrieve it, the approved purpose, retention conditions, and processing environment. A data or control owner can then assess a concrete proposal instead of responding to a broad request for access to “banking data.”
The transaction-screening application may need target identifiers, public filings, market prices, financial measures, deal-room documents, and prior screening outcomes. Each field can have a different authority and retention rule. A diligence document may contain confidential information that is unnecessary for an initial screen. A prior screening outcome may represent a senior judgement, a temporary exception, or a value copied from another system. Access to the field does not establish that its meaning is understood.
I would maintain important definitions with the application. The record should identify authoritative sources, field meanings, effective dates, transformations, known gaps, and unresolved interpretations. Tests should cover the meanings that affect the screening decision. When a valuation method or source-system definition changes, the application behaviour and relevant evaluation cases should be reviewed.
The manufacturing comparison reinforces the point. A certificate can be accessible and readable while still applying to the wrong subsidiary or product category. The application needs source identity, issue date, expiry date, and document status. Procurement experts must define how those facts determine eligibility. The FDE team translates those definitions into application behaviour and evidence that the controls work.
Restrictions can also reveal a better design. The banking application may need selected filings rather than every document in a deal workspace. It may retrieve data under the analyst's identity rather than a broad service account. It may process records inside an approved environment. Each alternative still needs assessment against the actual use. Less access can reduce exposure, but it does not automatically establish that the proposed processing is allowed.
Define acceptable behaviour, evaluation, and outcome measurement at the start
The workflow, authority map, funding plan, and data definitions provide the basis for deciding what the application may do. The banking application may propose comparisons, show evidence, identify missing records, and route uncertain cases. It should not issue a trading instruction, approve a valuation, or communicate with a client merely because a model produces a confident explanation. Application logic should control state transitions, and access enforcement should sit outside model judgement.
Human review needs a concrete design as well. The banking analyst must see the source documents behind a proposed comparison, the reporting period and entity, the reason for an exception, and the information still missing. The reviewer must know when to accept, reject, request clarification, or escalate to valuation or risk. If the application produces an attractive summary that hides omissions, it can preserve a formal review step while making meaningful judgement harder.
Evaluation should be designed before the first experiment, because it determines whether the practice can tell if the application helps. Domain experts define acceptable and unacceptable screening outcomes using real cases. Engineers create repeatable checks for entity matching, state transitions, retrieval permissions, and failure recovery. Banking and control owners decide which errors matter most. The release authority decides whether the evidence supports the proposed scope.
The outcome comparison belongs in the same early plan. For comparable screening cases, record manual preparation effort, analyst review effort, clarification requests, correction effort, completion time, and the rate of unresolved exceptions. Include cases that do not use the application and cases that return to the manual process. The purpose is to measure the complete workflow, not only the fraction that produces an attractive automated output.
How to use: one manual screening case is 100 units of analyst time, 85 collecting and comparing and 15 writing the brief. That split and the starting values are assumptions. Set how much of the collecting and comparing the application removes, then add the three things a demonstration never shows. The readout gives the fallback rate at which the workflow costs more than it did by hand.
Different failures require different evidence. A comparison based on the wrong target is a business correctness failure. Retrieval of another deal team's document is an access failure. A duplicate case after a retry is a workflow-state failure. An analyst who cannot verify the evidence is a review-design failure. A single average score would conceal these differences. The practice should report each class and identify which release decision it affects.
Evaluation cases also require continuing ownership. Valuation methods change, new targets enter the process, and source systems add fields. The service owner should maintain the collection and rerun relevant checks when those conditions change. A universal sample size or timetable cannot establish readiness across every banking workflow. The needed evidence depends on variation, consequence, expected frequency, and the decision it must support.
Decide who will operate the application after the FDE engagement
The future operating arrangement should be decided while the application is being built. For the transaction-screening workflow, the business owner is accountable for the screening procedure and the definition of a useful result. The technical service owner is accountable for deployment, support, dependencies, model and retrieval changes, and incident response. Both need a named lead, backup capacity, access, and funding.
The receiving investment-banking technology team should participate in real delivery work. It can implement an integration, run an evaluation, deploy a controlled change, and diagnose a deliberately introduced retrieval failure. Those exercises reveal whether the team can operate the service with the agreed documentation and access. A signed handover record cannot demonstrate that capability.
The enterprise can choose continuing specialist support when the application requires expertise that would be inefficient to recreate. It can transfer operation to the investment-banking technology team when that team has suitable skills and capacity. It can retain a shared evaluation or retrieval service with the platform team. Each choice is valid when its responsibilities and funding are explicit. The problem is leaving the original FDE engineers as an invisible support team because no one made the operating decision.
The same test applies when external specialists participate. A partner may provide model expertise or implementation capacity, but the enterprise should decide which capabilities it must retain. Those may include source access, evaluation methods, deployment ability, control evidence, and the ability to change providers. The receiving team should practise those responsibilities before the engagement ends.
The practice now has a proposed result, decision owners, funding, evidence requirements, and an operating destination. The final setup step is to authorise a bounded commitment that can test those choices.
Authorise the first commitment and learn from it
Establish the practice through a sequence of explicit commitments
The setup is now a connected set of commitments. The enterprise has identified a repeated AI delivery gap, defined a transaction-screening workflow, named its dependencies and decision owners, chosen an organisational home, separated local rules from shared primitives, funded the business result and future service, selected an engagement, staffed for engineering judgement, and designed evaluation and operation.
The first commitment should authorise discovery for the transaction-screening workflow. It should name the banking owner, the FDE lead, the platform and control participants, the supported cases, the principal uncertainty, the budget, and the evidence required for further investment. The commitment is to answer a bounded question, not to promise a production launch before the evidence exists.
If discovery supports delivery, the next commitment should build the limited workflow and evaluate it against the agreed comparison. Leaders should review the application and the delivery arrangement separately. An application can produce useful proposals through an unsustainable amount of FDE intervention. A well-run investigation can also conclude that a different process or conventional software would be better. Both findings improve the next decision.
A second coverage group can test whether the first engagement's retrieval changes, evaluation method, and operating practices generalise. A manufacturing workflow can test whether the practice can adapt its approach when documents and procurement rules differ. These engagements should reveal which parts of the practice repeat reliably and which depended on unusual access, expert relationships, or local workarounds.
The practice should be willing to change its form as those patterns become visible. Repeated work around one integration may justify a platform investment. Concentrated demand within investment banking may justify a permanent product team. Mature workflow families may need standard components and enablement rather than intensive FDE delivery. The practice's success should be measured by the enterprise's ability to complete valuable workflows and sustain them, rather than by permanent growth in FDE headcount.
The next article follows this arrangement during delivery. It explains how an FDE team observes the work, chooses experiments, evaluates failures, releases within evidence, transfers responsibility, measures value, and decides whether to expand, reorganise, or stop.