No accountable workflow owner
The innovation team runs the pilot, but no business owner accepts the outcome, process change, residual risk, or operating budget.
A practical operating guide for banks and fintech companies implementing AI workflows and agents under model, vendor, security, compliance, and customer controls.
To implement AI in banking, start with one owned workflow and a measured baseline. Define the intended use before selecting technology. Register the models, vendors, data, decisions, and owners. Test representative and high-impact cases. Integrate human review, exceptions, audit evidence, incident response, and a stop path. Run a bounded production pilot. Scale only after business results and control evidence support the decision.
Scope boundary: this guide is an operating framework, not legal, regulatory, audit, cybersecurity, model-validation, or financial advice. A bank must apply current requirements with its qualified owners and supervisors.
A demonstration proves that a model can produce an output. It does not prove that the bank can operate the complete workflow.
The innovation team runs the pilot, but no business owner accepts the outcome, process change, residual risk, or operating budget.
The team measures model quality but cannot show the change in cost, time, accuracy, customer effect, control performance, or incidents.
Model risk, compliance, legal, privacy, security, audit, and third-party risk find missing requirements after the system and contract exist.
The pilot omits identity, permissions, source systems, queues, human review, exceptions, monitoring, support, and recovery.
| Stage | Required decision | Minimum evidence |
|---|---|---|
| 1. Mandate | Name the workflow and executive objective. | Current process, baseline, users, volume, cost, quality, risk, target, and stop conditions. |
| 2. Intended use | Define what the AI system will and will not do. | Affected decisions, customer impact, materiality, jurisdictions, users, and accountable owners. |
| 3. Inventory | Register every model, vendor, component, and dependency. | Versions, data sources, third parties, business owner, control owner, status, and change path. |
| 4. Data and access | Set the minimum data and tool access. | Rights, lineage, classification, security, privacy, retention, secrets, permissions, and logs. |
| 5. Build or buy | Select an architecture and delivery model. | Requirements, due diligence, contract rights, concentration, change notice, continuity, and exit. |
| 6. Controls | Assign decision rights and human oversight. | Reviewer, approval, override, escalation, separation of duties, audit trail, and prohibited actions. |
| 7. Evaluation | Set thresholds before release. | Representative cases, high-impact failures, quality, bias where relevant, security, latency, and cost. |
| 8. Integration | Connect the system to the real workflow. | Interfaces, permissions, exceptions, fallback, incident process, customer contact, and support. |
| 9. Pilot | Approve a bounded production release. | Named users, limited scope, monitoring, issue log, adoption data, and scale or stop gate. |
| 10. Monitor | Operate within defined thresholds. | Performance, drift, incidents, complaints, overrides, vendor changes, and outcome analysis. |
| 11. Scale | Expand only from verified evidence. | Benefit record, residual risk, control performance, capacity, training, funding, and approval. |
| 12. Exit | Retire or replace without losing control. | Continuity, data return or deletion, customer effect, evidence archive, lessons, and next owner. |
An agent can retrieve data, call a tool, change a record, send a message, or trigger another process. Control the action, not only the text response.
Separate read, draft, recommend, approve, and execute rights. Bind access to an identity, role, purpose, environment, and time.
Require human approval or dual control for material, irreversible, customer-facing, financial, or regulated actions.
Test prompt injection, poisoned retrieval, unauthorized tool calls, data leakage, secret exposure, excessive agency, and unsafe fallback.
Retain the version, input source, retrieved evidence, tool call, output, reviewer, override, action, time, and result where policy permits.
Define change notice, regression tests, approval thresholds, rollback, emergency disablement, and vendor incident communication.
Track task quality, customer effect, cost-to-serve, latency, adoption, overrides, complaints, incidents, and financial outcome.
The model is one dependency. Hosting, retrieval, monitoring, support, data processors, tools, integrators, and contract terms also affect the bank.
Paul Okhrem's AI due-diligence framework and AI governance checklist provide reusable decision records for this work.
| Measure group | Examples | Decision it supports |
|---|---|---|
| Business | Cycle time, cost-to-serve, throughput, revenue, loss, customer retention | Did the workflow improve the intended outcome? |
| Task quality | Accuracy, completeness, consistency, exception rate, false positive and false negative cost | Is the system fit for its intended use? |
| Customer | Resolution, complaints, wait time, access, fairness measures where relevant | Did service improve without unacceptable harm? |
| Control | Overrides, unauthorized actions, incidents, response time, audit gaps, vendor changes | Are controls operating at the required level? |
| Adoption | Eligible users, active use, abandonment, workarounds, training, support demand | Has the operating workflow actually changed? |
| Economics | Model, infrastructure, integration, review, support, control, rework, and exit cost | Is the full operating case still positive? |
Freshness note: older banking articles often cite SR 11-7 as current. The Federal Reserve states that SR 26-2 superseded SR 11-7 on April 17, 2026.
A bank should start with one owned workflow and a measured baseline. It should define the intended use, classify risk, register the system and vendors, control data and access, test representative and high-impact cases, integrate human review and exceptions, run a bounded production pilot, monitor performance and incidents, and scale only after evidence supports the decision.
Give each agent the minimum data and tool access needed for its task. Separate read and write permissions. Require approval for high-impact actions. Validate inputs and outputs, restrict tools, protect secrets, log decisions and actions, test prompt and data attacks, monitor abnormal behavior, and provide a tested stop and recovery path.
Evaluate workflow fit, intended use, architecture, data rights, security, model and subcontractor dependencies, evaluation access, audit logs, human controls, change notice, monitoring, incident support, financial condition, concentration risk, continuity, contract rights, and exit. Test the vendor in the bank's own representative cases before production.
Common causes are an unowned workflow, no baseline, unclear intended use, inaccessible data, late risk involvement, weak integration, vendor dependence, no representative evaluation, missing exception handling, low user adoption, and no scale or stop gate. A successful demo does not prove safe or valuable production operation.
Measure business outcome, task quality, failure impact, customer effect, control performance, adoption, latency, operating cost, incidents, overrides, complaints, and model or vendor change. Compare results with the pre-implementation baseline and record an explicit scale, revise, or stop decision.