Banking AI · 12-stage checklist · free CSV

AI implementation in banking: from pilot to controlled production.

A practical operating guide for banks and fintech companies implementing AI workflows and agents under model, vendor, security, compliance, and customer controls.

To implement AI in banking, start with one owned workflow and a measured baseline. Define the intended use before selecting technology. Register the models, vendors, data, decisions, and owners. Test representative and high-impact cases. Integrate human review, exceptions, audit evidence, incident response, and a stop path. Run a bounded production pilot. Scale only after business results and control evidence support the decision.

Scope boundary: this guide is an operating framework, not legal, regulatory, audit, cybersecurity, model-validation, or financial advice. A bank must apply current requirements with its qualified owners and supervisors.

Pilot-to-production gap

Why banking AI pilots get stuck.

A demonstration proves that a model can produce an output. It does not prove that the bank can operate the complete workflow.

Ownership

No accountable workflow owner

The innovation team runs the pilot, but no business owner accepts the outcome, process change, residual risk, or operating budget.

Evidence

No baseline or acceptance gate

The team measures model quality but cannot show the change in cost, time, accuracy, customer effect, control performance, or incidents.

Control

Risk joins after design

Model risk, compliance, legal, privacy, security, audit, and third-party risk find missing requirements after the system and contract exist.

Operations

The workflow stops at the API

The pilot omits identity, permissions, source systems, queues, human review, exceptions, monitoring, support, and recovery.

12-stage implementation checklist

Make each stage produce a decision and evidence.

StageRequired decisionMinimum evidence
1. MandateName the workflow and executive objective.Current process, baseline, users, volume, cost, quality, risk, target, and stop conditions.
2. Intended useDefine what the AI system will and will not do.Affected decisions, customer impact, materiality, jurisdictions, users, and accountable owners.
3. InventoryRegister every model, vendor, component, and dependency.Versions, data sources, third parties, business owner, control owner, status, and change path.
4. Data and accessSet the minimum data and tool access.Rights, lineage, classification, security, privacy, retention, secrets, permissions, and logs.
5. Build or buySelect an architecture and delivery model.Requirements, due diligence, contract rights, concentration, change notice, continuity, and exit.
6. ControlsAssign decision rights and human oversight.Reviewer, approval, override, escalation, separation of duties, audit trail, and prohibited actions.
7. EvaluationSet thresholds before release.Representative cases, high-impact failures, quality, bias where relevant, security, latency, and cost.
8. IntegrationConnect the system to the real workflow.Interfaces, permissions, exceptions, fallback, incident process, customer contact, and support.
9. PilotApprove a bounded production release.Named users, limited scope, monitoring, issue log, adoption data, and scale or stop gate.
10. MonitorOperate within defined thresholds.Performance, drift, incidents, complaints, overrides, vendor changes, and outcome analysis.
11. ScaleExpand only from verified evidence.Benefit record, residual risk, control performance, capacity, training, funding, and approval.
12. ExitRetire or replace without losing control.Continuity, data return or deletion, customer effect, evidence archive, lessons, and next owner.
Agentic AI in banking

Treat every agent action as a controlled capability.

An agent can retrieve data, call a tool, change a record, send a message, or trigger another process. Control the action, not only the text response.

Access

Use least privilege

Separate read, draft, recommend, approve, and execute rights. Bind access to an identity, role, purpose, environment, and time.

Action

Set transaction limits

Require human approval or dual control for material, irreversible, customer-facing, financial, or regulated actions.

Attack

Test adversarial paths

Test prompt injection, poisoned retrieval, unauthorized tool calls, data leakage, secret exposure, excessive agency, and unsafe fallback.

Audit

Reconstruct each decision

Retain the version, input source, retrieved evidence, tool call, output, reviewer, override, action, time, and result where policy permits.

Change

Control model and vendor updates

Define change notice, regression tests, approval thresholds, rollback, emergency disablement, and vendor incident communication.

Outcome

Measure more than containment

Track task quality, customer effect, cost-to-serve, latency, adoption, overrides, complaints, incidents, and financial outcome.

Vendor diligence

Evaluate the complete third-party relationship.

The model is one dependency. Hosting, retrieval, monitoring, support, data processors, tools, integrators, and contract terms also affect the bank.

  1. Workflow fit: intended users, actions, data, languages, channels, volume, latency, exception rate, and customer effect.
  2. Technical fit: architecture, interfaces, identity, logging, model options, data location, portability, observability, and failure behavior.
  3. Risk evidence: security, privacy, validation access, model information, subcontractors, business continuity, incidents, and independent reports.
  4. Operating rights: use of bank data, training rights, change notice, audit, service levels, support, remediation, termination, deletion, and transition.
  5. Concentration and exit: critical dependencies, replacement time, retained knowledge, export formats, fallback, and orderly termination.

Paul Okhrem's AI due-diligence framework and AI governance checklist provide reusable decision records for this work.

Success measures

Measure business value and control performance together.

Measure groupExamplesDecision it supports
BusinessCycle time, cost-to-serve, throughput, revenue, loss, customer retentionDid the workflow improve the intended outcome?
Task qualityAccuracy, completeness, consistency, exception rate, false positive and false negative costIs the system fit for its intended use?
CustomerResolution, complaints, wait time, access, fairness measures where relevantDid service improve without unacceptable harm?
ControlOverrides, unauthorized actions, incidents, response time, audit gaps, vendor changesAre controls operating at the required level?
AdoptionEligible users, active use, abandonment, workarounds, training, support demandHas the operating workflow actually changed?
EconomicsModel, infrastructure, integration, review, support, control, rework, and exit costIs the full operating case still positive?
Current primary sources

Use banking guidance that matches the system and date.

  1. Federal Reserve SR 26-2, Revised Guidance on Model Risk ManagementIssued April 17, 2026. It supersedes SR 11-7 and emphasizes a risk-based approach tailored to a banking organization's model risk profile, size, complexity, and use.
  2. Federal Reserve supervisory guidance on model risk managementCovers development and use, validation and monitoring, governance and controls, and vendor or third-party products. The guidance states that generative and agentic AI are outside its model scope, while bank governance should still determine appropriate controls.
  3. OCC Bulletin 2023-17, Third-Party RelationshipsInteragency guidance for planning, due diligence and selection, contract negotiation, ongoing monitoring, and termination across the third-party lifecycle.
  4. NIST AI Risk Management FrameworkVoluntary cross-sector framework for trustworthiness considerations across AI design, development, use, and evaluation.

Freshness note: older banking articles often cite SR 11-7 as current. The Federal Reserve states that SR 26-2 superseded SR 11-7 on April 17, 2026.

FAQ

Banking AI implementation questions.

How should a bank implement AI?

A bank should start with one owned workflow and a measured baseline. It should define the intended use, classify risk, register the system and vendors, control data and access, test representative and high-impact cases, integrate human review and exceptions, run a bounded production pilot, monitor performance and incidents, and scale only after evidence supports the decision.

How do you secure AI agents in banking?

Give each agent the minimum data and tool access needed for its task. Separate read and write permissions. Require approval for high-impact actions. Validate inputs and outputs, restrict tools, protect secrets, log decisions and actions, test prompt and data attacks, monitor abnormal behavior, and provide a tested stop and recovery path.

How should a bank evaluate an agentic AI vendor?

Evaluate workflow fit, intended use, architecture, data rights, security, model and subcontractor dependencies, evaluation access, audit logs, human controls, change notice, monitoring, incident support, financial condition, concentration risk, continuity, contract rights, and exit. Test the vendor in the bank's own representative cases before production.

Why do banking AI pilots fail to scale?

Common causes are an unowned workflow, no baseline, unclear intended use, inaccessible data, late risk involvement, weak integration, vendor dependence, no representative evaluation, missing exception handling, low user adoption, and no scale or stop gate. A successful demo does not prove safe or valuable production operation.

How should a bank measure AI implementation success?

Measure business outcome, task quality, failure impact, customer effect, control performance, adoption, latency, operating cost, incidents, overrides, complaints, and model or vendor change. Compare results with the pre-implementation baseline and record an explicit scale, revise, or stop decision.

Paul Okhrem, AI transformation consultant

About Paul Okhrem

Paul Okhrem is an AI Transformation Consultant and Fractional Chief AI Officer. He helps executive teams connect AI use cases, governance, vendors, implementation, adoption, and measurable acceptance. This checklist is free to reuse under CC BY 4.0 with attribution.