Skip to content
LAUNCH FILM — LIVEGREENLIGHT · PATENT-PENDINGRUNTIME AUTHORIZATION FOR AI AGENTSMCP CONNECTORS — IN DESIGNISO/IEC 27001 — ROADMAPSOC 2 TYPE II — ROADMAPISO/IEC 42001 — ROADMAPMAS FEAT — DESIGN-ALIGNEDDETERMINISTIC · EXPLAINABLE · TAMPER-EVIDENTAI ACTS · CRELIS DECIDESLAUNCH FILM — LIVEGREENLIGHT · PATENT-PENDINGRUNTIME AUTHORIZATION FOR AI AGENTSMCP CONNECTORS — IN DESIGNISO/IEC 27001 — ROADMAPSOC 2 TYPE II — ROADMAPISO/IEC 42001 — ROADMAPMAS FEAT — DESIGN-ALIGNEDDETERMINISTIC · EXPLAINABLE · TAMPER-EVIDENTAI ACTS · CRELIS DECIDES
All posts
Accountability 11 July 2026

AI Agent Accountability Framework: Financial Services

A bank that deploys an AI agent inherits a narrower obligation than most governance programmes assume. It is not to prove the model is safe. It is to show, for one specific action the agent took, the authority it was operating under, the decision that let that action through, and a record of both that does not depend on the agent's own account of itself. Everything else in an accountability framework is scaffolding around that artefact.

What is an AI agent accountability framework for financial services?

An AI agent accountability framework for financial services is a control layer that evaluates each proposed agent action before it executes, checks it against the authority that agent was actually granted, and writes an independent record of the decision. It governs actions, not models. Its purpose is evidentiary: to make a single past action reconstructable by someone who was not there and does not trust the agent.

That framing is not a vendor invention. It is the structure the industry itself converged on when it wrote down what runtime governance of agents in finance would have to do.

For each agentic system and action type, institutions need to establish what the system is authorised to do, how proposed actions are assessed at runtime before execution, and what records are retained to support review, accountability, and remediation when outcomes diverge from intent. — Monetary Authority of Singapore and industry contributors, Safeguards for Agentic Finance at Runtime

Three obligations, in one sentence: define the authority, assess the action before it happens, keep the record. A framework that does not do all three is a description of good intentions.

One caution before going further, because it is the most common error in this category of writing. SAFR is an industry white paper published under a regulator's initiative. It is not a rule, and it says so plainly.

It does not constitute regulatory guidance or supervisory expectations, nor does it prescribe or anticipate future directions for such guidance or expectations — Safeguards for Agentic Finance at Runtime

Treat it as the clearest available statement of where industry practice is heading, not as an obligation with a commencement date. A compliance officer who is told otherwise stops reading. We cover the Singapore picture in more depth in our note on verifiable compliance for Singapore financial services.

What does an AI accountability framework for banks have to contain?

Four things, and they are ordered. Each one is useless without the one before it.

Identity. The framework has to know which agent is asking. Not which application, and not which service account — which agent, resolved against a registry that says what that agent is. An action attributed to "the payments system" is not attributable to anything.

Authority. Somewhere there must be an explicit, machine-readable statement of what this agent may do, within what limits, and for how long. This is the part institutions most often skip, because the equivalent for a human employee is a job description and an approval matrix that already exist on paper. Paper does not travel with a request.

A decision at the point of action. Before execution, the proposed action is tested against that authority and receives a binding outcome. In SAFR's model the outcomes are four: reject it, hold it for a person, let it run, or let it run while flagging it for review. Four outcomes rather than two matters more than it sounds, because a binary allow-or-block forces every ambiguous case into whichever answer is cheaper.

A record written by something other than the agent. This is the artefact the other three exist to produce.

The United States has been assembling the same picture from a different direction. In February the Treasury released the Financial Services AI Risk Management Framework alongside an AI lexicon, adapting the NIST risk framework to financial services.

the U.S. Department of the Treasury today released two new resources to guide AI use in the financial sector, a shared Artificial Intelligence Lexicon and the Financial Services AI Risk Management Framework (FS AI RMF) — U.S. Department of the Treasury

Note the verb. Released, to guide. Resources, not requirements. Anyone telling you this framework is "in full effect" and that your institution is now in breach of it has not read the announcement. That does not make it unimportant — supervisors read what industry publishes, and a control objective you ignored is an awkward thing to explain later. It makes it a benchmark rather than a deadline.

Who is liable when an AI agent makes a payment?

The institution that deployed the agent. This is the one part of the question that is not genuinely open, and it does not depend on any framework published this year.

Singapore's FEAT principles predate agentic systems entirely, and they already answer it:

Firms using AIDA are accountable for both internally developed and externally sourced AIDA models. — Monetary Authority of Singapore, Principles to Promote Fairness, Ethics, Accountability and Transparency

The same document requires clear ownership of the decision inside the firm, with an internal authority that approved the use in the first place. Externally sourced changes nothing. Buying the model does not move the accountability to the vendor, and neither does buying the agent framework, the orchestration layer, or the governance tool.

So the practical question is never "who is liable". It is "when the liability lands, what can you produce". Those are different questions and only the second one has an engineering answer. We have written separately on why the autonomous black box is a poor defence.

There is a second-order point worth stating, because it changes what you build. An agent's own logs are the agent's testimony. If the record of what happened is produced by the same component whose behaviour is in question, it is evidence of what the agent reported, not evidence of what occurred. The record has to be written by something with no stake in the answer.

AI agent controls for financial institutions: what actually stops an agent?

A control either changes what happens or it does not. Most of what is filed under AI governance does not.

A policy stating that agents must not initiate payments above a threshold does not stop an agent from initiating one. A model card does not. A quarterly attestation does not. These are worth having and they are not controls; they are descriptions of controls you may or may not have built. The distinction becomes visible only under examination, which is the worst moment to discover it.

An enforcing control has three properties. It sits between the agent and the system it acts on, so the action cannot route around it. It evaluates deterministically, so the same inputs give the same outcome and the outcome can be explained without reference to a model's confidence. And it produces a record whether the answer was yes or no — denials are the evidence that the control was live.

There is a subtler property that multi-step agents make essential, and SAFR states it directly:

An Auto-Execute or Observe outcome at one step carries no authority into the next — Safeguards for Agentic Finance at Runtime

Permission granted for step one is not permission for step four. An agent that adapts to intermediate results is, by the end of a workflow, doing something nobody authorised at the start. Frameworks that authorise a session rather than an action miss this entirely, and it is exactly where money moves.

How confident is the sector that its controls would survive inspection? Grant Thornton's banking survey puts a number on it.

Half of banking executives say governance and compliance are already limiting AI performance, yet only 18% are sure they could pass an independent audit of AI controls. — Grant Thornton, Banking insights: 2026 AI Impact Survey

Two caveats belong with that number, because a bank's own reviewers will find them. The banking cut is 50 of the 950 leaders surveyed — the smallest cell in the study — and Grant Thornton marks findings within it as directional. The same report separately puts 20% in the explicitly not-confident column, which leaves around three in five who are neither: not confident, not worried, simply untested.

Read the figures together rather than separately. Governance is already expensive enough to slow the work down, and it is still not producing confidence. That is the signature of documentation without enforcement: the cost is fully incurred and the assurance is not. The untested majority is the more interesting group anyway — they are not disputing that they could produce evidence, they have just never been asked to.

What evidence does a bank need for an AI agent decisions review?

This is the part most articles gesture at and never specify. The white paper specifies it.

Each log entry should capture the essential elements of a governance decision: the governance envelope as submitted, the mandate against which the action was checked, the outcome produced by the Disposition Engine, the specific rules applied, the basis for that outcome, and the time elapsed at each stage. — Safeguards for Agentic Finance at Runtime

Translated into what an examiner asks for, one action at a time:

  • What was proposed. The action, its parameters, and the steps the agent took to arrive at it — the tools it called, the data it read, the checks it performed.
  • Under what authority. The specific grant of permission the action was tested against, and where that grant came from.
  • What was decided. The outcome, the rules that produced it, and the reason. A denial needs a reason recorded as much as an approval does.
  • Who, if anyone, intervened. If a person reviewed it, that person's decision is part of the record, not a separate ticket in a separate system.
  • When. Time at each stage, so the sequence can be reconstructed rather than inferred.

If you can produce that set for an action chosen at random by someone else, you have an accountability framework. If you can produce it only for actions you selected in advance, you have a demonstration.

The gap between having a policy and having that record is wide and measurable. IBM's breach research found the policy layer largely absent:

Additionally, among the 600 organizations researched by the independent Ponemon Institute, 63% revealed they have no AI governance policies in place to manage AI or prevent workers from using shadow AI. — IBM, Cost of a Data Breach Report

That figure is about policies, not records. The population with a working evidentiary trail is necessarily smaller than the population with a policy, because the record is the harder artefact. Any institution that treats a written policy as the finish line should assume it is in the larger, weaker group.

What this means if you have to produce evidence

Five things you can check this week, none of which require buying anything.

  1. Pick one agent action from last month at random. Not a representative one. Ask for the four items above. Time how long it takes and count how many systems the answer has to be assembled from. If the answer is "we would need to reconstruct it", that is your finding.
  1. Find the authority statement. For one agent, locate the explicit, machine-readable limits on what it may do. If the only statement of its limits is a paragraph in a design document, the limits are not enforced.
  1. Check who writes the record. If the audit trail is produced by the agent framework itself, note it as a dependency. Testimony from the component under examination is the weakest form of evidence you can offer.
  1. Look for the denials. A control that has never refused anything is either perfectly calibrated or not connected. Denial records distinguish the two.
  1. Check whether authority carries across steps. Take a multi-step workflow and establish whether each action was evaluated, or whether one approval covered the sequence.

These are the questions worth asking a vendor as well, in the same order. Whether the vendor is us or anyone else, the answers should be demonstrable rather than described. You can read what GREENLIGHT does and does not do, and our security page sets out its own limits rather than glossing them. In the interest of the same standard this article argues for: Crelis does not currently hold SOC 2, ISO 27001, or ISO 42001 certification, and that page says so in those words.

More work on this cluster sits under financial services.

The uncomfortable version of this article's argument is short. Most institutions can describe their AI controls and cannot evidence them, the difference is invisible until an examiner picks an action, and the fix is not a larger policy. It is a record that someone else can read.

Want the full story?

Explore GREENLIGHT