Skip to content
LAUNCH FILM — LIVEGREENLIGHT · PATENT-PENDINGRUNTIME AUTHORIZATION FOR AI AGENTSMCP CONNECTORS — IN DESIGNISO/IEC 27001 — ROADMAPSOC 2 TYPE II — ROADMAPISO/IEC 42001 — ROADMAPMAS FEAT — DESIGN-ALIGNEDDETERMINISTIC · EXPLAINABLE · TAMPER-EVIDENTAI ACTS · CRELIS DECIDESLAUNCH FILM — LIVEGREENLIGHT · PATENT-PENDINGRUNTIME AUTHORIZATION FOR AI AGENTSMCP CONNECTORS — IN DESIGNISO/IEC 27001 — ROADMAPSOC 2 TYPE II — ROADMAPISO/IEC 42001 — ROADMAPMAS FEAT — DESIGN-ALIGNEDDETERMINISTIC · EXPLAINABLE · TAMPER-EVIDENTAI ACTS · CRELIS DECIDES
All posts
Audit Trail 19 August 2026

Shadow Mode: How to Prove AI Governance Works Before You Let It Block Anything

What happens if your policy engine is wrong on a Tuesday afternoon? Until a governance pilot has an answer to that, it does not get approved, because the risk of the control looks larger than the risk it was bought to manage.

Shadow mode removes that objection by deciding nothing. It watches every action an agent proposes, evaluates it against your policies, records what the decision would have been, and then stays out of the way. The agent proceeds exactly as it did the day before. What you gain is evidence about your own estate, gathered before anyone has to defend a decision to block something.

That is the pattern we keep seeing described. This is the pilot plan that answers it. What gets switched on, what changes in the execution path, what you hold at the end of it, and the part most descriptions skip: what shadow mode cannot tell you.

AI agent shadow mode

Shadow mode is an evaluation mode in which every AI agent action is checked against policy and recorded, but no action is ever stopped. The decision that would have been made is written to the record; the action itself proceeds untouched. It runs on one small configuration change with nothing else in your stack modified, and it produces an audit-ready record of every action your agents attempt alongside what policy would have said about each one .

It is the first rung of a ladder. SHADOW observes, ADVISORY informs, ENFORCE governs. You move one business line at a time, with no flag-day integration and no proxy sitting in your execution path .

How to pilot AI governance without blocking?

You pilot it by separating two things that are usually sold as one: evaluating an action and controlling it.

Evaluation is cheap and reversible. It answers "what would we have said about this?" Control is expensive and political. It answers "who is accountable when this stops a payment run at 2am?" Most governance pilots fail because they ask a change board to approve both at once, on the strength of a slide.

Shadow mode buys the first without the second. For the duration of the pilot, the honest answer to "what happens if your policy engine is wrong?" is that nothing happens. A wrong decision in shadow mode is a row in a report somebody reads on Friday. That is a very different conversation from the one where a wrong decision is an interrupted transaction.

The practical consequence is that the approval needed to start is much smaller than the approval needed to enforce. You are not asking to be in the critical path. You are asking to watch.

AI policy dry run production

A dry run is only useful if it runs against the traffic you actually care about.

Start on traffic that is not yet load-bearing, because it costs nothing and catches the obvious configuration mistakes. Do not stop there. The decisions that matter are the ones your agents make against real systems, with real amounts and real customers attached, and those are exactly the decisions a test copy of your systems never sees. A dry run that only ever sees synthetic traffic tells you your policies parse. It does not tell you what your agents are doing.

What makes this safe to point at your own live systems is also what makes it worth doing: nothing in the execution path changes. Systems that do not know about the governance layer keep working exactly as they do today, while every action is still checked and recorded. Watching and advising require zero changes anywhere .

If observing were risky, it would not be observation.

Test AI guardrails without breaking workflows

There is a second mode worth knowing about, because it is the one most teams actually want and few ask for by name.

Advisory mode informs without governing. A policy fires, the decision is recorded, the relevant person is told, and the action still proceeds. It is the middle rung, and it is where a team learns whether its policies are right rather than merely whether they are enforceable.

That distinction matters more than it sounds. A policy that is technically correct and operationally intolerable will pass every test you write for it and then fail on contact with a Tuesday. Advisory mode surfaces that before enforcement does, at the cost of a notification instead of an outage.

Move to enforcement per system, when that system's own evidence justifies it. Full enforcement is a choice you make deliberately for one business line, not a switch that flips an estate .

AI governance proof of concept plan

A pilot that produces an opinion is a wasted month. A pilot that produces evidence is a decision. This is the shape that produces evidence.

Before you start, write down the number you expect. How many actions a day do you think your agents take against systems of record? Teams are routinely wrong about this, and the gap between the guess and the count is often the single most useful output of the whole exercise.

Week one: shadow, on traffic that is not load-bearing. Confirm the configuration is right and the policies parse. Nothing here is a finding. This is plumbing.

Weeks two and three: shadow, against your own live systems, one business line. This is the actual pilot. At the end you hold a record of every action your agents attempted, what policy would have said, and which actions would have been held. Read it looking for three things: actions you did not know were happening, decisions where the policy was wrong, and decisions where the policy was right and the outcome would still have been unacceptable. The third category is the important one.

Week four: advisory on the same line. Now a person is in the loop and nothing is blocked. You are testing whether the notification reaches someone who can act on it, which is an organisational question no amount of policy testing answers.

Then decide. Enforcement on one line, or another cycle. The evidence supports either, and the choice is now a business decision rather than a bet.

What shadow mode cannot tell you

This is the section most descriptions leave out, and it is the reason to trust the rest of them.

It cannot tell you whether enforcement would have been survivable. A held action in shadow mode costs nothing. The same action held under enforcement has a person waiting on it, a service level attached to it, and a queue behind it. Shadow mode measures the frequency of holds. It does not measure the cost of holds, and those are different numbers with different owners.

It cannot find the actions your agents did not attempt. A record of proposed actions is a record of what the agent tried, not of what it could try. If a policy would have caught a class of action your agents never happened to take during the pilot window, the pilot is silent about it. Absence of evidence in a four-week window is not evidence of a safe estate.

It cannot validate a policy nobody exercised. A policy that never fired during the pilot is untested, not proven. Read the list of policies that never triggered and treat it as unfinished work rather than a clean bill of health.

It cannot tell you whether the policy is the right policy. It tells you whether the policy behaves as written. Whether it should have been written that way is a judgement, and judgement does not come out of a log.

What this means if you have to produce evidence

The output of a shadow-mode pilot is not a dashboard. It is a record.

Every decision is chained and can be replayed, and the record is tamper-evident by construction rather than by assurance . What that leaves you with is a document you can hand to somebody who was not there: this is what our agents did, this is what our policy said about each action, this is who would have been asked, and here is the proof the record has not been altered since.

That is a better artefact than a pilot report, because it does not require anyone to trust the person who wrote it. It is also the thing you will eventually be asked for — by an auditor, a regulator, or a customer's risk function — and building the habit of producing it during a pilot is considerably cheaper than building it during an incident.

Start in shadow mode where nothing is load-bearing. See every decision the policy layer would have made, before you let it make one.

Want the full story?

Explore GREENLIGHT