MAS's AI Risk Management Proposals: What a Financial Institution Would Have to Evidence
The consequence of MAS’s AI risk management work is evidence pressure. MAS says its Artificial Intelligence (AI) Model Risk Management information paper sets out good practices for AI and Generative AI model risk management, with attention to governance, oversight, risk management systems and processes, and development and deployment of AI. MAS also says it is reviewing responses to an earlier public consultation on Guidelines on AI Risk Management. A financial institution reading that material should not stop at the policy wording.
Key Takeaways
- MAS position. MAS encourages financial institutions to reference good practices when developing and deploying AI.
- Model risk scope. The MAS information paper is framed around governance and oversight, risk management systems and processes, and the development and deployment of AI.
- Guidelines status. MAS says it is reviewing responses to an earlier public consultation on Guidelines on AI Risk Management.
- Toolkit direction. MAS refers to an Artificial Intelligence (AI) Risk Management Toolkit for the financial services sector.
- Evidence gap. A control that records only the final outcome does not show who authorised the action, which policy allowed it, or why escalation did not occur.
MAS AI risk management guidelines in context
The cleanest reading starts with the documents and stops where they stop. MAS’s information paper is not described in the quoted wording as a rulebook. It is described as an information paper that sets out good practices observed during a review. MAS’s own wording also says that all financial institutions are encouraged to reference these good practices when developing and deploying AI.
"This information paper sets out good practices for AI and Generative AI model risk management that were observed during the review, focusing on those relating to governance and oversight, key risk management systems and processes, and development and deployment of AI."
That wording matters. “Good practices” and “encourages” are not the same as a direct command. A compliance team that treats the paper as if every sentence were already a binding requirement overstates the source. A technology team that treats “encourages” as if it meant “optional and safe to ignore” makes the opposite mistake.
The better operational question is narrower. If a reviewer asks how the institution applied the good practice, what artefact can the institution produce. A policy PDF is one artefact. A model approval note is another. Neither is the same as a record showing that a particular AI agent was allowed to release a payment, alter a customer record, or change a credit limit at the moment it acted.
MAS has also said that it is reviewing responses to an earlier public consultation on Guidelines on AI Risk Management. That statement does not give a final text for those guidelines. It does tell risk owners that the direction of travel is not confined to voluntary experimentation.
"3 MAS is presently reviewing responses to an earlier public consultation on a set of Guidelines on AI Risk Management."
— MAS Partners Industry to Develop AI Risk Management Toolkit for the Financial Sector
This is where loose AI governance writing usually becomes unhelpful. It turns governance into a noun and then walks away. Governance only becomes inspectable when it produces records: who approved the use case, what limits applied, what changed before deployment, what happened when the action was attempted, and what evidence remains after the action completed.
The page on MAS Guidelines is useful background for that reason. It should not be read as a substitute for the MAS text. The live question for a regulated firm is not whether it has an AI governance story. The question is whether the story survives contact with one disputed action.
AI model risk management at MAS
MAS frames its Artificial Intelligence (AI) Model Risk Management information paper around AI and Generative AI model risk management. The quoted scope includes governance and oversight, risk management systems and processes, and development and deployment of AI. Those categories are broad enough to touch the full lifecycle of an AI use case, but the quoted wording does not list every record an institution must keep.
That limit is important. A vendor should not pretend that the information paper, standing alone, prescribes a particular software architecture. It does not. The quoted MAS wording also does not say that buying any product satisfies MAS expectations. Anyone claiming that should be asked to point to the exact words.
Even so, the evidence burden is not abstract. Governance and oversight require more than a committee name. If an AI system proposes a credit-limit increase and an agent executes it, the institution needs a way to show the authority chain behind the increase. That chain is not the same as the model’s explanation. MAS’s SAFR wording points to how agent actions are authorised and what is recorded at the point of every decision, not to the ticket that approved the project. It is the record that says this actor was permitted to do this thing under these conditions.
Risk management systems and processes have the same problem. They can be strong on design and weak at runtime. A control can require approval for a payment above a threshold, but if the approval is stored as a generic workflow event, later review may not show whether the AI action was within scope, whether it should have escalated, or whether the approver saw the relevant facts.
Development and deployment controls also need an audit trail that survives deployment. Pre-release testing can show that a system behaved acceptably before go-live. It cannot, by itself, prove that a later deleted record was permitted by the rule that existed when deletion occurred. Post-event evidence has to be captured at the point of decision, not reconstructed from memory and dashboards.
MAS’s material on the AI Risk Management Toolkit points in the same operational direction. MAS says the MindForge AI Risk Management Toolkit features an AI Risk Management Operationalisation Handbook that provides practical guidance on implementing AI risk management frameworks. MAS also says the toolkit provides resources for managing AI-related risks across traditional AI, generative AI, and emerging agentic AI technologies.
That does not remove the need for judgement. “Operationalisation” can become another label pasted onto a document library. It only helps if the institution can map each control to an artefact that would answer a later challenge.
A simple test is brutal and fair. Pick one AI action that touched something important. Ask what policy allowed it. Ask who had authority over the policy. Ask whether the decision was automatic, escalated, approved, or blocked. Ask whether the record shows what was recorded at the point of decision, because MAS’s SAFR wording points to that point. If the answer depends on a Slack thread, a model log, and a recollection, the evidence is weak.
The related Crelis piece on MAS SAFR and evidence takes the same stance for agentic systems. It is not enough to say that an agent complied. The institution has to show the permission that made the action legitimate.
MAS AI guidelines: banks and evidence records
Banks should be especially careful with the word “guidelines”. MAS’s environmental risk pages show that MAS uses guidelines to set expectations for financial institutions in other risk domains. That does not prove the final content of the AI risk management guidelines. It does show why a bank should not treat MAS guideline language as mere commentary.
For banks, the evidence problem is practical. An AI assistant that drafts a response is different from an agent that changes a fee waiver, releases a payment file, or amends a customer limit. The second class of activity needs a record of authority, not only a record of output.
The phrase “model risk management” can mislead engineers here. It can make the problem sound as if it lives only inside the model. The MAS information paper is wider than that, because its quoted focus includes governance and oversight, risk management systems and processes, and development and deployment of AI. Those words reach the controls around the model, not just the model artefact.
This is where an AI inventory becomes useful, but only if it is tied to action. A spreadsheet listing use cases may satisfy a first request. It will not answer why a particular agent was allowed to refund a customer after the customer had already disputed the transaction. The inventory needs links to ownership, allowed actions, escalation paths, and runtime records.
A lifecycle control has the same weakness when it stops at stage gates. Approval to develop is not approval to execute every future action. Approval to deploy is not approval to delete any record the agent can reach. MAS’s SAFR wording is about agent actions being authorised and recorded at the point of every decision, so an evidence file should not collapse project approval into action authorization.
That distinction matters when the institution has to defend itself. A bank can have a clean model development file and still fail to show why one action was allowed. It can have access logs and still fail to show the business authority for a sensitive change. It can have monitoring alerts and still fail to show that the action met the policy at the moment of execution.
MAS also says these examples offer insights into challenges, approaches and risk management practices when using AI in different organisational contexts. Those examples can help teams compare approaches, but they are not the bank’s own evidence. The bank still needs its own record for its own action.
Financial institutions can use the wider material collected under Financial Services to compare the pattern across posts. The recurring issue is not whether AI governance exists on paper. It is whether the authority for an action is captured before the action reaches money, customers, records, or infrastructure.
What this means if you have to produce evidence
Evidence starts at the point where a proposed action becomes an authorised action. That is a smaller unit than a model. It is also smaller than a use case. The unit that matters in a dispute is the payment release, the deleted record, the changed credit limit, or the blocked attempt.
A governance pack can show that oversight exists. It may show that the board or a committee received a status update. It may show that a risk function reviewed a use case. It will not necessarily show that the agent’s action was within authority at the moment it touched a system.
That is the gap between outcome logging and authorization evidence. An outcome log says what happened. Authorization evidence says why the action was allowed or stopped. The first helps reconstruct an incident. The second helps defend the decision.
MAS’s SAFR wording makes the runtime point explicit. MAS says SAFR proposes a framework for governance of AI agents in financial services, defining how agent actions are authorised, how human oversight is activated, and what is recorded at the point of every decision.
"Safeguards for Agentic Finance at Runtime (SAFR) proposes a framework for the governance of AI agents in financial services, defining how agent actions are authorised, how human oversight is activated, and what is recorded at the point of every decision."
That sentence does more work than most governance summaries. It ties agent governance to authorisation, oversight, and recording at the point of decision. It does not say that a generic application log is enough. It does not say that a model trace is enough. It points to the moment where the institution must decide whether the action can proceed.
The uncomfortable implication is that evidence design has to be part of control design. If the control says “escalate when the payment is unusual”, the record should show whether escalation happened. If the control says “block changes to restricted records”, the record should show that the attempted change was blocked. If the control says “allow only within approved limits”, the record should show the limit that applied.
There is a second implication. Evidence has to be readable by someone who was not in the delivery team. Internal dashboards often assume local knowledge. They use shorthand. They show fragments. They may be adequate for operations and poor for accountability.
That is why the record should not depend on the agent’s own story about itself. The model may explain why it proposed an action. The institution still needs the separate authority record that allowed, escalated, approved, or blocked the action. A fluent explanation is not a permission.
Practical steps for an evidence file
- Separate the use case from the action. Record the approved AI use case, but do not treat that approval as proof that every later action was allowed. A customer-limit change needs its own authority record.
- Map each sensitive action to a policy. For each payment release, record change, account amendment, or customer communication, identify the rule that allowed the action or required escalation. If the rule cannot be named, the control will be hard to defend.
- Record the decision before the system is touched. The useful evidence is the allow, approve, escalate, or block decision at the point of action. A later summary is weaker because it asks the reviewer to trust reconstruction.
- Keep approval evidence distinct from identity evidence. Knowing which agent acted is necessary. It is not sufficient. The record must show what the agent was allowed to do at that moment.
- Preserve the reason for non-escalation. If no human approval was required, the evidence should show why. That is often the first question after a loss, a complaint, or a disputed change.
- Make the evidence readable outside engineering. A risk owner should be able to follow the record without translating system internals. If only the implementation team can explain the decision, the artefact is not yet fit for review.
- Test the file against a hostile scenario. Choose one completed action and assume it is challenged. The file should show the actor, the policy, the decision, the escalation state, and the final outcome.
Crelis exists for the narrow part of this problem: runtime authorization and evidence for AI agent actions. Crelis describes GREENLIGHT as a permission step in which agent requests for sensitive actions are routed through Crelis, and Crelis says systems can refuse sensitive actions that arrive without a valid Crelis permission. For teams comparing that model with the MAS material, the practical next step is to examine GREENLIGHT against one real action path and, if it fits the risk, use the product demonstration to test whether the evidence file would answer a reviewer.
Crelis is patent-pending. It holds no SOC 2, ISO 27001 or ISO 42001 certification, and claims none.
Related reading
- MAS Guidelines — related Crelis posts on MAS guidance and evidence.
- Verifiable AI Compliance: Singapore Financial Services — a closer look at AI compliance evidence in Singapore financial services.
- AI Risk Management Platform: Verifiable Oversight — how oversight records differ from operational logs.
Want the full story?
Explore GREENLIGHT