Skip to content
BIZENIUS

What is AI governance? A practical guide for regulated institutions

BIZENIUS Advisory Team · Last updated: 25 August 2026

Written and reviewed by the BIZENIUS advisory practice — senior practitioners from risk, treasury, finance and supervision.

Not an ethics statement and not a policy document — the decision rights, risk tiering, validation standards and evidence that let an institution answer for what its models did. What a working framework contains and what a decorative one looks like.

In short

  • AI governance is the set of decision rights, standards and records through which an institution controls what its artificial-intelligence systems are allowed to do, establishes that they work, and remains able to answer for their outputs afterwards. It is a control framework rather than a statement of values.
  • An institution that cannot produce the model inventory on request has not, in any meaningful sense, got a framework — whatever the policy says.
  • Where a system genuinely cannot be validated to the depth its tier requires, the honest output is a restriction on what it may be used for rather than an approval with a caveat nobody reads.
  • Human oversight has to be specified as a mechanism, because as a principle it is unfalsifiable.
  • Monitoring matters more for AI than for classical models because the failure mode is silent.
On this page
  1. What AI governance actually is
  2. Why existing model governance is not enough
  3. Scope, defined by consequence rather than technique
  4. Risk tiering keeps the framework proportionate
  5. The inventory is the framework’s spine
  6. Validation shifts from mechanism to behaviour
  7. Human oversight has to be a mechanism
  8. Monitoring, because the failure mode is silent
  9. Third-party systems and what to contract for
  10. Where frameworks fail

What AI governance actually is#

AI governance is the set of decision rights, standards and records through which an institution controls what its artificial-intelligence systems are allowed to do, establishes that they work, and remains able to answer for their outputs afterwards. It is a control framework rather than a statement of values.

The test of one is not whether it expresses the right principles but whether, when a specific automated decision is questioned months later, the institution can say:

  • which model made it
  • on what version
  • using what data
  • who approved that model for that purpose
  • what testing supported the approval
  • whether a human could have intervened

Why existing model governance is not enough#

The reason it cannot be inherited wholesale from existing model risk management is that AI systems break several assumptions that framework was built on. A classical credit or capital model has a stated functional form, a fixed set of inputs, parameters someone estimated deliberately, and behaviour that changes only when the model is redeveloped.

Many AI systems have none of those properties: the relationship between input and output is learned rather than specified, the input space is far wider and often unstructured, performance can degrade as the world moves without any code changing, and a growing share of what institutions deploy is a third-party system whose internals are not available for inspection at all.

Existing model governance is the right foundation. It is not the whole building.

Scope, defined by consequence rather than technique#

The first component of a working framework is scope, and it is where most frameworks quietly fail before they begin. A definition of “AI system” narrow enough to exclude the vendor tool that ranks alerts, the spreadsheet macro that has become a decision rule, and the model somebody in a business line built without telling anyone, is a definition that governs almost nothing.

The workable approach defines scope by consequence rather than by technique: any system whose output materially shapes a decision about a customer, a counterparty, a financial figure or a regulatory obligation falls in, however it was built and whoever built it.

Technique determines the controls; consequence determines whether the framework applies at all.

Risk tiering keeps the framework proportionate#

Risk tiering follows, and it is what keeps the framework proportionate enough to be obeyed.

A model that drafts internal summaries and a model that declines credit applications should not carry the same validation burden, and a framework that pretends otherwise will be complied with on the low-risk cases and quietly bypassed on the high-risk ones, because the cost of compliance is felt where the pressure to ship is greatest.

Tiering is normally driven by the consequence of an incorrect output, whether the decision affects a customer directly, whether a human reviews it before it takes effect, the reversibility of the outcome, and the scale of exposure. Each tier then carries its own approval level, validation depth, monitoring frequency and documentation standard.

The inventory is the framework’s spine#

The inventory is the framework’s spine and, in practice, the deliverable that produces the most immediate surprise. Building it means asking every function what makes or shapes decisions on its behalf, which reliably surfaces systems nobody had counted: a vendor module inside a core platform, a scoring rule embedded in a workflow tool, a procured analytics service, a departmental model that outlived the person who built it.

Each entry needs an owner in the business rather than in technology, a stated purpose, a tier, a validation date and a review date. An institution that cannot produce this list on request has not, in any meaningful sense, got a framework — whatever the policy says.

Validation shifts from mechanism to behaviour#

Validation is where the difference from classical model risk is sharpest. Where the internals cannot be inspected — because the system is procured, or because the learned relationship is not meaningfully readable — validation shifts from examining the mechanism to examining behaviour:

  • performance measured on data the developer never saw
  • testing across the customer segments and edge cases where failure would matter most
  • comparison against a simpler benchmark the institution understands
  • adversarial testing of inputs designed to break it
  • stability testing over time

Independence still means what it always meant — the validator is not the builder, and has the standing to withhold approval. Where a system genuinely cannot be validated to the depth its tier requires, the honest output is a restriction on what it may be used for rather than an approval with a caveat nobody reads.

Human oversight has to be a mechanism#

Human oversight has to be specified as a mechanism, because as a principle it is unfalsifiable. Everyone agrees a human should be in the loop; almost nobody defines what the human is expected to do, what information they receive, how much time they have, and what happens when they disagree.

Oversight that consists of a reviewer confirming a recommendation on a screen with no contrary information and a queue behind them is not oversight, and a framework should be able to distinguish the two.

The useful specifications are concrete: which decisions require review before they take effect, what the reviewer must be shown including the basis for the recommendation, how an override is recorded, and — the measure that reveals everything — what the override rate actually is. An override rate of zero means the control is decorative.

Monitoring, because the failure mode is silent#

Monitoring matters more for AI than for classical models because the failure mode is silent. A model that was accurate at approval can become inaccurate without any change to its code, simply because the population it scores has shifted, a channel mix has changed, or the behaviour it learned belongs to conditions that no longer hold. Detecting this requires:

  • tracking outcome accuracy against realised results rather than only tracking whether the system is running
  • watching the distribution of inputs for drift away from the development population
  • monitoring outcome rates across customer segments where fairness is at stake
  • setting in advance the threshold at which the model is retrained, restricted or withdrawn

The threshold must be defined before it is approached; a decision to keep running a degrading model, taken under commercial pressure and without a pre-agreed trigger, is the most predictable failure in the whole framework.

Third-party systems and what to contract for#

Third-party systems deserve their own treatment because they are the majority of what most institutions actually deploy, and because the accountability does not transfer with the software. A supervisor asking why a customer was declined will not accept that the vendor knows.

What is obtainable, and should be contracted for before purchase rather than requested after an incident, is:

  • documentation of intended use and known limitations
  • evidence of the testing the vendor performed and on what population
  • a right to test independently on the institution’s own data
  • notification before the model changes
  • access to logs sufficient to reconstruct a decision
  • an exit that does not strand the institution mid-process

Procurement is the only moment at which these terms are cheap.

Where frameworks fail#

The frameworks that fail do so in recognisable ways.

  • A policy expressing principles with no decision rights attached, so nobody can say who approves what.
  • Scope defined by technique, leaving the vendor tools and the departmental models outside.
  • Uniform controls that are too heavy for trivial uses and therefore evaded on serious ones.
  • An inventory compiled once for an audit and never maintained.
  • Human oversight asserted but never measured.
  • Monitoring that confirms the system is available rather than that it is right.
  • most commonly, a framework written by a function that does not own any of the models, approved by a committee that has never rejected one, and referenced by nobody building anything.
A governance framework that has never stopped a project is not governing.

Frequently asked

What is AI governance?

AI governance is the set of decision rights, standards and records through which an institution controls what its artificial-intelligence systems may do, establishes that they work, and remains able to answer for their outputs afterwards. It is a control framework rather than a statement of values, and the practical test is whether a specific automated decision can be reconstructed months later: which model made it, on what version, using what data, who approved that model for that purpose, what testing supported the approval, and whether a human could have intervened. Its core components are scope, risk tiering, a maintained model inventory, independent validation, specified human oversight, ongoing monitoring and third-party controls.

How is AI model risk different from traditional model risk?

Traditional model risk management assumes a stated functional form, a fixed set of inputs, parameters someone estimated deliberately, and behaviour that changes only on redevelopment — assumptions many AI systems break. The input-output relationship is learned rather than specified, the input space is wider and often unstructured, performance can degrade silently as the underlying population shifts without any code changing, and a large share of deployed systems are third-party products whose internals cannot be inspected. The consequence is that validation moves from examining the mechanism to examining behaviour, and that ongoing monitoring against realised outcomes becomes a primary control rather than a periodic check.

What should an AI model inventory contain?

An AI model inventory should list every system whose output materially shapes a decision about a customer, a counterparty, a financial figure or a regulatory obligation — including vendor modules inside purchased platforms, scoring rules embedded in workflow tools, procured analytics services and departmental models that outlived their builders. Each entry needs an owner in the business rather than in technology, a stated purpose and permitted use, a risk tier, the date and outcome of its last validation, its monitoring arrangements and a next review date. Scoping the inventory by consequence rather than by technique is what prevents the systems that matter most from falling outside it.

What does meaningful human oversight of an AI system look like?

Meaningful human oversight is specified as a mechanism rather than asserted as a principle: it states which decisions require review before they take effect, what the reviewer must be shown including the basis for the recommendation and any contrary indicators, how much time the review realistically allows, how an override is recorded and justified, and who is accountable for the final outcome. The measure that reveals whether it is real is the override rate — a reviewer confirming recommendations on a screen with no contrary information and a queue behind them is performing a formality, and an override rate at or near zero indicates the control is decorative rather than operative.

More where this came from

Browse the full resources hub, or subscribe in the footer for occasional substantial pieces.

BIZENIUS

Speak to an expert

Tell us where you stand — an expert replies within one business day.

Phone *
Area of interest
+ Add a message or details (optional)

We only use your details to respond to your enquiry. See our Privacy Policy.