AI Governance Maturity Models for Large Organizations
Organizations claim AI governance they can't actually demonstrate or enforce.

Ask a large enterprise to show proof that its AI systems stayed inside policy last quarter, and see what happens. The policies arrive quickly: ethics board charters, vendor risk assessments, framework adoption memos, all properly filed and version-controlled. What doesn't arrive is proof. The organization can describe what its AI is supposed to do. It cannot show, with evidence a regulator would accept, what the AI actually did, who was accountable when it drifted, or what stopped it from drifting further. The research has a name for this: the attestation deficit, the space between what a document claims and what a system can demonstrate.
The gap exists for a structural reason, not a motivational one. A governance framework only produces governance once its language turns into something that routes through an organization: roles that carry real authority, evidence that gets captured automatically, correction mechanisms that fire without a human remembering to check. A policy document does none of that on its own. It sits in a binder, or a wiki, or a slide deck from the AI ethics committee's Q3 offsite, and it waits for someone to operationalize it. Most large organizations never finish that second step, which is a separate and harder build than the first.
Part of the reason it never gets finished is that no one owns it. Responsibility for AI governance tends to scatter across IT, risk management, a cross-functional working group, and sometimes a dedicated AI governance office, none of which holds enough authority on its own to act when something goes wrong. A violation can be fully visible, logged, flagged, discussed in a meeting, and still go uncorrected because four different functions each believe someone else has the mandate to intervene. These organizations aren't negligent. They built governance instruments designed for a human-paced world and pointed them at a machine-paced problem, and the mismatch is doing exactly what mismatches do.
What a maturity model measures
A maturity model asks a narrow question: do governance activities exist, and how far have they spread across the organization? That's a process maturity measurement, and it's useful for what it is. It is not a measurement of whether governance actually stops bad AI behavior or enforces a policy at the moment an AI system acts. An organization can score well on a maturity assessment while having no runtime mechanism that would catch a single policy violation as it happens.
The fair objection to these models is that they reward documentation that looks like governance without confirming it functions as governance. At the policy-driven middle levels in particular, a predictable pattern appears: policies exist, compliance gates exist, and teams route around them because the gates slow down work without controlling what the AI actually does. Compliance becomes a toll booth, not a checkpoint. People pay it and keep driving the same way they were driving before.
None of this makes maturity models useless. It means they need to be read as a description of infrastructure, not as a finish line. Climbing from one level to the next only produces governance in the real sense when what gets added at each step is an enforcement capability, not another policy. The levels are scaffolding for a building that still needs to get built. McKinsey's update to its AI Trust Maturity Model makes the same point institutionally: assessed across roughly 500 organizations, the framework added agentic AI governance and controls as its own fifth dimension, separate from strategy, risk management, data and technology, and general governance. The field is now treating agent-level enforcement as its own lagging capability, not a line item inside governance that was already supposed to cover it.
Where large organizations sit on the maturity curve today
The space between what organizations claim about their AI governance and what they can actually demonstrate is wide. Most large organizations report active AI governance initiatives. Far fewer can show measurable maturity, and only a small share have governance that can hold up when someone really scrutinizes it.
The specific practices that separate governance on paper from governance that functions, model inventories, bias and fairness testing, red-team exercises against AI systems, appear in only a minority of organizations. Red teaming is the rarest of the three, which tracks with intuition: it's the practice that actually tries to break the system. Ownership matters more than the policy text itself. McKinsey's survey found that if an organization assigns an explicit AI governance role, it scores measurably higher on maturity than one that doesn't. The accountability structure moves the number, not the quality of the prose in the policy document.
Agentic AI is where the curve bends hardest. Close to three-quarters of organizations plan to deploy agents within two years, but only a small fraction can now point to a mature governance model for agents specifically. The gap isn't evenly spread across every dimension of governance either. A minority of organizations have a genuinely comprehensive AI policy, and a meaningful share have no AI policy at all, so a portion of the enterprise population hasn't cleared Level 1 of the ladder yet. The shortfall concentrates at the enforcement layer: the practices that require building something, rather than writing something, are the practices that are missing.
Why agentic AI makes the enforcement gap dangerous in a way earlier AI did not
Access control for enterprise systems was built around how humans actually behave: people get granted far more permission than they use, and the gap between granted access and exercised access sits there harmlessly for years. Agentic AI breaks that assumption at the root. An agent inherits human-scale permissions and then exercises them at machine speed, acting on access that a human in the same role would have left dormant indefinitely. The access model worked for humans because permissions could sit unused for years, but it fails for something that doesn't get tired, doesn't get distracted, and doesn't leave permissions unused out of habit or caution.
Agents also don't act in isolation. They call other systems, trigger downstream processes, and chain actions together, and that compounds quickly if oversight is thin. Governance infrastructure built to review static outputs, a model's prediction, a chatbot's response, has no mechanism for watching a sequence of interconnected, autonomous actions unfold across systems in real time.
The exposure already exists. A growing share of organizations run agents in production, and more than two-thirds cannot clearly tell AI agent actions apart from human activity, based on the Cloud Security Alliance's survey conducted with Aembit. That exposure is a condition already running inside production environments that most organizations are managing without the visibility to know what they're managing. Enforcement, in this setting, requires attributing a specific action to a specific agent, catching the moment an agent's behavior steps outside its authorized scope, and applying policy without waiting for a human to review every single action an agent takes. None of that comes from a document. Ownership gets worse here too: agents don't stay inside one department. They move through engineering, product, finance, and sales, so the same fragmented accountability that already stalls governance for human-driven AI spreads across more functions and gets thinner at each one.
What the maturity levels look like when enforcement is the lens, not documentation
The five-level maturity ladder is sound as a description of progress. The mistake large organizations make is treating each step up as a documentation milestone, a new policy, a new committee, a new training module, rather than an enforcement capability they now have and didn't before.
Level 1 is ad hoc: AI runs somewhere in the organization without repeatable process, clear accountability, or discipline around the data feeding it. Governance doesn't exist here in practice or on paper. The first task is discovery: finding out what AI is actually running and where.
Level 2 is policy-driven, where the documents first appear. The bottleneck at this stage is behavioral: compliance gates move slowly enough that teams work around them or quietly stop engaging with them. No enforcement infrastructure exists yet to check whether the policy is being followed in practice, so attestation, the document itself, becomes the only evidence an organization has to offer.
Level 3 is defined and repeatable, where governance processes get standardized across teams. The enforcement question changes shape here: does the same rule apply the same way regardless of which business unit or function is running the AI, and does anyone log and review the exceptions when it doesn't?
Level 4 is managed and measured, and the enforcement capability that separates it from Level 3 is real-time detection of anomalous AI behavior rather than a quarterly audit that finds problems months after they occurred. Organizations at this level can spot a risk crossing between functions while it's still forming, not after it's already caused damage.
Level 5 is optimizing and proactive. The defining enforcement trait is the ability to stop a violation before it completes. Policy gets enforced at the moment of action, and every AI action leaves a complete audit trail behind it. Each level, read this way, describes what an organization can prove under scrutiny, not what it has written down.
What enforcement-layer governance requires organizations to build
Closing the attestation deficit takes infrastructure, not better writing. What regulators can verify depends on runtime controls that exist independently of the policy document, not on how thorough that document is.
Discovery comes first, because an organization cannot govern a system it doesn't know exists. So you need a live inventory of every AI model and every agent running inside the organization, continuously updated, not a snapshot pulled together for last quarter's review. CMMI AIM's Data domain treats data lineage and data quality as a governance foundation, and the same logic extends to models and agents: knowing where each one runs and what access it holds is the starting condition for everything that follows.
Monitoring has to run continuously and has to be attributable to a specific actor. A governance process built around periodic audits cannot catch an agent that steps outside its authorized scope between one review and the next. Runtime monitoring means logging every agent action with enough detail to tell agent behavior apart from human behavior, which is the exact gap exposed by the finding that most organizations currently can't make that distinction.
Enforcement has to happen before a violation finishes, not after. Detecting a problem without the ability to halt or constrain the action that caused it leaves an organization reactive no matter how good its detection is. Whether an organization is at Level 4 or Level 5 depends on whether it can apply a rule the moment an agent acts, instead of flagging the action for review once it's already done.
Audit logs need to be complete and resistant to tampering. Regulatory frameworks including the EU AI Act require accountability that can be demonstrated, meaning audit trails that allow the system's operation to be traced, risk situations to be identified, and post-market monitoring to actually happen, with identification of the people who verified results required specifically for biometric identification systems. A summary report doesn't meet that bar. The Decision Evidence Maturity Model for Agentic AI gets at this directly: it evaluates whether the evidence an organization already has is sufficient to answer specific governance questions about an agent's decisions, and it names the trap most organizations fall into as the "container fallacy," the assumption that having logs or traces at all is the same as having logs sufficient for audit. That distinction is the enforcement-layer version of the policy most organizations already have sitting in a binder somewhere.
How the CMMI AIM model and the McKinsey AI Trust framework treat enforcement maturity differently
CMMI AIM and McKinsey's AI Trust Maturity Model start from different places on enforcement, so you gain more from understanding both than from adopting either one as a template.
CMMI AIM launched in July 2026 from CMMI Institute, and it's built around outcomes and integration. It applies AI-specific content across all 31 existing CMMI Practice Areas, maps directly to the EU AI Act, and lets organizations work toward DORA and NIS2 compliance through one integrated structure. Its enforcement contribution is appraisability: it gives organizations a standardized way to produce evidence of governance capability that an outside party can check. Government Technical Services Corporation took part in the CMMI AIM pilot specifically because its AI adoption had outpaced the governance, process, and controls needed to support it, a rationale that doubles as a description of the exact gap the model is built to close. IBM Consulting and Infosys Limited were part of the same pilot, with KPMG Assurance and Consulting Services LLP contributing appraisal expertise. The model's development drew on a working group of more than 25 industry experts and structured review from 100 companies, giving it a broad base of practical input.
McKinsey's framework, updated for 2026, takes a different route to the same problem by adding agentic AI governance and controls as a standalone fifth dimension. Its contribution is a forcing function: agentic AI now has to be assessed as its own domain with its own control requirements, rather than getting folded into a general governance score that could mask how far behind agent-specific controls actually are.
Microsoft's agentic AI adoption maturity model adds a third angle, tracking progress from early experimentation to full agent-first operation across five pillars: AI strategy and experience, business strategy and value, governance and security, technology and data, and organization and culture. Its usefulness is in pairing governance with technology explicitly. What an organization can enforce at any given stage is constrained by what its technology stack can actually support, and Microsoft's model treats those two pillars as advancing together. The practical move for a large organization is mapping the specific enforcement capabilities, discovery, continuous monitoring, attribution, runtime enforcement, tamper-evident audit trails, against whichever framework is already in use, and treating any shortfall in those capabilities as a maturity gap no matter what level the framework happens to assign.
Why the regulatory environment rewards enforcement maturity now, not attestation
A policy document is no longer enough proof of control for regulators. DORA and NIS2 push the same expectation into operational resilience and incident reporting, where a committee charter has never been able to help.
Frameworks like CMMI AIM build toward appraisable evidence, and McKinsey now assesses agentic controls as their own category. Regulators are no longer asking whether an organization has a position on AI governance. They ask whether the organization can produce the record that proves the position held up when an AI system actually acted. An organization that has spent years perfecting its policy language but never built the infrastructure to enforce it will find that the paperwork, however thorough, answers a question no one in the room is asking anymore.


