Est.

ISO 42001 Certification Readiness for Enterprise AI Programs

Enterprise buyers now require ISO 42001 certification or a credible roadmap from AI vendors.

Staff Writer · · 10 min read
Cover illustration for “ISO 42001 Certification Readiness for Enterprise AI Programs”
AI Risk Frameworks · October 8, 2026 · 10 min read · 2,294 words

ISO 42001 certification has gone from a nice-to-have badge on a vendor's website to a line item procurement teams will not skip. Three forces pushed it there, and together they make sitting still the costliest option on the table. At the same time, enterprise buyers spent 2025 adding "ISO 42001 certified or roadmap" language to vendor questionnaires, so certification questions now show up in a substantial share of enterprise AI vendor RFPs across both EU and North American markets. If a vendor cannot show a certificate or a credible plan, the deal is gone before anyone even talks about price.

The third force is proof of concept, and there is a growing list of it. IBM certified its Granite models, Anthropic certified Claude, and Microsoft's 365 Copilot was first certified in March 2025 and recertified in March 2026 with zero non-conformities. KPMG Australia became the first organization in the world to achieve ISO 42001 certification, through BSI Australia. Singapore's Changi Airport holds SGS's first-ever accredited ISO/IEC 42001 certificate. Anaplan announced its certification on September 10, 2026, after being certified by Bay Mountain Security on August 13, 2026, and Aurigo Software followed with its own announcement on September 22, 2026. Each of these certifications means an independent auditor checked the organization's governance, risk management, deployment, and ongoing oversight against a fixed standard and signed off.

The infrastructure behind those certifications has caught up with the demand for them. Schellman was accredited by ANAB to issue ISO 42001 certifications in September 2024, and UKAS granted BSI its first accreditation for the same purpose on January 15, 2026. A standard is only as credible as the bodies qualified to audit against it, and that bench is no longer thin. What was theoretical in 2023 is now something named companies have already walked through, audited and certified, which is a different kind of pressure than a regulation on paper.

Mapping ISO 42001's structure onto an enterprise AI program

Diagram: ISO 42001's Ten Clauses at a Glance. Visualizes: Show the ten-clause structure of ISO/IEC 42001 as a numbered linear sequence, making clear which clauses are foundational framing versus the ones where real operational work lives.

ISO/IEC 42001 is the first international standard for artificial intelligence management systems, published in December 2023. It uses the same High-Level Structure as ISO 27001 and ISO 9001. Any organization that has gone through an information security or quality certification will recognize the skeleton immediately, even if the content inside it is new.

The standard runs ten clauses. Clauses 1 through 3 set scope, references, and AI-specific terms such as AI system, AI lifecycle, intended use, provider, deployer, and user. Clause 4 covers organizational context, defining the AIMS scope, identifying internal and external issues, and mapping interested parties. Clause 5 covers leadership: top management commitment, AI policy, and roles and responsibilities. Clause 6 is planning, and it carries the standard's most distinctly AI-specific addition, the AI risk assessment and the AI system impact assessment. Clause 7 covers support: resources, competence, awareness, documented information. Clause 8 is operation, where lifecycle controls get put into practice. Clause 9 covers performance evaluation through monitoring, internal audit, and management review. Clause 10 covers improvement: nonconformity handling, corrective action, continual improvement.

Annex A is where the AI-specific controls live, organized into categories covering AI policies, internal organization, resources for AI systems, impact assessment, the AI system lifecycle, data for AI systems, information for interested parties, use of AI systems, and third-party relationships. No organization has to implement every control in Annex A. Instead, each produces a Statement of Applicability documenting which controls are in scope, which are excluded, and why. If an organization only deploys third-party AI models rather than building its own, it can exclude several lifecycle controls but still has to keep every third-party supplier control, because the risk it carries sits with its vendors, not its own training pipelines. That document becomes the map for everything that follows: the clauses define the process, Annex A defines the controls, and the Statement of Applicability defines which of those controls apply to a given organization's actual footprint.

Stage one: scoping the AIMS and building a complete AI inventory

The stage most organizations underestimate is the first one: figuring out what is actually inside the management system before trying to govern it. Self-reported AI inventories, the kind built from interviews and surveys, routinely miss unsanctioned tools, free-tier model usage, and AI features quietly embedded inside SaaS applications that nobody had flagged as an AI system.

Clause 4 requires defining the AIMS scope directly: which AI systems, which business units, which geographies, and which AI roles (provider, deployer, user, partner) the organization occupies for each one. That last part matters because the same organization can be a provider of one AI system and a deployer of another, with different obligations attached to each role.

Shadow AI, unsanctioned AI use outside any tracked system, produces this stage's governance gap. A large share of organizations using AI have no policies governing third-party AI use, and the gap between what gets self-reported and what is actually running tends to be wide. An engineering team piping customer data into a code assistant nobody vetted, a sales team running prompts through a free consumer chatbot, a finance team using an AI-powered spreadsheet plug-in: none of it shows up on an org chart, and all of it is now something the AIMS has to account for.

Scoping decisions carry cost consequences downstream. A narrow, well-defined scope keeps audit complexity and timeline down. A scope that is narrow because systems were missed, rather than deliberately excluded, creates expensive rework when those systems appear during the audit itself, which they tend to do. The practical fix combines automated discovery, pulling from network traffic, SaaS integration logs, and expense data for AI subscriptions, with structured interviews across engineering, product, finance, and sales, the functions where AI adoption moves fastest and gets governed least. The formal output of this stage is the Statement of Applicability: a documented position on which of the 38 Annex A controls apply, which are excluded, and the stated rationale for each exclusion. What gets built in later stages depends on that document being accurate, not convenient.

Stage two: AI risk assessment and the system impact assessment

Once the inventory is settled, the question shifts from what exists to what it could do wrong. This stage builds the risk register that drives which controls get selected and how much weight each one carries.

ISO 42001's risk work goes further than a conventional cyber risk assessment, because you have to evaluate failure modes that traditional security frameworks were never built to catch. Two distinct processes sit inside Clause 6. The AI risk assessment identifies and evaluates risks to the organization itself: system failures, misuse, and unintended behavior. The AI system impact assessment, covered under Clause 6 and Clause 8 and widely regarded as one of the highest-effort controls in the standard (A.6.1), evaluates potential harm to the people affected by the system: customers, employees, third parties, anyone downstream of a decision the system makes.

The risks that fall outside conventional cyber assessment include prompt injection, hallucinations, model manipulation, data poisoning, AI agent abuse, privacy leakage, and model drift. When auditors check this stage, they look for a documented scoring methodology that covers likelihood, severity, and stakeholder impact. They look for a direct line from each high-risk scenario to a specific Annex A control addressing it, executive approval records for any residual risk the organization has decided to accept rather than mitigate, and refresh cycles scheduled to trigger automatically when a new dataset arrives, a regulation changes, or an incident teaches the organization something it didn't know before.

You cannot write the risk register once and then file it away. It has to update on defined triggers: a new model deployed, a new dataset introduced, an incident detected, a regulatory shift. AI agents deserve particular attention here, because they inherit human-scale access permissions and act at machine speed. A misconfigured or compromised agent can do in minutes what would take a careless employee weeks to replicate, and the risk register has to treat that speed differential as a first-class risk factor.

Stage three: building the governance structure and leadership accountability

Clause 5 asks for more than a signed policy sitting in a shared drive. It requires named individuals who hold real authority and the operational ability to step into a live AI system and change what it's doing. That requirement exposes how many enterprises have governance structures that exist on paper but hold no actual power over the systems they claim to oversee.

Top management has to show documented commitment to the AIMS. An AI policy that gets written by the compliance function and never reaches an executive signature does not satisfy Clause 5, no matter how well it reads. Roles and responsibilities have to be assigned by name, not by department: someone owns the AI risk register, someone approves impact assessments, someone authorizes the deployment of a new AI system. Most critically, someone holds the mandate and the technical means to pause, stop, or amend a system that is already running in production.

Human oversight is one of the standard's more demanding provisions precisely because it asks for mechanism, not assertion. Saying a human is "in the loop" satisfies nobody. The standard wants defined triggers for intervention: specific conditions under which a human steps in, and a system built so that stepping in is actually possible. High-impact decisions have to stay contestable, with escalation paths and override mechanisms that function end to end. Auditors test whether those mechanisms work in practice, not whether they're described accurately in a policy document.

An AI governance committee can exist, meet, and produce minutes at this stage while having no decision rights over what the AI systems in production actually do. A governance body that can recommend but not act is a discussion group, not an oversight function, and ISO 42001 auditors are trained to tell the two apart.

Stage four: implementing lifecycle controls and data governance

Clause 8 and the Annex A lifecycle controls are where the standard gets technically demanding, and where the distance between a written policy and operational evidence is hardest to close. Most of the implementation effort, and most of the audit findings, concentrate here.

Lifecycle control runs across five phases, and each one needs documented procedures that have evidence behind them. Development and training requires data-quality gates, bias testing, and reproducibility requirements. Deployment and change management requires approval workflows, human-oversight checkpoints, and release documentation. Monitoring and incident response means tracking model performance, bias, and security signals, with predefined escalation paths that fire when thresholds get breached. For continual improvement, you feed what monitoring turns up back into retraining cycles, policy updates, and the risk register. Decommissioning requires retirement procedures, data disposition plans, and notification to the parties affected when a system is retired.

Data governance controls under Annex A.7.x are among the costliest to implement well: data lineage, quality validation, provenance documentation, and protection requirements all need to be in place and auditable, not just described. Teams have to trace where data came from, document how it was processed, and validate that it represents the population the system serves, because auditors test whether an organization's transparency statements match what actually happens day to day, not what the policy says should happen.

Model cards are the standard's documentation artifact for explainability, and they need enough detail on architecture, training parameters, and limitations that a stakeholder outside the engineering team can understand how a given output got produced. Drift detection has to be an operational capability, not a reporting habit. The monitoring system has to be able to flag anomalies in data or concept drift when they cross a threshold, not surface them in a quarterly report after the fact. If an organization runs AI agents, behavioral monitoring works as the agent equivalent of drift detection. An agent that strays from its approved action patterns is a live governance event, and the AIMS has to be built to catch it and escalate it while it's happening, not after.

Stage five: third-party supplier controls and the foundation model problem

Annex A.10, and A.10.3 in particular, asks for supplier controls that go well past what ISO 27001's vendor management provisions cover. Organizations that assume their existing vendor management program already handles this are wrong often enough that it counts as a pattern rather than an exception.

The AI-specific additions break into three pieces. A.10.2 requires you to allocate responsibility across the AI supply chain, so it's clear which party owns which obligation when something goes wrong. A.10.3 requires you to assess and manage AI suppliers against responsible AI expectations, including fairness and explainability, and not just uptime and data security. A.10.4 requires factoring in customer obligations and expectations: an organization has to understand what its own customers are owed before it can confirm its suppliers are meeting that bar.

An organization can pass its ISO 27001 audit cleanly, then fail ISO 42001 because its AI model explainability logs weren't retained under the same retention policy as its incident records. Nobody flagged it as a problem until an auditor asked for the logs and found a gap that had been invisible to the organization itself.

Foundation model vendors, the API-based providers of large language models, run into a structural version of the same problem. The documentation an organization needs to pass its own audit, training data provenance, bias evaluation results, model cards, is often not something the vendor can produce, or it exists in a form that doesn't satisfy what an auditor is asking for. That's not a gap an organization can fix after the fact by asking nicely post-signature. The fix is building supplier questionnaires and contract language that specify what AI documentation a vendor has to provide, set before the contract is signed rather than discovered after the audit scope is already locked in.

Sources

  1. ISO 42001: Practical Implementation Guide
  2. ISO 42001 Certification: The New Benchmark for AI ... - SAP Community

More in AI Risk Frameworks