Est.

NIST AI RMF Implementation for Autonomous Agents

NIST's framework needs rethinking for AI systems that act, not just write.

Contributing Editor · · 10 min read
Cover illustration for “NIST AI RMF Implementation for Autonomous Agents”
AI Risk Frameworks · October 5, 2026 · 10 min read · 2,196 words

Autonomous agents did not break the NIST AI Risk Management Framework. They exposed what it was never built to see: the gap between a system that answers and a system that acts. Implementing the AI RMF for agents means extending each of its four functions, Govern, Map, Measure, Manage, past the assumptions baked into a 2023 document written for a different kind of machine. This piece walks through what each function demands when an AI system stops producing text and starts producing consequences.

Why the base NIST AI RMF was not built for agents

Picture a language model asked to draft a vendor invoice email. It writes the paragraph, hands it back, and its job ends there. Now picture an agent given the same task: it reads an inbox, decides a message means "send the invoice," calls the payment API, and the money moves. Nothing stopped between the thought and the transaction. That is the entire shift agentic AI represents, and the NIST AI RMF 1.0, published January 2023, was conceived for the first case: a system that takes an input, produces an output, and stops.

The RMF's four functions, Govern, Map, Measure, Manage, organize risk management as a continuous lifecycle, and that architecture is sound. Nothing about adding agents to the picture requires tearing it down. What the base framework lacks is any accounting for what happens when a model gains the ability to call tools and act on its own, unsupervised, in a live production system. The 2024 companion document, NIST AI 600-1, the Generative AI Profile, pushed the framework further into risks like confabulation, harmful content, and intellectual property exposure. But AI 600-1 was written for systems that generate content, not systems that take action. It says nothing about delegation chains, nothing about multi-agent coordination, nothing about runtime permission boundaries.

That silence is not a minor omission to patch with a footnote. Agentic systems carry risk properties that sit outside what either RMF 1.0 or AI 600-1 was built to describe, so you close that gap by extending the framework rather than discarding it. Four problems recur across every agentic deployment and will resurface in each of the sections that follow: privilege escalation, where an agent ends up holding more access than anyone intended to grant it; multi-step action chains, where a single task becomes dozens or hundreds of tool calls, each one a fresh chance for something to go wrong; stale authorization, where permissions granted for one context quietly persist into contexts where they no longer belong; and machine-speed consequence, where the damage from a bad decision outruns any human's ability to notice and intervene.

How the NIST standards ecosystem is responding to the agent governance gap

NIST has not ignored this. The agency's response spans several active workstreams, some finished, some in draft, some still ahead, and knowing which is which matters for anyone scoping a governance program around them.

What is published and usable now starts with NIST AI 100-2, updated in March 2025. That update names AI agents as a threat surface for the first time inside the AI 100-2 taxonomy, a category entirely absent from the 2023 edition, and it introduces specific categories for both direct and indirect prompt injection. Separately, NIST AI 100-5 covers NIST's plan for global engagement on AI standards. It is not an agentic AI profile, despite what the number might suggest to anyone skimming a document list.

Further along is a preliminary draft of NIST IR 8596, published December 2025, and it maps cybersecurity framework functions onto AI-specific risks, agentic threats among them. In February 2026 a federally affiliated cybersecurity center released a concept paper on software and AI agent identity and authorization, and it proposes policy-based access controls for delegation chains plus fine-grained authorization constraints on what tools an agent can reach.

Ahead of both sits the clearest signal of where NIST is heading. On February 17, 2026, NIST's Center for AI Standards and Innovation launched the AI Agent Standards Initiative, one of the first US government programs built specifically around interoperability and security standards for autonomous agents. The initiative is organized around three tracks: industry-led standards development, community-led open-source protocol work, and research into agent security and identity. A NIST AI Agent Interoperability Profile is planned for the fourth quarter of 2026.

Practitioners do not have to wait on NIST alone. The CSA MAESTRO framework supplies a threat modeling methodology built specifically for agentic systems. Where the AI RMF sets organizational governance functions, MAESTRO identifies what agents are actually at risk from and how those risks play out across multi-agent architectures. Singapore's IMDA launched its Model AI Governance Framework for Agentic AI on January 22, 2026, and updated it on May 20, 2026 with industry feedback covering multi-agent systems, third-party agents, and automation bias. A university research center built its Agentic AI Risk-Management Standards Profile explicitly to complement and extend the NIST AI RMF, organized around the same four functions. None of these frameworks compete for the title of best answer. They are the available scaffolding while NIST finishes building its own.

Govern: defining accountability before agents act, not after

The base Govern function assigns accountability to models and systems. That assignment made sense when one model mapped to one deployment with one clear purpose. Agentic deployments break that mapping immediately: accountability now has to reach individual agents, the delegation relationships between them, and the orchestration layer that coordinates them, at a level of granularity most governance programs have never had to operate at before.

A single model can power a dozen agents, and each one has its own tool access, its own data scope, and its own blast radius if something fails. A customer support agent with read-only access to a knowledge base poses a different order of risk than a DevOps agent with write access to production infrastructure, even if both run on the same underlying model. Governing the model tells an organization nothing useful about either one.

The fix the CSA Agentic Profile proposes is autonomy tier classification: each agent gets assigned a tier based on how much it can do without human sign-off, and each tier carries its own oversight obligations. That classification needs to exist as policy before a higher-tier agent ever goes live, because the oversight it requires has to be built in ahead of deployment, not bolted on after an incident forces the question. Every deployed agent needs a named owner, a documented purpose, and a risk tier of its own, even when five other agents share its underlying model.

Delegation needs its own policy layer. Govern has to specify, before anything ships, which agents may invoke which other agents, whether an orchestrating agent can hand out permissions it does not hold itself, and how the audit trail captures the full chain of actions rather than just the last step taken. The NCCoE concept paper's proposal for policy-based access controls on delegation chains gives this a concrete shape: fine-grained constraints on what each agent in a chain can actually touch. Treat the orchestration layer as its own governed system with its own accountable owner. If you govern each agent individually and assume the orchestration layer takes care of itself, you end up with gaps nobody notices until something breaks.

The sharpest version of this problem is privilege escalation. When an agent spawns a sub-agent, calls an API, or reaches into a data store on its own initiative, the access token issued to the human who started the task can get passed along, delegated, or quietly assumed in ways nobody signed off on. Govern policy has to settle, in advance, whether an agent may ever hold more permission than the human who initiated it, and what stops that permission from compounding as it moves down a delegation chain. Get these rules settled at the Govern layer, and the next question becomes practical: which agents in the organization are actually subject to them.

Map: inventorying agents and tracing how trust flows between them

Govern sets the rules. Map is where an organization finds out who is actually playing by them. The uncomfortable answer, in most organizations, starts with an inventory nobody has ever actually built. Before any meaningful Map-function work can happen, you need a current, accurate list of every deployed agent and every tool each one can call. Shadow agents and undocumented tool connections are the most common gap, and they leave the rest of the Map function governing only the agents the organization already knew it had.

Individual teams standing up agents without central review is one source of this problem. Undocumented connections to Model Context Protocol servers are another. Anthropic introduced MCP in late 2024 as a standard interface for agents to reach external tools, and within months hundreds of enterprise MCP server integrations had followed, each one a connection that may or may not have made it onto anyone's list. Teams that skip the inventory step and jump straight to writing controls end up governing a fraction of what they actually run. A living inventory records each agent's purpose, its tool and data access, its owner, and its risk tier, and gets updated continuously rather than compiled once and filed away.

Once the inventory exists, Map has to answer a harder question than "what does this agent do." It has to answer "what can this agent cause." The CSA Agentic Profile calls this tool-use risk modeling and action-consequence mapping: every tool an agent can invoke gets characterized by how reversible its effects are, how wide its access reaches, and what downstream systems it touches. A tool that sends an email and a tool that initiates a financial transaction can share the exact same API call structure and carry entirely different consequences. Action-consequence mapping asks, for every tool an agent can reach: if this call happens in this context, what is the worst plausible outcome, and can anyone undo it afterward?

Trust has to be traced the same way access is. If Agent A can call Agent B, and Agent B holds credentials to a financial system, then Agent A's effective access includes those credentials, whether or not any human ever explicitly granted them. Mapping where an agent's own authorized scope ends and where inherited or delegated access begins is what gives the authorization controls set at the Govern layer anything to actually enforce.

This same mapping exposes where prompt injection becomes dangerous rather than theoretical. NIST AI 100-2's March 2025 update introduces specific categories for direct and indirect prompt injection, the technique by which adversarial instructions buried in a web page, a document, or a tool's own output hijack an agent's reasoning and push it toward an action nobody intended. In a multi-agent chain, this gets worse with distance rather than better: a cascaded injection can be built to target an agent sitting far down the chain, so a query that looks harmless to Agent A arrives as a fully weaponized instruction by the time it reaches Agent N. Map-function work has to treat every external data source an agent reads as a potential injection surface and work out what a successful injection would cost at each point along the chain.

Measure: tracking agent behavior at the speed and scale agents operate

Map tells an organization where its risks live. Measure is supposed to tell it whether those risks appear in production, but the base Measure function assumes that pre-deployment testing and periodic monitoring are enough, and this is where the base RMF runs into its clearest limit. For agentic systems, they are not, because risk in an agent compounds across an entire session rather than surfacing in any single inference. Stale authorization and behavioral drift can both build quietly for weeks, and they cross a threshold before anyone was watching for it.

Periodic review was built for a world where the thing being reviewed does not change much between reviews. An agent's behavior can shift gradually, call by call, until it is operating well outside its original design intent, a drift that a monitoring cycle set to run weekly or monthly will miss until the accumulated effect is already large enough to matter. The same logic applies to audit logs: reviewing them after the fact works when the volume of actions is something a human can reasonably read. It stops working when an agent can take hundreds of consequential actions in a single minute, and oversight has to operate at something closer to the speed the agent itself is moving at, not the speed a quarterly audit was designed around.

Stale authorization follows the same pattern as drift. A permission granted for a specific task, in a specific context, does not expire automatically once that context changes, and an agent carrying forward access it no longer needs is a liability sitting quietly until something triggers it. None of this is reason to abandon the Measure function's core premise, that risk has to be tracked continuously rather than assumed away after a single evaluation. It is reason to rebuild what continuous means for a system that does not pause between decisions the way the frameworks of 2023 assumed it would.

Sources

  1. Agentic AI Governance: NIST Standards for Autonomous Systems
  2. NIST AI Risk Management Framework: Agentic Profile
  3. Agentic AI Governance: NIST Standards for Autonomous Systems
  4. Agentic AI Risk-Management Standards Profile - CLTC

More in AI Risk Frameworks