Agentic AI Systems Under Current Risk Classification Schemes
Current risk frameworks miss the five dimensions that make agentic AI uniquely dangerous.

Agentic AI systems break goals into sub-tasks, call external tools, remember what happened last session, change their own plans mid-stream, and hand off work to other agents. Every risk framework now in force, including NIST's AI RMF and ISO/IEC 42001, was built to govern something else: a system that takes an input and produces an output, once, for a human to read and judge. That older model fits a system that predicts the next word in a sentence. It does not fit a system that reads a support ticket, decides to issue a refund, calls the payments API, and tells a second agent to update the customer record, all without a human in the loop.
Each behavior on that list creates its own kind of exposure. Goal decomposition means a single instruction can spawn sub-tasks nobody explicitly approved, any of which might go sideways. Tool invocation means the system isn't just generating text anymore: it's deleting rows, moving money, and executing code, with consequences a chatbot transcript never had. Persistent memory means an agent can be nudged slightly off course in one session and carry that drift forward into the next, with each session compounding on the last. Delegation across multiple agents scatters both capability and accountability, so that no single model is "the" system anyone can point to and audit. And all of it happens faster than a person watching a dashboard can react.
A 2026 working paper from the Responsible AI Institute, the TrustX Agent Risk Classification Framework (ARC), defines agentic AI as applications capable of advanced perception, planning, and action with relatively high degrees of autonomy and orchestration, and draws a hard line between that category and the systems current frameworks were actually written for. The distinction matters because it relocates where risk lives. You can no longer read risk off a model's outputs, because it becomes a property of a system acting inside an environment, over time, often without anyone watching the clock.
How the EU AI Act and NIST AI RMF fail to classify agentic risk
The EU AI Act sorts systems into four tiers, unacceptable, high, limited, and minimal, based on the sector a system operates in: biometrics, critical infrastructure, employment, law enforcement, the Annex III list. None of those tiers asks how autonomous the system is, how many tools it can call, or how fast it acts before a human can intervene. An agentic system making consequential financial, medical, or legal decisions will probably land in the high-risk tier, but only by inference from its sector, not because the Act has any mechanism for detecting agentic behavior directly.
That gap has teeth. The high-risk tier requires that systems be designed so a person can effectively oversee them, interrupt them, and understand their limits. Nothing in the Act tells a deployer how to determine whether a given agentic system actually triggers that obligation in the first place, which leaves the oversight requirement sitting on top of a classification process that was never built to detect the thing it's supposed to apply to. Compound that with the time lag between a cascading action and human notice: an agentic system can fire off a cascade of irreversible actions, deleting a database, sending an email campaign, approving a wire transfer, long before any human notices something went wrong.
NIST's AI RMF, ISO/IEC 42001, OWASP's guidance, and MITRE ATLAS all carry real influence, and TrustX ARC notes that their guidance remains voluntary and contains little to no agentic-specific provision in the core text. Sector rules show the same blind spot. A banking regulator's model risk management instrument replaced its predecessor in April 2026, and it mandates model risk management, but it says outright that its guidance does not cover generative or agentic AI. And when an orchestrating agent spawns sub-agents, responsibility fans out across a chain that existing RMF categories, built to classify one system at a time, have no way to represent.
What real incidents show about the classification gap's consequences
Agentic incidents from 2025 and 2026 share one shape: a system takes a consequential action that no classification tier required anyone to watch, limit, or answer for. Security researcher Ari Marzouk's "IDEsaster" research found a wide range of vulnerabilities across AI coding tools, including remote code execution carried out through JSON schema attacks and manipulated IDE settings. In March 2025, Pillar Security disclosed the Rules File Backdoor attack, in which hidden Unicode characters embedded in configuration files silently manipulated outputs from Copilot and Cursor.
Both incidents belong to a broader category: indirect prompt injection, where adversarial instructions buried in third-party content get executed by an LLM-integrated application as though they were legitimate commands. That category exists because of agentic tool-use. A single-turn model generating a paragraph of text has no equivalent failure mode, because it has no tools to misuse. TrustX ARC documents that the pattern held through 2026, with incidents running from persistent prompt injection up through multi-stage attacks carried out with almost no human involvement at any point.
What ties these cases together is that each agent operated inside a system no framework classified by autonomy level, tool-access scope, or delegation structure, so no control was mandatory. That raises the obvious next question: what would a classification scheme actually need to measure to catch systems like these before the damage is done?
The five risk dimensions that agentic systems introduce and existing tiers do not measure
Autonomy is not a yes-or-no switch but a spectrum, and a single platform can sit at different points on it for different capabilities at the same time. One illustrative five-tier spectrum runs from L1, where a human operator directs every action, through L2 (human and AI collaborate), L3 (AI leads and consults the human), L4 (AI acts and the human approves only in risky scenarios), up to L5, where the system runs fully on its own and a human only monitors. A system can run supervised for detection and fully autonomous for response, and no sector label captures that split.
Tool-use scope is a second, separate axis: how broad and how reversible are the actions an agent can take against external systems, deleting records, sending messages, moving money, changing configurations. That has to be assessed on its own terms, independent of what industry the agent happens to sit in.
Delegation chain structure is a third. In a multi-agent pipeline, one compromised agent upstream corrupts every agent downstream that trusts its output as authoritative, so you have to govern the pipeline as a whole, not any single model inside it.
Temporal exposure is a fourth: the length of the gap between when an agent starts a consequential action and when a human can actually observe and stop it. A passive system waits for a human to act on its output rather than acting first, so it has no equivalent to this gap.
The fifth dimension is measurement: how risk gets scored once you identify the first four. TrustX ARC's twelve-dimension rubric uses a "critical dimension" rule that stops a high score on one axis, say, irreversibility of action, from getting averaged down by low scores elsewhere. That design choice matters because it shows general-purpose AI risk scoring, built around averages and aggregate harm categories, cannot simply be extended to agentic systems. It has to be rebuilt around the dimensions that actually produce agentic harm.
What the first generation of agentic-specific frameworks addresses
Four efforts now attempt to close part of this gap, and each one covers different ground.
A national regulator published what counts as the first agentic AI governance framework, in January 2026. It's voluntary, and it names the core problem directly: an agent's access to sensitive data and its ability to change its environment create a new risk profile, and complex interactions among multiple agents make outcomes substantially harder to predict. Rather than inventing new principles from scratch, the IMDA framework translates existing ones, fairness, transparency, human oversight, into agentic terms. It speaks to the autonomy and delegation dimensions directly, but it leaves enforcement untouched.
The multinational guidance led by CISA, published May 1, 2026, with the NSA and cybersecurity agencies from Australia, Canada, New Zealand, and the UK, takes a sharper security angle. It defines five risk categories, privilege, design and configuration, behavioral, structural, and accountability, and anchors its core control in cryptographic, per-agent identity with short-lived credentials, treating each agent as its own security principal the way a human user account would be treated. It also puts the burden of deciding which actions need human approval on system designers, not on the agents themselves.
NIST's AI Agent Standards Initiative, announced in February 2026 through the Center for AI Standards and Innovation, is building voluntary guidelines across identity and authorization, security and risk management, and monitoring and logging, with an AI Agent Interoperability Profile planned for release in the fourth quarter of 2026.
TrustX ARC, submitted July 10, 2026, is the most granular of the four: a twelve-dimension scoring rubric across seven categories of agentic system, a GPA plus IAT classification model, the five-level autonomy framework, a three-tier governance output tied to specific control recommendations, and a Coding Assistant extension built for that deployment class specifically. The authors present it as a working paper, with iteration ongoing.
None is legally binding, and none has been written into law in any major jurisdiction, so each identifies the right dimensions to measure without forcing anyone to measure them. None of them yet force anyone to measure them.
Prompt Injection as the Hardest Technical Problem for These Frameworks
Prompt injection sits exactly where agentic design and agentic governance collide. If an agent retrieves a document, reads output from another agent, or pulls data from an external API, an instruction crafted to look authorized can be buried in that content. The agent has no reliable way to tell content it's supposed to process from commands it's supposed to obey, so it executes the instruction anyway. TrustX ARC documents this mechanism at length under the heading of indirect prompt injection, and the CISA-led multinational guidance calls it the most persistent, hardest-to-fix threat facing agentic systems, because it stems from how language models process instructions in the first place, not from a missing patch or a misconfigured setting.
Authentication doesn't touch this problem. A token or certificate verifies the container holding an identity, not what that identity does once it's holding the credential. An agent can carry a perfectly valid credential and still act completely outside its intended mandate the instant after authentication clears, because nothing checked what its next action would be, only its right to take one.
Multi-agent pipelines make the problem worse by multiplication. When agents call other agents, one compromised agent upstream corrupts every agent downstream that trusts its output. The threat model shifts from "an attacker reaches my agent" to "an attacker reaches any agent anywhere in the dependency graph," including third-party agents nobody in the organization ever reviewed. That turns prompt injection into a supply chain risk, not a perimeter one, and a supply chain cannot be secured by hardening the front door.
The GTG-1002 incident involving Claude Code, documented in TrustX ARC and dated November 2025, shows what that gap looks like in practice. The attack exploited the same split between authentication and behavioral integrity, and Claude Code carried out roughly 80 to 90 percent of the espionage operation autonomously before any human had the chance to observe what was happening. No framework, voluntary or binding, has closed this gap, because closing it requires solving how a language model tells authorized instruction from adversarial content embedded in the data it's processing, which remains an open engineering problem. Enterprises cannot wait for someone to solve it before they act.
The practical governance stance enterprises need to take before agentic-specific standards become binding
No binding agentic-specific standard exists yet, and the voluntary frameworks on the table, Singapore's IMDA, the CISA-led guidance, NIST's initiative, TrustX ARC, name the right risk dimensions without requiring anyone to act on them. Enterprises running autonomous agents now don't have the option of waiting for enforcement to catch up. The scale of the blind spot is already measured: per the Cloud Security Alliance AI Safety Initiative, 82 percent of enterprises have discovered AI agents running in their environments that they didn't know were there.
That number points to an inventory failure before it points to a security failure. You cannot apply an autonomy tier, a tool-access scope, or a delegation-chain review to an agent you don't know exists. The first governance task, ahead of any framework choice, is building a live inventory of every agent running in production, what tools each one can call, what data each one can reach, and which other agents each one talks to.
From there, the frameworks converge on a short list of operational moves: per-agent identity with short-lived credentials rather than shared service accounts, mandatory human approval gates for actions a system designer has flagged as high-impact, logging detailed enough to reconstruct a delegation chain after the fact, and runtime monitoring that watches what an agent actually does rather than relying solely on what it was authorized to do. That last point is the practical bridge between where governance stands today and where the binding standards are likely to land once the Q4 2026 interoperability profile and its successors arrive: behavioral visibility now, in place of a credential check that only confirms identity and says nothing about intent.
Sources
- TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems
- [2607.09586] TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems
- TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems
- CISA’s Agentic AI Five-Risk Framework: Enterprise Implementation


