Est.
Shadow AILong read

Measuring Shadow AI Prevalence in Enterprise Workforces

Different measurement methods reveal vastly different shadow AI prevalence numbers.

Contributing Editor · · 12 min read
Cover illustration for “Measuring Shadow AI Prevalence in Enterprise Workforces”
Shadow AI · October 1, 2026 · 12 min read · 2,606 words

Security leaders are being asked to put a number on shadow AI, and the instruments they have cannot agree with each other. Shadow AI has moved from an emerging risk to a baseline operating condition in most enterprises, so the measurement question is no longer academic, it is operationally urgent. Yet the figures circulating in board decks and vendor reports contradict each other: an August 2026 audit from StationX found self-reported prevalence ranging from roughly half to four in five employees depending on how the survey was built, while the same underlying IBM figure keeps reappearing under different names across different databases.

That absence of noise is structure. Each measurement method captures a different slice of the phenomenon, and any organization that treats a single number as the answer to "how much shadow AI do we have?" is measuring something real, just not the thing it thinks it is measuring. Surveys measure belief and behavior as employees report it. Network telemetry measures traffic to known destinations. Data-movement monitoring measures what content actually left the building. None of the three measures the other two, and none of them, alone or combined, measures the fastest-growing category of the problem: autonomous agents acting without a human in the loop.

What follows is a dissection of each method, in order, tracing what it can see, what it cannot, and why the gaps do not close just because you are running more sensors.

Shadow AI's expanding surface area

The thing being counted has changed shape faster than the counting methods have. Older definitions of shadow AI assumed a single behavior: an employee opens a browser tab to an unapproved chatbot and pastes something in. That behavior still happens constantly, but by 2026 the surface had grown to include personal AI accounts used on corporate laptops, desktop assistants, coding copilots, models running locally on a machine with no network call to intercept, direct API calls to model providers, AI features quietly embedded inside SaaS tools the company already approved, autonomous agents, and MCP connections linking one system to another.

The growth in the SaaS layer alone tells the story. Netskope tracked a rapidly expanding count of distinct generative AI SaaS applications by mid-2025, up from roughly 317 at the start of the year, nearly a fivefold jump in a matter of months, with the average organization running meaningfully more genAI apps in May than it had three months prior. Grammarly, of all things, received more enterprise data than ChatGPT in 2025. That detail matters because it wrecks the assumption baked into most policies, which name the obvious offender (a chatbot) and ignore the grammar checker quietly ingesting every internal memo an employee drafts. The biggest data destination is rarely the tool a policy bothers to name.

Agents are the newer and stranger addition. A shadow SaaS app is passive: it sits there and receives whatever a human decides to paste into it. A shadow AI agent is not passive. It initiates connections to outside services on its own, executes code, and can hold onto credentials long after the task that spawned it is finished, which turns the risk from data leaving the building into something acting inside it. Methods built to catch a person typing into a browser were never built to catch a process that never asks permission to run.

Self-report survey measurements and their inflated numbers

Surveys are where most of the headline numbers come from, and they consistently produce higher prevalence figures than anything governance teams can independently verify. That gap is baked into how the survey questions get asked rather than reflecting a flaw in any single survey.

Denominators do most of the damage. Second Talent's analysis points out that Microsoft and Salesforce report figures as a share of employees who already use AI, dropping non-users out of the count, while BlackFog, Microsoft UK, and EY count every employee surveyed, users and non-users alike. Running those two approaches on similar populations produces very different-looking numbers, even before anyone lies to a pollster. That methodological split explains much of the distance between a figure like the 66% of office professionals PagerDuty found using unauthorized AI tools and the narrower shares reported elsewhere. Phrasing does the rest of the damage: ask about "unapproved tools" and you get one answer, ask about "banned tools" and you get another, from the same people, on the same day.

Citation decay compounds the confusion. StationX's own provenance audit found that three of the nine shadow AI figures in its database, about a third, could not be traced back to whoever actually produced them, and two more turned out to be the same IBM figure wearing a different label. The numbers that get repeated most often are the ones that survived the copy-paste chain intact, not necessarily the ones closest to reality.

None of this makes surveys useless. They are the only instrument that captures intent and context, the gap between what employees know about policy and what they actually do. Second Talent, citing UpGuard's State of Shadow AI report, notes that employees who said they understood the security rules were more likely to use unauthorized tools regularly, a finding no firewall log could ever surface. A survey cannot tell leadership which specific tool an employee used, what data went into it, or whether the use was a one-off or a daily habit. What it gives leadership is a behavioral baseline, a sense of appetite and awareness across the workforce, a starting point for writing policy rather than a technical inventory of exposure.

What network telemetry and CASB tools can see

Network telemetry is the instrument most security teams reach for once they stop trusting survey numbers, and it earns that trust only partway. Gateway and network tools see traffic headed to known AI domains, but they cannot tell a sanctioned enterprise login from an employee's personal account, cannot read anything inside encrypted traffic, and cannot see a model running locally on a laptop with no network call to catch.

A gateway only sees what someone bothered to route through it, and the network layer records hostnames, not the people behind them, not the prompts they typed, not whether the interaction was approved. That means telemetry answers one question well: did someone in this organization contact an AI endpoint? It does not answer whether that contact was governed. Teramind's report adds that AI traffic increasingly runs encrypted through the browser, which makes plain URL blocking both too blunt and too fragile: binary allow-or-deny rules ignore the context, the user's role, the sensitivity of the data, and the business reason that determines whether an interaction is risky.

CASB tools inherit the same blind spot at the credential layer. They can flag a visit to a known AI domain, but if that visit ran on a personal login instead of a corporate one, the whole interaction disappears from governance controls, functionally invisible even though it happened in plain sight on the corporate network. That is not a hypothetical crack in the glass. StationX's research, drawing on LayerX telemetry, found that roughly half of enterprise AI conversations run on personal identities even inside organizations that have already rolled out sanctioned AI tools. Buying the enterprise license does not stop employees from logging in with their own.

None of this makes telemetry a bad investment. It produces a population-level count of AI traffic that no survey can replicate, it surfaces AI domains nobody had cataloged, and it can flag data volumes that look abnormal against a baseline. It is the right tool for answering "what endpoints are employees reaching," and the wrong tool for answering what they are actually doing once they get there. Data-movement signals pick up exactly where that question is left hanging.

What data-movement and endpoint signals reveal

Browser-layer DLP and endpoint monitoring shift the question from whether a session happened to what moved inside it, and that shift is the difference between measuring activity and measuring exposure. It is the first method in this survey that puts a number on what actually left the organization's control, rather than just confirming that a door was opened.

Airia's analysis of Cyberhaven data found that a meaningful share of what employees paste into AI tools qualifies as sensitive, with source code the single most common category leaked. Samsung's own well-known incidents involving engineers pasting proprietary code into a public chatbot are the template case everyone in this field already has in mind, and the Cyberhaven finding says that template is closer to a norm than an outlier. Teramind's 2026 report extends the picture past copy-paste behavior: employees upload whole files, PDFs, slide decks, spreadsheets, logs, code repositories, and those uploads routinely carry proprietary IP, financials, or personal data that a tool watching only pasted text would never catch. Customer records buried in a spreadsheet do not announce themselves the way a pasted code block does.

Data-movement monitoring answers a sharper question than network telemetry ever could: not which AI endpoints are being reached, but what information left the perimeter and what kind of information it was. That specificity is valuable and expensive to get right, and it still has holes. Browser-layer DLP cannot see a model running locally, since there is no network traffic to intercept in the first place. It cannot see AI features baked into an approved SaaS product, because that traffic heads to a sanctioned domain and gets filtered out before anyone looks twice. It cannot see anything happening on a personal device that never joined the managed fleet. Even the tools built specifically for this job have calibration limits: StationX notes that Cyberhaven's own testing of a free Presidio-based detector caught only a fraction of sensitive items straight out of the box, improving only after deliberate tuning. Purpose-built detection is not automatically accurate detection.

Stacking a survey, a network monitor, and a DLP tool on top of each other makes the coverage look close to complete, but gaps remain.

Why the measurement gap persists across all three methods

Running all three methods at once does not add up to a full picture. It produces three partial pictures, drawn from different populations, over different time windows, using different working definitions of what counts as shadow AI in the first place. A survey samples people. Telemetry samples traffic. DLP samples content. None of the three samples the same universe, so stitching their outputs together does not triangulate a true number, it just produces a number that feels more authoritative because three sources agree to disagree in the same report.

No government agency or national statistics office measures shadow AI, academia has produced roughly one small study on the topic, and every figure carrying real scale comes from vendor research. Companies selling AI detection products are the ones generating the data on how big the AI detection problem is. That does not mean the numbers are fabricated, but the measurement infrastructure has a commercial stake in the problem looking large and urgent. IBM's breach-cost figures, among the most widely quoted in the industry, come specifically from organizations that already suffered a breach, a narrow slice rather than a representative one of enterprises generally. StationX makes the point directly: that data describes companies having a bad year, not the market as a whole.

Read together with everything above, that is not a statement about which vendor an organization picked. It is a statement that the available tools cannot deliver full coverage no matter which combination you buy, because each one was designed to see a different slice of the surface.

One genuine complication cuts against the doom-and-gloom framing. Netskope telemetry, cited by StationX, found the share of the enterprise population using personal SaaS genAI apps declining as sanctioned alternatives expanded, the first measured evidence that governance investment displaces shadow use rather than just running alongside it. Displacement means measurement programs need to start tracking behavior over time, not just a single prevalence snapshot, because the presence of a sanctioned tool appears to change behavior rather than simply adding another data point to ignore.

How agentic AI breaks the measurement assumptions

Every method surveyed so far assumes a human is the one initiating the AI interaction, someone typing into a browser, someone pasting a file, someone answering a survey question about their own behavior. Autonomous agents remove that assumption. They act persistently on their own, initiating connections, executing code, and holding credentials without a person clicking anything in the moment.

Surveys cannot capture agent behavior because the employee being surveyed may not know an agent is even running, may not think to classify it as "an AI tool I use," or may have set it up through a developer workflow that never crosses the policy-awareness questions a survey is built to ask. Network telemetry fares no better. Tools tuned to catch human-initiated browser sessions will not reliably flag traffic originating from application servers, CI/CD pipelines, or MCP connections, since that traffic reads as ordinary system-to-system communication rather than anything resembling shadow AI. Airia's analysis found that MCP adoption and agentic AI architecture have opened an entire new tier of ungoverned risk that most existing security stacks were never built to handle.

The market has started reacting. Checkmarx launched an AI Inventory capability in June 2026 built specifically to identify models, agents, MCP servers, AI libraries, and SDKs sitting inside application repositories, which amounts to an admission that agent discovery needs its own instrument, distinct from whatever caught shadow SaaS apps a year earlier. The reason that distinct instrument matters comes down to what an unmanaged agent can actually do. A shadow SaaS app just sits and receives data someone feeds it. A shadow agent can persist access credentials, open outbound connections with no human triggering them, and take actions in external systems on its own. That moves the risk from data leaving quietly to systems acting without anyone watching, which is a different category of exposure and needs a fourth kind of measurement: real-time behavioral monitoring, catching the action as it happens rather than reconstructing it from a log afterward.

What a defensible measurement program requires

A single number will never describe shadow AI honestly, because the phenomenon itself does not fit inside one instrument. A defensible program starts by naming what each method is actually for. Surveys establish workforce sentiment and the size of the knowledge-behavior gap: how many people know the policy and ignore it anyway. Network telemetry establishes population-level traffic to known AI destinations and flags domains nobody had catalogued. Browser-layer DLP and endpoint monitoring establish what content, source code, financial data, customer records, actually crossed the perimeter. Runtime behavioral monitoring, still the least mature of the four, is the only method built to catch an agent acting on its own rather than a human typing into a box.

None of these four methods substitutes for another, and none of them, run in isolation, deserves to be quoted as "the" shadow AI number in a board presentation. A program with any rigor reports all four as separate lines, tracks each one over time rather than as a single snapshot, and pays particular attention to displacement, whether sanctioned tools are pulling usage away from personal accounts, since that trend line says more about whether governance is working than any static prevalence figure ever could. Treating the four instruments as complementary lenses on the same object rather than four attempts to reach the same number still leaves the picture incomplete. It will at least be honest about which parts are missing, and honest incompleteness beats a confident number that was never measuring what it claimed to measure.

Sources

  1. Shadow AI Statistics 2026: Adoption, Cost, and Real Risk
  2. Shadow AI Statistics: Key Data Points Every CISO Needs in 2026 | Airia
  3. Shadow AI Report 2026 - Teramind
  4. Top 50 Shadow AI Statistics 2026: Who Uses It and What It Costs - Second Talent
  5. PagerDuty Report Finds Two-Thirds (66%) of Office Professionals Have Used Unauthorized AI Tools at Work | PagerDuty
  6. State of Shadow AI 2026: ShadowLock Research Report
Filed underShadow AI

More in Shadow AI