Get a demo — 30 minutes →
← Back to blog
Anomity robot illustrating GTIG's AI Vulnerability Data: Half of 2026's AI-Software CVEs Hit Agent Frameworks
Research

GTIG's AI Vulnerability Data: Half of 2026's AI-Software CVEs Hit Agent Frameworks

TL;DR
  • On October 1, 2026, the Google Threat Intelligence Group (GTIG) published its analysis of vulnerability discovery and exploitation in the AI era, covering January 2025 to August 2026.
  • Monthly CVE disclosures doubled in 2026, from 5,045 in January to 10,740 in August. Only 0.23% of 2026 disclosures, roughly 1 in 431, were seen exploited in the wild.
  • Vulnerabilities found by AI skew dangerous: 50% lead to remote code execution, against 26% for everything else, and 58% rate Medium threat risk against 28%.
  • GTIG tracked 2,076 CVEs in AI software, more than 1,500 of them in 2026. 782 hit agent orchestration frameworks such as Flowise, Langflow, LangChain and Dify, 230 hit AI web apps, and 212 hit inference servers such as vLLM, Ollama and LiteLLM.
  • A further 97 sit under the frontier vendors themselves, and the vectors GTIG lists for them are coding-agent bugs: shell interpolation in CLIs, untrusted workspace configs, sandbox escapes through Git worktrees.
  • Most of that software is installed by individual developers and data scientists, not deployed by IT. GTIG's advice to move from mass-patching to threat-informed triage assumes you know where it runs. For AI software, most organizations do not.

Google Threat Intelligence Group's October 1, 2026 report on vulnerability discovery in the AI era has two headline findings. The first is that AI is changing which vulnerabilities get found: the ones AI finds are disproportionately the ones attackers want. The second, further down, is that AI software has become one of the fastest-growing sources of vulnerabilities in its own right.

This post is mostly about the second finding, because it changes a practical question for security teams. It is no longer only whether AI makes attackers faster. It is whether you know where your own AI software runs.

The overall picture

The volume numbers are stark. Monthly vulnerability disclosures rose from 5,045 in January 2026 to 10,477 in July and 10,740 in August. Exploitation rose too: 141 vulnerabilities were disclosed and exploited between January and August 2026, against 127 in all of 2025, for an average of 18 a month against 10.5. Zero-days made up 62% of exploited vulnerabilities this year, peaking at 22 in August alone.

Against that, only 0.23% of 2026's disclosed vulnerabilities, roughly 1 in 431, were ever observed exploited. That ratio is the core of GTIG's argument: patching everything is no longer feasible, and choosing what to patch first is the job.

What AI finds is worse

MeasureNot discovered by AIDiscovered by AI
Leads to remote code execution26%50%
Low threat risk69%39%
Medium threat risk28%58%
High threat risk3%4%

GTIG's illustration is CVE-2026-1731, an unauthenticated OS command injection in BeyondTrust's remote access products that a third-party research agent found on its own. Within four days of disclosure, GTIG observed one threat cluster exploiting it; within seven days, five more. The window between an AI-found bug becoming public and it being used is short, and it is the same window defenders have to find their exposed assets. We looked at the offensive side of this shift in autonomous hackbots and agent-layer visibility.

AI software is now its own vulnerability category

GTIG tracked 2,076 CVEs in AI-related software from January 2025 to August 2026, more than 1,500 of them this year. Orchestration middleware alone accounts for about half, and GTIG reports a 347% surge in its disclosures in 2026.

Category (GTIG)Examples GTIG names2026 count
AI orchestration and agent frameworksFlowise, Langflow, LangChain, Dify, LlamaIndex, AutoGen, CrewAI, MCP782
AI web apps and portalsOpen WebUI, AnythingLLM, LibreChat, RAGFlow, Gradio, Streamlit230
Inference and serving infrastructurevLLM, Ollama, LiteLLM, llama.cpp, Triton, Ray, LocalAI212
Model security advisoriesPrompt injection, guardrail bypass, system prompt exfiltration106
ML frameworks and hubsPyTorch, Transformers, ONNX Runtime, Safetensors, Keras99
Frontier model vendorsAnthropic, Gemini, OpenAI97
MLOps and experiment trackingMLflow, ClearML, Weights & Biases, Langfuse, LangSmith39
Vector databases and searchMilvus, Qdrant, ChromaDB, Weaviate, LanceDB19

Two rows deserve a closer look. The frontier vendor row is small, but the vectors GTIG lists for it are coding-agent bugs, not model bugs: command injection through CLI shell interpolation, implicit execution of untrusted workspace configs, sandbox escape through Git worktree confusion, and data exfiltration through injected Markdown images. Those land on developer endpoints.

And the named exploited examples are all familiar: CVE-2026-42271 in LiteLLM, a command injection in MCP preview endpoints that we covered in our LiteLLM MCP-preview RCE advisory; CVE-2026-5027 in Langflow, a path traversal file write that drops cron jobs or SSH keys onto the host; and CVE-2025-3248 in Langflow, unauthenticated code injection already in CISA's Known Exploited Vulnerabilities catalog.

Why AI software escapes vulnerability management

Look at the examples in the three biggest rows. Almost none of it arrives through procurement. Langflow and Flowise are started with a single command from a README. Ollama is a desktop install. A LiteLLM proxy is a container someone runs to share an API key with their team. Open WebUI sits in front of a local model on a workstation under a desk. The securing AI agent frameworks guide and the LLM gateways and proxies guide cover the controls for each; this is the step before them.

Vulnerability management matches advisories to assets. Threat-informed triage, which GTIG rightly recommends, ranks the matches. Both depend on an asset list. For conventional infrastructure that list exists. For AI software it usually does not, so a 782-advisory year in orchestration frameworks produces 782 advisories and very few matches, and the exposure is still there.

You cannot triage an advisory against software you do not know you run. For AI tooling, the inventory is the bottleneck, not the patch.

What GTIG recommends, and what it assumes

  • Move from mass-patching to threat-informed triage, combining targeted edge defense with automated, agentic remediation. This assumes an asset inventory that includes AI software.
  • Contain and sandbox autonomous agentic workloads. This assumes you know which agentic workloads exist and where they run.
  • Risk-based vulnerability management for AI infrastructure. This assumes the AI infrastructure is in scope to begin with.
  • Pre-release AI code review for software providers. This one is self-contained, and GTIG argues it could eventually slow disclosure growth. We cover one implementation in the claude-code-security-review guide.

How Anomity closes the inventory gap

The Endpoint Sensor inventories AI software on every managed endpoint: 144 tracked AI tools, plus the MCP servers, coding agents, CLIs and local LLM runtimes that make up GTIG's largest categories, with versions. That turns an advisory for Langflow, Ollama or LiteLLM into a list of machines in minutes rather than an investigation, and it includes the installs nobody filed a ticket for.

The Browser Sensor adds the hosted side, across 311 tracked AI web services, and cloud discovery adds the OAuth grants AI applications hold against Google Workspace and GitHub. Findings route to SIEM, Slack, email or Jira, so the AI inventory feeds the vulnerability process you already run rather than standing up a separate one. Anomity complements vulnerability management and EDR; it does not replace them.

GTIG's data says AI software is now a top-tier source of vulnerabilities, and AI-found bugs are more likely to be the dangerous kind. The response it recommends is sound. It starts with knowing what you run, which is the step most organizations skip for AI. For a practical walkthrough, see how to build an AI agent inventory. To see the AI software actually installed across your fleet, book a 30-minute demo.

Frequently asked questions

What did the GTIG report find?

Four things matter most for security teams. Monthly vulnerability disclosures doubled across 2026. Exploitation in the wild rose too, to an average of 18 exploited vulnerabilities a month in 2026 against 10.5 in 2025, though only 0.23% of disclosed vulnerabilities were ever seen exploited. Vulnerabilities discovered by AI are disproportionately severe, with half leading to remote code execution. And AI software itself has become a large and fast-growing vulnerability category, with 2,076 CVEs tracked from January 2025 to August 2026 and roughly half of this year's falling in agent orchestration frameworks.

Which AI software categories have the most vulnerabilities?

By GTIG's 2026 counts: AI orchestration and agent frameworks, 782, including Flowise, Langflow, LangChain, Dify, LlamaIndex, AutoGen, CrewAI and MCP; AI web apps and portals, 230, including Open WebUI, AnythingLLM, LibreChat and Gradio; inference and serving infrastructure, 212, including vLLM, Ollama, LiteLLM and llama.cpp; model security advisories, 106; ML frameworks and hubs, 99; frontier model vendors, 97; MLOps and experiment tracking, 39; and vector databases, 19.

Are AI-discovered vulnerabilities really more dangerous?

On GTIG's data, yes. Exactly 50% of AI-discovered vulnerabilities result in remote code execution, compared with 26% across the broader CVE ecosystem. 58% qualify for Medium threat risk, more than double the 28% baseline, while low-risk findings drop from 69% to 39%. GTIG's example is CVE-2026-1731, an unauthenticated command injection in BeyondTrust remote access products found autonomously by a third-party research agent: one threat cluster exploited it within four days of disclosure, and five more within seven days.

Which exploited AI-software vulnerabilities does GTIG name?

Three stand out. CVE-2026-42271 in LiteLLM, a command injection in MCP server preview endpoints that leads to host takeover and API credential theft. CVE-2026-5027 in Langflow, a path traversal file write in the upload handler that lets attackers drop cron jobs or SSH keys onto the host. And CVE-2025-3248 in Langflow, unauthenticated Python code injection that CISA added to its Known Exploited Vulnerabilities catalog in May 2025.

Why is AI software harder to patch than other software?

Because nobody deployed it. A web server or a VPN appliance arrives through procurement and lands in an asset inventory. Langflow, Ollama, Open WebUI or a local LiteLLM proxy typically arrives through pip install or docker run on a developer's laptop or a team's spare VM, started from a README. Vulnerability management works by matching advisories against known assets. Software that never became a known asset never matches, however good the triage process is.

What does GTIG recommend?

Its central recommendation is to move from unprioritized mass-patching to threat-intelligence-driven triage, combining targeted edge defense with automated, agentic remediation. For AI infrastructure specifically, it calls for immediate containment strategies, sandboxing autonomous agentic workloads, and risk-based vulnerability management. It also recommends that software providers run AI-enhanced code review before release, on the reasoning that if pre-release AI review becomes standard, disclosure growth could slow.

How does Anomity help with this?

The Endpoint Sensor inventories AI software across the fleet: 144 tracked AI tools, plus the MCP servers, CLIs, coding agents and local LLM runtimes that GTIG's largest categories describe, with versions. That is the asset list a GTIG-style triage process needs before it can prioritize anything, built from what is actually installed rather than what was procured. Findings route to SIEM, Slack, email or Jira, and Anomity complements existing vulnerability management rather than replacing it, by supplying the AI-specific inventory it has been missing.

Ask AI about Anomity
ChatGPT Claude Perplexity Google AI Grok