Analysis · all datasets

What all the data says, taken together

Every other page carries one dataset. This page carries the standing questions, asked across all of them — the CVE decade, 103 months of Patch Tuesday, the CWE type rankings, the AI-discovery and attacker evidence, and the breach series. It's rebuilt whenever a source is added; the build fails if a dataset isn't covered here.

Q1 — Are we seeing more breaches because of AI?

More breaches, yes. Because of AI — the data can't support that claim, and the timing actively argues against it. US data compromises stepped up 78% in 2023 and have held at record levels since (3,205 → 3,158 → 3,322). But that step-up landed before capable AI vulnerability tooling existed publicly — ITRC attributes it to mass supply-chain exploitation (one MOVEit campaign generated hundreds of downstream "breaches") and to notification laws surfacing more of what was always happening. No breach dataset attributes incidents to AI at all, and it's getting worse, not better: 70% of 2025 breach notices disclosed no attack vector whatsoever. The honest answer is that the breach data isn't instrumented to answer this question yet.

2023

The step-up predates the tools

The breach surge's inflection is 2023 — Big Sleep's first find was late 2024, XBOW mid-2025. Whatever caused 2023, it wasn't AI vulnerability discovery.

70%

The attribution vacuum

Most breach notices now say nothing about how the attacker got in. A dataset that mostly reports "unknown" cannot confirm or deny an AI effect.

1

Confirmed AI-driven breaches remain countable

Arup's $25.6M deepfake fraud is independently verified. The claimed AI-run espionage campaign (GTG-1002) shipped no IOCs. Individual cases exist; a dataset-scale signal does not.

Q2 — If breaches are rising, what are the actual reasons?

Ranked by how well the data supports each mechanism — documented first, asserted last.

MechanismEvidence classWhat the data shows
Vulnerability exploitation as entrydocumentedThe DBIR vector series is the cleanest trend in any of these datasets: ~5% → 14% → 20% → ~31% of breaches starting with an exploited vuln, #1 vector by 2026. Attackers moved toward the front door this site counts.
The exposure window wideneddocumentedMedian time-to-exploit collapsed to ~5 days (negative in 2026 — before a patch exists) while median time-to-patch rose 32 → 43 days. A growing gap between those clocks is more breaches, mechanically, with zero new attacker capability.
Supply-chain concentrationdocumentedOne exploited platform now yields hundreds of breach notices (MOVEit 2023; the five ≥100M-notice mega-breaches of 2024). Third-party involvement in DBIR breaches doubled to 30%, ~48% by 2026.
More complete reportingstructuralNotification laws and disclosure culture inflate the reported series over time — same confounder class as CNA expansion on the CVE side. Some of the rise is measurement, not intrusion.
AI attacker upliftassertedVendor and agency consensus (NCSC, Microsoft, Google) is "productivity, not new capability" — evolutionary. Plausibly a contributor to the speed story above; not isolated or quantified in any public dataset.

Q3 — What do the data actually tell us?

Supported by the data

  • Disclosure and breach counts both broke pattern, but on different clocks. Breaches stepped up in 2023; the CVE surge is 2024–2026; Microsoft's records are 2026. Different mechanisms, not one wave.
  • Disclosure grew ~7.5× in a decade; breaches roughly 3×. The found-vs-getting-through curves are not proportional — most new disclosure never becomes a breach.
  • The vuln→breach channel is widening. DBIR's vector series is the one place the two halves of this site measurably connect.
  • Speed changed more than volume or type. ~5-day exploitation, 43-day patching, stable CWE mix across eight years of Patch Tuesday.

What the data cannot answer

  • Whether AI has caused any measurable share of breaches — attribution is absent from every breach dataset (and 70% of notices carry no vector at all).
  • Whether AI drove the time-to-exploit collapse, or coincided with commodity exploit pipelines maturing.
  • Whether the 2026 disclosure records reflect AI-assisted research, tooling changes, or reporting shifts — the confounders are unquantified.
  • Net effect: AI demonstrably finds real bugs for defenders AND compresses attacker timelines. Which side nets out ahead is not answerable from public data in mid-2026.

The cross-dataset reads

Tempo

The AI signal is speed, everywhere it's real

Across the discovery corpus (Big Sleep, AIxCC), the attacker corpus (TTE collapse), and the breach series (vuln-exploitation #1), the same shape recurs: AI-era evidence concentrates on faster, not more or different. Volume rises trace to structural causes; the type mix is stable.

Found ≠ breached

The CVE firehose and the breach series are routinely conflated in public argument. A decade of both on one site shows they move on different clocks for different reasons — citing rising CVE counts as evidence of rising compromise is the specific mistake the combined data refutes.

32%

The measurement is degrading as the stakes rise

Only ~32% of 2025 CVEs got NVD enrichment; 70% of 2025 breach notices carry no vector. Both instruments are losing resolution exactly when the AI question needs them sharpest — the biggest obstacle to answering Q1 is data quality, not data absence.

How this page stays honest

These are standing questions, re-asked every time a dataset is added or updated — the build fails if a data source exists without coverage here. Every claim above traces to the site's data files, which are public in the Substrate repo. Where the answer is "the data can't say," that's the answer that ships. Sources → · Methodology →