Security Stagflation

By Sean Atkinson, Chief Information Security Officer (CISO) at Center for Internet Security® (CIS®)

ciso blog graphic

The vulnerability account is overdrawn, and the currency is losing value. The credit market is freezing up all at once for a single reason: security has always relied on debt. No system is perfect, and we often assume that adversaries will take time to discover the next bug. But those assumptions have now expired.

Our vulnerability account is in the red. The Common Vulnerabilities and Exposures (CVE) system serves as the currency we used to assess this debt, and it is diminishing in value in real time. The credit market that allows us to manage this debt safely is tightening. Artificial intelligence (AI) has reduced the cost of identifying bugs, while the expense of resolving them remains unchanged.

We are entering a period of security stagflation. Vulnerability signal is inflating, while remediation productivity remains comparatively stagnant. The disclosure and triage system is the credit channel transmitting that imbalance into an accumulating backlog of security debt. Software quality isn't deteriorating. The problem is capacity. Our ability to find flaws has been structurally transformed, and our ability to validate, coordinate, and deploy fixes has not.

The Asymmetry that Started It

Finding a bug used to cost about the same as fixing one. Both discovery and remediation required skilled professionals who could carefully read and understand code. Discovering vulnerabilities was challenging, and fixing them was equally difficult. The costs of both processes remained roughly in balance because the same talent pool was needed for each.

In the 2025 AI Cyber Challenge final, seven autonomous systems analyzed more than 54 million lines of code, discovered 54 of the competition's 63 synthetic vulnerabilities, and surfaced 18 additional real-world flaws at an average cost of roughly $152 per completed task. The asymmetry is visible inside the same result set: of the 18 real vulnerabilities found, 11 received patches. Even in a controlled harness built to reward end-to-end automation, discovery outran remediation.

AI is also beginning to accelerate patch generation, but patch generation is not the same as remediation. Engineers still must validate the proposed fix, test for regressions, coordinate disclosure, schedule releases, deploy the change, and confirm that downstream users actually adopted it. Discovery and patch drafting are becoming elastic; validation, governance, and deployment remain bounded by human and organizational capacity.

FIG01

FIG.01: The discovery–remediation asymmetry. Discovery and remediation once cost about the same; after the AI inflection, the cost of discovery collapses toward zero, while remediation holds at its human-coordination floor. The widening shaded gap is the engine of everything that follows.

The Inflation

A CVE is not valuable because it is rare. It is valuable because it gives vendors, scanners, defenders, regulators, and intelligence systems a common identifier for the same condition. What inflation erodes is not the identifier itself but the analyst attention and contextual intelligence available for each record.

That dilution is accelerating. The mid-year update from Forum of Incident Response and Security Teams (FIRST) revises the 2026 projection to approximately 66,000 CVEs — averaging 181 disclosures per day — after actual volume ran 46% above the February forecast. FIRST attributes the surge to three drivers: AI-assisted discovery, a 449% year-over-year increase in GitHub Security Advisory volume, and a 3,119% increase in VulnCheck's activity as CVE numbering authority of last resort, absorbing a backlog of previously unassigned vulnerabilities. Not all of that is new risk; some is reclassification of old risk into a countable form. The defensible conclusion is narrower and more uncomfortable: intake is rising faster than the ecosystem's ability to validate, enrich, prioritize, and remediate.

FIG02

FIG.02: Currency debasement as a supply shock. Demand for the CVE-as-signal stays roughly fixed, while the discovery supply curve floods rightward at near-zero cost, dragging the equilibrium from scarce-and-valued (E₀) to abundant-and-debased (E₁). The fall in value per CVE is the inflation, the same currency, printed without limit.

The central clearinghouse has already changed its operating model. Effective April 15, 2026, NIST announced that the National Vulnerability Database (NVD) enrichment would prioritize vulnerabilities in the Known Exploited Vulnerabilities (KEV) catalog maintained by the U.S. Cybersecurity and Infrastructure Security Agency (CISA) with a target of enrichment within one business day along with software used within the federal government and critical software as defined under Executive Order 14028. Everything else remains listed but is categorized as lowest priority, not scheduled for immediate enrichment. Unenriched backlog records published before March 1, 2026, were moved into that category, KEV entries excepted. NIST cites a 263% increase in CVE submissions between 2020 and 2025, with the first quarter of 2026 running roughly a third above the prior year.

The number worth sitting with is not the policy change but the productivity figure behind it. NVD analysts enriched nearly 42,000 CVEs in 2025, about 45% more than in any previous year, and still lost ground. That is what stagflation looks like from inside an institution: record output, falling coverage. The vulnerability ecosystem is moving from comprehensive rating toward explicit triage. "Critical" used to have a specific meaning. When everything is deemed critical, nothing truly is.

The Credit Crunch

No team patches every vulnerability, and no team ever has. Instead, we carry the gaps as security debt, aiming to pay it down on a schedule we hope will outpace our adversaries. The disclosure ecosystem functioned like a credit market, making this debt manageable. Bug bounties helped price the risk involved. NVD enhanced and rated these obligations, while the CISA KEV list acted as a margin call, indicating that certain debts needed immediate settlement.

The disclosure market is showing visible signs of congestion. For instance, as reported by The Register, cURL ended its paid HackerOne bounty after low-quality and AI-assisted reports overwhelmed the economics of human triage although it continued to accept disclosures through other channels. Similarly, HackerOne paused new submissions to its Internet Bug Bounty (IBB) program as the imbalance between discovery and maintainer remediation widened. NVD shifted toward risk-based enrichment rather than universal enrichment. At the same time, Google’s M-Trends 2026 report estimated mean time to exploit at negative seven days relative to patch availability. In other words, exploitation is increasingly observed before a patch exists. The debt is being called before the ledger entry is complete.

FIG_03

FIG.03: The disclosure market stops clearing. Human triage capacity is effectively vertical — it cannot scale — so when AI-generated reports shift demand sharply right, the volume that would clear at a sustainable price overruns capacity. That excess demand is the backlog that forces programs to shut their doors (cURL, HackerOne IBB). The crunch is a quantity problem the prevailing price can no longer absorb.

Why Both at Once Matter

Classic stagflation combines rising prices with stagnant output. The security equivalent is signal inflation combined with stagnant remediation productivity, and unlike the inflation side, the stagnation side is measurable. Verizon's 2026 Data Breach Investigations Report (DBIR) found the median organization facing roughly 50% more critical vulnerabilities than the year before, while mean time to full resolution moved in the wrong direction from 32 days to 43. In the same report, exploitation of vulnerabilities became the most common initial access vector for the first time in the report's history. Volume rose, remediation slowed, and the consequence showed up in breach data.

That is why the usual responses fail. Producing more findings worsens the signal problem. Simply adding scanners or telling teams to patch faster does not change the coordination and deployment constraints. The production process itself must change how findings are validated, how applicability is proven, how attack paths are interrupted, and how scarce remediation capacity is allocated.

The most serious objection comes from the same source as the volume figure. FIRST's mid-year analysis notes that when 2026 disclosures are filtered against actual exploitability, KEV membership or an EPSS score above 10% only about 7% clear the threshold, and the actionable patching burden has stayed close to flat. Their conclusion is that teams practicing exploitability-based triage should not need to scale headcount with raw CVE counts.

That objection is largely right, and it is not a rebuttal. It is the same argument arriving from the other direction. If 93% of the disclosure stream is background noise, then the scarce resource was never patching throughput; it was the capacity to distinguish. Stagflation in this system is not a shortage of fixes but a shortage of trustworthy triage, and it binds hardest on organizations that still measure themselves against total CVE count.

FIG04

FIG.04: One root cause forks into two failure legs that arrive together. The credit crunch is best read as the transmission mechanism that turns the asymmetry into a system-wide, non-clearing market, which is why neither monetary nor fiscal analogues (print faster / scale harder) resolve it.

Do We Get to Shift Left out of This?

Project Glasswing represents the optimistic case: use AI to find defects while code is still under the vendor's control. Anthropic's initial report describes substantial expansion and also that under 1% of discovered vulnerabilities have been patched. The best-case version of this still produced a backlog. That is not an argument against shift left. It is the clearest available evidence that discovery and remediation have decoupled.

Shift-left remains necessary, but it is not sufficient. Every shift-left investment needs a shift-right counterpart: runtime compartmentalization, trustworthy asset and component inventories, exposure management, attack-path analysis, and the ability to determine applicability within minutes.

We are now in the proof race. For every advisory, a mature organization should be able to answer five questions quickly:

  • Is the affected component present?
  • Is the vulnerable function reachable?
  • Is the interface exposed?
  • Are the required exploit conditions present?
  • Is a compensating control already breaking the path?

A defender that can prove non-applicability or isolate the exposed path can neutralize much of the negative-seven-day window without pretending it can patch everything first.

Two key realities impact the race against attackers:

  1. You must fix every vulnerability on your surface and deploy the necessary fixes. However, an adversary needs only to find one vulnerability you have not yet addressed and be ready to exploit it.
  2. For vendors, every patch released effectively discloses information about your systems. Patch-diff engineering treats your update as a labeled before-and-after pair, enabling AI to identify differences more quickly than your team can deploy fixes. In a “negative-seven-day world,” releasing updates teaches attackers which timeframes matter.

Therefore, while you should shift left, this strategy should apply only to code you control, can monitor, and can deploy faster than any disclosure can leak. While shift-left changes involve addressing security issues, they do not alter the imbalance between creating and addressing vulnerabilities. Each shift-left investment requires an equivalent shift-right investment, such as runtime compartmentalization, rapid asset verification, and the ability to determine within minutes whether a security advisory affects you. You may not win the patch race, but you can succeed in the proof race.

Are These New Vulnerability Families?

Three tiers, each failing our taxonomies in different ways.

First, the long tail collapsed into the discoverable. The 27-year-old OpenBSD bug, which includes flaws that have survived millions of automated tests, is not new at the Common Weakness Enumeration (CWE) level. It involves a signed-integer-overflow condition in OpenBSD's TCP SACK implementation allowing a remote attacker to crash any host responding over TCP, leading to a Denial of Service (DoS). The cost of discovering these issues has significantly decreased, making the entire long tail of vulnerabilities economically accessible all at once. From an actuarial perspective, this long tail behaves like a newly identified family of vulnerabilities for analysis given its distinct statistics. There’s less clustering around commonly used code, more presence in hardened components, and a poor correlation with the heuristics that defenders typically use to prioritize threats. Historical base rates are outdated, and the tail of the loss distribution has become even fatter. Consequently, the previous assumptions and priors no longer hold.

Second, the combinatorial chain became a first-class object. Mythos’s distinguishing capability is chaining many individually low-severity flaws into one high-severity path. CWE describes atomic weaknesses, and MITRE ATT&CK describes attacker behaviors. The synthesized chain lives in the seam between atomic weakness taxonomies and behavioral taxonomies. Attack graphs and exposure-management tools can represent portions of this composition, but the ecosystem still lacks a standardized, portable object that identifies and scores the chain as a whole.

A compositional vulnerability chain can be defined as a typed sequence of individually low-severity conditions whose combined impact is non-additive. CVSS scores the atoms; the missing measurement is the composition gap between the most severe individual condition and the expected impact of the complete path. AI makes these chains operationally important because it can search combinations that humans rarely have the time to enumerate.

FIG05

FIG.05: The synthesized chain is your most original unit of analysis. Each node is individually low-severity; composed, they reach a critical path that CVSS never scores. It sits in the seam — CWE scores the atoms above, MITRE ATT&CK names the behaviors below, and neither captures the composition.

Third, AI-native classes arrived. As AI increasingly writes code and identifies bugs, we encounter failure modes specific to probabilistic and agent-based systems. This includes plausible-but-wrong generated cryptography, hallucinated dependencies in the build chain, and prompt injection as a new class of vulnerabilities. Additionally, we see confused-deputy patterns in the use of agentic tools and trust-boundary failures in the context of model-protocol plumbing.

Mythos has already provided us with an example. In a controlled red-team evaluation, an earlier Mythos version was instructed to attempt a sandbox escape and contact a researcher. It found an unintended path to broader internet access and took additional public-facing actions beyond the requested task. The vulnerability was not due to a buffer overflow; instead, it stemmed from emergent agentic overreach, a new category of risk, one for which we are currently ill-equipped to identify.

AI did not create new classes of defect. Rather, AI has made historically rare vulnerabilities economically discoverable, operationalized the combinatorial chain of code, and co-evolved with AI-generated code to produce classes of issues that our existing taxonomies cannot yet adequately describe.

There is an actuarial consequence here that deserves its own treatment. Every historical loss and vulnerability dataset we calibrate against was generated under a discovery regime that no longer exists back when the arrival of new flaws could reasonably be treated as something that happened to a codebase rather than something an actor decided to cause. Once discovery effort becomes a dial someone can turn, the count processes underneath most cyber risk models, and the frequency inputs underneath FAIR and Gordon-Loeb, are calibrated against a world that has moved. Lloyd's has already called cyber rate adequacy marginal with a weakening forward trajectory and is refreshing its disaster scenario suite while explicitly setting AI to one side in doing so. A follow-up blog post will review into these areas in more detail.

What to Do about Security Stagflation

The answer is not to patch faster in the abstract. Against a negative patch window, “patch faster” is not a strategic approach to build a secure program in response to stagflation but an aspirational element we have been unsuccessfully dealing with for decades. Organizations need a harder prioritization currency:

  • Exploitability weighted: Is anyone actually using this?
  • Exposure weighted: Can it be reached from where an attacker stands?
  • Reachability aware: Is the vulnerable code path executed in your build and configuration?
  • KEV anchored: Has the ecosystem already confirmed exploitation in the wild?
  • Attack path aware: Does it advance a chain that ends somewhere that matters?

The first job of every CISO is to decide explicitly what the organization will not patch immediately and to document the controls, isolation boundaries, and risk decisions that make those defaults tolerable. That decision is uncomfortable, which is exactly why it needs structure: a named accountable owner for each deferral, a written compensating-control basis, an expiry date, and an automatic trigger that reopens the decision the moment a deferred item enters KEV or acquires a public exploit. A deferral without an expiry date is not a risk decision. It is an unrecorded liability.

Then measure the decision cycle rather than the patch cycle. Track four things:

  1. Mean time to applicability decision: How long it takes to answer "Does this affect us?" reported separately from mean time to remediate. These are different capabilities, and conflating them hides which one is failing.
  2. The share of KEV-listed and critical advisories receiving an applicability decision within the hour, within the shift, and within the business day: One business day is a defensible outer bound. It is the target NIST set for enriching KEV entries, and an organization should be able to reach a decision about its own estate at least as fast as a national database reaches one about everyone's.
  3. The share proven non-applicable through machine-verifiable evidence rather than analyst assertion: Evidence that a human asserts does not survive an audit or a bad quarter.
  4. The ratio of findings closed by proof of non-applicability to findings closed by patching: This is the one that matters most. It is the clearest single indicator of whether a program has actually entered the proof race, and it is the only number on this list that cannot be improved by working harder at the old model.

Finally, re-baseline the backlog itself: weight exposed findings by expected loss rather than counting raw CVEs. A queue measured in CVE counts will grow at 181 a day no matter how well you perform.

The currency changed. The winning programs will not be those that claim to patch everything. They will be those that can prove what matters, isolate what they cannot fix in time, and spend scarce remediation capacity where the loss path is real.


Sources

As of June 23, 2025, the MS-ISAC has introduced a fee-based membership. Any potential reference to no-cost MS-ISAC services no longer applies.