A new class of arms race is reshaping the economics of blockchain security. In February 2026, three landmark studies — from Anthropic, OpenAI/Paradigm, and AI security firm Cecuro — independently demonstrated that frontier AI models can now autonomously detect, exploit, and in some cases patch sm...
A new class of arms race is reshaping the economics of blockchain security. In February 2026, three landmark studies — from Anthropic, OpenAI/Paradigm, and AI security firm Cecuro — independently demonstrated that frontier AI models can now autonomously detect, exploit, and in some cases patch smart contract vulnerabilities at scale. The cost: $1.22 per contract. The capability doubling time: 1.3 months.
These findings arrive against a backdrop of $3.4 billion in crypto stolen during 2025, with North Korea's Lazarus Group alone responsible for $2.02 billion — including the $1.5 billion Bybit heist, the largest single crypto theft in history. The question is no longer whether AI will transform blockchain security. It is whether the defenders can deploy AI fast enough to outpace the attackers who already are.
This report examines the three converging research breakthroughs, quantifies the economic asymmetry between AI-powered offense and defense, and maps the value redistribution already underway in the $8.4 billion blockchain security market.
Within a span of eight weeks in late 2025 and early 2026, three independent research efforts converged on the same conclusion: AI agents have crossed the threshold of autonomous smart contract exploitation.
Anthropic's Red Team (December 2025): Anthropic tested Claude Opus 4.5, Claude Sonnet 4.5, and GPT-5 against 405 benchmark smart contract problems. The models collectively produced turnkey exploits for 207 of them — a 51.1% success rate — yielding $550.1 million in simulated stolen funds. More critically, when tested on vulnerabilities exploited after the models' training cutoffs, they still cracked 19 of 34 targets (55.8%), producing $4.6 million in simulated theft. These were not memorized solutions. The models were reasoning their way to novel exploits. Two genuine zero-day vulnerabilities were also discovered across 2,849 recently deployed contracts.
OpenAI & Paradigm's EVMbench (February 2026): The two firms launched EVMbench, an open-source benchmark built from 120 curated vulnerabilities across 40 audits. GPT-5.3-Codex, the top performer, exploited 72% of vulnerabilities and patched 41.5%. Claude Opus 4.6 led in raw detection at 45.6%. The performance gap between model generations is itself accelerating — GPT-5.3-Codex represents a dramatic leap over GPT-5, which scored just 33.3% and was released only six months prior.
Cecuro's Domain-Specific Agent (February 2026): AI security startup Cecuro tested a purpose-built security agent against 90 real-world exploited DeFi contracts representing $228 million in verified losses. The specialized agent detected vulnerabilities in 92% of contracts ($96.8 million in exploit value), compared with just 34% ($7.5 million) for a baseline GPT-5.1 coding agent running on the same underlying model. The 2.7x performance multiplier came entirely from domain-specific methodology — structured review phases and DeFi-focused security heuristics — not superior AI.
The implication is stark: the raw capability exists in every frontier model. The differentiator is now the application layer.
The economics of AI-powered exploitation are alarming in their simplicity. Anthropic's research quantified the cost of an exhaustive AI-driven vulnerability scan at $1.22 per contract. At that price, scanning every verified contract on Ethereum mainnet costs less than a business-class flight.
More troubling is the trajectory. Anthropic found that AI exploit revenue potential has been doubling every 1.3 months, while the cost of inference tokens has been falling by approximately 22% every two months. This creates a compounding curve where the economic incentive to deploy offensive AI agents grows exponentially, even as the barrier to entry collapses.
The defensive picture is grimmer than the headline numbers suggest. Cecuro's study revealed that several contracts in its benchmark had previously passed professional human audits before being exploited in the wild. Traditional security reviews, priced at $60,000–$250,000 per engagement and taking weeks to complete, are being outperformed by AI agents running for minutes at negligible cost.
There is, however, a crucial nuance. EVMbench revealed that AI agents' biggest weakness is not exploitation or patching — it is finding vulnerabilities in large codebases. When agents were given hints about where a vulnerability was located, exploit success rates jumped from 63% to 96%. This means the current bottleneck is search, not reasoning — a problem that scales away with more compute and better tooling.
Detection timing matters enormously. Cecuro's data showed that immediate vulnerability detection yields an 86–89% success probability, but success drops to just 6–21% when detection is delayed by a week. In smart contract security, speed is not an advantage — it is the entire game.
These AI breakthroughs arrive at the worst possible moment for an industry hemorrhaging capital to theft. Chainalysis reported that $3.4 billion in digital assets were stolen in 2025, with North Korea's state-backed hackers responsible for $2.02 billion — a 51% increase over 2024.
The concentration of losses is extreme. The top three hacks of 2025 accounted for 69% of total stolen funds. The Bybit heist alone — $1.5 billion extracted through a supply-chain attack on Safe{Wallet}'s infrastructure — represented 44% of all crypto theft for the year. The ratio between the largest hack and the median incident crossed the 1,000x threshold for the first time, signaling that the industry's security model is failing at the tail risk that matters most.
The Lazarus Group's methodology has itself evolved. The Bybit attack did not exploit a smart contract bug — it compromised the signing interface, showing legitimate transaction details to multi-sig signers while routing funds to attacker-controlled addresses. Three of Bybit's signers approved what they believed was a routine internal transfer. This is social engineering at infrastructure scale, a category of attack where AI-powered code analysis alone cannot provide protection.
With DPRK's cumulative crypto theft now at $6.75 billion and capabilities continuing to advance, the 2026 outlook Chainalysis describes as "uncertain" is more accurately described as ominous.
The industry's attempt to standardize AI security evaluation has already run into trouble. OpenZeppelin, the most prominent smart contract security framework, audited EVMbench shortly after its release and found significant methodological problems.
The most damaging finding: training data contamination. The best-performing models — Claude Opus 4.6 (training cutoff May 2025) and GPT-5.2 (cutoff August 2025) — were likely exposed to the benchmark's vulnerability reports during pretraining. This means the headline exploitation rates may reflect memorization rather than genuine reasoning.
OpenZeppelin also identified at least four issues classified as high-severity that are not actually exploitable in practice, calling into question the precision of the benchmark's vulnerability labels.
These are serious flaws, but they do not invalidate the broader conclusion. Even with contamination concerns, the trajectory across model generations is clear and accelerating. And Anthropic's post-knowledge-cutoff testing — where models cracked 55.8% of previously unseen vulnerabilities — provides cleaner evidence of genuine exploitation capability.
The lesson is institutional: the blockchain security industry needs rigorous, contamination-resistant benchmarks before AI auditing tools can be deployed with confidence. EVMbench is version 0.1 of a standard that must evolve rapidly.
The blockchain security market is projected at $8.41 billion in 2026, growing at a 66.4% CAGR to $495 billion by 2034 according to industry estimates. Within this, smart contract auditing is a high-margin, labor-intensive service — a mid-complexity DeFi protocol audit runs $60,000–$120,000, with top-tier engagements exceeding $250,000.
AI is about to compress these economics violently. If a specialized agent can detect 92% of exploitable vulnerabilities in minutes for negligible compute cost, the value proposition of a six-figure, multi-week human audit changes fundamentally. This does not mean human auditors disappear — the Cecuro data shows that domain expertise in the application layer is what multiplies AI performance by 2.7x. But it means the audit industry's pricing power shifts from labor hours to methodology and domain knowledge.
The value chain is redistributing along three tiers:
Commodity scanning: Baseline AI agents running continuous vulnerability checks on deployed contracts. Near-zero marginal cost. This becomes table stakes — a protocol without continuous AI monitoring will be uninsurable.
Specialized AI auditing: Domain-specific agents like Cecuro's, combining frontier models with structured security methodologies. This is where the 92% detection rate lives. Priced at a fraction of traditional audits but requiring proprietary methodology.
Human-AI hybrid review: Senior security researchers working alongside AI agents, focusing on the attack surfaces AI cannot yet cover — infrastructure compromise, social engineering, economic design flaws. This is the premium tier, and it is where the Bybit-class attacks live.
The firms that will capture the most value are those building proprietary methodology layers — not those relying on raw model capability that is commoditizing in real time.
The uncomfortable truth is that offense currently holds a structural advantage. Attackers need to find one vulnerability. Defenders need to find all of them. AI amplifies both sides, but the asymmetry favors the attacker when capability is scaling faster than deployment.
Three factors could shift the balance toward defense:
Speed of deployment matters more than capability. Cecuro's finding that detection probability drops from 89% to 6–21% within a week means the race is won by whoever deploys AI monitoring first, not whoever has the most powerful model. Continuous, automated scanning of deployed contracts — even with imperfect models — dramatically reduces the attack surface.
Methodology is the new moat. The 2.7x performance gap between a raw frontier model and a domain-specialized agent means that open-source model capability alone is insufficient for either offense or defense. The teams building structured security frameworks on top of commodity AI will outperform those relying on model capability alone.
Regulation may force adoption. As FATF travel rule enforcement intensifies — 85 of 117 jurisdictions have now passed implementation legislation, with Australia's March 31 deadline imminent — the compliance infrastructure for blockchain is expanding. Security standards and mandatory AI auditing requirements are a logical next step, particularly for protocols seeking institutional capital.
AI agents can now autonomously exploit smart contracts at $1.22 per scan, with capability doubling every 1.3 months — the economics overwhelmingly favor attackers at current defensive adoption rates.
Domain-specific AI security agents detect 92% of exploitable vulnerabilities, a 2.7x improvement over raw frontier models, proving that methodology — not model size — is the critical variable.
The $3.4 billion stolen in 2025 (44% from a single Bybit heist) demonstrates that the industry's current security model fails catastrophically at the exact tail risks that matter most.
EVMbench establishes a benchmark standard but has training data contamination issues that must be resolved before AI auditing tools can be deployed with institutional confidence.
The blockchain security market ($8.4B in 2026) is being restructured into three tiers: commodity scanning, specialized AI auditing, and human-AI hybrid review — with value accruing to proprietary methodology, not raw AI capability.
Speed of defensive deployment, not capability, will determine outcomes. Continuous automated monitoring at imperfect accuracy beats periodic perfect audits against adversaries operating at machine speed.
The convergence of Anthropic's red-team results, OpenAI's EVMbench, and Cecuro's specialized agent research marks a phase transition in blockchain security. The question facing every protocol, exchange, and institutional allocator is no longer whether AI can find vulnerabilities — it is whether your AI finds them before someone else's AI exploits them.
At $1.22 per scan and a 1.3-month capability doubling time, the window for passive security postures is closing. The protocols and institutions that survive the next generation of AI-powered attacks will be those that have already deployed AI-powered defense — not as an experiment, but as critical infrastructure.
The $3.4 billion lost in 2025 was stolen by humans aided by increasingly sophisticated tools. The next $3.4 billion will be contested by AI agents on both sides of the battle. The arms race is no longer theoretical. It is the defining economic contest in blockchain security for 2026 and beyond.