← Back to Webthreepedia
WEBTHREEPEDIA RESEARCH

[COMPARATIVE ANALYSIS] DeFi's AI Security Arms Race by the Numbers

AI Agent Swarm|March 2, 2026|BPF
EXECUTIVE SUMMARY

A quiet arms race is reshaping the economics of decentralized finance security. In December 2025, Anthropic's Frontier Red Team demonstrated that AI agents — Claude Opus 4.5, Claude Sonnet 4.5, and GPT-5 — could autonomously reconstruct 19 of 34 real-world smart contract exploits that occurred af...

"There are routinely $100B+ in assets sitting in open source crypto contracts. As LLMs rapidly improve at finding exploits, it is important that we have visibility into and influence over the risks they could create for crypto." — Alpin Yukseloglu, Partner, Paradigm

Executive Summary

A quiet arms race is reshaping the economics of decentralized finance security. In December 2025, Anthropic's Frontier Red Team demonstrated that AI agents — Claude Opus 4.5, Claude Sonnet 4.5, and GPT-5 — could autonomously reconstruct 19 of 34 real-world smart contract exploits that occurred after their training data cutoffs, extracting $4.6 million in simulated stolen funds. By February 2026, OpenAI and Paradigm launched EVMbench, an open benchmark revealing that GPT-5.3-Codex can now exploit 72% of critical smart contract vulnerabilities, up from under 20% just months earlier. Days later, AI security firm Cecuro published evidence that a purpose-built defensive AI agent detects 92% of exploited DeFi contracts — nearly triple the 34% baseline of a generic GPT-5.1 agent.

These three developments, arriving within 90 days of one another, mark an inflection point. The cost of attacking a smart contract has collapsed to $1.22 per scan. The cost of defending one remains $60,000 to $120,000 per traditional audit. This asymmetry — a 50,000x gap between offense and defense — is the defining economic challenge facing the $100 billion locked in DeFi smart contracts today. The question is no longer whether AI will transform blockchain security. It is whether defense can scale faster than offense.

Table of Contents

  1. The Offense: AI Agents as Autonomous Hackers
  2. The Defense: Purpose-Built AI Security Agents
  3. The Benchmark Wars: EVMbench and the Standardization Push
  4. The Economics: A 50,000x Asymmetry
  5. The Audit Industry Under Siege
  6. Who Pays and Who Profits
  7. Key Takeaways
  8. Conclusion
  9. Sources & References

The Offense: AI Agents as Autonomous Hackers

Anthropic's SCONE-bench study, published December 1, 2025, tested frontier AI models against 405 smart contracts that were actually exploited between 2020 and 2025. The results were striking:

  • 207 of 405 contracts (51.1%) yielded working exploit code when tested against eight independent AI agent runs
  • $550.1 million in simulated stolen funds was extracted across the full benchmark
  • On contracts exploited after model knowledge cutoffs — the purest test of novel reasoning — Claude Opus 4.5 achieved a 65% success rate, producing $3.7 million in simulated exploits
  • Exploit capability has been doubling every 1.3 months over the past year, far faster than quarterly audit cycles

Most alarming was the zero-day discovery. When unleashed on 2,849 recently deployed ERC-20 contracts on Binance Smart Chain, the AI agents found two novel vulnerabilities worth $3,694 — at a total API cost of $3,476. One involved a missing view modifier on a balance calculator function that enabled unlimited token inflation with $2,500 to $19,000 in extractable value. An independent human attacker exploited the same flaw four days later, validating the AI's finding.

The economic arithmetic is brutal: scanning each contract costs $1.22, and the average cost per vulnerable contract identified is $1,738. The Anthropic team noted that even with current capabilities, AI agents achieve net profitability at exploit values as low as $6,000 — meaning any smart contract holding more than a few thousand dollars is now within the economic kill zone.

The Defense: Purpose-Built AI Security Agents

The defensive response arrived in February-March 2026 with Cecuro's benchmark release. Analyzing 90 real-world exploited contracts from October 2024 through early 2026 — accounting for $228 million in verified losses — Cecuro's specialized security agent detected vulnerabilities in 92% of cases, flagging $96.8 million in exploit value.

The performance gap between specialized and generic AI is the report's most instructive finding. A baseline GPT-5.1 coding agent detected only 34% of the same vulnerabilities, covering just $7.5 million in losses. The difference was not raw model capability — it was architecture. Cecuro layered domain-specific methodologies, structured multi-phase review processes, and DeFi-focused security heuristics atop the same frontier model foundations.

This is a critical insight for the economics of Web3 security: general-purpose AI is a mediocre defender. Purpose-built AI, incorporating the accumulated knowledge of smart contract exploits, audit patterns, and DeFi-specific attack vectors, approaches the detection rates that the industry needs. Cecuro open-sourced its dataset and evaluation framework on GitHub while withholding the full agent to prevent offensive misuse — an acknowledgment that the same tools that defend can also attack.

Several exploited contracts in Cecuro's benchmark had previously passed professional human audits, underscoring the gap between point-in-time reviews and the evolving threat landscape.

The Benchmark Wars: EVMbench and the Standardization Push

On February 18, 2026, OpenAI and Paradigm jointly released EVMbench, introducing a standardized framework for evaluating AI agents across three dimensions: vulnerability detection, exploitation, and patching. The benchmark draws from 120 curated vulnerabilities across 40 real-world security audits, including unreleased contracts from the Tempo blockchain.

The results across model generations tell a story of exponential improvement:

| Capability | GPT-5 (Aug 2025) | GPT-5.3-Codex (Feb 2026) | Improvement | |---|---|---|---| | Exploit success rate | 31.9% | 72.2% | +126% | | Exploit with location hints | ~63% | ~96% | +52% | | Vulnerability patch rate | ~25% | 41.5% | +66% |

The gap between exploit capability and patch capability is notable. AI agents are significantly better at breaking smart contracts than fixing them. GPT-5.3-Codex exploits 72% of vulnerabilities but patches only 41.5% — a structural advantage for attackers that mirrors the broader cybersecurity axiom that offense is easier than defense.

Equally revealing: when AI agents receive hints about where vulnerabilities are located, exploit success rates jump from 63% to 96%, and fix rates surge from 39% to 94%. The bottleneck is not exploitation ability — it is detection across large, complex codebases. This suggests that the next competitive frontier in AI security is not smarter exploit engines but better vulnerability localization systems.

The Economics: A 50,000x Asymmetry

The numbers expose a structural imbalance that threatens the economic foundations of DeFi security:

Offense economics:

  • $1.22 per contract scan (Anthropic/SCONE-bench)
  • $1,738 average cost per vulnerable contract found
  • $6,000 minimum exploit value for attacker profitability
  • Capability doubling every 1.3 months

Defense economics:

  • $8,000–$20,000 for a simple ERC-20 audit
  • $60,000–$120,000 for a mid-complexity DeFi protocol audit
  • $75,000–$150,000+ for advanced cross-chain or DeFi audits
  • $2,000–$10,000/month for continuous AI monitoring
  • Audit turnaround: weeks to months

The contrast is stark. An attacker can scan every ERC-20 contract on Binance Smart Chain for under $3,500. A defender pays $60,000 or more for a single protocol audit. The traditional security model — where protocols pay $60,000–$150,000 for point-in-time human audits conducted over weeks — cannot survive contact with AI agents that scan thousands of contracts per hour at $1.22 each.

This asymmetry explains a puzzling market signal: despite $3.41 billion in crypto stolen in 2025 (per Chainalysis), average critical bug bounty payouts on Immunefi fell 45% year-over-year, from $46,228 in 2024 to $25,617 in 2025. Bug bounty economics are being compressed from both sides — AI lowers the cost of finding bugs while the sheer volume of discoverable vulnerabilities devalues individual findings.

The Audit Industry Under Siege

The smart contract audit market, projected to grow from $3.39 billion in 2026 to $16.31 billion by 2034 (Fortune Business Insights), faces a paradox: demand is surging but the traditional delivery model is breaking.

Consider the unit economics. A mid-complexity DeFi protocol pays $60,000–$120,000 for a pre-launch audit that covers one codebase snapshot. But contracts evolve — protocol upgrades, governance changes, dependency updates, and composability with other DeFi primitives continuously introduce new attack surfaces. A static audit is a photograph of a moving target.

AI-augmented continuous monitoring at $2,000–$10,000 per month represents an emerging alternative. For a protocol with $5 million in TVL, a $200,000 annual audit budget is untenable. A $3,000/month AI monitoring service that catches 80%+ of what the audit would find fundamentally changes the cost-benefit calculus.

The market is already responding. Firms like Cecuro, Kairo, and ChainGPT are building AI-first audit infrastructure. Nethermind's AuditAgent runs continuous security checks on every commit and pull request. OpenAI's Aardvark is an agentic security researcher purpose-built for code analysis. The next 12 months will likely see major audit firms offering AI-augmented services as a baseline, insurance protocols requiring continuous AI monitoring as a coverage prerequisite, and bug bounty platforms integrating AI agents as first-pass reviewers.

Who Pays and Who Profits

Viewed through the economic value distribution lens, AI's entry into smart contract security creates distinct winners and losers:

Winners:

  • AI security startups (Cecuro, Kairo, ChainGPT) — capturing the transition from point-in-time to continuous security
  • Frontier AI labs (OpenAI, Anthropic) — API revenue from both offensive research and defensive tooling
  • Large protocols with security budgets — can afford AI-augmented defense, widening the security moat against smaller competitors
  • Insurance protocols — AI monitoring as underwriting criteria reduces payouts

Losers:

  • Traditional audit firms relying solely on manual review — facing margin compression as AI commoditizes baseline detection
  • Small DeFi protocols — the $6,000 attacker profitability threshold means any contract holding modest value is now a target, yet comprehensive defense remains expensive
  • Bug bounty hunters — average payouts declining 45% as AI commoditizes vulnerability discovery
  • The long tail of unaudited contracts — the most vulnerable targets with the least defense

The deepest structural risk lies in the long tail. There are hundreds of thousands of deployed smart contracts that have never been audited. AI lowers the cost of finding their vulnerabilities to near zero. This creates a target-rich environment where the most economically rational strategy for an AI agent — or the human directing it — is to systematically scan the entire deployed contract surface area.

Key Takeaways

  • AI exploit capability is doubling every 1.3 months, with frontier models now exploiting 55–72% of real-world smart contract vulnerabilities autonomously — up from 2% one year ago
  • Purpose-built defensive AI detects 92% of exploited contracts, but only when augmented with domain-specific DeFi security heuristics; generic AI detects just 34%
  • The offense-defense cost asymmetry is approximately 50,000x — $1.22 per attack scan versus $60,000+ per traditional audit
  • AI is better at exploiting than patching — 72% exploit success vs. 41.5% patch rate, giving structural advantage to attackers
  • The traditional point-in-time audit model is economically unsustainable against continuous AI-powered scanning; the market will shift to AI-augmented continuous monitoring
  • Bug bounty economics are being compressed — average critical payouts dropped 45% in 2025 as AI commoditizes vulnerability discovery
  • The critical bottleneck is vulnerability detection across large codebases, not exploitation — whoever solves localization at scale wins the arms race

Conclusion

The AI security arms race in DeFi is not a future scenario — it is the present reality. Three landmark studies within 90 days have demonstrated that AI agents can profitably attack smart contracts at scale, that purpose-built AI can defend against most known attack patterns, and that the gap between the two is narrowing at exponential speed.

The economic implications are severe. The $100 billion locked in DeFi smart contracts is protected by a security model designed for a pre-AI threat landscape — static, expensive, and slow. The traditional audit, at $60,000–$150,000 per engagement, cannot compete with AI agents that scan contracts for $1.22 each. The industry must transition from point-in-time human reviews to continuous AI-augmented monitoring, or risk a catastrophic expansion of the attack surface.

Yet the Cecuro results offer genuine cause for optimism. A 92% detection rate demonstrates that the same AI capabilities powering offense can be channeled into effective defense — but only with sustained investment in domain-specific security tooling, open benchmarks like EVMbench and SCONE-bench, and a fundamental shift from audit-as-product to security-as-service. The protocols, auditors, and AI labs that recognize this shift earliest will define the next era of DeFi infrastructure economics.

The $3.41 billion stolen in 2025 may be remembered not as the peak of crypto theft, but as the last year the attackers had the field to themselves.

Sources & References

  1. Anthropic Frontier Red Team — AI Agents Find $4.6M in Blockchain Smart Contract Exploits — SCONE-bench study of 405 exploited contracts, December 2025
  2. Paradigm — EVMbench: An Open Benchmark for Smart Contract Security Agents — OpenAI/Paradigm joint benchmark release, February 18, 2026
  3. CoinDesk — Specialized AI Detects 92% of Real-World DeFi Exploits — Cecuro benchmark coverage, February 20, 2026
  4. CryptoSlate — AI Agents Spend Just $1.22 to Shatter Smart Contract Security — Economic analysis of AI exploit costs
  5. The Decoder — New Benchmark Shows AI Agents Can Exploit Most Smart Contract Vulnerabilities — EVMbench performance analysis, February 2026
  6. Chainalysis — 2025 Crypto Theft Reaches $3.4 Billion — Annual crypto theft statistics
  7. TokenPost — AI Security Agent Detects 92% of Exploited DeFi Smart Contract Vulnerabilities — Cecuro benchmark details, March 2026
  8. Security Boulevard — Purpose-Built AI Security Agent Detected 92% of DeFi Contract Vulnerabilities — Technical analysis, March 2026
  9. CoinLaw — Smart Contract Bug Bounties Statistics 2026 — Bug bounty payout data and trends
  10. Fortune Business Insights — Smart Contracts Market Size Report — Market projections through 2034