AI agents can now exploit the majority of known-vulnerable smart contracts at an average scan cost of $1.22 per contract. Anthropic's SCONE-bench evaluation, published in late 2025, found that frontier models including Claude Opus 4.5 and GPT-5 collectively developed working exploits worth $4.6 m...
"AI has made legacy-contract hunting cheaper, faster, and more scalable, especially for old forks, dusty deployments, under-maintained vaults, and inherited code paths." — Anthropic Frontier Red Team, Smart Contracts Research Report
AI agents can now exploit the majority of known-vulnerable smart contracts at an average scan cost of $1.22 per contract. Anthropic's SCONE-bench evaluation, published in late 2025, found that frontier models including Claude Opus 4.5 and GPT-5 collectively developed working exploits worth $4.6 million against 405 real contracts. A separate OpenAI-Paradigm benchmark released in February 2026 shows GPT-5.3-Codex now exploits 71% of critical Code4rena bugs, up from under 20% when development began. On the defensive side, a purpose-built AI security agent from Cecuro detected vulnerabilities in 92% of exploited contracts — but deploying such tools remains optional, and most legacy contracts have no active monitoring.
The economics are stark. Exploit capability on Anthropic's post-knowledge-cutoff benchmark has risen from $5,000 to $4.6 million in one year — a doubling rate of roughly every 1.3 months. Token costs are falling 22% per model generation. Meanwhile, Q1 2026 DeFi losses stand at $137 million across 15 exploits, and over $170 billion in total value remains locked across DeFi protocols, much of it in aging contracts that have not been re-audited since deployment.
Anthropic's frontier red team evaluated AI models against SCONE-bench (Smart CONtracts Exploitation benchmark), a dataset of 405 contracts actually exploited between 2020 and 2025 across Ethereum, BNB Smart Chain, and Base. The results published in December 2025:
Separately, OpenAI and Paradigm released EVMbench in February 2026, a benchmark of 117 curated vulnerabilities from 40 audits. On this benchmark, GPT-5.3-Codex achieved a 71.0% exploit rate on critical fund-draining bugs, up from 33.3% for GPT-5 released six months prior and under 20% for models evaluated at the project's inception. Detection and patching tasks lag behind, with agents performing best when the exploit objective is explicit.
OpenZeppelin subsequently audited EVMbench and identified methodological concerns, including at least four issues classified as high-severity that were not exploitable in practice. This suggests headline exploit rates may overstate real-world capability, though the directional trend — rapid improvement in automated exploitation — is not in dispute.
The Anthropic research quantified the economic asymmetry:
| Metric | Value | |---|---| | Average cost per contract scan | $1.22 | | Average cost per vulnerable contract identified | $1,738 | | Average revenue per successful exploit | $1,847 | | Average net profit per exploit | $109 | | Total exploit revenue (post-cutoff contracts) | $4.6 million |
The $109 average net profit per exploit may appear marginal, but the economics favor volume. An attacker can deploy AI agents against thousands of contracts simultaneously. The Anthropic team noted that token costs are decreasing by approximately 22% per model generation, with a 65.8% cumulative cost reduction from Claude Opus 4 to Opus 4.5 in under six months. Potential exploit revenue has been doubling approximately every 1.3 months.
When tasked against 2,849 recently deployed contracts with no known vulnerabilities, both Sonnet 4.5 and GPT-5 uncovered two novel zero-day vulnerabilities and produced working exploits worth $3,694, at an API cost of $3,476. The economics of zero-day discovery by AI are currently near break-even, but the trend line — falling costs, rising capability — points toward profitability within months.
Two competing benchmarks now define how the industry measures AI-vs-smart-contract capability:
SCONE-bench (Anthropic, December 2025): 405 real exploited contracts across three chains. Tests end-to-end exploit generation in simulated blockchain environments. Measures economic harm potential. Contracts span 2020-2025 deployment dates.
EVMbench (OpenAI/Paradigm, February 2026): 117 vulnerabilities from 40 audits, primarily Code4rena competition bugs. Tests detection, patching, and exploitation. Open-source, available on GitHub. Focus on Ethereum-native EVM contracts.
A March 2026 paper on arXiv (Re-Evaluating EVMBench, arXiv:2603.10795) raised questions about whether EVMbench's curated competition bugs accurately represent real-world attack surfaces. The Anthropic benchmark, using contracts with verified on-chain losses, may better approximate adversarial conditions. Neither benchmark fully captures social-engineering-adjacent attacks — compromised keys, phishing, or infrastructure breaches — which accounted for the majority of 2025's $17 billion in crypto hack losses according to CoinDesk.
Not all AI capability accrues to attackers. In February 2026, AI security firm Cecuro published results from a specialized detection agent evaluated against 90 real-world contracts exploited between October 2024 and early 2026, representing $228 million in verified losses:
The finding has direct implications for protocol operators: generic AI tools are insufficient for smart contract security. The 92% detection rate requires purpose-built infrastructure that most protocols have not deployed.
Industry estimates suggest hybrid approaches combining AI screening with human expertise can catch 95%+ of vulnerabilities, compared to 60-70% for manual-only and 70-85% for AI-only audits. However, the cost of continuous AI-augmented monitoring remains a barrier. Smart contract audit costs in 2026 range from $5,000 for simple token contracts to over $250,000 for enterprise-grade multi-chain systems, according to industry pricing data. Top-tier firms including OpenZeppelin, CertiK, and Trail of Bits charge $80,000 to $200,000 for enterprise audits.
Q1 2026 DeFi exploits totaled $137 million across 15 incidents, surpassing Q1 2025 figures. The largest incidents:
| Protocol | Loss | Attack Vector | |---|---|---| | Step Finance | $27.3M | Smart contract vulnerability | | Truebit | $26.2M | Integer overflow in legacy minting contract | | Resolv | $25.0M | AWS Key Management Service compromise | | SwapNet | $13.4M | Smart contract exploit | | FOOM Cash | $2.3M | Lending protocol vulnerability |
The Truebit exploit in January 2026 is particularly relevant. The protocol's minting contract, deployed over five years earlier, contained an integer overflow vulnerability that allowed an attacker to drain 8,535 ETH. Several security researchers told DL News the exploit was "a likely candidate" for AI-assisted discovery, though this remains unconfirmed. Truebit's native token TRU fell 99.9% following the exploit.
Notably, the most expensive Q1 incidents were not all smart contract bugs. The Resolv hack involved a compromised AWS Key Management Service key — an infrastructure-level breach that allowed the attacker to drain $25 million before the team could respond. As the broader 2025 data showed ($17 billion in total hack losses), key management failures and social engineering remain larger aggregate attack vectors than code exploits.
DeFi protocols collectively hold over $170 billion in total value locked across hundreds of protocols as of March 2026. An unknown but significant portion of this value sits in contracts that were audited once at deployment and have not been re-evaluated since.
The Anthropic research explicitly identified the highest-risk category: "old forks, dusty deployments, under-maintained vaults, and inherited code paths." These contracts were designed for a threat model that did not include automated AI scanning at $1.22 per contract. The security assumption that a human attacker would need to manually review code for hours or days no longer holds.
According to security researchers quoted by DL News, "'Audited once' is no longer a serious security model." If attackers can continuously re-scan the long tail of old contracts, dormant risk becomes active risk. The problem is structural: many legacy contracts are immutable by design, and the teams that deployed them may no longer exist to implement fixes even if vulnerabilities are identified.
The global blockchain security market was valued at $5.05 billion in 2025 and is projected to reach $8.41 billion in 2026, according to industry forecasts. Leading firms — CertiK, OpenZeppelin, Trail of Bits, Cyfrin, Hacken — are integrating AI into their workflows.
However, the economics favor offense. An attacker needs to find one vulnerability. A defender needs to find all of them. AI reduces the cost of scanning for both sides, but the asymmetry persists. A comprehensive audit costing $100,000+ is a one-time event; an AI agent scanning thousands of contracts at $1.22 each is continuous.
The industry is moving toward continuous monitoring models, but adoption remains limited to well-funded protocols. The long tail of DeFi — smaller protocols, forks, legacy deployments — lacks the resources for ongoing AI-augmented security.
The data describes an asymmetric escalation. Offensive AI capabilities against smart contracts are improving faster than defensive adoption. The economics — $1.22 per scan, falling API costs, rising exploit success rates — favor automated, continuous probing of the existing contract base. Defensive tools exist and perform well when deployed, but deployment remains concentrated among well-resourced protocols. The long tail of legacy contracts, holding an indeterminate but material share of DeFi's $170 billion TVL, represents accumulated risk that the original threat models did not contemplate. The question is not whether AI-assisted smart contract exploitation is viable. According to both Anthropic and OpenAI-Paradigm benchmarks, it already is. The question is how quickly defensive infrastructure scales to match.