← Back to Webthreepedia
WEBTHREEPEDIA RESEARCH

[DEEP DIVE] AI Exploits Smart Contracts for $1.22 Each

Zephyra|March 1, 2026|BPF
EXECUTIVE SUMMARY

The cost of autonomously exploiting a vulnerable smart contract has fallen to $1.22. According to Anthropic's Frontier Red Team, frontier AI models including Claude Opus 4.5 and GPT-5 autonomously reconstructed 19 of 34 post-knowledge-cutoff exploits, extracting $4.6 million in simulated value. A...

"A growing portion of audits in the future will be done by agents." — Paradigm, EVMbench Technical Report, February 2026

Executive Summary

The cost of autonomously exploiting a vulnerable smart contract has fallen to $1.22. According to Anthropic's Frontier Red Team, frontier AI models including Claude Opus 4.5 and GPT-5 autonomously reconstructed 19 of 34 post-knowledge-cutoff exploits, extracting $4.6 million in simulated value. Across a broader benchmark of 405 real exploits from 2020–2025, ten frontier models produced working proof-of-concepts for 207 contracts, with simulated stolen funds totaling $550.1 million.

The defensive side is scaling in parallel. Cecuro's specialized AI security agent detected vulnerabilities in 92% of 90 exploited DeFi contracts, flagging $96.8 million in exploit value — compared with 34% detection for a baseline GPT-5.1 coding agent. On February 18, OpenAI and Paradigm released EVMbench, an open benchmark for evaluating AI agents across vulnerability detection, patching, and exploitation. GPT-5.3-Codex now exploits over 70% of critical Code4rena bugs, up from less than 20% when the project began.

DeFi protocols hold over $100 billion in total value locked. The question is no longer whether AI will reshape smart contract security — it is whether defensive tooling can outpace offensive capability that doubles every 1.3 months.

Table of Contents

  1. The Offense: $1.22 Per Contract
  2. The Defense: 92% Detection and Rising
  3. EVMbench: The Industry's First Shared Ruler
  4. The Economics of AI Auditing vs. Human Auditing
  5. Capital Flows Into AI Security
  6. Zero-Day Discovery: The Inflection Point
  7. Key Takeaways
  8. Conclusion
  9. Sources & References

The Offense: $1.22 Per Contract

Anthropic's Frontier Red Team, in research published December 1, 2025 and authored by Winnie Xiao, Cole Killian, Henry Sleight, Alan Chan, Nicholas Carlini, and Alwin Peng (via the MATS and Anthropic Fellows programs), introduced SCONE-bench — a benchmark of 405 smart contracts exploited between 2020 and 2025 across Ethereum, Binance Smart Chain, and Base.

The numbers are specific. The average cost per agent run was $1.22. The average revenue per successful exploit was $1,847, yielding an average net profit of $109 per exploit. Token costs declined 70.2% across four generations of Claude models. The research describes exploit revenue doubling every 1.3 months — a pace that, if sustained, compresses the timeline for economically viable autonomous exploitation of live contracts.

On post-knowledge-cutoff contracts (exploited after the models' training data ended), Claude Opus 4.5 reconstructed 13 of 20 attacks on contracts exploited after June 2025 — a 65% success rate. GPT-5 and Sonnet 4.5 performed comparably. The agents operated with 60-minute time limits and access to Foundry, bash, Python, and Uniswap routing via Model Context Protocol.

Across the full 405-contract benchmark, 207 contracts were exploited by at least one of ten frontier models (51.1% success rate), representing $550.1 million in simulated stolen funds. These are not hypothetical vulnerabilities. They are reconstructions of actual on-chain exploits.

The Defense: 92% Detection and Rising

On February 20, 2026, AI security firm Cecuro published results from a study evaluating 90 real-world smart contracts exploited between October 2024 and early 2026, representing $228 million in verified losses. Cecuro's domain-specific AI security agent detected vulnerabilities in 83 of 90 contracts (92%), flagging $96.8 million in exploit value.

A baseline GPT-5.1 coding agent, running on the same underlying model, detected vulnerabilities in only 34% of the same contracts, flagging $7.5 million. According to Cecuro, the performance gap stems from domain-specific security methodology — structured review phases and DeFi-focused heuristics — layered on top of the frontier model, not from differences in core AI capability.

Cecuro open-sourced the benchmark dataset, evaluation framework, and baseline agent on GitHub. The full security agent was withheld, citing concerns about offensive repurposing. This is a recurring pattern in the field: defensive tooling is powerful enough to be dangerous if inverted.

The implication is direct. General-purpose AI tools miss the majority of high-value, complex vulnerabilities. Specialized systems do not. For protocol teams relying on generic AI-assisted audits, the gap between 34% and 92% detection represents substantial unmitigated risk.

EVMbench: The Industry's First Shared Ruler

On February 18, 2026, OpenAI and Paradigm released EVMbench — an open benchmark for evaluating AI agents' ability to detect, patch, and exploit smart contract vulnerabilities. OtterSec contributed frontend support. The benchmark draws on 120 curated vulnerabilities from 40 audits, with most sourced from open code audit competitions on Code4rena.

EVMbench operates in three modes. Detect mode tests whether an AI agent can audit code and identify known bugs. Patch mode evaluates whether the agent can propose fixes without breaking contract logic. Exploit mode tests whether the agent can chain attacks to drain funds in a controlled local environment.

According to Paradigm, when the project began, top models exploited less than 20% of critical Code4rena bugs. GPT-5.3-Codex now exploits over 70%. The rate of improvement matches the trajectory observed in Anthropic's SCONE-bench data.

Alongside the release, OpenAI announced $10 million in API credits to support open-source security and infrastructure protection. EVMbench is positioned both as a measurement tool and as an accelerant: standardized benchmarks enable rapid iteration on defensive agents, but they also provide attackers with a training environment.

The three benchmarks now active — SCONE-bench (Anthropic, offensive capability), Cecuro's evaluation framework (defensive detection), and EVMbench (OpenAI/Paradigm, full-spectrum) — constitute the first generation of standardized evaluation infrastructure for AI-driven smart contract security. Prior to 2025, no such shared measurement existed.

The Economics of AI Auditing vs. Human Auditing

Traditional human smart contract audits cost $50,000–$100,000 for a standard DeFi protocol, require 2–3 auditors, and take 3–6 weeks. A simple protocol with 1,000–3,000 lines of code costs $6,000–$12,000. A comprehensive audit with formal verification and economic attack-surface modeling runs $100,000 or more, adding a fourth auditor and continuous monitoring.

AI-assisted continuous monitoring costs $2,000–$10,000 per month. According to industry reporting, a DeFi protocol with $5 million in TVL cannot justify a $200,000 audit but can justify $3,000/month for continuous AI monitoring that catches an estimated 80% or more of what the full audit would find.

OpenZeppelin's AI tools reportedly cut auditing time by 50%. The emerging industry standard is a hybrid model: AI-assisted scanning for pattern detection (2–3 days), followed by human validation of high-risk areas (3–5 weeks), with formal verification applied selectively. AI accelerates detection of known vulnerability patterns, reducing audit scope by an estimated 15–25%, but cannot reliably catch economic exploits, governance vulnerabilities, or architectural issues requiring domain expertise.

The economics create a two-tier market. Well-funded protocols can afford hybrid audits plus continuous AI monitoring. Smaller protocols — the long tail of DeFi — face a choice between AI-only monitoring with known blind spots and no monitoring at all. Given that Anthropic's research shows autonomous exploitation costs $1.22 per contract, unmonitored protocols face asymmetric risk.

Capital Flows Into AI Security

Venture capital is responding to the threat-defense dynamic. Andreessen Horowitz led a $20 million funding round for AuditAI on February 10, 2026. According to Tracxn, the top-funded AI companies in the smart contracts sector — CUBE3, Phala Network, and BlockSec — have collectively raised $51.8 million.

CertiK announced a partnership with IBM on February 14, 2026 to integrate Watson AI into its audit framework. ConsenSys Diligence developed algorithms for code anomaly detection prior to human review. PwC and Deloitte both launched blockchain-focused AI audit divisions during this period.

On the protocol side, Ethereum-based DeFi platforms including Uniswap and Compound have begun requiring AI-assisted audits for protocol upgrades. This represents a shift from AI auditing as supplementary to AI auditing as prerequisite.

The total addressable market is substantial. Over $100 billion sits in open-source crypto contracts, according to Paradigm. Crypto losses in 2025 reached $3.4 billion according to Chainalysis, the highest since 2022, with the $1.4 billion Bybit hack accounting for 69% of total thefts. February 2026 saw approximately $37.7 million in hack losses — the lowest monthly figure since March 2025, though attribution to improved tooling versus seasonal variation remains unclear.

Zero-Day Discovery: The Inflection Point

The most consequential finding in Anthropic's research is not exploit reconstruction — it is zero-day discovery. When Anthropic's agents were pointed at 2,849 recently deployed contracts on Binance Smart Chain (April–October 2025, filtered for ERC-20 standard, verified source code, and $1,000+ liquidity), Claude Sonnet 4.5 and GPT-5 independently discovered two novel vulnerabilities.

Vulnerability one: an unprotected calculator function enabling token inflation, worth approximately $2,500 in simulated profit at the snapshot block, with potential value of $19,000 at peak liquidity. Vulnerability two: a missing fee recipient validation worth approximately $1,000 — which was independently exploited by a real attacker four days later.

Combined simulated revenue from these two zero-days: $3,694. GPT-5's API cost for the full evaluation: $3,476. Net profit: $218. The margin is thin, but the cost curve is moving in one direction. Token efficiency improved 70.2% across four model generations. At current improvement rates, the economics of autonomous zero-day discovery cross into clear profitability within months, not years.

Anthropic open-sourced SCONE-bench explicitly for defenders. Protocol teams can plug their own agents into the harness and test contracts on forked chains before deployment. The stated intent is to give defenders the same tools attackers will inevitably build.

Key Takeaways

  • Autonomous exploitation costs $1.22 per contract. Anthropic's research shows frontier AI models reconstruct real on-chain exploits at 51% success rate across 405 contracts, with exploit capability doubling every 1.3 months.
  • Specialized defense detects 92% of exploited contracts. Cecuro's domain-specific AI agent outperformed general-purpose tools by 2.7x (92% vs. 34% detection) on the same underlying model.
  • EVMbench establishes a shared measurement standard. GPT-5.3-Codex exploits over 70% of critical Code4rena bugs, up from less than 20% at project inception.
  • AI-only monitoring costs 94–97% less than full human audits. $3,000/month vs. $50,000–$100,000 one-time, with estimated 80%+ coverage.
  • Zero-day discovery by AI is now marginally profitable. Two novel vulnerabilities discovered autonomously yielded $3,694 in simulated revenue against $3,476 in API costs.
  • The long tail of DeFi is unprotected. Protocols that cannot afford $50,000+ audits face $1.22-per-contract scanning by autonomous agents.

Conclusion

Smart contract security is entering an arms race with measurable velocity on both sides. Offensive AI capability doubled approximately every 1.3 months through late 2025. Defensive tooling demonstrated 92% detection rates in controlled benchmarks. Three standardized evaluation frameworks now exist where none did 18 months ago.

The structural concern is asymmetry. Exploitation is cheap, automated, and scales horizontally. Defense requires specialized methodology, domain expertise, and continuous monitoring — all of which cost more than $1.22. The protocols most at risk are not the well-funded DeFi blue chips with six-figure audit budgets. They are the thousands of smaller contracts holding $1,000 to $100,000 in liquidity — enough to be profitable targets at current exploitation costs, too small to justify comprehensive security.

The market's response — $51.8 million in venture funding, partnerships between OpenAI and Paradigm, mandated AI audits at major protocols — indicates recognition of the problem. Whether defensive tooling scales faster than offensive capability is an empirical question that the next 12 months will answer. The benchmarks now exist to measure it.

Sources & References

  1. Anthropic Frontier Red Team — Smart Contract Exploitation Research — SCONE-bench methodology, $4.6M simulated exploit findings, zero-day discovery results
  2. Specialized AI Detects 92% of Real-World DeFi Exploits — CoinDesk — Cecuro study on AI-powered vulnerability detection
  3. EVMbench: An Open Benchmark for Smart Contract Security Agents — Paradigm — OpenAI/Paradigm benchmark methodology and results
  4. AI Agents Spend Just $1.22 to Shatter Smart Contract Security — CryptoSlate — Economics of autonomous exploitation
  5. 2025 Crypto Theft Reaches $3.4 Billion — Chainalysis — Annual crypto loss statistics
  6. What Smart Contract Audits Actually Cost in 2026 — Zealynx Security — Audit cost breakdown and comparison
  7. AI Slashes Smart Contract Audit Times by Half — The Currency Analytics — AI-assisted audit efficiency data
  8. OpenAI and Paradigm Partner on AI Agent Tool for Smart Contract Security — The Block — EVMbench partnership details
  9. Cyber Valuations Climb as Capital Concentrates, AI Security Expands — Help Net Security — VC funding trends in AI security