On February 18, 2026, OpenAI and Paradigm quietly released EVMbench — an open-source benchmark that measures how well AI agents can detect, patch, and exploit critical vulnerabilities in Ethereum smart contracts. The headline result should alarm every protocol team in DeFi: GPT-5.3-Codex now succ...
"It's now clear to us that a growing portion of audits in the future will be done by agents." — Alpin Yukseloglu, Partner, Paradigm
On February 18, 2026, OpenAI and Paradigm quietly released EVMbench — an open-source benchmark that measures how well AI agents can detect, patch, and exploit critical vulnerabilities in Ethereum smart contracts. The headline result should alarm every protocol team in DeFi: GPT-5.3-Codex now successfully exploits 72.2% of critical, fund-draining smart contract bugs autonomously, up from less than 20% when the project began and 31.9% just six months ago with GPT-5.
This is not a theoretical exercise. EVMbench draws on 120 real vulnerabilities from 40 actual security audits, most sourced from Code4rena competitions — the same bugs that drained hundreds of millions of dollars from live protocols. With over $100 billion locked in open-source smart contracts today, the implications are existential: AI has crossed the threshold from a supplementary auditing tool to an autonomous exploit machine that can drain funds faster and cheaper than any human attacker.
The economic consequences are already reshaping the $2.69 billion smart contract security industry. Traditional audits cost $50,000–$250,000 per engagement and take 3–6 weeks to complete, yet still miss 15–30% of critical vulnerabilities. AI agents can now replicate — and in many cases exceed — human exploit capabilities in minutes for a fraction of the cost. The question is no longer whether AI will transform smart contract security. It is whether the defensive deployment of these tools can outpace their weaponization.
EVMbench is not another synthetic coding benchmark. Built by OpenAI in collaboration with Paradigm and OtterSec, it evaluates AI agents against three distinct operational modes that mirror the real-world smart contract security lifecycle:
Detect: AI agents audit an entire smart contract repository and are scored on their recall of ground-truth vulnerabilities — the same bugs that earned bounty payouts in live audit competitions. This mode tests whether AI can find needles in a codebase-sized haystack.
Patch: Agents must modify vulnerable contracts to eliminate exploitability while preserving intended functionality. This is the hardest task because fixing subtle vulnerabilities requires understanding deeper design assumptions in the code — not just spotting what's broken, but knowing what "correct" looks like.
Exploit: The most realistic test. AI agents interact with a fully deployed contract on a local blockchain and attempt to autonomously drain funds end-to-end. No hints, no guardrails — just an agent, a vulnerable contract, and the objective to steal everything.
The benchmark includes 120 curated vulnerabilities from 40 real audits, supplemented by custom scenarios from the Tempo blockchain — a purpose-built L1 for stablecoin payments — ensuring the dataset extends beyond DeFi lending into payment-oriented contract code. Every task is containerized with answer keys to verify solvability, and agents operate in realistic environments that mirror actual deployment conditions.
The EVMbench results reveal a technology curve that is accelerating faster than most security professionals anticipated:
| Model | Exploit Rate | Patch Rate | Detection Rate | |-------|-------------|------------|----------------| | GPT-5.3-Codex (via Codex CLI) | 72.2% | 41.5% | — | | Claude Opus 4.6 | — | — | 45.6% | | GPT-5 (6 months prior) | 31.9% | — | — | | Top models at project start | <20% | — | — |
The exploit improvement curve is staggering: from sub-20% to 72.2% in the span of a single development cycle. GPT-5.3-Codex more than doubled its predecessor's exploit success rate in just six months.
Perhaps more revealing is what happens when agents receive locational hints — information about roughly where in a codebase a vulnerability exists. With medium-level hints, exploit success jumps from 63% to 96%, and patch rates climb from 39% to 94%. This 30+ percentage-point improvement confirms that the bottleneck is not exploitation capability but detection — finding where the bugs are in the first place.
The asymmetry in EVMbench's results tells a critical story: AI agents are already devastatingly effective attackers but only moderately effective defenders.
In exploit mode, the objective is explicit and binary — keep iterating until funds are drained. Agents excel here because the feedback loop is immediate: either the exploit worked or it didn't. In detection mode, agents face an open-ended search problem. The researchers observed that agents frequently stop after identifying a single vulnerability rather than exhaustively auditing the entire codebase — a behavior pattern that would be catastrophic in a real audit engagement.
Claude Opus 4.6 led detection at 45.6%, meaning even the best AI auditor misses more than half of the critical bugs in a given codebase. Compare this to GPT-5.3-Codex's 72.2% exploit rate: an attacker only needs to find one exploitable bug, while a defender must find all of them.
This asymmetry — attackers need one hit, defenders need perfection — has always defined security. But AI is dramatically compressing the attacker's cost curve while the defender's challenge remains fundamentally hard. The economic implications are severe.
The smart contract security industry is caught between two brutal realities.
The loss side is accelerating. Crypto lost approximately $17 billion to hacks, scams, and exploits in 2025 — the worst year on record. Smart contract-specific exploits accounted for $3.1 billion in the first half of 2025 alone. The Bybit hack in February 2025, attributed to North Korea's Lazarus Group, drained $1.5 billion in a single incident by compromising the transaction approval process of a Safe{Wallet} multi-signature setup. In January 2026 alone, $127 million was lost to exploits — a pace that projects to exceed 2025's totals.
The audit market cannot scale to meet the threat. A realistic pre-launch security budget for a mid-complexity DeFi protocol in 2026 runs $60,000–$120,000 for a single audit cycle. Complex protocols with cross-chain components can spend $100,000–$250,000. Annual security budgets for protocols with meaningful TVL routinely hit $150,000–$500,000. Despite these costs, traditional firm-led audits that assign a dedicated team over 3–6 weeks still miss 15–30% of critical vulnerabilities.
The math is damning. DeFi currently secures approximately $105–149 billion in TVL across all chains, with Ethereum alone holding roughly $52.8 billion. The total addressable audit market is estimated at $2.69 billion (2025), growing at 22% annually. Yet the losses from successful exploits — $17 billion in 2025 — dwarf the entire security industry's revenue by a factor of six.
This is the structural gap that AI-native security promises to close — or, if deployed offensively first, to catastrophically widen.
The EVMbench results crystallize a race condition that has been building for two years. Consider the asymmetric economics:
Offensive AI: An attacker using GPT-5.3-Codex can scan a target protocol's codebase, identify candidate vulnerabilities, and generate working exploit code at near-zero marginal cost. The 72.2% success rate against real-world critical bugs means that a motivated attacker running the model against 10 target protocols will, statistically, find exploitable bugs in 7 of them. The cost: API credits measured in dollars, not the $50,000–$250,000 that a human audit team charges.
Defensive AI: A protocol deploying AI-assisted auditing gets faster turnaround and broader coverage, but still faces the detection gap. At 45.6% recall on the best-performing detection model, AI auditing is a powerful supplement but not a replacement for human expertise. The real value is in continuous integration — running AI checks on every commit, flagging high-risk patterns before deployment, and reducing the surface area that human auditors must cover.
The danger zone exists in the gap between these two curves. Offensive AI capabilities are improving faster (exploit rates more than doubled in six months) than defensive capabilities (detection rates remain below 50%). Every month this gap persists, the expected value of attacking DeFi protocols increases relative to the cost of defense.
Critically, this analysis extends beyond Ethereum. EVMbench specifically includes vulnerability scenarios from Tempo, a payment-focused L1, demonstrating that the exploit surface encompasses not just DeFi lending and trading protocols but stablecoin infrastructure and payment systems — the sectors that institutional capital is actively entering.
The market is responding, though unevenly:
Sherlock AI launched in September 2025, built by experienced auditors Bernhard Mueller and 0x52. The system is trained on data from Sherlock's own audit contests and exploit reports, offering native GitHub integration that runs automated vulnerability checks on every commit and pull request. Sherlock positions AI as a pre-audit layer that catches issues during development rather than after deployment.
Olympix has focused on DevSecOps integration, embedding AI-powered security checks into the continuous integration pipeline. This approach treats security as a build-time constraint rather than a post-facto review — a model borrowed from traditional software engineering that is only now reaching smart contract development.
Almanax has differentiated by targeting complex logical vulnerabilities with open training datasets, addressing the class of bugs that are hardest for both humans and AI to detect — business logic errors, governance attack vectors, and cross-contract interaction flaws.
Uniswap Labs took a different approach on February 20, 2026, releasing seven open-source AI Skills for Uniswap v4 that include hook security foundations — structured interfaces designed to help AI agents audit and deploy hook contracts safely. This signals that major protocols are internalizing AI security tooling rather than relying solely on external audit firms.
Code4rena and contest-based platforms are evolving hybrid models where AI assistants pre-screen code and surface high-risk areas for focused human review. This 100–500 researcher competition model, augmented by AI triage, may prove more effective than either pure-AI or pure-human approaches.
AI can now autonomously exploit 72.2% of critical smart contract bugs — a capability that has more than doubled in six months and shows no signs of plateauing.
The attacker-defender asymmetry is widening. Exploit rates (72.2%) far exceed detection rates (45.6%), meaning AI is currently a more effective weapon than shield in smart contract security.
With locational hints, exploit success hits 96% and patching hits 94% — proving the technology works when pointed at the right target. The bottleneck is finding the vulnerabilities, not exploiting or fixing them.
The audit market is structurally broken. $2.69 billion in annual security spending cannot defend against $17 billion in annual losses, and traditional 3–6 week audit cycles cannot match the deployment velocity of modern DeFi.
AI-native security is already being built by Sherlock, Olympix, Almanax, and protocol teams like Uniswap, but adoption remains fragmented and early-stage.
The economic value at risk is massive. Over $100 billion in smart contract TVL and the institutional capital flowing into tokenized assets and stablecoin infrastructure sit directly in the exploit path that AI agents have learned to walk.
EVMbench is not just a research paper — it is a wake-up call priced in billions. The benchmark proves that frontier AI models have crossed from theoretical code-analysis capabilities into practical, autonomous exploit execution against real-world smart contracts. The 72.2% exploit success rate represents a phase transition: the point at which AI-assisted attacks become more reliable than most human-led security reviews.
For the $100+ billion locked in DeFi protocols, the implications are immediate. Every protocol that has not undergone a recent audit is now vulnerable to an attacker class that didn't exist 18 months ago. Every protocol that has been audited faces the uncomfortable reality that traditional reviews miss 15–30% of critical bugs — and AI agents can now find and exploit many of those misses.
The path forward requires a fundamental rethinking of smart contract security economics. Continuous AI-assisted monitoring must replace point-in-time audits. Security budgets must scale with TVL, not with headcount. And the industry must grapple with the dual-use reality that the same models that protect protocols can be weaponized against them.
OpenAI and Paradigm have open-sourced EVMbench, making the benchmark freely available. The question now is whether the DeFi ecosystem treats this as a research curiosity or as the strategic threat it actually is. The next six months — as exploit capabilities continue their exponential improvement curve — will determine which protocols survive the age of autonomous AI exploitation and which become its casualties.