On February 18, 2026, OpenAI and Paradigm released EVMbench — an open-source benchmark that measures how well AI agents can detect, patch, and exploit high-severity smart contract vulnerabilities. The results are striking: GPT-5.3-Codex now successfully exploits 72.2% of critical, fund-draining b...
"As LLMs rapidly improve at finding exploits, it is important that we have visibility into and influence over the risks they could create for crypto." — Alpin Yukseloglu, Partner, Investing & Research at Paradigm
On February 18, 2026, OpenAI and Paradigm released EVMbench — an open-source benchmark that measures how well AI agents can detect, patch, and exploit high-severity smart contract vulnerabilities. The results are striking: GPT-5.3-Codex now successfully exploits 72.2% of critical, fund-draining bugs from Code4rena audit competitions, up from less than 20% when the project began. In just over a year, AI has gone from failing most exploit tasks to outperforming the majority of human auditors on standardized vulnerability sets.
This report examines the collision between AI-powered security tooling and the $456 million smart contract auditing industry. The economic implications cut both ways: AI promises to dramatically reduce the cost of securing $100 billion in on-chain assets, but it simultaneously hands attackers the same exploit-discovery capabilities. For protocols, investors, and the audit firms caught in the middle, this is not a future concern — it is reshaping security economics right now.
The comparative analysis reveals a market in rapid transition. Andreessen Horowitz led a $20 million round for AuditAI on February 10. CertiK partnered with IBM Watson on February 14. PwC and Deloitte both launched blockchain-focused AI audit divisions in January 2026. Meanwhile, human auditors command $60,000–$120,000 per engagement for mid-complexity DeFi protocols, with wait times stretching months. AI is compressing what took weeks of human review into hours — but it still cannot catch novel economic exploits, governance logic flaws, or cross-protocol composability risks that define the most devastating attacks.
EVMbench represents the first serious attempt to standardize the measurement of AI capabilities in smart contract security. Built from 120 curated vulnerabilities across 40 real-world security audits — most sourced from Code4rena competitions, with additional scenarios from Paradigm's proprietary Tempo audit process — the benchmark tests AI agents across three distinct modes:
Each task runs in a containerized environment that mirrors real-world conditions, with deterministic on-chain state changes serving as the success criterion for exploit tasks. The benchmark includes verified "answer keys" for each vulnerability, ensuring solvability is confirmed before measuring model performance.
The collaboration is notable for who built it. Paradigm is crypto's most influential venture firm, with portfolio companies representing a significant share of total value locked in DeFi. OpenAI is the world's most capitalized AI company. OtterSec contributed frontend implementation. That these three entities jointly invested in building an open-source security benchmark signals that AI's role in smart contract security has crossed from experimental to strategic.
The headline figures from EVMbench's initial evaluation are remarkable for both what AI can do and what it cannot:
Exploit Mode:
The rate of improvement is the critical signal. AI exploit capability more than doubled in six months and nearly quadrupled from the project's inception. At this trajectory, near-complete exploit coverage on standardized vulnerability sets is likely within 12–18 months.
Detection and Patching: Performance in detect and patch modes lags significantly behind exploit capability. OpenAI and Paradigm noted that patching remains "a major weakness" — fixing contract vulnerabilities requires preserving correct behavior across edge cases, which demands understanding deeper design assumptions in the code. Detection similarly requires reasoning about intended versus actual behavior, a task that demands context AI still struggles to maintain across complex codebases.
This asymmetry has profound implications. AI is learning to break smart contracts faster than it is learning to fix them.
The global smart contract audit market was valued at $456 million in 2024 and is projected to reach $3.42 billion by 2033, growing at a CAGR of 24.5%. But the structure of this market is being rewritten in real time.
Current pricing benchmarks (2026):
| Protocol Complexity | Cost Range | Timeline | |---|---|---| | Simple ERC-20 token | $5,000–$20,000 | 3–5 days | | Mid-complexity DeFi | $40,000–$100,000 | 2–6 weeks | | Enterprise multi-chain | $150,000–$250,000+ | 2–4 months | | ZK circuit audit | +80–120% premium | Extended |
Language premiums add further costs: Rust/Solana audits carry a 25–40% premium over Solidity baseline, while Cairo (StarkNet) and Move (Sui/Aptos) command 30–45% premiums. Urgency surcharges add 20–40%.
Three dominant structures compete for this revenue:
AI is now emerging as a fourth structure — and it threatens to compress the economics of the first three.
The data from EVMbench and industry deployments reveals a clear division of capabilities:
Where AI excels (today):
Where humans remain essential:
The Bybit lesson: The industry's worst hack of 2025 — $1.4 billion stolen from Bybit — did not exploit a smart contract flaw. Attackers compromised a developer's machine and injected malicious JavaScript into the Safe{Wallet} UI, causing multi-sig signers to unknowingly authorize a malicious transaction. No AI smart contract auditor would have caught this. No human smart contract auditor would have either. The attack surface was the human layer, not the code layer — illustrating that even perfect AI auditing solves only part of the security problem.
The same capabilities that make AI a powerful defensive tool also make it a powerful weapon. EVMbench's results demonstrate this dual-use reality explicitly: the benchmark was designed to measure exploit capability alongside detection and patching capability.
If GPT-5.3-Codex can exploit 72.2% of critical Code4rena bugs in a standardized environment, then any actor with API access — or any sufficiently capable open-source model — can replicate similar attack reconnaissance at scale. The economics flip dramatically:
The OWASP 2026 framework for smart contract security already identifies structural governance and access control failures — not coding bugs — as the dominant risk vector. This aligns with the EVMbench data: as AI eliminates the low-hanging fruit of known vulnerability patterns, the remaining attack surface shifts toward governance, social engineering, and supply chain compromises that AI cannot yet model.
Total crypto theft reached $3.4 billion in 2025, with the Bybit hack alone accounting for $1.4 billion (69% of total). Access control exploits were the largest category at $1.63 billion. Smart contract bugs accounted for approximately $263 million in H1 2025 — a figure that AI auditing is specifically positioned to reduce.
The capital flowing into AI security tells the story of where the market is heading:
OpenZeppelin reported that its AI tools cut auditing time by 50%. Sherlock launched its AI auditing product trained on verified audit findings, contest submissions, and exploited codebases. The competitive audit platform model is evolving: Sherlock AI now provides continuous analysis on pull requests and code changes, using multi-step reasoning to trace state transitions.
For protocols, the economic calculus is shifting. A mid-complexity DeFi audit at $60,000–$120,000 with a multi-week timeline may increasingly be supplemented — or partially replaced — by continuous AI monitoring at a fraction of the cost. The realistic budget for a 2026 security program is moving from "one-time audit plus bug bounty" toward "continuous AI surveillance plus targeted human review for novel risk vectors."
But this efficiency gain creates a paradox: if auditing becomes cheap and fast, the barrier to deploying unaudited code drops. More code deployed means more attack surface. The net security outcome depends on whether defenders adopt AI faster than attackers — a race with no guaranteed winner.
AI exploit capability is doubling every six months. GPT-5.3-Codex exploits 72.2% of critical smart contract bugs, up from 31.9% for GPT-5 six months earlier and less than 20% at project inception. The trajectory points toward near-complete coverage of known vulnerability classes within 12–18 months.
AI breaks contracts faster than it fixes them. Exploit performance dramatically outpaces detection and patching capabilities. This asymmetry favors attackers in the short term and demands that the industry invest disproportionately in AI patching research.
The audit market is being restructured, not replaced. Human auditors remain essential for novel economic exploits, governance logic, and cross-protocol risk — the categories responsible for the largest losses. AI compresses the commodity tier of auditing while increasing demand for elite human judgment.
The $263 million addressable problem is only part of the story. Smart contract bugs caused $263 million in losses in H1 2025, but access control failures caused $1.63 billion. AI auditing addresses the smaller slice. The industry's largest losses come from attack surfaces AI cannot yet model.
Dual-use risk is real and measurable. EVMbench explicitly benchmarks attack capability. Any improvement in AI defense is simultaneously an improvement in AI offense. Open-source models will eventually match frontier model exploit rates, democratizing attack capability.
EVMbench marks a turning point not because it reveals something unknown, but because it quantifies what the industry suspected: AI is approaching parity with human auditors on standardized vulnerability detection and already surpasses most humans on exploit generation for known bug classes. The $456 million audit industry built on human expertise is entering a structural transition.
The winners will not be pure AI shops or pure human auditoriums. They will be firms that master the handoff — deploying AI agents for continuous, broad-surface scanning while directing expensive human attention to the novel, the complex, and the compositional. Paradigm's framing is instructive: "a growing portion of audits in the future will be done by agents." Not all audits. A growing portion.
For protocols managing $100 billion in on-chain value, the message is urgent. AI-powered attackers are already scanning for the vulnerability patterns that EVMbench demonstrates AI can find at 72% accuracy. Defenders who rely solely on point-in-time human audits are bringing a clipboard to a machine-speed arms race. Continuous AI monitoring is no longer a luxury — it is table stakes for any protocol that expects to survive 2026.