On February 18, 2026, OpenAI and Paradigm jointly released EVMbench — an open-source benchmark that measures how well AI agents can detect, patch, and exploit high-severity smart contract vulnerabilities. The results are striking: GPT-5.3-Codex now successfully exploits 72.2% of critical vulnerab...
On February 18, 2026, OpenAI and Paradigm jointly released EVMbench — an open-source benchmark that measures how well AI agents can detect, patch, and exploit high-severity smart contract vulnerabilities. The results are striking: GPT-5.3-Codex now successfully exploits 72.2% of critical vulnerabilities drawn from real Code4rena audit competitions, up from under 20% when the project began and 31.9% for GPT-5 just six months ago.
This is not an incremental improvement. It represents a phase change in the relationship between artificial intelligence and blockchain security. Smart contracts currently secure more than $100 billion in open-source crypto assets, and the industry lost $3.4 billion to hacks in 2025 alone. The question EVMbench forces the industry to confront is uncomfortable but unavoidable: if AI can exploit 70% of critical bugs, who controls that capability — and what does it mean for the $2.69 billion smart contract security market?
The answer is reshaping the entire audit value chain, from how protocols get reviewed before launch, to how bug bounties are priced, to whether human auditors remain the final line of defense.
EVMbench is not a marketing exercise. It is a rigorous, containerized evaluation framework built on 120 curated vulnerabilities from 40 real security audits — primarily sourced from Code4rena audit competitions, supplemented by scenarios from Paradigm's own audit process for the Tempo blockchain, a purpose-built L1 for stablecoin payments.
The benchmark tests AI agents across three distinct tasks:
Detect Mode: Agents audit a smart contract repository and are scored on recall of ground-truth vulnerabilities identified by professional human auditors. This simulates the core work of a security review.
Patch Mode: Agents modify vulnerable contracts to eliminate exploitability while preserving intended functionality, verified through automated tests and exploit checks. This tests whether AI can fix what it finds.
Exploit Mode: Agents execute end-to-end fund-draining attacks against deployed contracts in a sandboxed blockchain environment. Success is measured by on-chain state changes — did the funds actually move?
Each task runs in an isolated container, simulating realistic operating conditions. Crucially, EVMbench includes "answer keys" for each task to verify solvability, and it only uses previously disclosed vulnerabilities. The entire framework — tasks, tooling, and documentation — is freely available on GitHub.
What makes EVMbench significant is its provenance. This is not an academic exercise from a university lab. It was built by the world's leading AI company (OpenAI) in partnership with one of crypto's most technically rigorous venture firms (Paradigm). OpenAI has committed $10 million in API credits for open-source and infrastructure protection, signaling this is a strategic priority, not a side project.
The headline number — GPT-5.3-Codex exploiting 72.2% of critical vulnerabilities — is dramatic enough. But the full performance breakdown reveals something more nuanced and, in some ways, more concerning:
| Task | GPT-5.3-Codex | GPT-5 (6 months prior) | |------|---------------|----------------------| | Exploit | 72.2% | 31.9% | | Detect | 41% | — | | Patch | 28% | — |
Three observations stand out:
1. The exploit-detect gap is alarming. AI is dramatically better at attacking contracts than defending them. GPT-5.3-Codex can exploit 72% of critical bugs but only detect 41% and patch just 28%. This asymmetry mirrors a broader truth in cybersecurity, but the magnitude of the gap — AI being nearly 2.5x better at offense than defense — creates real systemic risk.
2. The rate of improvement is exponential. Moving from under 20% to 72.2% exploit capability in the time it took to develop the benchmark suggests that the next generation of models will be even more capable. At this trajectory, near-universal exploit capability for known vulnerability classes is not a question of if, but when.
3. Patching is the hardest problem. At just 28%, patch mode performance reveals that fixing vulnerabilities requires understanding edge cases, design assumptions, and cross-contract dependencies that current AI struggles with. This is precisely the kind of contextual reasoning that human auditors bring — and it suggests their role will evolve rather than disappear.
The benchmarks acknowledged their own limitations: they don't fully reflect real-world conditions, and certain timing-based and multi-chain attacks fall outside scope. But the directional signal is unmistakable.
The global smart contracts market was valued at $2.69 billion in 2025, projected to reach $16.31 billion by 2034 at a 26.3% CAGR. Within that, the security audit segment commands premium pricing: $8,000–$20,000 for simple ERC-20 token audits, $25,000–$100,000 for standard DeFi protocol reviews, and $75,000–$250,000+ for complex cross-chain or advanced DeFi systems.
These prices reflect scarcity. There are perhaps a few hundred world-class smart contract auditors, and the top firms — Trail of Bits, OpenZeppelin, Consensys Diligence — have months-long waitlists. Protocols routinely delay launches or go live unaudited because they cannot secure a review slot.
AI is now disrupting this bottleneck from two directions simultaneously:
Speed compression. OpenZeppelin's new AI-powered tools have cut audit times by approximately 50%. What previously took weeks to months can now be completed in days to weeks. This does not eliminate the need for human expertise, but it fundamentally changes the throughput equation.
Platform democratization. Code4rena and Sherlock are integrating AI assistants that pre-screen code before human auditors engage, expanding coverage while maintaining competitive incentive structures. Immunefi, the largest Web3 bug bounty platform with over $110 million in total payouts, is similarly incorporating AI tooling into its workflow.
The economic implications cascade. If AI halves the time for a $70,000 audit, the per-engagement price compresses — but total market volume may expand as previously unaudited protocols can now afford reviews. The audit market may grow in aggregate while individual firm revenues face pressure. This is the classic technology disruption pattern: better, faster, cheaper — eventually.
The 2025 crypto hack data provides the grim backdrop against which EVMbench must be evaluated. According to Chainalysis, $3.4 billion was stolen in crypto hacks in 2025. When including scams and fraud, the figure balloons to approximately $17 billion. The Bybit breach alone accounted for $1.5 billion — 44% of all service losses.
But here's the critical insight from the data: smart contract bugs specifically accounted for approximately $263 million of H1 2025's $3.1 billion in losses. The vast majority of losses — centralized exchange compromises, social engineering, access control failures — stemmed from what CoinDesk aptly called "a people problem, not a code problem."
This distinction matters enormously for evaluating AI's impact. EVMbench tests AI against code-level vulnerabilities — the $263 million problem. AI is rapidly becoming capable of finding and exploiting these bugs. But the $3+ billion problem — compromised credentials, insider threats, social engineering — lies largely outside the scope of what smart contract auditing AI can address.
The attacker-defender asymmetry cuts both ways:
For defenders: AI audit tools can now cover 10x more ground in half the time, systematically finding the kinds of reentrancy, overflow, and logic bugs that have historically drained DeFi protocols. OpenZeppelin, Trail of Bits (with Slither, Echidna, and Medusa), and emerging AI-native security firms are building what amounts to a continuous audit capability.
For attackers: The same AI that finds bugs for auditors can find bugs for exploitation. GPT-5.3-Codex's 72.2% exploit rate on real Code4rena vulnerabilities demonstrates that offensive AI capability is already formidable. AI-driven exploits surged by 1,025% in 2025, largely from insecure APIs and inference setups.
The race is asymmetric but not hopeless. The key advantage defenders have is access: legitimate auditors get full source code, test suites, and protocol context. Attackers typically work with deployed bytecode and limited information. The question is whether that advantage holds as AI models become more capable at decompilation and pattern recognition.
Winners:
Protocols that couldn't afford audits. The most immediate beneficiary of AI-powered security is the long tail of DeFi projects that previously launched without formal review. If AI reduces a $70,000 audit to $20,000 in cost (and a few days in time), the addressable market for security services expands dramatically. Given that the average smart contract exploit costs approximately $1.9 million, the ROI math becomes trivial.
Security platforms with AI integration. Code4rena, Sherlock, and Immunefi are well-positioned to layer AI tooling onto their existing marketplace models. The human-AI hybrid — where AI does the initial scan and human auditors focus on complex logic and edge cases — is the most likely near-term model for high-stakes audits.
OpenAI and Paradigm. By releasing EVMbench as an open benchmark, they establish the evaluation standard for AI security agents. Whoever defines the benchmark defines the competition — and both firms have clear strategic interest in the crypto-AI intersection.
Losers:
Mid-tier audit firms. Firms competing primarily on thoroughness rather than specialized expertise face existential pressure. If AI replicates the quality of a competent but not exceptional auditor, the mid-market collapses.
Solo auditors without AI tooling. The era of the individual security researcher manually reviewing contracts is ending. Not because their skills are obsolete, but because AI-augmented competitors will cover more ground faster at lower cost.
Protocols relying on "security theater." Projects that treated a single audit report as a permanent stamp of approval will find that continuous, AI-driven monitoring becomes the new baseline expectation.
EVMbench establishes the first rigorous benchmark for AI smart contract security. Built on 120 real vulnerabilities from 40 audits, it provides a credible, reproducible evaluation framework that the industry has lacked.
AI is dramatically better at attacking than defending. GPT-5.3-Codex scores 72.2% on exploit, 41% on detect, and just 28% on patch — a 2.5x offense-defense gap that creates systemic risk.
The rate of improvement is exponential. From under 20% to 72.2% exploit capability in months. The next model generation will narrow remaining gaps further.
The audit market faces structural disruption. AI tools are halving audit times and compressing costs. The total market may grow, but pricing power shifts from providers to protocols.
Human auditors evolve, not disappear. The 28% patch rate shows AI still cannot replace contextual reasoning. The winning model is human-AI collaboration, not replacement.
The biggest losses in crypto remain human, not code. Of $3.4 billion stolen in 2025, smart contract bugs accounted for roughly $263 million. AI audit tools address the code problem, not the people problem.
EVMbench represents a watershed moment — not because AI has solved smart contract security, but because it has made the problem measurable. For the first time, we have a rigorous framework to track how quickly AI is closing the gap between finding bugs and fixing them, between offense and defense.
The uncomfortable truth is that AI is advancing faster on the attack side than the defense side. A 72.2% exploit rate against real, critical vulnerabilities is not a proof of concept — it is a capability that exists today, in public, for anyone with API access. The security implications are profound: every unaudited contract is now more vulnerable than it was six months ago, because the tools available to attackers have improved faster than the tools available to defenders.
But the economic opportunity is equally profound. If AI can make comprehensive security reviews accessible to every protocol — not just the ones that can afford $100,000+ engagements — the total value protected increases dramatically. The audit market may compress on price while expanding on volume. Human auditors transition from line-by-line code review to architectural review, threat modeling, and the kind of contextual analysis that AI still cannot replicate.
The next 12 months will determine whether the crypto industry can absorb this capability shift constructively — deploying AI for defense faster than attackers can weaponize it. EVMbench gives us the scoreboard. The game is already underway.