On February 18, 2026, OpenAI and Paradigm quietly released EVMbench — an open-source benchmark that tests how well AI agents can detect, exploit, and patch smart contract vulnerabilities. The timing was not accidental. Days earlier, a lending protocol called Moonwell lost $1.78 million because AI...
"Smart contracts routinely secure $100B+ in open-source crypto assets. As AI agents improve at reading, writing, and executing code, it becomes increasingly important to measure their capabilities in economically meaningful environments." — OpenAI, EVMbench Launch Announcement
On February 18, 2026, OpenAI and Paradigm quietly released EVMbench — an open-source benchmark that tests how well AI agents can detect, exploit, and patch smart contract vulnerabilities. The timing was not accidental. Days earlier, a lending protocol called Moonwell lost $1.78 million because AI-generated code contained a critical oracle pricing error that passed both human review and a professional security audit. The irony was painful: the same technology that created the vulnerability is now being weaponized to prevent the next one.
EVMbench is built on 120 curated vulnerabilities from 40 real-world audits, most sourced from competitive Code4rena-style contests. The results are striking. OpenAI's GPT-5.3-Codex now exploits over 70% of critical smart contract bugs in controlled environments — up from less than 20% when the project started. But the benchmark also reveals a troubling asymmetry: AI agents are dramatically better at breaking contracts than fixing them. This gap defines the next chapter of DeFi security, and the $890 million audit industry that services it.
The economic stakes are enormous. Crypto theft reached $3.4 billion in 2025. January 2026 alone saw $370 million drained across 40 incidents. The audit market is projected to grow to $3.4 billion by 2033. And now, with Big Four firms PwC and Deloitte launching blockchain-focused AI audit divisions, the question is no longer whether AI will reshape smart contract security — it is who captures the economic value of doing so.
The Moonwell exploit of February 2026 has become the defining cautionary tale for what Dragonfly managing partner Haseeb Qureshi calls crypto's fundamental "human misalignment" problem. During the activation of governance proposal MIP-X43 — designed to integrate Chainlink's Oracle Extractable Value (OEV) wrapper contracts — AI-generated code introduced a devastating logic error. Rather than multiplying the cbETH/ETH exchange rate by the ETH/USD price feed, the code used the raw exchange ratio as if it were already denominated in dollars.
The result: cbETH, trading at approximately $2,200, was suddenly valued by the oracle at $1.12. Trading bots pounced. Liquidation cascades ripped through positions collateralized in cbETH, with attackers repaying approximately $1 of debt to seize 1,096 cbETH tokens per transaction. Net loss: $1.78 million.
What made the Moonwell incident uniquely alarming was not the loss amount — it was small by DeFi standards. It was the failure mode. The AI-generated code was syntactically perfect. It compiled cleanly. It passed basic unit tests. It even survived a professional audit from Halborn, a respected security firm. The vulnerability was a semantic error — the kind that looks correct at a glance but collapses under the adversarial conditions of live DeFi markets. This is the signature risk of "vibe coding": plausible-looking output that encodes fundamental misunderstandings of protocol logic.
OpenAI's response came within days. EVMbench was not a coincidence — it was a countermove.
EVMbench, developed collaboratively by OpenAI, Paradigm, and OtterSec, is the first standardized framework for evaluating AI agents on practical smart contract security tasks. It draws on 120 curated vulnerabilities from 40 audits, including competitive audit contests and scenarios from Paradigm's proprietary Tempo audit process.
The benchmark evaluates three distinct capabilities:
Detection Mode: Can the AI identify known vulnerabilities documented in professional audits? Scored on recall — the agent must locate security flaws that human auditors previously flagged. This measures the AI's ability to serve as a first-pass screening tool.
Exploit Mode: Can the AI construct a working proof-of-concept that drains funds in a sandboxed EVM environment? Measured by deterministic on-chain state changes — drained balances, triggered failure conditions. This is the most adversarial test.
Patch Mode: Can the AI fix vulnerable code without breaking the contract's intended functionality? This requires preserving correct behavior across edge cases and demands deep understanding of design assumptions. It is, by a significant margin, the hardest task.
The benchmark is open-source, available on GitHub with full documentation, harness tooling, and reproducible task definitions. OpenAI has committed $10 million in API credits through its Cybersecurity Grant Program specifically to support defensive smart contract research using EVMbench.
EVMbench's initial results reveal a competitive landscape where no single model dominates across all three tasks:
Detection Rankings (by average detect award):
| Rank | Model | Average Detect Award | |------|-------|---------------------| | 1st | Anthropic Claude Opus 4.6 | $37,824 | | 2nd | OpenAI GPT-5.2 | $31,623 | | 3rd | Google Gemini 3 Pro | $25,112 |
Exploit Performance:
| Model | Exploit Success Rate | |-------|---------------------| | GPT-5.3-Codex (via Codex CLI) | 72.2% | | GPT-5 (baseline) | 31.9% | | Earlier models (project start) | <20% |
Cross-Task Performance: GPT-5.3-Codex achieved the highest results in both patching and exploiting smart contracts, while Claude Opus 4.6 dominated detection. The improvement trajectory is steep — exploit success more than doubled between GPT-5 and GPT-5.3-Codex, suggesting rapid capability gains with each model generation.
The irony was not lost on the industry: Anthropic's Claude Opus 4.6 — the same model blamed for generating the vulnerable Moonwell code — scored highest in detecting vulnerabilities on OpenAI's own benchmark. This underscores a critical nuance that the industry must grapple with: AI models are simultaneously the source of new risks and the most powerful tool for mitigating them.
EVMbench's most important finding is not any single score — it is the structural asymmetry between AI's offensive and defensive capabilities.
As OpenAI stated in their research paper: "Agents perform best in the exploit setting, where the objective is explicit: continue iterating until funds are drained." In exploitation, the success condition is binary and unambiguous. The agent can brute-force creative attack vectors until one works. The feedback loop is immediate.
Patching, by contrast, requires something fundamentally different: understanding intent. Fixing a vulnerable contract means preserving correct behavior across every edge case while eliminating the flaw. A patch that breaks legitimate functionality is as dangerous as the original vulnerability. This demands the kind of contextual reasoning — understanding what a contract is supposed to do, not just what it does — that remains a frontier challenge for large language models.
This asymmetry has immediate economic consequences. If attackers can deploy AI agents that exploit 70%+ of known vulnerability classes, while defensive AI can only partially patch them, the advantage shifts decisively toward offense. The $3.4 billion in 2025 crypto theft could accelerate, not decline, as AI tools become more accessible.
OpenAI acknowledged this risk directly, cautioning that EVMbench "doesn't capture the true challenge of securing smart contracts" given its limited vulnerability sample. The benchmark cannot reliably distinguish false positives from genuine discoveries. It is a starting point, not a solution.
The smart contract audit market — valued at approximately $890 million in 2024 and projected to reach $3.4 billion by 2033 at a 24.5% CAGR — is entering a phase of rapid structural change. Three forces are converging:
1. AI-Augmented Auditing Is Already Here
OpenZeppelin has reported that its AI-powered audit tools cut auditing time by 50%. Polygon's development team documented a 40% reduction in pre-launch security testing cycles after implementing hybrid AI-human audit workflows. The average DeFi audit currently costs $50,000–$100,000 and takes 3–6 weeks. AI compression of these timelines threatens the pricing model of every major audit firm.
2. Traditional Finance Is Entering
In a signal that smart contract security is becoming a mainstream financial concern, both PwC and Deloitte launched blockchain-focused AI audit divisions in early 2026. This represents the first serious incursion by Big Four accounting firms into what was previously a crypto-native cottage industry dominated by firms like CertiK, Trail of Bits, and Halborn.
3. The "Continuous Audit" Model Is Emerging
EVMbench's open-source nature points toward a future where AI agents perform continuous, real-time security monitoring rather than one-time pre-deployment audits. This shifts the economic model from project-based consulting ($50K–$100K per engagement) to subscription-based monitoring — a fundamentally different business.
From an economic value perspective, the audit market represents one of the few genuinely revenue-generating sectors in the blockchain economy. Unlike validator rewards or token inflation — which comprise 85–90% of all blockchain value flows — audit fees are paid from real protocol budgets for measurable risk reduction. AI's disruption of this market is therefore not about speculation. It is about the reallocation of real economic value.
The January 2026 hack data paints a sobering picture of the threat landscape that EVMbench is designed to address:
The dominance of social engineering over code exploitation is itself significant. As Haseeb Qureshi observed: "Crypto's failure modes, which always made it feel broken for humans, in retrospect were never bugs. They were simply signs that we humans were the wrong users." His vision of "self-driving wallets" — AI agents that manage transactions, verify contracts, and reject suspicious interactions on behalf of human users — represents the logical endpoint of the EVMbench thesis: if humans cannot safely interact with smart contracts, and if AI is better at both attacking and defending them, then AI mediation becomes not a luxury but a necessity.
The $3.4 billion in 2025 theft losses represents approximately 25% of the blockchain industry's total on-chain fee revenue of $13.7 billion. In other words, for every $4 the industry earns, $1 is stolen. No traditional financial system would tolerate this ratio. AI-powered security is not an innovation narrative — it is an existential requirement.
EVMbench establishes the first standardized measure of AI smart contract security capability, creating accountability and comparability across models from OpenAI, Anthropic, and Google.
AI exploit success rates have more than tripled (from <20% to 72.2%) in recent model generations, meaning both attackers and defenders are racing up the same capability curve.
Detection and exploitation outpace patching, creating a structural asymmetry where AI can find and break contracts faster than it can fix them — a dynamic that favors offense.
The $890M+ audit market faces compression as AI tools cut audit times by 40–50%, while Big Four firms enter the space and subscription-based continuous monitoring emerges.
84% of January 2026 crypto losses came from social engineering, not code exploits, suggesting that AI-mediated wallets — not just AI-audited contracts — are the missing layer.
OpenAI's $10 million in API credits for defensive blockchain research signals that major AI labs view smart contract security as a strategic priority, not a side project.
EVMbench arrives at an inflection point. The Moonwell exploit demonstrated that AI-generated code can pass professional audits while harboring critical vulnerabilities. The benchmark's results demonstrate that AI can also detect, exploit, and — to a lesser degree — fix those same vulnerabilities. The technology is simultaneously the disease and the cure.
The economic question is not whether AI will transform smart contract security — the audit time compression and Big Four market entry have already answered that. The question is whether the defensive application of AI can outpace its offensive use before the next billion-dollar exploit. EVMbench's 72.2% exploit rate and its patching limitations suggest the race is far from won.
For an industry that generates $13.7 billion in on-chain revenue while losing $3.4 billion annually to theft, the margin for error is razor-thin. The protocols, auditors, and AI labs that close the exploit-patch asymmetry first will not just prevent losses — they will capture a disproportionate share of the trust premium that has always been blockchain's most valuable and most elusive asset.
The arms race between AI attackers and AI defenders has begun. EVMbench is the scoreboard.