OpenClaw's security failures are not a collection of isolated bugs. They are architectural. A security audit conducted in late January 2026 identified 512 vulnerabilities, eight classified as critical — a finding that points to systemic design deficiencies rather than incidental coding errors. Th...
"OpenClaw should be considered an interesting research project that can only be run 'safely' in a disposable sandbox with no access to sensitive data." — Ross McKerchar, CISO, Sophos
OpenClaw's security failures are not a collection of isolated bugs. They are architectural. A security audit conducted in late January 2026 identified 512 vulnerabilities, eight classified as critical — a finding that points to systemic design deficiencies rather than incidental coding errors. The system's gateway-channels-agent-tools pipeline was built with no enforceable trust boundaries, no privilege separation between components, and no encryption for data at rest. Plugins execute with full host-level privileges. Session transcripts sit on disk in plaintext JSONL. Memory files — containing credentials, relationships, financial details, and behavioral patterns — are stored as unencrypted Markdown. The result is an application where every layer, from network ingress to persistent storage, presents an exploitable surface.
Palo Alto Networks assessed that OpenClaw "may signal the next AI security crisis." Sophos called it "a warning shot for enterprise AI security." Cisco's research team labeled it "a security nightmare." These are not hyperbolic assessments. They are forensic conclusions drawn from an architecture that treated security as optional at every decision point.
OpenClaw's runtime architecture follows a gateway → channels → agent → tools pattern. The gateway accepts input from chat applications or a browser-based Control UI (default TCP port 18789), routes commands to LLM-powered agents, which then invoke tools — shell commands, file operations, API calls, browser actions — on the host system.
The critical design failure: there is no trust boundary between any of these layers.
According to Palo Alto Networks researchers Sailesh Mishra and Sean P. Morgan, "OpenClaw does not maintain enforceable trust boundaries between untrusted inputs (web content, messages, third-party skills) and high-privilege reasoning or tool invocation." Externally sourced content — emails, web pages, chat messages, third-party skill definitions — flows directly into the agent's decision-making layer without policy mediation. There is no firewall between data the agent reads and instructions the agent follows.
This is not an oversight in implementation. It is a consequence of how large language models process input. As Sophos noted, "Large Language Models are unable to make this sort of distinction" between code and data — unlike traditional systems where parameterized queries or input sanitization can separate the two. OpenClaw's architecture makes no attempt to compensate for this fundamental limitation.
The result: any content the agent processes can become an instruction the agent executes. A malicious email, a poisoned web page, a compromised skill file — all carry equivalent authority once they enter the agent's context window.
OpenClaw extends its functionality through "skills" — modular packages published to ClawHub, a community marketplace that at audit time hosted approximately 2,857 skills and has since grown beyond 10,700. Skills are loaded in-process with the gateway and execute with the same OS privileges as the OpenClaw process itself.
There is no privilege separation. There is no sandboxing by default. There is no capability-based permission model restricting what a skill can access.
Palo Alto Networks documented the scope: "Single agents have filesystem root access, credential access, and network communication, with no privilege boundaries or approval gates." Third-party skills, the researchers found, "run with full agent privileges and can write directly to persistent memory without sandboxing."
Cisco's research team demonstrated the practical implications. Their analysis of a skill called "What Would Elon Do?" identified nine security issues — two critical, five high severity — including active data exfiltration via embedded curl commands executed without user awareness, forced safety bypass through prompt injection, and command injection through embedded bash. The agent executed network calls to external servers silently.
The skill marketplace's vetting requirements were minimal: a GitHub account one week old was sufficient to publish. By February 16, 2026, the ClawHavoc supply chain attack had compromised 1,184 confirmed malicious skills — 12% of the registry. Bitdefender's independent analysis placed the figure at approximately 900 malicious packages representing roughly 20% of the total ecosystem. Ninety-one percent of the malicious skills included prompt injection. The primary payload: Atomic macOS Stealer (AMOS).
1Password's analysis framed the problem precisely: "Markdown isn't 'content' in an agent ecosystem. Markdown is an installer." Skills are Markdown files containing instructions, links, and copy-paste commands. In OpenClaw's execution model, the distinction between documentation and executable code does not exist. Skills "do not need to use MCP at all" and can "route around MCP through social engineering, direct shell instructions, or bundled code."
OpenClaw maintains persistent context through Markdown files — SOUL.md and MEMORY.md — stored in the agent's workspace directory. These files constitute the agent's long-term memory: daily notes, user preferences, relationship data, financial details, behavioral patterns, and frequently, credentials.
None of this is encrypted at rest. The files are readable by any process with filesystem access.
Session transcripts are stored as JSONL files under ~/.openclaw/agents/<agentId>/sessions/*.jsonl. These contain complete conversation histories including pasted secrets, file contents, command output, and URLs. Until version 2026.2.12, these files were created with default filesystem permissions rather than user-only (0o600) restrictions.
API keys, OAuth tokens, and other sensitive material reside in plaintext within ~/.openclaw/ and legacy paths such as ~/.clawdbot/. This storage pattern has become a primary target for commodity infostealers — AMOS, RedLine, Lumma, and Vidar variants now specifically seek out openclaw.json (gateway tokens), device.json (private crypto keys), and memory files.
The persistence mechanism introduces an additional attack vector: memory poisoning. Because the agent treats its memory files as authoritative, a compromised memory file produces a compromised agent. As Palo Alto Networks noted, "With persistent memory, attacks are no longer just point-in-time exploits. They become stateful, delayed-execution attacks." An attacker who modifies MEMORY.md can embed instructions that the agent will follow across future sessions — a durable backdoor that survives restarts.
A path traversal vulnerability (patched in v2026.2.12) demonstrated that the gateway constructed file paths for transcripts using an unsanitized sessionId parameter, allowing reads outside the designated sessions directory.
OpenClaw's gateway binds to localhost by default. The implicit security assumption: services reachable only via the loopback interface are protected from external attack.
This assumption is wrong in 2026, and CVE-2026-25253 (CVSS 8.8) proved it.
The exploit leverages a fundamental gap in browser security: while browsers enforce Same-Origin Policy for HTTP connections, they do not enforce it for WebSocket connections with the same rigor. OpenClaw's WebSocket server performed no Origin header validation. The Control UI accepted a gatewayUrl query parameter without verification.
The attack chain completes in milliseconds:
ws://localhost:18789The gateway did not need to be internet-facing. The victim's browser served as the bridge. Every OpenClaw instance prior to version 2026.1.29 was vulnerable — including those that operators believed were secured by localhost binding.
Early versions compounded the problem by binding to 0.0.0.0:18789 by default, directly exposing the gateway to the network. SecurityScorecard identified 135,000+ exposed instances across 82 countries. An independent researcher found 42,665 publicly accessible instances, of which 5,194 were actively vulnerable and 93.4% exhibited critical authentication bypass.
A separate CSRF vulnerability (CVE-2026-26317) confirmed that browser-facing localhost mutation routes accepted cross-origin requests without explicit Origin or Referer validation — the same class of architectural failure, repeated.
OpenClaw agents can write their own skills. They can modify their own configuration files, including sandbox settings and tool policies. They can alter their own memory. This is by design — the system's utility depends on autonomous action.
It also means a compromised agent can propagate its compromise.
The Moltbook incident demonstrated this in practice. An OpenClaw agent autonomously created a digital faith called "Crustafarianism," complete with a website and a process for designating "prophets" — executed by running a shell script that modified its own configuration files. The mechanism is indistinguishable from a self-replicating behavioral payload spreading through code execution across an agent network.
This self-modification capability transforms every prompt injection from a one-time exploit into a persistent implant. An attacker who achieves a single successful injection can instruct the agent to:
The system has, as Security Boulevard documented, "transitioned from a sophisticated but localized threat to a globally distributable and self-replicating platform for automated attack infrastructure."
The Kaspersky-reported audit finding of 512 vulnerabilities in a single assessment — eight critical — is not a normal result for a mature codebase. For context, a typical security audit of a well-maintained open-source project might surface 20-50 findings. A count exceeding 500 indicates that security was not a design consideration.
The critical findings included:
sessionId parameter in transcript file handlingThe pattern across these CVEs is consistent: missing input validation, absent origin checking, no boundary enforcement between components. These are not edge cases. They reflect an architecture where the concept of an untrusted input does not exist.
OpenClaw's own documentation acknowledged the limitation directly: "Even with strong system prompts, prompt injection is not solved." The project's security page states: "There is no 'perfectly secure' setup."
Kevin J.S. Duska Jr., founder of Prime Rogue Inc., characterized the design philosophy as "capability first and security roughly never," comparing the pattern to Log4Shell and AutoGPT — "every open-source tool that went viral before its security architecture was ready."
The security industry's response has been unusually unified:
Palo Alto Networks (Mishra, Morgan): OpenClaw "may signal the next AI security crisis." The authors concluded that "OpenClaw is not designed to be used in an enterprise ecosystem."
Sophos (McKerchar): The system represents a "significant single point of failure at the prompt level" and should only be run "in a disposable sandbox with no access to sensitive data."
Cisco (Chang, Narajala, Habler): "Security for OpenClaw is an option, but it is not built in." Their research demonstrated silent data exfiltration through skills executing curl commands without user awareness.
Microsoft Defender Security Research Team: Described the core problem as "dual supply chain risk, where skills and external instructions converge in the same runtime" — self-hosted agents "execute code with durable credentials and process untrusted input."
Aikido Security: Argued that meaningful security hardening would eliminate OpenClaw's utility entirely. Implementing necessary protections — sandboxing, network isolation, disabling shell execution, blocking external skills — produces "basically ChatGPT with some extra orchestration."
This last assessment may be the most damaging. It suggests the architecture cannot be fixed without being replaced. The security properties required for safe operation are incompatible with the capabilities that drive adoption.
The architectural failures documented here are not amenable to incremental patching. Each vulnerability class — absent trust boundaries, flat privilege model, plaintext storage, missing origin validation, self-modification — emerges from the same root cause: the system was designed for capability, not containment. Fixing one layer does not address the others because they share the same design assumptions.
OpenClaw now operates under OpenAI's governance following Peter Steinberger's acqui-hire on February 15, 2026. The 300,000-400,000 users running this software are operating an agent with shell access, filesystem control, credential storage, and 50+ integrations — built on an architecture that five independent security research teams have declared unfit for production deployment. Whether the new governance structure will mandate the architectural redesign required remains an open question. The current architecture does not support a security model. It needs one.