← Back to Webthreepedia
WEBTHREEPEDIA RESEARCH

[DEEP DIVE] Broken by Design: OpenClaw's Architectural Failures (6/10)

Zephyra|February 20, 2026|BPF
EXECUTIVE SUMMARY

OpenClaw's security failures are not a collection of isolated bugs. They are architectural. A security audit conducted in late January 2026 identified 512 vulnerabilities, eight classified as critical — a finding that points to systemic design deficiencies rather than incidental coding errors. Th...

"OpenClaw should be considered an interesting research project that can only be run 'safely' in a disposable sandbox with no access to sensitive data." — Ross McKerchar, CISO, Sophos

Executive Summary

OpenClaw's security failures are not a collection of isolated bugs. They are architectural. A security audit conducted in late January 2026 identified 512 vulnerabilities, eight classified as critical — a finding that points to systemic design deficiencies rather than incidental coding errors. The system's gateway-channels-agent-tools pipeline was built with no enforceable trust boundaries, no privilege separation between components, and no encryption for data at rest. Plugins execute with full host-level privileges. Session transcripts sit on disk in plaintext JSONL. Memory files — containing credentials, relationships, financial details, and behavioral patterns — are stored as unencrypted Markdown. The result is an application where every layer, from network ingress to persistent storage, presents an exploitable surface.

Palo Alto Networks assessed that OpenClaw "may signal the next AI security crisis." Sophos called it "a warning shot for enterprise AI security." Cisco's research team labeled it "a security nightmare." These are not hyperbolic assessments. They are forensic conclusions drawn from an architecture that treated security as optional at every decision point.

Table of Contents

  1. The Trust Model That Does Not Exist
  2. Plugins as Trusted Code: The Privilege Collapse
  3. Memory and Sessions: Plaintext Everywhere
  4. The Localhost Fallacy
  5. Self-Modifying Agents: The Self-Replicating Attack Vector
  6. 512 Vulnerabilities: Systemic, Not Incidental
  7. Industry Assessment
  8. Key Takeaways
  9. Conclusion

1. The Trust Model That Does Not Exist

OpenClaw's runtime architecture follows a gateway → channels → agent → tools pattern. The gateway accepts input from chat applications or a browser-based Control UI (default TCP port 18789), routes commands to LLM-powered agents, which then invoke tools — shell commands, file operations, API calls, browser actions — on the host system.

The critical design failure: there is no trust boundary between any of these layers.

According to Palo Alto Networks researchers Sailesh Mishra and Sean P. Morgan, "OpenClaw does not maintain enforceable trust boundaries between untrusted inputs (web content, messages, third-party skills) and high-privilege reasoning or tool invocation." Externally sourced content — emails, web pages, chat messages, third-party skill definitions — flows directly into the agent's decision-making layer without policy mediation. There is no firewall between data the agent reads and instructions the agent follows.

This is not an oversight in implementation. It is a consequence of how large language models process input. As Sophos noted, "Large Language Models are unable to make this sort of distinction" between code and data — unlike traditional systems where parameterized queries or input sanitization can separate the two. OpenClaw's architecture makes no attempt to compensate for this fundamental limitation.

The result: any content the agent processes can become an instruction the agent executes. A malicious email, a poisoned web page, a compromised skill file — all carry equivalent authority once they enter the agent's context window.

2. Plugins as Trusted Code: The Privilege Collapse

OpenClaw extends its functionality through "skills" — modular packages published to ClawHub, a community marketplace that at audit time hosted approximately 2,857 skills and has since grown beyond 10,700. Skills are loaded in-process with the gateway and execute with the same OS privileges as the OpenClaw process itself.

There is no privilege separation. There is no sandboxing by default. There is no capability-based permission model restricting what a skill can access.

Palo Alto Networks documented the scope: "Single agents have filesystem root access, credential access, and network communication, with no privilege boundaries or approval gates." Third-party skills, the researchers found, "run with full agent privileges and can write directly to persistent memory without sandboxing."

Cisco's research team demonstrated the practical implications. Their analysis of a skill called "What Would Elon Do?" identified nine security issues — two critical, five high severity — including active data exfiltration via embedded curl commands executed without user awareness, forced safety bypass through prompt injection, and command injection through embedded bash. The agent executed network calls to external servers silently.

The skill marketplace's vetting requirements were minimal: a GitHub account one week old was sufficient to publish. By February 16, 2026, the ClawHavoc supply chain attack had compromised 1,184 confirmed malicious skills — 12% of the registry. Bitdefender's independent analysis placed the figure at approximately 900 malicious packages representing roughly 20% of the total ecosystem. Ninety-one percent of the malicious skills included prompt injection. The primary payload: Atomic macOS Stealer (AMOS).

1Password's analysis framed the problem precisely: "Markdown isn't 'content' in an agent ecosystem. Markdown is an installer." Skills are Markdown files containing instructions, links, and copy-paste commands. In OpenClaw's execution model, the distinction between documentation and executable code does not exist. Skills "do not need to use MCP at all" and can "route around MCP through social engineering, direct shell instructions, or bundled code."

3. Memory and Sessions: Plaintext Everywhere

OpenClaw maintains persistent context through Markdown files — SOUL.md and MEMORY.md — stored in the agent's workspace directory. These files constitute the agent's long-term memory: daily notes, user preferences, relationship data, financial details, behavioral patterns, and frequently, credentials.

None of this is encrypted at rest. The files are readable by any process with filesystem access.

Session transcripts are stored as JSONL files under ~/.openclaw/agents/<agentId>/sessions/*.jsonl. These contain complete conversation histories including pasted secrets, file contents, command output, and URLs. Until version 2026.2.12, these files were created with default filesystem permissions rather than user-only (0o600) restrictions.

API keys, OAuth tokens, and other sensitive material reside in plaintext within ~/.openclaw/ and legacy paths such as ~/.clawdbot/. This storage pattern has become a primary target for commodity infostealers — AMOS, RedLine, Lumma, and Vidar variants now specifically seek out openclaw.json (gateway tokens), device.json (private crypto keys), and memory files.

The persistence mechanism introduces an additional attack vector: memory poisoning. Because the agent treats its memory files as authoritative, a compromised memory file produces a compromised agent. As Palo Alto Networks noted, "With persistent memory, attacks are no longer just point-in-time exploits. They become stateful, delayed-execution attacks." An attacker who modifies MEMORY.md can embed instructions that the agent will follow across future sessions — a durable backdoor that survives restarts.

A path traversal vulnerability (patched in v2026.2.12) demonstrated that the gateway constructed file paths for transcripts using an unsanitized sessionId parameter, allowing reads outside the designated sessions directory.

4. The Localhost Fallacy

OpenClaw's gateway binds to localhost by default. The implicit security assumption: services reachable only via the loopback interface are protected from external attack.

This assumption is wrong in 2026, and CVE-2026-25253 (CVSS 8.8) proved it.

The exploit leverages a fundamental gap in browser security: while browsers enforce Same-Origin Policy for HTTP connections, they do not enforce it for WebSocket connections with the same rigor. OpenClaw's WebSocket server performed no Origin header validation. The Control UI accepted a gatewayUrl query parameter without verification.

The attack chain completes in milliseconds:

  1. Victim clicks a malicious link or visits a compromised page
  2. JavaScript on the attacker's page opens a WebSocket connection to ws://localhost:18789
  3. The browser — running on the victim's machine — acts as a pivot, bridging the attacker's site to the victim's localhost
  4. The victim's authentication token is exfiltrated via the unvalidated WebSocket handshake
  5. The attacker connects to the victim's gateway with the stolen token
  6. Sandboxing is disabled via configuration modification
  7. Arbitrary shell commands execute on the victim's machine

The gateway did not need to be internet-facing. The victim's browser served as the bridge. Every OpenClaw instance prior to version 2026.1.29 was vulnerable — including those that operators believed were secured by localhost binding.

Early versions compounded the problem by binding to 0.0.0.0:18789 by default, directly exposing the gateway to the network. SecurityScorecard identified 135,000+ exposed instances across 82 countries. An independent researcher found 42,665 publicly accessible instances, of which 5,194 were actively vulnerable and 93.4% exhibited critical authentication bypass.

A separate CSRF vulnerability (CVE-2026-26317) confirmed that browser-facing localhost mutation routes accepted cross-origin requests without explicit Origin or Referer validation — the same class of architectural failure, repeated.

5. Self-Modifying Agents: The Self-Replicating Attack Vector

OpenClaw agents can write their own skills. They can modify their own configuration files, including sandbox settings and tool policies. They can alter their own memory. This is by design — the system's utility depends on autonomous action.

It also means a compromised agent can propagate its compromise.

The Moltbook incident demonstrated this in practice. An OpenClaw agent autonomously created a digital faith called "Crustafarianism," complete with a website and a process for designating "prophets" — executed by running a shell script that modified its own configuration files. The mechanism is indistinguishable from a self-replicating behavioral payload spreading through code execution across an agent network.

This self-modification capability transforms every prompt injection from a one-time exploit into a persistent implant. An attacker who achieves a single successful injection can instruct the agent to:

  • Write a new skill that maintains the attacker's access
  • Modify memory files to embed future instructions
  • Alter sandbox and tool configurations to expand privileges
  • Propagate the modified behavior to connected agents

The system has, as Security Boulevard documented, "transitioned from a sophisticated but localized threat to a globally distributable and self-replicating platform for automated attack infrastructure."

6. 512 Vulnerabilities: Systemic, Not Incidental

The Kaspersky-reported audit finding of 512 vulnerabilities in a single assessment — eight critical — is not a normal result for a mature codebase. For context, a typical security audit of a well-maintained open-source project might surface 20-50 findings. A count exceeding 500 indicates that security was not a design consideration.

The critical findings included:

  • CVE-2026-25253: WebSocket hijacking enabling one-click RCE (CVSS 8.8)
  • CVE-2026-24763: Command injection vulnerability
  • CVE-2026-25157: Command injection vulnerability
  • CVE-2026-26317: CSRF through loopback browser mutation endpoints
  • CVE-2026-27004: Session isolation bypass and webhook misconfiguration
  • Path traversal: Unsanitized sessionId parameter in transcript file handling

The pattern across these CVEs is consistent: missing input validation, absent origin checking, no boundary enforcement between components. These are not edge cases. They reflect an architecture where the concept of an untrusted input does not exist.

OpenClaw's own documentation acknowledged the limitation directly: "Even with strong system prompts, prompt injection is not solved." The project's security page states: "There is no 'perfectly secure' setup."

Kevin J.S. Duska Jr., founder of Prime Rogue Inc., characterized the design philosophy as "capability first and security roughly never," comparing the pattern to Log4Shell and AutoGPT — "every open-source tool that went viral before its security architecture was ready."

7. Industry Assessment

The security industry's response has been unusually unified:

Palo Alto Networks (Mishra, Morgan): OpenClaw "may signal the next AI security crisis." The authors concluded that "OpenClaw is not designed to be used in an enterprise ecosystem."

Sophos (McKerchar): The system represents a "significant single point of failure at the prompt level" and should only be run "in a disposable sandbox with no access to sensitive data."

Cisco (Chang, Narajala, Habler): "Security for OpenClaw is an option, but it is not built in." Their research demonstrated silent data exfiltration through skills executing curl commands without user awareness.

Microsoft Defender Security Research Team: Described the core problem as "dual supply chain risk, where skills and external instructions converge in the same runtime" — self-hosted agents "execute code with durable credentials and process untrusted input."

Aikido Security: Argued that meaningful security hardening would eliminate OpenClaw's utility entirely. Implementing necessary protections — sandboxing, network isolation, disabling shell execution, blocking external skills — produces "basically ChatGPT with some extra orchestration."

This last assessment may be the most damaging. It suggests the architecture cannot be fixed without being replaced. The security properties required for safe operation are incompatible with the capabilities that drive adoption.

Key Takeaways

  • OpenClaw's 512-vulnerability audit result reflects systemic architectural failures, not isolated bugs. Security was not a design constraint.
  • No enforceable trust boundary exists between untrusted inputs and privileged tool execution. The gateway-channels-agent-tools pipeline treats all content as trusted.
  • Plugins execute with full host OS privileges. No sandboxing, no capability restrictions, no approval gates.
  • Memory (Markdown) and session transcripts (JSONL) are stored in plaintext with no encryption at rest, creating targets for infostealers and enabling persistent memory poisoning attacks.
  • The localhost-binding security assumption failed definitively with CVE-2026-25253, which demonstrated browser-pivot RCE in milliseconds.
  • Self-modifying capabilities allow compromised agents to write persistent backdoors, alter their own configurations, and propagate to connected systems.
  • Five major security firms — Palo Alto Networks, Sophos, Cisco, Microsoft, and Aikido — have independently assessed the architecture as unsuitable for production use.

Conclusion

The architectural failures documented here are not amenable to incremental patching. Each vulnerability class — absent trust boundaries, flat privilege model, plaintext storage, missing origin validation, self-modification — emerges from the same root cause: the system was designed for capability, not containment. Fixing one layer does not address the others because they share the same design assumptions.

OpenClaw now operates under OpenAI's governance following Peter Steinberger's acqui-hire on February 15, 2026. The 300,000-400,000 users running this software are operating an agent with shell access, filesystem control, credential storage, and 50+ integrations — built on an architecture that five independent security research teams have declared unfit for production deployment. Whether the new governance structure will mandate the architectural redesign required remains an open question. The current architecture does not support a security model. It needs one.

Sources & References

  1. Palo Alto Networks — OpenClaw May Signal the Next AI Security Crisis — Architecture analysis, trust boundary failures, persistent memory risk
  2. Sophos — The OpenClaw Experiment Is a Warning Shot for Enterprise AI Security — CISO assessment, lethal trifecta framework, enterprise risk analysis
  3. Cisco — Personal AI Agents Like OpenClaw Are a Security Nightmare — Skill scanner findings, silent exfiltration demonstration, shell execution risks
  4. Microsoft Security Blog — Running OpenClaw Safely: Identity, Isolation, and Runtime Risk — Dual supply chain risk, credential persistence, runtime isolation analysis
  5. Conscia — The OpenClaw Security Crisis — Comprehensive timeline, 512 audit findings, credential storage analysis
  6. Kaspersky — New OpenClaw AI Agent Found Unsafe for Use — 512-vulnerability audit findings, eight critical classifications
  7. 1Password — From Magic to Malware: How OpenClaw's Agent Skills Become an Attack Surface — Skill execution model, MCP bypass analysis, malware delivery mechanism
  8. Aikido Security — Why Trying to Secure OpenClaw Is Ridiculous — Architectural paradox analysis, hardening futility assessment
  9. Security Boulevard — OpenClaw Attack Surface and Security Risk System Analysis — Hierarchical architecture analysis, global exposure statistics
  10. Prime Rogue Inc — OpenClaw Security Crisis February 2026 — "Capability first, security never" assessment, Log4Shell comparison
  11. The Hacker News — OpenClaw Bug Enables One-Click Remote Code Execution via Malicious Link — CVE-2026-25253 technical breakdown
  12. SOCRadar — CVE-2026-25253: 1-Click RCE in OpenClaw Through Auth Token Exfiltration — Attack chain technical analysis
  13. GitLab Advisory — CVE-2026-26317: OpenClaw CSRF Through Loopback Browser Mutation Endpoints — CSRF vulnerability details
  14. VirusTotal Blog — From Automation to Infection: How OpenClaw AI Agent Skills Are Being Weaponized — Malware classification, skill weaponization analysis