OpenClaude Sandbox Bypass via Model-Controlled dangerouslyDisableSandbox Parameter
LLM-settable flag in the BashTool schema flips sandbox enforcement off when combined with the permissive allowUnsandboxedCommands: true default, yielding host-level RCE from prompt injection
OpenClaude, the open-source coding-agent CLI from Gitlawb, ships versions prior to 0.5.1 with a critical sandbox-escape flaw. Its BashTool exposes a dangerouslyDisableSandbox parameter directly inside the schema the LLM itself populates — meaning the model, not just the operator, can turn sandboxing off. Combined with a default allowUnsandboxedCommands: true policy, a single prompt-injected tool call lets an attacker achieve full host-level command execution: reading credentials, writing arbitrary files, opening reverse shells, and pivoting into connected infrastructure. No operator error is required — only default configuration. Fixed in OpenClaude 0.5.1.
Technical Breakdown
OpenClaude Sandbox Bypass via Model-Controlled dangerouslyDisableSandbox Parameter
OpenClaude's BashTool input schema exposes a dangerouslyDisableSandbox boolean that the LLM itself can populate in its own tool_use response — a security-governing control placed inside the untrusted model's output surface rather than behind a trusted, operator-only control plane. shouldUseSandbox() returns false whenever that model-supplied flag is true AND areUnsandboxedCommandsAllowed() also passes — which it does by default. A prompt-injected model needs only emit one JSON field to escape the sandbox entirely.
sandbox.allowUnsandboxedCommands: false (inferred stopgap, not officially blessed — prioritize the upgrade)
Root Cause
The BashTool schema OpenClaude hands to the model includes a dangerouslyDisableSandbox parameter the LLM can set on its own tool_use calls. Sandbox enforcement is gated on this model-writable flag combined with settings?.sandbox?.allowUnsandboxedCommands ?? true — a default that resolves to permissive unless an operator explicitly overrides it. There is no trust boundary between "what the model can ask for" and "what actually disables isolation."
Vulnerable Pattern
Any agent framework that exposes a security-governing boolean inside a tool schema the model itself populates — rather than enforcing it exclusively from operator-side config — inherits this class of bug. The model is, by design, subject to prompt injection from any content it ingests; putting sandbox control in its hands makes that content attacker-controlled infrastructure policy.
Why AI Services Are High-Value
Coding agents like OpenClaude routinely hold SSH keys, cloud IAM credentials, kubeconfig, and CI/CD tokens in their execution environment because that's what makes them useful. A sandbox escape in this class of tool isn't a contained bug — it's a direct line from a single injected instruction to every credential the host can reach.
The Fix
OpenClaude 0.5.1 removes the model-writable escape hatch — dangerouslyDisableSandbox is no longer honored from the LLM's tool_use input, closing the bypass regardless of the allowUnsandboxedCommands setting. This is the only complete remediation.
Attack Chain
Injection vector placement
Attacker plants adversarial instructions in content the agent will ingest: a README, fetched web page, issue/PR body, API response, or repo file.
Defensive note: Treat all agent-ingested content as untrusted input; log and diff what content sources the agent consumed each session so injected content is forensically traceable.
Model coerced into emitting the bypass flag
Injected instructions manipulate the model's reasoning so its next BashTool tool_use call includes "dangerouslyDisableSandbox": true, disguised with a plausible justification.
Defensive note: Monitor/alert on any BashTool invocation carrying dangerous* flags — these should be rare-to-never in normal operation, a high-value detection signal.
Permissive default silently approves the bypass
shouldUseSandbox() evaluates the model-supplied flag against areUnsandboxedCommandsAllowed(), which defaults true unless explicitly overridden. Both conditions met with zero operator involvement.
Defensive note: Audit deployed configs for sandbox.allowUnsandboxedCommands — if unset or true, the instance is exposed regardless of other hardening.
Unsandboxed command executes with full host permissions
The Bash command runs outside the container/sandbox boundary, in the same process/user context as OpenClaude, giving filesystem read/write, env var access, network egress from the host.
Defensive note: Run agent processes under a least-privilege OS user, not a developer's own account or privileged service account, so a bypass is contained to a low-value identity.
Credential exfiltration, persistence, and lateral movement
Attacker harvests SSH keys, cloud tokens (AWS/GCP/Azure), kubeconfig, CI/CD secrets reachable from the host, optionally drops a reverse shell/scheduled task, pivots to reachable systems.
Defensive note: Rotate any credential readable from the host the moment compromise is suspected; egress-filter outbound connections from agent hosts so reverse-shell callbacks are blocked/alerted at the network layer.
Impact
Agentic coding tools sit on top of exactly the credentials an attacker wants — SSH keys, cloud IAM, CI/CD tokens — and this flaw turns any content the agent reads into a potential command-execution vector, with no operator misconfiguration required beyond accepting the shipped defaults.
Credential Theft
- SSH private keys enable lateral movement to any reachable server
- Cloud credentials (AWS/GCP/Azure) grant whatever IAM permissions the host identity holds
- Kubernetes kubeconfig can hand over cluster-admin access
- CI/CD and package-registry tokens enable supply-chain follow-on attacks
Remote Code Execution / Host Compromise
- Arbitrary shell commands with full process permissions, no privilege escalation needed since the sandbox WAS the boundary
- Reverse shells/C2 beacons for persistent access
- File read/write outside any project directory means the whole host filesystem is in scope
- If OpenClaude runs in shared CI, blast radius extends to every other job/tenant
Supply Chain / Trust Boundary Failure
- Letting an untrusted LLM control a security-critical config flag is a pattern likely repeated in other agent tooling
- Prompt injection is not hypothetical — any content the agent summarizes/reviews/builds against becomes a code-execution vector
- Orgs treating "the agent runs in a sandbox" as their sole isolation guarantee had it silently invalidated by default config
Defensive Tutorial
Upgrade to OpenClaude 0.5.1+
ImmediateRun npm ls openclaude @gitlawb/openclaude across repos/CI images to inventory exposed instances, then npm install openclaude@latest (pin >=0.5.1) and rebuild container images. The only complete fix — removes the model-writable flag entirely.
Set sandbox.allowUnsandboxedCommands: false explicitly
ImmediateSet this in every config as an immediate compensating control if you can't upgrade same-day.
Grep session/audit logs for the bypass flag
ImmediateSearch BashTool call logs for dangerouslyDisableSandbox or "dangerouslyDisableSandbox":true to check for prior exploitation.
Rotate credentials reachable from any host that ran pre-0.5.1 OpenClaude
ImmediateSSH keys, cloud IAM tokens, kubeconfig, CI/registry tokens.
Run OpenClaude under a dedicated least-privilege OS/service account
ImportantNot a developer's personal account or privileged CI runner identity.
Egress-filter agent hosts
ImportantRestrict outbound network access to an allowlist of required destinations (package registries, git remotes).
Add a config/CI gate that fails builds on unsafe agent settings
ImportantA script/policy check grepping deployed config for allowUnsandboxedCommands: true or absence of the key, failing the pipeline until explicitly false.
Treat all model-writable tool schemas as untrusted input in design review
ImportantNo security-governing parameter should ever appear in a schema the LLM itself populates; such controls belong exclusively in operator-side config.
Add OpenClaude to SBOM/dependency-tracking with CVE alerting
ImportantThis vendor had two other CVEs (CVE-2026-35570, CVE-2026-42073) disclosed in the same release cycle.
References
- NVD Entry CVE-2026-42074 — National Vulnerability Database record, CVSS 3.1 and 4.0 scoring.
- GitHub Security Advisory GHSA-m77w-p5jj-xmhg — Vendor disclosure with technical details and fix confirmation.
- Fix Pull Request Gitlawb/openclaude#778 — Patch removing the model-writable sandbox-disable flag.
- CWE Reference CWE-284 — Improper Access Control; CWE-306 — Missing Authentication for Critical Function