SGLang SafeUnpickler Bypass — Unauthenticated RCE
Builtin Allowlist Gap in Pickle Deserialization Enables Code Execution on the Inference Server
SGLang, an open-source LLM inference framework, exposes an unauthenticated HTTP endpoint (/update_weights_from_tensor) that deserializes attacker-supplied pickle data when no API key is configured. Its SafeUnpickler policy is meant to restrict which Python builtins a payload can call, but the allowlist permits builtins.__import__ and builtins.getattr — letting an attacker chain the two to reach denied functions like os.system anyway, achieving full remote code execution. This is a bypass of a prior fix (CVE-2025-10164), and a sibling endpoint (/load_lora_adapter_from_tensors) shares the same flaw. No released version currently contains the fix; a PR addressing the allowlist gap is open upstream.
Technical Breakdown
SGLang Unauthenticated Pickle Deserialization via /update_weights_from_tensor
SGLang allows unauthenticated pickle deserialization through /update_weights_from_tensor when no auth keys are configured, and the SafeUnpickler policy can be bypassed because builtins.__import__ and builtins.getattr are resolvable, enabling code execution via pickle REDUCE.
--api-key configured on the SGLang server
Root Cause
SafeUnpickler restricts which Python builtins a deserialized pickle payload may reference, using an allow/deny-list. The list was too permissive: builtins.__import__ and builtins.getattr both resolve successfully, even though functions like os.system are individually denied. This CVE is tagged CWE-94 (Code Injection) in some vulnerability databases and CWE-502 (Deserialization of Untrusted Data) in others describing the same underlying pattern in a sibling SGLang CVE — both are defensible; this advisory treats the root cause as unsafe deserialization, with code injection as the resulting impact.
Vulnerable Pattern
An attacker crafts a pickle payload that calls __import__("os"), then getattr(os, "system"), reconstructing a blocked function reference from two individually-allowed builtins. Any allowlist-based sandboxing of pickle's REDUCE opcode is vulnerable to this class of composition attack unless it blocks the building blocks, not just the end targets.
Why AI Services Are High-Value
Inference servers like SGLang run with access to GPU infrastructure, model weights, and often the same network segment as other internal AI tooling. Unauthenticated RCE here isn't just a data leak — it's a foothold directly onto compute infrastructure that's expensive to compromise-detect and often under-monitored compared to traditional web-facing services.
Temporary Mitigation
No patched release exists yet. Configure an API key (or admin API key) on any SGLang deployment, and restrict network access to /update_weights_from_tensor and /load_lora_adapter_from_tensors to trusted callers only. Track sgl-project/sglang PR #30423 for the upstream fix.
Attack Chain
Recon: Identify an Unauthenticated SGLang Deployment
Attacker scans for SGLang inference servers with no API key configured — a common state in development, research, or fast-moving internal deployments.
Any inference server reachable without authentication should be treated as a critical finding in an asset inventory or exposure scan.
Craft a Malicious Pickle Payload
Attacker builds a pickle payload using the REDUCE opcode to call builtins.__import__("os") followed by builtins.getattr(os, "system") — reconstructing a denied function from two allowed primitives.
Static or dynamic scanning for outbound pickle payloads referencing __import__ or getattr in this pattern can flag exploitation attempts.
Deliver via /update_weights_from_tensor
Attacker sends the crafted payload directly to the unauthenticated endpoint, which deserializes it without validation.
Network-level restriction of this endpoint (firewall/proxy allowlisting) stops this even without an application-level fix.
Code Execution on the Inference Host
The reconstructed os.system call executes with the privileges of the SGLang process, giving the attacker a shell on the host running inference.
Running inference services under a least-privilege service account limits blast radius even if this step succeeds.
Post-Exploitation on AI Infrastructure
From an inference host, an attacker can pivot to GPU cluster credentials, model weights, or other internal services sharing the network segment.
Network segmentation between inference infrastructure and other internal systems limits how far a single compromised host can reach.
Impact
Unauthenticated RCE on an inference server is especially damaging in AI environments because these hosts sit closer to expensive, under-monitored compute infrastructure than a typical web application does.
Compute & Credential Theft
- GPU cluster credentials reachable from the compromised host
- Cloud provider metadata/credentials if running on managed infrastructure
- API keys for upstream model providers or internal services
- SSH keys or service account tokens stored on the host
Model & Data Integrity
- Ability to tamper with model weights being served
- Access to any training or inference data cached on the host
- Potential to poison responses served to downstream applications
- Loss of trust in any output produced while compromised
Operational Impact
- Full compromise of the inference host, not just the SGLang process
- Lateral movement risk to other AI infrastructure on the same network
- Difficulty detecting compromise given limited monitoring on inference-specific infrastructure
- Service disruption if the attacker uses access for denial-of-service
Defensive Tutorial
Configure an API key on every SGLang deployment
ImmediateSet --api-key (or the admin API key equivalent) on any SGLang server. This alone closes the unauthenticated access path this CVE depends on.
Restrict network access to the affected endpoints
ImmediateFirewall or reverse-proxy rule blocking external access to /update_weights_from_tensor and /load_lora_adapter_from_tensors except from trusted internal callers.
Audit for signs of prior exploitation
ImportantReview SGLang server logs for POST requests to /update_weights_from_tensor or /load_lora_adapter_from_tensors from unexpected sources, and check for unexpected child processes spawned by the SGLang process.
Rotate credentials reachable from inference hosts
ImportantAny GPU cluster credentials, cloud metadata tokens, or API keys accessible from the SGLang host should be rotated as a precaution, since exploitation prior to mitigation can't be ruled out retroactively.
Track and apply the upstream fix
ImportantMonitor sgl-project/sglang PR #30423 and upgrade as soon as a release containing the fix is available.
Run inference services under least privilege
ImportantService accounts for inference processes should have no more access to credentials, secrets, or network segments than the inference workload itself requires — limiting blast radius for this and future deserialization-class bugs.
References
- GitHub Security Advisory GHSA-m675-99qc-9233 — Official SGLang project advisory for this vulnerability.
- Upstream Fix PR sgl-project/sglang#30423 — Fix denying getattr/__import__ in SafeUnpickler builtins, addressing the root cause.
- CWE Reference CWE-502 — Deserialization of Untrusted Data