Critical · CVSS 9.8 September 11, 2026 · SGLang Project

SGLang SafeUnpickler Bypass — Unauthenticated RCE

Builtin Allowlist Gap in Pickle Deserialization Enables Code Execution on the Inference Server

SGLang, an open-source LLM inference framework, exposes an unauthenticated HTTP endpoint (/update_weights_from_tensor) that deserializes attacker-supplied pickle data when no API key is configured. Its SafeUnpickler policy is meant to restrict which Python builtins a payload can call, but the allowlist permits builtins.__import__ and builtins.getattr — letting an attacker chain the two to reach denied functions like os.system anyway, achieving full remote code execution. This is a bypass of a prior fix (CVE-2025-10164), and a sibling endpoint (/load_lora_adapter_from_tensors) shares the same flaw. No released version currently contains the fix; a PR addressing the allowlist gap is open upstream.

1 CVE
9.8 CVSS Score
RCE Class
No-Auth Access Required

Technical Breakdown

CVE-2026-86793 Critical · CVSS 9.8 Unauthenticated RCE

SGLang Unauthenticated Pickle Deserialization via /update_weights_from_tensor

SGLang allows unauthenticated pickle deserialization through /update_weights_from_tensor when no auth keys are configured, and the SafeUnpickler policy can be bypassed because builtins.__import__ and builtins.getattr are resolvable, enabling code execution via pickle REDUCE.

CWE: CWE-502 — Deserialization of Untrusted Data  |  Fixed in: No released version yet — fix pending in sgl-project/sglang PR #30423  |  Workaround: --api-key configured on the SGLang server

Root Cause

SafeUnpickler restricts which Python builtins a deserialized pickle payload may reference, using an allow/deny-list. The list was too permissive: builtins.__import__ and builtins.getattr both resolve successfully, even though functions like os.system are individually denied. This CVE is tagged CWE-94 (Code Injection) in some vulnerability databases and CWE-502 (Deserialization of Untrusted Data) in others describing the same underlying pattern in a sibling SGLang CVE — both are defensible; this advisory treats the root cause as unsafe deserialization, with code injection as the resulting impact.

Vulnerable Pattern

An attacker crafts a pickle payload that calls __import__("os"), then getattr(os, "system"), reconstructing a blocked function reference from two individually-allowed builtins. Any allowlist-based sandboxing of pickle's REDUCE opcode is vulnerable to this class of composition attack unless it blocks the building blocks, not just the end targets.

Why AI Services Are High-Value

Inference servers like SGLang run with access to GPU infrastructure, model weights, and often the same network segment as other internal AI tooling. Unauthenticated RCE here isn't just a data leak — it's a foothold directly onto compute infrastructure that's expensive to compromise-detect and often under-monitored compared to traditional web-facing services.

Temporary Mitigation

No patched release exists yet. Configure an API key (or admin API key) on any SGLang deployment, and restrict network access to /update_weights_from_tensor and /load_lora_adapter_from_tensors to trusted callers only. Track sgl-project/sglang PR #30423 for the upstream fix.

Attack Chain

1

Recon: Identify an Unauthenticated SGLang Deployment

Attacker scans for SGLang inference servers with no API key configured — a common state in development, research, or fast-moving internal deployments.

Any inference server reachable without authentication should be treated as a critical finding in an asset inventory or exposure scan.

2

Craft a Malicious Pickle Payload

Attacker builds a pickle payload using the REDUCE opcode to call builtins.__import__("os") followed by builtins.getattr(os, "system") — reconstructing a denied function from two allowed primitives.

Static or dynamic scanning for outbound pickle payloads referencing __import__ or getattr in this pattern can flag exploitation attempts.

3

Deliver via /update_weights_from_tensor

Attacker sends the crafted payload directly to the unauthenticated endpoint, which deserializes it without validation.

Network-level restriction of this endpoint (firewall/proxy allowlisting) stops this even without an application-level fix.

4

Code Execution on the Inference Host

The reconstructed os.system call executes with the privileges of the SGLang process, giving the attacker a shell on the host running inference.

Running inference services under a least-privilege service account limits blast radius even if this step succeeds.

5

Post-Exploitation on AI Infrastructure

From an inference host, an attacker can pivot to GPU cluster credentials, model weights, or other internal services sharing the network segment.

Network segmentation between inference infrastructure and other internal systems limits how far a single compromised host can reach.

Impact

Unauthenticated RCE on an inference server is especially damaging in AI environments because these hosts sit closer to expensive, under-monitored compute infrastructure than a typical web application does.

Compute & Credential Theft

  • GPU cluster credentials reachable from the compromised host
  • Cloud provider metadata/credentials if running on managed infrastructure
  • API keys for upstream model providers or internal services
  • SSH keys or service account tokens stored on the host

Model & Data Integrity

  • Ability to tamper with model weights being served
  • Access to any training or inference data cached on the host
  • Potential to poison responses served to downstream applications
  • Loss of trust in any output produced while compromised

Operational Impact

  • Full compromise of the inference host, not just the SGLang process
  • Lateral movement risk to other AI infrastructure on the same network
  • Difficulty detecting compromise given limited monitoring on inference-specific infrastructure
  • Service disruption if the attacker uses access for denial-of-service

Defensive Tutorial

IMMEDIATE · 0–24 HRS

Configure an API key on every SGLang deployment

Immediate

Set --api-key (or the admin API key equivalent) on any SGLang server. This alone closes the unauthenticated access path this CVE depends on.

IMMEDIATE · 0–24 HRS

Restrict network access to the affected endpoints

Immediate

Firewall or reverse-proxy rule blocking external access to /update_weights_from_tensor and /load_lora_adapter_from_tensors except from trusted internal callers.

SHORT-TERM · 1–7 DAYS

Audit for signs of prior exploitation

Important

Review SGLang server logs for POST requests to /update_weights_from_tensor or /load_lora_adapter_from_tensors from unexpected sources, and check for unexpected child processes spawned by the SGLang process.

SHORT-TERM · 1–7 DAYS

Rotate credentials reachable from inference hosts

Important

Any GPU cluster credentials, cloud metadata tokens, or API keys accessible from the SGLang host should be rotated as a precaution, since exploitation prior to mitigation can't be ruled out retroactively.

LONG-TERM

Track and apply the upstream fix

Important

Monitor sgl-project/sglang PR #30423 and upgrade as soon as a release containing the fix is available.

LONG-TERM

Run inference services under least privilege

Important

Service accounts for inference processes should have no more access to credentials, secrets, or network segments than the inference workload itself requires — limiting blast radius for this and future deserialization-class bugs.

References

  • GitHub Security Advisory GHSA-m675-99qc-9233 — Official SGLang project advisory for this vulnerability.
  • Upstream Fix PR sgl-project/sglang#30423 — Fix denying getattr/__import__ in SafeUnpickler builtins, addressing the root cause.
  • CWE Reference CWE-502 — Deserialization of Untrusted Data
Advisory analysis by Spectreworks AI. Original research by the SGLang project. All defensive recommendations are based on publicly available disclosure information. Verify patch applicability against your specific deployment before production changes.