Keras Archive Extraction Path Traversal
CWD-Relative Validation Bypass in filter_safe_tarinfos / filter_safe_zipinfos Enables Arbitrary File Write in AI Pipelines
Keras versions prior to 3.14.0 contain a path traversal flaw in keras/src/utils/file_utils.py where filter_safe_tarinfos() and filter_safe_zipinfos() validate archive member paths against the process current working directory rather than the actual extraction destination. In Docker containers, CI/CD runners, and Jupyter notebooks where CWD is /, the validation boundary becomes the filesystem root — making every traversal path appear safe. A secondary bug causes the ZIP filter to crash silently on blocked entries. An attacker who controls the model archive delivered via keras.utils.get_file() can write files anywhere on the container filesystem. Fixed in Keras 3.14.0.
Technical Breakdown
Keras Archive Extraction CWD-Relative Path Validation Bypass (GHSA-hqp4-2352-xf5r)
In keras/src/utils/file_utils.py, both filter_safe_tarinfos() and filter_safe_zipinfos() resolve archive member paths against os.getcwd() — the process current working directory — rather than the actual extraction base_dir. When the process CWD is / (standard in Docker containers, CI/CD runners, and Jupyter kernels), the containment check resolved.startswith(base) becomes trivially true for all absolute paths, granting traversal access to the entire filesystem. An independent bug in the ZIP filter additionally throws an AttributeError when attempting to block a traversal entry, causing the filter to crash silently and extract all entries unchecked.
Root Cause
The filter functions call os.path.realpath(os.path.join(".", member_path)) — hardcoding "." as the base instead of receiving the caller-supplied base_dir. In any environment where os.getcwd() returns /, the containment check resolved.startswith("/") is unconditionally true, and the filter passes every archive entry regardless of traversal depth.
Vulnerable Pattern
Any call to keras.utils.get_file(origin=..., extract=True) against an attacker-controlled archive triggers the flaw. For ZIP archives the situation is worse: the broken AttributeError means traversal entries are never even evaluated — the filter fails open. Python 3.11 and earlier additionally lack the tarfile filter="data" fallback, leaving those runtimes entirely dependent on the broken Keras-level check.
Why AI Pipelines Are High-Value
Keras is the canonical deep learning frontend for TensorFlow, JAX, and PyTorch backends. keras.utils.get_file() is the standard mechanism for fetching pre-trained weights, embeddings, and datasets — commonly invoked in training scripts and fine-tuning pipelines running in Docker containers where CWD is /. Arbitrary file writes in these contexts allow an attacker to silently backdoor the Python environment, corrupt training data, or exfiltrate cloud credentials before a single training step runs.
The Fix
Keras 3.14.0 adds a base_dir parameter to both filter functions and passes the actual extraction destination from the call site. The ZIP AttributeError is also resolved. Upgrade with pip install "keras>=3.14.0". No workaround exists for prior versions — the fix requires source changes to the filter functions.
Attack Chain
Craft the Malicious Archive
The attacker crafts a tar.gz or ZIP file containing members with traversal paths — e.g., ../../usr/lib/python3/dist-packages/keras/__init__.py — or symlinks pointing outside the extraction root. The archive otherwise looks like a legitimate model weights file (correct filename, plausible size, matching the expected content type). For ZIP archives no crafting sophistication is needed at all — the AttributeError means any traversal entry silently extracts.
Audit every keras.utils.get_file() call in your codebase. Any call with extract=True that pulls from a URL you do not fully control is exploitable on Keras < 3.14.0.
Deliver via Model Hub or Compromised Artifact Source
The malicious archive is served from an attacker-controlled host, a compromised model hub mirror, a poisoned S3 bucket, or a dependency confusion attack against an internal artifact registry. Keras fetches the archive over HTTPS and optionally validates a caller-supplied file_hash — but if the training script omits the hash (the common case), there is no integrity check whatsoever before extraction begins.
Always supply file_hash in every keras.utils.get_file() call and pin it to a known-good SHA-256. Use a private artifact registry rather than pulling directly from public URLs in production training jobs.
Bypass Extraction Filter via CWD=/
Keras calls filter_safe_tarinfos() or filter_safe_zipinfos(). Both resolve paths against os.getcwd(). In Docker (WORKDIR defaults to /), CI/CD runners, and Jupyter kernels, CWD is /. The check resolved.startswith("/") is always True — every path, including /etc/passwd or /usr/lib/python3/..., passes. For ZIPs, the filter crashes before it can block anything.
Set an explicit WORKDIR /app (or equivalent non-root path) in all Docker images that run Keras. This limits the traversal base to /app, reducing blast radius even on unpatched versions.
Arbitrary File Write — Code Injection or Pivot
With traversal unrestricted the attacker writes to any path the process user can reach: overwriting a Python module in site-packages to persist a backdoor across container restarts (when shared volumes are mounted), injecting a malicious __init__.py into a frequently-imported package, appending to ~/.ssh/authorized_keys for persistent shell access, reading the Kubernetes service account token at /var/run/secrets/kubernetes.io/serviceaccount/token, or corrupting training datasets before the training loop begins.
After any get_file(extract=True) call on an untrusted archive, verify no files landed outside the cache dir: find /root/.keras -name "*.py" 2>/dev/null | xargs grep -l "import subprocess\|__import__\|exec(" 2>/dev/null. Any hit is a compromised environment — rebuild from a clean image.
Silent Persistence in MLOps Pipeline
Because the injected code executes as part of a legitimate training run, it blends into normal CI/CD output. Backdoored weights or poisoned dataset files can propagate through model versioning systems (MLflow, W&B, SageMaker Model Registry) to production inference endpoints. A compromised container image layer can be committed to an ECR/GCR registry if the CI/CD step rebuilds and pushes after training, propagating the compromise downstream to every deployment that pulls the image.
Treat any model artifact produced by a training run on a host that loaded an unverified archive as untrusted until a clean re-run is completed. Enforce image provenance checks (Cosign/Sigstore) in your model deployment pipeline.
Impact
Path traversal in model-loading infrastructure is especially damaging in AI contexts because the same process that loads model weights also holds cloud credentials, dataset access tokens, and container runtime privileges — one malicious archive can pivot across the entire training and deployment pipeline.
AI Pipeline & Model Integrity
- Pre-trained weights can be silently replaced with backdoored versions before training begins
- Injected code executes under the training process identity, with full access to all loaded secrets
- Dataset files on mounted volumes can be overwritten pre-training to poison model outputs
- Inference servers loading cached model weights may serve trojaned predictions post-compromise
Container & Cluster Compromise
- Writes to
site-packagespersist across container restarts on shared volumes - Kubernetes service account tokens at
/var/run/secrets/become readable via traversal - SSH key injection enables persistent shell access that survives credential rotation
- CI/CD runners compromised during training can pivot to production deployments
Affected Sectors
- Organizations running Keras-based fine-tuning or training in Docker or Kubernetes
- MLOps platforms auto-downloading public model weights (Hugging Face, TF Hub, Keras Hub)
- Jupyter-based research environments where CWD is always
/ - IBM Watson Discovery and any product embedding keras 3.x as a dependency
Defensive Tutorial
Upgrade to Keras 3.14.0
ImmediateRun pip install "keras>=3.14.0" in all environments and rebuild any Docker images that bake in a pinned Keras version. Verify with python -c "import keras; print(keras.__version__)". There is no workaround for prior versions — the fix requires changes to file_utils.py that cannot be replicated via configuration. Check transitive dependencies: IBM Watson Discovery, TensorFlow bundles, and any ML platform that ships its own Keras wheel may need separate updates.
Audit all keras.utils.get_file() calls with extract=True
Immediate
Find every vulnerable call site across your codebase: grep -rn "get_file" --include="*.py" . | grep "extract=True". For each result, verify the origin URL is fully controlled by your organization. Any call pulling from a public URL without a hardcoded file_hash is a live attack surface on unpatched Keras.
Pin file_hash on every archive download
Immediate
Add a SHA-256 hash to every get_file() call: keras.utils.get_file(origin=url, file_hash="sha256:abc123...", extract=True). Compute the hash of known-good archives with sha256sum modelweights.tar.gz. This ensures that even if the remote source is compromised, Keras will refuse to extract a tampered archive. Store hashes in version control alongside the training script — treat any hash update as a security review event.
Set explicit WORKDIR in all Docker images
Important
Replace any Docker image that uses the default WORKDIR / with an explicit non-root path: WORKDIR /workspace. This is a defense-in-depth measure that limits the blast radius of traversal bugs by ensuring os.getcwd() returns a subdirectory rather than the filesystem root — an attacker still bypasses the check, but can only reach paths within /workspace, not /usr/lib or /etc. This does not fix the vulnerability — upgrade to 3.14.0.
Route model weight downloads through a private artifact registry
ImportantReplace public-URL downloads with pulls from a controlled registry (AWS CodeArtifact, Artifactory, or a private S3 bucket with bucket policy). Mirror approved model archives to the registry once, scan them, record their SHA-256, and reference only the internal URL in training code. This eliminates the delivery vector entirely — a compromised public host cannot reach your training environment. Block outbound HTTPS from training containers to public model hubs at the network level.
Mount dataset and model volumes read-only where possible
ImportantIn Kubernetes, set readOnly: true on any volume mount containing training datasets or base model weights. In Docker, use -v /host/models:/models:ro. This prevents traversal writes from reaching those volumes even if the extraction bypass succeeds. The Keras cache directory (~/.keras/datasets/) still needs to be writable, but model source directories should not.
Implement content-addressable model storage with signed manifests
ImportantAdopt a model registry that enforces content-addressing (every artifact stored and referenced by its SHA-256 digest) and requires cryptographic signatures on model manifests before any training job can consume them. Tools: MLflow Model Registry with artifact signing, DVC with S3 content-addressed storage, or Hugging Face Hub with commit signing enabled. This makes supply chain tampering visible — any modification to a model archive changes its digest and breaks the signature, failing the download before extraction occurs.
Add Keras to SBOM tracking and subscribe to keras-team security advisories
ImportantCVE-2026-11816 is the third significant path traversal in Keras in 18 months (after CVE-2025-12060 and CVE-2025-12638). Generate SBOMs for all ML training images: pip install cyclonedx-bom && cyclonedx-py environment -o sbom.json. Ingest into Dependency-Track and configure alerts for any new keras advisory. Watch the keras-team security advisories feed directly.
References
- GitHub Security Advisory GHSA-hqp4-2352-xf5r — Keras Archive Extraction Path Traversal — Official advisory with affected versions and patch reference
- NVD Entry CVE-2026-11816 — NVD — CVSS 8.1 (AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:N), awaiting enrichment
-
Patch Commit
keras-team/keras@2465b66 — Fix: pass
base_dirto filter functions; resolve ZIPAttributeError - Bug Bounty Disclosure Huntr — Bounty a07e3983 — Original researcher disclosure on the Huntr AI/ML bug bounty platform
- Related Advisory GHSA-hjqc-jx6g-rwp9 — CVE-2025-12060 — Prior Keras TAR path traversal via symlinks, fixed in 3.12.0
- CWE Reference CWE-22 — Improper Limitation of a Pathname to a Restricted Directory ('Path Traversal')