Bleeding Llama in Ollama
Defending Local LLM Runtimes from Unauthenticated Process-Memory Disclosure
Cyera Research disclosed an unauthenticated process-memory disclosure vulnerability in Ollama, one of the most widely deployed local LLM runtimes. An attacker with network access to an exposed Ollama instance can upload a crafted GGUF model file to trigger a heap out-of-bounds read, leaking unintended process memory — including user prompts, system prompts, API keys, and environment variables. With roughly 170,000 GitHub stars, more than 100 million Docker Hub downloads, and an estimated 300,000 servers globally, the exposure surface is significant.
"Local LLM runtimes can expose high-value secrets even when the model itself is not compromised."
— Cyera Research
Technical Breakdown
The vulnerability lies in Ollama's GGUF model file parser, written in unsafe Go code. Crafted model files trigger a heap out-of-bounds read during quantization, surfacing unintended process memory in the API response.
Ollama GGUF Model Parsing Heap Out-of-Bounds Read
Ollama's unsafe Go code creates a heap out-of-bounds read condition during GGUF model parsing and quantization. An unauthenticated attacker with access to the blob upload API can upload a crafted GGUF file, then trigger model creation via /api/create to cause the server to read and return unintended heap memory. Since Ollama runs without authentication by default and frequently binds to 0.0.0.0, network-exposed instances are reachable without any credentials.
Root Cause: Unsafe Go Parsing
The GGUF model parser is implemented with unsafe Go code that does not validate model file structure before heap allocation. Attacker-controlled fields in the GGUF header control memory read offsets, creating a classic out-of-bounds read primitive.
Default Exposure: No Auth, Any Interface
Ollama ships with authentication disabled by default and frequently binds to all interfaces (0.0.0.0) when installed in server or Docker configurations. This means any network-accessible Ollama instance is an unauthenticated attack target without additional configuration.
What Leaks
Heap memory at the time of model processing may contain anything the Ollama process has handled: user chat prompts, system prompts, API keys passed via environment variables, cloud credentials, internal service tokens, and configuration data from connected agents or automation systems.
Scale of Exposure
Ollama has approximately 170,000 GitHub stars and more than 100 million Docker Hub downloads. Cyera Research estimated the issue could affect 300,000 servers globally — many of which are developer workstations, internal AI services, and enterprise LLM deployments exposed on internal or cloud networks without authentication.
Attack Path
Identify Exposed Ollama Instance
Attacker locates a network-accessible Ollama API service — on a cloud host, internal network, or developer workstation — bound to a non-loopback interface with no authentication required.
Ollama instances are trivially discoverable via Shodan, internal network scanning, or misconfigured cloud security groups. Treat all Ollama deployments as network-accessible until proven otherwise.
Upload Crafted GGUF Model via Blob API
Attacker uploads a specially crafted GGUF model file through Ollama's blob upload endpoint. No authentication is required. The malicious file contains header fields that control heap memory read offsets during subsequent parsing.
GGUF files are treated as trusted model artifacts. Restrict who can upload models via the blob API — especially in shared or production environments.
Trigger Model Creation to Cause Memory Read
Attacker triggers model creation or conversion via /api/create, causing Ollama to parse the crafted GGUF file. The out-of-bounds read fires during quantization, surfacing heap memory outside the intended model buffer.
Disable /api/create and blob-upload workflows on instances that do not require model ingestion from untrusted sources.
Retrieve Leaked Process Memory
The API response contains leaked heap data — user prompts, system prompts, API keys, environment variables, cloud credentials, and internal service tokens that have passed through the Ollama process. Attacker exfiltrates this material for immediate or downstream use.
Even a single successful read may expose credentials sufficient to compromise connected SaaS platforms, cloud infrastructure, or internal services. Rotate all secrets reachable from the affected host.
Impact
Process memory disclosure from an LLM runtime is particularly high-value because inference processes handle a wide variety of sensitive material in the normal course of operation — prompts, credentials, configuration, and context from connected agents and automation pipelines.
Directly Leaked
- User chat prompts
- System prompts and instructions
- API keys (OpenAI, Anthropic, etc.)
- Environment variables
Secondary Exposure
- Cloud credentials (AWS, GCP, Azure)
- Internal service tokens
- Repository access tokens
- Agent configuration and context
Downstream Risk
- SaaS platform compromise
- Cloud infrastructure access
- Source code repository access
- Lateral movement via stolen credentials
No authentication required: The default Ollama configuration ships without authentication and frequently binds to all network interfaces. An attacker needs only network reachability — no credentials, no prior access, no social engineering. The ~300,000 globally estimated exposed servers represent a large unauthenticated attack surface.
Defensive Tutorial
0–24 Hours
Upgrade all Ollama deployments to the patched version immediately. Restrict network exposure — bind Ollama to loopback only or enforce firewall rules preventing external access. Disable /api/create and blob-upload endpoints on instances that do not need them. Rotate all secrets reachable from affected hosts — API keys, cloud credentials, repository tokens, and service tokens.
1–7 Days
Inventory all Ollama deployments across workstations, labs, services, and containers. Isolate instances with access to sensitive credentials or high-value systems. Audit environment variables loaded into Ollama's process — remove broad credentials and replace with scoped, short-lived alternatives. Implement logging for model uploads and creation activity to detect future exploitation attempts.
Long-Term Governance
Treat local LLM runtimes the same way you treat databases, CI runners, and API gateways — as infrastructure with defined ownership, exposure policies, authentication requirements, logging standards, and patch timelines. Adopt model-artifact intake policies that treat all external GGUF files as untrusted input until validated. Do not allow Ollama processes to inherit broad developer credentials or production secrets. Use workload identity and scoped credentials.
Response Checklist
Apply the patch across all environments — production, staging, development, and developer workstations. A CVSS 9.1 unauthenticated disclosure requires same-day treatment regardless of environment classification.
Ensure all Ollama instances bind exclusively to 127.0.0.1, not 0.0.0.0. Set OLLAMA_HOST=127.0.0.1 in the environment. Verify with netstat -tlnp — binding address is not always visible in config files.
Instances that serve inference only (no model ingestion from external sources) should block the /api/create and blob-upload workflows at the reverse proxy or firewall level. This removes the primary exploitation vector without affecting inference availability.
Rotate API keys (OpenAI, Anthropic, Hugging Face, etc.), cloud credentials (AWS IAM, GCP service accounts, Azure managed identities), repository tokens, and internal service tokens accessible from any unpatched or exposed Ollama instance. Review audit logs for unusual API usage in the period before patching.
Search developer workstations, internal servers, CI/CD pipelines, Kubernetes clusters, and container registries for Ollama deployments. Include unofficial and experimental instances — developer laptops running Ollama exposed on a home or office network carry the same risk as production.
Review what secrets are loaded as environment variables in each Ollama deployment. Remove broad credentials — production API keys, cloud IAM credentials, database passwords — from Ollama's environment. Replace with scoped, short-lived credentials or workload identity where available.
Search Ollama and reverse proxy logs for unexpected blob upload or /api/create calls, especially from unfamiliar source IPs or at unusual times. Any suspicious upload activity before the patch was applied should be treated as a potential exploitation attempt.
Apply the same security posture to Ollama that you apply to databases, CI runners, and internal APIs: named ownership, defined exposure policies, authentication requirements, patch timelines, and logging standards. Local LLM runtimes are now first-class attack targets.
Treat all external GGUF and model files as untrusted input until validated. Establish an intake process that verifies model provenance and file integrity before loading into any inference runtime — particularly in shared or multi-tenant environments.
References
[1] Cyera Research
Bleeding Llama: Unauthenticated Process-Memory Disclosure in Ollama — May 5, 2026
[2] NVD
CVE-2026-7482 — Ollama GGUF Model Parsing Heap Out-of-Bounds Read