Critical · CVSS 9.1 May 5, 2026 · Cyera Research

Bleeding Llama in Ollama

Defending Local LLM Runtimes from Unauthenticated Process-Memory Disclosure

Cyera Research disclosed an unauthenticated process-memory disclosure vulnerability in Ollama, one of the most widely deployed local LLM runtimes. An attacker with network access to an exposed Ollama instance can upload a crafted GGUF model file to trigger a heap out-of-bounds read, leaking unintended process memory — including user prompts, system prompts, API keys, and environment variables. With roughly 170,000 GitHub stars, more than 100 million Docker Hub downloads, and an estimated 300,000 servers globally, the exposure surface is significant.

1 CVE
9.1 CVSS Score
~300K Servers Affected
Patched Status

"Local LLM runtimes can expose high-value secrets even when the model itself is not compromised."

— Cyera Research

Technical Breakdown

The vulnerability lies in Ollama's GGUF model file parser, written in unsafe Go code. Crafted model files trigger a heap out-of-bounds read during quantization, surfacing unintended process memory in the API response.

CVE-2026-7482 Critical · 9.1 Memory Disclosure

Ollama GGUF Model Parsing Heap Out-of-Bounds Read

Ollama's unsafe Go code creates a heap out-of-bounds read condition during GGUF model parsing and quantization. An unauthenticated attacker with access to the blob upload API can upload a crafted GGUF file, then trigger model creation via /api/create to cause the server to read and return unintended heap memory. Since Ollama runs without authentication by default and frequently binds to 0.0.0.0, network-exposed instances are reachable without any credentials.

Affected: Ollama < patched version  ·  Fixed: See Ollama security advisory
Vector: AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:H  ·  Impact: Unauthenticated heap memory disclosure — user prompts, system prompts, API keys, tokens, credentials, and environment variables.

Root Cause: Unsafe Go Parsing

The GGUF model parser is implemented with unsafe Go code that does not validate model file structure before heap allocation. Attacker-controlled fields in the GGUF header control memory read offsets, creating a classic out-of-bounds read primitive.

Default Exposure: No Auth, Any Interface

Ollama ships with authentication disabled by default and frequently binds to all interfaces (0.0.0.0) when installed in server or Docker configurations. This means any network-accessible Ollama instance is an unauthenticated attack target without additional configuration.

What Leaks

Heap memory at the time of model processing may contain anything the Ollama process has handled: user chat prompts, system prompts, API keys passed via environment variables, cloud credentials, internal service tokens, and configuration data from connected agents or automation systems.

Scale of Exposure

Ollama has approximately 170,000 GitHub stars and more than 100 million Docker Hub downloads. Cyera Research estimated the issue could affect 300,000 servers globally — many of which are developer workstations, internal AI services, and enterprise LLM deployments exposed on internal or cloud networks without authentication.

Attack Path

1

Identify Exposed Ollama Instance

Attacker locates a network-accessible Ollama API service — on a cloud host, internal network, or developer workstation — bound to a non-loopback interface with no authentication required.

Ollama instances are trivially discoverable via Shodan, internal network scanning, or misconfigured cloud security groups. Treat all Ollama deployments as network-accessible until proven otherwise.

2

Upload Crafted GGUF Model via Blob API

Attacker uploads a specially crafted GGUF model file through Ollama's blob upload endpoint. No authentication is required. The malicious file contains header fields that control heap memory read offsets during subsequent parsing.

GGUF files are treated as trusted model artifacts. Restrict who can upload models via the blob API — especially in shared or production environments.

3

Trigger Model Creation to Cause Memory Read

Attacker triggers model creation or conversion via /api/create, causing Ollama to parse the crafted GGUF file. The out-of-bounds read fires during quantization, surfacing heap memory outside the intended model buffer.

Disable /api/create and blob-upload workflows on instances that do not require model ingestion from untrusted sources.

4

Retrieve Leaked Process Memory

The API response contains leaked heap data — user prompts, system prompts, API keys, environment variables, cloud credentials, and internal service tokens that have passed through the Ollama process. Attacker exfiltrates this material for immediate or downstream use.

Even a single successful read may expose credentials sufficient to compromise connected SaaS platforms, cloud infrastructure, or internal services. Rotate all secrets reachable from the affected host.

Impact

Process memory disclosure from an LLM runtime is particularly high-value because inference processes handle a wide variety of sensitive material in the normal course of operation — prompts, credentials, configuration, and context from connected agents and automation pipelines.

Directly Leaked

  • User chat prompts
  • System prompts and instructions
  • API keys (OpenAI, Anthropic, etc.)
  • Environment variables

Secondary Exposure

  • Cloud credentials (AWS, GCP, Azure)
  • Internal service tokens
  • Repository access tokens
  • Agent configuration and context

Downstream Risk

  • SaaS platform compromise
  • Cloud infrastructure access
  • Source code repository access
  • Lateral movement via stolen credentials

No authentication required: The default Ollama configuration ships without authentication and frequently binds to all network interfaces. An attacker needs only network reachability — no credentials, no prior access, no social engineering. The ~300,000 globally estimated exposed servers represent a large unauthenticated attack surface.

Defensive Tutorial

0–24 Hours

Upgrade all Ollama deployments to the patched version immediately. Restrict network exposure — bind Ollama to loopback only or enforce firewall rules preventing external access. Disable /api/create and blob-upload endpoints on instances that do not need them. Rotate all secrets reachable from affected hosts — API keys, cloud credentials, repository tokens, and service tokens.

1–7 Days

Inventory all Ollama deployments across workstations, labs, services, and containers. Isolate instances with access to sensitive credentials or high-value systems. Audit environment variables loaded into Ollama's process — remove broad credentials and replace with scoped, short-lived alternatives. Implement logging for model uploads and creation activity to detect future exploitation attempts.

Long-Term Governance

Treat local LLM runtimes the same way you treat databases, CI runners, and API gateways — as infrastructure with defined ownership, exposure policies, authentication requirements, logging standards, and patch timelines. Adopt model-artifact intake policies that treat all external GGUF files as untrusted input until validated. Do not allow Ollama processes to inherit broad developer credentials or production secrets. Use workload identity and scoped credentials.

Response Checklist

STEP 01 Upgrade Ollama to the Patched Version Immediate

Apply the patch across all environments — production, staging, development, and developer workstations. A CVSS 9.1 unauthenticated disclosure requires same-day treatment regardless of environment classification.

STEP 02 Restrict Ollama to Loopback Only Immediate

Ensure all Ollama instances bind exclusively to 127.0.0.1, not 0.0.0.0. Set OLLAMA_HOST=127.0.0.1 in the environment. Verify with netstat -tlnp — binding address is not always visible in config files.

STEP 03 Disable /api/create and Blob Upload Endpoints Where Not Required Immediate

Instances that serve inference only (no model ingestion from external sources) should block the /api/create and blob-upload workflows at the reverse proxy or firewall level. This removes the primary exploitation vector without affecting inference availability.

STEP 04 Rotate All Secrets Reachable from Affected Hosts Immediate

Rotate API keys (OpenAI, Anthropic, Hugging Face, etc.), cloud credentials (AWS IAM, GCP service accounts, Azure managed identities), repository tokens, and internal service tokens accessible from any unpatched or exposed Ollama instance. Review audit logs for unusual API usage in the period before patching.

STEP 05 Inventory All Ollama Deployments Urgent

Search developer workstations, internal servers, CI/CD pipelines, Kubernetes clusters, and container registries for Ollama deployments. Include unofficial and experimental instances — developer laptops running Ollama exposed on a home or office network carry the same risk as production.

STEP 06 Audit Environment Variables in Ollama Processes Urgent

Review what secrets are loaded as environment variables in each Ollama deployment. Remove broad credentials — production API keys, cloud IAM credentials, database passwords — from Ollama's environment. Replace with scoped, short-lived credentials or workload identity where available.

STEP 07 Review Recent Model Upload and Creation Activity Urgent

Search Ollama and reverse proxy logs for unexpected blob upload or /api/create calls, especially from unfamiliar source IPs or at unusual times. Any suspicious upload activity before the patch was applied should be treated as a potential exploitation attempt.

STEP 08 Treat LLM Runtimes as Infrastructure Important

Apply the same security posture to Ollama that you apply to databases, CI runners, and internal APIs: named ownership, defined exposure policies, authentication requirements, patch timelines, and logging standards. Local LLM runtimes are now first-class attack targets.

STEP 09 Adopt a Model Artifact Intake Policy Important

Treat all external GGUF and model files as untrusted input until validated. Establish an intake process that verifies model provenance and file integrity before loading into any inference runtime — particularly in shared or multi-tenant environments.

References

[1] Cyera Research

Bleeding Llama: Unauthenticated Process-Memory Disclosure in Ollama — May 5, 2026

[2] NVD

CVE-2026-7482 — Ollama GGUF Model Parsing Heap Out-of-Bounds Read

Primary source: Cyera Research, May 5, 2026. This advisory is an independent defensive guide produced by Spectreworks AI for educational purposes only and is not affiliated with Cyera Research or the Ollama project.