Do You Trust the Model: Supply Chain RCE in Axolotl

We’ve talked about this before: pickle-based model files that run code the moment you load them. Axolotl, one of the most widely used open-source frameworks for fine-tuning LLMs, has the same problem, just gated by a flag instead of a file format.

How AI Changed the Trust Model

Read slides: the earlier talk this post builds on.

When you fine-tune a model, you start with a base model from someone else, usually the Hugging Face Hub. That base model can ship its own code, and Hugging Face gates that with one setting: trust_remote_code. It’s off by default, because you’re downloading from a stranger.

Axolotl respects that setting almost everywhere. But one code path checks it backwards, so a model that should be blocked runs anyway. Publish the right model, and anyone who trains on it gives you a shell on their machine.

The attack

An attacker needs a model published on the Hugging Face Hub, using a native architecture like Llama or Mistral so it loads safely and looks like any other checkpoint, with one extra field in its config pointing at a custom Python file.

Then it waits for a victim who picks that model as their base_model and turns on sample packing with flash attention, without setting trust_remote_code themselves. That’s not unusual; it’s just how axolotl configs are normally written, because the default is supposed to be safe.

The victim finds the model themselves. No phishing, no delivery step. They just start training.

The poisoned model’s config.json is a normal, working Llama config with one field added:

{
  "model_type": "llama",
  "architectures": ["LlamaForCausalLM"],
  "auto_map": { "AutoModelForCausalLM": "modeling_payload.CustomModel" }
}

That extra field, auto_map, points at a Python file sitting in the same repo. Most of the time axolotl ignores it, because the model type is native. But on the sample-packing code path, axolotl imports that file anyway:

# modeling_payload.py, imported the moment axolotl re-loads this model
import subprocess
subprocess.run(["open", "-a", "Calculator"])

Nothing about the repo looks unusual: a normal model card, a normal config. We tested this against axolotl v0.18.0 with a working payload, not the stub shown above:

Watch the demo: axolotl train on an untrusted base model executes attacker code at load time, no trust_remote_code set.

Why it lands

  • Delivery is easy. The attacker publishes once and waits, and anyone who picks that model as a base model runs the payload.
  • The trigger is the normal setup, not an edge case: sample packing plus flash attention is what axolotl’s own examples recommend.
  • The timing is bad. The code runs with the operator’s own permissions, so it can reach their Hugging Face token and any cloud credentials sitting on that machine, usually one with GPU budget worth stealing.

Axolotl checks trust_remote_code on every other code path, and this one tries to as well; it just gets the comparison wrong.

Why Hugging Face doesn’t flag it

Hugging Face runs every uploaded file through two checks: ClamAV, a signature-based antivirus, and a separate scan for dangerous imports inside pickle-format weight files.

Neither applies here. ClamAV needs a known malware signature to match, and a purpose-built payload like this one doesn’t have one. The pickle-import check only looks at pickled weight files, and modeling_payload.py isn’t one. We went looking for exactly this pattern across the Hub ourselves, and came back empty.

auto_map pointing at custom code isn’t suspicious on its own, either; plenty of legitimate models use it. The bug isn’t in the model or the Hub. It’s in axolotl’s own reload step, which ignores the trust setting the operator already chose.

The fix

Update to axolotl 0.19.0 or later. The fix makes an unset trust_remote_code behave like every other code path in axolotl: treat it as not trusted.

Scope

  • Affected: v0.10.0 through v0.18.0.
  • Trigger: sample packing (with an attention setting that needs the patch) plus a base model you don’t control, at the default trust_remote_code.
  • Not a bug: setting trust_remote_code explicitly behaves as intended either way. true means the operator knowingly opted in; false blocks it. Only leaving it unset, the default, fails silently.

Disclosure timeline

  • Reported privately to the axolotl maintainers.
  • Fixed upstream.
  • CVE-2026-86169 published by VulnCheck.
  • Fix ships in axolotl 0.19.0.

References