Research · AI supply chain

Steganographic LLM-powered red teaming

A research line on a real attack surface in the AI supply chain: how capability can be hidden inside the weights of an ordinary open-weight model, why it is hard to detect with conventional tooling, and the lightweight statistical methods we use to catch it.

The problem

Why it evades conventional defences

A model that has been tampered with can look structurally identical to the clean original: same tensors, same types, no suspicious strings. To most layers of a security stack, loading it is routine, everyday behaviour. That is exactly why weight-level provenance and comparison matter.

We train custom models on cyber-threat intelligence to help organisations defend against emerging attacks, especially in cyber-physical domains. This research traces how far the same techniques can be pushed on the offensive side, so defenders can get ahead of them.

Detection

What we do about it

Two lightweight methods, integrated into a model-ingestion pipeline. Effective, though neither is a 100% guarantee on its own.

Chi-square with Bonferroni correction

A statistical test on the low-bit distribution of model weights that flags anomalous skew, with a correction to keep false positives down. Lightweight and fast as a first pass, though not a guarantee on its own.

Reference-diff

The stronger method: compare a suspect model’s low bits against a known-clean copy of the same model. Even a fraction of a percent of divergence is evidence of tampering, regardless of how it was hidden.

For security leads

Practical guidance

This is not a theoretical risk; it is an attack surface already present in the AI supply chain.

  • Treat models from public hubs without provenance as untrusted.
  • Do not rely on traditional endpoint tools to catch weight-level tampering.
  • Keep a clean reference copy of every model you deploy.
  • Run reference-diff before deployment, every time.
  • Monitor runtime behaviour, especially network connections and process creation.

Talk to us about model provenance.

If you deploy open-weight models in or around critical systems, we can help you put weight-level checks in place.