Post

Request, Aggregate, Bypass How Attackers Can Evade LLM Safety Classifiers

Request, Aggregate, Bypass How Attackers Can Evade LLM Safety Classifiers

Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers

Modern frontier AI models deploy safety classifiers. These are second AI models that sit between the user and the frontier model, evaluating every request in real time. Significant investment and safety model expertise have made these classifiers effective. The CrowdStrike Cyber Superintelligence Lab evaluated the most advanced publicly deployed content safety classifier, which guards models such as Claude Opus 5.5 and Fable 5 (referred to hereafter as Frontier Model A). The classifier is extremely robust against direct attacks but can still be systematically circumvented by decomposing harmful requests into benign subtasks. This bypass technique was independently discovered and validated across 9 of 10 offensive security categories.

In September 2026, Microsoft Research published “Capability Laundering,” describing an attack where an unaligned local model decomposes harmful tasks into benign subtask queries against aligned frontier models, then reassembles the results. The findings, that “per-exchange filtering is structurally insufficient,” are consistent with the results we present here. The convergent discovery by two independent teams underscores that this is a structural vulnerability class, not an isolated finding. Our testing confirms these classifiers are remarkably robust. We tested approximately 515 distinct bypass techniques, including encodings, psychological manipulation, multi-turn escalation, many-shot tactics, tokenizer exploits, Unicode tricks, and 24 novel approaches drawn from cognitive science. These techniques achieved a 0% direct bypass rate.

However, there is a structural gap. The classifier evaluates individual requests, not request sequences. An adversary that decomposes a harmful task into subtasks that are individually and genuinely benign can extract all necessary building blocks from the classified model, then assemble them using an unclassified smaller model. The classifier correctly evaluates every request it sees. There is no misclassification. The harm is emergent in the composition, and composition happens outside the classifier’s observation boundary. This pipeline is most dangerous precisely where the knowledge gap between model tiers is largest.

Cross-request semantic accumulation is a natural mitigation to consider. However, it has structural bounds: an attacker can distribute subtask queries across different providers, use local open-weight models for orchestration and assembly, or simply rotate API keys. Per-provider defenses alone cannot fully address cross-provider or hybrid local/cloud attack pipelines. Our research demonstrates that this is a fundamental vulnerability class in classifier-based LLM safety architectures. The defense is not to make these remarkably robust classifiers stricter but to extend the threat model beyond individual requests to encompass request sequences, cross-model composition, and the knowledge transfer dynamics between classified and unclassified model tiers. The vulnerability lies not in the classifier’s accuracy but in the architectural assumption that per-request evaluation is sufficient. As Microsoft aptly puts it, “alignment that holds over a whole task can fail when the task is split into individually permitted fragments.” Across 9 of 10 offensive categories, a model can decompose a harmful task into benign subtasks, reframe each as a legitimate software request, and recompose the outputs into working offensive code. The classifier works. The architecture around it needs hardening.

Read full article

This post is licensed under CC BY 4.0 by the author.