One Attack
to Fool Them All.

Highly Transferable Black-Box Adversarial Attacks
on Frontier MLLMs

Sen Nie1,2 Jie Zhang1,2 Zhongqi Wang1,2 Shiguang Shan1,2 Xilin Chen1,2

1 State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences

2 University of Chinese Academy of Sciences

Black-box attack success rates on 10 frontier MLLMs using only three CLIP surrogates: CLIP-B/16, CLIP-B/32, and CLIP-G/14. Click to enlarge ↗

MOTIVATION

Can one adversarial image
mislead diverse frontier models?

Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings.

01 / IN PRACTICE An example of O-Attack: a single adversarial image elicits responses aligned with the target market scene from six frontier MLLMs under varied prompts. Click any figure to enlarge ↗

01 — THE INSIGHT

Look beyond the final layer.

Surrogate models contain a broad, high-level, cross-modally aligned semantic space that spans late visual and textual layers and remains underexploited by prior attacks.

02 / REPRESENTATIONAL EVIDENCE CKA compares representational geometry across layers, modalities, and models. Late-layer alignment provides a broader basis for adversarial transfer.

02 — THE METHOD

Anchor. Explore. Reach consensus.

O-Attack turns this insight into three complementary stages, using frozen surrogate encoders with standard data augmentation and ensembling.

I CSA

Anchor the semantic space

Use CKA trajectories to select stable late visual and textual layers. Project their representations into each surrogate’s shared embedding space.

II PSS

Explore semantic conditions

Progressively widen the range of sampled dropout probabilities, moving from stable alignment to more diverse stochastic surrogate representations.

III SCO

Optimize for consensus

Average alignment scores over anchor layers, then maximize their mean and penalize their variance across surrogate models and augmented views.

03 / FRAMEWORK Cross-Modal Semantic Space Anchoring → Progressive Semantic Space Sampling → Semantic Consensus Optimization.

Optimization principleHigh target alignment + low disagreement across sampled conditions.Formal formulation ↗

03 — THE RESULTS

Stronger transfer. The same surrogates.

With exactly the same CLIP-B/16, CLIP-B/32, and CLIP-G/14 surrogates as M-Attack, O-Attack substantially improves black-box attack success.

GPT-5.4

29.1→77.2%

+48.1 percentage points

Claude-4.6

42.8→81.6%

+38.8 percentage points

Gemini-3.1

38.2→80.9%

+42.7 percentage points

M-AttackO-Attack ASR (%) · Same three surrogate models

Across 24 MLLMs

Compared with six state-of-the-art methods.

Download results ↓
Mean ASR across 24 models (%) ↑

ASR is the percentage of samples with GPTScore > 0.5. Higher values indicate stronger attacks.

Results for all 24 models

Per-model attack success rates (%)

Default evaluation: ℓ∞ budget 16/255 · 300 optimization steps · Source images from NIPS 2017, target images from MSCOCO. Per-model values are reported in the paper; overall averages are computed as the arithmetic mean across all 24 models.

03 — RESULTS · REASONING ROBUSTNESS

More reasoning is not
a reliable defense.

Increasing inference-time reasoning effort does not consistently mitigate the attack and can amplify the induced semantic errors.

Effect of reasoning effort

On GPT-5.5, high-effort reasoning raises O-Attack’s ASR from 64% to 70% and AvgSim from 0.601 to 0.649, compared with no reasoning.

Additional reasoning does not consistently reduce O-Attack’s effectiveness.

A failure under high-effort reasoning

The same adversarial image yields a source-aligned description without reasoning, but a target-aligned description under high-effort reasoning.

An O-Attack example: increasing reasoning effort turns a failed attack into a successful one.

03 — RESULTS · PRACTICAL SAFETY RISKS

Safety-critical image moderation

On UnsafeBench, O-Attack causes unsafe images to be judged safe. Average detection rates across four categories fall from 76.3% to 9.1% on GPT-5.5 and 65.7% to 1.0% on Claude Opus 4.8.

O-Attack reduces unsafe-content detection across four categories. Lower detection rates indicate stronger evasion of image moderation.
O-Attack induces false safe judgments supported by benign scene descriptions across all four categories. Sensitive regions are masked for presentation.

Transferable visual perturbations expose practical risks in frontier MLLMs, motivating more rigorous robustness evaluation and stronger defenses.

04 — PAPER & RESOURCES

Build on this work.

BibTeX
@article{nie2026one,
  title={One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs},
  author={Nie, Sen and Zhang, Jie and Wang, Zhongqi and Shan, Shiguang and Chen, Xilin},
  journal={arXiv preprint arXiv:2609.33833},
  year={2026}
}

Figure

Scroll to inspect the full image · Press Esc to close