Anchor the semantic space
Use CKA trajectories to select stable late visual and textual layers. Project their representations into each surrogate’s shared embedding space.
Highly Transferable Black-Box Adversarial Attacks
on Frontier MLLMs
MOTIVATION
Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings.
01 — THE INSIGHT
Surrogate models contain a broad, high-level, cross-modally aligned semantic space that spans late visual and textual layers and remains underexploited by prior attacks.
02 — THE METHOD
O-Attack turns this insight into three complementary stages, using frozen surrogate encoders with standard data augmentation and ensembling.
Use CKA trajectories to select stable late visual and textual layers. Project their representations into each surrogate’s shared embedding space.
Progressively widen the range of sampled dropout probabilities, moving from stable alignment to more diverse stochastic surrogate representations.
Average alignment scores over anchor layers, then maximize their mean and penalize their variance across surrogate models and augmented views.
Optimization principleHigh target alignment + low disagreement across sampled conditions.Formal formulation ↗
03 — THE RESULTS
With exactly the same CLIP-B/16, CLIP-B/32, and CLIP-G/14 surrogates as M-Attack, O-Attack substantially improves black-box attack success.
29.1→77.2%
+48.1 percentage points
42.8→81.6%
+38.8 percentage points
38.2→80.9%
+42.7 percentage points
M-AttackO-Attack ASR (%) · Same three surrogate models
Compared with six state-of-the-art methods.
ASR is the percentage of samples with GPTScore > 0.5. Higher values indicate stronger attacks.
Default evaluation: ℓ∞ budget 16/255 · 300 optimization steps · Source images from NIPS 2017, target images from MSCOCO. Per-model values are reported in the paper; overall averages are computed as the arithmetic mean across all 24 models.
03 — RESULTS · REASONING ROBUSTNESS
Increasing inference-time reasoning effort does not consistently mitigate the attack and can amplify the induced semantic errors.
On GPT-5.5, high-effort reasoning raises O-Attack’s ASR from 64% to 70% and AvgSim from 0.601 to 0.649, compared with no reasoning.
The same adversarial image yields a source-aligned description without reasoning, but a target-aligned description under high-effort reasoning.
03 — RESULTS · PRACTICAL SAFETY RISKS
On UnsafeBench, O-Attack causes unsafe images to be judged safe. Average detection rates across four categories fall from 76.3% to 9.1% on GPT-5.5 and 65.7% to 1.0% on Claude Opus 4.8.
Transferable visual perturbations expose practical risks in frontier MLLMs, motivating more rigorous robustness evaluation and stronger defenses.
04 — PAPER & RESOURCES
@article{nie2026one,
title={One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs},
author={Nie, Sen and Zhang, Jie and Wang, Zhongqi and Shan, Shiguang and Chen, Xilin},
journal={arXiv preprint arXiv:2609.33833},
year={2026}
}