Hidden State Poisoning Attacks against Mamba-based Language Models

📰 ArXiv cs.AI

arXiv:2601.01972v4 Announce Type: cross Abstract: State space models (SSMs) like Mamba offer efficient alternatives to Transformer-based language models, with linear time complexity. Yet, their adversarial robustness remains critically unexplored. This paper studies the phenomenon whereby specific short input phrases induce a partial amnesia effect in such models, by irreversibly overwriting information in their hidden states, referred to as a Hidden State Poisoning Attack (HiSPA). Our benchmark

Published 16 May 2026
Read full paper → ← Back to Reads