Shifting the Gradient: Understanding How Defensive Training Methods Protect Language Model Integrity

📰 ArXiv cs.AI

arXiv:2604.16423v1 Announce Type: cross Abstract: Defensive training methods such as positive preventative steering (PPS) and inoculation prompting (IP) offer surprising results through seemingly similar processes: both add trait-inducing objects to large language models (LLMs) during training, and both defend the LLM against acquiring the trait. The surprising success of these methods comes with the question: how do they work? Are PPS and IP doing the same thing? We provide behavioral and mecha

Published 21 Apr 2026
Read full paper → ← Back to Reads