Shifting the Gradient: Understanding How Defensive Training Methods Protect Language Model Integrity
📰 ArXiv cs.AI
arXiv:2604.16423v1 Announce Type: cross Abstract: Defensive training methods such as positive preventative steering (PPS) and inoculation prompting (IP) offer surprising results through seemingly similar processes: both add trait-inducing objects to large language models (LLMs) during training, and both defend the LLM against acquiring the trait. The surprising success of these methods comes with the question: how do they work? Are PPS and IP doing the same thing? We provide behavioral and mecha
DeepCamp AI