Non-linear Interventions on Large Language Models

📰 ArXiv cs.AI

arXiv:2605.14749v1 Announce Type: cross Abstract: Intervention is one of the most representative and widely used methods for understanding the internal representations of large language models (LLMs). However, existing intervention methods are confined to linear interventions grounded in the Linear Representation Hypothesis, leaving features encoded along non-linear manifolds beyond their reach. In this work, we introduce a general formulation of intervention that extends naturally to non-linear

Published 16 May 2026
Read full paper → ← Back to Reads