Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
📰 ArXiv cs.AI
Mechanistic interpretability reveals that LLMs have dissociable affect reception and emotion categorization mechanisms
Action Steps
- Use mechanistic interpretability to analyze LLMs' internal representations of emotion
- Examine the difference between affect reception and emotion categorization in LLMs
- Investigate how LLMs respond to emotional stimuli without explicit emotion keywords
Who Needs to Know This
AI researchers and engineers working with LLMs can benefit from understanding how these models process emotions, and product managers can use this knowledge to improve AI-powered products
Key Insight
💡 LLMs' emotion detection mechanisms can be dissociated from their reliance on explicit emotion keywords
Share This
💡 LLMs have separate mechanisms for detecting emotional meaning and categorizing emotions
Key Takeaways
Mechanistic interpretability reveals that LLMs have dissociable affect reception and emotion categorization mechanisms
Full Article
Title: Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs
Abstract:
arXiv:2603.22295v1 Announce Type: cross Abstract: Large language models appear to develop internal representations of emotion -- "emotion circuits," "emotion neurons," and structured emotional manifolds have been reported across multiple model families. But every study making these claims uses stimuli signalled by explicit emotion keywords, leaving a fundamental question unanswered: do these circuits detect genuine emotional meaning, or do they detect the word "devastated"? We present the first
Abstract:
arXiv:2603.22295v1 Announce Type: cross Abstract: Large language models appear to develop internal representations of emotion -- "emotion circuits," "emotion neurons," and structured emotional manifolds have been reported across multiple model families. But every study making these claims uses stimuli signalled by explicit emotion keywords, leaving a fundamental question unanswered: do these circuits detect genuine emotional meaning, or do they detect the word "devastated"? We present the first
DeepCamp AI