Sessa: Selective State Space Attention
📰 ArXiv cs.AI
arXiv:2604.18580v2 Announce Type: cross Abstract: Modern sequence modeling is dominated by two families: Transformers, whose self-attention can access arbitrary elements of the visible sequence, and structured state-space models, which propagate information through an explicit recurrent state. These mechanisms face different limitations on long contexts: when attention is diffuse, the influence of individual tokens is diluted across the effective support, while recurrent state propagation can lo
DeepCamp AI