MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios
📰 ArXiv cs.AI
arXiv:2604.14158v1 Announce Type: cross Abstract: Current evaluations of long-term memory in LLMs are fundamentally static. By fixating on simple retrieval and short-context inference, they neglect the multifaceted nature of complex memory systems, such as dynamic state tracking and hierarchical reasoning in continuous interactions. To overcome these limitations, we propose MemGround, a rigorous long-term memory benchmark natively grounded in rich, gamified interactive scenarios. To systematical
DeepCamp AI