Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs

📰 ArXiv cs.AI

arXiv:2505.11556v4 Announce Type: replace-cross Abstract: Multi-agent systems built on large language models (LLMs) are expected to enhance decision-making by pooling distributed information, yet systematically evaluating this capability has remained challenging. We introduce HiddenBench, a 65-task benchmark grounded in the Hidden Profile paradigm, which isolates collective reasoning under distributed information from individual reasoning ability. Evaluating 15 frontier LLMs, we find that multi-

Published 14 May 2026
Read full paper → ← Back to Reads