CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
📰 ArXiv cs.AI
arXiv:2601.14952v2 Announce Type: replace-cross Abstract: While large language models now handle million-token contexts, their capacity for reasoning across entire document repositories remains largely untested. Existing benchmarks are inadequate, as they are mostly limited to single long texts or rely on a "sparse retrieval" assumption-that answers can be derived from a few relevant chunks. This assumption fails for true corpus-level analysis, where evidence is highly dispersed across hundreds
Full Article
Title: CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
Abstract:
arXiv:2601.14952v2 Announce Type: replace-cross Abstract: While large language models now handle million-token contexts, their capacity for reasoning across entire document repositories remains largely untested. Existing benchmarks are inadequate, as they are mostly limited to single long texts or rely on a "sparse retrieval" assumption-that answers can be derived from a few relevant chunks. This assumption fails for true corpus-level analysis, where evidence is highly dispersed across hundreds
Abstract:
arXiv:2601.14952v2 Announce Type: replace-cross Abstract: While large language models now handle million-token contexts, their capacity for reasoning across entire document repositories remains largely untested. Existing benchmarks are inadequate, as they are mostly limited to single long texts or rely on a "sparse retrieval" assumption-that answers can be derived from a few relevant chunks. This assumption fails for true corpus-level analysis, where evidence is highly dispersed across hundreds
DeepCamp AI