Large Language Models Could Be Rote Learners

📰 ArXiv cs.AI

arXiv:2504.08300v5 Announce Type: replace-cross Abstract: Benchmark-based evaluation, e.g., multiple-choice questions (MCQs) and open-ended questions (OEQs), is widely used for evaluating Large Language Models (LLMs), yet their reliability is undermined by benchmark contamination. When pre-exposed to the testing benchmark during training, less capable LLMs have been found to achieve inflated performance, thereby yielding erroneous results in LLM evaluation. In this study, we reframe contaminatio

Published 18 May 2026
Read full paper → ← Back to Reads