ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
📰 ArXiv cs.AI
arXiv:2510.18941v2 Announce Type: replace-cross Abstract: Evaluating progress in large language models (LLMs) is often constrained by the challenge of verifying responses, limiting assessments to tasks like mathematics, programming, and short-form question-answering. However, many real-world applications require evaluating LLMs in processing professional documents, synthesizing information, and generating comprehensive reports in response to user queries. We introduce ProfBench: a set of over 70
DeepCamp AI