Estimating Item Difficulty with Large Language Models as Experts

📰 ArXiv cs.AI

arXiv:2605.18562v1 Announce Type: cross Abstract: Accurate estimates of item difficulty are essential for valid assessment and effective adaptive learning. However, for newly created tasks, response data are typically unavailable. Pretesting and expert judgement can be costly and slow, while machine learning methods often require large labelled training datasets. Recent work suggests that large language models (LLMs) may help. However, there is limited evidence on the elicitation procedures and

Published 19 May 2026
Read full paper → ← Back to Reads