LLMs show limited ability to predict item difficulty in educational assessments
Zero-shot GPT-4.1 achieved the highest accuracy among LLMs but still underperformed encoder-only models like ConvBERT, with all models struggling to label hard items.
1 source · cross-referenced
- LLMs were evaluated for predicting item difficulty in a large-scale Reading and Writing test, with zero-shot GPT-4.1 achieving a QWK of 0.578, the highest among LLMs tested.
A new arXiv preprint evaluates how large language models (LLMs) perform in predicting item difficulty levels using items from a large-scale Reading and Writing test. The study compared LLM performance across various prompting strategies and parameter settings, including zero-shot and temperature settings, with zero-shot GPT-4.1 at temperature 0 achieving the highest item difficulty prediction accuracy among LLMs, with a quadratic weighted kappa (QWK) of 0.578.
The study found that LLMs' prediction accuracy was lower than that of encoder-only language models like ConvBERT, which achieved a QWK of 0.625. ConvBERT also outperformed the best feature-based supervised machine learning model in the comparison.
Further analysis revealed that all LLMs struggled to label hard items accurately, with the advanced GPT-5.4 model tending to underestimate item difficulty levels. The authors note that item embeddings from different difficulty levels were mixed together, indicating semantic information alone may be insufficient for predicting item difficulty.
The findings suggest caution in using LLMs for automated item generation with targeted difficulty levels, as empirical evidence indicates they may not yet reliably understand item difficulty.
- Aug 3, 2026 · arXiv cs.AI
Researchers propose LLM pipeline to generate and formally validate mathematical conjectures
Trust76 - Aug 3, 2026 · arXiv cs.AI
Benchmark finds FARS-generated papers outperform other AI Scientist frameworks by more than twofold
Trust79 - Aug 1, 2026 · Apple — Machine Learning Research
Apple researchers propose graph-based sensemaking workflows using UMAP’s internal kNN graph
Trust79