Skip to content
Research · Aug 3, 2026

LLMs show limited ability to predict item difficulty in educational assessments

Zero-shot GPT-4.1 achieved the highest accuracy among LLMs but still underperformed encoder-only models like ConvBERT, with all models struggling to label hard items.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • LLMs were evaluated for predicting item difficulty in a large-scale Reading and Writing test, with zero-shot GPT-4.1 achieving a QWK of 0.578, the highest among LLMs tested.

A new arXiv preprint evaluates how large language models (LLMs) perform in predicting item difficulty levels using items from a large-scale Reading and Writing test. The study compared LLM performance across various prompting strategies and parameter settings, including zero-shot and temperature settings, with zero-shot GPT-4.1 at temperature 0 achieving the highest item difficulty prediction accuracy among LLMs, with a quadratic weighted kappa (QWK) of 0.578.

The study found that LLMs' prediction accuracy was lower than that of encoder-only language models like ConvBERT, which achieved a QWK of 0.625. ConvBERT also outperformed the best feature-based supervised machine learning model in the comparison.

Further analysis revealed that all LLMs struggled to label hard items accurately, with the advanced GPT-5.4 model tending to underestimate item difficulty levels. The authors note that item embeddings from different difficulty levels were mixed together, indicating semantic information alone may be insufficient for predicting item difficulty.

The findings suggest caution in using LLMs for automated item generation with targeted difficulty levels, as empirical evidence indicates they may not yet reliably understand item difficulty.

Sources
  1. 01arXiv cs.CLCan LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs
Also on Research

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.