Evaluation design alters measured gap between expert and automatic MeSH terms in systematic review classifiers
A controlled study finds that the apparent advantage of expert-assigned MeSH terms over automatic ones in systematic review classifiers depends heavily on evaluation choices, with the gap shrinking from +0.096 to about +0.02 WSS@95% under stricter designs.
1 source · cross-referenced
- A new arXiv preprint shows that the measured performance gap between expert-assigned and automatic MeSH terms in systematic review classifiers varies widely depending on evaluation design choices.
A preprint on arXiv reports that the commonly observed advantage of expert-assigned Medical Subject Headings (MeSH) over automatic MeSH in systematic review classifiers is highly sensitive to evaluation design. Using the Cohen et al. (2006) drug-class benchmark across three topics (Statins, Opioids, ADHD), the author compares a bag-of-words logistic regression classifier (seven reruns) against BiomedBERT (five seeds) under multiple evaluation regimes. Under the canonical 5-fold full-corpus design, the bag-of-words model shows a +0.096 WSS@95% gap favoring expert MeSH over automatic MeSH on the Statins topic. However, when the corpus size is reduced to match smaller topics (n = 803), the gap falls to +0.033, with the 95% bootstrap confidence interval including zero. Using 10-fold cross-validation at full size further reduces the gap to +0.021, with the confidence interval narrowly excluding zero. Under the same canonical evaluation, BiomedBERT yields a +0.020 gap, which is within sampling noise of the bag-of-words 10-fold result. The author notes that a Statins-sized effect would not have been detectable at the variance levels observed for Opioids or ADHD, indicating those null results are design-limited rather than informative. The study also identifies a representation asymmetry: 15.1% of Statins inputs exceed BiomedBERT’s 512-token limit when expert MeSH terms are appended, which may contribute to the smaller transformer gap, although the effect cannot be separated from training volume differences in this analysis. Overall, the paper concludes that in screening pipelines using transformers or 10-fold bag-of-words, the expert-vs-auto MeSH gap on the tested topics is about 0.02 WSS@95%, with confidence intervals spanning zero on at least one bound, and that benchmark conclusions about feature sources can change substantially under reasonable changes to evaluation design.
- Jul 27, 2026 · arXiv cs.CL
Researchers propose Humanly, a platform to document and audit human-AI collaborative writing sessions
Trust79 - Jul 27, 2026 · Apple — Machine Learning Research
Apple proposes GH-ESD to systematically identify grounded error slices in instance-level vision tasks
Trust79 - Jul 27, 2026 · arXiv cs.CL
Study probes whether Qwen2.5-7B-Instruct infers Colombian identity from linguistic cues
Trust79