Skip to content
Research · Jul 27, 2026

Evaluation design alters measured gap between expert and automatic MeSH terms in systematic review classifiers

A controlled study finds that the apparent advantage of expert-assigned MeSH terms over automatic ones in systematic review classifiers depends heavily on evaluation choices, with the gap shrinking from +0.096 to about +0.02 WSS@95% under stricter designs.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • A new arXiv preprint shows that the measured performance gap between expert-assigned and automatic MeSH terms in systematic review classifiers varies widely depending on evaluation design choices.

A preprint on arXiv reports that the commonly observed advantage of expert-assigned Medical Subject Headings (MeSH) over automatic MeSH in systematic review classifiers is highly sensitive to evaluation design. Using the Cohen et al. (2006) drug-class benchmark across three topics (Statins, Opioids, ADHD), the author compares a bag-of-words logistic regression classifier (seven reruns) against BiomedBERT (five seeds) under multiple evaluation regimes. Under the canonical 5-fold full-corpus design, the bag-of-words model shows a +0.096 WSS@95% gap favoring expert MeSH over automatic MeSH on the Statins topic. However, when the corpus size is reduced to match smaller topics (n = 803), the gap falls to +0.033, with the 95% bootstrap confidence interval including zero. Using 10-fold cross-validation at full size further reduces the gap to +0.021, with the confidence interval narrowly excluding zero. Under the same canonical evaluation, BiomedBERT yields a +0.020 gap, which is within sampling noise of the bag-of-words 10-fold result. The author notes that a Statins-sized effect would not have been detectable at the variance levels observed for Opioids or ADHD, indicating those null results are design-limited rather than informative. The study also identifies a representation asymmetry: 15.1% of Statins inputs exceed BiomedBERT’s 512-token limit when expert MeSH terms are appended, which may contribute to the smaller transformer gap, although the effect cannot be separated from training volume differences in this analysis. Overall, the paper concludes that in screening pipelines using transformers or 10-fold bag-of-words, the expert-vs-auto MeSH gap on the tested topics is about 0.02 WSS@95%, with confidence intervals spanning zero on at least one bound, and that benchmark conclusions about feature sources can change substantially under reasonable changes to evaluation design.

Sources
  1. 01arXiv cs.CLEvaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark
Also on Research

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.