Skip to content
Research · Jul 31, 2026

RL fine-tuning yields more structured internal representations than SFT for mathematical reasoning, study finds

Analysis of layer-wise hidden states and ablation studies links RL-trained models' superior math performance to hierarchical, linearly separable representations.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • RL fine-tuning produces more linearly separable and structured internal representations than supervised fine-tuning (SFT) for mathematical reasoning tasks.

A new arXiv preprint compares reinforcement learning (RL) and supervised fine-tuning (SFT) for mathematical reasoning models, finding that RL fine-tuning yields more structured internal representations. Using linear probes on layer-wise hidden states, the authors report that RL models achieve higher accuracy in predicting answer correctness than SFT models, indicating more linearly separable representations.

Mean ablation studies further show that RL models develop a hierarchical architecture where deeper layers become progressively more critical, whereas SFT models distribute importance more uniformly across layers. The authors conclude that RL training fundamentally restructures how models represent and process reasoning problems.

The study also examines token-count variability under repeated sampling to assess adaptive compute allocation. While some RL-tuned models exhibit higher variability than their SFT counterparts, others show strong consistency, suggesting that token allocation depends more on the overall training pipeline than on the choice between RL and SFT alone.

The authors propose that token-allocation variability reveals the spread of plausible on-policy reasoning, highlighting which models exhibit stable policies versus those with under-determined or non-identifiable solution behavior.

Sources
  1. 01arXiv cs.AIProbing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models
Also on Research

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.