Skip to content
Research · Aug 8, 2026

Apple demonstrates few-step text generation with a 1.7B-parameter flow model trained on 2.1 trillion tokens

A new paper introduces Categorical Flow Maps to accelerate text generation while maintaining near-data-level token entropy, scaling to a 1.7B-parameter model trained on 2.1T tokens.

Trust79
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • Apple’s ML Research proposes Categorical Flow Maps (CFMs) as an alternative to autoregressive language modeling, enabling accelerated sampling with as few as 4 inference steps.
  • A 1.7B-parameter base flow model was trained on 2.1 trillion tokens and self-distilled into a CFM that generates high-quality text while preserving token entropy.
  • The work introduces a likelihood bound for CFMs in the semi-discrete setting and shows competitive results on standard language modeling benchmarks compared to discrete diffusion methods.
  • Researchers report challenges at scale and provide prescriptive insights on loss weighting and time scheduling.

Apple’s Machine Learning Research group describes Categorical Flow Maps (CFMs) as a non-autoregressive alternative to traditional language modeling, arguing that continuous diffusion and flow matching methods can offer accelerated sampling and other benefits previously limited to continuous modalities like images.

The team trained a 1.7-billion-parameter base flow model on 2.1 trillion tokens and then self-distilled it into a CFM capable of generating diverse, high-quality text in as few as four inference steps while maintaining near-data-level token entropy.

To evaluate CFMs rigorously, the researchers introduce a likelihood bound for CFMs in the semi-discrete setting and demonstrate that CFMs can be used to score models on standard language modeling benchmarks, achieving results comparable to discrete diffusion baselines.

The paper also details practical challenges encountered during large-scale training—such as loss weighting and time scheduling—and provides prescriptive guidance for practitioners aiming to replicate or extend these results.

The authors note that prior work had only evaluated CFMs at modest scales (less than 1 billion parameters), leaving scalability an open question. This study addresses that gap by pushing into the multi-billion-parameter and multi-trillion-token regime.

Sources
  1. 01Apple — Machine Learning ResearchScaling Categorical Flow Maps
Also on Research

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.