Daily PaperJul 14, 2026
Is a dLLM's freedom really an advantage?
The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models
01/09
Does the arbitrary generation order of diffusion LLMs really enable more diverse reasoning?
The flexibility trap
On math and coding reasoning, the paper shows that dLLMs tend to defer uncertain branching tokens and generate the easy tokens first — which can actually shrink the space of solutions they explore.
More freedom does not automatically translate into more diversity.
JustGRPO
The authors propose JustGRPO, which restricts generation to left-to-right order only during RL training.
Without any complex diffusion-specific reinforcement learning, it retains both strong reasoning performance and parallel decoding.
Tags
- #ai-papers
- #paper-review
- #diffusion-llm
- #reinforcement-learning
- #grpo
- #reasoning








