Skip to content
Daily PaperJul 14, 2026

← Archive

Is a dLLM's freedom really an advantage?

The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models

Is a dLLM's freedom really an advantage? — 1 / 9
01
Is a dLLM's freedom really an advantage? — 2 / 9
02
Is a dLLM's freedom really an advantage? — 3 / 9
03
Is a dLLM's freedom really an advantage? — 4 / 9
04
Is a dLLM's freedom really an advantage? — 5 / 9
05
Is a dLLM's freedom really an advantage? — 6 / 9
06
Is a dLLM's freedom really an advantage? — 7 / 9
07
Is a dLLM's freedom really an advantage? — 8 / 9
08
Is a dLLM's freedom really an advantage? — 9 / 9
09
01/09

Does the arbitrary generation order of diffusion LLMs really enable more diverse reasoning?

The flexibility trap

On math and coding reasoning, the paper shows that dLLMs tend to defer uncertain branching tokens and generate the easy tokens first — which can actually shrink the space of solutions they explore.

More freedom does not automatically translate into more diversity.

JustGRPO

The authors propose JustGRPO, which restricts generation to left-to-right order only during RL training.

Without any complex diffusion-specific reinforcement learning, it retains both strong reasoning performance and parallel decoding.

Tags

  • #ai-papers
  • #paper-review
  • #diffusion-llm
  • #reinforcement-learning
  • #grpo
  • #reasoning
View on Instagram