[Paper] Rethinking the Divergence Regularization in LLM RL

Published: 3 days ago (June 8, 2026 at 01:58 PM EDT)

2 min read

Source: arXiv

Source: arXiv - 2606.09821v1

Overview

Reinforcement learning (RL) has become a key component of post-training large language models (LLMs). In practice, LLM RL is often off-policy because of training-inference mismatch and policy staleness, making trust-region control essential for stable optimization. Mainstream methods such as PPO and GRPO approximate this control with a ratio-clipping mechanism, but the importance ratio can be a poor proxy for distributional shift in long-tailed vocabularies. Recent work such as DPPO addresses this mismatch by replacing ratio-based clipping with a divergence-based mask, yielding a trust region defined by the sampled token’s absolute probability shift. However, DPPO still relies on a hard mask: once a token crosses the trust-region boundary in a harmful direction, its gradient is discarded rather than corrected. To address this, we propose Divergence Regularized Policy Optimization (DRPO), which replaces the hard mask with a smooth advantage-weighted quadratic regularizer on policy shift. DRPO preserves the same trust-region geometry as DPPO while inducing bounded, continuous gradient weights that attenuate diverging updates and provide corrective signals beyond the boundary. Experiments across model scales, architectures, and precision settings show that DRPO improves the stability and efficiency of LLM RL training.

Key Contributions

This paper presents research in the following areas:

cs.LG

Methodology

Please refer to the full paper for detailed methodology.

Practical Implications

This research contributes to the advancement of cs.LG.

Authors

Jiarui Yao
Xiangxin Zhou
Penghui Qi
Wee Sun Lee
Liefeng Bo
Tianyu Pang

Paper Information

arXiv ID: 2606.09821v1
Categories: cs.LG
Published: June 8, 2026
PDF: Download PDF

[Paper] Rethinking the Divergence Regularization in LLM RL

Overview

Key Contributions

Methodology

Practical Implications

Authors

Paper Information

Related posts

[Paper] Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models

[Paper] Context-Driven Incremental Compression for Multi-Turn Dialogue Generation

[Paper] FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning

[Paper] DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?