Staleness-Adaptive Trust Region cuts asynchronous RL performance loss to 3% at 8× policy lag
Researchers propose SAT, a trust-region method that adapts PPO clipping to policy staleness, achieving 34.79 AIME24 avg@8 on Qwen3-30B when rollouts lag by 8 steps.

Asynchronous reinforcement learning trades stability for throughput by decoupling rollout generation from policy updates, but the resulting staleness—compounded by policy lag, engine delays, and routing mismatches—leaves high-staleness rollouts weakly controlled under standard PPO clipping. A new preprint introduces Staleness-Adaptive Trust Region (SAT), which adapts the PPO clip interval per token based on a staleness proxy derived from the detached sampled log-ratio. Rather than applying uniform clipping, SAT identifies high-mismatch tails within each batch via kernel scaling and contracts only the sign-selected endpoint of the nominal PPO interval, preserving baseline behavior on ordinary tokens while enforcing conservative updates on newly intercepted outward bands.
Evaluated on Qwen3-30B-A3B-Base in a decoupled asynchronous setup with SGLang inference and Megatron training, SAT-GSPO with R3 routing replay reached 35.83 AIME24 avg@8 at lag 1 and 34.79 at lag 8—retaining 96.8% of lag-1 performance. The authors prove local interval containment and pointwise pessimism relative to PPO, showing how the adaptive rule reshapes update geometry under heterogeneous staleness. The preprint was posted to arXiv on July 22, 2026.

