Distilled RL transfers knowledge across model families without unconditional imitation
A new preprint proposes Distilled Reinforcement Learning, which combines fine-grained teacher guidance with RL objectives to transfer knowledge across model families more effectively than standard RL or on-policy distillation.







