Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Paper • 2608.00782 • Published Aug 1 • 17
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 171
view post Post 1777 Who wants a TRL sticker? 🙋https://github.com/huggingface/trl See translation 1 reply · 🤗 5 5 ❤️ 3 3 🚀 2 2 🔥 2 2 + Reply