ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF Image-Text-to-Text • 177B • Updated 12 days ago • 4.19M • 785
Base Models Can Reason By Taking a Cue From Training Data Paper • 2610.06851 • Published 6 days ago • 25
ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation Paper • 2609.09076 • Published Sep 8 • 24
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published Sep 8 • 83
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published Aug 27 • 154
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization Paper • 2608.20281 • Published Aug 20 • 15
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Paper • 2608.16590 • Published Aug 17 • 153
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published Aug 13 • 48
SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models Paper • 2608.10538 • Published Aug 11 • 17