Learning to Solve Hard Problems in RL for LLMs by Never Giving Up Paper • 2609.13443 • Published 24 days ago • 13
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 18 days ago • 222
Revisiting Complete Reasoning Traces for Post-Training Paper • 2609.07103 • Published 28 days ago • 23
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published Sep 3 • 104
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published Aug 31 • 316
NVIDIA Ising Collection NVIDIA Ising is a new Model Family to enable building useful Quantum Computers with AI. • 7 items • Updated Aug 11 • 27
VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks Paper • 2511.04662 • Published Nov 6, 2025 • 37
A Survey of Data Agents: Emerging Paradigm or Overstated Hype? Paper • 2510.23587 • Published Oct 27, 2025 • 67
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning Paper • 2510.19338 • Published Oct 22, 2025 • 116
Fine-Tuning Large Language Models on Quantum Optimization Problems for Circuit Generation Paper • 2504.11109 • Published Apr 15, 2025 • 3
ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory Paper • 2509.25140 • Published Sep 29, 2025 • 15
Cache-to-Cache: Direct Semantic Communication Between Large Language Models Paper • 2510.03215 • Published Oct 3, 2025 • 99
Less is More: Recursive Reasoning with Tiny Networks Paper • 2510.04871 • Published Oct 6, 2025 • 520
DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search Paper • 2509.25454 • Published Sep 29, 2025 • 148