Wringer: Fill, then Wring — a 4B Reasoning Model at 2.655 Bits per Weight with 93.7% of Its Benchmark Scores wcamon • about 2 hours ago
ChordStream: a key-relative, time-anchored chord token stream for learning cadence PsiPi • about 13 hours ago
One sandbox per rollout, or how labs run RL for agents in 2026 sergiopaniego • about 22 hours ago • 3
From Barge-In to Floor Control: Engineering Voice Interruptions Under Uncertainty ericmey • 1 day ago • 2
Coordinating a real-time voice model with a separate task-execution agent: a practical architecture NatalieY • 2 days ago • 1
Trained 210M text-to-image model from scratch on one GPU: what actually mattered ivanmikhnenkov • 3 days ago • 4
Introducing Terminal-Bench-LILT: Multilingual Agentic Benchmark Grounded in Language, Region, and Culture Lilt-org • 4 days ago