The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images Paper • 2608.06270 • Published Aug 6 • 8
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents Paper • 2608.11878 • Published Aug 12 • 12
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection Paper • 2608.11562 • Published Aug 12 • 10
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Paper • 2608.08160 • Published Aug 8 • 30
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization Paper • 2608.12314 • Published Aug 12 • 28
Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models Paper • 2608.10708 • Published Aug 11 • 16
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 265
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses Paper • 2608.12307 • Published Aug 12 • 116