JEPA-Anything: Learning Predictive Models across Different Worlds Paper • 2609.20800 • Published 5 days ago • 58
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control Paper • 2609.19554 • Published 5 days ago • 33
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning Paper • 2609.13318 • Published 12 days ago • 7
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 12 days ago • 171
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization Paper • 2608.25864 • Published 27 days ago • 9
Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans Paper • 2608.24909 • Published Jul 22 • 6
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Paper • 2607.25895 • Published Jul 28 • 95
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Paper • 2607.23782 • Published Jul 26 • 80
N_0-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation Paper • 2607.23783 • Published Jul 26 • 48
Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning Paper • 2606.07436 • Published Jun 5 • 27
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research Paper • 2606.07591 • Published May 28 • 106
LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer Paper • 2602.10556 • Published Feb 11 • 2
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation Paper • 2601.22153 • Published Jan 29 • 76
AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security Paper • 2601.18491 • Published Jan 26 • 127
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding Paper • 2512.19693 • Published Dec 22, 2025 • 68
EgoX: Egocentric Video Generation from a Single Exocentric Video Paper • 2512.08269 • Published Dec 9, 2025 • 124
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics Paper • 2512.13660 • Published Dec 15, 2025 • 37
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework Paper • 2512.03041 • Published Dec 2, 2025 • 65