Abstract
Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and its growing history. We introduce RobotUse, a robot agent harness that organizes computation, context, and decisions around specifying and revising physical actions. Agents visually select targets and poses, while the backend handles geometry, motion planning, and control. Subagents retain detailed interactions within each subgoal and return the information needed for subsequent decisions. Continual harnessing lets agents learn from execution by updating a persistent playbook. On RoboLab, RobotUse achieves 45% task success, outperforming CaP-X by 6.7 percentage points while maintaining compact decision contexts and reducing reliance on predefined action abstractions. Furthermore, we show that RobotUse learns from real-world execution despite imperfect feedback and transfers what it learns to subsequent tasks. Project page is available at https://robotuse-team.github.io/.
Community
RobotUse is a robot agent harness that connects language-level goals to visual decisions and physical execution. Robot tasks require agents to plan toward an overall goal while selecting targets, choosing gripper poses, and revising actions based on observed outcomes. RobotUse organizes computation, context, and decisions across a Main agent, a Subagent, and a Backend.
The Main agent plans in language, delegating subgoals with relevant constraints. The Subagent makes visual action decisions: it selects targets and locations in images, inspects candidate gripper poses, and adjusts them using execution feedback. The Backend performs geometric computation, motion planning, and control, returning fresh observations and diagnostics. The Subagent then reports the outcome, its causes, and the current state to the Main agent, informing the next decision.
Detailed observations and local recovery attempts remain within each Subagent’s context, while the Main agent receives the information needed to continue planning. RobotUse also refines a persistent playbook for continual harness , allowing accumulated guidance to improve subsequent decisions without changing model parameters or robot control code.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation (2026)
- Recursive Harness Distillation across Agents for Robot Manipulation (2026)
- World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal (2026)
- Agent as Policy for Robotic Manipulation (2026)
- In-Context Robot Learning with VLM Agents (2026)
- Know Your Body: A Harness for Direct and Self-Improving Robot Control with VLMs (2026)
- MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.04929 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper