RoboPrompt.

Intuitive Robot Policy Steering with Sparse Human Input

† Corresponding authors

A trace, a point, or a rough direction.

Bread

Cup

Maze

Full video

An intuitive interface for everyday users to steer robot policies is essential for large-scale deployment.

RoboPrompt overview: points, traces, and directional guidance steer the robot; successful rollouts are filtered and mixed with offline data for continued policy improvement.

In laboratory settings, we typically use teleoperation devices to correct robot policies when they encounter failures. When robots are deployed in homes, however, the additional hardware costs and the skills required to use these devices become significant barriers.

RoboPrompt lets users guide the robot with a point, a trace, or a coarse direction. A lightweight module translates this sparse intent into an action draft; the robot's own policy refines it into fine-grained movement.

No specialized teleoperation hardware. No architectural changes or steerability fine-tuning of the base policy.

Intuitive prompts

A trace becomes
a correction.

A short trajectory drawn on the camera view guides the robot toward the toaster. The base policy refines the motion for insertion.

Decouple sparse intention translation and action refinement.

PHASE I

From prompt to action draft

A reusable 0.77B-parameter steering model translates points, traces, and coarse directions into a continuous 6D action draft.

PHASE II

From draft to fine movement

The base policy refines the draft through diffusion or flow matching. Progressive denoising gradually returns control to the policy.

Two-phase architecture: visual and directional prompts enter Evo-1 to produce a 3D action draft, then progressive policy refinement across action chunks produces guided actions.

Learning steerability

Collect the motion

The Phase I steering model is trained on 500 play-data episodes collected on the same robot platform. Each training window pairs the current observation with the robot's executed future motion.

Train Phase I model

The 0.77B Evo-1 model receives the prompted camera image, a separate prompt-only image, and the task instruction. Directional prompts are appended as text. Training combines action prediction with point, trace, and direction supervision; position and scale perturbations accommodate imprecise human input.

TOT filtering for better training data.

Not every successful rollout is suitable for online training. A robot may finish the task after long hesitation, too many retries, or repeated back-and-forth motions. Learning from these trajectories can introduce inconsistent actions.

We use Time-weighted Optimal Transport (TOT) to filter these rollouts. TOT compares their policy-feature sequences with expert demonstrations and penalizes excessive duration, retaining trajectories that are closer to expert behavior for the next training round.

Better policies. Fewer interventions.

Co-training with the original offline demonstrations and filtered steered DAgger rollouts leads to better policies and fewer human interventions.

+15.5 pp

Average task progress gain

80.0% to 95.5% on average for π0.5 across three tasks after DAgger: a 15.5 percentage-point increase.

44.0%

Fewer human interventions

2.86 to 1.60 on average for π0.5 across three tasks after DAgger.

Policy improvement with π0.5 across three real-world tasks. Round 0 is the original policy.

One plug-and-play module for various policies, without fine-tuning the base policy.

Policy improvement on the same Insert Bread task. Round 0 is each original base policy.

Cite this work

@misc{zou2026robopromptintuitiverobotpolicy,
      title={RoboPrompt: Intuitive Robot Policy Steering with Sparse Human Input},
      author={Yanwen Zou and Chenyang Shi and Guoxuan Xu and Wenye Yu and Wendi Chen and Ye Pan and Cewu Lu and Chuan Wen},
      year={2026},
      eprint={2610.10534},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2610.10534},
}