PAPER DEEP DIVE
RoboPrompt: Intuitive Robot Policy Steering with Sparse Human Input
End-to-end robot policies trained through imitation learning remain constrained by limited data diversity, making reliable zero-shot deployment in real-world settings challenging. Shared-autonomy methods enable human correction through teleoperation, but specialized hardware and operator training hinder deployment at scale. Other approaches incorporate human guidance as additional policy inputs, often requiring architectural changes and dedicated training for steerability, which limits their applicability across policies. We present RoboPrompt, a general-purpose, lightweight robot policy steering system that enables users to guide policy behavior through intuitive, sparse inputs, including drawn traces, target points, and coarse directional instructions. RoboPrompt decouples human-intention translation from the underlying policy: a reusable module converts human guidance into action drafts, which are refined through the diffusion or flow-matching dynamics of the base policy. By controlling action generation in noise space, RoboPrompt balances human intent with the policy prior without modifying the base policy architecture or fine-tuning it for steerability. Experiments demonstrate effective steering across Diffusion Policy, π_0.5, and FastWAM. We further use steered rollouts for online policy improvement through DAgger. After 2-3 rounds of iteration, average success rates increase by 15.5% for π_0.5 across three tasks and by 21.3% across three policies(Diffusion Policy, π_0.5, FastWAM) on the Insert Bread task, while average human intervention counts decrease by 44.0% (2.86 to 1.60) and 81.9% (2.60 to 0.47), respectively.
RoboPrompt: Intuitive Robot Policy Steering with Sparse Human Input
Yanwen Zou, Chenyang Shi, Guoxuan Xu, Wenye Yu, Wendi Chen, Ye Pan, Cewu Lu†, Chuan Wen† (Shanghai Jiao Tong University / Shanghai Innovation Institute / Noematrix)
arXiv:2610.10534v1 [cs.RO], submitted 2026-10-07 · arXiv:2610.10534 · Project page · Code: github.com/yanwen-zou/Roboprompt
IEEE index terms: Robot Policy Steerability, Robot Manipulation, Imitation Learning
One-sentence summary
Imitation-learned policies fail zero-shot, and fixing them in the field currently costs either teleoperation hardware with expert operators, or architecture changes plus steerability-specific fine-tuning. RoboPrompt takes a third path: it does not touch the base policy at all. A lightweight 0.77B VLA (Evo-1) translates a drawn point, a sketched trace, or a coarse directional sentence into an action draft, and the draft is then refined inside the base policy's own diffusion / flow-matching denoising process. The control knob is the denoising level $\gamma$: at $\gamma=0$ the output is exactly the human draft, at $\gamma=1$ it falls back to the unguided base policy. Across Diffusion Policy, $\pi_{0.5}$, and FastWAM, prompt alignment reaches 99.44% / 97.1% / 100.0%, and feeding filtered steered rollouts to DAgger raises $\pi_{0.5}$ success by +15.5% over three tasks in 2-3 rounds while cutting human interventions from 2.86 to 1.60 (-44.0%).
1. Problem: every existing correction route is expensive
Fig. 1 (paper Fig. 1): system vision. Offline data and online rollouts co-train the policy; humans steer at inference time with three sparse inputs (point, trace, direction sliders paired with e.g. "Move left a little bit."); successful steered rollouts are filtered and fed back as a self-improvement loop.
The capability ceiling of an end-to-end visuomotor policy is set by the diversity of its training data, which on real hardware shows up as unreliable zero-shot deployment: the policy fails on unseen object placements, lighting, or distractors. The paper sorts existing correction mechanisms into three families and prices each one.
Family one: shared autonomy / teleoperation correction. A human takes over low-level control directly. It works, but it requires a teleoperation device (Sigma.7 in this paper's experiments) and a trained operator, which turns "correcting the policy" into "hiring a person to drive the robot" and blocks deployment at scale.
Family two: human guidance as an extra policy conditioning input. Steerable Policies translate prompts into language conditioning; $\pi_{0.7}$ prompts a world model to turn subtask-level language into goal images. Two problems: language is a poor carrier of fine-grained spatial and motion information ("a bit to the left" is 2 cm or 5 cm?), and these methods require architectural changes or fine-tuning for steerability, so every new base policy means starting over.
Family three: training-free intervention at inference. Inference-Time Policy Steering biases the sampling process of a diffusion policy with visual guidance objectives, no training needed. But reweighting can only move outputs inside the policy's learned action manifold; when failure recovery needs an action outside that manifold, the route dead-ends.
RoboPrompt's positioning follows: decouple human-intention translation from policy execution. The translator is an independent small model, trained once and reused everywhere; the base policy stays frozen and only lends its denoising dynamics for refinement. No teleoperation hardware, no per-policy steerability retraining, and corrective actions outside the learned manifold remain reachable.
2. Phase I: translating sparse prompts into spatiotemporally complete action drafts
2.1 Three prompt types and their hindsight labels
Humans teach each other by sketching, pointing, or giving a coarse direction. RoboPrompt defines exactly these three forms as its prompt interface: point and trajectory prompts are overlaid visually on the environment-camera observation, while the global action prompt is appended to the task instruction as text. Formally, an episode is $\tau=\{(O_t, Q_t, a_t)\}_{t=0}^{T-1}$ with observation $O_t$, TCP pose $Q_t$, low-level action $a_t$, and action horizon $H$; a projection function
$$u_t = f_{\mathrm{proj}}(Q_t) \in \mathbb{R}^2 \tag{1}$$
maps the TCP onto the image plane of the fixed environment camera. The point prompt hindsight target is the projected TCP endpoint of the short-horizon motion:
$$p_t^{\mathrm{pt}} = u_{\bar{t}}, \qquad \bar{t} = \min(t+H-1,\ T-1) \tag{2}$$
and the model input is the observation with that point drawn on it, $\tilde{O}_t^{\mathrm{pt}} = \mathrm{Overlay}_{\mathrm{pt}}(O_t, p_t^{\mathrm{pt}})$. The trajectory prompt target is the projected future TCP sequence $P_t^{\mathrm{traj}} = (u_t, u_{t+1}, \ldots, u_{\bar{t}})$, connected into a line over the observation to form $\tilde{O}_t^{\mathrm{traj}}$. The global action prompt describes the coarse Cartesian displacement in the base frame; its hindsight target accumulates raw XYZ actions over the same horizon and normalizes:
$$g_t = \mathrm{Norm}_g\!\left(\sum_{k=0}^{H-1} a_{t+k}^{xyz}\right) \in [-1,1]^3 \tag{3}$$
then verbalizes it into text (e.g. "Cook bread, move x:0.128, y:0.465, z:-0.140 in global frame." in Fig. 2 below). The paper notes that VLMs can produce similar labels but prove unreliable under semantic ambiguity, visual occlusion, or fine-grained motion reasoning, hence the geometric hindsight pipeline over play data. Because human prompts are inherently imprecise, random position and scale perturbations are applied to prompts during training.
Fig. 2 (paper Fig. 3): prompt label generation. Connected TCP projections and the final TCP projection over the next n frames provide ground truth for trajectory and point prompts; summed and normalized actions supervise the global prompt. The policy actually receives two visual streams: the prompted camera image and a prompt-only image with no background.
2.2 The 0.77B drafting model Evo-1
Phase I uses a 0.77B-parameter VLA, Evo-1. For point/trajectory prompts the prompt is rendered on the environment-camera image as the primary visual input, and a second prompt-only image without camera background is fed alongside so the steering signal is not diluted by scene appearance. For global prompts the normalized displacement is verbalized into the task text. The model predicts a temporally continuous action chunk:
$$a_{\mathrm{I}} = \pi_\phi(O_t^{(1)}, O_t^{(2)}, q_t) \tag{4}$$
Training uses only 500 play-data trajectories collected on the same robot platform, with a main action loss plus three auxiliary supervisions:
$$\mathcal{L}_{\mathrm{I}} = \mathcal{L}_{\mathrm{act}}(a_{\mathrm{I}}, a_{t:t+H-1}) + \lambda_{\mathrm{aux}}(\mathcal{L}_{\mathrm{pt}} + \mathcal{L}_{\mathrm{traj}} + \mathcal{L}_{\mathrm{global}}) \tag{5}$$
where $\mathcal{L}_{\mathrm{pt}}$ and $\mathcal{L}_{\mathrm{traj}}$ supervise the target point and trajectory at pixel level in the visual input space, and $\mathcal{L}_{\mathrm{global}} = \operatorname{MSE}(\hat{g}_t, g_t)$. The design intent is explicit: Phase I does not aim for execution precision, only for a draft that is consistent with human intent, compatible with the current observation, and complete in space and time. Precision is Phase II's job.
3. Phase II: balancing human intent against the policy prior in noise space
Fig. 3 (paper Fig. 2): system architecture. In Phase I the prompt inputs pass through the lightweight VLA to produce a temporally continuous 6D action chunk draft. In Phase II the base policy refines the draft; the noise fraction $\gamma$ controls the balance between human guidance and the base policy, and $\gamma$ grows across replanning steps (0.3, 0.6, 0.9), handing control back progressively.
3.1 Progressive Biased Denoising (PBD)
Mainstream robot policies generate actions through diffusion or flow matching. RoboPrompt exploits the mathematical structure of iterative denoising: let $m$ be the replanning chunk index after a prompt is issued, $N$ the base policy's total denoising steps, and $k^m$ the current denoising step updated by a fixed increment $\Delta k$:
$$k^m = \min(k^0 + m\,\Delta k,\ N), \qquad \gamma^m = k^m / N \tag{6}$$
The basic formulation perturbs the Phase I draft at level $\gamma^m$ and denoises it with the base policy:
$$z_{\gamma^m} = (1-\gamma^m)\,a_{\mathrm{I}} + \gamma^m \epsilon, \quad \epsilon \sim \mathcal{N}(0, I), \qquad a_{\mathrm{II}} = \mathcal{D}_\theta(z_{\gamma^m},\ \gamma^m \rightarrow 0) \tag{7}$$
$\gamma$ is a tunable control ratio: at $\gamma=0$ the Phase II output is exactly the Phase I draft; at $\gamma=1$ it is the unguided base-policy action. A single human sketch usually corresponds to motion longer than one action chunk, and the monotone schedule of $\gamma^m$ keeps the same human input effective across multiple replanning chunks: the first chunk follows the human strongly, later chunks gradually reintroduce the base-policy prior, so control is handed over smoothly.
3.2 Noise-Augmented Flow Reversal Steering (NA-FRS)
The paper's measurements expose a key fact: the action change induced by denoising is not uniform over denoising time. A few denoising steps already move the action a lot, while the final steps at large noise levels are near convergence and barely move it. So $\gamma$ with naive random noise is not a linear knob, and cross-chunk handover looks abrupt. Building on Flow Reversal Steering, NA-FRS first integrates the learned denoising dynamics backward from the clean Phase I action to the target level:
$$z_{\gamma^m}^{\mathrm{rev}} = \mathcal{R}_\theta(a_{\mathrm{I}},\ 0 \rightarrow \gamma^m), \qquad \epsilon_{\mathrm{rev}} = \big(z_{\gamma^m}^{\mathrm{rev}} - (1-\gamma^m)a_{\mathrm{I}}\big)/\gamma^m \tag{8}$$
The reverse path preserves Phase I information but caps the original policy's performance, so the reverse noise is mixed with pure Gaussian noise, gradually corrupting the Phase I distribution as $\gamma^m$ grows:
$$\epsilon_{\mathrm{mix}} = \sqrt{1-\rho^2}\,\epsilon_{\mathrm{rev}} + \rho\,\epsilon, \qquad z_{\gamma^m}^{\mathrm{NA}} = (1-\gamma^m)a_{\mathrm{I}} + \gamma^m \epsilon_{\mathrm{mix}} \tag{9}$$
with $\rho \in [0,1]$ controlling the random-noise ratio; $z_{\gamma^m}^{\mathrm{NA}}$ is then denoised back to $t=0$. The reverse component anchors the output near the draft while the random component recovers pure-noise denoising behavior at large $\gamma$, yielding the smooth, near-linear transition across denoising levels that cross-chunk steering needs.
flowchart TD
H[Sparse human input
point / trace / directional text] --> P1[Phase I: Evo-1 0.77B VLA
prompted image + prompt-only image + task text]
PLAY[500 play-data trajectories
geometric hindsight labels + random perturbations] --> P1
P1 --> DRAFT[Action draft a_I
temporally continuous 6D chunk]
DRAFT --> GAMMA[Per-chunk denoising level
gamma^m = min(k0 + m*dk, N)/N]
GAMMA --> NA[NA-FRS
reverse integration gives eps_rev
mixed with Gaussian noise by rho]
NA --> DEN[Base policy denoising D_theta
pi0.5 / Diffusion Policy / FastWAM
frozen, no fine-tuning, no arch change]
DEN --> ACT[Refined action a_II executed]
ACT --> ROLL[Steered rollouts]
ROLL --> TOT[TOT filtering
OT distance to experts + duration penalty]
TOT --> MIX[Mixed 50/50 with offline data]
MIX --> DAG[Next DAgger round]
DAG --> DEN
Laid out this way the cost structure is clear: the only component that needs training is the 0.77B Phase I model (500 trajectories), and it is reused across all base policies; Phase II borrows the frozen base policy's existing denoising process. The DAgger loop then turns human corrections from a one-time expense into a compounding asset.
4. Steered rollouts are not free data: TOT filtering
Successful steered rollouts can still contain oscillations, repeated gripper toggles, or long suboptimal paths. Such data pairs similar observations with inconsistent or delayed actions and degrades Markovian imitation-learning policies during DAgger-style training. The paper introduces Time-weighted Optimal Transport (TOT) as an admission gate: for a steered rollout $\tau^r = \{x_i^r\}_{i=1}^{T_r}$ and an expert trajectory $\tau^e = \{x_j^e\}_{j=1}^{T_e}$, take policy embeddings $\phi(x)$ and define a timestep-level cost
$$C_{ij} = 1 - \frac{\phi(x_i^r)^\top \phi(x_j^e)}{\lVert \phi(x_i^r)\rVert_2\, \lVert \phi(x_j^e)\rVert_2} \tag{10}$$
then compute the OT distance to the closest expert trajectory while penalizing unnecessarily long executions. With $\bar{T}_E = \mathrm{median}_{\tau \in \mathcal{D}_E}|\tau|$:
$$D_{\mathrm{OT}}(\tau^r, \tau^e) = \min_{\Pi \in \mathcal{U}(T_r, T_e)} \sum_{i=1}^{T_r}\sum_{j=1}^{T_e} \Pi_{ij} C_{ij} \tag{11}$$
$$S(\tau^r) = \min_{\tau^e \in \mathcal{D}_E}\Big[D_{\mathrm{OT}}(\tau^r, \tau^e) + \lambda \max\big(0,\ T_r/\bar{T}_E - 1\big)\Big] \tag{12}$$
where $\mathcal{U}(T_r, T_e)$ is the set of transport plans with uniform temporal marginals. Low scores mean "semantically close to expert behavior and not excessively long". Only rollouts with $S(\tau^r)$ below a threshold are kept, and following prior work they are mixed with the original offline data at 50%/50% for the next DAgger round, pushing the policy toward the corrected deployment distribution while retaining the base-policy prior.
5. Experiments: three base policies, three targets, six first-time users
5.1 Platform and steering headline results
The hardware is a Flexiv Rizon4 arm with a Robotiq 2F-85 gripper; teleoperation for data collection uses a Sigma.7 device, with 100 demonstrations per task for the base policies. The steering task is "pick bread from the bowl and place it into one of three targets" (pot / left toaster / right toaster), with 46 demonstrations per target and the same text description across all three targets, so the base policy is deliberately multimodal and picks targets by its own preference when unsteered (pie charts in Fig. 4: $\pi_{0.5}$ 40%/25%/20%, DP 40%/25%/5%, FastWAM 60%/25%/15%; the missing mass is trials that failed before reaching any target).
Fig. 4 (paper Fig. 4): steering setup and natural policy distributions. Left: the robot picks bread from the bowl and places it into the pot, left toaster, or right toaster. Right: each base policy's unsteered preference over the three targets.
| Base policy | Target | Progress w/o steering | Progress w/ steering | Avg. steering counts | of which visual | of which global | Prompt alignment |
|---|---|---|---|---|---|---|---|
| $\pi_{0.5}$ | Left | 59.6% | 77.5% | 2.50 | 1.45 | 1.90 | 99.44% |
| Right | 72.4% | 2.95 | 2.50 | 2.50 | |||
| Pot | 89.8% | 2.35 | 1.80 | 1.25 | |||
| Diffusion Policy | Left | 46.5% | 71.3% | 4.80 | 3.10 | 3.40 | 97.1% |
| Right | 64.75% | 3.45 | 1.85 | 1.95 | |||
| Pot | 57.95% | 4.40 | 3.95 | 1.85 | |||
| FastWAM | Left | 46.2% | 57.4% | 3.15 | 2.15 | 2.30 | 100.0% |
| Right | 62.3% | 3.10 | 1.90 | 2.10 | |||
| Pot | 52.3% | 2.65 | 1.70 | 1.55 |
Table I of the paper reads as two blocks. The left block answers "does steering help": across all three policies and nine policy-target combinations, steered progress exceeds the unsteered baseline everywhere, with the largest single gains on $\pi_{0.5}$ Pot (59.6% to 89.8%, +30.2 points) and DP Left (46.5% to 71.3%, +24.8 points). The right block answers "is steering faithful and cheap": prompt alignment (the fraction of interventions whose resulting motion follows the direction specified by the prompt, counting only the dimensions the prompt specifies) stays above 95% for all three policies, reaching 100.0% on FastWAM; average steering counts range from 2.35 to 4.80, cheapest on $\pi_{0.5}$ (2.35-2.95) and most expensive on DP (3.45-4.80), consistent with DP having the weakest baseline (46.5%): the weaker the base policy, the more often a human must step in.
5.2 DAgger: turning corrections into training data
Fig. 5 (paper Fig. 5a): the three real-world DAgger tasks. Insert Bread: pick bread from a bowl and insert it into the toaster's right slot (unlike the steering experiment, training demos contain only the right slot); Hang Cup: hang a cup by its handle on the second-highest peg; Push Ball: push a polyhedral ball across an uneven platform until it touches the flag.
Fig. 6 (paper Fig. 5b): the Push Ball training set covers five maze layouts (20 demonstrations x 5 routes), while evaluation uses another unseen layout. Steered rollouts complete the task in the new layout and simultaneously supply data for successive DAgger rounds.
| Setting | Metric | Result |
|---|---|---|
| $\pi_{0.5}$, Insert Bread / Hang Cup / Push Ball, 2-3 DAgger rounds | Avg. success-rate gain | +15.5% |
| DP / $\pi_{0.5}$ / FastWAM, Insert Bread, same protocol | Avg. success-rate gain | +21.3% |
| $\pi_{0.5}$, three tasks | Avg. human interventions | 2.86 to 1.60 (-44.0%) |
| Three policies, Insert Bread | Avg. human interventions | 2.60 to 0.47 (-81.9%) |
| 6 first-time users, 5-minute intro, 5 bread-insertion trials each | Progress / interventions | Progress comparable to experienced users and well above the unsteered baseline; slightly more interventions |
| 20 successful rollouts, 1 DAgger round, TOT-filtered vs random | Progress on 3 tasks | Filtered higher on all three tasks, with fewer interventions |
The trend matters more than any single number: every DAgger round raises success while lowering the steering counts needed in the next round, i.e. a policy with moderate initial performance and frequent human intervention slides toward near-autonomous execution within a few rounds. On Insert Bread across three policies, interventions drop to 0.47 per episode, meaning most episodes need no human at all. In the user study (Q4), six volunteers with no prior exposure received only a 5-minute introduction and matched experienced users' task progress, which is the evidence behind the word "intuitive"; experienced users intervene less because familiarity lets them prevent failures rather than recover from them.
5.3 Two ablations: denoising dynamics and data quality
The denoising-dynamics experiment (Q3, paper Fig. 8) is NA-FRS's reason to exist: for both flow-matching and diffusion policies, action changes distribute unevenly over denoising steps, with early steps moving the action a lot and high-noise steps nearly converged. Plain FRS keeps outputs close to the Phase I action but caps base-policy performance; NA-FRS's reverse-plus-random mixture produces the smooth, approximately linear transition across denoising levels. The filtering ablation (Q5, paper Fig. 10) shows that "successful" does not mean "trainable": unfiltered successful rollouts inject motion jitter that causes failures or extra interventions, while TOT keeps the smooth, intent-consistent subset.
6. Code and reproducibility
The repository yanwen-zou/Roboprompt shipped with the paper and mirrors the method: Evo-1/ holds the Phase I model, prompt-conditioned dataset loader, and training implementation; openpi/, diffusion_policy/, and FastWAM/ are the three Phase II base policies with configs; steering/ implements the Phase I + Phase II composition and policy adapters; scripts/labeling/ contains the hindsight labeling that projects TCP trajectories into prompt labels; scripts/realworld/train/evo1/ provides two-stage launchers (freeze InternVL3 and train the action head for 5,000 steps, then resume and train VLM plus head for 30,000 steps); web_steer/ is a browser-based prompt interface; hardware/ covers robot/camera integration and calibration. Training and evaluation data live on HuggingFace as ywzou/Roboprompt_Play_Data (bread training set, maze observations, example checkpoints), and offline evaluation needs no robot: download the checkpoints, open the interactive visualization window, draw points or traces on the observation or move the direction sliders, and compare unprompted Phase II samples against the Phase I draft and the Phase I+II refinement. --phase2-steps accepts fractional step counts; smaller values preserve more of the Phase I proposal and larger values grant more refinement to the downstream policy, which is exactly the $\gamma$ knob exposed as an engineering parameter.
7. Limitations: two from the authors, four of my own
Authors' own. First, steering currently acts on end-effector pose only and does not support dexterous, multi-fingered end-effectors. Second, although the Phase I module is not tied to a specific embodiment or environment in principle, it is trained on a single platform's data today; zero-shot cross-embodiment steering would need larger and more diverse datasets.
My read. First, the tension between 500 play-data trajectories and the "general steering" claim is never quantified. Phase I is described as task-agnostic, but its training data comes from the same platform and a closely related task family; there is no train-on-A, steer-on-B experiment, so "reusable translator" is currently an architectural claim rather than an experimental result. Second, the prompt-alignment metric is lenient by construction: it counts only the dimensions the prompt specifies and measures whether motion follows the prompted direction, not whether the target is reached; the coexistence of 99.44%/100.0% alignment with only 57.4%-89.8% steered progress shows that execution precision sits between alignment and success, and alignment should not be read as success rate. Third, the DAgger headline numbers lack a per-round, per-task table: +15.5% and +21.3% come from curves in Fig. 6/Fig. 7, with no per-round per-task success rates, no standard deviations, and no seed counts in the text; the -81.9% intervention drop (2.60 to 0.47) likewise carries no variance information, and at real-robot sample sizes that magnitude can fluctuate a lot. Fourth, TOT's embedding $\phi(x)$, its admission threshold, and $\lambda$ are unspecified: Eq. (12)'s cost depends on cosine similarity of "policy embeddings", but the paper never says which model or layer, nor gives sensitivity analyses for the threshold and duration penalty, which are exactly the knobs that decide how much data is filtered out and which bias remains.
8. Conclusion and outlook
RoboPrompt's contribution compresses into one sentence: split "human intent to action" into "human intent to draft" and "draft to action", with a reusable small model for the first and the frozen base policy's denoising dynamics for the second. The value of this split is not a single performance number but an interface: steering capability decouples from the base policy, so $\pi_{0.5}$, Diffusion Policy, and FastWAM share one Phase I and one semantics of $\gamma$; and human corrections stop being consumables and become DAgger training assets, with interventions driven near zero in 2-3 rounds. PBD's monotone $\gamma$ schedule and NA-FRS's reverse-plus-random mixture address the same engineering problem, smooth cross-chunk handover of control, one supplying the semantics and the other the mathematical property.
Looking forward, the authors' roadmap is "larger, more diverse Phase I data for zero-shot cross-embodiment steering" plus dexterous end-effector support. From the evidence in this paper, two intermediate steps matter more to me: a cross-domain transfer experiment for Phase I (how much alignment drops on a new platform or camera layout), which decides whether "train once, reuse everywhere" holds; and shipping TOT's embedding and thresholds as reproducible defaults, since the repository currently provides labeling and training scripts while the filtering details remain in the paper's equations. With those two filled in, the "training-free steering plus steered-data self-improvement" route becomes a genuinely practical paradigm for operating real-robot policies in deployment.
Golden line
Humans do not need to know how to drive the robot, only how to point the way; translating the pointing into actions is the small model's job, and making those actions precise is the base policy's own denoising process's job. RoboPrompt's entire design is letting each of these two jobs live at the noise level where it is best.



