Pick a moving object
Grasp the moving piggy bank and place it in the basin.
Swipe to compare methods →
Korea Advanced Institute of Science & Technology (KAIST) * Equal contribution
Static-trained policies can lose responsiveness to moving targets: target-response collapse.
WAMs trained on static demonstrations often fail on moving targets even though they have the skill. As execution advances, the policy follows the continuation of what it was already doing and stops responding to where the target is: target-response collapse.
Replanning from the latest frame pairs a late robot state with a relocated target, a combination the static demonstrations never contained.
As execution advances, the same policy responds less to target relocation.
Restore the robot after k executed actions, relocate the target across a 9×7 grid, and ask the same frozen FastWAM where it would grasp. The response shrinks as execution advances.
ImageWAM shows the same decline (mean gain 0.885 → 0.063). On dual-bottle manipulation, relocation success after 24 actions drops from 100% to 66.7% (FastWAM) and 72.9% (ImageWAM).
After 0 actions
DPP plans from the initial robot context, with the target at its predicted interaction position.
Instead of asking the policy to recover from an unfamiliar robot–target pair, DPP plans from a familiar one: the initial robot state, with the target placed where the interaction will happen.
DPP never updates the WAM’s weights. It changes the observation the WAM is asked about, then connects the plan to the robot’s actual state.
WAM predictive rollout
The WAM rolls the learned skill out beyond its native action chunk and finds the first predicted grasp or release.
Observed target motion
SAM 2 segments the target, FoundationPose tracks it, and the fitted motion is evaluated at that time.
Existing static skill
A counterfactual observation places the target there in a familiar robot context, so the frozen WAM generates a skill it already has.
Predict → Prepay → Plan → Bridge → Refine.
Motion prepayment shifts the initial actions using estimated target motion while canonical planning runs. An activation bridge connects the moving robot to the canonical trajectory; fresh observations then support current-state refinement. These are overlapping processes, not five serial waits.
τ is a skill-derived timing prior, not an optimized interaction time. The diagram and timeline are schematic, not to scale. After the interaction, continuation follows the task's subtask structure.
Keep the robot and scene. Move the target to its predicted interaction position.
High-resolution method illustration: recorded robot joints are restored in a re-rendered scene, and the target is edited using DPP RGB-D reprojection. This is not a new evaluation frame. Rendering details.
Same robot, same scene, target where it will be. This edited observation is the counterfactual query that retrieves the WAM’s existing skill.
The Oracle-free pipeline uses no known object CAD, simulator actor mask, or simulator object pose. The same edit is carried into the wrist views.
DPP improves dynamic manipulation with a frozen FastWAM.
Across six simulated task families, DPP with a frozen FastWAM outperforms every comparator, including methods trained on dynamic data.
Success (%). Trials per method: 300 (T1), 192 (T2), 200 (T3), 100 (T4, T5, T6). PUMA and DynamicWAM are omitted on T3 and T6, which are absent from DOMINO.
Compare success, contact, and route completion on grasping and pick-and-place.
Bimanual Grasp (300 trials per method) and Pick-and-Place (192 trials: 144 co-motion, 48 counter-motion), with both DPP backbones and every baseline.
PUMA and DynamicWAM use additional training on dynamic data; every other method uses frozen pretrained checkpoints.
DPP leads across all tested target speeds.
Task 1 success by nominal object speed, 60 scenes per bin. DPP leads in every bin, including the fastest.
Hover or focus a speed bin for every value.
Dynamic manipulation on a real robot, using a policy trained on static demonstrations.
An I2RT-based platform running FastWAM trained only on static demonstrations, with learned inference on a single RTX 4090. 30 trials per method and task.
Choose a task on the left. In Task B, DPP and Reactive (Hybrid) share a policy trained on anticipation-oriented static demonstrations.
Grasp the moving piggy bank and place it in the basin.
Swipe to compare methods →
Separate trials, shared playback. Shorter clips hold on their final frame.
Object speed: 5.2–7.3 cm/s · Original-speed playback (1×)
Position the goalpost where the moving ball will arrive. Static demonstrations place the goalpost at a ball-blocking location; DPP transfers that geometry to the ball’s predicted position.
Swipe to compare methods →
Both methods use a policy trained on anticipation-oriented static demonstrations. Separate trials, shared playback.
Object speed: 5.2–7.3 cm/s · Original-speed playback (1×)
DPP also handles acceleration, circular motion, and position jumps.
Acceleration, circular motion, and position jumps that model intervention by a human or another robot. 100 trials per method and motion family.
DPP uses FastWAM with the optional nonlinear extension and a declared motion-family prior: late canonical planning and spatial warping of the current-state plan.
Choose a task and compare methods side by side.
One episode per task with every method on a shared clock. Choose a task, then press Play together.
Episodes are selected for illustration; labels mark the outcome of the displayed episode only.
Selected matched-scene examples where all three baselines fail and DPP succeeds; they do not replace the aggregate results. Playback follows each recording, and DPP recording begins after about 1 s of startup target motion.
Keep the target moving until the intended interaction.
Our benchmark uses task-specific interaction events to hand control from prescribed motion to physics, so incidental contact alone does not end the target’s motion.
Grab Roller · episode 24 · ALOHA · clean · Level 1
Our evaluation scene 206 · DOMINO-style · Success
Same evaluation scene 206 · Task-specific handoff · Failure
DOMINO explicitly identifies stop-on-contact simplification as a limitation (Appendix F.2). We distinguish incidental contact from task-specific interaction events and additionally reject qualifying unintended contact disconnected from grasping by more than one second.
The middle and right videos use the same PUMA checkpoint, scene, and seed under different handoff rules. The training demonstration on the left is a separate scene. The middle video applies DOMINO’s contact-stop rule to our fixed evalset, rather than reproducing the full official DOMINO benchmark. This selected comparison is not an aggregate result or evidence of a training–evaluation mismatch. Initial target positions can vary slightly with startup inference latency.
Training source: official DOMINO archive. Evaluation: scene stable_xy_068_r2, seed 4300412, 3.747 mm/step, RT+OIDN head view. Middle video: frame=completed policy step, 10 fps. Right video: 128 recorded frames at 10 fps; playback uses video frames because the exact policy-step alignment has not been verified. Training contact/activation timestamps and evaluation planner wall-time timestamps are unavailable; no synchronized physical clock is implied.
Initial execution and canonical-state conditioning drive the largest gains.
Removing initial execution or canonical-state conditioning costs about a third of all successes; oracle perception shows the remaining headroom.
FastWAM on the same 300 Task 1 scenes. Arrow colors follow Figure 3’s planning stages.
Dynamic manipulation failures can reflect limited access to learned skills rather than missing capabilities. DPP separates the planning context from the execution state, using WAM predictions to construct future conditions that elicit existing skills: an implicit retrieval query for the model’s learned manipulation capabilities.
@misc{park2026dpp,
title = {Dynamic Manipulation with World-Action Models via Counterfactual Planning},
author = {Park, Sunwoo and Lee, Wonbin and Jin, Seonghyun and Kim, Youngmin and Park, Jangho and Ye, Jong Chul},
year = {2026},
eprint = {2609.33172},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.33172}
}