Predict, don’t chase

Dynamic Manipulation with World-Action Models
via Counterfactual Planning

  • Sunwoo Park*
  • Wonbin Lee*
  • Seonghyun Jin*
  • Youngmin Kim*
  • Jangho Park
  • Jong Chul Ye

Korea Advanced Institute of Science & Technology (KAIST) * Equal contribution

arXiv Code soon BibTeX
Compare reactive planning and DPP

Reactive Planning

0% success · 30 trials

Dynamic Predictive Planning Ours

76.7% success · 30 trials
Real robot, Task A: pick a moving piggy bank and place it in the basin · separate trials, real-time playback, object speed: 5.2–7.3 cm/s
  • 83.3% T1 Bimanual Grasp vs. DynamicWAM 52.0%
  • 80.7% T2 Pick-and-Place vs. 53.1%
  • 76.7% Real robot, Task A vs. reactive 0%
  • 0 dynamic training · one consumer GPU

Why chasing fails

Static-trained policies can lose responsiveness to moving targets: target-response collapse.

Why latest-frame replanning can fail

WAMs trained on static demonstrations often fail on moving targets even though they have the skill. As execution advances, the policy follows the continuation of what it was already doing and stops responding to where the target is: target-response collapse.

Replanning from the latest frame pairs a late robot state with a relocated target, a combination the static demonstrations never contained.

Static demonstrations: trajectories from a shared start state fan out to many targets S₀ later states: one target each start: many targets demonstrated targets
What static demonstrations coverNear the start state S₀, one robot configuration is paired with many targets. Later states become specific to the trajectory toward one target.
A moving target: the robot is late on the way to the old target position while the target has moved elsewhere S₀ Sₜ target was here target now never demonstrated (Sₜ, new target) keeps going
What a moving target asks forThe robot has reached a late state Sₜ on the way to one target while the target moved. That pairing was never demonstrated, so the policy keeps heading for the familiar subgoal.

Target-response collapse, measured

As execution advances, the same policy responds less to target relocation.

Additional measurements

Restore the robot after k executed actions, relocate the target across a 9×7 grid, and ask the same frozen FastWAM where it would grasp. The response shrinks as execution advances.

ImageWAM shows the same decline (mean gain 0.885 → 0.063). On dual-bottle manipulation, relocation success after 24 actions drops from 100% to 66.7% (FastWAM) and 72.9% (ImageWAM).

Prefix 0: predicted grip points spread widely and follow the target grid. Prefix 8: predicted grip points still spread across the grid. Prefix 16: the response hull shrinks. Prefix 32: predicted grip points cluster near the center. Prefix 48: predicted grip points collapse to a small region. After 0 actions
Executed actions
Target offset Predicted grip response Response hull

Predict, don’t chase

DPP plans from the initial robot context, with the target at its predicted interaction position.

Model and execution details

Instead of asking the policy to recover from an unfamiliar robot–target pair, DPP plans from a familiar one: the initial robot state, with the target placed where the interaction will happen.

DPP never updates the WAM’s weights. It changes the observation the WAM is asked about, then connects the plan to the robot’s actual state.

Figure 1. Predict, don’t chase. Click to view at full resolution.

When

WAM predictive rollout

The WAM rolls the learned skill out beyond its native action chunk and finds the first predicted grasp or release.

Where

Observed target motion

SAM 2 segments the target, FoundationPose tracks it, and the fitted motion is evaluated at that time.

How

Existing static skill

A counterfactual observation places the target there in a familiar robot context, so the frozen WAM generates a skill it already has.

How DPP works

Predict → Prepay → Plan → Bridge → Refine.

Timing and execution details

Motion prepayment shifts the initial actions using estimated target motion while canonical planning runs. An activation bridge connects the moving robot to the canonical trajectory; fresh observations then support current-state refinement. These are overlapping processes, not five serial waits.

τ is a skill-derived timing prior, not an optimized interaction time. The diagram and timeline are schematic, not to scale. After the interaction, continuation follows the task's subtask structure.

Moving target x̂(τ) Imagined rollout First interaction → τ Motion prepayment Canonical plan Activation bridge Current-state plan S₀ Initial context + x̂(τ) Fresh observation + x̂(τ) Interaction at τ S₀
Planning
Initial Canonical Current-state
Execution
Prepaid prefix Bridge Suffix Align Refined suffix τ

Counterfactual Observation

Original observation: the target at its observed position. Counterfactual observation: the same robot pose with the target moved to the predicted grip position.
OriginalCounterfactual

Keep the robot and scene. Move the target to its predicted interaction position.

How the observation is constructed

High-resolution method illustration: recorded robot joints are restored in a re-rendered scene, and the target is edited using DPP RGB-D reprojection. This is not a new evaluation frame. Rendering details.

Same robot, same scene, target where it will be. This edited observation is the counterfactual query that retrieves the WAM’s existing skill.

  1. aRemove the visible target while protecting the robot and manipulated objects.
  2. bInpaint RGB and restore depth from aligned background memory.
  3. cReproject the target’s RGB-D geometry at the predicted position, with a z-buffer resolving occlusion.

The Oracle-free pipeline uses no known object CAD, simulator actor mask, or simulator object pose. The same edit is carried into the wrist views.

Simulation results

DPP improves dynamic manipulation with a frozen FastWAM.

Details

Across six simulated task families, DPP with a frozen FastWAM outperforms every comparator, including methods trained on dynamic data.

Success (%). Trials per method: 300 (T1), 192 (T2), 200 (T3), 100 (T4, T5, T6). PUMA and DynamicWAM are omitted on T3 and T6, which are absent from DOMINO.

  • R: FastWAM Refresh-24
  • F: Oracle Future-cond.
  • DW: DynamicWAM
  • PU: PUMA
  • π: π0.5
  • DPP: FastWAM

Task 1 and Task 2 in detail

Compare success, contact, and route completion on grasping and pick-and-place.

Details

Bimanual Grasp (300 trials per method) and Pick-and-Place (192 trials: 144 co-motion, 48 counter-motion), with both DPP backbones and every baseline.

PUMA and DynamicWAM use additional training on dynamic data; every other method uses frozen pretrained checkpoints.

Success rate (%) ↑0–100 · Higher is better
Task 1Bimanual Grasp
Task 2Pick-and-Place
050100
050100
FastWAM + VP · DPPOurs · visual prediction
83.33
80.73
ImageWAM + VP · DPPOurs · visual prediction
78.67
91.15
DynamicWAMAdditional dyn. training
52.00
53.13
PUMAAdditional dyn. training
50.33
9.38
π0.5Pretrained policy
9.00
13.02
AHA-WAMPretrained policy
7.33
14.58
Refresh & Future-Conditioned baselines 8 methods
FastWAMRefresh-24 (Default)
10.00
33.33
FastWAMRefresh-8
13.00
38.54
FastWAMFuture-Conditioned (DPP bootstrap)
18.33
34.38
FastWAMFuture-Conditioned (Oracle)
27.00
48.96
ImageWAMRefresh-16 (Default)
11.33
39.58
ImageWAMRefresh-8
12.67
46.88
ImageWAMFuture-Conditioned (DPP bootstrap)
18.33
48.96
ImageWAMFuture-Conditioned (Oracle)
22.67
60.42

Success stays high as targets speed up

DPP leads across all tested target speeds.

Measurement details and full table

Task 1 success by nominal object speed, 60 scenes per bin. DPP leads in every bin, including the fastest.

Hover or focus a speed bin for every value.

  • DPP · FastWAM
  • DPP · ImageWAM
  • DynamicWAM
  • PUMA
  • FC-FastWAM (Oracle)
  • FastWAM Refresh-24

Real-world evaluation

Dynamic manipulation on a real robot, using a policy trained on static demonstrations.

Evaluation setup and paper figure

An I2RT-based platform running FastWAM trained only on static demonstrations, with learned inference on a single RTX 4090. 30 trials per method and task.

Choose a task on the left. In Task B, DPP and Reactive (Hybrid) share a policy trained on anticipation-oriented static demonstrations.

Task A · Pick (Move) & Place

Pick a moving object

Grasp the moving piggy bank and place it in the basin.

Reactive 0% success
Replans from the current observation
Future-conditioned 0% success
Fixed 16-step prediction horizon
DPP Ours 76.7% success
Plans for predicted interaction

Swipe to compare methods →

0:00 / 0:00
Task B · Pick & Place (Move)

Place to intercept a moving ball

Position the goalpost where the moving ball will arrive. Static demonstrations place the goalpost at a ball-blocking location; DPP transfers that geometry to the ball’s predicted position.

Hybrid 40.0% success
Reactive inference with a learned anticipation prior
DPP Ours 86.7% success
Predictive planning with the same learned prior

Swipe to compare methods →

0:00 / 0:00

Complex target motion

DPP also handles acceleration, circular motion, and position jumps.

Nonlinear extension and paper figure

Acceleration, circular motion, and position jumps that model intervention by a human or another robot. 100 trials per method and motion family.

DPP uses FastWAM with the optional nonlinear extension and a declared motion-family prior: late canonical planning and spatial warping of the current-state plan.

  • R: FastWAM Refresh-24
  • F: Oracle Future-cond.
  • DW: DynamicWAM
  • DPP: FastWAM
  • RC: route completion

Simulation videos

Choose a task and compare methods side by side.

Details

One episode per task with every method on a shared clock. Choose a task, then press Play together.

Episodes are selected for illustration; labels mark the outcome of the displayed episode only.

Beyond contact-stop simplification

Keep the target moving until the intended interaction.

Our benchmark uses task-specific interaction events to hand control from prescribed motion to physics, so incidental contact alone does not end the target’s motion.

DOMINO training demonstration

Grab Roller · episode 24 · ALOHA · clean · Level 1

Stored training observations, played at 10 fps. Frame timing is a preview, not measured physical time.

PUMA with contact-stop

Our evaluation scene 206 · DOMINO-style · Success

At step 43, contact releases prescribed motion while both grippers are open (L=1.0, R=1.0). The policy later completes the grasp.

PUMA under our protocol

Same evaluation scene 206 · Task-specific handoff · Failure

The same PUMA policy fails on the matching scene when incidental contact does not trigger motion termination. The original gripper-based handoff and 75-step motion cap remain in place.
Protocol & sources

DOMINO explicitly identifies stop-on-contact simplification as a limitation (Appendix F.2). We distinguish incidental contact from task-specific interaction events and additionally reject qualifying unintended contact disconnected from grasping by more than one second.

The middle and right videos use the same PUMA checkpoint, scene, and seed under different handoff rules. The training demonstration on the left is a separate scene. The middle video applies DOMINO’s contact-stop rule to our fixed evalset, rather than reproducing the full official DOMINO benchmark. This selected comparison is not an aggregate result or evidence of a training–evaluation mismatch. Initial target positions can vary slightly with startup inference latency.

Training source: official DOMINO archive. Evaluation: scene stable_xy_068_r2, seed 4300412, 3.747 mm/step, RT+OIDN head view. Middle video: frame=completed policy step, 10 fps. Right video: 128 recorded frames at 10 fps; playback uses video frames because the exact policy-step alignment has not been verified. Training contact/activation timestamps and evaluation planner wall-time timestamps are unavailable; no synchronized physical clock is implied.

Ablations

Initial execution and canonical-state conditioning drive the largest gains.

Experimental setup and full table

Removing initial execution or canonical-state conditioning costs about a third of all successes; oracle perception shows the remaining headroom.

FastWAM on the same 300 Task 1 scenes. Arrow colors follow Figure 3’s planning stages.

  • w/o initial execution
  • w/o canonical-state conditioning
  • w/o current-state planning
  • Oracle substitution
@misc{park2026dpp,
  title  = {Dynamic Manipulation with World-Action Models via Counterfactual Planning},
  author = {Park, Sunwoo and Lee, Wonbin and Jin, Seonghyun and Kim, Youngmin and Park, Jangho and Ye, Jong Chul},
  year   = {2026},
  eprint = {2609.33172},
  archivePrefix = {arXiv},
  primaryClass = {cs.RO},
  url    = {https://arxiv.org/abs/2609.33172}
}

Matched-prefix probe

Directional grip response gain and matched-cell changes across execution prefixes 0, 8, 16, 32, and 48
Target-response collapse under matched target relocation. Directional grip-response gain (top) and its change from prefix 0 (bottom), evaluated on the same static-support mask. Means use 33 noncentral cells; gray denotes undefined gains. Color scales differ by row.
Prefix (actions)Mean gainDirection cosineArea retentionXY error (mm)Success
00.8970.999100.00%23.495.77%
80.8060.99877.82%34.484.13%
160.6360.99746.24%58.367.20%
320.2250.9306.12%118.030.16%
480.0780.6491.47%136.119.58%

FastWAM on grab_roller, three paired repeats per cell, target offsets from −240 to 240 mm (x) and −120 to 240 mm (y). Gain, cosine, area, and error use the static-support set; success uses all 63 cells under a 400-action budget. Unshifted controls remain successful at every tested prefix.

Figure 3. DPP pipeline

DPP pipeline: initial planning, canonical planning with activation, and current-state planning.
An initial rollout predicts interaction timing while available actions begin execution. The canonical planning stage uses the familiar robot context and predicted target position. An activation bridge connects its trajectory to ongoing execution, after which current-state planning uses fresh observations and current robot-state input while retaining the predicted interaction position as its planning target. Click to view at full resolution.

Counterfactual observation synthesis

Counterfactual target removal, background restoration, and reprojection
Depth-aware counterfactual observation synthesis. Editing starts from the selected observation. Source-target removal is followed by RGB inpainting and depth restoration using aligned background memory where available. Target reprojection uses scene depth to resolve visibility while preserving the selected robot context.

Success rate (%) on tasks T3–T6

MethodT3 · BottlesT4 · BellT5 · HandoverT6 · Stacking
FastWAM (Refresh-24)1.503.0011.000.00
FC-FastWAM (Oracle)24.0011.0058.008.00
DynamicWAM—5.0024.00—
PUMA—4.0019.00—
π0.53.500.0012.000.00
DPP (FastWAM)72.0063.0091.0064.00

T1 and T2 values are in the main table (slide 7). T3 pools the Dual and Diverse Bottle Picking variants (100 trials each).

Evaluation notes

Success follows the official RoboTwin criterion; Pick-and-Place requires both subtasks. Contact is a separate physical-contact diagnostic: for T1 it requires both gripper values below 0.5 with bilateral contact for two consecutive policy steps; T2 uses its required-arm criterion. RC is route completion on a 0–100 scale.

PUMA and DynamicWAM use additional training on dynamic data; every other method, including DPP, uses frozen pretrained checkpoints. +VP denotes video prediction during DPP planning. The Refresh baselines assume zero inference latency. Both Future-Conditioned variants refresh every 8 actions with a fixed 16-step look-ahead: Oracle uses simulator-rendered future observations, while DPP bootstrap synthesizes them from motion estimated at steps 0, 2, 4, 6, and 8 and current RGB-D. Neither uses DPP’s skill-dependent interaction timing or full controller.

All methods use the same predeclared scenes and simulator seeds inside empirically validated static-support regions.

Task 1 success (%) by object speed

Method1.0–1.81.8–2.62.6–3.43.4–4.24.2–5.0All
FastWAM · Refresh-24 (Default)25.0013.338.333.330.0010.00
FastWAM · Refresh-826.6715.0013.336.673.3313.00
FastWAM · Future-Conditioned (DPP bootstrap)40.0016.6716.678.3310.0018.33
FastWAM · Future-Conditioned (Oracle)56.6723.3320.0015.0020.0027.00
ImageWAM · Refresh-16 (Default)26.6711.6711.676.670.0011.33
ImageWAM · Refresh-830.0015.0011.675.001.6712.67
ImageWAM · Future-Conditioned (DPP bootstrap)33.3321.6715.008.3313.3318.33
ImageWAM · Future-Conditioned (Oracle)43.3323.3315.0013.3318.3322.67
π0.528.3310.005.001.670.009.00
AHA-WAM21.6710.005.000.000.007.33
PUMA71.6765.0045.0033.3336.6750.33
DynamicWAM53.3366.6750.0046.6743.3352.00
FastWAM + VP · DPP95.0081.6781.6778.3380.0083.33
ImageWAM + VP · DPP96.6781.6771.6770.0073.3378.67

Nominal object speed in mm per control step. Each bin contains 60 scenes; All aggregates the 300 T1 scenes.

Figure 4. Real-world evaluation

Real-robot results and DPP intermediate observations.
Each method is evaluated over 30 trials per task. Reactive replans every 10 steps; Future-Cond. uses DPP-generated counterfactual observations with a fixed +16-step horizon. In Task B, starred methods share a policy trained on anticipation-oriented static demonstrations; Hybrid denotes reactive inference with this learned prior. Click to view at full resolution.

Figure 5. Complex target motion

Acceleration, circular motion, and position-jump evaluations with reference trajectories and simulation overlays.
Each method is evaluated over 100 trials per motion family. R uses FastWAM Refresh-24; F receives +16-step oracle observations. DPP uses FastWAM with the optional nonlinear extension and a declared motion-family prior. Click to view at full resolution.

Ablations and oracle diagnostics

VariantSR (%) ↑Δ SRContact (%) ↑Δ ContactRC ↑Δ RC
Full DPP83.33—94.67—94.21—
Component ablations
w/o initial execution50.33−33.0046.67−48.0083.96−10.25
w/o canonical-state conditioning51.67−31.6747.33−47.3377.96−16.25
w/o current-state planning80.00−3.3392.33−2.3392.82−1.39
Oracle diagnostics
Oracle target velocity92.33+9.0095.33+0.6797.22+3.01
Oracle target velocity & observation94.00+10.6798.00+3.3397.90+3.69

Without initial execution, mean bilateral gripper closure shifts from 69.1 to 89.1 steps and the fraction of canonical-plan targets outside the statically evaluated workspace rises from 27.0% to 65.0%.

SettingSuccess (%) ↑Contact (%) ↑RC ↑
FastWAM + VP
Oracle-free83.3394.6794.21
Oracle-V92.3395.3397.22
Oracle-Obs94.0098.0097.90
ImageWAM + VP
Oracle-free78.6793.3393.43
Oracle-V90.0092.6796.11
Oracle-Obs93.6794.6796.89

Oracle-V substitutes ground-truth target velocity; Oracle-Obs additionally uses simulator-rendered counterfactual observations. The Oracle-free pipeline uses off-the-shelf SAM 2 and FoundationPose with direct geometric RGB-D editing, without task-specific perception training.