R²-WAMarXiv ↗

WORLD ACTION MODELS · POST-TRAINING

R²-WAM

Repair-and-Reject Post-Training for World Action Models

R²-WAM repairs action-conditioned video predictions and uses imagined consequences to selectively reject inferior actions during policy post-training.

Ruiyan Xu1 · Haisheng Su2 · Sixu Lin1 · Zhaokun Yue3 · Chengming Hu2 · Xin Jin4 · Guiliang Liu1,*

1 CUHK-Shenzhen   2 Manifold AI   3 Southeast University   4 Chang’an University

* Corresponding author

01 / DEMONSTRATIONS

Real-world demonstrations

93.8%

RoboTwin 2.0 average success

87.5%

Real-world Fold Shirt average success

2 stages

Predictive repair → Action rejection

02 / OVERVIEW

Consequence-based policy refinement

Visually plausible video predictions can mislead policy refinement when they fail to reflect the input actions. R²-WAM first improves the consistency of predicted futures with input actions, then uses these futures to select inferior action samples for negative fine-tuning.

R²-WAM extends video prediction from representation learning to consequence feedback. Results and execution sequences illustrate its performance in simulation and on real robots.

03 / METHOD

Predictive repair and action rejection

R²-WAM is a two-stage post-training framework, without additional environment interaction or changes to the inference procedure.

01

Predictive repair

Given the same observation and demonstrated action, compare the predicted future with the observed robot behavior. A kinematic alignment score selects predictions with substantial motion discrepancies for negative fine-tuning.

Video expert refinement
02

Action rejection

With the repaired video expert frozen, compare predicted task progress for sampled and demonstrated actions from the same initial observation. Select sampled actions that fall below the demonstrated reference by a prescribed margin for negative fine-tuning.

Action policy refinement

Both stages retain flow matching on demonstrations and apply negative fine-tuning to selected generated samples.

Predictive repair uses observed futures to identify action-inconsistent predictions. Action rejection uses the repaired video expert to select inferior actions by comparing their predicted progress with the demonstrated reference.

04 / RESULTS

Experimental results

Simulation performance across 50 RoboTwin 2.0 tasks, alongside real-world evaluation over seen and unseen conditions.

Simulation results

RoboTwin 2.0

Average success · Clean and randomized settings

Action-conditioned Fast-WAM 90.5%
Fast-WAM 91.9%
LingBot-VA 92.3%
R²-WAM 93.8%

REAL-WORLD FOLD SHIRT

87.5%

Average success across four evaluation conditions.

Fast-WAM0%

20 autonomous trials per condition per method. Four conditions vary initial position and object instance.

R²-WAM achieves 93.8% success in clean environments and 93.7% under randomization, improving over action-conditioned Fast-WAM by 3.2 and 3.3 percentage points, respectively.

Real-world results and qualitative analysis

On the long-horizon Fold Shirt task, R²-WAM achieves 87.5% average success, while Fast-WAM and Motus complete none of the evaluation trials. It also outperforms the VLA baselines across all four conditions. On Clean Table, R²-WAM achieves 70.0% success when both object position and instance vary, compared with 40.0% for the strongest baseline, π₀.₅.

(a, b) Success rates on Fold Shirt and Clean Table. (c) Fast-WAM executions and predicted consequences of actions from Fast-WAM and R²-WAM during shirt folding. Human intervention enables later-stage analysis in (c); these rollouts are excluded from success rates.

During shirt folding, Fast-WAM grasps or lifts the garment but fails to complete the subsequent fold or pull. The repaired video expert predicts a similar lack of progress for these actions, whereas predictions for R²-WAM actions show forward folding, garment pulling, and lateral folding. These examples suggest that predicted consequences can help identify stalled actions for rejection during post-training.

BibTeX

@misc{xu2026r2wam,
  title         = {{R$^2$-WAM}: Repair-and-Reject Post-Training
                   for World Action Models},
  author        = {Ruiyan Xu and Haisheng Su and Sixu Lin and
                   Zhaokun Yue and Chengming Hu and Xin Jin and
                   Guiliang Liu},
  year          = {2026},
  eprint        = {2610.04913},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2610.04913}
}

Enlarged paper figure