Open Frameworks and Self‑Distillation Boost World‑Action Models' Real‑Robot Success
Two new arXiv papers present methods that raise benchmark scores by fixing a mismatch between model‑generated video and executable robot actions.
Overview
- The two papers were posted Oct. 7–8, 2026 and introduce complementary advances: OPENWAM builds an open, causal robot–video pretraining platform and AutodidactWAM supplies a pipeline to recover actions from generated video.
- OPENWAM reports large-scale causal robot–video pretraining on over 10,000 hours and a Mixture‑of‑Transformers interaction design that lifts LIBERO‑Long visual‑then‑action (VTA) success from 68.4% to 97.8%.
- AutodidactWAM shows a clear video–action asymmetry when adapting a pretrained WAM to a Unitree G1 with BrainCo hands, and recovers actions from generated video via a hand‑pose estimator plus inverse kinematics to raise pre‑grasp, grasp, and pick‑and‑place success from about 17%/10%/7% to roughly 75%/47%/42%.
- Long‑WAM finds that longer visual context matters most when the video backbone is pretrained autoregressively, reports 107.4 ms per action chunk on an RTX 5090 for real‑time control, and demonstrates 95% success on a dynamic cup‑stacking task in real robots.
- Across papers, counterfactual supervision, frozen inverse dynamics, and hybrid training objectives (preference+DPO, supervised fitting, trajectory anchoring) improve closed‑loop performance while single‑objective contrastive adaptation can produce strong validation metrics yet fail in execution, pointing to the need for combined objectives and real‑robot evaluation.