Particle.news

OpenWAM and AutodidactWAM Push World-Action Models Toward Real-Robot Control

Large-scale causal pretraining that raises benchmark performance alongside self-distillation adaptation that recovers actions on real robots still leave full-task success limited.

Overview

  • Both preprints were posted on Wednesday, Oct. 7, 2026, reporting complementary advances in world-action models that jointly predict future video and generate robot actions.
  • OPENWAM introduces an open, configurable WAM framework and performs causal robot-video pretraining on more than 10,000 hours of video, reporting large benchmark gains such as raising VTA success on LIBERO-Long from 68.4% to 97.8%.
  • OPENWAM uses a shared Mixture-of-Transformers to integrate an action expert and shows that counterfactual transitions improve inverse and forward dynamics, with a frozen local-context inverse model reaching 84.0% mean success on held-out LIBERO tasks versus much lower baselines.
  • AutodidactWAM documents a video-action asymmetry when a Cosmos 3 WAM is adapted to a Unitree G1 with BrainCo hands, finding plausible generated video but very low native action success (about 17%, 10%, and 7% at pre-grasp, grasp, and pick-and-place).
  • AutodidactWAM recovers actions by running a hand-pose estimator on generated video and producing inverse-kinematics targets for action-layer fine-tuning, and shows a hybrid objective (DPO+SFT+DTW) markedly improves closed-loop success though full end-to-end task success remains modest and varies by object.