- Submitted to ICLR 2027
World-action models typically inherit a complete pretrained video generator along with its predictive knowledge. V-JEPA Policy instead builds one on the frozen latent space of a V-JEPA 2.1 encoder, jointly learning an instruction-conditioned future-latent predictor and a flow-matching action expert from scratch. At 0.9B parameters, it stays competitive with world-action and vision-language-action baselines on LIBERO, LIBERO-Plus, and RoboCasa-GR1, and transfers to real-world bimanual manipulation. Action-free predictor pretraining on DROID videos further improves out-of-distribution generalization.