DAMO Academy DAMO Academy (dark) CUHK Hong Kong Embodied AI Lab
Auto

RynnWorld-Latent

A Cross-Embodiment World Model Conditioned on Universal Latent Actions

Haoyu Zhao* Yinzhou Tang* Zixuan Wang* Minghao Zhu Dongchi Huang Xingyue Zhao Qin Zhao Kehan Li Siteng Huang† Xin Li Zhongyu Li†
DAMO Academy, Alibaba Group · Hong Kong Embodied AI Lab · CUHK · Hupan Lab
* Core contributors † Corresponding authors

Abstract

We present RynnWorld-Latent, a cross-embodiment video world model designed not only for future prediction but also for scalable robot data generation. RynnWorld-Latent addresses the action-label bottleneck by using continuous latent actions as a unified representation of world-state transitions, so human egocentric videos and robot data can be jointly modeled without embodiment-specific control commands. Built on a pretrained video diffusion transformer, it is text-free and introduces explicit frame-aligned action conditioning, enabling precise control over future dynamics while preserving the pretrained video prior. Beyond prediction, it serves as a data-generation engine: conditioned on target robot actions it supports interactive digital teleoperation; conditioned on latent actions it synthesizes new rollouts by editing a source video's first frame and reusing its latent actions.

Key Highlights

  • Universal latent-action interface:
    One embodiment-agnostic action space lets human, robot and simulation video pretrain a single world model.
  • Prior-preserving frame-aligned conditioning:
    Identity-initialized, frame-aligned injection gives precise action control without overwriting the pretrained video prior.
  • World model as data engine:
    Interactive digital teleoperation and first-frame-editing + latent-action synthesis produce paired action-video data without physical execution.

Training Data

Pre-training corpus of RynnWorld-Latent. Human egocentric video, real-robot teleoperation and simulation rollouts are all relabelled by RynnLAM into one shared latent-action space, giving a single cross-embodiment conditioning signal. Click the figure to enlarge.

RynnWorld pretraining corpus

Action-Conditioned Video Generation

Given an initial frame and a supplied action sequence, RynnWorld-Latent generates the future rollout. Three embodiments × three rollouts.

GR-1
GR-1 — generated
GR-1 — generated
GR-1 — generated
Marvin-WUJI
Marvin-WUJI — generated
Marvin-WUJI — generated
Marvin-WUJI — generated
Astribot-S1
Astribot-S1 — generated
Astribot-S1 — generated
Astribot-S1 — generated

Action Following

Each strip is one pair: GT-A | Gen(A frame + A action) | Gen(A frame + B action) | GT-B. Both generations start from the SAME first frame (A's); only the action differs, so the third panel diverging from the second — and matching GT-B's motion — is direct evidence of action following.

GR-1
GR-1 pair 1 — GT-A | Gen(A+A) | Gen(A+B) | GT-B
GR-1 pair 2 — GT-A | Gen(A+A) | Gen(A+B) | GT-B
Marvin-WUJI
Marvin-WUJI pair 1 — GT-A | Gen(A+A) | Gen(A+B) | GT-B
Marvin-WUJI pair 2 — GT-A | Gen(A+A) | Gen(A+B) | GT-B
Astribot-S1
Astribot-S1 pair 1 — GT-A | Gen(A+A) | Gen(A+B) | GT-B
Astribot-S1 pair 2 — GT-A | Gen(A+A) | Gen(A+B) | GT-B

First-Frame Editing + Latent-Action Synthesis

Left: the source human-hand video. Right: the rollout generated after editing the first frame to the robot embodiment and reusing the source latent actions.

Human source | edited-first-frame robot rollout (pan washing)
Human source | edited-first-frame robot rollout (stir-fry)

Results

Open-loop simulation quality on our own real-robot evaluation sets, recorded as ground-truth rollouts on both platforms and held out from post-training (higher PSNR/SSIM and lower LPIPS are better).

MethodMarvin-WUJI EvalAstribot-S1 Eval
PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓
Cosmos3-Edge20.030.780.3113.990.420.54
Cosmos3-Nano20.090.800.2614.220.530.45
RynnWorld-Teleop20.890.760.19N/AN/AN/A
RynnWorld-Latent21.860.800.1215.960.600.34

Open-loop rollout quality on the DreamDojo GR-1 evaluation sets (higher PSNR/SSIM and lower LPIPS is better).

ModelIn-lab EvalEgoDex EvalDreamDojo-HV Eval
PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓PSNR↑SSIM↑LPIPS↓
Cosmos-Predict2.520.580.770.2219.950.790.2218.270.750.24
DreamDojo-2B21.110.770.2220.410.780.2318.810.750.24
DreamDojo-14B21.410.790.2120.530.790.2118.920.750.23
RynnWorld-Latent21.420.820.2020.890.810.1919.220.790.21

Closed-loop success rate (%) on real-world manipulation tasks, 35 trials each. Bold marks the best within each training-setting block; N/A = not applicable to Astribot-S1 (mobile-base motion not modelled).

MethodTraining SettingMarvin-WUJIAstribot-S1
Plate PlacementPlate WipingSteak ServingPlate Retrieval
Digital teleoperation
π0Real only80.0071.4337.1445.71
π0Real + Digital Teleop88.5794.2951.4351.43
π0.5Real only82.8677.1440.0045.71
π0.5Digital Teleop only57.1454.2928.5731.43
π0.5Real + RynnWorld-Teleop85.7182.86N/AN/A
π0.5Real + Digital Teleop91.4391.4354.2957.14
Large-scale synthetic mid-training
π0RynnWorld-80M + Real85.7182.8648.5754.29
π0.5Real + Ego2Robot88.5791.43N/AN/A
π0.5RynnWorld-80M + Real94.2997.1460.0068.57

BibTeX

@article{rynnworld2026latent,
  title   = {RynnWorld-Latent: A Cross-Embodiment World Model Conditioned on Universal Latent Actions},
  author  = {Haoyu Zhao and Yinzhou Tang and Zixuan Wang and Minghao Zhu and Dongchi Huang and Xingyue Zhao and Qin Zhao and Kehan Li and Siteng Huang and Xin Li and Zhongyu Li},
  year    = {2026}
}