DAMO Academy and Hupan Laboratory

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

A general-purpose value model that understands how far a robot is from completing an instruction—and turns that signal into progress estimation, failure detection, and rewards for real-world learning.

Dongchi Huang* Hongyin Zhang* Bohan Hou* Siteng Huang Zhian Su Hang Guo Tong Lu Zhaofeng Xu Jiahao Tang Jianfei Yang Donglin Wang Peixi Peng Mingxiu Chen Deli Zhao Xin Li

DAMO Academy, Alibaba Group Hupan Lab * Equal contribution · † Corresponding authors

RynnValue learns value through temporal distance across diverse robot embodiments and viewpoints.

01 · Method

Value, grounded in time.

A multimodal backbone predicts absolute and relative value while a language head verifies instruction alignment. Value-isolation attention keeps temporal supervision focused across long clips.

RynnValue architecture, heterogeneous training data, value heads, and downstream robot tasks
RynnValue architecture—from heterogeneous robot data to progress, failure, and reward signals.
RynnValue temporal sampling, shuffling, instruction mismatch augmentation, and value-isolation attention
Training recipe: temporal sampling, order shuffling, instruction mismatch augmentation, and value-isolation attention.

02 · Trajectory-ranking results

Generalizes beyond the training distribution.

RynnValue is evaluated on six out-of-distribution robot datasets. The 8B model achieves the strongest average trajectory-ranking result, while the ablation isolates the contribution of each design component.

Main result 0.675 average Kendall’s τ
RynnValue trajectory-ranking results on RBM-EVAL-OOD
Trajectory ranking across six out-of-distribution robot datasets.
Ablation Full model all components enabled
RynnValue ablation results for shuffling, isolation, language supervision, and random sampling
Component study on temporal-order shuffling, isolation, language supervision, random temporal sampling, and relative modeling.

03 · Reinforcement learning results

RynnValue as a real-world reward.

Online and offline robot learning consistently benefit from the richer value signal, with RynnValue highlighted in purple.