Egocentric world-action modeling

EgoGenesis

Controllable video-action generation with anchored 3D scene memory and camera-aware action geometry.

Zexuan Yan1,4 Yuzhou Wu2 Yue Ma3 Zonghang He1 Kaibo Yin7 Xiaobing Tu4 Yinggui Wang4 Jinkui Ren4 Xiantao Zhang4 Shijian Wang5 Jinghong Liu6 Linfeng Zhang1,†

1 Shanghai Jiao Tong University 2 Tianji KernalMind Co., Ltd. 3 The Hong Kong University of Science and Technology 4 Alibaba Group 5 Southeast University 6 Renmin University of China 7 The University of Tokyo

Paper coming soon Code coming soon

The idea

Generate more than pixels.
Preserve the action.

Real robot trajectories are expensive to collect and narrow in coverage. EgoGenesis turns limited demonstrations into controllable video-action pairs while keeping the scene, object identity, camera motion, and end-effector trajectory coherent over time.

EgoGenesis augments demonstrations by changing end-effector, object, and environment, then uses generated trajectories to tune a world-action model.
From scarce demonstrations to diverse, action-aligned data for downstream world-action model tuning.
210Ksource-balanced training clips
79.3% ↓Depth Error at frame 80 vs. RoPE
78.3% ↓Camera Error at frame 80 vs. RoPE
+17 ptsdual-arm OOD success with synthetic data

Generated rollouts

One model, diverse embodiments.

Long-horizon egocentric generation across human hands, dexterous hands, parallel grippers, and robot-arm end-effectors.

Generated
Skeleton
AgiBot · dual-arm robotTabletop object manipulation
Generated
Skeleton
AgiBot · dual-arm robotManipulate and fold a cloth
Generated
Skeleton
EgoDex · human handsAdd and remove container lids
Generated
Skeleton
EgoDex · human handsAssemble and disassemble Lego blocks
Generated
Skeleton
RoboTwin · FrankaClick a bell with Franka grippers
Real RGB
Skeleton
Real scene · hand motionSingle-arm pick and place

Method

Stable worlds.
Metric actions.

EgoGenesis decouples world memory from action geometry. OAPM protects persistent scene content while tracking the current interaction state; A3D-RoPE places end-effector motion into a shared camera-aware 3D coordinate system.

EgoGenesis architecture with autoregressive diffusion transformer, Action-3D RoPE, and Online Anchored Projective Memory.
Autoregressive generation with two geometry-aware conditioning mechanisms.
01

Online Anchored Projective Memory

An immutable first-frame anchor preserves scene layout and identity. A replace-only recent slot refreshes the current state without overwriting that stable reference.

OAPM ablation comparing full anchored and recent memory with anchor-only and no-memory variants.
02

Action-3D Rotary Position Embedding

Skeleton and end-effector joints are unprojected with camera trajectories, then encoded as metric rotary phases in action-to-video cross-attention.

79.3% lower Depth Error vs. RoPE 78.3% lower Camera Error vs. RoPE
A3D-RoPE reduces accumulated depth and camera error over an 80-frame rollout.

Evaluation

Coherent interactions, not just appearance.

Across appearance, geometry, action alignment, physical faithfulness, and temporal consistency, EgoGenesis delivers broad gains rather than optimizing a single perceptual score.

Qualitative comparison of EgoGenesis against RynnWorld-TeleOp, Cosmos3, EgoHOI, and Wan2.2 on shorts flattening and square-table assembly.
Qualitative comparison on unseen embodiments and contact-rich tasks. EgoGenesis remains close to the reference without morphology or object drift.
Generation quality on test trajectories under identical scene and action conditioning.
CategoryMethodPSNR ↑SSIM ↑LPIPS ↓Kpt. Err ↓Phys. Faith ↑Subj. Cons. ↑Bg. Cons. ↑
Generic videoWan2.1-14B-InP19.99100.82880.33560.12950.74440.87860.9437
Wan2.2-5B-Control16.58960.75420.43230.08130.72960.77610.8956
Cosmos3-Nano19.09960.80520.38100.11150.80930.89050.9392
Egocentric videoEgoHOI20.22300.78830.33260.18840.75370.85900.9193
Mask2IV18.68210.81020.37890.20860.60180.87860.9353
EgoSim-14B19.41140.83170.27500.08110.77590.91010.9491
CosHand18.75490.80330.40390.12020.46670.80890.9104
RynnWorld-TeleOp18.82470.82150.39130.21070.78700.88870.9466
OursEgoGenesis21.86090.85090.23990.05010.82780.89230.9546

Bold denotes best; underline denotes second-best.

Real-robot transfer

Synthetic experience that transfers.

Adding 400 generated trajectories to 400 real trajectories improves out-of-distribution success across four single-arm and four bimanual Tianji M6 skills.

Complete dual-arm robot bottle-handoff rollout.
Bimanual policy
Bottle handoff8× playback

OOD average success

53% 70%
Complete single-arm robot cube-stacking rollout.
Single-arm policy
Cube stacking8× playback

OOD average success

77% 84%
Real-robot rollout sequences for four bimanual and four single-arm tasks.
Eight real-robot skills used for downstream in-distribution and out-of-distribution evaluation.

Summary

Egocentric world-action modeling with geometry at its core.

Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present EgoGenesis, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data.

EgoGenesis builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms. Online Anchored Projective Memory preserves a first-frame 3D scene anchor while periodically refreshing a recent state during autoregressive generation. Action-3D Rotary Position Embedding encodes end-effector motion with camera-aware 3D rotary coordinates, injecting action geometry into skeleton-to-video cross-attention for precise control.

Together, these components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Augmenting 400 real trajectories with 400 EgoGenesis-generated trajectories improves out-of-distribution real-robot success from 77% to 84% on single-arm tasks and from 53% to 70% on dual-arm tasks.