Egocentric world-action modeling
EgoGenesis
Controllable video-action generation with anchored 3D scene memory and camera-aware action geometry.
The idea
Generate more than pixels.
Preserve the action.
Real robot trajectories are expensive to collect and narrow in coverage. EgoGenesis turns limited demonstrations into controllable video-action pairs while keeping the scene, object identity, camera motion, and end-effector trajectory coherent over time.
Generated rollouts
One model, diverse embodiments.
Long-horizon egocentric generation across human hands, dexterous hands, parallel grippers, and robot-arm end-effectors.
Method
Stable worlds.
Metric actions.
EgoGenesis decouples world memory from action geometry. OAPM protects persistent scene content while tracking the current interaction state; A3D-RoPE places end-effector motion into a shared camera-aware 3D coordinate system.
Online Anchored Projective Memory
An immutable first-frame anchor preserves scene layout and identity. A replace-only recent slot refreshes the current state without overwriting that stable reference.
Action-3D Rotary Position Embedding
Skeleton and end-effector joints are unprojected with camera trajectories, then encoded as metric rotary phases in action-to-video cross-attention.
Evaluation
Coherent interactions, not just appearance.
Across appearance, geometry, action alignment, physical faithfulness, and temporal consistency, EgoGenesis delivers broad gains rather than optimizing a single perceptual score.
| Category | Method | PSNR ↑ | SSIM ↑ | LPIPS ↓ | Kpt. Err ↓ | Phys. Faith ↑ | Subj. Cons. ↑ | Bg. Cons. ↑ |
|---|---|---|---|---|---|---|---|---|
| Generic video | Wan2.1-14B-InP | 19.9910 | 0.8288 | 0.3356 | 0.1295 | 0.7444 | 0.8786 | 0.9437 |
| Wan2.2-5B-Control | 16.5896 | 0.7542 | 0.4323 | 0.0813 | 0.7296 | 0.7761 | 0.8956 | |
| Cosmos3-Nano | 19.0996 | 0.8052 | 0.3810 | 0.1115 | 0.8093 | 0.8905 | 0.9392 | |
| Egocentric video | EgoHOI | 20.2230 | 0.7883 | 0.3326 | 0.1884 | 0.7537 | 0.8590 | 0.9193 |
| Mask2IV | 18.6821 | 0.8102 | 0.3789 | 0.2086 | 0.6018 | 0.8786 | 0.9353 | |
| EgoSim-14B | 19.4114 | 0.8317 | 0.2750 | 0.0811 | 0.7759 | 0.9101 | 0.9491 | |
| CosHand | 18.7549 | 0.8033 | 0.4039 | 0.1202 | 0.4667 | 0.8089 | 0.9104 | |
| RynnWorld-TeleOp | 18.8247 | 0.8215 | 0.3913 | 0.2107 | 0.7870 | 0.8887 | 0.9466 | |
| Ours | EgoGenesis | 21.8609 | 0.8509 | 0.2399 | 0.0501 | 0.8278 | 0.8923 | 0.9546 |
Bold denotes best; underline denotes second-best.
Real-robot transfer
Synthetic experience that transfers.
Adding 400 generated trajectories to 400 real trajectories improves out-of-distribution success across four single-arm and four bimanual Tianji M6 skills.
OOD average success
53% 70%
OOD average success
77% 84%
Summary
Egocentric world-action modeling with geometry at its core.
Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present EgoGenesis, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data.
EgoGenesis builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms. Online Anchored Projective Memory preserves a first-frame 3D scene anchor while periodically refreshing a recent state during autoregressive generation. Action-3D Rotary Position Embedding encodes end-effector motion with camera-aware 3D rotary coordinates, injecting action geometry into skeleton-to-video cross-attention for precise control.
Together, these components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Augmenting 400 real trajectories with 400 EgoGenesis-generated trajectories improves out-of-distribution real-robot success from 77% to 84% on single-arm tasks and from 53% to 70% on dual-arm tasks.