Overview
Nowadays there is a hot trend for using video as policy in
1. Causality as Policy
2. From Visual Signals to Action
Introducing the model implementantation
2.1 Joint Video–Action Architecture
CWAM extends the Causal World Model [1] video backbone, itself built on LTX-2.3 [2], with a dedicated action stream to jointly generate future visual observations and robot actions. Conditioned on the visual observation, proprioceptive state, and a language instruction, the model predicts future video and robot actions through a mixture-of-transformers [3] architecture.
The video branch inherits the pretrained weights of the Causal World Model backbone and processes compact spatiotemporal latents produced by its video autoencoder. The action branch embeds continuous control vectors through a learned MLP. At each transformer layer, modality-specific queries, keys, and values are combined for joint attention, allowing video and action tokens to exchange information while retaining separate projections, feed-forward networks, and output heads. Language conditions both branches through text cross-attention, and proprioception enters as a dedicated conditioning token accessible to both streams. The current visual observation remains fixed as a clean latent throughout denoising, anchoring the prediction of future video and actions.
The video branch keeps the full LTX-2.3 width (4096, ~13B parameters). The action branch is far narrower (residual width 1024, ~2B parameters), which brings the total to roughly 15B. Modality-specific projections map these representations into a shared attention space, enabling interaction between branches of different widths.
2.2 Training Data Collection and Unified Representation
Data Filtering Pipeline
Noisy trajectories and misaligned video–action pairs can undermine joint video–action learning. Inspired by Qwen-RobotManip [13] and LingBot-VLA 2.0 [14], we apply a quality-control pipeline to all data sources before training.
We screen each source for structural errors, motion discontinuities, stale sensor readings, physical-limit violations, and inconsistencies between actions, states, and instructions. Statistical outliers are flagged only when supported by temporal, physical, or independent evidence, preserving legitimate pauses and fast movements. Deterministic unit and coordinate errors are corrected; invalid end-effector poses are reconstructed through forward kinematics where possible.
Confirmed bad frames are marked without altering the original timeline. During conversion, these flags are expanded by one frame on either side to define clean segments. A training sample is retained only when its full prediction horizon lies within a clean segment and has a matching instruction. Episodes with unrecoverable action or state dimensions are excluded, and normalization statistics omit confirmed bad frames.
Data curation retains 15.8k trainable hours from 19.5k raw hours. At load time, we further require valid wrist-camera action representations and source frame rates of 15, 30, or 60 Hz, enabling integer-stride video sampling at 7.5 fps.
Figure X. Data filtering pipeline. Raw trajectories undergo quality checks and deterministic corrections before conversion into clean, instruction-aligned training windows. Load-time admission enforces frame-rate compatibility and wrist-action validity before dataset mixture weights are computed.
Data
CWAM is trained in three stages, each on a different kind of data (Table 1). The video backbone first learns from robot and egocentric video without actions. Joint video–action pretraining then covers a broad, cross-embodiment action corpus. Fine-tuning finally targets each benchmark's own demonstrations.
Table 1. Pretraining data by stage.
| Stage | Data | Scale | Views | Actions |
|---|---|---|---|---|
| Video pretraining I | Robot + egocentric video (46 subsets) | ~19.8M video–caption pairs | single | – |
| Video pretraining II | AgiBot, DROID, Galaxea, RoboCoin, RoboCasa, RoboMIND, RoboTwin | 0.9M scenes / ~1.45M multi-view clips | 2–4 | – |
| Video–action pretraining | 15 sources: real robot, UMI and simulation (Table 2) | 1.75M episodes, 15.8k hours | 1–4 | ✓ |
Table 2. Video–action pretraining mixture. Sampling is first balanced across three categories (UMI 0.2 / Real robot 0.6 / Simulation 0.2), then proportional to the number of valid training windows within each category. Each window is indexed by a starting timestep with a complete action horizon and its corresponding video window; adjacent windows may overlap.
| Category | Source | Valid training windows (M) | Sampling share |
|---|---|---|---|
| UMI (0.20) | GenRobot Gendas dual-gripper [3] | 590.1 | 20.0% |
| Real robot (0.60) | AgiBot World G1 [4] | 265.7 | 23.2% |
| ABC130K, i2rt YAM bimanual [5] | 243.9 | 21.3% | |
| AgiBot World G2 (real-world episodes) [6] | 43.9 | 3.8% | |
| RealSource RS-02 [7] | 36.2 | 3.2% | |
| Galaxea Open-World, R1 Lite [8] | 27.0 | 2.4% | |
| RoboCOIN, AgileX dual Piper [9] | 26.2 | 2.3% | |
| RoboMIND 2, Franka [10] | 16.4 | 1.4% | |
| RoboCOIN, Galaxea R1 Lite [9] | 10.2 | 0.9% | |
| DROID [11] | 9.8 | 0.9% | |
| RoboCOIN, RealMan [9] | 6.5 | 0.6% | |
| Simulation (0.20) | InternData-A1, lift2 [12] | 153.3 | 7.9% |
| InternData-A1, split_aloha [12] | 150.6 | 7.8% | |
| InternData-A1, Franka [12] | 65.4 | 3.4% | |
| InternData-A1, Genie-1 [12] | 16.8 | 0.9% | |
| AgiBot World G2 (simulated episodes) [6] | 1.1 | 0.1% | |
| Total | 1,663 | 100% |
2.3 Training Strategy and Objectives
CWAM is pretrained in three stages (Table 3). The first two teach a video generator how robot scenes evolve, first from a single camera and then from several cameras at once. The third adds the action branch and trains video and actions together on the cross-embodiment corpus of §2.2.
Table 3. Pretraining stages.
| Video pretraining I | Video pretraining II | Video–action pretraining | |
|---|---|---|---|
| Initialization | LTX-2.3 22B [2] | Stage I, step 60k | Video branch: Stage II, step 120k · Action branch: random |
| Data | Single-view robot + egocentric video | Multi-view robot video (2–4 views) | Video + actions, 15 sources (Table 2) |
| Trained | Video transformer | Video transformer | Both branches + state / extrinsics encoders |
| Frozen | VAE, text connector | VAE, text connector | VAE, text encoder |
| Steps | 60k | 120k | 276k |
| Global batch | - | 64 | 512 |
| Hardware | - | 32 × H200 | 128 × H100 |
Stages I–II: learning to continue a video
Both video stages fine-tune all transformer weights of LTX-2.3 with a flow-matching objective on 65-frame clips (480×640 per view, ~16 fps), conditioned on a caption. Two choices shape what the backbone learns:
- Prediction from a clean past. With probability 0.5, a random-length prefix of the clip is kept clean: 1 latent frame half of the time, and 2–8 frames otherwise. Only the remaining frames are noised and supervised. The model therefore learns to continue an observed history instead of only generating a clip from scratch, which is the capability the policy relies on at test time.
- Caption dropout (0.2) keeps an unconditional mode available and stops the model from leaning on text alone.
Stage II keeps the same recipe but feeds two to four synchronized views side by side, so the backbone learns to keep multiple cameras geometrically and temporally consistent before it ever sees an action.
Stage III: joint video–action pretraining
Objective. Video latents and action chunks are trained with rectified flow. For a clean sample $x_0$, noise $\epsilon$ and noise level $\sigma$, the model sees $x_\sigma=(1-\sigma),x_0+\sigma,\epsilon$ and predicts the velocity $v=\epsilon-x_0$:
$$ \mathcal{L}=\lambda_{\text{video}},\big|\hat v_{\text{video}}-v_{\text{video}}\big|^2_{\text{future frames}} +\lambda_{\text{action}},\big|M\odot(\hat v_{\text{action}}-v_{\text{action}})\big|^2, \qquad \lambda_{\text{video}}=\lambda_{\text{action}}=1 $$
The current observation stays clean ($\sigma=0$) and is excluded from the loss. The mask $M$ removes padded action dimensions and padded timesteps. One noise level is drawn per sample and shared by both modalities, from LTX's resolution-shifted logit-normal schedule, so video and actions are always equally corrupted.
Causal coupling. Throughout pretraining, video tokens cannot attend to action tokens. The video loss must therefore be solved from the observation, state and instruction alone, and the action branch learns to read the motion out of the predicted future.
Conditioning dropout. Proprioceptive state is dropped with probability 0.5, so the policy cannot simply extrapolate the arm's current motion and has to ground its actions in vision. The embodiment and control-rate clauses of the prompt are dropped independently (p = 0.15).
Optimization. The pretrained video branch and the freshly initialized action branch are optimized together but at different speeds. The action branch uses a 10× higher learning rate, so it can catch up without eroding the video prior. All downstream models start from the EMA weights at step 276k.
Table 4. Video–action pretraining hyperparameters.
| Value | |
|---|---|
| Objective | Rectified flow, velocity target, λ_video = λ_action = 1 |
| Noise schedule | LTX-2 resolution-shifted logit-normal, shared across modalities |
| Optimizer | Muon (matrices) + AdamW (rest), momentum 0.95, Nesterov, 5 Newton–Schulz steps |
| Learning rate (video / action) | Muon 2e-4 / 2e-3 · AdamW 1e-5 / 1e-4 |
| Schedule | 5k warmup, cosine (300k horizon), stopped at 276k |
| Weight decay · grad clip | 0.01 · 1.0 |
| EMA | Power EMA, s = 0.1 |
| Global batch · steps · samples | 512 · 276k · 141M |
| State dropout · prompt-clause dropout | 0.5 · 0.15 |
| Attention mask | Video does not attend to actions |
3. Emergent Zero-Shot Generalization
Here we brief show the case of some zero-shot ability.
4. Downstream Adaptation and Evaluation
4.1 SFT strategy
How we SFT the downstream task, show the difference between pretrain and SFT. We need a chart to show the difference here.
4.2 Libero & Libero Plus
4.3 Robotwin & Robotwin Memory
5. Bridging Action Priors and High-Level Reasoning: A RoboDojo Case Study
Briefly talk about GPT6-Astra's ability and how it impacts the embodied AI. Explain how a high-level planner works with CWAM to execute more complex tasks. Trace a representative episode from the planner's decisions to the model's actions, with the responsibilities of each component made explicit.