Aether AI

Causal World Action Model

Technical blog outline

Overview

Nowadays there is a hot trend for using video as policy in

1. Causality as Policy

2. From Visual Signals to Action

Introducing the model implementantation

2.1 Joint Video–Action Architecture

CWAM extends the Causal World Model [1] video backbone, itself built on LTX-2.3 [2], with a dedicated action stream to jointly generate future visual observations and robot actions. Conditioned on the visual observation, proprioceptive state, and a language instruction, the model predicts future video and robot actions through a mixture-of-transformers [3] architecture.

The video branch inherits the pretrained weights of the Causal World Model backbone and processes compact spatiotemporal latents produced by its video autoencoder. The action branch embeds continuous control vectors through a learned MLP. At each transformer layer, modality-specific queries, keys, and values are combined for joint attention, allowing video and action tokens to exchange information while retaining separate projections, feed-forward networks, and output heads. Language conditions both branches through text cross-attention, and proprioception enters as a dedicated conditioning token accessible to both streams. The current visual observation remains fixed as a clean latent throughout denoising, anchoring the prediction of future video and actions.

The video branch keeps the full LTX-2.3 width (4096, ~13B parameters). The action branch is far narrower (residual width 1024, ~2B parameters), which brings the total to roughly 15B. Modality-specific projections map these representations into a shared attention space, enabling interaction between branches of different widths.

2.2 Training Data Collection and Unified Representation

Data Filtering Pipeline

Noisy trajectories and misaligned video–action pairs can undermine joint video–action learning. Inspired by Qwen-RobotManip [13] and LingBot-VLA 2.0 [14], we apply a quality-control pipeline to all data sources before training.

We screen each source for structural errors, motion discontinuities, stale sensor readings, physical-limit violations, and inconsistencies between actions, states, and instructions. Statistical outliers are flagged only when supported by temporal, physical, or independent evidence, preserving legitimate pauses and fast movements. Deterministic unit and coordinate errors are corrected; invalid end-effector poses are reconstructed through forward kinematics where possible.

Confirmed bad frames are marked without altering the original timeline. During conversion, these flags are expanded by one frame on either side to define clean segments. A training sample is retained only when its full prediction horizon lies within a clean segment and has a matching instruction. Episodes with unrecoverable action or state dimensions are excluded, and normalization statistics omit confirmed bad frames.

Data curation retains 15.8k trainable hours from 19.5k raw hours. At load time, we further require valid wrist-camera action representations and source frame rates of 15, 30, or 60 Hz, enabling integer-stride video sampling at 7.5 fps.

Figure X. Data filtering pipeline. Raw trajectories undergo quality checks and deterministic corrections before conversion into clean, instruction-aligned training windows. Load-time admission enforces frame-rate compatibility and wrist-action validity before dataset mixture weights are computed.

Data

CWAM is trained in three stages, each on a different kind of data (Table 1). The video backbone first learns from robot and egocentric video without actions. Joint video–action pretraining then covers a broad, cross-embodiment action corpus. Fine-tuning finally targets each benchmark's own demonstrations.

Table 1. Pretraining data by stage.

Stage Data Scale Views Actions
Video pretraining I Robot + egocentric video (46 subsets) ~19.8M video–caption pairs single –
Video pretraining II AgiBot, DROID, Galaxea, RoboCoin, RoboCasa, RoboMIND, RoboTwin 0.9M scenes / ~1.45M multi-view clips 2–4 –
Video–action pretraining 15 sources: real robot, UMI and simulation (Table 2) 1.75M episodes, 15.8k hours 1–4 ✓

Table 2. Video–action pretraining mixture. Sampling is first balanced across three categories (UMI 0.2 / Real robot 0.6 / Simulation 0.2), then proportional to the number of valid training windows within each category. Each window is indexed by a starting timestep with a complete action horizon and its corresponding video window; adjacent windows may overlap.

Category Source Valid training windows (M) Sampling share
UMI (0.20) GenRobot Gendas dual-gripper [3] 590.1 20.0%
Real robot (0.60) AgiBot World G1 [4] 265.7 23.2%
ABC130K, i2rt YAM bimanual [5] 243.9 21.3%
AgiBot World G2 (real-world episodes) [6] 43.9 3.8%
RealSource RS-02 [7] 36.2 3.2%
Galaxea Open-World, R1 Lite [8] 27.0 2.4%
RoboCOIN, AgileX dual Piper [9] 26.2 2.3%
RoboMIND 2, Franka [10] 16.4 1.4%
RoboCOIN, Galaxea R1 Lite [9] 10.2 0.9%
DROID [11] 9.8 0.9%
RoboCOIN, RealMan [9] 6.5 0.6%
Simulation (0.20) InternData-A1, lift2 [12] 153.3 7.9%
InternData-A1, split_aloha [12] 150.6 7.8%
InternData-A1, Franka [12] 65.4 3.4%
InternData-A1, Genie-1 [12] 16.8 0.9%
AgiBot World G2 (simulated episodes) [6] 1.1 0.1%
Total 1,663 100%

2.3 Training Strategy and Objectives

CWAM is pretrained in three stages (Table 3). The first two teach a video generator how robot scenes evolve, first from a single camera and then from several cameras at once. The third adds the action branch and trains video and actions together on the cross-embodiment corpus of §2.2.

Table 3. Pretraining stages.

Video pretraining I Video pretraining II Video–action pretraining
Initialization LTX-2.3 22B [2] Stage I, step 60k Video branch: Stage II, step 120k · Action branch: random
Data Single-view robot + egocentric video Multi-view robot video (2–4 views) Video + actions, 15 sources (Table 2)
Trained Video transformer Video transformer Both branches + state / extrinsics encoders
Frozen VAE, text connector VAE, text connector VAE, text encoder
Steps 60k 120k 276k
Global batch - 64 512
Hardware - 32 × H200 128 × H100

Stages I–II: learning to continue a video

Both video stages fine-tune all transformer weights of LTX-2.3 with a flow-matching objective on 65-frame clips (480×640 per view, ~16 fps), conditioned on a caption. Two choices shape what the backbone learns:

  • Prediction from a clean past. With probability 0.5, a random-length prefix of the clip is kept clean: 1 latent frame half of the time, and 2–8 frames otherwise. Only the remaining frames are noised and supervised. The model therefore learns to continue an observed history instead of only generating a clip from scratch, which is the capability the policy relies on at test time.
  • Caption dropout (0.2) keeps an unconditional mode available and stops the model from leaning on text alone.

Stage II keeps the same recipe but feeds two to four synchronized views side by side, so the backbone learns to keep multiple cameras geometrically and temporally consistent before it ever sees an action.

Stage III: joint video–action pretraining

Objective. Video latents and action chunks are trained with rectified flow. For a clean sample $x_0$, noise $\epsilon$ and noise level $\sigma$, the model sees $x_\sigma=(1-\sigma),x_0+\sigma,\epsilon$ and predicts the velocity $v=\epsilon-x_0$:

$$ \mathcal{L}=\lambda_{\text{video}},\big|\hat v_{\text{video}}-v_{\text{video}}\big|^2_{\text{future frames}} +\lambda_{\text{action}},\big|M\odot(\hat v_{\text{action}}-v_{\text{action}})\big|^2, \qquad \lambda_{\text{video}}=\lambda_{\text{action}}=1 $$

The current observation stays clean ($\sigma=0$) and is excluded from the loss. The mask $M$ removes padded action dimensions and padded timesteps. One noise level is drawn per sample and shared by both modalities, from LTX's resolution-shifted logit-normal schedule, so video and actions are always equally corrupted.

Causal coupling. Throughout pretraining, video tokens cannot attend to action tokens. The video loss must therefore be solved from the observation, state and instruction alone, and the action branch learns to read the motion out of the predicted future.

Conditioning dropout. Proprioceptive state is dropped with probability 0.5, so the policy cannot simply extrapolate the arm's current motion and has to ground its actions in vision. The embodiment and control-rate clauses of the prompt are dropped independently (p = 0.15).

Optimization. The pretrained video branch and the freshly initialized action branch are optimized together but at different speeds. The action branch uses a 10× higher learning rate, so it can catch up without eroding the video prior. All downstream models start from the EMA weights at step 276k.

Table 4. Video–action pretraining hyperparameters.

Value
Objective Rectified flow, velocity target, λ_video = λ_action = 1
Noise schedule LTX-2 resolution-shifted logit-normal, shared across modalities
Optimizer Muon (matrices) + AdamW (rest), momentum 0.95, Nesterov, 5 Newton–Schulz steps
Learning rate (video / action) Muon 2e-4 / 2e-3 · AdamW 1e-5 / 1e-4
Schedule 5k warmup, cosine (300k horizon), stopped at 276k
Weight decay · grad clip 0.01 · 1.0
EMA Power EMA, s = 0.1
Global batch · steps · samples 512 · 276k · 141M
State dropout · prompt-clause dropout 0.5 · 0.15
Attention mask Video does not attend to actions

3. Emergent Zero-Shot Generalization

Here we brief show the case of some zero-shot ability.

4. Downstream Adaptation and Evaluation

4.1 SFT strategy

How we SFT the downstream task, show the difference between pretrain and SFT. We need a chart to show the difference here.

4.2 Libero & Libero Plus

4.3 Robotwin & Robotwin Memory

5. Bridging Action Priors and High-Level Reasoning: A RoboDojo Case Study

Briefly talk about GPT6-Astra's ability and how it impacts the embodied AI. Explain how a high-level planner works with CWAM to execute more complex tasks. Trace a representative episode from the planner's decisions to the model's actions, with the responsibilities of each component made explicit.

5.1 Reasoning and Action: A Division of Labor

5.2 An Integrated Agent Architecture

5.3 Robodojo Task Performance and Costs

5.4 Causally Self-Evolving Physical Agent

6. Limitations and the Road Ahead


Refernces