Aether AI

Causal World Action Model

Technical blog outline

Overview

Nowadays there is a hot trend for using video as policy in

1. Causality as Policy

2. From Visual Signals to Action

Introducing the model implementantation

2.1 Joint Video–Action Architecture

CWAM extends Causal World Model [//]: # (这里补充具体的 LTX 版本) video backbone which inherits from the LTX-2.3 with a dedicated action stream to jointly generate future visual observations and robot actions. Conditioned on the current observation, robot state, and a language instruction, the model represents an upcoming interaction through two complementary sequences: how the scene evolves and how the robot moves.

2.2 Training Data Collection and Unified Representation

Our data collection(need a table) how we unify the action representation how we process video data

2.3 Training Strategy and Objectives

2.4 Inference and Execution

3. Emergent Zero-Shot Generalization

Here we brief show the case of some zero-shot ability.

4. Downstream Adaptation and Evaluation

4.1 SFT strategy

How we SFT the downstream task, show the difference between pretrain and SFT. We need a chart to show the difference here.

4.2 Libero & Libero Plus

4.3 Robotwin & Robotwin Memory

5. Bridging Action Priors and High-Level Reasoning: A RoboDojo Case Study

Briefly talk about GPT6-Astra's ability and how it impacts the embodied AI. Explain how a high-level planner works with CWAM to execute more complex tasks. Trace a representative episode from the planner's decisions to the model's actions, with the responsibilities of each component made explicit.

5.1 Reasoning and Action: A Division of Labor

5.2 An Integrated Agent Architecture

5.3 Robodojo Task Performance and Costs

5.4 Causally Self-Evolving Physical Agent

6. Limitations and the Road Ahead


Refernces