Overview
Nowadays there is a hot trend for using video as policy in
1. Causality as Policy
2. From Visual Signals to Action
Introducing the model implementantation
2.1 Joint Video–Action Architecture
CWAM extends Causal World Model [//]: # (这里补充具体的 LTX 版本) video backbone which inherits from the LTX-2.3 with a dedicated action stream to jointly generate future visual observations and robot actions. Conditioned on the current observation, robot state, and a language instruction, the model represents an upcoming interaction through two complementary sequences: how the scene evolves and how the robot moves.
2.2 Training Data Collection and Unified Representation
Our data collection(need a table) how we unify the action representation how we process video data
2.3 Training Strategy and Objectives
2.4 Inference and Execution
3. Emergent Zero-Shot Generalization
Here we brief show the case of some zero-shot ability.
4. Downstream Adaptation and Evaluation
4.1 SFT strategy
How we SFT the downstream task, show the difference between pretrain and SFT. We need a chart to show the difference here.
4.2 Libero & Libero Plus
4.3 Robotwin & Robotwin Memory
5. Bridging Action Priors and High-Level Reasoning: A RoboDojo Case Study
Briefly talk about GPT6-Astra's ability and how it impacts the embodied AI. Explain how a high-level planner works with CWAM to execute more complex tasks. Trace a representative episode from the planner's decisions to the model's actions, with the responsibilities of each component made explicit.