训练数据策略
- 01
Train all perceptual modalities in one backbone instead of separate image, video, and audio models.
- 02
Spend most compute on video prediction to learn dynamics, causality, contact, and motion.
- 03
Reuse the learned video representation for action prediction with limited task-specific data.
阶段与操作
Jointly learn image, video, and audio from the beginning, with video prediction consuming over 95 percent of training compute.
joint-multimodal-trainingself-flowAdd action prediction to the existing video curriculum and train a lightweight decoder over intermediate world-model features.
curriculum-expansionaction-decoder-training数据引用
Exact sources, filtering, and rights composition are not named.
没有数据卡:一手资料没有披露可识别的数据集。Focused on human and robot manipulation tasks for visual intelligence.
没有数据卡:一手资料没有披露可识别的数据集。Used to add action prediction and adapt the FLUX-mimic action decoder.
没有数据卡:一手资料没有披露可识别的数据集。模型更新:内容版本监控 / 核心监控。监控源变化时进入每周复核队列。
仍然未知
- Exact image, video, and audio sources and their licenses.
- Data mixture ratios, deduplication, safety filters, and captioning pipeline.
- Final model size, open-weight Dev license, image endpoint availability, and post-training recipe.