生成式 AI 数据情报

追踪模型,
追到它的数据源头。

不止告诉你“有哪些数据集”。这里持续拆解最新 AIGC 模型用了什么数据、怎样清洗与训练,以及官方仍未披露什么。

最近核验核验 2026/08/16Genesis

发布 2026/08/16部分披露开放权重

查模型,也查它背后的数据逻辑。

目录直接由仓库中的 YAML 数据卡生成;更新事实源,页面随之更新。

应用场景
120 条结果发布时间:新 → 旧

训练数据策略

  1. 01

    Train all perceptual modalities in one backbone instead of separate image, video, and audio models.

  2. 02

    Spend most compute on video prediction to learn dynamics, causality, contact, and motion.

  3. 03

    Reuse the learned video representation for action prediction with limited task-specific data.

阶段与操作

pretraining规模已披露

Jointly learn image, video, and audio from the beginning, with video prediction consuming over 95 percent of training compute.

joint-multimodal-trainingself-flow
action-adaptation规模未知

Add action prediction to the existing video curriculum and train a lightweight decoder over intermediate world-model features.

curriculum-expansionaction-decoder-training

数据引用

General video training corpuspretraining · 未披露 · tens of millions of hours

Exact sources, filtering, and rights composition are not named.

没有数据卡:一手资料没有披露可识别的数据集。
Human and robot manipulation video corpusfine-tuning · 未披露 · hundreds of thousands of hours

Focused on human and robot manipulation tasks for visual intelligence.

没有数据卡:一手资料没有披露可识别的数据集。
Robot action demonstrationsfine-tuning · 未披露 · 规模未披露

Used to add action prediction and adapt the FLUX-mimic action decoder.

没有数据卡:一手资料没有披露可识别的数据集。

模型更新:内容版本监控 / 核心监控。监控源变化时进入每周复核队列。

仍然未知

  • Exact image, video, and audio sources and their licenses.
  • Data mixture ratios, deduplication, safety filters, and captioning pipeline.
  • Final model size, open-weight Dev license, image endpoint availability, and post-training recipe.

我们把“没说”也写下来。

01

一手来源优先

官方发布、论文、模型卡和代码仓库交叉核验,不用能力猜训练数据。

02

数据与策略分层

区分可下载数据集、未发布语料、合成数据和人类反馈,以及它们所在的训练阶段。

03

权利边界显式化

元数据许可不等于媒体可商用;访问方式、商用和再分发分别记录。

04

持续复核

活跃模型 14 天、普通模型 45 天、数据集 90 天触发过期检查。