Sekai2: From World Exploration to Interactive World Modeling
作者: Kang He, Wenshuo Peng, Zihui Gao, Jiaming Tan, Kaipeng Zhang, Yongtao Ge
分类: cs.CV
发布日期: 2026-08-10
备注: Sekai2 dataset technical report. Developed at Alaya Lab
💡 一句话要点
提出Sekai2以解决长视频生成与交互建模问题
🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)
关键词: 长视频生成 交互建模 多源数据集 相机轨迹 层次化注释 场景表示 空间记忆
📋 核心要点
- 现有视频数据集缺乏长时间、相机轨迹和时间对齐的语义信息,限制了视频世界模型的训练效果。
- Sekai2数据集通过整合长视频、相机轨迹和层次化注释,提供了丰富的训练资源,支持交互式世界建模。
- 实验结果表明,Sekai2在姿态和描述覆盖、地理与语义多样性等方面表现优异,显著提升了模型的生成能力。
📝 摘要(中文)
视频世界模型必须捕捉场景随时间和视角的演变。为了实现长时间生成和相机控制,模型训练需要配备长视频、相机轨迹和时间对齐的语义信息。然而,现有数据集通常缺乏这三者的结合。为此,本文提出了Sekai2,一个多源真实世界视频数据集,包含来自113个国家或地区的128,892个片段,总计2,826小时,重点关注持续观察。每个片段都附带相机轨迹和层次化注释,提供649,597个时间对齐的片段。此外,数据集中还引入了982个沿非线性轨迹捕获的全景序列,提供了对同一地点的重复观察。这些特性使Sekai2成为长时间视频生成、相机可控合成和交互世界模型预训练的可扩展资源。
🔬 方法详解
问题定义:本论文旨在解决现有视频数据集中缺乏长时间、相机轨迹和时间对齐语义信息的问题。这些缺失导致视频世界模型在生成和控制方面的能力受限。
核心思路:Sekai2通过整合来自多个来源的真实世界视频,提供长时间观察和丰富的注释信息,以支持交互式世界建模。这样的设计旨在增强模型对场景动态和相机行为的理解。
技术框架:Sekai2数据集包含128,892个视频片段,配备相机轨迹和层次化注释。数据集的构建过程包括视频采集、注释生成和数据整合,确保了数据的多样性和覆盖面。
关键创新:最重要的创新在于引入了982个沿非线性轨迹捕获的全景序列,这些序列提供了对同一地点的重复观察,促进了模型对持久场景表示和长期空间记忆的学习。
关键设计:数据集中每个片段的注释包括主体运动、环境动态、静态场景内容和相机行为,采用层次化结构进行组织,确保了信息的清晰和可用性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,Sekai2在姿态和描述覆盖方面实现了完全覆盖,展现出广泛的地理和语义多样性。数据集中的43,594个片段达到了完整的两分钟,显著提升了模型在长时间视频生成和相机控制方面的性能。
🎯 应用场景
Sekai2数据集的潜在应用领域包括长时间视频生成、相机可控合成和交互式世界模型的预训练。这些应用可以广泛用于虚拟现实、增强现实、机器人导航和自动驾驶等领域,推动相关技术的发展与创新。
📄 摘要(原文)
Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefore benefits from long videos paired with camera trajectories and temporally grounded semantics. Existing corpora rarely offer the three together: large-scale web video provides broad visual diversity but no trajectories or time-aligned text, while pose-annotated datasets are typically short-range or reconstruction-oriented. We introduce Sekai2, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling. The release contains 128,892 clips totaling 2,826 hours from 10,428 source videos across 113 countries or regions, and is deliberately weighted toward sustained observation: under a common 120-second decomposition, 43,594 segments reach the full two minutes and account for 51.4% of all footage. Every clip includes a released camera trajectory and hierarchical annotations disentangling subject motion, environment dynamics, static scene content, and camera behavior, resulting in 649,597 temporally grounded segments. Crucially, we further introduce 982 panoramic sequences captured along non-linear trajectories with loops and revisits. These revisits provide repeated observations of the same locations across time and viewpoints, offering essential supervision for learning persistent scene representations, long-term spatial memory, and geometrically consistent world models. Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant temporal descriptions. Together, these properties make Sekai2 a scalable resource for long-horizon video generation, camera-controllable synthesis, and interactive world-model pre-training.