ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

📄 arXiv: 2607.28625v1 📥 PDF

作者: Yukang Cao, Haozhe Xie, Beichen Wen, Runmao Yao, Yinghao Liu, Yue Huang, Zhichao Liao, Yunxiang Wang, Haiheng Liu, Xingshun Tian, Dawei Su, Long Zhuo, Dacheng Tao, Xiaogang Wang, Liang Pan, Ziwei Liu

分类: cs.CV

发布日期: 2026-07-30

备注: Project Page: https://ace-data-engine.github.io/ACE-Data-0/


💡 一句话要点

提出ACE-Data-0以解决人类中心数据捕获的瓶颈问题

🎯 匹配领域: 支柱一:机器人控制 (Robot Control) 支柱二:RL算法与架构 (RL & Architecture) 支柱五:交互与反应 (Interaction & Reaction) 支柱六:视频提取与匹配 (Video Extraction) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 人类中心数据捕获 多模态融合 环境捕获引擎 模仿学习 具身人工智能 长时间序列理解 家庭场景交互

📋 核心要点

  1. 现有数据集在捕获人类行为时存在视角、模态和空间尺度的碎片化,导致感知-行动循环的完整性不足。
  2. 本文提出的ACE通过将家庭环境转化为同步的录音室,能够全面捕获人类的感知和行为数据,解决了数据瓶颈问题。
  3. ACE-Data-0数据集包含150小时的多模态数据,展示了在接触、遮挡和长时间序列下的显著性能提升,推动了模仿学习和具身AI的研究。

📝 摘要(中文)

在现有的数据集中,第一人称感知、全身运动、灵巧操作等信息往往被分割,导致感知-行动循环的观察不完整。为此,本文提出了环境捕获引擎(ACE),将真实家庭环境转化为空间校准和时间同步的录音室。ACE在桌面和房间两个尺度上进行操作,记录了多种感官信号,包括视频、运动轨迹、音频和触觉信号。基于ACE构建的ACE-Data-0数据集包含150小时、1700万帧视频,涵盖200个任务类别,提供了同步的人类演示,支持模仿学习和具身人工智能的发展。

🔬 方法详解

问题定义:本文旨在解决现有数据集在捕获人类行为时的碎片化问题,导致感知-行动循环的观察不完整,限制了具身智能的发展。

核心思路:通过引入环境捕获引擎(ACE),将真实家庭环境转化为空间校准和时间同步的录音室,全面记录人类的感知和行为数据。

技术框架:ACE在桌面和房间两个尺度上操作,分别捕获手-物体操作和全身运动。系统记录多种感官信号,包括第一人称和多视角视频、全身及手部运动、物体几何形状、音频和触觉信号。

关键创新:ACE的核心创新在于其能够同步捕获多模态数据,提供全面的感知-行动循环视图,这与传统方法的局限性形成鲜明对比。

关键设计:ACE采用了高精度的空间校准技术,确保不同模态数据的时间同步,同时设计了适应不同任务的录制配置,以支持多样化的行为捕获。该系统的灵活性和全面性为后续的研究提供了坚实基础。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

在对比现有最先进方法的评估中,ACE-Data-0展示了在接触、遮挡和长时间序列下的显著性能提升,尤其是在复杂的家庭场景中,表现出更好的鲁棒性和准确性,填补了现有研究中的重要空白。

🎯 应用场景

ACE-Data-0数据集在模仿学习、世界模型、视觉-语言-行动系统和具身人工智能等领域具有广泛的应用潜力。通过提供高质量的同步数据,该研究为智能体的训练和评估提供了新的基础,推动了人机交互和智能系统的进步。

📄 摘要(原文)

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.