HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction

📄 arXiv: 2608.16222v1 📥 PDF

作者: Jiahao Ji, Ji Ma, Runhan Zhang, Runyi Yu, Wenjia Wang, Weiheng Chi, Qianqian Peng, Weichao Yan, Yongfei Gu, Ye Tian, Ting Wu, Longwei Li, Chun Yuan, Ruoli Dai, Lei Han

分类: cs.RO, cs.AI

发布日期: 2026-08-17


💡 一句话要点

提出HiPHI数据集以解决人类运动与物体交互的高精度学习问题

🎯 匹配领域: 支柱一:机器人控制 (Robot Control) 支柱二:RL算法与架构 (RL & Architecture) 支柱七:动作重定向 (Motion Retargeting) 支柱八:物理动画 (Physics-based Animation)

关键词: 人类运动数据集 物体交互 高保真度 光学运动捕捉 FrameNet 人形智能 数据基准 运动覆盖

📋 核心要点

  1. 现有的嵌入式数据集在物理状态和交互基础方面存在显著不足,限制了人形智能的学习能力。
  2. HiPHI数据集通过光学运动捕捉技术,系统性地覆盖人类运动和交互的多样性,解决了现有数据集的局限性。
  3. 实验结果表明,HiPHI在运动覆盖和交互质量上显著优于现有数据集,为人形政策的训练和评估提供了新的基础。

📝 摘要(中文)

人形智能需要在多样化的全身运动和物理交互空间中进行学习。然而,现有的嵌入式数据集存在根本性限制:互联网规模的视频数据缺乏精确的物理状态和交互基础,而实验室运动数据集虽然提供高保真度,但行为覆盖面狭窄。为了解决这一瓶颈,本文提出了HiPHI,一个超过600小时的高保真全身人类运动数据集,旨在系统性地最大化人类运动和交互的覆盖。HiPHI基于FrameNet理论框架,采用光学运动捕捉技术,提供亚毫米级的空间标记跟踪精度。我们还引入了一个基准套件,评估运动空间的多样性、交互基础、一致性及物理AI应用。分析表明,HiPHI显著扩展了运动覆盖范围,同时保持高保真的交互质量,为训练、评估和推广人形政策奠定了可扩展的数据基础。

🔬 方法详解

问题定义:本文旨在解决现有嵌入式数据集在物理状态和交互基础方面的不足,导致人形智能学习的瓶颈。现有数据集要么缺乏精确性,要么覆盖面狭窄。

核心思路:HiPHI数据集通过光学运动捕捉技术,系统性地最大化人类运动和交互的覆盖,基于FrameNet理论框架设计,确保数据的多样性和高保真度。

技术框架:HiPHI的整体架构包括数据采集、运动捕捉、数据标注和基准评估四个主要模块。数据采集通过高精度设备进行,运动捕捉确保亚毫米级的精度,数据标注则依赖于FrameNet框架进行系统化处理。

关键创新:HiPHI的最大创新在于其高保真度和广泛的运动覆盖,显著优于现有数据集,提供了一个可扩展的基础用于训练和评估人形政策。

关键设计:在数据采集过程中,采用了高精度的光学运动捕捉设备,确保运动数据的准确性;同时,设计了针对多样性和一致性的损失函数,以优化数据集的质量。通过这些设计,HiPHI能够提供丰富的运动和交互数据。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,HiPHI在运动覆盖方面显著优于现有数据集,提供了超过600小时的高保真运动数据,提升了交互质量,成为人形政策训练的关键数据基础。

🎯 应用场景

HiPHI数据集的潜在应用领域包括机器人控制、虚拟现实、动画制作等。其高保真度和广泛的运动覆盖为人形智能的训练和评估提供了坚实的基础,未来可能推动相关领域的技术进步和应用创新。

📄 摘要(原文)

Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions. However, existing embodied datasets remain fundamentally limited: internet-scale video data lack precise physical states and interaction grounding, while laboratory motion datasets provide high fidelity but only narrow behavioral coverage. This mismatch creates a critical bottleneck for scalable humanoid policy learning. We present HiPHI, a 600+ hour scale high-fidelity whole-body human motion dataset designed to systematically maximize coverage of the human motion and interaction manifold. HiPHI is theoretically guided by FrameNet, a linguistic framework organizing human primitives. Created using an optical motion capture pipeline, HiPHI provides sub-millimeter spatial marker tracking accuracy for full-body human motion and mesh-level object trajectories. We further introduce a benchmark suite evaluating motion-space diversity, interaction grounding, object consistency, and physical AI applications. Our analyses demonstrate that HiPHI significantly expands motion coverage compared to existing motion datasets while maintaining high-fidelity interaction quality, and establishes a scalable data foundation for training, evaluating, and generalizing humanoid policies in real-world embodied tasks, where similar extensions are also applicable to motion prior models in computer graphics.