H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
作者: Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao, Zongrui Wang, Zeyu Zhang, Yao Mu, Guangtao Zhai, Ning Liu
分类: cs.RO, cs.CV
发布日期: 2026-08-13
💡 一句话要点
提出H2R-Bench以解决人机操作视频生成的跨体现问题
🎯 匹配领域: 支柱一:机器人控制 (Robot Control) 支柱二:RL算法与架构 (RL & Architecture) 支柱六:视频提取与匹配 (Video Extraction) 支柱七:动作重定向 (Motion Retargeting)
关键词: 人机交互 视频生成 机器人学习 跨体现转移 操作视频 基准评估 功能接触 任务执行
📋 核心要点
- 现有方法在将人类操作视频转化为机器人操作视频时面临体现一致性和功能交互的挑战。
- 论文提出H2R-Bench基准,通过系统评估模型在跨体现转移中的表现,推动人机操作视频生成的研究。
- 实验结果表明,现有视频世界模型在任务执行和体现一致性方面存在显著不足,亟需改进。
📝 摘要(中文)
大规模的操作数据对于机器人学习至关重要,但收集机器人演示仍然昂贵且难以扩展。同时,丰富的自我中心人类操作视频提供了丰富的行为经验,但由于人手与机器人末端执行器之间的差异,跨体现的转移仍然具有挑战性。为此,我们提出了H2R-Bench,这是一个用于评估跨体现人机操作视频生成的基准,模型将自我中心的人类演示转化为指定体现下的机器人操作视频。每个基准实例包含一个人类演示视频、目标体现约束和覆盖任务目标、动作事件、功能接触和对象响应的源基础注释。H2R-Bench通过五个维度评估生成的视频,包括目标状态完成、动作事件完成、功能接触转移、体现正确性和视频质量。我们在六个操作家族和两个机器人体现上基准测试了十一种最先进的视频生成模型,结果显示当前视频世界模型在跨体现转移方面仍然有限。
🔬 方法详解
问题定义:本论文旨在解决人类操作视频向机器人操作视频的跨体现转移问题。现有方法在体现一致性、功能交互和任务执行方面存在显著不足,限制了其应用。
核心思路:论文的核心思路是通过H2R-Bench基准,系统评估和比较不同模型在跨体现转移中的表现,推动视频生成技术的发展。通过将人类演示转化为机器人操作视频,提供有效的训练资源。
技术框架:H2R-Bench的整体架构包括数据收集、模型训练和评估三个主要阶段。每个基准实例由人类演示视频、目标体现约束和源基础注释组成,评估维度涵盖目标状态、动作事件、功能接触等。
关键创新:H2R-Bench的最大创新在于提供了一个系统的诊断框架,能够评估视频世界模型在跨体现转移中的有效性,填补了现有研究的空白。
关键设计:在模型设计中,采用了多维度评估指标,包括目标状态完成度、动作事件完成度等,确保生成视频的质量和功能性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,当前最先进的视频生成模型在H2R-Bench基准上表现有限,尤其在体现一致性和功能交互方面,领先模型的成功率低于50%。这一发现强调了跨体现转移研究的重要性和紧迫性。
🎯 应用场景
该研究的潜在应用领域包括机器人学习、自动化操作和人机协作等。通过有效的跨体现转移,机器人能够更好地学习和执行复杂的操作任务,从而提升其在实际环境中的应用价值和效率。
📄 摘要(原文)
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce H2R-Bench, a benchmark for evaluating cross-embodiment human-to-robot manipulation video generation, where models transform egocentric human demonstrations into robot manipulation videos under specified embodiments. Each benchmark instance contains a human demonstration video, target embodiment constraints, and source-grounded annotations covering task goals, action events, functional contacts, and object responses. H2R-Bench evaluates generated videos through five dimensions, including goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. We benchmark eleven state-of-the-art video generation models across six manipulation families and two robot embodiments. Our evaluation reveals that current video world models remain limited in human-to-robot manipulation transfer: even leading models often fail in embodiment consistency, functional interaction, and task execution. H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.