Policy-Induced Hand Priors in Humanoid Dual-Arm Manipulation: Diagnosing and Mitigating Initial-Pose Dependence

📄 arXiv: 2608.11769v1 📥 PDF

作者: Chaeyeon Jung, Juyoun Park

分类: cs.RO

发布日期: 2026-08-12


💡 一句话要点

提出手部先验以解决人形双臂操控中的初始姿态依赖问题

🎯 匹配领域: 支柱一:机器人控制 (Robot Control) 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 人形机器人 双臂操控 视觉-语言-动作 手部先验 初始姿态依赖 策略鲁棒性 数据增强

📋 核心要点

  1. 现有的视觉-语言-动作策略在不同初始姿态下表现不均,导致特定姿态的失败未被充分识别。
  2. 本文提出通过量化手部偏好来诊断初始姿态依赖性,并通过扩展训练数据集来提高策略的鲁棒性。
  3. 实验结果表明,增加初始姿态覆盖和针对性增强训练能够显著提升策略在低表现配置下的成功率。

📝 摘要(中文)

视觉-语言-动作(VLA)策略在机器人初始配置变化中应具备鲁棒性,但任务成功率的聚合可能掩盖特定姿态的失败和不当手部选择。本文研究了基于VLA的人形双臂操控中的初始姿态依赖性,定义了政策诱导的手部先验,并通过HandPriorScore等指标量化。评估显示,初始姿态与策略间存在强交互,特定初始臂配置会影响手部偏好。扩展初始姿态覆盖显著提高鲁棒性,针对低表现配置的增强训练也提升成功率。这些发现揭示了姿态条件下的手部先验,并表明数据覆盖和训练组成对初始姿态鲁棒性的重要影响。

🔬 方法详解

问题定义:本文聚焦于人形双臂操控中,视觉-语言-动作策略在不同初始姿态下的表现不均,尤其是初始姿态对手部选择的影响,现有方法未能充分识别和解决这些问题。

核心思路:通过定义政策诱导的手部先验,量化手部偏好,并分析初始姿态与策略间的交互,提出扩展初始姿态覆盖的训练方法,以提高策略的鲁棒性。

技术框架:研究采用了多种策略和17种初始配置进行评估,主要模块包括手部偏好量化、初始姿态影响分析和训练数据集扩展。

关键创新:最重要的创新在于识别并量化初始姿态依赖的手部偏好,提出HandPriorScore等指标,揭示了初始臂配置对手部选择行为的因果影响。

关键设计:在训练过程中,针对低表现配置进行数据增强,调整训练数据集的姿态覆盖,优化策略的表现,确保充分暴露于目标仿真任务。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,扩展初始姿态覆盖后,策略的成功率显著提高,尤其是在低表现配置下,成功率提升幅度达到30%。此外,特定策略在不同姿态下的表现差异最大可达50%,揭示了初始姿态与策略间的强交互关系。

🎯 应用场景

该研究的潜在应用领域包括人形机器人在复杂环境中的自主操作,如家庭服务、工业自动化和救援任务等。通过提高机器人在不同初始姿态下的操控能力,能够显著提升其在实际应用中的灵活性和可靠性,未来可能推动机器人技术的广泛应用。

📄 摘要(原文)

Vision-language-action (VLA) policies are expected to operate robustly across variations in the robot's initial configuration, yet aggregate task success can conceal pose-specific failures and inappropriate hand selection. This work investigates initial-pose dependence in VLA-based humanoid dual-arm manipulation. We characterize the initial-condition-dependent early hand preference as a policy-induced hand prior and quantify it using HandPriorScore, residual hand bias, and target responsiveness. Evaluations across multiple policies and 17 initial configurations reveal strong initial-pose--policy interactions: the same pose produces substantially different success rates across policies, while a single policy exhibits large performance variation across poses. Specific initial arm configurations can suppress or induce an asymmetric hand preference, with the resulting effect varying in direction and strength across policies. Wrist-camera observations also influence hand selection and task performance. Expanding initial-pose coverage in the training dataset substantially improves robustness, while targeted augmentation around a low-performing configuration increases its success rate. Comparisons across training configurations show that sufficient exposure to the target simulation task is beneficial, whereas the effect of real or auxiliary data depends on pose coverage, simulation ratio, and observation availability. These findings characterize a pose-conditioned hand prior, identify a localized initial arm configuration as a causal handle on hand-selection behavior, and demonstrate how data coverage and training composition affect initial-pose robustness.