ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation
作者: Mi Yan, Wenhao Zhang, Zhiqi Zhang, Yu Peng, Tangxinyu Wang, Lingfei Zhai, Jiayi Su, Shengliang Deng, Lin Peng, Yaowei Liu, Yuxing Chen, Zhiyuan Wei, Jilong Wang, Jiayi Chen, Jiangran Lyu, Zhizheng Zhang, He Wang
分类: cs.RO
发布日期: 2026-09-02
💡 一句话要点
提出ZETA以解决零-shot跨实体VLA转移问题
🎯 匹配领域: 支柱一:机器人控制 (Robot Control) 支柱七:动作重定向 (Motion Retargeting) 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 零-shot学习 跨实体转移 视觉-语言-动作 机器人操作 受控基准
📋 核心要点
- 核心问题:现有方法在零-shot泛化到未见实体时缺乏系统理解,且缺乏统一的转移定义和评估设置。
- 方法要点:本文通过区分严格零-shot转移和预训练暴露的零-shot转移,提出了一个涵盖多种实体的受控基准。
- 实验或效果:实验结果显示,局部末端执行器表示和源实体多样性显著提升了跨实体转移的性能。
📝 摘要(中文)
零-shot泛化到未见实体对于可泛化的视觉-语言-动作(VLA)模型至关重要,尤其是在机器人硬件不断演变和任务特定数据收集成本高昂的背景下。然而,现有文献对这一问题的系统理解仍然有限,部分原因在于缺乏统一的零-shot转移定义和能够隔离实体变化的受控评估设置。为此,本文首先区分严格的零-shot转移与预训练暴露的零-shot转移。接着,我们引入了一个涵盖14个保留目标实体的受控基准,并对状态-动作表示、预训练实体多样性、辅助共同训练目标和目标实体暴露等四个因素进行了分析。实验结果表明,局部末端执行器状态-动作表示、源实体多样性和辅助共同训练分别提高了跨实体转移约15、18和7个百分点。此外,仅在预训练中添加5%的目标实体数据就能将平均目标实体进展提高13.4个百分点,表明严格和预训练暴露的零-shot转移是不同的,应该分别报告。
🔬 方法详解
问题定义:本文旨在解决零-shot泛化到未见实体的挑战,现有方法在定义和评估上存在不足,无法有效隔离实体变化与任务、场景或协议的差异。
核心思路:论文通过明确区分严格零-shot转移和预训练暴露的零-shot转移,提出了一个新的受控基准,以便更好地评估和理解跨实体转移的性能。
技术框架:整体架构包括四个主要模块:状态-动作表示、预训练实体多样性、辅助共同训练目标和目标实体暴露。通过对这些因素的系统分析,评估其对跨实体转移的影响。
关键创新:最重要的创新在于引入了严格零-shot转移与预训练暴露的零-shot转移的区分,并提供了一个涵盖多种实体的受控基准,这在现有文献中尚属首次。
关键设计:在实验中,局部末端执行器状态-动作表示被证明是有效的,源实体的多样性和辅助共同训练目标的设置也显著提升了转移性能,具体参数和损失函数的设计在文中有详细描述。
🖼️ 关键图片
📊 实验亮点
实验结果显示,局部末端执行器状态-动作表示、源实体多样性和辅助共同训练目标分别提升了跨实体转移约15、18和7个百分点。此外,仅添加5%的目标实体数据就能将平均目标实体进展提高13.4个百分点,显示出显著的性能提升。
🎯 应用场景
该研究的潜在应用领域包括机器人操作、智能家居和工业自动化等场景。通过提升机器人在不同实体上的操作能力,能够降低任务特定数据收集的成本,并提高机器人在动态环境中的适应性,具有重要的实际价值和未来影响。
📄 摘要(原文)
Zero-shot generalization to unseen embodiments is important for generalizable vision-language-action (VLA) models as robot hardware evolves and task-specific data collection remains costly. However, a systematic understanding of this problem remains limited, in part because the literature lacks a unified zero-shot transfer definition and controlled evaluation settings that isolate embodiment changes from differences in tasks, scenes, or protocols. To address this gap, we first distinguish strict zero-shot transfer, where the target embodiment is absent from all training data, from pretrain-exposed zero-shot transfer, where it appears only during pretraining. We then introduce a controlled benchmark spanning 14 held-out target embodiments across simulation and real-world validation. Within this framework, we conduct a controlled analysis of four factors: state-action representations, pretraining embodiment diversity, auxiliary co-training objectives, and target-embodiment exposure. Experimental results show that local end-effector (EEF) state-action representations, the source embodiment diversity, and auxiliary co-training improve cross-embodiment transfer by around 15, 18, and 7 percentage points, respectively. We further find that adding only 5% target-embodiment data during pretraining improves average target-embodiment progress by 13.4 percentage points, showing that strict and pretrain-exposed zero-shot transfer are distinct and should be reported separately. Together, these findings provide practical guidance for evaluating and improving cross-embodiment VLA transfer in stationary tabletop manipulation with two-finger grippers, while motivating future investigation of broader settings including mobile-base control, dexterous hands, and long-horizon tasks.