PAMoR: Parameterized Affective Motion Generation in Real Time for Humanoid Robots
作者: Yan Pan, Lingfan Bao, Tianhu Peng, Chengxu Zhou
分类: cs.RO
发布日期: 2026-08-28
备注: Under Review of IEEE Robotics and Automation Letters, 8 pages
💡 一句话要点
提出PAMoR以实时生成类人机器人情感运动
🎯 匹配领域: 支柱一:机器人控制 (Robot Control) 支柱四:生成式动作 (Generative Motion) 支柱七:动作重定向 (Motion Retargeting) 支柱八:物理动画 (Physics-based Animation)
关键词: 类人机器人 情感运动 实时生成 效价-唤醒 动作先验 情感先验 人机交互 机器人技术
📋 核心要点
- 现有方法在生成类人机器人情感运动时,缺乏定量参数化,限制了其在社交场合的应用效果。
- 本文提出PAMoR,通过将情感转化为基于机器人运动学的V-A坐标,实现实时的情感运动生成。
- 实验结果表明,生成的运动在情感识别上优于基线,命令情感的识别率达到0.38,接近人类表演的0.44。
📝 摘要(中文)
人们在社交场合中解读类人机器人的运动时,不仅关注其执行的动作,还关注所传达的情感。以往生成情感运动的方法主要针对人类虚拟形象,缺乏定量参数化的能力。本文提出PAMoR,将情感转化为可测量的控制参数:基于机器人运动学计算的效价-唤醒(V-A)坐标。该坐标通过姿态扩展和运动能量的闭式形式获得,直接作为生成条件,无需人工标注。通过在共享潜在空间中训练的动作先验和两个情感先验,在每个去噪步骤中组合生成运动,最终实现实时的全身运动生成,且动作和情感均可编辑。生成的运动能够全范围跟踪命令的V-A,同时文本到运动的保真度与仅基于文本的基线相匹配。
🔬 方法详解
问题定义:本文旨在解决类人机器人在社交场合中情感运动生成的不足,现有方法无法定量参数化情感,影响了机器人的社交表现。
核心思路:PAMoR通过将情感转化为效价-唤醒(V-A)坐标,利用机器人运动学进行实时运动生成,避免了人工标注的需求。
技术框架:整体架构包括动作先验和情感先验的组合,生成过程在每个去噪步骤中进行,最终实现29自由度的Unitree G1机器人全身运动的自回归生成。
关键创新:最重要的创新在于将情感定量化为V-A坐标,并通过姿态扩展和运动能量的闭式计算直接作为生成条件,显著提升了生成运动的情感表达能力。
关键设计:在模型设计中,采用共享潜在空间训练动作和情感先验,确保生成运动的动作和情感可编辑,且生成的运动能够准确跟踪命令的V-A坐标。
🖼️ 关键图片
📊 实验亮点
实验结果显示,生成的运动在情感识别上表现优异,命令情感的识别率达到0.38,超过了基线并接近人类表演的0.44,表明PAMoR在情感运动生成方面的有效性和创新性。
🎯 应用场景
该研究具有广泛的应用潜力,尤其在社交机器人、虚拟助手和娱乐领域。通过提升类人机器人在情感表达上的能力,可以增强人机交互的自然性和有效性,推动智能机器人在日常生活中的应用。未来,该技术可能会影响教育、医疗和客户服务等多个行业。
📄 摘要(原文)
People read a humanoid robot's motion in social settings not only for the action performed but for the affect conveyed. Motion carrying that affect has so far been generated for human avatars, where style is taken from a reference clip or an emotion word, neither of which can be quantitatively parameterized. We present PAMoR, which turns affect into a measured control parameter: a valence-arousal (V-A) coordinate computed natively on robot kinematics. It is obtained in closed form from postural expansion and movement energy, and these measurements serve directly as generation conditions, with no human annotation. An action prior and two affect priors, trained in a shared latent space, are composed at each denoising step: the action prior fixes what is performed, the affect priors modulate how. Whole-body motion rolls out autoregressively on a 29-DoF Unitree G1 in real time, with action and affect both editable. Generated motion tracks the commanded V-A over its full range while text-to-motion fidelity still matches text-only baselines. In a perceptual study, raters identify the commanded emotion on 0.38 of trials, above both baselines and approaching the 0.44 reported for acted human bodies.