StableMimic: Smooth Human-Like Recovery for Humanoid Motion Tracking - Learning Beyond the Tracking Distribution for Structured Post-Fall Behavior
作者: Weihao Wu, Ming Huang, Ruofei Liu, Jinglei Nie, Shuxiang Guo, Chunying Li
分类: cs.RO
发布日期: 2026-08-03
备注: 8 pages, 7 figures. Preprint, not formally peer-reviewed
💡 一句话要点
提出StableMimic以解决类人机器人跌倒后的恢复问题
🎯 匹配领域: 支柱一:机器人控制 (Robot Control) 支柱八:物理动画 (Physics-based Animation)
关键词: 类人机器人 运动追踪 跌倒恢复 自主恢复 安全性提升
📋 核心要点
- 现有的类人机器人运动追踪方法在跌倒后难以恢复,容易导致不稳定的肢体动作和安全隐患。
- StableMimic通过训练超越标准追踪分布,采用多种人类起身参考,形成结构化的恢复策略。
- 在100次推倒实验中,StableMimic实现了100%的恢复成功率,并在多项后续运动和负载指标上表现最佳。
📝 摘要(中文)
类人运动追踪器在学习的追踪分布内表现可靠,但跌倒可能使机器人进入低高度、接触丰富的状态,导致无法及时执行前进指令。仅依赖追踪的策略可能会追逐不可行的参考,产生快速且幅度大的肢体修正,增加机器人及其周围环境的风险。本文提出StableMimic,一个超越标准追踪分布的统一追踪器,通过对多个人类起身参考的扰动重置,塑造结构化的恢复过程,使机器人返回可追踪区域。StableMimic使用专门的专家处理不同的状态-动作分布,并通过本体感知门持续融合其动作。隐藏的后继状态目标教会人类参考形状的恢复,而无需暴露参考身份或阶段,部署时无需起身参考、恢复指令、轨迹检索或外部策略切换。在完整的重定向LAFAN1舞蹈子集上,StableMimic在五种方法中实现了最低的追踪误差。
🔬 方法详解
问题定义:本文旨在解决类人机器人在跌倒后恢复能力不足的问题。现有方法在跌倒后容易导致机器人进入不可追踪的状态,造成安全隐患和不稳定的肢体动作。
核心思路:StableMimic的核心思路是通过扰动重置训练超越标准追踪分布,利用多个人类起身参考来塑造结构化的恢复过程,从而使机器人能够有效返回可追踪区域。
技术框架:StableMimic的整体架构包括多个模块:专门的专家用于处理不同的状态-动作分布,以及一个本体感知门用于融合不同专家的动作。通过隐藏的后继状态目标,系统能够在不暴露参考身份或阶段的情况下进行恢复。
关键创新:StableMimic的主要创新在于其训练方式和恢复策略的设计,能够在不依赖具体的起身参考或外部策略切换的情况下,实现高效的自主恢复,这与现有方法的依赖性形成鲜明对比。
关键设计:在技术细节上,StableMimic采用了专门设计的损失函数和网络结构,以确保在不同状态下的动作融合效果最佳,同时优化了参数设置以提高恢复效率。
🖼️ 关键图片
📊 实验亮点
在实验中,StableMimic在100次推倒实验中实现了100%的恢复成功率,并在六项后续运动和负载指标上表现出最低值,显示出其在安全性和稳定性方面的显著提升。
🎯 应用场景
该研究的潜在应用领域包括服务机器人、救援机器人和娱乐机器人等。通过提升机器人在跌倒后的恢复能力,能够显著提高其在复杂环境中的交互安全性和自主性,未来可能在家庭、医疗和公共服务等多个领域产生深远影响。
📄 摘要(原文)
Humanoid motion trackers perform reliably within learned tracking distributions, but falls can move the robot into low-height, contact-rich states from which an advancing command is temporarily unreachable. Tracking-only policies may chase infeasible references, producing rapid, large-amplitude limb corrections that increase risk to the robot and its surroundings. We present StableMimic, a unified tracker trained beyond the nominal tracking distribution. Perturbed resets around multiple human get-up references expose prone, supine, off-balance, and intermediate ground-contact states, shaping structured recovery that returns the robot to the trackable region. Because tracking and recovery occupy markedly different state--action distributions, StableMimic uses dedicated experts for each regime and a proprioceptive gate that continuously blends their actions. A hidden successor-state objective teaches human-reference-shaped recovery without exposing reference identity or phase to the deployed Actor; deployment requires no get-up reference, recovery command, trajectory retrieval, or external policy switch. On the complete retargeted LAFAN1 dance subset, StableMimic achieves the lowest errors on all four tracking metrics among five methods. Across 100 matched push-to-fall trials per method, it recovers in 100/100 and attains the lowest values on six of seven post-fall motion and load measures, supporting improved interaction safety under this protocol. Real Unitree G1 dance and standing-reference deployments qualitatively demonstrate bounded limb motion, autonomous recovery, and command resumption.