Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration

📄 arXiv: 2608.02285v1 📥 PDF

作者: Aoru Xue, Yujing Sun, Yiming Ren, Kwok-Yan Lam, Mao Ye, Yuexin Ma

分类: cs.CV

发布日期: 2026-08-03

备注: 16 pages, 8 figures, 4 tables. Accepted at ECCV 2026. Aoru Xue and Yujing Sun contributed equally. Yuexin Ma is the corresponding author


💡 一句话要点

提出Sen-Cap以解决多模态传感器对齐及噪声鲁棒性问题

🎯 匹配领域: 支柱七:动作重定向 (Motion Retargeting) 支柱八:物理动画 (Physics-based Animation)

关键词: 人类动作捕捉 多模态传感器 抗噪声技术 实时处理 灵活部署 机器人技术 体育分析

📋 核心要点

  1. 现有多模态传感器方法在对齐和噪声处理上存在显著不足,限制了其在动态环境中的应用。
  2. Sen-Cap通过引入统一的跨传感器运动估计器和抗噪声轨迹跟踪器,解决了传感器间校准和噪声鲁棒性的问题。
  3. Sen-Cap在多个数据集上实现了最先进的性能,显示出其在真实场景中的应用潜力和优势。

📝 摘要(中文)

我们提出了Sen-Cap,一个灵活且抗噪声的3D人类动作捕捉框架,集成了LiDAR和相机的多模态数据。现有方法在多模态传感器的对齐和噪声处理上存在挑战,通常需要显式校准并在噪声或部分传感器故障下性能下降。Sen-Cap引入了统一的跨传感器运动估计器,无需传感器间的校准,支持灵活的传感器数量,并通过迭代优化保持在严重点云噪声下的鲁棒性。Sen-Cap在Human-M3和FreeMotion等主要指标上实现了最先进的性能,并在LiDARHuman26M和RELI11D上展现了强大的跨域性能。这种灵活性和鲁棒性为运动捕捉在体育分析、场地机器人和大规模沉浸式环境等实际应用中开辟了新机会。

🔬 方法详解

问题定义:本论文旨在解决多模态传感器在动态环境中对齐和噪声处理的挑战。现有方法通常依赖显式校准,导致在视角变化时误差传播,并且在噪声或传感器故障下性能显著下降。

核心思路:Sen-Cap的核心思路是通过统一的跨传感器运动估计器来重建人类中心空间中的局部姿态和形状,避免了传感器间的校准问题,同时引入抗噪声轨迹跟踪器以增强鲁棒性。

技术框架:Sen-Cap的整体架构包括两个主要模块:统一的跨传感器运动估计器和噪声抗性轨迹跟踪器。前者负责姿态和形状的重建,后者则在点云噪声下进行迭代优化以保持跟踪精度。

关键创新:Sen-Cap的主要创新在于其无需校准的跨传感器运动估计能力和在高噪声环境下的鲁棒性,这与现有方法的依赖显式校准和对噪声敏感的特性形成鲜明对比。

关键设计:在设计中,Sen-Cap采用了迭代优化策略来处理噪声,并通过灵活的传感器配置支持多种部署方式。具体的损失函数和网络结构细节尚未明确,但其设计目标是最大化在复杂环境中的鲁棒性和灵活性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

Sen-Cap在Human-M3和FreeMotion数据集上实现了最先进的性能,并在LiDARHuman26M和RELI11D上展现了强大的跨域性能,具体指标提升幅度显著,表明其在真实场景中的有效性和可靠性。

🎯 应用场景

Sen-Cap的研究成果具有广泛的应用潜力,尤其是在体育分析、场地机器人和大规模沉浸式环境等领域。其灵活的传感器配置和抗噪声能力使得在复杂和动态的真实场景中进行高效的人类动作捕捉成为可能,推动了相关技术的进步和应用。

📄 摘要(原文)

We propose Sen-Cap, a Sensor-Flexible and Noise-Resilient 3D human motion Capture framework that integrates multi-modal data from LiDAR and camera. While multi-modal sensors provide richer information than single-modal sensors, existing approaches still suffer from two core challenges. First, multi-modal alignment/matching across arbitrarily deployed sensors is typically handled by explicit calibration, which propagates errors under changing viewpoints and in turn constrains deployment to fixed, highly overlapped layouts. Second, prior methods degrade under severe noise or partial sensor failures, which are common in real-world environments. To address these challenges, Sen-Cap introduces a Unified Across-Sensor Motion Estimator that reconstructs local pose and shape in a human-centric space without calibrations between sensors, supporting a flexible number of sensors, as well as a Noise-Resistant Trajectory Tracker that maintains robustness under severe point cloud noise through iterative refinement. These sensor-flexible and noise-resilient features make Sen-Cap more practical in real-world deployment. Notably, operating in real time, Sen-Cap achieves state-of-the-art performance on major metrics on Human-M3 and FreeMotion, as well as strong cross-domain performance on LiDARHuman26M and RELI11D. This combination of flexibility and robustness opens new opportunities for motion capture in real-world scenarios, e.g. sports analytics, field robotics, and large-scale immersive environments.