MS-MEM: Multi-Skill Manipulation-Enhanced Mapping via Uncertainty- and Disturbance-Aware Action Selection

📄 arXiv: 2609.02493v1 📥 PDF

作者: Yitian Shi, Jesper Mücke, Nils Dengler, Sicong Pan, Rania Rayyes, Maren Bennewitz

分类: cs.RO

发布日期: 2026-09-02

备注: under review


💡 一句话要点

提出MS-MEM以解决服务机器人在复杂环境中的物体定位问题

🎯 匹配领域: 支柱一:机器人控制 (Robot Control) 支柱三:空间感知与语义 (Perception & Semantics)

关键词: 服务机器人 不确定性感知 主动视角选择 物体抓取 映射技术 多技能操作 场景理解

📋 核心要点

  1. 现有方法在复杂环境中物体定位面临严重遮挡和可达性限制,导致映射准确性不足。
  2. MS-MEM框架通过集成主动视角选择、物体推送和抓取,采用不确定性感知的映射策略,提升场景理解能力。
  3. 实验结果显示,MS-MEM在映射准确性上显著优于单技能和无约束基线,同时有效减少场景干扰。

📝 摘要(中文)

在狭窄且杂乱的空间中,准确的场景理解对服务机器人至关重要,因为许多日常任务需要它们可靠地定位和获取物体。然而,由于严重的遮挡、受限的可达性以及避免过度场景变化的需求,这一任务仍然具有挑战性。本文提出了一种多技能操作增强映射(MS-MEM)框架,该框架集成了主动视角选择、物体推送和抓取的能力,旨在实现不确定性感知的映射。MS-MEM结合了场景级的度量-语义证据信念估计器与不确定性感知的抓取表示,后者通过一种新颖的全证据抓取估计器进行学习,该估计器同时建模抓取适应性和方向不确定性。实验结果表明,与忽略场景干扰的单技能和无约束基线相比,MS-MEM在提高映射准确性和显著减少场景干扰方面表现优异,突显了主动视角选择、推送和抓取动作的协同效应。

🔬 方法详解

问题定义:本文旨在解决服务机器人在复杂、杂乱环境中物体定位的准确性问题。现有方法往往忽视了场景的干扰因素,导致映射效果不佳。

核心思路:MS-MEM框架通过结合主动视角选择、物体推送和抓取,采用不确定性感知的映射策略,旨在有效减少映射不确定性,同时限制对场景的干扰。

技术框架:MS-MEM的整体架构包括三个主要模块:场景级度量-语义证据信念估计器、全证据抓取估计器和统一的动作选择管道。通过这些模块,系统能够评估感知和操作动作的有效性。

关键创新:最重要的创新在于引入了不确定性感知的抓取表示和协同干扰约束(CDC),使得系统能够在选择操作时考虑场景的信心区域,避免过度干扰。

关键设计:在设计中,抓取表示通过全证据抓取估计器进行学习,模型同时考虑抓取适应性和方向不确定性。此外,统一的动作选择管道使用共同的信息增益标准来评估候选动作。该设计确保了操作的有效性与场景信念的稳定性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果表明,MS-MEM在映射准确性上相比于单技能和无约束基线提高了显著的性能,具体表现为映射准确性提升了XX%,同时场景干扰减少了YY%,验证了主动视角选择、推送和抓取动作的协同效应。

🎯 应用场景

该研究的潜在应用领域包括服务机器人在家庭、仓储和医疗等环境中的物体识别与抓取。通过提高机器人在复杂环境中的操作能力,MS-MEM能够显著提升服务机器人的实用性和可靠性,未来可能推动智能家居和自动化物流的发展。

📄 摘要(原文)

Accurate scene understanding in confined, cluttered spaces such as shelves is essential for service robots, as many everyday tasks require them to locate and retrieve objects reliably. Yet, it remains challenging due to severe occlusions, restricted accessibility, and the need to avoid excessive scene changes. In this paper, we propose Multi-Skill Manipulation-Enhanced Mapping (MS-MEM), an evidential framework for uncertainty-aware mapping that integrates active viewpoint selection, object pushing, and grasping. MS-MEM combines scene-level metric-semantic evidential belief estimators with an uncertainty-aware grasp representation. This representation is learned using a novel full-evidential grasp estimator that models both grasp affordance and orientation uncertainty. In our framework, candidate perception and manipulation actions are evaluated within a unified action selection pipeline using a common information gain criterion. For manipulation actions, we further introduce a collateral disturbance constraint (CDC) that discourages excessive changes to confident regions of the scene belief. This enables MS-MEM to select actions that effectively reduce map uncertainty while limiting collateral scene changes. Experimental results show that, compared with single-skill and unconstrained baselines that ignore scene disturbance, MS-MEM achieves higher mapping accuracy while substantially reducing scene disturbance, highlighting the synergistic effects of active viewpoint selection, push, and grasp actions.