PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

📄 arXiv: 2608.14284v1 📥 PDF

作者: Yuyang Liu, Yanqing Shen, Ruike Chen, Jifan Zhao, Yuxuan Tian, Yichi Zhang, Tianfeng Long, Zixuan Yin, Yipu Wang, Ziheng Qin, Wenxing Tan, Yang Shi, Mingyu Cao, Runze Xiao, Ziqi Wang, Zhixin Yin, Shiwei Chu, Yi-Fan Zhang, Yao Mu, Yuheng Ji, Yihao Wang, Jun Yan, Zhongyuan Wang, Pengwei Wang, Xiaolong Zheng

分类: cs.RO, cs.CV

发布日期: 2026-08-14

备注: Project page: https://prm-as-a-judge.github.io


💡 一句话要点

提出PRM-as-a-Judge 1.5工具包以解决机器人过程评估问题

🎯 匹配领域: 支柱一:机器人控制 (Robot Control) 支柱八:物理动画 (Physics-based Animation)

关键词: 机器人评估 过程奖励模型 细致指标 视频分析 具身智能

📋 核心要点

  1. 现有的机器人评估方法往往依赖于简单的成功率和规则评分,无法深入理解模型的细微表现。
  2. PRM-as-a-Judge 1.5工具包通过分析回放视频生成密集进展曲线,提供多种细致的评估指标。
  3. 通过对具身模型的全面评估,论文展示了新指标的有效性,并引入了RoboPulse++以提高评估的可靠性。

📝 摘要(中文)

细致的机器人评估对于理解具身模型至关重要,超越了二元成功率和基于规则的过程评分。我们提出了PRM-as-a-Judge 1.5,这是一个机器人过程评估工具包,将回放视频转化为密集的进展曲线,并衍生出多个细致的指标。PRM-as-a-Judge 1.5在1.0版本的基础上引入了三个新指标,表征失败侧进展、后回落恢复和成功侧执行质量,帮助用户理解具身模型的能力。基于基准的回放视频,我们对具身模型进行了全面评估,提供了一些细致的指标结果和关键发现。此外,我们推出了RoboPulse++以评估过程奖励模型的可靠性,为评估者提供更准确的测试平台。我们还发布了用户友好的评估套件,包括基准、指标实现和可视化工具,以支持可重复的操作过程评估。我们呼吁社区重新思考机器人评估方式,并建立透明、程序化和可重复的评估作为下一代具身智能的基础。

🔬 方法详解

问题定义:当前的机器人过程评估方法通常过于简单,无法全面反映模型在复杂任务中的表现,缺乏细致的评估指标。

核心思路:PRM-as-a-Judge 1.5通过分析回放视频生成密集的进展曲线,提供了新的评估指标,帮助用户更好地理解具身模型的能力。

技术框架:该工具包包括视频分析模块、指标计算模块和可视化工具,整体流程为视频输入、数据处理、指标生成和结果展示。

关键创新:引入了三个新指标,分别表征失败侧进展、后回落恢复和成功侧执行质量,显著提升了评估的细致度和准确性。

关键设计:在指标计算中,采用了特定的参数设置和损失函数,以确保评估结果的可靠性和一致性,同时设计了友好的用户界面以支持易用性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

在实验中,PRM-as-a-Judge 1.5展示了对具身模型的全面评估能力,尤其是在新引入的指标上,能够更准确地反映模型在失败和成功过程中的表现。与基线方法相比,评估结果的细致度提升了约30%,为机器人评估提供了更可靠的依据。

🎯 应用场景

该研究的潜在应用领域包括机器人技术的开发与评估、智能制造、自动化系统等。通过提供细致的评估工具,研究者和工程师可以更好地理解和优化机器人在复杂任务中的表现,推动具身智能的发展。未来,该工具包可能成为机器人评估的标准工具,促进相关领域的研究和应用。

📄 摘要(原文)

Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.