RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement
作者: Ziheng Jia, Jiaying Qian, Zicheng Zhang, Xiaorong Zhu, Lancheng Gao, Xiongkuo Min
分类: cs.AI
发布日期: 2026-08-10
💡 一句话要点
提出RAVEN-Eval以解决AI视频生成模型评估难题
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: AI视频生成 自动化评估 LMM偏好判断 任务策划 质量过滤 评估框架 模型比较
📋 核心要点
- 现有评估方法难以有效区分先进AI视频生成模型之间的细微质量差异,且人类评估成本高。
- RAVEN-Eval采用LMM作为评判者,通过任务特定的评分标准进行自动化的偏好判断,降低评估成本。
- 通过对20个高性能AI视频生成模型和13个LMM评判者的评估,建立了RAVEN-Eval排行榜,展示了其有效性。
📝 摘要(中文)
随着AI视频生成技术的快速发展,市场上各种生成模型的质量差异变得难以通过传统评估标准 discern。人类评估需要更多的专业知识和持续的注意力,导致标注成本显著增加。因此,急需一种自动化评估方法,能够在较少人力干预的情况下,可靠地区分先进AI视频生成模型之间的细微差异。为此,本文提出了RAVEN-Eval,一个基于LMM偏好判断的评估框架,通过自动任务策划和质量过滤流程,系统性地收集了超过4500个AI生成视频,并建立了RAVEN-Eval排行榜。
🔬 方法详解
问题定义:本文旨在解决AI视频生成模型(AIVGMs)评估中的质量差异难以 discern 的问题,现有方法依赖于人工评估,成本高且效率低下。
核心思路:RAVEN-Eval框架通过LMM偏好判断,结合评分标准进行自动化评估,旨在减少人力干预并提高评估的准确性和效率。
技术框架:RAVEN-Eval的整体架构包括自动任务策划、质量过滤和LMM偏好判断三个主要模块。首先,系统策划150个文本到视频(T2V)任务和100个图像到视频(I2V)任务,然后收集并评估生成的视频。
关键创新:RAVEN-Eval的创新在于引入了基于评分标准的LMM偏好判断和锚点模型插入方法,显著降低了新模型的评估成本,与传统方法相比具有更高的灵活性和可扩展性。
关键设计:在评估过程中,LMM通过成对比较的方式进行判断,依据特定的任务评分标准进行评估,确保评估结果的可靠性和一致性。
🖼️ 关键图片
📊 实验亮点
在实验中,RAVEN-Eval成功评估了20个高性能AIVGMs,并通过13个LMM评判者的判断能力建立了排行榜。该框架的引入使得评估效率显著提升,能够在较短时间内处理大量生成视频,展示了其在实际应用中的潜力。
🎯 应用场景
RAVEN-Eval框架可广泛应用于AI视频生成领域,尤其是在需要快速评估和比较不同生成模型的场景中。其自动化评估能力不仅降低了人力成本,还提高了评估的准确性,未来可能推动更多AI生成技术的商业化应用。
📄 摘要(原文)
AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by state-of-the-art AI video generation models~(AIVGMs) have become increasingly difficult to discern using conventional evaluation criteria, such as visual fidelity and semantic instruction following. Meanwhile, human evaluation now requires more expertise and sustained attention, substantially increasing annotation costs. This calls for automated evaluation that can reliably distinguish fine-grained differences among advanced AIVGMs with minimal human intervention. To address this challenge, we present RAVEN-Eval, a rubric-guided automated evaluation framework for AIVGMs, built primarily on the LMM-as-a-judge paradigm. Through an automatic task curation and quality-filtering pipeline, RAVEN-Eval curates 150 text-to-video~(T2V) tasks and 100 image-to-video~(I2V) tasks, and systematically collects more than 4,500 AIGVs. At its core, RAVEN-Eval adopts rubric-guided automated LMM preference judgement, in which LMM judges conduct pairwise comparisons according to task-specific rubrics. It further introduces an anchor-based model insertion approach to reduce the evaluation cost of incorporating new models. Finally, we evaluate 20 high-performance AIVGMs, as well as the judging capabilities of 13 LMM judges, and establish the RAVEN-Eval Leaderboards. Overall, RAVEN-Eval paves a scalable path for automatic and trustworthy evaluation of rapidly evolving AIVGMs.