WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

📄 arXiv: 2608.02603v1 📥 PDF

作者: Yuxue Yang, Shuyao Shang, Jiahe Wang, Zitong Zhou, Liang Tan, Junhan Zeng, Ruizhi Li, Junyan Li, Yu Liu, Xiao Yang, Yong Li, Jun Zhu, Hongsheng Li, Tieniu Tan, Lue Fan, Zhaoxiang Zhang

分类: cs.CV

发布日期: 2026-08-03

备注: Project Website: https://WorldExam.github.io


💡 一句话要点

提出WorldExam基准以评估世界模型的内在反应能力

🎯 匹配领域: 支柱二:RL算法与架构 (RL & Architecture)

关键词: 世界模型 视频生成 内在反应能力 基准评估 多任务学习 智能系统 人机交互

📋 核心要点

  1. 现有基准主要关注视觉质量和明确指令的实现,未能充分评估模型的内在反应能力。
  2. 本文提出WorldExam基准,涵盖视觉质量、控制遵循、空间一致性和世界反应性四个层面,提供全面评估。
  3. 评估结果显示,现有模型在不同任务上的表现存在明显差异,且高视觉质量并不保证内在反应能力的强大。

📝 摘要(中文)

随着可控视频生成模型的不断发展,评估其作为世界模型的能力不仅要关注生成视频的表面质量,还需考虑其内在反应能力,即从场景状态推断世界应如何反应并生成合理的后果。然而,现有基准主要评估视觉质量或明确指令的实现,忽视了内在反应能力。为此,本文提出了WorldExam,一个涵盖视觉质量、控制遵循、空间一致性和世界反应性的分层诊断基准,包含1474个案例和八个专门任务,支持对不同模型范式的统一评估。通过对20个代表性模型的评估,发现模型在能力上存在明显差异,且没有模型能够在任务覆盖和性能上实现良好的平衡。

🔬 方法详解

问题定义:本文旨在解决现有可控视频生成模型在内在反应能力评估上的不足,现有方法主要集中于视觉质量和指令实现,忽视了模型如何从场景状态推断反应。

核心思路:提出WorldExam基准,分层评估模型的视觉质量、控制遵循、空间一致性和世界反应性,特别关注模型在未明确指定的情况下的反应能力。

技术框架:WorldExam基准包含四个层面,分别是视觉质量、控制遵循、空间一致性和世界反应性。每个层面下设有多个任务,共1474个案例,支持对不同类型模型的统一评估。

关键创新:WorldExam的创新在于其分层结构和对内在反应性的专门评估,填补了现有基准的空白,使得评估更加全面和深入。

关键设计:在设计中,采用了多种评估指标和任务设置,以确保对模型在不同场景下的反应能力进行全面评估,具体参数和损失函数的选择旨在提升评估的准确性和可靠性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,20个代表性模型在不同层面的表现存在明显差异。相机驱动模型在相机控制上表现优异,但缺乏动态交互能力;动作驱动模型在控制精度上较好,但世界反应不足;语言驱动模型在交互上表现较强,但复杂控制的遵循性较差。没有模型能够在任务覆盖和性能上实现全面的优势。

🎯 应用场景

该研究的潜在应用领域包括智能机器人、自动驾驶、虚拟现实等,能够帮助开发更具反应能力的智能系统。通过提升模型的内在反应能力,可以实现更自然的交互和更复杂的任务执行,未来可能对人机交互和自动化领域产生深远影响。

📄 摘要(原文)

Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.