PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation

📄 arXiv: 2608.16717v1 📥 PDF

作者: Yuji Wang, Yuheng Chen, Teng Hu, Ran Yi, Yijia Hong, Han Feng, Weijian Cao, Chengjie Wang, Lizhuang Ma, Jiangning Zhang

分类: cs.CV

发布日期: 2026-08-17


💡 一句话要点

提出PersonaShot以解决多镜头视频生成中的叙事连续性问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 多镜头视频生成 叙事连续性 物理连续性 情感动态 电影语法 基准评估 人类对齐评估者

📋 核心要点

  1. 现有视频生成基准主要关注角色外观和单镜头质量,缺乏对镜头间物理和情感状态连贯性的评估。
  2. 论文提出PersonaShot基准,通过引入16个指标,系统评估多镜头视频中的叙事连续性。
  3. 实验结果显示,当前模型在视觉质量与叙事连续性之间存在明显差距,且常出现物理状态重置和情感突变。

📝 摘要(中文)

视频生成正迅速从单镜头片段发展到多镜头叙事,其中人类角色作为核心叙事锚点。然而,现有基准主要评估角色外观或单镜头质量,未能衡量物理和情感状态在镜头间的连贯性。为了解决这些局限性,我们提出了PersonaShot,这是首个针对多镜头视频生成叙事连续性的人物中心基准。PersonaShot包含约1000个多镜头片段和16个涵盖物理连续性、情感动态和电影语法的指标。

🔬 方法详解

问题定义:本论文旨在解决多镜头视频生成中叙事连续性评估的不足,现有方法未能有效衡量镜头间的物理和情感状态连贯性。

核心思路:我们提出PersonaShot基准,系统性地评估多镜头视频中的叙事连续性,涵盖物理连续性、情感动态和电影语法等多个维度。

技术框架:PersonaShot包含约1000个多镜头片段,评估分为三个时间层次:镜头内状态、镜头间过渡和序列级轨迹。我们还引入了人类对齐的专业评估者,基于视觉、时间或关系证据进行评估。

关键创新:最重要的创新在于引入了针对叙事连续性的多维度评估指标,填补了现有基准的空白,并通过人类专家的判断进行对齐。

关键设计:我们设计了16个指标,分别针对物理连续性、情感动态和电影语法,确保评估的全面性和准确性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,当前最先进的模型在叙事连续性方面存在明显不足,尽管视觉效果出色,但仍频繁出现物理状态重置和情感突变。我们的评估者与专家判断之间达成了高度一致,验证了评估方法的有效性。

🎯 应用场景

该研究的潜在应用领域包括电影制作、游戏开发和虚拟现实等,能够帮助创作者生成更具连贯性和情感深度的多镜头视频内容。未来,该基准可能推动视频生成技术的进一步发展,提升人机交互体验。

📄 摘要(原文)

Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks mainly assess character appearance or individual-shot quality, without measuring whether physical and emotional states remain coherent across cuts. They also rarely provide criterion-specific evaluation methods, although physical continuity, facial dynamics, and cinematic relations require different visual, temporal, and relational evidence. To address these limitations, we introduce PersonaShot, the first person-centric benchmark for narrative continuity in multi-shot video generation. PersonaShot contains approximately 1,000 multi-shot segments and 16 metrics spanning physical continuity, affective dynamics, and cinematic grammar. \textbf{\textit{1)} Narrative Continuity Benchmark:} We evaluate character coherence across three temporal levels: within-shot states, cross-shot transitions, and sequence-level trajectories. \textbf{\textit{2)} Human-Aligned Specialist Evaluators:} We distill reasoning from a large multimodal teacher into lightweight criterion-specific evaluators, each grounded in the visual, temporal, or relational evidence required by its metric, and align them with expert human judgments. \textbf{\textit{3)} Systematic Evaluation and Insights:} Our evaluation reveals distinct capability profiles across state-of-the-art models and a clear gap between perceptual quality and cross-shot narrative continuity. Even visually compelling videos frequently exhibit physical-state resets, abrupt affective shifts, and broken cinematic relations across shots. Human studies further demonstrate strong agreement between our evaluators and expert judgments.