CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

📄 arXiv: 2608.01942v1 📥 PDF

作者: Xianjing Han, Yuhan Su, Yang Deng, Dong Ma, Wee Peng Tay, Bin Zhu

分类: cs.CV, cs.CL, cs.MM

发布日期: 2026-08-03

备注: Project page:https://hanxjing.github.io/CultureVidBench/


💡 一句话要点

提出CultureVidBench以解决文本到视频生成中的文化理解问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 文化理解 文本到视频生成 多模态评估 文化基准 动态表现

📋 核心要点

  1. 现有的文本到视频生成模型在捕捉文化特定细节方面存在不足,尤其是对多样文化背景的理解和表现不够全面。
  2. 论文提出了CultureVidBench基准,专门用于评估T2V生成中的文化理解,涵盖多个国家和文化方面,强调动态和多模态表现。
  3. 通过对七个T2V模型的评估,发现尽管模型在语义一致性和视觉质量上表现良好,但在细致的文化细节捕捉上仍有待提升。

📝 摘要(中文)

文本到视频生成(T2V)模型发展迅速,但其在表现多样文化背景方面的能力仍未得到充分探索。现有基准主要关注感知质量、物理合理性和文本视频对齐,而未直接评估生成视频是否捕捉到文化特定的物体、动作、仪式、可见文本或音频线索。我们引入CultureVidBench,这是一个综合基准,用于评估T2V生成中的文化理解。CultureVidBench包含1000个经过精心策划的提示,涵盖12个国家、6个大洲、8个文化区域和14个文化方面,分为物质文化、社会实践与表演、仪式与典礼三大类。该基准强调动态和多模态的文化表现,评估七个代表性T2V模型在文化忠实性、多模态文化呈现、语义一致性和感知质量方面的表现。结果显示,尽管当前模型在语义一致性和视觉质量上表现良好,但在捕捉细致的文化细节方面,尤其是对于代表性不足的地区、仪式和多模态文化线索时,仍存在不足。

🔬 方法详解

问题定义:论文要解决的问题是现有T2V生成模型在文化理解方面的不足,尤其是对文化特定对象、动作和仪式的表现不够准确。现有方法主要关注视觉质量和文本对齐,而忽视了文化细节的捕捉。

核心思路:论文的核心解决思路是引入CultureVidBench基准,通过系统评估T2V模型在文化理解方面的表现,强调动态和多模态的文化表现形式,以更全面地评估生成视频的文化忠实性。

技术框架:CultureVidBench的整体架构包括三个主要模块:物质文化、社会实践与表演、仪式与典礼。每个模块下又细分为多个文化方面,涵盖了丰富的文化提示,旨在评估生成视频的多样性和准确性。

关键创新:最重要的技术创新点在于创建了一个专注于文化理解的基准,填补了现有T2V生成评估中的空白,特别是在多模态文化表现的评估上,与传统的视觉质量评估方法有本质区别。

关键设计:在设计上,CultureVidBench包含1000个提示,覆盖12个国家和14个文化方面,采用了人类用户研究和基于MLLM的自动评估方法,确保评估的全面性和准确性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,尽管当前T2V模型在语义一致性和视觉质量上表现良好,但在捕捉细致的文化细节方面存在明显不足,尤其是在代表性不足的地区和仪式方面,提示了未来改进的方向。

🎯 应用场景

该研究的潜在应用领域包括文化教育、影视制作和虚拟现实等。通过提升T2V生成模型的文化理解能力,能够更好地满足多样化文化内容的需求,促进跨文化交流与理解,具有重要的实际价值和未来影响。

📄 摘要(原文)

Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, and text-video alignment, but do not directly assess whether generated videos capture culturally specific objects, actions, rituals, visible text, or audio cues. We introduce CultureVidBench, a comprehensive benchmark for evaluating cultural understanding in T2V generation. CultureVidBench contains 1,000 curated prompts covering 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects organized into three categories: material culture, social practice & performance, and ritual & ceremony. Designed specifically for video generation, CultureVidBench emphasizes dynamic and multimodal cultural representation, including social interactions, ritual procedure, and culturally appropriate visible text and audio. We evaluate seven representative T2V models through human user studies and MLLM-based automatic assessment across cultural faithfulness, multimodal cultural rendering, semantic adherence, and perceptual quality. Results show that although current models achieve strong semantic adherence and visual quality, they often fail to faithfully capture fine-grained cultural details, particularly for underrepresented regions, rituals, and multimodal cultural cues.