Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
作者: Shuo Liang, Yixing Ma, Pengfei Zhou, Xingyan Chen, Zihan Mei, Manting Li, Feihan Chen, Zhiwen Wang, Bin Xu, Haotian Zhang, Jiajun Song, Shiya Su, Run Liu, Zhenghang Ni, Yifa Yu, Jintao Hong, Bolong Feng, Yifei Liu, Zirui Zhang, Jingxuan Zhang, Songlin Zhao, Yifan Bai, Kang Tan, Yizhe Liu, Junhao Du, Yongtao Ge, Zhaopan Xv, Xinyuan Zhang, Mengru Ma, Chunhua Shen, Wei Wang, Yang You, Zheng Zhu, Kaipeng Zhang, Wangbo Zhao
分类: cs.CV, cs.AI
发布日期: 2026-08-14
备注: 63 pages, 20 figures, 32 tables
💡 一句话要点
提出RA-Bench以解决AI生成视频检测的挑战
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 视频生成 检测器 社交传播 误信息 基准评估 多模态模型 真实视频锚点
📋 核心要点
- 现有检测器在面对AI生成的视频时表现不佳,尤其是在社交传播过程中,检测准确性下降。
- 本文提出RA-Bench基准,系统评估检测器和生成器的行为,分析生成条件对检测能力的影响。
- 实验结果显示,当前检测器在真实感AI生成视频的检测上普遍存在困难,且社交传播加剧了这一问题。
📝 摘要(中文)
近年来,视频生成技术能够制造出逼真的战争、灾难和公共危机场景,带来了严重的误信息风险。然而,现有基准对检测器和生成器在这些场景中的表现提供的证据有限。为此,本文提出了RA-Bench,一个用于AI生成视频检测的基准,包含17,886个视频,涵盖10个社会风险类别的1,830个真实视频锚点和来自多个生成器的16,056个生成片段。通过RA-Bench,我们评估了七种传统检测器和十种零样本多模态模型的泛化能力,发现当前方法在检测真实感AI生成视频时存在显著困难,强调了对抗不断演变的生成器的检测器的需求。
🔬 方法详解
问题定义:本文旨在解决AI生成视频在真实世界危机事件中的检测问题,现有方法在不同生成条件下的检测能力不足,且社交传播对检测的影响未被充分研究。
核心思路:通过引入RA-Bench基准,提供一个包含真实视频锚点和生成视频的综合数据集,以系统评估检测器的性能和生成器的行为。
技术框架:RA-Bench包含17,886个视频,分为真实视频锚点和生成视频,评估过程包括对七种传统检测器和十种零样本多模态模型的测试,分析生成质量和条件信息对检测的影响。
关键创新:RA-Bench的提出是一个重要创新,它为AI生成视频检测提供了标准化的评估框架,填补了现有基准的空白。
关键设计:在实验中,使用了多种检测器和模型,评估了不同生成条件下的检测能力,特别关注生成质量、条件信息和采样种子对检测结果的影响。实验还考察了人类对视频真实性的判断与检测器的可靠性之间的关系。
🖼️ 关键图片
📊 实验亮点
实验结果表明,当前的检测器在面对真实感AI生成视频时普遍表现不佳,尤其是在社交传播环境中,检测准确性显著下降。不同检测器在不同生成条件下的表现差异明显,强调了对抗新型生成器的检测器的需求。
🎯 应用场景
该研究的潜在应用领域包括社交媒体平台、新闻机构和公共安全部门,能够帮助识别和防范AI生成的误导性视频,提升信息传播的真实性和可靠性。未来,随着生成技术的不断进步,研究成果将为开发更强大的检测工具提供基础。
📄 摘要(原文)
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.