MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos

📄 arXiv: 2607.27895v1 📥 PDF

作者: Jinpeng Hu, Erqiang Wang, Shan Wang, Zhuo Li, Peipei Song, Xun Yang, Meng Wang

分类: cs.AI, cs.CV

发布日期: 2026-07-30


💡 一句话要点

提出MMHBench以解决长视频心理健康理解问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 长视频理解 心理健康 多模态基准 问题生成 深度学习 社会角色模拟 多视角推理

📋 核心要点

  1. 现有基准方法在心理健康理解任务中多采用粗粒度分类,无法深入评估模型对心理现象的真实理解能力。
  2. 本文提出MMHBench基准,结合多模态数据和多视角推理,设计了多智能体问题生成框架以提升问题质量。
  3. 对22个多模态大语言模型的评估结果显示,长视频心理健康理解任务依然面临重大挑战,模型表现不尽如人意。

📝 摘要(中文)

心理健康理解在长视频中需要对可观察行为、人际背景和潜在心理状态进行细致推理。现有基准大多将此任务简化为粗粒度分类,限制了模型对心理现象的真实理解。为了解决这一局限性,本文提出了MMHBench,这是一个全面的多模态基准,包含268个长视频和2184个精心策划的问题。MMHBench将评估分为两个互补设置:第三人称评估和第一人称视角推理。我们还提出了多智能体问题生成框架(MAQG),模拟多种社会角色以从多个视角生成问题,并通过多角色反馈和专家验证确保问题的质量和有效性。对22个代表性多模态大语言模型的广泛评估表明,长视频心理健康理解仍然具有高度挑战性。

🔬 方法详解

问题定义:本文旨在解决长视频中的心理健康理解问题,现有方法在此任务中多采用粗粒度分类,缺乏对心理现象的深入理解。

核心思路:提出MMHBench基准,通过多模态数据和多视角推理,设计多智能体问题生成框架(MAQG),以生成高质量的问题,促进心理健康理解。

技术框架:整体架构包括视频数据收集、问题生成、反馈优化和专家验证四个主要模块。首先收集长视频数据,然后通过MAQG框架生成问题,接着进行多角色反馈优化,最后通过专家验证确保问题的有效性。

关键创新:最重要的创新点在于引入多智能体问题生成框架,能够模拟多种社会角色,从而生成多样化的问题,提升了评估的深度和广度。

关键设计:在问题生成过程中,设置了多角色反馈机制和迭代优化流程,确保生成的问题既具挑战性又符合心理健康理解的需求。

🖼️ 关键图片

img_0
img_1
img_2

📊 实验亮点

实验结果表明,22个多模态大语言模型在长视频心理健康理解任务中的表现仍然不理想,整体准确率未达到预期,显示出该领域的研究仍需进一步深入,尤其是在模型的推理能力和理解深度方面。

🎯 应用场景

该研究的潜在应用领域包括心理健康评估、教育培训和社交媒体内容分析等。通过深入理解长视频中的心理健康信息,能够为心理健康干预和支持提供数据驱动的依据,提升相关领域的研究和实践效果。

📄 摘要(原文)

Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological states. Existing benchmarks largely reduce this task to coarse-grained classification, providing limited insight into whether models truly understand psychological phenomena or rely on superficial correlations. To address this limitation, we introduce MMHBench, a comprehensive multimodal benchmark for multi-perspective mental health understanding, comprising 268 long-form videos and 2,184 carefully curated questions. MMHBench organizes the evaluation into two complementary settings: (1) third-person assessment, consisting of 605 questions that focus on the interpretation of observable behaviors and multimodal evidence, and (2) first-person perspective-taking, comprising 1,579 questions that require perspective-conditioned reasoning to identify the interpretation of the mental state supported by the available multimodal evidence. We propose a Multi-Agent Question Generation (MAQG) framework that simulates diverse social roles to synthesize questions from multiple perspectives. The generated questions are refined through multi-role feedback and iterative optimization, followed by expert-guided verification to ensure quality and validity. Extensive evaluation of 22 representative multimodal large language models (MLLMs), spanning both open-source and leading closed-source models, demonstrates that long-form video mental health understanding remains highly challenging.