Predictors of Loneliness in Older Adults Using Multimodal Analysis of Speech and Language
作者: Vinmay Khandode, Sai Karthik Kosuri, Neil K. R. Sehgal, Adam Greene, Elif Alpoge, Elana Duffy, Matthew Lee Smith, Thomas K. M. Cudjoe, Sharath Chandra Guntuku
分类: cs.CL
发布日期: 2026-09-02
💡 一句话要点
提出多模态分析方法以识别老年人孤独感
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 孤独感识别 多模态分析 老年人心理健康 语言特征 声学特征 心理评估 自然对话
📋 核心要点
- 现有方法在自然对话中检测老年人孤独感的能力有限,缺乏客观、可扩展的检测手段。
- 本文提出了一种多模态框架,结合语言特征和声学特征,分析老年人的孤独感表现。
- 实验结果显示,多模态模型在识别孤独感方面的相关性达到0.298,明显优于单一模型。
📝 摘要(中文)
孤独感是老年人面临的重大公共健康问题,与抑郁、认知衰退和死亡风险增加相关。现有的检测方法在自然对话环境中仍然有限。本文分析了310名老年人的语音和语言特征,利用半结构化电话访谈,探讨孤独感的语言表现及其差异。研究结合了语言特征和声学特征,发现孤独感与否定词、负面语调和冲突相关语言呈正相关,而与社交参考、动机驱动和情感丰富度呈负相关。多模态模型的表现优于单一文本或音频模型,表明语音分析在心理评估中的潜力。
🔬 方法详解
问题定义:本文旨在解决老年人孤独感的检测问题,现有方法在自然对话中缺乏有效性和客观性。
核心思路:通过结合语言和声学特征,分析老年人孤独感的表现,旨在提供更全面的评估工具。
技术框架:研究采用多模态分析框架,包含语言特征(心理语言学词典、n-gram、主题模型)和声学特征(音调、语调、响度),通过半结构化访谈收集数据。
关键创新:最重要的创新在于将语言和声学特征结合,形成多模态模型,显著提高了孤独感的识别能力。
关键设计:使用了预定义和数据驱动的方法来捕捉语言内容和声音表现的模式,关键参数包括声学特征的提取和语言特征的选择。
🖼️ 关键图片
📊 实验亮点
实验结果表明,多模态模型在孤独感识别中的相关性达到0.298,显著优于文本模型和音频模型,后者的相关性分别为0.11和0.12。这一提升表明多模态分析在心理评估中的重要性和有效性。
🎯 应用场景
该研究的潜在应用领域包括老年人心理健康评估、社交服务和医疗干预等。通过早期识别孤独感,可以帮助制定个性化的干预措施,改善老年人的生活质量。未来,该方法有望与现有评估工具结合,提升心理健康管理的有效性。
📄 摘要(原文)
Loneliness is a critical public health issue among older adults, linked to higher risks of depression, cognitive decline, and mortality. Scalable, objective methods for its detection remain limited, particularly in natural conversational contexts. We analyzed speech and language markers of loneliness in 310 older adults using semi-structured telephone interviews to help understand how they process feeling lonely and how their language differs at different levels of feeling loneliness. Our multimodal framework combined linguistic features (psycholinguistic dictionaries, n-grams, and topic models) with acoustic features (pitch, tone, loudness) to examine associations with self-reported loneliness scores. Both predefined and data-driven methods captured patterns in verbal content and vocal delivery. Higher loneliness was associated with negations(r = 0.11), negative tone(r = 0.12), and conflict-related language. Lower loneliness was linked to social references(r = -0.18), motivational drives(r = -0.11), and emotional richness in speech(r = -0.12). We also found that the multimodal model (r = 0.298) outperforms the text-only and audio-only models. Findings suggest that loneliness manifests through both linguistic and acoustic cues, supporting the potential of speech-based analysis in psychological assessments and as an early indicator of emotional loneliness when used alongside existing assessments, rather than as standalone diagnostic tools.