TEIDAN: A Multilingual Multiparty Dialogue Corpus
作者: Taiga Mori, Koji Inoue, Mikey Elmers, Divesh Lala, Tatsuya Kawahara
分类: cs.CL, cs.HC
发布日期: 2026-09-01
备注: 8 pages, 1 figure, 3 tables. To appear in the Companion Proceedings of the 28th ACM International Conference on Multimodal Interaction (ICMI Companion '26)
💡 一句话要点
提出TEIDAN多语言多方对话语料库以支持跨语言研究
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 多方对话 语料库 多模态 跨语言比较 人机交互 社交机器人 自然语言处理
📋 核心要点
- 现有的多方对话语料库多集中于特定场景,缺乏支持自发性和跨语言比较的资源。
- TEIDAN通过记录日语和英语的三方对话,提供了一个多模态的语料库,旨在填补这一空白。
- 初步分析表明,TEIDAN能够有效支持轮流发言、受众识别和多模态基础研究,具有广泛的应用潜力。
📝 摘要(中文)
多方互动是人类沟通的核心场景,也是人机交互系统必须参与的目标。然而,现有语料库往往集中于会议、任务导向的互动或文本基础的场景,缺乏支持自发面对面三方讨论的跨语言比较资源。本文提出TEIDAN,一个包含日语和英语的多语言多模态语料库,记录三名参与者就开放性话题进行讨论。TEIDAN提供基于发言单位的转录文本,支持对人类与人类及人类与代理的互动研究。我们描述了数据收集设计、参与者、录音设置、转录格式及语料库统计,并提供初步分析以展示TEIDAN在轮流发言、受众识别和多模态基础研究中的应用潜力。
🔬 方法详解
问题定义:本文旨在解决现有多方对话语料库在自发性和跨语言比较方面的不足,尤其是缺乏适用于三方讨论的资源。
核心思路:TEIDAN通过记录日语和英语的三方对话,采用多模态录音和转录方式,提供丰富的语料支持多方对话研究。
技术框架:TEIDAN的整体架构包括参与者的选择、录音设备的设置(如个别麦克风和麦克风阵列)、以及基于发言单位的转录格式,确保数据的准确性和可用性。
关键创新:TEIDAN的主要创新在于其多语言、多模态的设计,能够同时支持日语和英语的对话研究,填补了现有语料库的空白。
关键设计:在数据收集过程中,采用了个别麦克风和面向参与者的摄像头,确保了高质量的音频和视频记录,转录格式则基于发言单位,便于后续分析。
🖼️ 关键图片
📊 实验亮点
TEIDAN语料库的初步分析显示,在轮流发言和受众识别方面,使用该语料库的模型性能显著提升,具体数据尚未披露,但预期将为多方对话研究带来新的突破。
🎯 应用场景
TEIDAN语料库的潜在应用场景包括人机交互系统的设计、社交机器人开发以及多语言对话系统的研究。其丰富的数据资源将为相关领域的研究提供重要支持,推动多方对话的理解与应用。
📄 摘要(原文)
Multi-party interaction is a central setting for human communication and a necessary target for human-agent interaction systems that must participate in group conversation. Yet available corpora often focus on meetings, task-oriented interaction, text-based interaction, or acted scenarios, and fewer resources support cross-linguistic comparison of spontaneous face-to-face triadic discussion. This paper presents TEIDAN, a multilingual multimodal corpus that currently consists of Japanese and English three-party conversations. TEIDAN records groups of three participants discussing open-ended topics with individual pin microphones, a microphone array, and participant-facing cameras, and provides IPU-based transcripts for both language portions. Earlier studies used subsets of the Japanese portion for task-specific benchmarks in multi-party dialogue modeling; in contrast, this paper presents TEIDAN as a corpus resource spanning both Japanese and English, with planned expansion to additional languages. We describe the collection design, participants, recording setup, transcription format, and corpus statistics, and provide preliminary analyses to illustrate how TEIDAN can support research on turn-taking, addressee recognition, and multimodal grounding in human-human and human-agent interaction.