DuplexWorld: Can voice agents help you get through the day?
作者: Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli, Asif Shaik, Abhishek Mukherji, Dinesh Manocha
分类: cs.SD, cs.AI, cs.CL
发布日期: 2026-08-11
💡 一句话要点
提出DuplexWorld以全面评估语音助手在日常生活中的应用
🎯 匹配领域: 支柱一:机器人控制 (Robot Control)
关键词: 语音助手 对话系统 评估标准 人工智能 用户体验 多样性对话 自然语言处理
📋 核心要点
- 现有语音助手评估方法未能充分考虑日常活动中的对话多样性,导致其在实际应用中的表现不足。
- DuplexWorld通过构建六个特定场景,系统性地评估语音助手在不同对话类型下的表现,提供更全面的评估标准。
- 实验结果表明,当前最佳语音助手在代理性、对话性和语音自然性方面均有待提升,展示了未来改进的方向。
📝 摘要(中文)
语音助手在企业客户服务和消费者日常生活中越来越普遍,然而现有评估标准未能全面衡量其在多样化对话中的表现。DuplexWorld引入六个特定场景(如银行、保险、旅行等),并在156个场景中测试语音助手的对话和分析能力。通过对代理性、对话性和语音自然性等指标的评估,结果显示即使是最优秀的语音助手在这三方面仍有显著提升空间(Pass@1: 0.490,轮流对话: 0.653,DNSMOS: 3.378)。
🔬 方法详解
问题定义:本论文旨在解决现有语音助手评估方法未能全面考虑日常对话多样性的问题,导致评估结果无法真实反映其在实际应用中的表现。
核心思路:论文提出DuplexWorld,通过构建六个特定场景,系统性地评估语音助手在不同对话类型下的表现,强调对话的多样性和复杂性。
技术框架:整体架构包括六个应用场景,针对每个场景设计了156个对话场景,评估指标涵盖代理性、对话性和语音自然性等多个维度。
关键创新:最重要的创新在于引入了多样化的场景和对话类型,使得评估不仅限于数据库操作,而是涵盖了更广泛的实际应用场景。
关键设计:在实验中,设置了11种不同类型的对话,使用了350小时以上的对话数据,评估指标包括Pass@1、轮流对话率和DNSMOS等,确保评估的全面性和准确性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,当前最佳语音助手在代理性、对话性和语音自然性方面的表现分别为Pass@1: 0.490,轮流对话: 0.653,DNSMOS: 3.378,表明在所有三个评估维度上均有显著的改进空间,强调了未来研究的必要性。
🎯 应用场景
该研究的潜在应用领域包括客户服务、医疗咨询、旅行规划等,能够为企业提供更智能的语音助手解决方案,提升用户体验。未来,随着语音助手技术的进步,可能会在更多日常生活场景中得到广泛应用,改变人们的生活方式。
📄 摘要(原文)
Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation. To tackle this DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.