StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
作者: Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding, Qingshui Gu, Yongjie Zhong, Jinxiang Meng, Yuan Gao, Yunqiu Zhou, Hao Zhu, Jifeng He, Yongzhi Liao, Xinyi Zhang, Chaoxin Li, Yi Zhu, Xi Lin, Duju Zeng, Xiang Gao, Wen Zhang, Yunyang Wang, Duo Wang, Huan Zhou, Zuo Wang, Jin Chen, Kaiyuan Zhang, Chuqian Yu, Tianhao Yu, Longxiang Liu, Jianbo Xue, Huimin Che, Jiahao Wang, Yujia Qin, Jiaheng Liu, Shen Yan, Xiaolong Chang, Wenhao Huang
分类: cs.AI
发布日期: 2026-08-18
💡 一句话要点
提出StartupBench以评估通用智能体在市场验证工作流中的表现
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 智能体评估 市场验证 工作流分析 大型语言模型 真实用户需求
📋 核心要点
- 现有基准测试主要依赖研究者选择的任务,无法反映真实用户的需求和市场验证的工作流。
- 提出StartupBench,通过系统研究市场验证的AI产品及其工作流,识别真实世界的任务需求。
- 实验结果显示,即使是最强模型在StartupBench上的成功完成率仅约30%,揭示了当前智能体的局限性。
📝 摘要(中文)
近年来,大型语言模型和智能体的进展显著提升了AI系统执行复杂任务的能力。然而,现有基准测试主要依赖研究者选择的任务,无法确定这些进展是否适用于真实用户的需求。本文提出了StartupBench,一个基于市场验证的AI初创产品的端到端智能体基准。我们系统研究了已被采用的AI产品及其工作流,识别出AI在多专业领域的实际需求任务。通过细致的评估标准,我们发现即使是最强模型在StartupBench上的成功完成率仅约30%。分析表明,复杂指令遵循和领域特定专业知识是主要失败原因,揭示了当前通用智能体在满足市场验证工作流方面的局限性。
🔬 方法详解
问题定义:本文旨在解决现有基准测试无法反映真实用户需求的问题,现有方法未能有效评估智能体在实际工作流中的表现。
核心思路:通过分析市场验证的AI产品及其工作流,识别出真实的任务需求,并将其转化为可交付的任务进行评估。
技术框架:整体架构包括市场验证产品的研究、工作流的识别、任务的定义和评估标准的制定,形成一个系统化的评估流程。
关键创新:StartupBench作为一个基于市场验证的基准,首次将真实用户需求与智能体能力评估结合,填补了现有方法的空白。
关键设计:在任务评估中,采用细致的评分标准,关注复杂指令遵循和领域特定知识,确保评估的全面性和准确性。
🖼️ 关键图片
📊 实验亮点
实验结果显示,尽管许多模型在StartupBench上取得了一定的部分进展,但最强模型的成功完成率仅为30%。这一结果强调了当前通用智能体在复杂任务执行中的显著局限性,尤其是在复杂指令遵循和领域专业知识方面。
🎯 应用场景
该研究的潜在应用领域包括AI产品开发、智能体设计和用户需求分析。通过提供一个基于市场验证的评估标准,StartupBench可以帮助开发者更好地理解用户需求,从而提升AI系统的实用性和可靠性,推动行业进步。
📄 摘要(原文)
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.