SSP: An Event-Matched Syn2Sim2Phy Cross-Domain Evaluation Framework for Autonomous Driving VLA Models

📄 arXiv: 2608.14024v1 📥 PDF

作者: Haojie Feng, Peizhi Zhang, Xinrui Zhang, Zhuoren Li, Junpeng Huang, Xiurong Wang, Dongxiao Yin, Yuxiang Zhang, Junfan Zhu, Lu Xiong

分类: cs.CV

发布日期: 2026-08-14


💡 一句话要点

提出SSP框架以解决跨域评估VLA模型的性能差异问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 跨域评估 自动驾驶 VLA模型 事件匹配 性能评估 多模态融合 安全关键交互

📋 核心要点

  1. 现有的VLA模型评估方法常常使用独立选择的合成、模拟和物理数据,导致性能差异可能受到场景内容变化的影响。
  2. SSP框架通过构建事件匹配的评估标准,确保在不同域间进行一致的交互评估,从而消除内容变化的干扰。
  3. 实验结果表明,SSP在合成、模拟和物理域的宏观平均综合VLA能力得分分别为0.259、0.291和0.325,展示了不同场景下的性能差异。

📝 摘要(中文)

本文提出了一种名为SSP(Synthetic-Simulation-Physical)的事件匹配跨域评估框架,旨在解决现有评估方法中因场景内容变化而导致的性能差异混淆问题。SSP通过从合成长尾视频出发,构建一个验证的事件规范,确保在合成、模拟和物理域中进行一致的安全关键交互评估。该框架将来自不同平台的输出映射到共同的语义槽和1秒的轨迹窗口,以评估输出的有效性、语义准确性、关键交互识别、轨迹质量和风险响应。实验结果显示,在不同场景下,SSP能够提供可重复的场景转移链和证据合格的VLA行为评估。

🔬 方法详解

问题定义:本文旨在解决现有VLA模型评估中因场景内容变化而导致的性能差异混淆问题。现有方法往往依赖于独立选择的数据集,缺乏一致性和可比性。

核心思路:SSP框架通过构建事件匹配的评估标准,确保在合成、模拟和物理域中进行一致的安全关键交互评估,从而消除内容变化的干扰。

技术框架:SSP的整体架构包括事件规范的构建、平台特定实现的构建(在CARLA和封闭测试场上),以及在转移审核后进行的评估。主要模块包括事件规范、平台实现和评估模块。

关键创新:SSP的主要创新在于其事件匹配的评估方法,能够在不同域之间进行有效的比较,而不假设物理域总是优于其他域。与现有方法相比,SSP提供了更为可靠和一致的评估标准。

关键设计:在设计中,SSP确保了事件属性的保留,并通过将不同平台的输出映射到共同的语义槽和轨迹窗口来进行评估,关注输出的有效性、语义准确性和风险响应等关键指标。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,SSP在合成、模拟和物理域的宏观平均综合VLA能力得分分别为0.259、0.291和0.325,表明不同场景下的性能差异显著。Alpamayo-R1、OpenEMMA和LLaViDA的得分分别为0.405、0.338和0.131,展示了SSP在跨域评估中的有效性。

🎯 应用场景

该研究的潜在应用领域包括自动驾驶系统的性能评估、智能交通系统的安全性验证以及多模态AI系统的开发。通过提供一致的评估框架,SSP能够帮助研究人员和工程师更好地理解和优化VLA模型的行为,从而提升自动驾驶技术的安全性和可靠性。

📄 摘要(原文)

Vision-language-action (VLA) models for autonomous driving jointly produce scene interpretation, language-based reasoning, and driving trajectories. Existing evaluations often use independently selected synthetic, simulated, and physical data, so measured performance gaps can be confounded by changes in scenario content rather than genuine domain sensitivity. We propose SSP (Synthetic-Simulation-Physical), an event-matched Syn2Sim2Phy evaluation framework that anchors cross-domain comparison to the same safety-critical interaction. Starting from a synthetic long-tail video, SSP builds a validated event specification that preserves road topology, participant roles, relative motion, conflict evolution, passing order, response constraints, and event phases. Platform-specific realizations are then constructed in CARLA and on a closed proving ground and are evaluated only after transfer audits confirm preservation of mandatory event properties. SSP maps heterogeneous outputs from OpenEMMA, LLaViDA, and Alpamayo-R1 into common semantic slots and a 1 s trajectory window to assess output validity, semantic accuracy, critical-interaction recognition, trajectory quality, and risk response. Across Cut-in and vulnerable-road-user crossing cases, the macro-averaged Integrated VLA Capability Scores are 0.259, 0.291, and 0.325 in the Synthetic, Simulation, and Physical domains, respectively, while the best domain varies by scenario. Alpamayo-R1, OpenEMMA, and LLaViDA obtain scores of 0.405, 0.338, and 0.131. SSP provides a reproducible scene-transfer chain and an evidence-qualified evaluation of VLA behavior without assuming that the Physical domain is universally superior.