Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning
作者: Moloud Damandeh, Meead Saberi
分类: cs.LG, cs.CV
发布日期: 2026-08-07
💡 一句话要点
提出多模态深度学习框架以解决步行可达性感知的主观差异问题
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 步行可达性 多模态深度学习 用户条件模型 视觉感知 个体差异 城市规划 交通管理
📋 核心要点
- 现有研究通常将步行可达性感知简化为统一评分,忽视了个体差异,导致结果的局限性。
- 本文提出了一种用户条件的多模态深度学习框架,结合视觉特征与个体属性,以更好地捕捉步行可达性感知的主观性。
- 实验结果显示,该模型在与观察评分的一致性上提升了65%,表明个体特征对步行可达性感知的影响显著。
📝 摘要(中文)
步行可达性的视觉感知因个人特征、经历和偏好而异,现有研究通常将这些多样化的判断简化为聚合评分,假设感知是一致的,并且常依赖于车辆安装的街景图像,无法反映行人的视觉体验。本文引入了一个包含29,870个步行可达性评分的数据集,链接了来自1196名受访者的步行道视图图像与个体属性,并提出了首个用户条件的多模态深度学习框架,融合了视觉特征与受访者级别的表示。研究表明,步行道视图图像的步行可达性评分显著高于匹配的街景图像,表明图像来源在感知调查中是一个重要的设计决策。该模型在与观察评分的排名一致性上提高了65%,显示出评估环境的个体特征超越了图像内容本身。
🔬 方法详解
问题定义:本文旨在解决步行可达性感知的主观性问题,现有方法往往忽视个体差异,导致评分的统一性和准确性不足。
核心思路:提出用户条件的多模态深度学习框架,通过结合视觉特征与个体属性,来更全面地理解和评估步行可达性感知。
技术框架:该框架包括数据收集、特征提取、模型训练和评估四个主要模块,利用步行道视图和个体属性进行深度学习训练。
关键创新:首次引入用户条件的多模态学习,强调个体在步行可达性评估中的重要性,突破了传统方法的局限。
关键设计:模型采用了融合视觉特征和个体属性的网络结构,使用了特定的损失函数来优化评分一致性,并进行了参数调优以提高模型性能。
🖼️ 关键图片
📊 实验亮点
实验结果表明,用户条件模型在与观察评分的一致性上提升了65%,具体表现为二次加权kappa值从0.29提升至0.47,显示出个体特征在步行可达性评估中的重要性。
🎯 应用场景
该研究的潜在应用领域包括城市规划、交通管理和步行环境评估等。通过更准确地理解步行可达性感知,能够为政策制定者提供更具包容性的评估工具,促进行人友好型城市环境的建设。
📄 摘要(原文)
Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and preferences. Existing studies, however, often reduce these diverse judgements to aggregated scores, implicitly assuming uniform perception, and commonly rely on vehicle-mounted street-view imagery that does not reflect the pedestrian's visual experience. This paper introduces a dataset of 29,870 walkability ratings from 1,196 respondents, linking sidewalk-view imagery across urban, suburban, and regional Australian environments with individual rater attributes, and proposes the first user-conditioned multimodal deep learning framework for walkability perception, fusing visual features with respondent-level representations. A viewpoint-comparison study shows that sidewalk-view images receive significantly higher walkability ratings than matched street-view images, indicating that imagery source is a substantive design decision in perception surveys. The user-conditioned model improves rank agreement with observed ratings by 65% over an image-only baseline (quadratic weighted kappa 0.47 vs. 0.29), demonstrating that who is evaluating an environment carries predictive indication beyond image content alone. These findings support moving from aggregated, observer-independent walkability scores toward models that represent diverse users, enabling more inclusive assessment of pedestrian environments.