Do 3D Medical Foundation Models See Through MRI Artifacts? A Controlled Study of Representation Robustness

📄 arXiv: 2608.06613v1 📥 PDF

作者: Julia Anna Mielcarz, Daniel Klaaby, Mostafa Mehdipour Ghazi

分类: cs.CV, cs.AI

发布日期: 2026-08-06

备注: Accepted at the ECCV 2026 Workshop on Artificial Intelligence for Medical 3D Vision (AI4M3D)


💡 一句话要点

评估3D医学基础模型对MRI伪影的鲁棒性

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 3D医学影像 自监督学习 鲁棒性评估 MRI伪影 特征提取 模型评估 深度学习

📋 核心要点

  1. 现有的3D医学基础模型对MRI伪影的敏感性尚未得到充分理解,导致其在实际应用中的鲁棒性问题。
  2. 本文通过控制实验评估五种预训练3D编码器的表示鲁棒性,探讨不同模型对MRI伪影的响应。
  3. 实验结果表明,模型的鲁棒性与伪影类型密切相关,3DINO在多种条件下表现出最稳定的表示能力。

📝 摘要(中文)

自监督的3D医学基础模型越来越多地被用作通用特征提取器,但它们对MRI伪影的敏感性仍然不够了解。本文对五种不同架构、目标、预训练领域和数据集规模的预训练3D编码器进行了控制评估。通过使用BraTS-Africa案例生成七种伪影,评估结果显示鲁棒性强烈依赖于模型和伪影类型。3DINO表现出最稳定的表示,而BrainIAC对多种伪影高度敏感。研究表明,仅依靠大规模或领域特定的预训练并不能保证对伪影的不变性,强调了在异构MRI环境中部署3D基础模型前进行明确鲁棒性评估的必要性。

🔬 方法详解

问题定义:本文旨在解决自监督3D医学基础模型在面对MRI伪影时的鲁棒性问题。现有方法对伪影的敏感性尚未得到充分评估,影响了其在临床应用中的可靠性。

核心思路:通过对五种不同架构的预训练3D编码器进行控制评估,分析其在不同伪影条件下的表现,以揭示模型的鲁棒性特征。

技术框架:研究使用BraTS-Africa数据集生成七种伪影,采用线性中心核对齐(CKA)、RankMe和UMAP等方法评估鲁棒性,并进行独立的分割一致性分析。

关键创新:本文的创新在于系统性地评估不同3D编码器对MRI伪影的鲁棒性,发现鲁棒性不仅依赖于模型架构,还与伪影类型密切相关。

关键设计:实验中设置了五种预定义的伪影条件,使用多种评估指标(如CKA和RankMe)来量化模型在不同伪影下的表现,确保评估的全面性和准确性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,3DINO在多种伪影条件下表现出最稳定的表示能力,而BrainIAC对多种伪影高度敏感。CKA指标在多种条件下显著下降,而RankMe保持相对稳定,表明伪影会扭曲表示几何但不导致维度崩溃。

🎯 应用场景

该研究的潜在应用领域包括医学影像分析、临床诊断支持系统和医疗AI工具的开发。通过提高3D医学基础模型对MRI伪影的鲁棒性,可以增强其在临床环境中的应用价值,确保更可靠的诊断结果。

📄 摘要(原文)

Self-supervised 3D medical foundation models are increasingly used as general-purpose feature extractors, yet their sensitivity to MRI artifacts remains poorly understood. We present a controlled evaluation of representation robustness across five pretrained 3D encoders spanning different architectures, objectives, pretraining domains, and dataset scales. Using BraTS-Africa cases with four MRI sequences, we generate seven frequency- and image-domain artifacts at five predefined corruption settings. Robustness is assessed using linear centered kernel alignment (CKA), RankMe, and UMAP, complemented by an independent segmentation-consistency analysis. We find that robustness is strongly model- and artifact-dependent. 3DINO exhibits the most consistently stable representations, while BrainIAC is highly sensitive to several corruptions; NeuroVFM, BrainFM, and Neuro-SimCLR show intermediate but distinct artifact-specific profiles. Across many conditions, CKA decreases substantially while RankMe remains comparatively stable, indicating that artifacts often distort representation geometry without causing dimensional collapse. Segmentation consistency also degrades under corruption, particularly for ghosting and Rician noise, but aligns only partially with representation-level robustness. These findings show that larger-scale or domain-specific pretraining alone does not guarantee artifact invariance and motivate explicit robustness evaluation before deploying 3D foundation models in heterogeneous MRI settings.