Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

📄 arXiv: 2608.26090v1 📥 PDF

作者: Raphaël Bonnet-Guerrini, Johann Ioannou-Nikolaides, Inar Timiryasov, Vincenzo Piuri

分类: astro-ph.HE, cs.AI, cs.LG, hep-ex

发布日期: 2026-08-26


💡 一句话要点

提出稀疏自编码器以实现中微子模型的可解释性

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 粒子物理学 中微子模型 稀疏自编码器 可解释性 机器学习 角度重建 物理概念

📋 核心要点

  1. 现有方法在粒子物理学中缺乏有效的可解释性,难以理解模型内部表示的物理概念。
  2. 论文提出通过稀疏自编码器实现中微子模型的机制可解释性,识别并利用模型中的物理概念图谱。
  3. 实验结果显示,新的不确定性头显著提高了角度重建的中位分辨率,从20.2°降至3.2°,展示了方法的有效性。

📝 摘要(中文)

本文首次将基于稀疏自编码器的机制可解释性应用于粒子物理学。研究了在IceCube数据上预训练并针对方向重建微调的中微子基础模型,识别出模型表示中的物理概念图谱。通过严格的验证协议,使用因果干预显示方向头几乎不依赖于该图谱。为此,本文在相同事件级表示上训练了一个不确定性头,以预测模型的角度重建误差。该可解释估计器在20%的选择效率下,将中位角分辨率从20.2°提升至3.2°,表明机制可解释性能够揭示模型内部表示中编码的潜在物理信息,并帮助设计利用这些信息的下游任务。

🔬 方法详解

问题定义:本文旨在解决粒子物理学中模型可解释性不足的问题,现有方法难以揭示模型内部表示的物理概念,导致对模型预测的理解有限。

核心思路:通过稀疏自编码器实现机制可解释性,识别模型中的物理概念图谱,并利用这些信息改进模型的预测性能。

技术框架:整体架构包括预训练的中微子基础模型,稀疏自编码器用于提取物理概念,方向头和不确定性头分别用于角度重建和误差预测。

关键创新:最重要的创新在于通过稀疏自编码器识别出物理概念图谱,并利用该图谱训练不确定性头,从而显著提高了角度重建的精度。

关键设计:在模型训练中,采用了严格的验证协议,包括持出测试、匹配干扰控制和独立字典训练的复制,确保了结果的可靠性和可解释性。

🖼️ 关键图片

fig_0
img_1
img_2

📊 实验亮点

实验结果表明,新的不确定性头在20%的选择效率下,将中位角分辨率从20.2°显著提升至3.2°,显示出稀疏自编码器在提取物理概念方面的有效性和实用性。

🎯 应用场景

该研究在粒子物理学领域具有重要应用潜力,能够帮助科学家更好地理解中微子及其相互作用。此外,所提出的方法也可推广至其他需要可解释性的机器学习任务,提升模型的透明度和信任度。

📄 摘要(原文)

We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on the same event-level representation to predict the model's angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At $20\%$ selection efficiency, this interpretable estimator improves the median angular resolution from $20.2^\circ$ to $3.2^\circ$. These results suggest that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.