Using large language models to probe the limits of atom-centered structural descriptors
作者: Michelangelo Domina, Michele Ceriotti
分类: physics.chem-ph, cs.LG
发布日期: 2026-07-29
💡 一句话要点
利用大型语言模型探讨原子中心结构描述符的极限
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 原子中心结构 描述符简并 大型语言模型 机器学习 材料科学 跨学科研究
📋 核心要点
- 现有的原子中心结构描述符在较低层次上存在不完整性,导致不同结构可能具有相同的描述符。
- 论文提出通过考虑更大的邻居簇来构建描述符,从而解决描述符简并的问题。
- 实验结果表明,某些三维结构在考虑多达七个邻居时仍然无法区分,展示了AI在跨领域发现中的潜力。
📝 摘要(中文)
将原子结构映射到紧凑的几何描述符集是任何原子尺度建模的机器学习应用中的关键步骤。现有方法通过对对距、三角形等的直方图进行离散化,形成了一系列对称不变的原子中心描述符。然而,较低层次的描述符(如两、三、四邻居簇)被发现存在不完整性,导致对称无关的结构具有相同的描述符。本文通过考虑更大邻居簇来解决这些“描述符简并”问题,并展示了即使在考虑多达七个邻居的情况下,某些三维结构仍然无法区分。通过大型语言模型的帮助,模型能够识别相关文献并理解其在当前问题中的重要性,展示了AI在科学中的潜在应用价值。
🔬 方法详解
问题定义:本文旨在解决原子中心结构描述符在较低层次上存在的简并性问题,现有方法在描述符的构建上存在不足,导致不同结构可能产生相同的描述符。
核心思路:论文的核心思路是通过考虑更大范围的邻居簇来构建描述符,以此消除描述符的简并性,确保不同结构能够被有效区分。
技术框架:整体架构包括对原子结构的分析、邻居簇的选择、描述符的构建和最终的模型训练。主要模块包括数据预处理、特征提取和模型优化。
关键创新:最重要的技术创新在于利用大型语言模型识别和整合不同领域的文献,促进了对描述符构建方法的深入理解和应用,显著提升了描述符的区分能力。
关键设计:在参数设置上,考虑了邻居簇的大小和描述符的离散化程度,损失函数设计为优化描述符的区分能力,网络结构则采用了适合处理高维数据的深度学习模型。
🖼️ 关键图片
📊 实验亮点
实验结果显示,考虑多达七个邻居的描述符仍然无法区分某些三维结构,表明描述符的构建存在根本性挑战。通过大型语言模型的辅助,成功识别并整合了多个领域的相关文献,推动了描述符构建方法的创新。
🎯 应用场景
该研究的潜在应用领域包括材料科学、化学和生物分子建模等,能够为原子尺度的模拟提供更为精确的工具,推动新材料的发现和优化。未来,该方法有望在跨学科研究中发挥重要作用,促进不同领域之间的知识转化与应用。
📄 摘要(原文)
Mapping an atomic structure to a compact set of geometric descriptors is an essential step in any machine-learning application to atomic-scale modeling. A powerful and widely-used approach can be understood as a discretization of the histogram of pair distances, triangles, etc., that results in a hierarchy of symmetry-invariant atom-centered descriptors. Unfortunately, the lower rungs on this hierarchy (two, three, four-neighbor clusters) were found to be incomplete, with symmetry-unrelated pairs of structures having exactly the same descriptors. However, all the ``descriptor degeneracies'' reported so far are resolved by considering larger clusters of neighbors to build the descriptors. We report examples of 3D structures that are indistinguishable even if one considers clusters of up to seven neighbors, and to arbitrary order when considering a practical level of discretization of the descriptors, discovered with the assistance of large language models. The key ingredients in their construction can be traced to results that have been known for decades in different communities; the model was able to find the references and recognize their significance for the problem at hand. We believe this experiment exposes an extremely fruitful usage pattern for AI in science: translating results between different communities and application domains, accelerating the process by which serendipitous discoveries in a field become paradigm-shifting breakthroughs in another.