A Large Open Multi-Energy Corpus of Soil Compaction Tests, with Machine-Learning Baselines

📄 arXiv: 2609.03337v1 📥 PDF

作者: Sompote Youwai, Chana Phutthananon, Warat Kongkitkul

分类: cs.LG

发布日期: 2026-09-03


💡 一句话要点

提出一个大型开放的土壤压实测试数据集以解决数据稀缺问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 土壤压实 数据集 机器学习 Proctor测试 工程填料 密度预测 水分含量 开放数据

📋 核心要点

  1. 现有的土壤压实测试数据通常样本量小且来源单一,限制了研究的广泛性和可靠性。
  2. 本文通过整合来自多个公共来源的压实测试数据,构建了一个大型开放数据集,以提高数据的可用性和准确性。
  3. 实验结果表明,提出的基准模型在密度和水分含量预测上达到了较高的R2值,展示了新方法的有效性。

📝 摘要(中文)

每种工程填料都有其最大干密度和最佳含水量的规定,而这些测定通常需要完整的Proctor测试。现有的相关性研究基于的样本数量有限,且数据通常不公开。本文发布了一个包含2854个实验室压实测试的数据库,涵盖了162个来源组和四个Proctor能量水平,数据来源于六个公共渠道。所有记录均经过审核,确保数据的可靠性。研究结果显示,最佳饱和度为0.815,且提出了一种基准模型用于密度和水分含量的预测,展示了该领域数据的潜在价值与应用前景。

🔬 方法详解

问题定义:本文旨在解决土壤压实测试数据稀缺的问题,现有方法依赖于少量样本且通常来自单一实验室,导致结果的局限性和不可靠性。

核心思路:通过整合来自六个公共来源的2854个实验室压实测试数据,构建一个大型开放数据集,确保数据的多样性和可靠性,从而为土壤压实特性提供更全面的分析基础。

技术框架:研究首先收集和审核数据,确保符合Proctor测试标准,然后进行数据筛选,去除不符合零气泡条件的记录,最后利用机器学习模型进行密度和水分含量的预测。

关键创新:本研究的创新点在于发布了一个大规模的土壤压实测试数据集,打破了以往研究中样本量小和数据不公开的限制,为后续研究提供了丰富的基础数据。

关键设计:在模型设计中,采用了基于分类的预测方法,使用了多个参数设置和损失函数,确保模型在不同数据折叠下的鲁棒性,最终实现了密度和水分含量的有效预测。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,基于新数据集的基准模型在密度预测上达到了R2值0.824,而在水分含量预测上为0.784,显示出较强的预测能力。此外,使用修改后的Proctor记录进行预测时,密度的R2值为0.740,表明新方法在不同条件下的有效性。

🎯 应用场景

该研究的成果可广泛应用于土木工程、地质工程及环境科学等领域,尤其是在土壤压实特性分析和工程设计中。通过提供开放的数据集,研究人员和工程师可以更好地理解土壤行为,优化工程填料的使用,提升工程安全性和经济性。

📄 摘要(原文)

Every engineered fill is specified by a maximum dry density and an optimum moisture content. Each determination needs a full Proctor test. Published correlations rest on one to four hundred specimens, usually from one laboratory at one compactive energy, and are seldom released. This paper releases a corpus without those limits. It holds 2,854 laboratory compaction tests from six public sources, across 162 provenance groups and four Proctor energy levels, with fines from 1.5 to 100%. Every record is audited to the Proctor method its source names, and no energy is inferred. Screening on the zero-air-voids condition removed 11.8% of harmonised records, and 5.7% of those with a measured specific gravity. A material share of published compaction data is physically impossible. The optimum degree of saturation over the corpus is 0.815 at a coefficient of variation of 11%. That is a baseline, not a constant. Both parameters are then estimated from one classification suite and the compaction standard. A tabular foundation model reaches R2 0.824 for density and 0.784 for water content under random folds. It reaches 0.727 and 0.696 with folds drawn around provenance, and 0.520 and 0.614 with a whole source held out. Compactive energy is negligible marginally yet decisive conditionally. Density on the 66 modified-Proctor records is predicted at R2 0.740 with it and -0.651 without. Symbolic regression yields closed forms coupled through a phase relation. No predicted pair can then exceed the zero-air-voids line. The predictions are for screening, not acceptance.