Tree species mapping in Denmark: A comparison of spectral-temporal features with geospatial foundation model embeddings
作者: Alkiviadis Koukos, Spyros Kondylatos, Thomas Nord-Larsen, Lotte Nyborg, Christian Tøttrup, Kenneth Grogan
分类: cs.CV, cs.AI, cs.LG
发布日期: 2026-09-03
备注: Submitted to Remote Sensing of Environment. This preprint presents a national-scale tree species mapping framework for Denmark using Sentinel-1/2 time series, National Forest Inventory data, and EO foundation model embeddings. The resulted national map can be found here: https://zenodo.org/uploads/22108850
💡 一句话要点
提出基于光谱时序特征与基础模型嵌入的树种映射方法
🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)
关键词: 树种映射 光谱时序特征 基础模型 森林监测 生态研究 机器学习 遥感数据
📋 核心要点
- 现有的树种分类方法在大规模森林特征化中面临数据稀缺和分类精度不足的挑战。
- 论文提出结合光谱时序特征和基础模型嵌入的树种分类方法,以提高分类精度和适应性。
- 实验结果表明,基于STF的MLP模型在纯林分类中表现最佳,且TESSERA在训练样地不足时具有明显优势。
📝 摘要(中文)
本研究利用国家森林清查样地和遥感数据,绘制丹麦的树种分布图,并评估基础模型在大规模森林特征化中的潜力。比较了两种树种分类输入表示:手工设计的光谱时序特征(STF)和由基础模型生成的嵌入。研究显示,基于STF的多层感知机(MLP)在纯林和混交林中表现最佳,宏观F1分数分别为0.843和0.653。TESSERA嵌入在训练样地不足时表现优越,且多年的观测数据显著提高了分类精度。最终生成的10米分辨率树种地图为丹麦提供了首个高分辨率的国家级树种地图,具有重要的生态监测和土地管理价值。
🔬 方法详解
问题定义:本研究旨在解决大规模森林树种分类中的数据稀缺和分类精度不足的问题。现有方法往往依赖于有限的手工特征,导致分类效果不理想。
核心思路:论文提出结合手工设计的光谱时序特征(STF)与基础模型生成的嵌入,以提高树种分类的准确性和适应性,尤其是在训练样地不足的情况下。
技术框架:整体架构包括数据收集、特征提取、模型训练和评估四个主要阶段。首先,从Sentinel-1和Sentinel-2获取多时相数据,提取光谱时序特征,并生成基础模型嵌入。然后,使用随机森林、XGBoost和多层感知机(MLP)进行分类。
关键创新:最重要的创新在于将基础模型嵌入与传统的光谱时序特征结合,特别是在训练样地不足时,TESSERA嵌入表现出显著的优势。
关键设计:在模型训练中,采用了随机森林、XGBoost和MLP等多种分类器,MLP模型在超参数设置上进行了优化,损失函数选择了适合多类分类的交叉熵损失,确保了模型的稳定性和准确性。实验中还进行了消融实验,以验证不同特征的贡献。
🖼️ 关键图片
📊 实验亮点
实验结果显示,基于STF的MLP模型在纯林分类中取得了宏观F1分数0.843,而TESSERA嵌入在训练样地不足时表现出色,分类精度接近最佳模型。最终生成的树种地图在区域调整验证中显示出79.9%的整体准确率,具有重要的实用价值。
🎯 应用场景
该研究的树种映射方法可广泛应用于森林监测、生态研究和土地管理等领域。生成的高分辨率树种地图为政策制定者和研究人员提供了重要的数据支持,有助于实现可持续的森林管理和生态保护。
📄 摘要(原文)
We map tree species across Denmark using National Forest Inventory plots and EO data, while evaluating the potential of foundation models for large-scale forest characterization. We compare two alternative input representations for tree species classification: (i) manually engineered spectral-temporal features (STF) derived from multi-temporal Sentinel-1 and Sentinel-2 observations, and (ii) embeddings generated by the EO FMs TESSERA and AlphaEarth. Both representations are complemented with canopy height information. Random forest, XGBoost, and Multi-Layer Perceptron (MLP) classifiers are evaluated for all input representations, with separate assessments for pure and mixed forest stands. The STF-based MLP achieves the highest classification performance, yielding macro F1 scores of 0.843 and 0.653 for pure and mixed stands, respectively. The MLP trained on TESSERA embeddings delivers competitive performance for pure stands, achieving results within 1.1 percentage points of the best-performing model. TESSERA consistently outperforms STF-based models when fewer than approximately 25% of training plots are available, demonstrating a substantial advantage under limited training data. Multi-year observations systematically improve classification accuracy relative to single-year inputs, while ablation experiments reveal the complementary contributions of Sentinel-1 backscatter, spectral indices, and canopy height data. The best-performing model is subsequently applied at the national scale to generate a 10 m tree species map of Denmark. Area-adjusted validation indicates an overall map accuracy of 79.9%. The resulting map, released as an open-access product, is the first high-resolution national tree species map of Denmark and provides a valuable resource for forest monitoring, ecological research, and land management applications.