Evaluating Semantic and Spatial Guidance for Foundation Model Segmentation of Small-Scale PV in Remote Sensing Imagery

📄 arXiv: 2608.10801v1 📥 PDF

作者: Roni Blushtein-Livnon, Tal Svoray, Osher Rafaeli, Michael Dorman, Itay Fischhendler, Havazelet Yahel, Emir Galilee

分类: cs.CV

发布日期: 2026-08-11


💡 一句话要点

提出基于提示的模型以解决小规模光伏系统遥感影像分割问题

🎯 匹配领域: 支柱九:具身大模型 (Embodied Foundation Models)

关键词: 光伏系统 遥感影像 自动分割 提示模型 数据效率 空间引导 混合提示

📋 核心要点

  1. 现有的遥感影像分割方法在小规模光伏系统的识别上面临挑战,主要由于目标与背景的不平衡。
  2. 本文提出了一种基于提示的语义和空间引导方法,利用SAM3模型进行小规模光伏系统的自动分割。
  3. 实验结果表明,混合提示策略在准确性和稳定性上优于其他提示方式,且在少量标注样本下表现出强大的数据效率。

📝 摘要(中文)

时空光伏数据对于理解离网地区的采纳过程至关重要,但此类数据仍然稀缺。自动化遥感影像分割提供了一种解决方案,但由于住宅光伏系统体积小且分布稀疏,导致目标与背景严重失衡。本文系统评估了SAM3在小规模光伏分割中的表现,通过比较文本、几何和混合提示,分析了不同提示类型的相对贡献。研究表明,提示策略是影响模型行为的主要因素,空间引导显著提高了分割精度和鲁棒性,而混合提示则实现了最高的准确性和稳定性,展示了提示模型在数据受限的离网地区进行光伏映射的潜力。

🔬 方法详解

问题定义:本文旨在解决小规模光伏系统在遥感影像中分割的困难,现有方法在目标与背景失衡的情况下表现不佳。

核心思路:通过引入文本、几何和混合提示,系统评估不同提示类型对模型性能的影响,以优化小规模光伏系统的分割效果。

技术框架:研究采用SAM3模型,分为数据预处理、提示生成、模型训练和性能评估四个主要阶段,确保在不同条件下的全面评估。

关键创新:提出了混合提示策略,结合语义和空间信息,显著提升了模型的分割精度和鲁棒性,与传统方法相比具有本质的改进。

关键设计:在实验中,采用了不同的监督级别和训练策略,设置了多种空间分辨率,并对损失函数和网络结构进行了优化,以适应小规模光伏系统的特性。

🖼️ 关键图片

fig_0
fig_1
fig_2

📊 实验亮点

实验结果显示,混合提示策略在小规模光伏系统分割中实现了最高的准确率和稳定性,相较于文本提示,性能提升显著,且在仅使用几百个标注样本的情况下,展现出强大的数据效率。

🎯 应用场景

该研究的潜在应用领域包括离网地区的光伏系统监测与管理,能够为政策制定者和研究人员提供重要的时空光伏数据支持,促进可再生能源的推广与应用。未来,随着技术的进步,该方法有望扩展到更广泛的遥感影像分析任务中。

📄 摘要(原文)

Spatio-temporal PV data are essential for understanding adoption processes in off-grid regions, yet such data remain largely unavailable. Automated segmentation of remote sensing (RS) imagery offers a promising solution; yet, residential PV systems remain challenging targets because of their small size and sparse distribution, resulting in severe target-background imbalance. Vision-language foundation models (FMs) provide a data-efficient paradigm through prompt-based semantic and spatial guidance, but the relative contribution of different prompt types remains unclear. We systematically evaluate SAM3 for small-scale PV segmentation in RS imagery by comparing textual, geometric, and hybrid prompting, under varying supervision levels, training strategies, spatial resolutions, and imaging conditions. Multi-temporal aerial imagery from a large off-grid rural region serves as a study site, with findings validated across three additional datasets. Prompting strategy emerged as the dominant factor governing model behavior. Textual prompting consistently produced the lowest performance and showed the greatest sensitivity to supervision and imaging conditions. In contrast, spatial guidance substantially improved both segmentation accuracy and robustness. Hybrid prompting achieved the highest accuracy and stability, indicating that semantic and spatial guidance provide complementary information. Most performance gains were achieved with only a few hundred annotated samples, demonstrating strong data efficiency. Transfer learning had limited overall impact, with only modest improvements observed for textual prompting under limited supervision. Overall, our findings establish prompting strategy as a key determinant of SAM3 adaptation, robustness, and generalization, highlighting the potential of promptable FMs for scalable PV mapping in data-constrained off-grid regions.