浙江大学学报(理学版)2026,Vol.53Issue(4):511-520,10.DOI:10.3785/1008-9497.24042
结合超像素与Transformer的高分辨率遥感图像语义分割算法
Semantic segmentation method for high-resolution remote sensing images combining superpixel and Transformer
摘要
Abstract
A Transformer-based semantic segmentation model SegFormer-MSFF is proposed for automatic and accurate semantic segmentation of high-resolution remote sensing images.In order to solve the problem that when the existing Transformer-based method performs grid slicing and serializes the image,it may result in discontinuous semantic segmentation results between image slices,a sub-network with lightweight multi-scale superpixel information extraction is introduced based on convolutional neural network.The boundary information of the image is preserved through multi-scale dynamic superpixel information,and it is fused with the middle semantic features at all levels to enhance the Transformer's ability to capture the continuity between image slices.A joint loss function is proposed to achieve dynamic superpixel segmentation,which weights the superpixel segmentation loss and the semantic segmentation loss to obtain the objective function of the overall model,allowing the overall model to simultaneously explore the semantic and superpixel features of the image.Experiments on the high-resolution aerial remote sensing images datasets Vaihingen and Potsdam showed that the mean intersection over union of SegFormer-MSFF is 1.4%and 0.9%higher than that of SegFormer.关键词
高分辨率遥感图像/语义分割/超像素分割/Transformer/卷积神经网络Key words
high-resolution remote sensing images/semantic segmentation/superpixel segmentation/Transformer/convolutional neural network(CNN)分类
信息技术与安全科学引用本文复制引用
吴津锋,柴登峰..结合超像素与Transformer的高分辨率遥感图像语义分割算法[J].浙江大学学报(理学版),2026,53(4):511-520,10.基金项目
国家自然科学基金项目(42471346) (42471346)
浙江省自然科学基金项目(LY22D010003). (LY22D010003)