南京信息工程大学学报2026,Vol.18Issue(3):310-320,11.DOI:10.13878/j.cnki.jnuist.20250413001
融合场景多模态先验与稀疏注意力的文本图像超分辨率
Fusing scene multimodal prior and sparse attention for text image super-resolution
摘要
Abstract
Recovering High-Resolution(HR)images from Low-Resolution(LR)text images is extremely chal-lenging due to factors such as complex backgrounds,blurring,warping and distortion.Existing methods,which pre-dominantly rely on recurrent neural networks to extract textual context,often struggle with capturing long-distance dependencies and leveraging semantic information effectively.To address the above problems,this paper proposes a novel text image super-resolution approach that integrates scene multimodal priors and sparse attention.First,we in-novatively introduce a Scene Multimodal Prior Branch(SMPB),which leverages an advanced context parsing unit and a contour perception unit to fully explore and utilize both textual and visual information.Second,a Sparse At-tention-Based Super-Resolution(SABSR)enhancement module is designed to extract contextual information from text lines.By utilizing the global receptive field of the multi-head attention mechanism,it effectively constructs in-ter-character correlations,thereby alleviating the performance degradation commonly encountered with long text se-quences.Finally,the model's capability in extracting text contours and processing deformed text is significantly en-hanced by a joint loss function that combines gradient contour awareness with text structure perception.Experimen-tal results show that compared to the baseline model Text ATTention network(TATT),our method improves recogni-tion accuracy on the TextZoom test set by 4.3 percentage points on average.Furthermore,it achieves average PSNR and SSIM values of 21.4 dB and 0.790 9,respectively,demonstrating superior performance for text image super-res-olution in real-world scenarios.关键词
文本图像超分辨率/图像重建/多模态先验/稀疏注意力Key words
text image super-resolution/image reconstruction/multimodal prior/sparse attention分类
信息技术与安全科学引用本文复制引用
周颖,易尧华,余长慧,饶杨莉,王颖洁..融合场景多模态先验与稀疏注意力的文本图像超分辨率[J].南京信息工程大学学报,2026,18(3):310-320,11.基金项目
自然资源部西南山地自然资源遥感监测工程技术创新中心开放课题(RSMNR-SCM-2024-001) (RSMNR-SCM-2024-001)
新疆生产建设兵团重点研发项目(2024AB064) (2024AB064)