| 注册
首页|期刊导航|南京信息工程大学学报|融合场景多模态先验与稀疏注意力的文本图像超分辨率

融合场景多模态先验与稀疏注意力的文本图像超分辨率

周颖 易尧华 余长慧 饶杨莉 王颖洁

南京信息工程大学学报2026,Vol.18Issue(3):310-320,11.
南京信息工程大学学报2026,Vol.18Issue(3):310-320,11.DOI:10.13878/j.cnki.jnuist.20250413001

融合场景多模态先验与稀疏注意力的文本图像超分辨率

Fusing scene multimodal prior and sparse attention for text image super-resolution

周颖 1易尧华 1余长慧 1饶杨莉 2王颖洁2

作者信息

  • 1. 武汉大学遥感信息工程学院,武汉,430079
  • 2. 自然资源部西南山地自然资源遥感监测工程技术创新中心,成都,610000
  • 折叠

摘要

Abstract

Recovering High-Resolution(HR)images from Low-Resolution(LR)text images is extremely chal-lenging due to factors such as complex backgrounds,blurring,warping and distortion.Existing methods,which pre-dominantly rely on recurrent neural networks to extract textual context,often struggle with capturing long-distance dependencies and leveraging semantic information effectively.To address the above problems,this paper proposes a novel text image super-resolution approach that integrates scene multimodal priors and sparse attention.First,we in-novatively introduce a Scene Multimodal Prior Branch(SMPB),which leverages an advanced context parsing unit and a contour perception unit to fully explore and utilize both textual and visual information.Second,a Sparse At-tention-Based Super-Resolution(SABSR)enhancement module is designed to extract contextual information from text lines.By utilizing the global receptive field of the multi-head attention mechanism,it effectively constructs in-ter-character correlations,thereby alleviating the performance degradation commonly encountered with long text se-quences.Finally,the model's capability in extracting text contours and processing deformed text is significantly en-hanced by a joint loss function that combines gradient contour awareness with text structure perception.Experimen-tal results show that compared to the baseline model Text ATTention network(TATT),our method improves recogni-tion accuracy on the TextZoom test set by 4.3 percentage points on average.Furthermore,it achieves average PSNR and SSIM values of 21.4 dB and 0.790 9,respectively,demonstrating superior performance for text image super-res-olution in real-world scenarios.

关键词

文本图像超分辨率/图像重建/多模态先验/稀疏注意力

Key words

text image super-resolution/image reconstruction/multimodal prior/sparse attention

分类

信息技术与安全科学

引用本文复制引用

周颖,易尧华,余长慧,饶杨莉,王颖洁..融合场景多模态先验与稀疏注意力的文本图像超分辨率[J].南京信息工程大学学报,2026,18(3):310-320,11.

基金项目

自然资源部西南山地自然资源遥感监测工程技术创新中心开放课题(RSMNR-SCM-2024-001) (RSMNR-SCM-2024-001)

新疆生产建设兵团重点研发项目(2024AB064) (2024AB064)

南京信息工程大学学报

1674-7070

访问量0
|
下载量0
段落导航相关论文