| 注册
首页|期刊导航|电子学报|面向机器视觉的文本提示引导的图像编码

面向机器视觉的文本提示引导的图像编码

黄志勐 高峰 杨帆 马思伟

电子学报2026,Vol.54Issue(1):19-31,13.
电子学报2026,Vol.54Issue(1):19-31,13.DOI:10.12263/DZXB.20250778

面向机器视觉的文本提示引导的图像编码

Text Prompted Image Coding for Machine

黄志勐 1高峰 2杨帆 2马思伟1

作者信息

  • 1. 北京大学计算机学院,北京 100871
  • 2. 北京大学艺术学院,北京 100871
  • 折叠

摘要

Abstract

In recent years,with the rapid development of classic machine-to-machine(M2M)communication scenari-os such as the internet of things(IoT),semantic communication,and smart cities,the real-time transmission and efficient processing of massive visual data between devices have become a critical challenge.In this context,traditional image cod-ing methods,which are primarily optimized for human perceptual quality,often suffer from insufficient analysis accuracy when applied to machine vision tasks due to a fundamental mismatch between their optimization objectives and the require-ments of machine analysis.Consequently,image coding for machine(ICM)has emerged,aiming to maintain high analysis accuracy for downstream machine vision tasks(e.g.,classification,detection,segmentation)while achieving the lowest pos-sible bitrate,thereby better adapting to the bandwidth and storage constraints in M2M scenarios.However,existing ICM methods still face two major bottlenecks.First,their performance degrades sharply under extremely low bitrates.This is be-cause most current approaches rely on end-to-end nonlinear transformations to extract visual features,failing to fully exploit the compact representation of high-level semantic information within images,which leads to inefficient feature coding.Sec-ond,they exhibit weak generalization in open-set scenarios.Most methods are optimized for single tasks or single datasets,lacking the adaptability to unseen categories or cross-domain data,and thus struggle to maintain stable analytical perfor-mance in practical,dynamic environments.To overcome these limitations,this paper proposes a novel text-prompted image coding for machine(T-ICM)framework.The core idea is to decouple image information into two complementary compo-nents:semantic information and texture information.The semantic information is represented and encoded in the form of structured text prompts(e.g.,object categories,location descriptions),while the texture information is extracted and com-pressed as task-agnostic general visual features.At the encoder side,the text prompts,owing to their highly abstract and se-mantically compact nature,can significantly reduce the overall bitrate.The general features are efficiently compressed via our proposed grouped feature coding module.At the decoder side,the text prompts serve not only for direct parsing to ac-complish tasks like classification and detection but,more importantly,act as guidance signals.Through a prompt encoder and a mask decoder,they dynamically adjust the semantically relevant regions of the reconstructed general features,en-abling feature-level domain adaptation and task-specific adaptation,thereby significantly enhancing the model's robustness in open-set scenarios.The proposed T-ICM is comprehensively evaluated on multiple standard datasets and tasks.Experi-ments demonstrate that on dense prediction tasks such as semantic segmentation and instance segmentation,T-ICM can maintain analysis accuracy close to that of using the original uncompressed images even at very low bitrates,significantly outperforming H.266/VVC,learned image codecs,and other existing ICM methods.By migrating semantic information to the highly compressed text modality for transmission and utilizing it to guide feature reconstruction,T-ICM achieves a supe-rior trade-off between coding efficiency and task performance.This work provides a novel perspective and technical founda-tion for the future development of semantic communication,collaborative edge intelligence,and adaptive machine vision systems.

关键词

视频编码/智能编码/特征编码/面向机器视觉的特征编码/深度学习/信号处理

Key words

video coding/intelligent compression/feature coding/feature coding for machine/deep learning/signal processing

分类

信息技术与安全科学

引用本文复制引用

黄志勐,高峰,杨帆,马思伟..面向机器视觉的文本提示引导的图像编码[J].电子学报,2026,54(1):19-31,13.

基金项目

国家自然科学基金(No.62025101,No.62176006) (No.62025101,No.62176006)

中国博士后科学基金(No.2025M771511) National Natural Science Foundation of China(No.62025101,No.62176006) (No.2025M771511)

China Postdoc-toral Science Foundation(No.2025M771511) (No.2025M771511)

电子学报

0372-2112

访问量0
|
下载量0
段落导航相关论文