| 注册
首页|期刊导航|电子学报|基于对抗混合专家后训练机制的鲁棒AI生成图像检测方法

基于对抗混合专家后训练机制的鲁棒AI生成图像检测方法

张睿萱 刁云峰 陆智远 夏海峰 郭治卿 郝孝帅 汪萌

电子学报2026,Vol.54Issue(3):1178-1193,16.
电子学报2026,Vol.54Issue(3):1178-1193,16.DOI:10.12263/DZXB.20251196

基于对抗混合专家后训练机制的鲁棒AI生成图像检测方法

Adversarial Mixture of Experts Post-Training for Robust AI-Generated Image Detection

张睿萱 1刁云峰 1陆智远 1夏海峰 2郭治卿 3郝孝帅 4汪萌1

作者信息

  • 1. 合肥工业大学计算机与信息学院,安徽 合肥 230601
  • 2. 中山大学网络空间安全学院,广东 深圳 518107
  • 3. 新疆大学计算机科学与技术学院,新疆 乌鲁木齐 830017
  • 4. 小米汽车,北京 100085
  • 折叠

摘要

Abstract

AI-generated imagery(AIGI)technology has enabled the automated production of high-quality visual con-tent,demonstrating enormous application potential in fields such as artistic creation,digital entertainment,and virtual reali-ty.However,while empowering content production,this technology also brings serious security and ethical challenges.Gen-erative models can be maliciously used to forge real people or events,thereby creating false information,spreading deep-fake content,and even interfering with online public opinion.Therefore,how to effectively identify AI-generated images(AIGI detection)has become an important research topic for ensuring the credibility of digital content and maintaining cy-berspace security.However,existing AIGI detectors generally exhibit insufficient robustness against adversarial attacks.At-tackers only need to add subtle adversarial perturbations imperceptible to the human eye to the synthesized image to bypass detection,causing the synthesized content to be misclassified as a real image,and defense mechanisms against such attacks are still scarce.To address this issue,this paper first systematically evaluates the effectiveness of adversarial training in AI-GI detection tasks.Theoretical analysis and experimental results show that it is prone to inducing feature entanglement dur-ing training,leading to severe degradation or even collapse of detection performance.Therefore,there is an urgent need to develop a dedicated adversarial defense method effective for AIGI detection tasks.Unlike feature entanglement that occurs in adversarial training,this paper finds that adversarial perturbations in standard-trained detectors cause adversarial exam-ples to deviate significantly from clean examples in the feature space,resulting in significant separability.Based on this ob-servation,this paper proposes a strategy of modeling adversarial examples as independent categories and constructs a post-training defense framework:while keeping the pre-trained feature extractor fixed,it only learns new classification boundar-ies to fit the feature distribution of adversarial examples.To enhance the model's generalization ability to unknown attacks,this paper further proposes an adversarial hybrid expert post-training mechanism.This mechanism utilizes multiple expert modules to learn feature patterns for specific attack types and introduces shared experts to capture common representations among different attacks,thereby achieving efficient modeling and robust identification of multiple classes of adversarial ex-amples.Experimental results show that on mainstream AIGI datasets such as ProGAN and Stable Diffusion,facing various typical adversarial attack methods,the average adversarial accuracy is improved by 18.92%and 12.56%respectively com-pared to existing mainstream defense methods without sacrificing the detection accuracy of benign examples,demonstrating good practicality and application potential in real-world security scenarios.

关键词

AI生成图像检测/对抗样本/对抗攻击/对抗防御/混合专家/后训练策略

Key words

AI-generated image detection/adversarial example/adversarial attacks/adversarial defense/mixture of experts/post-training strategy

分类

信息技术与安全科学

引用本文复制引用

张睿萱,刁云峰,陆智远,夏海峰,郭治卿,郝孝帅,汪萌..基于对抗混合专家后训练机制的鲁棒AI生成图像检测方法[J].电子学报,2026,54(3):1178-1193,16.

基金项目

国家自然科学基金(No.62302139,No.62406068,No.62302427) (No.62302139,No.62406068,No.62302427)

中央高校基本科研业务费专项资金项目(No.JZ2025HGTB0227) National Natural Science Foundation of China(No.62302139,No.62406068,No.62302427) (No.JZ2025HGTB0227)

Fundamental Research Funds for the Central Universities(No.JZ2025HGTB0227) (No.JZ2025HGTB0227)

电子学报

0372-2112

访问量0
|
下载量0
段落导航相关论文