电子学报2026,Vol.54Issue(3):1178-1193,16.DOI:10.12263/DZXB.20251196
基于对抗混合专家后训练机制的鲁棒AI生成图像检测方法
Adversarial Mixture of Experts Post-Training for Robust AI-Generated Image Detection
摘要
Abstract
AI-generated imagery(AIGI)technology has enabled the automated production of high-quality visual con-tent,demonstrating enormous application potential in fields such as artistic creation,digital entertainment,and virtual reali-ty.However,while empowering content production,this technology also brings serious security and ethical challenges.Gen-erative models can be maliciously used to forge real people or events,thereby creating false information,spreading deep-fake content,and even interfering with online public opinion.Therefore,how to effectively identify AI-generated images(AIGI detection)has become an important research topic for ensuring the credibility of digital content and maintaining cy-berspace security.However,existing AIGI detectors generally exhibit insufficient robustness against adversarial attacks.At-tackers only need to add subtle adversarial perturbations imperceptible to the human eye to the synthesized image to bypass detection,causing the synthesized content to be misclassified as a real image,and defense mechanisms against such attacks are still scarce.To address this issue,this paper first systematically evaluates the effectiveness of adversarial training in AI-GI detection tasks.Theoretical analysis and experimental results show that it is prone to inducing feature entanglement dur-ing training,leading to severe degradation or even collapse of detection performance.Therefore,there is an urgent need to develop a dedicated adversarial defense method effective for AIGI detection tasks.Unlike feature entanglement that occurs in adversarial training,this paper finds that adversarial perturbations in standard-trained detectors cause adversarial exam-ples to deviate significantly from clean examples in the feature space,resulting in significant separability.Based on this ob-servation,this paper proposes a strategy of modeling adversarial examples as independent categories and constructs a post-training defense framework:while keeping the pre-trained feature extractor fixed,it only learns new classification boundar-ies to fit the feature distribution of adversarial examples.To enhance the model's generalization ability to unknown attacks,this paper further proposes an adversarial hybrid expert post-training mechanism.This mechanism utilizes multiple expert modules to learn feature patterns for specific attack types and introduces shared experts to capture common representations among different attacks,thereby achieving efficient modeling and robust identification of multiple classes of adversarial ex-amples.Experimental results show that on mainstream AIGI datasets such as ProGAN and Stable Diffusion,facing various typical adversarial attack methods,the average adversarial accuracy is improved by 18.92%and 12.56%respectively com-pared to existing mainstream defense methods without sacrificing the detection accuracy of benign examples,demonstrating good practicality and application potential in real-world security scenarios.关键词
AI生成图像检测/对抗样本/对抗攻击/对抗防御/混合专家/后训练策略Key words
AI-generated image detection/adversarial example/adversarial attacks/adversarial defense/mixture of experts/post-training strategy分类
信息技术与安全科学引用本文复制引用
张睿萱,刁云峰,陆智远,夏海峰,郭治卿,郝孝帅,汪萌..基于对抗混合专家后训练机制的鲁棒AI生成图像检测方法[J].电子学报,2026,54(3):1178-1193,16.基金项目
国家自然科学基金(No.62302139,No.62406068,No.62302427) (No.62302139,No.62406068,No.62302427)
中央高校基本科研业务费专项资金项目(No.JZ2025HGTB0227) National Natural Science Foundation of China(No.62302139,No.62406068,No.62302427) (No.JZ2025HGTB0227)
Fundamental Research Funds for the Central Universities(No.JZ2025HGTB0227) (No.JZ2025HGTB0227)