四川大学学报(自然科学版)2026,Vol.63Issue(2):241-258,18.DOI:10.19907/j.0490-6756.250250
面向优化的大语言模型黑盒越狱攻击研究综述
A review of black-box jailbreak attacks on large language models for optimization
摘要
Abstract
Large Language Models(LLMs)have demonstrated remarkable capabilities in Natural Language Processing(NLP).However,their security vulnerabilities,particularly jailbreak attacks,pose a critical chal-lenge.These attacks circumvent safety alignment mechanisms through carefully crafted adversarial prompts,revealing the limitations of alignment techniques like Reinforcement Learning from Human Feedback(RLHF).Template-based or manually crafted jailbreak methods typically exhibit low success rates and poor generalization,and they rapidly become obsolete as LLM safety mechanisms evolve.In contrast,optimization-based methods automatically generate adversarial prompts,leading to higher success rates and better stealthiness that effectively bypass common detection mechanisms.To overcome the limitations of white-box attacks,such as their reliance on gradients and limited transferability,the review investigates the black-box optimization paradigm and presents the first systematic taxonomy of jailbreak methods:Genetic Al-gorithm(GA)-based,Reinforcement Learning(RL)-based,Fuzzing-based,and LLM-based Adversarial Optimization.We delve into the core mechanisms,technical strengths,and limitations of each category.The primary contribution of this survey is proposing a novel taxonomy that critically examines existing defenses'deficiencies in real-time performance,generalizability,and attack-defense balance.It further advocates for dy-namic defense architectures and standardized benchmarks,thereby providing a theoretical foundation and prac-tical guidance for balancing the security and performance of LLMs in adversarial settings.关键词
大语言模型/优化/越狱攻击/越狱防御Key words
LLMs/optimization/jailbreak attack/jailbreak defence分类
信息技术与安全科学引用本文复制引用
陶佳玲,黄松,高心怡,方勇,曲豫宾,李瑞阳,陆江涛..面向优化的大语言模型黑盒越狱攻击研究综述[J].四川大学学报(自然科学版),2026,63(2):241-258,18.基金项目
国家自然科学基金(U24B20147) (U24B20147)