| 注册
首页|期刊导航|四川大学学报(自然科学版)|面向优化的大语言模型黑盒越狱攻击研究综述

面向优化的大语言模型黑盒越狱攻击研究综述

陶佳玲 黄松 高心怡 方勇 曲豫宾 李瑞阳 陆江涛

四川大学学报(自然科学版)2026,Vol.63Issue(2):241-258,18.
四川大学学报(自然科学版)2026,Vol.63Issue(2):241-258,18.DOI:10.19907/j.0490-6756.250250

面向优化的大语言模型黑盒越狱攻击研究综述

A review of black-box jailbreak attacks on large language models for optimization

陶佳玲 1黄松 1高心怡 2方勇 2曲豫宾 3李瑞阳 1陆江涛1

作者信息

  • 1. 中国人民解放军陆军工程大学指挥控制工程学院,南京 210001
  • 2. 四川大学网络空间安全学院,成都 610065
  • 3. 中国人民解放军陆军工程大学指挥控制工程学院,南京 210001||江苏工程职业技术学院信息工程学院,南通 226001
  • 折叠

摘要

Abstract

Large Language Models(LLMs)have demonstrated remarkable capabilities in Natural Language Processing(NLP).However,their security vulnerabilities,particularly jailbreak attacks,pose a critical chal-lenge.These attacks circumvent safety alignment mechanisms through carefully crafted adversarial prompts,revealing the limitations of alignment techniques like Reinforcement Learning from Human Feedback(RLHF).Template-based or manually crafted jailbreak methods typically exhibit low success rates and poor generalization,and they rapidly become obsolete as LLM safety mechanisms evolve.In contrast,optimization-based methods automatically generate adversarial prompts,leading to higher success rates and better stealthiness that effectively bypass common detection mechanisms.To overcome the limitations of white-box attacks,such as their reliance on gradients and limited transferability,the review investigates the black-box optimization paradigm and presents the first systematic taxonomy of jailbreak methods:Genetic Al-gorithm(GA)-based,Reinforcement Learning(RL)-based,Fuzzing-based,and LLM-based Adversarial Optimization.We delve into the core mechanisms,technical strengths,and limitations of each category.The primary contribution of this survey is proposing a novel taxonomy that critically examines existing defenses'deficiencies in real-time performance,generalizability,and attack-defense balance.It further advocates for dy-namic defense architectures and standardized benchmarks,thereby providing a theoretical foundation and prac-tical guidance for balancing the security and performance of LLMs in adversarial settings.

关键词

大语言模型/优化/越狱攻击/越狱防御

Key words

LLMs/optimization/jailbreak attack/jailbreak defence

分类

信息技术与安全科学

引用本文复制引用

陶佳玲,黄松,高心怡,方勇,曲豫宾,李瑞阳,陆江涛..面向优化的大语言模型黑盒越狱攻击研究综述[J].四川大学学报(自然科学版),2026,63(2):241-258,18.

基金项目

国家自然科学基金(U24B20147) (U24B20147)

四川大学学报(自然科学版)

0490-6756

访问量0
|
下载量0
段落导航相关论文