高技术通讯2026,Vol.36Issue(4):331-339,9.DOI:10.3772/j.issn.1002-0470.2026.04.001
推理时间约束的结构化剪枝
Inference time constrained structured pruning
摘要
Abstract
Structured pruning compresses and accelerates neural networks by removing groups of weights in a structured manner.Most existing pruning methods target a predefined sparsity level,i.e.,pruning a fixed proportion of weights,rather than directly optimizing inference latency.However,the relationship between sparsity and inference time is highly nonlinear,making sparsity-targeted pruning methods unsuitable for deployment scenarios with explicit latency constraints.To address this issue,we propose a novel structured pruning method,termed inference time constrained pruning(ITCP),which automatically searches for a pruning scheme that satisfies a desired inference-time budget while minimizing accuracy degradation.Specifically,ITCP formulates latency-constrained pruning as a constrained optimization problem,where the objective is to maximize a performance score under a given inference-time constraint,and solves it efficiently using dynamic programming.In addition,a performance model is devel-oped to rapidly estimate the inference time of models at different sparsity levels.Experimental results on CIFAR-10,CIFAR-100,and ImageNet demonstrate that,under the same acceleration requirements,ITCP consistently achieves higher accuracy than baseline pruning strategies.关键词
模型剪枝/模型压缩/性能模型/时间约束Key words
model pruning/model compression/performance models/time constraints引用本文复制引用
李晨昊,李琳,邱强,张志斌,郭嘉丰,程学旗..推理时间约束的结构化剪枝[J].高技术通讯,2026,36(4):331-339,9.基金项目
广东省科技计划(2023A1111120017),北京市科技新星计划(Z211100002121141)和国防基础科研(JCKY2022130C039)资助项目. (2023A1111120017)