| 注册
首页|期刊导航|硅酸盐学报|低碳熟料强度预测的数据增强建模及模型可解释性分析

低碳熟料强度预测的数据增强建模及模型可解释性分析

翟牧楠 郅晓 叶家元 任雪红 吴春丽 张洪滔 崔文娟 颜景华 张文生

硅酸盐学报2026,Vol.54Issue(8):2627-2643,17.
硅酸盐学报2026,Vol.54Issue(8):2627-2643,17.DOI:10.14062/j.issn.0454-5648.20260127

低碳熟料强度预测的数据增强建模及模型可解释性分析

Data-Augmented Modeling and Model Interpretability Analysis for Strength Prediction of Low-Carbon Clinker

翟牧楠 1郅晓 1叶家元 1任雪红 1吴春丽 2张洪滔 1崔文娟 1颜景华 1张文生1

作者信息

  • 1. 中国建筑材料科学研究总院有限公司,北京 100024
  • 2. 中存大数据科技有限公司,北京 100024
  • 折叠

摘要

Abstract

Introduction The clinker calcination process involves high-temperature solid/liquid reactions,and the controllable preparation of multiple clinker phase compositions,which further affect subsequent synergistic hydration reactions.The relationship between composition and 28 d compressive strength is not simply linear,but rather a complex mapping governed by the coupling of multiple processes.Conventional point-by-point experimental screening makes it difficult to cover a sufficiently broad compositional space within a limited timeframe.This restricts the exploration efficiency of novel low-carbon clinker.Machine learning models can be used to learn the composition and property relationships between clinker phase and strength,and to support the screening of candidate clinker compositions.However,under small data conditions,these models are susceptible to insufficient training distribution coverage,resulting in a unstable cross-system predictive performance.Therefore,this study was to develop a self-iterative data augmentation framework to address the above-mentioned issues.This framework was to expand effective supervised information to optimize the model's predictive performance in sparsely sampled compositional regions.Subsequently,model interpretability analysis was conducted to identify the key clinker phases affecting 28 d compressive strength prediction and their response patterns.This could provide a testable model basis for the optimization of low-carbon clinker compositions. Methods This study defined each simulated clinker as one sample.The dataset contained 264 simulated clinker samples,with the compressive strengths ranging from 0 to 70.3 MPa.The model inputs consisted of 12 features,including 11 clinker phases and gypsum,while the output was the 28 d compressive strength.In the self-iterative data augmentation strategy,an artificial neural network(ANN)model was first trained using the experimental data.Virtual data were then generated under compositional constraints,and their 28 d compressive strengths were predicted by the model from the previous iteration to assign pseudo-labels.The virtual and experimental samples were subsequently combined for training,and the surrogate model was updated iteratively.Two sampling patterns were adopted for virtual sample generation,i.e.,1)Pattern A:dense sampling within the experimental samples'compositional systems,and 2)Pattern B:generating virtual samples at a ratio of 1:1 from the in-and out-of-system regions of the experimental samples'compositional space.Before training,36 experimental samples were set aside as a test set,which remained isolated throughout training,iteration and hyperparameter selection.The validation set consisted exclusively of real experimental samples,whereas virtual samples were included only in the training set.Model hyperparameters were determined based on the Bayesian optimization,with the validation-set mean squared error(MSE)as the selection criterion.Finally,the model predictive performance was evaluated by MSE,the coefficient of determination(R2),mean absolute error(MAE),and mean absolute percentage error(MAPE).After the final surrogate model was determined,multilevel model interpretation was further performed via combining constrained permutation feature importance,accumulated local effects and SHapley additive explanations(SHAP). Results and discussion The self-iterative augmentation strategy effectively fills sparsely sampled regions,while maintaining feasible-domain constraints.This provides richer training information for subsequent accuracy evaluation and interpretability analysis of the surrogate model.Meanwhile,the data augmentation process does not generate a large number of near-duplicate samples,and the training set retains the necessary heterogeneity.The inclusion of virtual samples reduces the model's sensitivity to randomness in data partitioning.Moreover,the overall test set error and the in-/out-of-system errors decrease progressively over successive iterations,indicating an improved cross-system generalization performance.Pattern B in Iteration II achieves a favorable balance between predictive accuracy and extrapolation stability.Its overall test set MSE is 9.8,with in-system and out-of-system MSE values of 6.9 and 11.3,respectively.The R2,MAE,and MAPE are 0.79,2.36 and 5.11%,respectively.These results indicate that,under the current small data conditions,data augmentation can substantially improve the model's ability to identify performance trends.Subsequent interpretability analysis of the optimal surrogate model shows that different clinker phases can be classified as positive or negative driving phases.The analysis further indicates that when C3A content exceeds approximately 10%,its positive driving effect on strength decreases rapidly.Although CA and C12A7 act as positive driving phases,their benefits diminish when their contents exceed 40%,and even shift to strongly negative responses. Conclusions The self-iterative data augmentation framework proposed could expand a supervised training information for low-carbon clinker systems characterized by small datasets and sparse compositional spaces.The evaluation bias introduced by pseudo-labels could be also suppressed under a rigorous evaluation strategy.The results showed that data augmentation significantly improved the prediction accuracy and cross-system generalization capability of the ANN surrogate model for 28 d compressive strength.In Iteration II,Pattern B achieved the optimum overall performance under the current data conditions.Although further iterations reduced some in-system errors,an error rebound was observed for out-of-system samples.This indicated that diminishing marginal information could gain from pseudo-labels constrained the model's generalization performance.The combined use of constrained PFI,ALE,and SHAP revealed the decision-making basis of the surrogate model in terms of global sensitivity,average response trends,and sample-level contributions.The results indicated that C4A3$ and CA were the main positive driving clinker phases for compressive strength enhancement,whereas C3A,C5S2$ and C2S were the main negative factors.Moreover,C12A7 exhibited a pronounced threshold effect.These findings could provide a testable reference for subsequent composition screening,constraint-boundary definition and inverse design of low-carbon clinker.

关键词

低碳熟料/28 d抗压强度/预测模型/自迭代/数据增强/可解释性

Key words

low-carbon clinker/28 d compressive strength/prediction model/self-iterative/data augmentation/model interpretability

分类

化学化工

引用本文复制引用

翟牧楠,郅晓,叶家元,任雪红,吴春丽,张洪滔,崔文娟,颜景华,张文生..低碳熟料强度预测的数据增强建模及模型可解释性分析[J].硅酸盐学报,2026,54(8):2627-2643,17.

基金项目

"十四五"国家重点研发计划(2022YFC3803101) (2022YFC3803101)

国家自然科学基金双碳专项项目(52341202). (52341202)

硅酸盐学报

0454-5648

访问量0
|
下载量0
段落导航相关论文