计算机科学 ›› 2026, Vol. 53 ›› Issue (8): 437-445.doi: 10.11896/jsjkx.250600213

• 信息安全 • 上一篇    下一篇

融合动态可组合多头注意力的深度学习侧信道分析方法研究

蒋玲腊, 陈文, 孙伟, 赵奎   

  1. 四川大学网络空间安全学院 成都 610065
  • 收稿日期:2025-06-26 修回日期:2025-11-28 发布日期:2026-08-17
  • 通讯作者: 赵奎(zhaokui@scu.edu.cn)
  • 作者简介:(2279787117@qq.com)
  • 基金资助:
    国家重点研发计划(020YFB1805405)

Research on Deep Learning-based Side-channel Analysis Method with Dynamically ComposableMulti-head Attention

JIANG Lingla, CHEN Wen, SUN Wei, ZHAO Kui   

  1. School of Cyber Science and Engineering, Sichuan University, Chengdu 610065, China
  • Received:2025-06-26 Revised:2025-11-28 Online:2026-08-17
  • About author:JIANG Lingla,born in 2001,master.Her main research interests include cryptographic application security and side-channel analysis.
    ZHAO Kui,born in 1972,Ph.D,professor,master’s supervisor.His main research interests include network and information security,disaster backup and recovery,big data analysis and mining.
  • Supported by:
    National Key Research and Development Program of China(020YFB1805405).

摘要: 侧信道分析面临着从大量的功耗轨迹样本中提取密钥信息的特征这一难题。多头注意力机制(Multi-Head Attention,MHA)的多个头能够分别捕捉数据间的局部依赖和全局关联关系,具有较强的多特征学习能力。因此,近年来,MHA被广泛应用于侧信道分析的特征自动提取过程。然而,多头注意力机制在各个头同时进行特征学习时,容易受到低秩瓶颈的限制和冗余头的影响。该问题削弱了MHA捕捉数据中复杂时序依赖和全局特征关系的能力,进而导致训练过程缺乏稳定性,难以保证收敛至最优解。针对上述问题,提出了一种基于动态可组合多头注意力机制的深度侧信道分析方法。该方法引入动态可组合多头注意力机制,自适应组合不同注意力头的信息,以有效增强模型对关键信息的特征提取能力,确保训练过程的稳定性,从而持续提升攻击效果。在公开的ASCAD,AES_HD和CHES18数据集上进行了对比实验,结果表明,所提方法在模型训练稳定性和攻击效果方面均优于现有模型。例如,在AES_HD和CHES18数据集上,训练得到的模型在攻击时所需的功耗轨迹数量分别减少了56.2%和66.7%。

关键词: 侧信道分析, 深度学习, 特征提取, 多头注意力, 动态可组合多头注意力

Abstract: Side-channel analysis faces the challenge of extracting key-related features from a large number of power traces.Since MHA(Multi-Head Attention) mechanism enables multiple heads to capture both local dependencies and global correlations in data,it has strong multi-feature learning capabilities.Therefore,MHA has been widely applied to automatic feature extraction in side-channel analysis in recent years.However,when multiple heads learn features simultaneously,MHA is prone to the limitations of low-rank bottlenecks and redundant heads.This problem weakens the ability of MHA to capture complex temporal dependencies and global feature relationships in data,leading to instability during training and difficulty in converging to the optimal solution.To address this issue,this paper proposes a deep side-channel analysis method based on dynamically composable multi-head attention.The method introduces a dynamically composable multi-head attention mechanism that adaptively combines information from different attention heads,thereby effectively enhancing the model’s capability to extract key features,ensuring training stability,and continuously improving attack performance.Comparative experiments conducted on the public ASCAD,AES_HD,and CHES18 datasets demonstrate that the proposed method outperforms existing models in both training stability and attack effectiveness.For example,on the AES_HD and CHES18 datasets,the number of power traces required for successful attacks is reduced by 56.2% and 66.7%,respectively.

Key words: Side-channel analysi, Deep learning, Feature extraction, Multi-head attention, Dynamically composable multi-head attention

中图分类号: 

  • TP309
[1] KOCHER P,JAFFE J,JUN B.Differential power analysis[C]//Annual International Cryptology Conference.Berlin,Heidelberg:Springer Berlin Heidelberg,1999:388-397.
[2] GANDOLFI K,MOURTEL C,OLIVIER F.Electromagnetic analysis:Concrete results[C]//International Workshop on Cryptographic Hardware and Embedded Systems.Berlin:Springer,2001:251-261.
[3] EMMANUEL P,REMI S,RYAD B,et al.Study of deep learning techniques for side-channel analysis and introduction to ascad database[J].CoRR,2018,53:1-45.
[4] BRIER E,CLAVIER C,OLIVIER F.Correlation power analysis with a leakage model[C]//International Workshop on Cryptographic Hardware and Embedded Systems.Berlin,Heidelberg:Springer Berlin Heidelberg,2004:16-29.
[5] CHARI S,RAO J R,ROHATGI P.Template attacks[C]//International Workshop on Cryptographic Hardware and Embedded Systems.Berlin:Springer,2002:13-28.
[6] EL AABID M A,GUILLEY S,HOOGVORST P.Template at-tacks with a power model[J].IACR Cryptology ePrint Archive,2007,2007:443.
[7] VASWANI A,SHAZEER N,PARMAR N,et al.Attention is all you need[J].Advances in Neural Information Processing Systems,2017,30.
[8] HAJRA S,ALAM M,SAHA S,et al.On the Instability of Softmax Attention-Based Deep Learning Models in Side-Channel Analysis[J].IEEE Transactions on Information Forensics and Security,2023,19:514-528.
[9] BHOJANAPALLI S,YUN C,RAWAT A S,et al.Low-rankbottleneck in multi-head attention models[C]//International Conference on Machine Learning.PMLR,2020:864-873.
[10] VOITA E,TALBOT D,MOISEEV F,et al.Analyzing multi-head self-attention:Specialized heads do the heavy lifting,the rest can be pruned[J].arXiv:1905.09418,2019.
[11] CHARI S,JUTLA C S,RAO J R,et al.Towards sound approaches to counteract power-analysis attacks[C]//Annual International Cryptology Conference.Berlin:Springer,1999:398-412.
[12] AKKAR M L,GIRAUD C.An implementation of DES andAES,secure against some attacks[C]//International Workshop on Cryptographic Hardware and Embedded Systems.Berlin:Springer,2001:309-318.
[13] NASSAR M,SOUISSI Y,GUILLEY S,et al.RSM:A small and fast countermeasure for AES,secure against 1st and 2nd-order zero-offset SCAs[C]//2012 Design,Automation & Test in Europe Conference & Exhibition(DATE).IEEE,2012:1173-1178.
[14] CORON J S,KIZHVATOV I.An efficient method for random delay generation in embedded software[C]//International Workshop on Cryptographic Hardware and Embedded Systems.Berlin:Springer,2009:156-170.
[15] DIFFIE-HELLMAN RS A.Timing Attacks on Implementationsof Advances in Cryptology-CRYPTO’96[C]//16th Annual International Cryptology Conference.Springer,1996,1109:104.
[16] PERIN G,WU L,PICEKS.Exploring feature selection scenarios for deep learning-based side-channel analysis[J].IACR Transactions on Cryptographic Hardware and Embedded Systems,2022,2022:828-861.
[17] CHOUDARY O,KUHN M G.Efficient template attacks[C]//International Conference on Smart Card Research and Advanced Applications.Cham:Springer International Publishing,2013:253-270.
[18] SCHINDLER W,LEMKE K,PAAR C.A stochastic model for differential side channel cryptanalysis[C]//International Workshop on Cryptographic Hardware and Embedded Systems.Berlin:Springer,2005:30-46.
[19] LERMAN L,BONTEMPI G,MARKOWITCH O.A machinelearning approach against a masked AES:Reaching the limit of side-channel attacks with a learning model[J].Journal of Cryptographic Engineering,2015,5(2):123-139.
[20] HOSPODAR G,GIERLICHS B,DE MULDER E,et al.Ma-chine learning in side-channel analysis:a first study[J].Journal of Cryptographic Engineering,2011,1(4):293-302.
[21] HEUSER A,ZOHNER M.Intelligent machine homicide:Breaking cryptographic devices using support vector machines[C]//International Workshop on Constructive Side-Channel Analysis and Secure Design.Berlin:Springer,2012:249-264.
[22] PICEK S,HEUSER A,GUILLEY S.Template attack versusBayes classifier[J].Journal of Cryptographic Engineering,2017,7(4):343-351.
[23] MARTINASEK Z,ZEMAN V.Innovative method of the power analysis[J].Radioengineering,2013,22(2):586-594.
[24] MAGHREBI H,PORTIGLIATTI T,PROUFF E.Breakingcryptographic implementations using deep learning techniques[C]//International Conference on Security,Privacy,and Applied Cryptography Engineering.Cham:Springer International Publishing,2016:3-26.
[25] ZAID G,BOSSUET L,HABRARD A,et al.Methodology for efficient CNN architectures in profiling attacks[J].IACR Transactions on Cryptographic Hardware and Embedded Systems,2020,2020(1):1-36.
[26] LU X,ZHANG C,CAO P,et al.Pay attention to raw traces:A deep learning architecture for end-to-end profiling attacks[J].IACR Transactions on Cryptographic Hardware and Embedded Systems,2021,2021:235-274.
[27] CAGLI E,DUMAS C,PROUFF E.Convolutional neural net-works with data augmentation against jitter-based countermeasures:Profiling attacks without pre-processing[C]//International Conference on Cryptographic Hardware and Embedded Systems.Cham:Springer International Publishing,2017:45-68.
[28] BAHDANAU D.Neural machine translation by jointly learning to align and translate[J].arXiv:1409.0473,2014.
[29] LU X,ZHANG C,GU D.Attention-based non-profiled side-channel attack[C]//2021 Asian Hardware Oriented Security and Trust Symposium(AsianHOST).IEEE,2021:1-6.
[30] MANGARD S,OSWALD E,POPP T.Energy Analysis Attack[M].冯登国,周永彬,刘继业,等,译.Beijing:Science Press,2010.
[31] MICHEL P,LEVY O,NEUBIG G.Are sixteen heads really better than one?[J].Advances in Neural Information Processing Systems,2019,32.
[32] XIAO D,MENG Q,LI S,et al.Improving transformers with dynamically composable multi-head attention[J].arXiv:2405.08553,2024.
[33] ZHANG L,XING X,FAN J,et al.Multilabel deep learning-based side-channel attack[J].IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems,2020,40(6):1207-1216.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!