计算机科学 ›› 2026, Vol. 53 ›› Issue (7): 139-145.doi: 10.11896/jsjkx.250600038

• 人工智能 • 上一篇    下一篇

基于门控代理注意力机制的软提示迁移方法

张焱, 周剑, 韩磊, 程春玲   

  1. 南京邮电大学计算机学院、软件学院、网络空间安全学院 南京 210023
  • 收稿日期:2025-06-05 修回日期:2025-08-12 出版日期:2026-07-15 发布日期:2026-07-10
  • 通讯作者: 程春玲(chengcl@njupt.edu.cn)
  • 作者简介:(1223045241@njupt.edu.cn)
  • 基金资助:
    国家自然科学基金(62472232)

Gate-controlled Agent Attention Mechanism-based Soft Prompt Transfer Method

ZHANG Yan, ZHOU Jian, HAN Lei, CHENG Chunling   

  1. School of Computer Science,Nanjing University of Posts and Telecommunications,Nanjing 210023,China
  • Received:2025-06-05 Revised:2025-08-12 Published:2026-07-15 Online:2026-07-10
  • About author:ZHANG Yan,born in 2002,postgra-duate.His main research interests include deep learning and prompt lear-ning.
    CHENG Chunling,born in 1972,professor.Her main research interests include data mining and data management.
  • Supported by:
    National Natural Science Foundation of China(62472232).

摘要: 软提示迁移作为一种跨任务学习方法,通过学习源提示中的知识来引导目标任务模型学习更具泛化能力的提示向量。而现有的软提示迁移方法忽略了目标任务样本中的特有信息对模型训练的影响,导致目标提示训练偏移。对此,提出一种基于门控代理注意力机制的软提示迁移方法。首先,为分离和增强目标样本的全局和特有信息,提出双路径特征提取方法,通过两种轻量级特征提取路径,从目标样本中分别提取融合任务级语义信息的代理令牌和样本局部特征;其次,为衡量源提示的可迁移性,提出门控代理注意力机制,利用代理令牌和局部特征分别与提示特征计算注意力分布以表征任务相似度,同时,为保留样本特有信息的表达,利用门控单元过滤样本局部特征并整合两类注意力;最后,提出注意力扰动方法,利用源提示在目标任务上的预测分布损失来扰动注意力分布,从而降低那些在表示空间中与目标任务相似但实际预测效果较差的源提示在迁移过程中的权重。在公开数据集GLUE上与9个基线模型相比,所提方法在平均性能上表现最佳,验证了其有效性。

关键词: 提示微调, 软提示迁移, 特征提取, 门控代理注意力机制, 注意力分布扰动

Abstract: Soft prompt transfer,as a cross-task learning approach,aims to guide the target task model to learn more generalizable prompt representations by leveraging knowledge from source prompts.However,existing methods often ignore the impact of task-specific information embedded in target samples,which may result in biased target prompt training.To address this issue,a gate-controlled agent attention mechanism-based soft prompt transfer method(GPAPT) is proposed.Firstly,to separate and enhance both global and task-specific information in target samples,a dual-path feature extraction strategy is introduced.It employs two lightweight extraction paths to derive agent tokens that encode task-level semantics and local features that capture instance-specific characteristics.Secondly,to assess the transferability of source prompts,a gate-controlled agent attention mechanism is presented.It computes attention distributions between prompt features and both agent tokens and local features to model task similarity.To preserve task-specific information,a gating unit filters local features and integrates the two types of attention.Finally,a perturbation mechanism for attention distribution is introduced.It perturbs the attention weights based on the prediction loss of source prompts on the target task,thereby reduces the influence of source prompts that are similar in the representation space but perform poorly in prediction,and thus improving transfer robustness.Extensive experiments on the GLUE benchmark demonstrate that GPAPT achieves the best average performance compared to nine strong baseline methods,validating the effectiveness of the proposed approach.

Key words: Prompt tuning, Soft prompt transfer, Feature extraction, Gate-controlled agent attention mechanism, Attention distribution perturbation

中图分类号: 

  • TP391.1
[1]RAFFEL C,SHAZEER N,ROBERTS A,et al.Exploring the limits of transfer learning with a unified text-to-text transformer[J].Journal of Machine Learning Research,2020,21(140):1-67.
[2]HOULSBY N,GIURGIU A,JASTRZEBSKI S,et al.Parame-ter-efficient transfer learning for NLP[C]//Proceedings of the 36th International Conference on Machine Learning.2019,97:2790-2799.
[3]LIU P,YUAN W,FU J,et al.Pre-train,prompt,and predict:A systematic survey of prompting methods in natural language processing[J].ACM Computing Surveys,2023,55(9):1-35.
[4]BEN ZAKEN E,GOLDBERG Y,RAVFOGEL S.BitFit:Simpleparameter-efficient fine-tuning for transformer-based masked language models[C]//Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics.2022:1-9.
[5]LESTER B,AL-RFOU R,CONSTANT N.The power of scale for parameter-efficient prompt tuning[C]//Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing(EMNLP).2021:3045-3059.
[6]VU T,LESTER B,CONSTANT N,et al.SPoT:Better frozenmodel adaptation through soft prompt transfer[C]//Procee-dings of the 60th Annual Meeting of the Association for Computational Linguistics.2022:5039-5059.
[7]WANG A,SINGH A,MICHAEL J,et al.GLUE:A multi-task benchmark and analysis platform for natural language understanding[C]//Proceedings of the 2018 EMNLP Workshop BlackboxNLP:Analyzing and Interpreting Neural Networks for NLP.2018:353-355.
[8]HAN D,YE T,HAN Y,et al.Agent attention:On the integration of softmax and linear attention [C]//European Conference on Computer Vision.Cham:Springer,2024:124-140.
[9]WANG Z,PANDA R,KARLINSKY L,et al.Multitask prompt tuning enables parameter-efficient transfer learning[C]//The Eleventh International Conference on Learning Representations(ICLR).2023.
[10]ASAI A,SALEHI M,PETERS M,et al.ATTEMPT:Parameter-efficient multi-task tuning via attentional mixtures of soft prompts[C]//Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.2022:6655-6672.
[11]WU X,CHEN C L P,LI S,et al.Snapshot prompt ensemble for parameter-efficient soft prompt transfer[C]//ICASSP 2024-2024 IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP).2024:11946-11950.
[12]WU M,LIU W,XU J,et al.Parameter-efficient multi-task fine-tuning by learning to transfer token-wise prompts[C]//Findings of the Association for Computational Linguistics:EMNLP 2023.2023:8734-8746.
[13]ZHONG Q,DING L,LIU J,et al.Panda:Prompt transfer meets knowledge distillation for efficient model adaptation[J].IEEE Transactions on Knowledge and Data Engineering,2024,36:4835-4848.
[14]WANG A,PRUKSACHATKUN Y,NANGIA N,et al.Super GLUE:A stickier benchmark for general-purpose language understanding systems[C]//Proceedings of the 33rd International Conference on Neural Information Processing Systems.2019:3266-3280.
[15]KHOT T,SABHARWAL A,CLARK P.Sci Tail:A textual entailment dataset from science question answering [C]//Procee-dings of AAAI.2018:5189-5197.
[16]MAHABADI R K,RUDER S,DEHGHANI M,et al.Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks [C]//Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing.2021:565-576.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!