计算机科学 ›› 2026, Vol. 53 ›› Issue (7): 139-145.doi: 10.11896/jsjkx.250600038
张焱, 周剑, 韩磊, 程春玲
ZHANG Yan, ZHOU Jian, HAN Lei, CHENG Chunling
摘要: 软提示迁移作为一种跨任务学习方法,通过学习源提示中的知识来引导目标任务模型学习更具泛化能力的提示向量。而现有的软提示迁移方法忽略了目标任务样本中的特有信息对模型训练的影响,导致目标提示训练偏移。对此,提出一种基于门控代理注意力机制的软提示迁移方法。首先,为分离和增强目标样本的全局和特有信息,提出双路径特征提取方法,通过两种轻量级特征提取路径,从目标样本中分别提取融合任务级语义信息的代理令牌和样本局部特征;其次,为衡量源提示的可迁移性,提出门控代理注意力机制,利用代理令牌和局部特征分别与提示特征计算注意力分布以表征任务相似度,同时,为保留样本特有信息的表达,利用门控单元过滤样本局部特征并整合两类注意力;最后,提出注意力扰动方法,利用源提示在目标任务上的预测分布损失来扰动注意力分布,从而降低那些在表示空间中与目标任务相似但实际预测效果较差的源提示在迁移过程中的权重。在公开数据集GLUE上与9个基线模型相比,所提方法在平均性能上表现最佳,验证了其有效性。
中图分类号:
| [1]RAFFEL C,SHAZEER N,ROBERTS A,et al.Exploring the limits of transfer learning with a unified text-to-text transformer[J].Journal of Machine Learning Research,2020,21(140):1-67. [2]HOULSBY N,GIURGIU A,JASTRZEBSKI S,et al.Parame-ter-efficient transfer learning for NLP[C]//Proceedings of the 36th International Conference on Machine Learning.2019,97:2790-2799. [3]LIU P,YUAN W,FU J,et al.Pre-train,prompt,and predict:A systematic survey of prompting methods in natural language processing[J].ACM Computing Surveys,2023,55(9):1-35. [4]BEN ZAKEN E,GOLDBERG Y,RAVFOGEL S.BitFit:Simpleparameter-efficient fine-tuning for transformer-based masked language models[C]//Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics.2022:1-9. [5]LESTER B,AL-RFOU R,CONSTANT N.The power of scale for parameter-efficient prompt tuning[C]//Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing(EMNLP).2021:3045-3059. [6]VU T,LESTER B,CONSTANT N,et al.SPoT:Better frozenmodel adaptation through soft prompt transfer[C]//Procee-dings of the 60th Annual Meeting of the Association for Computational Linguistics.2022:5039-5059. [7]WANG A,SINGH A,MICHAEL J,et al.GLUE:A multi-task benchmark and analysis platform for natural language understanding[C]//Proceedings of the 2018 EMNLP Workshop BlackboxNLP:Analyzing and Interpreting Neural Networks for NLP.2018:353-355. [8]HAN D,YE T,HAN Y,et al.Agent attention:On the integration of softmax and linear attention [C]//European Conference on Computer Vision.Cham:Springer,2024:124-140. [9]WANG Z,PANDA R,KARLINSKY L,et al.Multitask prompt tuning enables parameter-efficient transfer learning[C]//The Eleventh International Conference on Learning Representations(ICLR).2023. [10]ASAI A,SALEHI M,PETERS M,et al.ATTEMPT:Parameter-efficient multi-task tuning via attentional mixtures of soft prompts[C]//Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.2022:6655-6672. [11]WU X,CHEN C L P,LI S,et al.Snapshot prompt ensemble for parameter-efficient soft prompt transfer[C]//ICASSP 2024-2024 IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP).2024:11946-11950. [12]WU M,LIU W,XU J,et al.Parameter-efficient multi-task fine-tuning by learning to transfer token-wise prompts[C]//Findings of the Association for Computational Linguistics:EMNLP 2023.2023:8734-8746. [13]ZHONG Q,DING L,LIU J,et al.Panda:Prompt transfer meets knowledge distillation for efficient model adaptation[J].IEEE Transactions on Knowledge and Data Engineering,2024,36:4835-4848. [14]WANG A,PRUKSACHATKUN Y,NANGIA N,et al.Super GLUE:A stickier benchmark for general-purpose language understanding systems[C]//Proceedings of the 33rd International Conference on Neural Information Processing Systems.2019:3266-3280. [15]KHOT T,SABHARWAL A,CLARK P.Sci Tail:A textual entailment dataset from science question answering [C]//Procee-dings of AAAI.2018:5189-5197. [16]MAHABADI R K,RUDER S,DEHGHANI M,et al.Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks [C]//Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing.2021:565-576. |
|
||