计算机科学 ›› 2026, Vol. 53 ›› Issue (7): 132-138.doi: 10.11896/jsjkx.250600021
车云力, 唐晋韬, 王挺, 张健
CHE Yunli, TANG Jintao, WANG Ting, ZHANG Jian
摘要: 针对长文本检索中存在的语义稀释与细粒度信息丢失问题,提出了一种基于显性知识提取的文本嵌入优化框架。传统方法依赖单一向量表征全文,难以有效捕捉长文档中多主题、多层次的局部语义关联。为此,设计了两种知识感知嵌入策略:1)独立知识点编码(KASE),通过大语言模型提取文本片段中的关键知识点并独立向量化,保留细粒度语义;2)知识点融合编码(KACE),将知识点拼接后进行整体编码,探索聚合效应。在中文(CMRC,DRCD)与英文(SQuAD,NewsQA)数据集上进行的对比实验结果表明,知识增强方法显著优于传统基线及问答对增强(QAEA-DR)、假设文档嵌入(HyDE)等方法,其中KASE表现最优,尤其在长文本场景下性能提升更为显著。消融实验表明,独立知识点表征对缓解信息丢失问题具有关键作用,而结合原始文本与知识点向量的混合策略可进一步优化检索效果。
中图分类号:
| [1]KARPUKHIN V,OĞUZ B,MIN S,et al.Dense Passage Retrieval for Open-Domain Question Answering[J].arXiv:2004.04906,2020. [2]LEWIS P,PEREZ E,PIKTUS A,et al.Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks [J].Advances in Neural Information Processing Systems,2021,34:9459-9474. [3]LIN J,NOGUEIRA R,YATES A.Pretrained Transformers for Text Ranking:BERT and Beyond[J].Foundations and Trends in Information Retrieval,2021,15(3/4):184-288. [4]TAN H M,ZHAN S X,LIN H,et al.QAEA-DR:A Unified Text Augmentation Framework for Dense Retrieval[J].IEEE Transactions on Knowledge & Data Engineering,2025,37(6):3669-3683. [5]GAO L Y,MA X G,LIN J,et al.Precise Zero-Shot Dense Retrieval without Relevance Labels[J].arXiv:2205.02335,2022. [6]FAN C X,YAN Z,WU Y X,et al.Span prompt dense passage retrieval for Chinese open domain question answering[J].Journal of Intelligent & Fuzzy Systems,2023,45(5):7285-7295. [7]LUAN Y,EISENSTEIN J,TOUTANOVA K,et al.Sparse,Dense,and Attentional Representations for Text Retrieval[J].Transactions of the Association for Computational Linguistics,2021,9:329-345. [8]LI H,MOURAD A,ZHUANG S,et al.Pseudo relevance feedback with deep language models and dense retrievers:successes and pitfalls[J].ACM Transactions on Information Systems,2023,41(3):1-40. [9]ZHAO P,ZHANG H,YU Q,et al.Retrieval-Augmented Generation for AI-Generated Content:A Survey [J].arXiv:2402.19473,2024. [10]XIANG W,CUI X Y,CHENG N,et al.Zero-shot information extraction via chatting with ChatGPT[J].arXiv:2302.10205,2023. [11]BONIFACIO L,ABONIZIO H,FADAEE M,et al.Inpars:Unsupervised dataset generation for information retrieval [C]//Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval.New York:ACM,2022:2387-2392. [12]HUANG H,LIU X,SHI G,et al.Event extraction with dynamic prefix tuning and relevance retrieval[J].IEEE Transactions on Knowledge and Data Engineering,2023,35(10):9946-9958. [13]LIU X Y,ZHENG Y X,DUAN Z Y,et al.Graph ERE:Jointly Multiple Event-Event Relation Extraction via Graph-Enhanced Event Embeddings [C]//Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.Stroudsburg:Association for Computational Linguistics,2021:6614-6624. [14]HEILMAN M,SMITH N A.Good question! Statistical ranking for question generation [C]//Human Language Technologies:The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics.Association for Computational Linguistics,2010:609-617. [15]GIRAY L.Prompt Engineering with Chat GPT:A Guide for Academic Writers[J].Annals of Biomedical Engineering,2023,51(12):2629-2633. [16]WANG X Y,GUI L,HE Y L.Document-level Multi-Event Extraction with Event Proxy Nodes and Hausdorff Distance Minimization[C]//Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics.Association for Computational Linguistics,2023:10118-10133. [17]XIAO S T,LIU Z,ZHANG P T,et al.C-Pack:packaged resources to advance general Chinese embedding [J].arXiv:2309.07597,2023. [18]XIE X.T2Ranking:A large-scale Chinese benchmark for pas-sage ranking[C]//Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval(SIGIR 2023).New York:ACM,2023:2681-2690. [19]ZHAN J T,MAO J X,LIU Y Q,et al.Jointly Optimizing Query Encoder and Product Quantization to Improve Retrieval Performance[C]//Proceedings of the 30th ACM International Conference on Information and Knowledge Management.New York:ACM,2021:2487-2496. [20]LEE K,HAN S,HWANG S,et al.When to Read Documents or QA History:On Unified and Selective Open-domain QA[C]//Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics.Association for Computational Linguistics,2023:1-15. [21]CUI Y,LIU T,CHE W,et al.A Span-Extraction Dataset for Chinese Machine Reading Comprehension[C]//Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing(EMNLP-IJCNLP 2019).Association for Computational Linguistics,2019:5883-5889. [22]RAJPURKAR P,ZHANG J,LOPYREV K,et al.SQuAD:100 000+ Questions for Machine Comprehension of Text[J].arXiv:1606.05250,2016. [23]TRISCHLER A,WANG T,YUAN X,et al.NewsQA:A Machine Comprehension Dataset [R].Redmond:Microsoft Research Technical Report,2017. [24]LI Z,ZHANG X,ZHANG Y,et al.Towards General Text Embeddings with Multi-stage Contrastive Learning[J].arXiv:2308.03241,2023. [25]DEEPSEEK A,LIU A X,FENG B,et al.Deep Seek-V3 Technical Report [J].arXiv:2412.19437,2024. [26]QWEN TEAM.Qwen2.5:A party of foundation models[EB/OL].(2024-09) [2025-08-07].https://qwenlm.github.io/blog/qwen2.5/. |
|
||