计算机科学 ›› 2026, Vol. 53 ›› Issue (7): 132-138.doi: 10.11896/jsjkx.250600021

• 人工智能 • 上一篇    下一篇

基于知识增强的文本关键信息提取嵌入优化研究

车云力, 唐晋韬, 王挺, 张健   

  1. 国防科技大学计算机学院 长沙 410073
  • 收稿日期:2025-06-03 修回日期:2025-12-25 出版日期:2026-07-15 发布日期:2026-07-10
  • 通讯作者: 唐晋韬(tangjintao@nudt.edu.cn)
  • 作者简介:(462824127@qq.com)

Knowledge-enhanced Text Embedding Optimization Based on Key Information Extraction

CHE Yunli, TANG Jintao, WANG Ting, ZHANG Jian   

  1. School of Computer Science,National University of Defense Technology,Changsha 410073,China
  • Received:2025-06-03 Revised:2025-12-25 Published:2026-07-15 Online:2026-07-10
  • About author:CHE Yunli,born in 1994,postgraduate.His main research interest is information extraction.
    TANG Jintao,born in 1981,Ph.D,professor,Ph.D supervisor,is a member of CCF(No.12145S).His main research interests include information extraction and natural language processing.

摘要: 针对长文本检索中存在的语义稀释与细粒度信息丢失问题,提出了一种基于显性知识提取的文本嵌入优化框架。传统方法依赖单一向量表征全文,难以有效捕捉长文档中多主题、多层次的局部语义关联。为此,设计了两种知识感知嵌入策略:1)独立知识点编码(KASE),通过大语言模型提取文本片段中的关键知识点并独立向量化,保留细粒度语义;2)知识点融合编码(KACE),将知识点拼接后进行整体编码,探索聚合效应。在中文(CMRC,DRCD)与英文(SQuAD,NewsQA)数据集上进行的对比实验结果表明,知识增强方法显著优于传统基线及问答对增强(QAEA-DR)、假设文档嵌入(HyDE)等方法,其中KASE表现最优,尤其在长文本场景下性能提升更为显著。消融实验表明,独立知识点表征对缓解信息丢失问题具有关键作用,而结合原始文本与知识点向量的混合策略可进一步优化检索效果。

关键词: 文本检索, 知识提取, 语义嵌入

Abstract: To address the challenges of semantic dilution and fine-grained information loss in long-text retrieval,this study proposes a text embedding optimization framework based on explicit knowledge extraction.Traditional methods that rely on single-vector representations of entire texts struggle to capture the multi-topic and hierarchical local semantic associations present in lengthy documents.It designs two knowledge-aware embedding strategies:1) Knowledge-aware separate embedding(KASE),which employs a large language model to extract key knowledge points from text segments and vectorizes them independently to preserve fine-grained semantics;2) Knowledge-aware concatenated embedding(KACE),which concatenates the extracted know-ledge points into a single passage and encodes it holistically to explore aggregation effects.Experimental results on Chinese datasets(CMRC,DRCD) and English datasets(SQuAD,NewsQA) demonstrate that the proposed knowledge-enhanced methods significantly outperform both traditional baselines and recent approaches such as QA-pair-augmented retrieval(QAEA-DR) and Hypothetical document embeddings(HyDE).KASE achieves the most substantial improvements,especially in long-text scenarios.Ablation studies reveal that independent knowledge-point representations are crucial for mitigating information loss,while hybrid strategies that combine original text embeddings with knowledge-aware vectors further optimize retrieval effectiveness.

Key words: Text retrieval, Knowledge extraction, Semantic embedding

中图分类号: 

  • TP391
[1]KARPUKHIN V,OĞUZ B,MIN S,et al.Dense Passage Retrieval for Open-Domain Question Answering[J].arXiv:2004.04906,2020.
[2]LEWIS P,PEREZ E,PIKTUS A,et al.Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks [J].Advances in Neural Information Processing Systems,2021,34:9459-9474.
[3]LIN J,NOGUEIRA R,YATES A.Pretrained Transformers for Text Ranking:BERT and Beyond[J].Foundations and Trends in Information Retrieval,2021,15(3/4):184-288.
[4]TAN H M,ZHAN S X,LIN H,et al.QAEA-DR:A Unified Text Augmentation Framework for Dense Retrieval[J].IEEE Transactions on Knowledge & Data Engineering,2025,37(6):3669-3683.
[5]GAO L Y,MA X G,LIN J,et al.Precise Zero-Shot Dense Retrieval without Relevance Labels[J].arXiv:2205.02335,2022.
[6]FAN C X,YAN Z,WU Y X,et al.Span prompt dense passage retrieval for Chinese open domain question answering[J].Journal of Intelligent & Fuzzy Systems,2023,45(5):7285-7295.
[7]LUAN Y,EISENSTEIN J,TOUTANOVA K,et al.Sparse,Dense,and Attentional Representations for Text Retrieval[J].Transactions of the Association for Computational Linguistics,2021,9:329-345.
[8]LI H,MOURAD A,ZHUANG S,et al.Pseudo relevance feedback with deep language models and dense retrievers:successes and pitfalls[J].ACM Transactions on Information Systems,2023,41(3):1-40.
[9]ZHAO P,ZHANG H,YU Q,et al.Retrieval-Augmented Generation for AI-Generated Content:A Survey [J].arXiv:2402.19473,2024.
[10]XIANG W,CUI X Y,CHENG N,et al.Zero-shot information extraction via chatting with ChatGPT[J].arXiv:2302.10205,2023.
[11]BONIFACIO L,ABONIZIO H,FADAEE M,et al.Inpars:Unsupervised dataset generation for information retrieval [C]//Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval.New York:ACM,2022:2387-2392.
[12]HUANG H,LIU X,SHI G,et al.Event extraction with dynamic prefix tuning and relevance retrieval[J].IEEE Transactions on Knowledge and Data Engineering,2023,35(10):9946-9958.
[13]LIU X Y,ZHENG Y X,DUAN Z Y,et al.Graph ERE:Jointly Multiple Event-Event Relation Extraction via Graph-Enhanced Event Embeddings [C]//Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.Stroudsburg:Association for Computational Linguistics,2021:6614-6624.
[14]HEILMAN M,SMITH N A.Good question! Statistical ranking for question generation [C]//Human Language Technologies:The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics.Association for Computational Linguistics,2010:609-617.
[15]GIRAY L.Prompt Engineering with Chat GPT:A Guide for Academic Writers[J].Annals of Biomedical Engineering,2023,51(12):2629-2633.
[16]WANG X Y,GUI L,HE Y L.Document-level Multi-Event Extraction with Event Proxy Nodes and Hausdorff Distance Minimization[C]//Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics.Association for Computational Linguistics,2023:10118-10133.
[17]XIAO S T,LIU Z,ZHANG P T,et al.C-Pack:packaged resources to advance general Chinese embedding [J].arXiv:2309.07597,2023.
[18]XIE X.T2Ranking:A large-scale Chinese benchmark for pas-sage ranking[C]//Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval(SIGIR 2023).New York:ACM,2023:2681-2690.
[19]ZHAN J T,MAO J X,LIU Y Q,et al.Jointly Optimizing Query Encoder and Product Quantization to Improve Retrieval Performance[C]//Proceedings of the 30th ACM International Conference on Information and Knowledge Management.New York:ACM,2021:2487-2496.
[20]LEE K,HAN S,HWANG S,et al.When to Read Documents or QA History:On Unified and Selective Open-domain QA[C]//Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics.Association for Computational Linguistics,2023:1-15.
[21]CUI Y,LIU T,CHE W,et al.A Span-Extraction Dataset for Chinese Machine Reading Comprehension[C]//Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing(EMNLP-IJCNLP 2019).Association for Computational Linguistics,2019:5883-5889.
[22]RAJPURKAR P,ZHANG J,LOPYREV K,et al.SQuAD:100 000+ Questions for Machine Comprehension of Text[J].arXiv:1606.05250,2016.
[23]TRISCHLER A,WANG T,YUAN X,et al.NewsQA:A Machine Comprehension Dataset [R].Redmond:Microsoft Research Technical Report,2017.
[24]LI Z,ZHANG X,ZHANG Y,et al.Towards General Text Embeddings with Multi-stage Contrastive Learning[J].arXiv:2308.03241,2023.
[25]DEEPSEEK A,LIU A X,FENG B,et al.Deep Seek-V3 Technical Report [J].arXiv:2412.19437,2024.
[26]QWEN TEAM.Qwen2.5:A party of foundation models[EB/OL].(2024-09) [2025-08-07].https://qwenlm.github.io/blog/qwen2.5/.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!