计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250300034-10.doi: 10.11896/jsjkx.250300034
顾显俊1, 覃思航1, 舒一峰2, 马宝新2, 刘飞雪2, 刘铭2
GU Xianjun1, QIN Sihang1, SHU Yifeng2, MA Baoxin2, LIU Feixue2, LIU Ming2
摘要: 近年来,基于大模型的网络安全漏洞风险感知逐渐成为研究热点。然而,现有方法在智能化运营效率与细粒度信息感知方面仍存在响应速度慢、语义质量低等不足。为此,文中提出了一种基于检索增强生成(Retrieval-Augmented Generation)的轻量级网络安全漏洞风险感知方法。首先,通过构建跨知识域的知识库并结合跨知识域向量化算法,实现了网络漏洞信息的高效向量化。然后,设计了一种多目标检索算法,可自适应地从知识库中提取高匹配度的细粒度漏洞信息。最后,结合元数据二次增强与本地大模型,完成智能化漏洞风险感知响应。实验结果表明,所提方法在语义质量与运行效率上显著优于现有方案,能够快速响应网络安全风险并提供高质量的应对建议,充分满足网络安全智能化运营场景的需求。
中图分类号:
| [1] GHIASI M,NIKNAM T,WANG Z,et al.A comprehensive review of cyber-attacks and defense mechanisms for improving security in smart grid energy systems:Past,presentand future[J].Electric Power Systems Research,2023,215:108975. [2] LIN F,MEI Y,ZHU Y H,et al.A Review of the Whole-Process Impact of Cyber Attacks on Typical Scenarios in Power Systems[J].Southern Power Grid Technology,2023,17(11):61-75. [3] ALTULAIHAN E A,ALISMAIL A,FRIKHA M.A survey on web application penetration testing[J].Electronics,2023,12(5):1229. [4] ZHANG J,BU H,WEN H,et al.When llms meet cybersecurity:A systematic literature review[J].arXiv:2405.03644,2024. [5] YAMIN M M,HASHMI E,ULLAH M,et al.Applications of llms for generating cyber security exercise scenarios[J].IEEE Access,2024(12):143806-143822. [6] MITRA S,NEUPANE S,CHAKRABORTY T,et al.Localin-tel:Generating organizational threat intelligence from global and local cyber knowledge[J].arXiv:2401.10036,2024. [7] XIA C S,PALTENGHI M,LE TIAN J,et al.Fuzz4all:Universal fuzzing with large language models[C]//Proceedings of the IEEE/ACM 46th International Conference on Software Engineering.2024:1-13. [8] SHESTOV A,CHESHKOV A,LEVICHEV R,et al.Finetuning large language models for vulnerability detection[J].arXiv:2401.17010,2024. [9] KOEHN P,KNOWLES R.Six challenges for neural machinetranslation[J].arXiv:1706.03872,2017. [10] RAUNAK V,MENEZES A,JUNCZYS-DOWMUNT M.The curious case of hallucinations in neural machine translation[J].arXiv:2104.06683,2021. [11] MAYNEZ J,NARAYAN S,BOHNET B,et al.On faithfulness and factuality in abstractive summarization[J].arXiv:2005.00661,2020. [12] JI Z,LEE N,FRIESKE R,et al.Survey of hallucination in natural language generation[J].ACM Computing Surveys,2023,55(12):1-38. [13] GAO Y,XIONG Y,GAO X,et al.Retrieval-augmented generation for large language models:A survey[J].arXiv:2312.10997,2023. [14] RAJAPAKSHA S,RANI R,KARAFILI E.A rag-based question-answering solution for cyber-attack investigation and attribution[J].arXiv:2408.06272,2024. [15] DU X,ZHENG G,WANG K,et al.Vul-rag:Enhancing llm-based vulnerability detection via knowledge-level rag[J].arXiv:2406.11147,2024. [16] DANESHVAR S S,NONG Y,YANG X,et al.Exploring rag-based vulnerability augmentation with llms[J].arXiv:2408.04125,2024. [17] XU Y M,HU L,ZHAO J Y,et al.Research Progress and Insights on Large Language Models and Multilingual Intelligence[J].Computer Applications,2023,43(S2):1-8. [18] KENTON J D M-W C,TOUTANOVA L K.Bert:Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of naacL-HLT,vol.1.Minneapolis,Minnesota,2019. [19] WANG Y Y.What Can ChatGPT Bring to the Healthcare Industry[J]China Health,2023(4):73-75. [20] GUO X Y.The ‘Domestic Medical Version of ChatGPT' Amazingly Debuts[J].Chinese Hospital CEO,2023,19(7):24-25. [21] ZHANG X F,ZHANG L P,YAN S,et al.Personalized Learning Recommendation through the Collaboration of Knowledge Graphs and Large Language Models[J].Computer Applications,1-15. [22] WANG J K,QIN D H,BAI F B,et al.A Survey on the Integration of Speech Recognition and Large Language Models[J].Computer Engineering and Applications,1-13. [23] AGHAJANYAN A,ZETTLEMOYER L,GUPTA S.Intrinsic dimensionality explains the effectiveness of language model fine-tuning[J].arXiv:2012.13255,2020. [24] VILLALOBOS P,SEVILLA J,HEIM L,et al.Will we run out of data? an analysis of the limits of scaling datasets in machine learning[J].arXiv:2211.04325,2022. [25] WANG L M.Fundamental Issues in the Protection of Sensitive Personal Information-Interpreted in the Context of the Civil Code and the Personal Information Protection Law[J].Contemporary Law,2022,36(1):3-14. [26] ZHANG L J.The ‘Twenty Data Measures' Released:How to Tap into Data as the ‘New Oil'?[J].China Report,2023(1):66-68. [27] KANDPAL N,DENG H,ROBERTS A,et al.Large languagemodels struggle to learn long-tail knowledge[C]//International Conference on Machine Learning.PMLR,2023:15696-15707. [28] LEWIS P,PEREZ E,PIKTUS A,et al.Retrieval-augmentedgeneration for knowledge-intensive nlp tasks[J].Advances in Neural Information Processing Systems,2020,33:9459-9474. [29] MA X,GONG Y,HE P,et al.Query rewriting for retrieval augmented large language models[J].arXiv:2305.14283,2023. [30] GLASS M,ROSSIELLO G,CHOWDHURY M F M,et al.Re2g:Retrieve,rerank,generate[J].arXiv:2207.06300,2022. [31] CHEN J,XIAO S,ZHANG P,et al.Bge m3-embedding:Multi-lingual,multi-functionality,multi-granularity text embeddings through self-knowledge distillation[J].arXiv:2402.03216,2024. [32] VAN DER MAATEN L,HINTON G.Visualizing data usingt-sne[J].Journal of Machine Learning Research,2008,9(11). [33] PAN J J,WANG J,LI G.Survey of vector database management systems[J].The VLDB Journal,2024,33(5):1591-1615. [34] GRATTAFIORI A,DUBEY A,JAUHRI A,et al.The llama 3 herd of models[J].arXiv:2407.21783,2024. [35] EDGE D,TRINH H,CHENG N,et al.From local to global:A graph rag approach to query-focused summarization[J].arXiv:2404.16130,2024. [36] LIU Q,SONG J,HUANG Z,et al.glide the,and liunux4odoo,“langchain-chatchat,”[OL].https://github.com/chatchat-space/Langchain-Chatchat,2024. [37] ZHANG Y P,CHEN M F,TIAN C H,et al.Multi-Strategy Retrieval-Augmented Generation Method for Knowledge-Based Question Answering Systems in the Military Domain[J].Computer Applications,2025,45(3):746-754. [38] ACHIAM J,ADLER S,AGARWAL S,et al.Gpt-4 technical report[J].arXiv:2303.08774,2023. |
|
||