计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250400113-8.doi: 10.11896/jsjkx.250400113
魏清, 张与鹏, 刘少勋, 张金锋, 张跃中, 陈昊阳
WEI Qing, ZHANG Yupeng, LIU Shaoxun, ZHANG Jinfeng, ZHANG Yuezhong, CHEN Haoyang
摘要: 随着软件系统的广泛应用,其安全性问题日益突出,模糊测试作为一种有效的漏洞检测手段,在软件开发中扮演着重要角色。然而,传统模糊测试工具依赖手工编写驱动程序,存在效率低、覆盖率不足等问题。为此,文中提出了一种基于大模型的自动化模糊测试驱动生成方法,通过智能代码解析模块提取函数接口和结构体定义,结合大模型的代码生成能力,自动生成符合Honggfuzz框架要求的驱动程序,并引入基于反馈的修正机制提升驱动生成成功率。实验结果表明,该方法在开源cJSON库和自研TBox项目中实现了100%的驱动生成成功率和模糊测试接口覆盖率;在开源Libtiff库中,驱动生成成功率为76.2%,模糊测试接口覆盖率为40.5%。针对Qwen2.5-coder(14B参数)和Qwen2.5-coder(32B参数)的消融实验表明,引入反馈修正机制能进一步优化驱动程序的生成成功率,对上述两个模型的反馈修正机制分别使驱动生成成功率提升了5.9%和2.4%。该方法显著提升了模糊测试的自动化程度和覆盖率,为复杂软件系统的漏洞检测提供了高效解决方案。未来可通过优化代码解析模块、改进提示词模板和增强大模型适应性,来进一步提升方法的通用性和漏洞发现能力。
中图分类号:
| [1] SCHILLER N,CHLOSTA M,SCHLOEGELM,et al.Drone Security and the Mysterious Case of DJI's DroneID[C]//Proceedings 2023 Network and Distributed System Security Symposium.2023. [2] ZHAO X Q,QU H P,XU J L,et al.A systematic review of fuzzing[J].Soft Computing,2023,28(6):5493-5522.. [3] YAN Q,HUANG M H,CAO H Y.A Survey of Human-ma-chine Collaboration in Fuzzing[C]//2022 7th IEEE International Conference on Data Science in Cyberspace(DSC).IEEE,2022:375-382. [4] ZALEWSKI M.American fuzzy lop[EB/OL].[2025-03-16] .https://lcamtuf.coredump.cx/afl/. [5] KRALEWSKI K.Honggfuzz:A security oriented,feedback-driven,evolutionary,easy-to-use fuzzer [EB/OL]. [2025-03-16] .https://github.com/google/honggfuzz. [6] SEREBRYANY K.LibFuzzer—A library for coverage-guided fuzz testing [EB/OL]. [2025-03-16] .http://llvm.org/docs/LibFuzzer.html. [7] LYU Y L,XIE Y,CHEN P,et al.Prompt Fuzzing for Fuzz Driver Generation[C]//ACM Conference on Computer and Communications Security.2024:1-15. [8] CHEN P,XIE Y X,LYU Y L,et al.Hopper:InterpretativeFuzzing for Libraries[C]//Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security.ACM,2023:1600-1614. [9] ZHANG W Y,ZHANG L,MAO J L,et al.Reverse analysis and automated testing of unknown protocols[J].Journal of Computers,2020,43(4):653-667. [10] LSPOGLOU K,AUSTIN D,MOHAN V,et al.FuzzGen:Automatic Fuzzer Generation[C]//29th Usenix Security Symposium(usenix Security 20),2020:2271-2287. [11] XIA C S,PALTENGHI M,TIAN J L,et al.Fuzz4All:Universal Fuzzing with Large Language Models[C]//Proceedings of the IEEE/ACM 46th International Conference on Software Engineering.ACM,2023:1-13. [12] ZHANG H X,RONG Y Y,HE Y F,et al.LLAMA FUZZ:Large Language Model Enhanced Greybox uzzing[C]//The 39th IEEE/ACM International Conference on Automated Software Engineering.2024:1-11. [13] LIU J H,JIANG H.DeepGenFuzz:An Efficient PDF Application Fuzzing Test Case Generation Framework Based on Deep Learning[J].Computer Science,2024,51(12):53-62. [14] PROTECT AI.Vulnhuntr:A tool to identify remotely exploitable vulnerabilities using LLMs and static code analysis.[EB/OL]. [2025-03-16] .https://github.com/protectai/vulnhuntr. [15] PEARCE H,TAN B,AHMAD B,et al.Examining zero-shotvulnerability repair with large language models[C]//IEEE Symposium on Security and Privacy.2023:2339-2356 [16] HAZIMEH A,HERRERA A,PAYER M,et al.Magma:AGround-Truth Fuzzing Benchmark[J].Proceedings of the Acm on Measurement and Analysis of Computing Systems,2020,4:1-29. [17] GAMBLE D.cJSON:Ultralightweight JSON parser in ANSI C[EB/OL].[2025-03-16] .https://github.com/DaveGamble/cJSON. [18] LI Y,YANG W Z,ZHANG Y,et al.Survey on Fuzzing Based on Large Language Model[J].Ruan Jian XueBao/Journal of Software,2025,36(6):1-28. [19] ALIBAB A.Qwen2.5-Coder:A code-specialized large language model[EB/OL].[2025-03-16] .https://ollama.com/library/qwen2.5-coder. [20] ZHANG Z,ZHANG Y Z,ZHANG J F,et al.An Endogenous Security Study of Telematics Box in Intelligent Connected Vehicles[J].IEEE Embedded Systems Letters,2024,16(4):501-504. [21] LEFFLER S.libtiff:TIFF Library and Utilities[EB/OL].[2025-03-16] .https://libtiff.gitlab.io/libtiff. [22] KANG J J,PAN W C,ZHANG T,et al.Correcting Factuality Hallucination in Complaint Large Language Model via Entity-Augmented[C]//2024 International Joint Conference on Neural Networks(IJCNN).IEEE,2024:1-8. |
|
||