计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250400092-7.doi: 10.11896/jsjkx.250400092
徐睿, 刘金, 刘旭东, 关健, 董伟
XU Rui, LIU Jin, LIU Xudong, GUAN Jian, DONG Wei
摘要: 近年来,大语言模型(Large Language Models,LLMs)凭借 Zero-/Few-Shot Prompt 在文本分类中取得显著进展,但不同模型、任务和语言环境下性能波动明显。对此,基于 DeepSeek,Qwen,GPT-4o 3个系列模型,在 AG News,THUCNews,IMDb,ChnSentiCorp 这4个中英文数据集上系统评估 Zero-/Few-Shot(1/3/5-Shot)效果。结果表明:模型规模越大,Few-Shot 稳定性越高;情感任务适合 Few-Shot,新闻任务更偏好 Zero-Shot;类别代表性、区分度决定 Prompt 示例效用。进而提出“四步决策流程”和语义特征驱动的 Prompt 设计准则,为 LLM 文本分类部署提供参考。
中图分类号:
| [1] MOHAJERI M M,DOUSTI M J,AHMADABADI M N.CoCoP:Enhancing Text Classification with LLM through Code Completion Prompt[J].arXiv:2411.08979,2024. [2] SUN X,LI X,LI J,et al.Text Classification via Large Language Models[C]//Findings of the Association for Computational Linguistics:EMNLP 2023.2023:8990-9005. [3] HE J,RUNGTA M,KOLECZEK D,et al.Does Prompt Formatting Have Any Impact on LLM Performance?[J].arXiv:2411.10541,2024. [4] ERRICA F,SIRACUSANO G,SANVITO D,et al.What did I do wrong? quantifying LLMs' sensitivity and consistency to prompt engineering[J].arXiv:2406.12334,2024. [5] ZHANG Y,WANG M,LI Q,et al.Pushing the limit of LLM capacity for text classification[C]//Companion Proceedings of the ACM on Web Conference 2025.2025:1524-1528. [6] GLAZKOVA A,ZAKHAROVA O.Evaluating llm prompts for data augmentation in multi-label classification of ecological texts[J].arXiv:2411.14896,2024. [7] ZHANG X,TALUKDAR N,VEMULAPALLI S,et al.Comparison of prompt engineering and fine-tuning strategies in large language models in the classification of clinical notes[J].AMIA Summits on Translational Science Proceedings,2024,2024:478. [8] SAKAI H,LAM S S.QUAD-LLM-MLTC:Large LanguageModels Ensemble Learning for Healthcare Text Multi-Label Classification[J].arXiv:2502.14189,2025. [9] GUO Y,OVADJE A,AL-GARADI M A,et al.Evaluating large language models for health-related text classification tasks with public social media data[J].Journal of the American Medical Informatics Association,2024,31(10):2181-2189. [10] LIU M,SHI G.Poliprompt:A high-performance cost-effectivellm-based text classification framework for political science[J].arXiv:2409.01466,2024. [11] PARIZI A H,LIU Y,NOKKU P,et al.A Comparative Study of Prompting Strategies for Legal Text Classification[C]//Proceedings of the Natural Legal Language Processing Workshop 2023.2023:258-265. [12] CRUICKSHANK I J,NG L H X.Prompting and fine-tuningopen-sourced large language models for stance classification[J].arXiv:2309.13734,2023. [13] YIN K,LIU C,MOSTAFAVI A,et al.Crisissense-llm:Instruction fine-tuned large language model for multi-label social media text classification in disaster informatics[J].arXiv:2406.15477,2024. [14] LIU M,BU C,BAI S,et al.Classification of Table Cells Based on LLM Prompts[C]//2024 IEEE International Conference on Systems,Man,and Cybernetics(SMC).IEEE,2024:2140-2145. [15] VAJJALA S,SHIMANGAUD S.Text Classification in theLLM Era-Where do we stand?[J].arXiv:2502.11830,2025. [16] KOSTINA A,DIKIAKOS M D,STEFANIDIS D,et al.Large Language Models For Text Classification:Case Study And Comprehensive Review[J].arXiv:2501.08457,2025. [17] XU H,LOU R,DU J,et al.LLMs' Classification Performance is Overclaimed[J].arXiv:2406.16203,2024. [18] FECHNER R,DÖRPINGHAUS J.No Train,No Pain? Assessing the Ability of LLMs for Text Classification with no Finetuning[C]//Proceedings of the Position Papers of the 19th Confe-rence on Computer Science and Intelligence Systems(FedCSIS).Belgrade,Serbia.2024:8-11. [19] WANG Z,PANG Y,LIN Y,et al.Adaptable and Reliable Text Classification using Large Language Models[C]//2024 IEEE International Conference on Data Mining Workshops(ICDMW) 2024:67-74. [20] LIU C,ZHANG H,ZHAO K,et al.LLMEmbed:RethinkingLightweight LLM's Genuine Function in Text Classification[C]//Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics(Volume 1:Long Papers).2024:7994-8004. [21] LIU Y,YANG T,HUANG S,et al.Calibrating LLM-BasedEvaluator[C]//Proceedings of the 2024 Joint International Conference on Computational Linguistics,Language Resources and Evaluation(LREC-COLING 2024).2024:2638-2656. [22] ZHAO Z,WALLACE E,FENG S,et al.Calibrate before use:Improving few-shot performance of language models[C]//International Conference on Machine Learning.PMLR,2021:12697-12706. [23] KAPLAN J,MCCANDLISH S,HENIGHAN T,et al.Scaling laws for neural language models[J].arXiv:2001.08361,2020. [24] HOFFMANN J,BORGEAUD S,MMENSCH A,et al.Training compute-optimal large language models[C]//Proceedings of the 36th International Conference on Neural Information Processing Systems.2022:30016-30030. |
|
||