计算机科学 ›› 2026, Vol. 53 ›› Issue (8): 336-356.doi: 10.11896/jsjkx.251100160

• 人工智能 • 上一篇    下一篇

基于多智能体协作的刑事案件裁判文书智能生成方法

韩林睿1,2, 宋高捷2, 郑日1,2, 李冰3,4, 崔衍5   

  1. 1 教育部哲学社会科学实验室——中国政法大学数据法治实验室 北京 100088
    2 中国政法大学数据法治研究院 北京 100088
    3 中国政法大学证据科学教育部重点实验室 北京 100088
    4 中国政法大学证据科学研究院 北京 100088
    5 中国司法大数据研究院有限公司 北京 100043
  • 收稿日期:2025-11-28 修回日期:2026-04-23 出版日期:2026-08-15 发布日期:2026-08-17
  • 通讯作者: 李冰(bingl@cupl.edu.cn)
  • 作者简介:(linrui_han@163.com)
  • 基金资助:
    广东省证据材料司法鉴定(南天)工程技术研究中心开放课题(ETRC202306);2022年国家重点研发计划“社会治理与智慧社会科技支撑”重点专项(2022YFC3303000);中央高校基本科研业务费专项资金

Automated Judicial Document Generation for Criminal Cases Based on Multi-agent Collaboration

HAN Linrui1,2, SONG Gaojie2, ZHENG Ri1,2, LI Bing3,4, CUI Yan5   

  1. 1 Ministry of Education Laboratory of Philosophy and Social Sciences-The CUPL Data Law Lab, China University of Political Science and Law, Beijing 100088, China
    2 Institute for Data Law, China University of Political Science and Law, Beijing 100088, China
    3 Key Laboratory of Evidence Law and Forensic Science, Ministry of Education, China University of Political Science and Law, Beijing 100088, China
    4 Institute of Evidence Law and Forensic Science, China University of Political Science and Law, Beijing 100088, China
    5 China Judicial Big Data Research Institute Co., Ltd., Beijing 100043, China
  • Received:2025-11-28 Revised:2026-04-23 Published:2026-08-15 Online:2026-08-17
  • About author:HAN Linrui,born in 2000,master,is a member of CCF(No.U9119G).His main research interests include data law,legal artificial intelligence and blockchain.
    LI Bing,born in 1982,Ph.D,associate professor,master’s supervisor.Her main research interests include judicial identification science,AI-assisted judicial identification application,and evidence evaluation.
  • Supported by:
    Open Fund of the Guangdong Engineering Technology Research Center for Forensic Evidence Identification(ETRC202306),2022 National Key R&D Program “Social Governance and Smart Society Technology Support” Key Special Project (2022YFC3303000) and Fundamental Research Funds for the Central Universities of Ministry of Education of China.

摘要: 针对刑事案件裁判文书自动生成中存在的量刑预测精度不足、复杂案情适应性差及格式规范性欠缺等挑战,提出一种基于多智能体协作的裁判文书智能生成方法(MAC-AG)。该方法以“理解-规划-执行-生成”为技术路线,构建由4个智能体组成的协作框架。案情判断智能体依托通用大语言模型,负责案件要素分析、争议焦点提取与案由分类等事实认定任务。案件分类智能体采用检索增强生成技术(Retrieval-Augmented Generation,RAG)完成法律规范与类案检索,实现罪名判定及繁简分流。判决预测智能体利用深度学习回归模型对大语言模型的初步判决预测进行修正,提升刑期与罚金的预测精度。文书生成智能体基于领域监督微调的大语言模型,确保生成文本的格式规范性。在面向中国法律体系的裁判文书生成基准评测集JuDGE上进行了实验,结果表明:1)MAC-AG显著提升了7种基线大语言模型的文书生成性能,其中Qwen3-8B模型应用MAC-AG后,在刑罚预测、罪名判定、法条引用、文本语义4项指标上提升最为显著,验证了方法的普适性;2)相较于传统文书生成最优方法MRAG,Qwen3-8B@MAC-AG的罪名预测F1值提升2.5%,法条引用F1值提升11.4%,裁判文书语义相似度综合提升14.01%;3)消融实验结果证实各智能体不可或缺,移除任一组件均导致性能显著下降,凸显了任务解耦与协同的有效性;4)人工评测中,Qwen3-8B@MAC-AG平均得分4.79(满分5分),在说理性、逻辑性、规范性、完整性、可读性、价值平衡性6个维度均优于基线。该研究为司法智能化提供了可复现的多智能体协作范式,提升了裁判文书生成中量刑预测的准确性、逻辑严谨度与司法适用性。

关键词: 多智能体协作, 裁判文书生成, 刑事案件, 数字法院, 思维链, 检索增强生成

Abstract: To address the challenges of automated judicial document generation for criminal cases,including limited accuracy in sentencing prediction,limited adaptability to complex case circumstances,and insufficient compliance with formal writing stan-dards,this paper proposes a multi-agent collaborative automated generation method for judicial documents(MAC-AG).Following an “understanding-planning-execution-generation” pipeline,MAC-AG is built on a collaborative framework composed of four specialized agents.The case fact analysis agent,powered by a general-purpose large language model(LLM),is responsible for factual determination tasks such as case element extraction,dispute focus identification,and case type classification.The case classification agent incorporates retrieval-augmented generation(RAG) to retrieve relevant legal provisions and similar cases,thereby supporting charge determination and simple-complex case routing.The judgment prediction agent employs a deep learning regression model to calibrate the LLM’s preliminary judgment predictions,improving the accuracy of sentence term and fine prediction.The document generation agent,built on a domain-supervised fine-tuned LLM,ensures that the generated text conforms to judicial writing conventions and formatting requirements.Experiments conducted on JuDGE,a benchmark for judicial document generation in the Chinese legal system,the results show that:1)MAC-AG consistently improves the document generation performance of seven baseline LLMs,with Qwen3-8B@MAC-AG achieving the most significant gains in sentencing prediction,charge determination,legal article citation,and semantic quality,demonstrating the generalizability of the proposed method;2)Compares with the state-of-the-art MRAG method,Qwen3-8B@MAC-AG improves the F1 score for charge prediction by 2.5%,the F1 score for legal article citation by 11.4%,and overall semantic similarity of judicial documents by 14.01%;3)Ablation experiments confirm that each agent is indispensable,as removing any component leads to a marked performance decline,highlighting the effectiveness of task decoupling and inter-agent collaboration;and 4)In human evaluation,Qwen3-8B@MAC-AG achieves an average score of 4.79/5.00,outperforming baseline methods across six dimensions:reasoning,logical consistency,norm compliance,completeness,readability,and value balancing.Overall,this study provides a reproducible multi-agent collaboration paradigm for judicial intelligence and substantially improves the accuracy of sentencing prediction,the rigor of legal reasoning,and the practical applicability of automated judicial document generation.

Key words: Multi-agent collaboration, Judicial document generation, Criminal cases, Digital court, Chain-of-Thought, Retrieval-augmented generation

中图分类号: 

  • TP183
[1] YU T Z.Analysis on Essentials of Criminal Adjudications[J].Journal of Law Application,2024(3):105-117.
[2] CHENG J H.The Empirical Evaluation and Coping Strategies of “More Cases and Fewer People” in Chinese Courts[J].China Legal Science,2022(6):238-261.
[3] SU W,YUE B,AI Q,et al.JuDGE:Benchmarking JudgmentDocument Generation for Chinese Legal System[C]//Procee-dings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval.New York,NY,USA:Association for Computing Machinery,2025:3573-3583.
[4] LI Z G.Designing and Realizing Comprehensive Digital Court[J].Peking University Law Journal,2022,34(1):5-24.
[5] SU X H,LIU P X.Model Construction for AI-Generated Court Decisions[J].Study & Exploration,2025(7):89-97.
[6] GAO S.The Basic Requirement and Ideal Model of Criminal Adjudication Reasoning—Based on the Experience of Three Cases of Commutation[J].Journal of Northeast Normal University(Philosophy and Social Sciences Edition),2022(4):91-104.
[7] Legal Application Research Center.Supreme People’s CourtCriminal Litigation Document Specifications:Compilation Standards and Legal Basis(People’s Courts Volume)[M].Beijing:China Legal Publishing House,2021.
[8] CUI J,SHEN X,WEN S.A Survey on Legal Judgment Prediction:Datasets,Metrics,Models and Challenges[J].IEEE Access,2023,11:102050-102071.
[9] SHU D,ZHAO H,LIU X,et al.LawLLM:Law Large Language Model for the US Legal System[C]//Proceedings of the 33rd ACM International Conference on Information and Knowledge Management.2024:4882-4889.
[10] CONG Y N,HAN L R,MA J Y,et al.Research on Intelligent Judgment of Criminal Cases Based on Large Language Models[J].Computer Science,2025,52(5):248-259.
[11] CUI J,NING M,LI Z,et al.ChatLaw:A Multi-Agent Collaborative Legal Assistant with Knowledge Graph Enhanced Mixture-of-Experts Large Language Model[J].arXiv:2306.16092,2023.
[12] WANG L M,HAN L R,DU Z W,et al.Research on Privacy Policy Compliance Detection Method for Mobile Application Based on Large Language Model[J].Computer Science,2025,52(8):1-16.
[13] WEI B.Judicial Applications and Regulation of Large LegalLanguage Models[J].Oriental Law,2024(5):57-73.
[14] LAI J,GAN W,WU J,et al.Large Language Models in Law:A Survey[J].AI Open,2024,5:181-196.
[15] DAHL M,MAGESH V,SUZGUN M,et al.Large Legal Fictions:Profiling Legal Hallucinations in Large Language Models[J].Journal of Legal Analysis,2024,16(1):64-93.
[16] LEWIS P,PEREZ E,PIKTUS A,et al.Retrieval-AugmentedGeneration for Knowledge-Intensive NLP Tasks[J].Advances in Neural Information Processing Systems,2020,33:9459-9474.
[17] HU E J,SHEN Y,WALLIS P,et al.LoRA:Low-Rank Adaptation of Large Language Models[C]//International Conference on Learning Representations.2022.
[18] KORT F.Predicting Supreme Court Decisions Mathematically:A Quantitative Analysis of the “Right to Counsel” Cases[J].American Political Science Review,1957,51(1):1-12.
[19] SERGOT M J,SADRI F,KOWALSKI R A,et al.The British Nationality Act as a Logic Program[J].Communications of the ACM,1986,29(5):370-386.
[20] LUO B,FENG Y,XU J,et al.Learning to Predict Charges for Criminal Cases with Legal Basis[C]//Proceedings of the 2017 Conference on Empirical Methods in Natural Language Proces-sing.2017:2727-2736.
[21] HU Z,LI X,TU C,et al.Few-Shot Charge Prediction with Discriminative Legal Attributes[C]//Proceedings of the 27th International Conference on Computational Linguistics.2018:487-498.
[22] ZHANG H,PAN B Z,TAN H Y,et al.Judgment Prediction Based on Legal Judgment Documents[J].Big Data Research,2021,7(5):164-175.
[23] ZHONG H,GUO Z,TU C,et al.Legal Judgment Prediction via Topological Learning[C]//Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing.2018:3540-3549.
[24] LIAO J C,YANG W Z,QIN Y B,et al.Method for Generating Judgment Documents Based on Trial Logic[J].Computer Science,2025,52(11):223-229.
[25] YU F,QUARTEY L,SCHILDER F.Exploring the Effective-ness of Prompt Engineering for Legal Reasoning Tasks[C]//Findings of the Association for Computational Linguistics:ACL 2023.2023:13582-13596.
[26] SONG Y,QIN Y,HUANG R,et al.Legal Text Summarization via Judicial Syllogism with Large Language Models[J].Journal of King Saud University Computer and Information Sciences,2025,37(5):111.
[27] SHI J,GUO Q,LIAO Y,et al.LegalGPT:Legal Chain ofThought for the Legal Large Language Model Multi-Agent Framework[C]//International Conference on Intelligent Computing.Singapore:Springer Nature Singapore,2024:25-37.
[28] WEI B,YU Y,GAN L,et al.An LLMs-based Neuro-Symbolic Legal Judgment Prediction Framework for Civil Cases[J/OL].Artificial Intelligence and Law,2025:1-35.https://link.springer.com/article/10.1007/s10506-025-09433-1.
[29] YAO R,WU Y,WANG C,et al.Elevating Legal LLM Responses:Harnessing Trainable Logical Structures and Semantic Knowledge with Legal Reasoning[J].arXiv:2502.07912,2025.
[30] WANG X,ZHANG X,HOO V,et al.LegalReasoner:A Multi-Stage Framework for Legal Judgment Prediction via Large Language Models and Knowledge Integration[J].IEEE Access,2024,12:166843-166854.
[31] QIN J H,MA Q C,LI M,et al.Recent Advances on Multi-Agent Collaboration:A Cross-Perspective of Game and Control Theory[J].Acta Automatica Sinica,2025,51(3):489-509.
[32] HEESS N,TB D,SRIRAM S,et al.Emergence of Locomotion Behaviours in Rich Environments[J].arXiv:1707.02286,2017.
[33] WANG L,MA C,FENG X,et al.A Survey on Large Language ModelBased Autonomous Agents[J].Frontiers of Computer Science,2024,18(6):186345.
[34] TRAN K T,DAO D,NGUYEN M D,et al.Multi-Agent Collaboration Mechanisms:A Survey of LLMs[J].arXiv:2501.06322,2025.
[35] HONG S,ZHUGE M,CHEN J,et al.MetaGPT:Meta Pro-gramming for a Multi-Agent Collaborative Framework[C]//Proceedings of the Twelfth International Conference on Lear-ning Representations(ICLR 2024).2024:35980-35999.
[36] CHEN W,SU Y,ZUO J,et al.AgentVerse:Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors[C]//Proceedings of the Twelfth International Conference on Lear-ning Representations(ICLR 2024).2024:6545-6588.
[37] WU Q,BANSAL G,ZHANG J,et al.AutoGen:Enabling Next-Gen LLM Applications via Multi-Agent Conversations[C]//First Conference on Language Modeling.2024:1-46.
[38] YUAN W,CAO J,JIANG Z,et al.Can Large Language Models Grasp Legal Theories? Enhance Legal Reasoning with Insights from Multi-Agent Collaboration[C]//Findings of the Association for Computational Linguistics:EMNLP 2024.2024:7577-7597.
[39] SUN J,DAI C,LUO Z,et al.LawLuo:A Multi-Agent Collaborative Framework for Multi-Round Chinese Legal Consultation[J].arXiv:2407.16252,2024.
[40] JIANG C,YANG X.AgentsBench:A Multi-Agent LLM Simulation Framework for Legal Judgment Prediction[J].Systems,2025,13(8):641.
[41] CHEN X,MAO M,LI S,et al.Debate-Feedback:A Multi-Agent Framework for Efficient Legal Judgment Prediction[J].arXiv:2504.05358,2025.
[42] YUAN W,SONG K,JIANG Z,et al.A Multi-Agent Frameworkwith Legal Event Logic Graph for Multi-Defendant Legal Judgment Prediction[J].Information Processing & Management,2026,63(1):104319.
[43] XU C,LIU Z S,ZHOU L Y,et al.Multi-Agent Collaborative Framework for Audit Case Qualification and Regulation Recommendation[J].Journal of Frontiers of Computer Science and Technology,2026,20(1):280-290.
[44] WANG Q,NI S,LIU H,et al.AutoPatent:A Multi-AgentFramework for Automatic Patent Generation[J].arXiv:2412.09796,2024.
[45] CHEN J,XIAO S,ZHANG P,et al.M3-Embedding:Multi-Linguality,Multi-Functionality,Multi-Granularity Text Embeddings Through Self-Knowledge Distillation[C]//Findings of the Association for Computational Linguistics:ACL 2024.2024:2318-2335.
[46] LI C,LIU Z,XIAO S,et al.Making Large Language Models aBetter Foundation for Dense Retrieval[J].arXiv:2312.15503,2023.
[47] WANG X Q.A Study on Distinguishing between Formal and Simplified Versions of Criminal Judgement Documents[J].The Jurist,2017(5):144-153,179-180.
[48] WILLMOTT C J,MATSUURA K.Advantages of the Mean Ab-solute Error(MAE) over the Root Mean Square Error(RMSE) in Assessing Average Model Performance[J].Climate Research,2005,30(1):79-82.
[49] POPESCU M C,BALAS V E,PERESCU-POPESCU L,et al.Multilayer Perceptron And Neural Networks[J].WSEAS Transactions On Circuits And Systems,2009,8(7):579-588.
[50] Tencent-Hunyuan/Hunyuan-7B[EB/OL].[2025-08-31].ht-tps://github.com/Tencent-Hunyuan/Hunyuan-7B.
[51] AI@Meta.Llama 3 Model Card[EB/OL].[2025-08-22].https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md.
[52] YANG A,LI A,YANG B,et al.Qwen3 Technical Report[J].arXiv:2505.09388,2025.
[53] GUO D,YANG D,ZHANG H,et al.DeepSeek-R1:Incentivizing Reasoning Capability in LLMs via Reinforcement Learning[J].arXiv:2501.12948,2025.
[54] GLM T,ZENG A,XU B,et al.ChatGLM:A Family of Large Language Models from GLM-130B to GLM-4 all Tools[J].ar-Xiv:2406.12793,2024.
[55] WU Y,LIU Y,LIU Y,et al.wisdom Interrogatory[EB/OL].(2024-03-18)[2025-08-22].https://github.com/zhihaiLLM/wisdomInterrogatory.
[56] DENG W,PEI J,KONG K,et al.Syllogistic Reasoning for Legal Judgment Analysis[C]//Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing.2023:13997-14009.
[57] YANG H,YUE S,HE Y.Auto-GPT for Online Decision Ma-king:Benchmarks and Additional Opinions[J].arXiv:2306.02224,2023.
[58] WEI J,WANG X,SCHUURMANS D,et al.Chain-of-Thought Prompting Elicits Reasoning in Large Language Models[J].Advances in Neural Information Processing Systems,2022,35:24824-24837.
[59] MIHALCEA R,TARAU P.TextRank:Bringing Order intoText[C]//Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing.2004:404-411.
[60] RAMOS J.Using TF-IDF to Determine Word Relevance in Document Queries[C]//Proceedings of the first Instructional Conference on Machine Learning.2003,242(1):29-48.
[61] LEWIS M,LIU Y,GOYAL N,et al.BART:Denoising Se-quence-to-Sequence Pre-training for Natural Language Generation,Translation,and Comprehension[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.2020:7871-7880.
[62] ZHANG Z,ZHANG H,CHEN K,et al.Mengzi:TowardsLightweight yet Ingenious Pre-trained Models for Chinese[J].arXiv:2110.06696,2021.
[63] LIU Y,LAPATA M.Text Summarization with Pretrained Encoders[C]//Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing(EMNLP-IJCNLP).2019:3730-3740.
[64] SHAO Y,GENG Z,LIU Y,et al.CPT:A Pre-trained Unba-lanced Transformer for both Chinese Language Understanding and Generation[J].Science China Information Sciences,2024,67(5):152102.
[65] BANERJEE S,LAVIE A.METEOR:An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments[C]//Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization.2005:65-72.
[66] ZHANG T,KISHORE V,WU F,et al.BERTScore:Evaluating Text Generation with BERT[J].arXiv:1904.09675,2019.
[67] JOSHI A,KALE S,CHANDEL S,et al.Likert Scale:Explored and Explained[J].British Journal of Applied Science & Technology,2015,7(4):396.
[68] KOO T K,LI M Y.A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research[J].Journal of Chiropractic Medicine,2016,15(2):155-163.
[69] MOSQUEIRA-REY E,HERNÁNDEZ-PEREIRA E,ALONSO-RÍOS D,et al.Human-in-the-Loop Machine Learning:A State of the Art[J].Artificial Intelligence Review,2023,56(4):3005-3054.
[70] LI J,LIU R,LI Y,et al.Tree of Reviews:A Tree-based Dyna-mic Iterative Retrieval Framework for Multi-hop Question Answering[J].arXiv:2404.14464,2024.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!