计算机科学 ›› 2026, Vol. 53 ›› Issue (8): 316-325.doi: 10.11896/jsjkx.260500095

• 人工智能 • 上一篇    下一篇

基于要素感知的普法案例筛选方法

毛逸潇1, 汪子霄2, 张柏礼3, 宗绍昊4   

  1. 1 中国政法大学证据科学教育部重点实验室 北京 100088
    2 东南大学软件学院 南京 211189
    3 东南大学计算机科学与工程学院 南京 211189
    4 中国政法大学数据法治研究院 北京 100088
  • 收稿日期:2026-05-18 修回日期:2026-07-23 出版日期:2026-08-15 发布日期:2026-08-17
  • 通讯作者: 毛逸潇(1196060588@qq.com)

Element-aware Screening Method for Popularization Cases

MAO Yixiao1, WANG Zixiao2, ZHANG Baili3, ZONG Shaohao4   

  1. 1 Key Laboratory of Evidence Science (China University of Political Science and Law), Ministry of Education, Beijing 100088, China
    2 College of Software Engineering, Southeast University, Nanjing 211189, China
    3 School of Computer Science and Engineering, Southeast University, Nanjing 211189, China
    4 Institute for Data Law, China University of Political Science and Law, Beijing 100088, China
  • Received:2026-05-18 Revised:2026-07-23 Published:2026-08-15 Online:2026-08-17
  • About author:MAO Yixiao,born in 1994,Ph.D,lectu-rer.His main research interests include evidence science and data science.

摘要: 普法案例筛选是智慧司法体系建设中的一项重要工作。目前,筛选工作主要依赖领域专家人工完成,成本高且效率低。从技术上看,普法案例筛选是一种普法价值判别的文本分类任务,但现有方法难以准确辨析裁判文书中与普法价值相关的关键事实、裁判结果和社会热点等细粒度要素,导致判别准确率不高。针对上述挑战,构建了普法案例数据集,并提出了一种能够实现普法要素精准感知的预训练模型。首先,通过跨平台数据采集与正则匹配方法构建数据集,明晰要素分布,并将其划分为事实、判决和标签3个互补视图,为模型训练提供数据支撑;然后,提出普法要素精准感知的预训练模型,分别用事实编码器表征争议焦点,判决编码器提取警示性特征,标签编码器捕捉案由的层级依赖,并设计涵盖独立路由和共享专家网络的特征融合模块,实现多视图的动态融合。在自建普法案例数据集上的实验结果表明,该模型能够有效提升普法价值判别的准确率与鲁棒性,并为普法案例筛选提供可行的辅助方案。

关键词: 自然语言处理, 普法案例, 预训练语言模型, 混合专家模型

Abstract: The screening of legal popularization cases is a critical mission in building a smart justice system.Traditional manual screening is costly and inefficient,while existing text classification methods struggle to accurately discern fine-grained elements related to educational value,such as key facts,judgment results,and social hot spots,thus limiting identification accuracy.To address this challenge,a legal popularization case dataset is constructed and an element-aware pre-trained model is proposed.Speci-fically,the dataset is built via cross-platform data collection and regular expression matching,clarifying the structured distribution of elements and partitioning them into three complementary semantic spaces:fact view,judgment view,and label view.A fact encoder exploring core disputes,a judgment encoder extracting warning features,and a label encoder capturing hierarchical dependen-cies are respectively utilized to extract deep features.Moreover,a multi-view feature fusion module is developed,employing independent routing and shared experts to achieve dynamic fusion of key elements across different views.Experimental results on the self-built dataset demonstrate that the proposed model effectively improves the accuracy and robustness of legal educational value identification,providing a feasible auxiliary solution for automated screening.

Key words: Natural language processing, Legal popularization cases, Pre-trained language models, Mixture-of-experts

中图分类号: 

  • TP391
[1] GUO J.Make good use of vivid cases to bridge the “last mile” of legal education[N].Qinghai Rule of Law Newspaper,2024.007.
[2] BELTAGY I,PETERS M E,COHAN A.Longformer:TheLong-Document Transformer[J].arXiv:2004.05150,2020.
[3] XIAO C,HU X,LIU Z,et al.Lawformer:A pre-trained lan-guage model for Chinese legal long documents[J].AI Open,2021,2:79-84.
[4] LI H,AI Q,CHEN J,et al.SAILER:Structure-aware Pre-trained Language Model for Legal Case Retrieval[C]//Procee-dings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval.2023:1035-1044.
[5] XIAO S,LIU Z,SHAO Y,et al.RetroMAE:Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder[C]//Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing.2022:538-548.
[6] WANG D S.Multi-defendant Legal Judgment Prediction withMulti-turn LLM and Criminal Knowledge Graph[J].Computer Science,2025,25(8):308-316.
[7] JI W D,WANG Y Q,SHEN Y C.Boosting Generative Rule Extraction via Negative-aware Approach[J].Computer Science,2026,53(5):276-285.
[8] YANG D,LI X.Legal Case Retrieval Method Based on the Integration of GraphStructure and Text Semantics[J/OL].Compu-ter Science:1-10.[2026-05-27].https://link.cnki.net/urlid/50.1075.tp.20260205.0904.002.
[9] SUN Z.A Short Survey of Viewing Large Language Models in Legal Aspect[J].arXiv:2303.09136,2023.
[10] YU Z,DONG Z,YU C,et al.A review on multi-view learning[J].Frontiers of Computer Science,2025,19(7):197334.
[11] DE MARTINO G,PIO G,CECI M.Multi-view overlappingclustering for the identification of the subject matter of legal judgments[J].Information Sciences,2023,638:118956.
[12] YANG S,TONG S,ZHU G,et al.MVE-FLK:A multi-task legal judgment prediction via multi-view encoder fusing legal keywords[J].Knowledge-Based Systems,2022,239:107960.
[13] YE F,LI S.MileCut:A Multi-view Truncation Framework for Legal Case Retrieval[C]//Proceedings of the ACM Web Conference 2024.2024:1341-1349.
[14] SHAZEER N,MIRHOSEINI A,MAZIARZ K,et al.Outra-geously Large Neural Networks:The Sparsely-Gated Mixture-of-Experts Layer[J].arXiv:1701.06538,2017.
[15] ZHANG Y,CAI J,WU Z,et al.Mixture of experts as representation learner for deep multi-view clustering[C]//Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence and Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence and Fifteenth Symposium on Educational Advances in Artificial Intelligence.2025:22704-22713.
[16] ZHU J,ZOU X,SUN J,et al.MoEGCL:Mixture of Ego-Graphs Contrastive Representation Learning for Multi-View Clustering[J].arXiv:2511.05876,2025.
[17] GARG S,PEITZ S,NALLASAMY U,et al.Jointly Learning to Align and Translate with Transformer Models[C]//Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP).2019:4453-4462.
[18] DEVLIN J,CHANG M W,LEE K,et al.BERT:Pre-training of Deep Bidirectional Transformers for Language Understanding[C]//Proceedings of the 2019 Conference of the North {A}meri-can Chapter of the Association for Computational Linguistics:Human Language Technologies.2019:4171-4186.
[19] LI H,AI Q,HAN X,et al.DELTA:pre-train a discriminativeencoder for legal case retrieval via structural word alignment[C]//Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence.2025:27072-27080.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!