Computer Science ›› 2026, Vol. 53 ›› Issue (9): 16-23.doi: 10.11896/jsjkx.260100147

• Research and Application of Large Language Model Technology • Previous Articles     Next Articles

Survey of Chinese Datasets for LLM Safety Alignment:Landscape and Prospects

LI Luozheng1, LI Lingbo1, YUAN Quan2   

  1. 1 Institute of Software,Chinese Academy of Science,Beijing 100190,China
    2 Beijing Institute of Tracking and Telecommunications Technology,Beijing 100080,China
  • Received:2026-01-25 Revised:2026-05-06 Online:2026-09-15 Published:2026-09-10
  • About author:LI Luozheng,born in 1990,Ph.D,assistant researcher.His main research interests include cognitive computing and social computing.

Abstract: Safety alignment is a critical technique for preventing large language models(LLMs) from generating harmful content and ensuring their outputs conform to human values and safety standards,whose effectiveness heavily relies on high-quality datasets.Current research primarily focuses on English contexts and Western cultural backgrounds,lacking specialized datasets for Chinese scenarios.This paper systematically reviews the technical framework of LLMs safety alignment and provides a focused overview of the current landscape of publicly available Chinese safety alignment datasets.Firstly,it clarifies the position of safety alignment within the post-training safety paradigm and its technical evolution.Then,it analyzes the characteristics and limitations of existing Chinese datasets in terms of safety risk taxonomy,data scale,and application purposes.Finally,this paper outlines future research directions from three perspectives:data scale,completeness,and orientation.It suggests that future work should prioritize data quality,mitigate the “safety tax” phenomenon,and explore synergistic optimization between pre-training and alignment stages.This survey aims to provide a data foundation and research direction for building more efficient,robust,and contextually appropriate safety alignment systems for the Chinese language.

Key words: Large language models, Safety alignment, Chinese datasets, Jailbreaking

CLC Number: 

  • TP316
[1] BROWN T B,MANN B,RYDER N,et al.Language Models are Few-Shot Learners[J].Advances in Neural Information Processing Systems,2020,33:1877-1901.
[2] BOMMASANI R,HUDSON D A,ADELI E,et al.On the Opportunities and Risks of Foundation Models[J].arXiv:2108.07258,2021.
[3] ZENG A,LIU X,DU Z,et al.GLM-130B:an open bilingual pre-trained model[J].arXiv:2210.02414,2022.
[4] BAO H,ZHOU P,WANG X,et al.Ethical and Social Risks of Large Language Models:A Survey[J].Journal of Artificial Intelligence Research,2024,79:1-55.
[5] WEI A,HAGHTALAB N,STEINHARDT J.Jailbro-ken:How Does LLM Safety Training Fail?[C] //Proceedings of the 40th International Conference on Machine Learning.2023:34567-34583.
[6] SHEN L,LI R,LIU Y,et al.Chain-of-Thought Hijack-ing:A NewJailbreaking Attack on Large Language Models[J].arXiv:2501.01289,2025.
[7] LI X T,WU J,ZHENG Q H,et al.Jailbreak attacks on large language models:models,root causes,and the evolution of attacks and defenses [J].Science China Information Sciences,2025,55:1372-1405.
[8] WANG Y,CHEN Z,LIU J,et al.Sugar-Coated Poison:Bypas-sing LLM Safety Filters through Benign-Content Prefixes[C] //Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security.2025:1-15.
[9] RÖTTGER P,PERNISI F,VIDGEN B,et al.Safety prompts:a systematic review of open datasets for evaluating and improving large language model safety[C] //Proceedings of the 2025 AAAI Conference on Artificial Intelligence.2025:27617-27627.
[10] OUYANG L,WU J,JIANG X,et al.Training Language Models to Follow Instructions with Human Feedback[J].Advances in Neural Information Processing Systems,2022,35:27730-27744.
[11] WANG K,ZHANG G B,ZHOU Z H,et al.A ComprehensiveSurvey in LLM(-Agent) Full Stack Safety:Data,Training and Deployment[J].arXiv:2504.15585,2025.
[12] CHUA J,LI Y,YANG S,et al.AI Safety in Generative AI Large Language Models:A Survey[J].arXiv:2407.18369,2024.
[13] ZHANG Z,LU Y,MA J,et al.ShieldLM:Empower-ing LLMs as Aligned,Customizable and Explainable Safety Detectors[C] //Findings of the Association for Computational Linguistics(EMNLP 2024).ACL,2024:10420-10438.
[14] LI Z J.Research on safety alignment techniques for large language models [D].Harbin:Harbin Institute of Technology,2025.
[15] DENG G,LIU Y,LI Y,et al.Masterkey:Automated Jailbrea-king of Large Language Model Chat-bots[C] //Proceedings 2024 Network and Distributed System Security Symposium.2024.
[16] HARTVIGSEN T,GABRIEL S,PALANGI H,et al.ToxiGen:A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection[C] //Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics.2022:3309-3326.
[17] ANTYPAS D,CAMACHO-COLLADOS J.Robust Hate Speech Detection in Social Media:A Cross-Dataset Empirical Evaluation[C] //The 7th Workshop on Online Abuse and Harms(WOAH).ACL,2023:231-242.
[18] ALMOHAIMEED S,ALMOHAIMEED S,SHAFIN A A,et al.THOS:A Benchmark Dataset for Targeted Hate and Offensive Speech[C] //Proceedings of Datacentric Machine Learning Research(DMLR) Workshop at ICML 2023.2023.
[19] INAN H,UPASANI K,CHI J,et al.Llama Guard:LLM-based Input-Output Safeguard for Human-AI Conversations[J].ar-Xiv:2312.06674,2023.
[20] ZHAO H,YUAN C,HUANG F,et al.Qwen3Guard Technical Report[J].arXiv:2510.14276,2025.
[21] HU E J,SHEN Y,WALLIS P,et al.LoRA:Low-Rank Adaptation of Large Language Models[C] //ICLR.2022:3.
[22] YI X Y,XIE X.An analysis of moral and value alignment issues in large models [J].Journal of Computer Research and Development,2023,60(9):1926-1945.
[23] BAI Y,KADAVATH S,KUNDU S,et al.Constitutional AI:Harmlessness from AI Feedback[J].arXiv:2212.08073,2022.
[24] HUANG Y,ZHANG H,WANG L,et al.Lazy Safety Align-ment:Efficient Safety Fine-tuning for Large Language Models[C] //Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics.2024:456-470.
[25] CAO Z,YANG Y,ZHAO H.SCANS:Mitigating the Exaggerated Safety for LLMs via Safety-Conscious Activation Steering[C] //Proceedings of the AAAI Conference on Artificial Intelligence.2025:23523-23531.
[26] XIE Y,YI J,SHAO J,et al.Defending ChatGPT Against Jailbreak Attack via Self-Reminders[J].Nature Machine Intelligence,2023,5(12):1486-1496.
[27] CHAO P,ROBEY A,DOBRIBAN E,et al.Jailbreaking Black Box Large Language Models in Twenty Queries[C] //R0-FoMo:Robustness of Few-shot and Zero-shot Learning in Large Foundation Models.2023.
[28] ZHANG Z,YANG J,KE P,et al.Defending Large LanguageModels Against Jailbreaking Attacks Through Goal Prioritization[C] //Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics.ACL,2024:8865-8887.
[29] WU X,WANG R,YANG Y,et al.Improving LLM Safety Alignment with Dual-Objective Optimiza-tion[C] //Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing.2024:5678-5693.
[30] DU T,WEI Z,CHEN Q,et al.Advancing LLM Safe Alignment with Safety Representation Ranking[C] //ICML 2025 Workshop on Reliable and Responsible Foundation Models.2025.
[31] SUN H,ZHANG Z,DENG J,et al.Safety Assessment of Chinese Large Language Models[J].arXiv:2304.10436,2023.
[32] YUAN X,LI J,WANG D,et al.S-Eval:Towards Automatedand Comprehensive Safety Evaluation for Large Language Mo-dels[C] //Proceedings of the ACM on Software Engineering.2025:2136-2157.
[33] XU L,ZHAO K,ZHU L,et al.SC-Safety:A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in Chinese[J].arXiv:2310.05818,2023.
[34] WANG Y,ZHAI Z,LI H,et al.A Chinese Dataset for Evaluating the Safeguards in Large Language Models[C] //Findings of the Association for Computational Linguistics:ACL 2024.2024:3106-3119.
[35] WANG Y,ZHAI Z,LI H,et al.A Chinese Dataset for Evaluating the Safeguards in Large Language Models[J].arXiv:2402.12193,2024.
[36] TAN Y,ZHENG B,ZHENG B,et al.Chinese SafetyQA:ASafety Short-form Factuality Benchmark for Large Language Models[J].arXiv:2412.15265,2024.
[37] ZHANG H,GAO H,HU Q,et al.ChineseSafe:A ChineseBenchmark for Evaluating Safety in Large Language Models[J].arXiv:2410.18491,2025.
[38] ZHANG W,LEI X,LIU Z,et al.CHiSafetyBench:A Chinese Hierarchical Safety Benchmark for Large Language Models[J].arXiv:2406.10311,2024.
[39] XU G,LIU J,YAN M,et al.CValues:Measuring the Values of Chinese Large Language Models from Safety to Responsibility[J].arXiv:2307.09705,2023.
[40] ZHANG M,PAN X,YANG M.Jade:A Linguistics-Based Safety Evaluation Platform for LLM[J].arXiv:2311.00286,2023.
[41] CHEN Z,YU H,WU X,et al.Libra:Large Chines Based Safeguard for AI Content[J].arXiv:2507.21929,2025.
[42] ZHANG Z,LEI L,WU L,et al.SafetyBench:Evaluating theSafety of Large Language Models[J].arXiv:2309.07045,2023.
[43] ROTTGER P,KIRK H,VIDGEN B,et al.XSTest:A Test Suitefor Identifying Exaggerated Safety Behaviours in Large Language Models[C] //Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies.ACL,2024:5377-5400.
[44] SHI C,WANG X,GE Q,et al.Navigating the Over-Kill inLarge Language Models[C] //Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics.ACL,2024:4602-4614.
[45] CUI J,CHIANG W L,STOICA I,et al.OR-Bench:An Over-Refusal Benchmark for Large Language Mod-els[J].arXiv:2405.20947,2024.
[46] ZHANG Z,XU W,WU F,et al.FalseReject:A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning[J].arXiv:2505.08054,2025.
[47] WANG Z,TU H,WANG Y,et al.STAR-1:Safer Alignment of Reasoning LLMs with 1K Data[J].arXiv:2504.01903,2025.
[48] HUANG T,HU S,ILHAN F,et al.Safety Tax:Safety Align-ment Makes Your Large Reasoning Models Less Reasonable[J].arXiv:2503.00555,2025.
[49] QI X,PANDA A,LYU K,et al.Safety Alignment Should BeMade More Than Just a Few Tokens Deep[J].arXiv:2406.05946,2024.
[50] CHEN J,WANG X,YAO Z,et al.Finding Safety Neurons in Large Language Models[J].arXiv:2406.14144,2024.
[51] XUE Y,MIRZASOLEIMAN B.LoRA is All You Need for Safety Alignment of Reasoning LLMs[J].arXiv:2507.17075,2025.
[52] JI J,WANG K,QIU T A,et al.Language Models Re-sist Alignment:Evidence from Data Compression[C] //Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics.2025:23411-23432.
[53] KANG D,PARK J,JO Y,et al.From values to opinions:predicting human behaviors and stances using value-injected large language models[C] //Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing.ACL,2023:15539-15559.
[54] YAO J,YI X,XIE X.CLAVE:an adaptive framework for evaluating values of LLM generated responses[J].Advances in Neural Information Processing Systems,2024,37:58868-58900.
[1] HE Jiaojun, LI Xin. Review of Graph Learning Based on Large Language Models:Methods,Benchmarks and Advances [J]. Computer Science, 2026, 53(9): 1-15.
[2] ZHANG Rongjie, PANG Xiongwen, WANG Fengling. Review of Large Language Model-based Time Series Modeling via Fine-tuning and AgentArchitecture [J]. Computer Science, 2026, 53(9): 55-70.
[3] GUO Yuyang, SHI Lei, LIU Huan, DONG Yixiang, LI Rui. Domain-adapted and Dynamically Retrieval-augmented Approach for Large-scale History Discipline Model Construction [J]. Computer Science, 2026, 53(9): 92-100.
[4] LI Zhennan, QIAN Jiayan, WANG Xinzhi, ZHANG Hui. Technology Risk Structure Recognition Based on Multi-granularity Semantic Dual Reflection [J]. Computer Science, 2026, 53(9): 395-404.
[5] LIU Jing. Review of Music Artificial Intelligence Driven by Large Language Models [J]. Computer Science, 2026, 53(8): 229-244.
[6] ZHANG Haoran, HAO Wenning, JIN Dawei, CHENG Kai, LIU Junyang. Agentic Retrieval Augmented Generation Framework Based on Retrieval Task Planning and Reflection Mechanism [J]. Computer Science, 2026, 53(8): 285-297.
[7] CHEN Zhixiang, XIE Zhipeng. Event Causal Data Augmentation Method Based on Large Language Model [J]. Computer Science, 2026, 53(7): 125-131.
[8] XU Rui, LIU Jin, LIU Xudong, GUAN Jian, DONG Wei. Exploring the Generalization Ability of Prompt-based Large Language Models for TextClassification [J]. Computer Science, 2026, 53(6A): 250400092-7.
[9] WEI Qing, ZHANG Yupeng, LIU Shaoxun, ZHANG Jinfeng, ZHANG Yuezhong, CHEN Haoyang. Fuzzing Driver Generation Based on Large Language Models [J]. Computer Science, 2026, 53(6A): 250400113-8.
[10] ZHANG Yongyu, GUO Chenjuan, FEI Xueqin, LI Feng. Study on Financial Text Sentiment Analysis Method Based on Large Language Models with Market Feedback Supervision [J]. Computer Science, 2026, 53(6A): 250500073-14.
[11] SHI Hongxu, LIU Yi, LIU Kun. Survey of Recommendation Systems Based on Large Language Models [J]. Computer Science, 2026, 53(6): 281-303.
[12] WANG Shenghui, LI Teng. Innovative Automated Scoring Based on Large Language Models [J]. Computer Science, 2026, 53(5): 90-98.
[13] HU Junjie, CHEN Yujie, HU Yikun, WEN Cheng, CAO Jialun, MA Zhi, SU Jie, SUN Weidi, TIAN Cong, QIN Shengchao. Formal Theorem Proving Empowered by Large Language Model:Survey and Perspectives [J]. Computer Science, 2026, 53(4): 1-23.
[14] LIU Suyi, LIU Qi, GAO Weibo. Agent4Stu:Efficient LLM-based Student Answer Behavior Simulation Agent [J]. Computer Science, 2026, 53(4): 347-355.
[15] XU Cheng, LIU Yuxuan, WANG Xin, ZHANG Cheng, YAO Dengfeng, YUAN Jiazheng. Review of Speech Disorder Assessment Methods Driven by Large Language Models [J]. Computer Science, 2026, 53(3): 307-320.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!