计算机科学 ›› 2026, Vol. 53 ›› Issue (8): 266-275.doi: 10.11896/jsjkx.260300033
李庆凯1,2, 群诺1,2,3, 倪胜巧1,4, 杨进1,5
LI Qingkai1,2, QUN Nuo1,2,3, NI Shengqiao1,4, YANG Jin1,5
摘要: 针对藏文命名实体识别任务中标注数据稀缺、分词困难及长实体识别效果不佳等挑战,提出一种基于字节级全局建模与门控特征融合机制的命名实体识别框架。该框架利用Byte Latent Transformer模型直接对原始UTF-8字节序列建模,利用其基于下一字节熵的动态分块机制获得不依赖固定词表的全局上下文表示,以规避传统分词器破坏实体边界并增强对未登录词的鲁棒性;同时设计门控线性条件融合模块,将对齐后的全局上下文特征以可控强度、按序列位置自适应地动态注入到保留精确边界信息的局部特征中,从而形成互补表示并提升跨度建模能力。在此基础上,模型结合BiLSTM-CRF完成序列建模与标签解码,兼顾全局语义建模与局部边界判别能力。在TibetanAI_NER和TibNER两个数据集上的实验结果表明,模型F1值相比各自最优模型分别提升11.5个百分点和4.53个百分点;消融实验进一步验证了全局特征与门控融合机制的协同贡献。结果表明,该框架在分词受限及低资源场景下能稳定提升实体边界判别能力,尤其对长实体识别更为有效,为低资源语言命名实体识别提供了一种可行的建模思路。
中图分类号:
| [1] MAYHEW S,BLEVINS T,LIU S,et al.Universal NER:Agold-standard multilingual named entity recognition benchmark[C]//Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies.2024:4322-4337. [2] GAO F,HUANG C,LIU Y,et al.Tlue:A tibetan language understanding evaluation benchmark[C]//Proceedings of the 2025 Conference on Empirical Methods in Natural Language Proces-sing.2025:35059-35085. [3] PATTNAYAK P,PATEL H,AGARWAL A.Tokenizationmatters:Improving zero-shotner for indic languages[C]//2025 IEEE International Conference on Electro Information Technology(eIT).IEEE,2025:456-462. [4] LI D,DU S,LI P,et al.LE-NER:A Chinese NER Model Based on Lexical Enhancement[M]//Advanced Data Mining and Applications.Singapore:Springer,2025:344-359. [5] PAGNONI A,PASUNURU R,RODRIGUEZ P,et al.Byte latent transformer:Patches scale better than tokens[C]//Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics.2025:9238-9258. [6] JIAYANG J,LI Y C,ZONG C Q,et al.Tibetan Person Name Recognition Based on Fusion of Maximum Entropy and Conditional Random Field Models[J].Journal of Chinese Information Processing,2014,28(1):107-112. [7] YUE J.Part-of-speech tagging of corpus based onBiLSTM-CRF model[C]//Proceedings of the 2025 4th International Confe-rence on Distributed Computing and Electrical Circuits and Electronics(ICDCECE).IEEE,2025:1-5. [8] ZHANG J,ZHANG Z,YESHI L,et al.Tibetan medical named entity recognition based on syllable-word-sentence embedding transformer[J].CAAI Transactions on Intelligence Technology,2025,10(4):1148-1158. [9] HU S,GUAN J,LI W.Chinese named entity recognition basedon multi-level information extraction[C]//Proceedings of the 2022 3rd International Conference on Electronic Communication and Artificial Intelligence(IWECAI).IEEE,2022:554-557. [10] DEVLIN J,CHANG M W,LEE K,et al.Bert:Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies.2019:4171-4186. [11] CONNEAU A,KHANDELWAL K,GOYAL N,et al.Unsupervised cross-lingual representation learning at scale[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.2020:8440-8451. [12] POLITOV A,SHKALIKOV O,JÄKEL R,et al.Revisiting Projection-based Data Transfer for Cross-Lingual Named Entity Recognition in Low-Resource Languages[C]//Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on HumanLanguage Technologies(NoDaLiDa/Baltic-HLT 2025).2025:499-507. [13] ZHU G,XIAO R,WANG H,et al.Large margin representation learning for robust cross-lingual named entity recognition[C]//Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics.2025:4270-4291. [14] AFRICA D D,SALHAN S,WEISS Y,et al.Meta-Pretrainingfor Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages[C]//Proceedings of the 5th Workshop on Multilingual Representation Learning(MRL 2025).2025:106-127. [15] RADCHENKO V,DRUSHCHAK N.Improving Named Entity Recognition for Low-Resource Languages Using Large Language Models:A Ukrainian Case Study[C]//Proceedings of the Fourth Ukrainian Natural Language Processing Workshop(UNLP 2025).2025:27-35. [16] ZHANG Z,LEE S Y M,ZHANG D,et al.Zero-shot Cross-lingual NER via Mitigating Language Difference:An Entity-aligned Translation Perspective[C]//Findings of the Association for Computational Linguistics:EMNLP 2025.2025:4541-4557. [17] CLARK J H,GARRETTE D,TURC I,et al.Canine:Pre-training an efficient tokenization-free encoder for language representation[J].Transactions of the Association for Computational Linguistics,2022,10:73-91. [18] XUE L,BARUA A,CONSTANT N,et al.ByT5:Towards a token-free future with pre-trained byte-to-byte models[J].Trans-actions of the Association for Computational Linguistics,2022,10:291-306. [19] LI S,XU J,ZHANG M.EMBYTE:Decomposition and Com-pression Learning for Small yet Private NLP[C]//Findings of the Association for Computational Linguistics:EMNLP 2025.2025:7182-7201. [20] ERONEN J,PTASZYNSKI M,MASUI F.Zero-shot cross-lingual transfer language selection using linguistic similarity[J].Information Processing & Management,2023,60(3):103250. [21] WANG P,CHEN Y,LIU C,et al.An attention-enhanced LSTM model for Chinese educational named entity recognition[C]//Proceedings of the 2024 5th International Conference on Information Science and Education(ICISE-IE).IEEE,2024:272-277. [22] RATCHATORN T,TANAKA M.Adaptive Adversarial Cross-Entropy Loss for Sharpness-Aware Minimization[C]//2024 IEEE International Conference on Image Processing(ICIP).IEEE,2024:479-485. [23] LIU J,SUN M,ZHANG W,et al.DAE-NER:dual-channel attention enhancement for Chinese named entity recognition[J].Computer Speech & Language,2024,85:101581. [24] VARSHNEY N,CHATTERJEE A,PARMAR M,et al.Investigating acceleration ofLLaMA inference by enabling interme-diate layer decoding via instruction tuning with ‘LITE’[C]//Findings of the Association for Computational Linguistics:NAACL 2024.2024:3656-3677. [25] LI J,FANG A,SMYRNIS G,et al.Datacomp-lm:In search ofthe next generation of training sets for language models[J].Advances in Neural Information Processing Systems,2024,37:14200-14282. [26] YU T,ZHANG Y,YONG C.Tibetan named entity recognition based on few-shot learning[J].Computer and Modernization,2023,39(5):13-19. [27] ZHOUM K,OUJIAN C R,DAOJI C D,et al.TibNER:a Tibetan named entity recognition dataset[J].China Scientific Data(Chinese & English Online Edition),2024,9(4):17-27. [28] LAMAOJIE,WANMAC D,YONGCUO,et al.Tibetan medical named entity recognition via adversarial training and iterative dilated convolution[J].Plateau Science Research,2025,9(1):105-118. |
|
||