计算机科学 ›› 2026, Vol. 53 ›› Issue (8): 29-39.doi: 10.11896/jsjkx.260300019

• 数据库 & 大数据 & 数据科学 • 上一篇    下一篇

面向持续葡萄糖监测解读的时间序列语言模型

王昱麒1,2, 张仰森1,3, 郭亚龙1,3, 亢静1,3, 王雅伦4   

  1. 1 北京信息科技大学智能信息处理研究所 北京 100192
    2 北京信息科技大学仪器科学与光电工程学院 北京 102206
    3 北京信息科技大学计算机学院 北京 102206
    4 首都医科大学附属北京佑安医院 北京 100069
  • 收稿日期:2026-03-04 修回日期:2026-05-29 出版日期:2026-08-15 发布日期:2026-08-17
  • 通讯作者: 张仰森(zys@bistu.edu.cn)
  • 作者简介:(wangyuqi2023@bistu.edu.cn)
  • 基金资助:
    北京市自然科学基金(L233008)

Time Series Language Model for Continuous Glucose Monitoring Interpretation

WANG Yuqi1,2, ZHANG Yangsen1,3, GUO Yalong1,3, KANG Jing1,3, WANG Yalun4   

  1. 1 Institute of Intelligent Information Processing, Beijing Information Science and Technology University, Beijing 100192, China
    2 School of Instrumentation Science and OPTO-Electronics Engineering, Beijing Information Science and Technology University, Beijing 102206, China
    3 College of Computer Science, Beijing Information Science and Technology University, Beijing 102206, China
    4 Beijing Youan Hospital, Capital Medical University, Beijing 100069, China
  • Received:2026-03-04 Revised:2026-05-29 Published:2026-08-15 Online:2026-08-17
  • About author:WANG Yuqi,born in 1999,postgra-duate,is a member of CCF(No.Q9032G).Her main research interests include trusted medical large model and multimodal information processing.
    ZHANG Yangsen,born in 1962,professor,Ph.D supervisor,is a distinguished member of CCF(No.16640D).His main research interests include information mining and language security go-vernance.
  • Supported by:
    Natural Science Foundation of Beijing,China(L233008).

摘要: 解读糖尿病患者的持续葡萄糖监测(CGM)序列有助于患者的血糖管理。但大语言模型(LLM)存在误判率高和因缺乏对序列内在规律的理解而产生与实际状况不符的内容的缺陷。为此,提出面向持续葡萄糖监测解读的时间序列语言模型(Time Series Language Model for CGM Interpretation,CGM-TSLM)和基于“血糖序列-语言描述”对的CGM解读数据集(GLiDCGM)。模型采用一维卷积神经网络的时间序列编码器捕捉CGM序列的关键量化特征,利用语言模型编码提示指令和生成与实际血糖状况匹配的自然语言描述,引入注意力机制进行特征的对齐融合。最后,基于多模态监督微调,完成CGM序列到文本的转换。数据集构建采用了基于模糊逻辑的标注方法,生成包含波动等血糖特征的初文,再借助LLM将其整合为简洁准确的描述。GLiDCGM数据集上的实验表明,CGM-TSLM模型生成的描述优于LLaMA,Qwen,T5,BART等基线模型,词汇重叠度和文本相似度平均提高了12.22个百分点和14.45个百分点,证明了CGM-TSLM提升模型生成CGM序列总结的有效性,为可穿戴设备生理数据分析提供了理论和数据支撑。

关键词: 持续葡萄糖监测解读, 时间序列语言模型, 多模态信息处理, 时间序列到文本, 可穿戴设备

Abstract: The interpretation of continuous glucose monitoring(CGM) signals for diabetic patients is crucial for glycemic management.However,large language model(LLM) for CGM is found the deficiency of high misinterpretation rates and the generation that conflicts with actual glucose due to the poor understanding of time series.To address these issues,the time series language model for CGM interpretation(CGM-TSLM) and a CGM interpretation dataset(GLiDCGM) of “glucose series-language description” pairs are proposed.A one-dimensional convolutional neural network time-series encoder is involved to capture the key quantitative features of the CGM series,a language model is employed to encode prompt instructions and generate matching language,and an attention mechanism is introduced to align and fuse features.Finally,the transport of the CGM series to text is completed by the multimodal supervised fine-tuning method.For the GLiDCGM dataset construction,the fuzzy logic text annotation method is adopted to generate initial descriptions of curve features,and optimized into the concise and accurate medical summaries by LLM.Experiments on the GLiDCGM dataset show that,the description generation of CGM-TSLM model is superior than baseline models such as LLaMA2-7B-Chat,Qwen,T5,and BART,with average improvements of 12.22 percentage points and 14.45 percentage points on word overlap and text similarity,respectively.Experimental results prove that CGM-TSLM is able to further enhance the generation ability of CGM series summaries and provide theoretical and data support for the analysis of wearable devices physiological data.

Key words: Continuous glucose monitoring, Time series language model, Multimodal information processing, Time series to text, Wearable devices

中图分类号: 

  • TP399
[1] WANG Y,ZHANG Y,LIU Z,et al.Large Language Model Reasoning Framework for Diabetes Prevention and Treatment via Personal Health Data[C]//Proceedings of the 2025 International Conference on Intelligent Computing.Springer,2025:374-385.
[2] BELLIDO V,AGUILERA E,CARDONA-HERNANDEZ R,et al.Expert Recommendations for Using Time-in-Range and Other Continuous Glucose Monitoring Metrics to Achieve Patient-Centered Glycemic Control in People with Diabetes[J].Journal of Diabetes Science and Technology.2023,17(5):1326-1336.
[3] Diabetes Nursing Committee of Chinese Nursing Association,National Innovation Center for Advanced Medical Devices.Expert consensus on nursing applications of ambulatory glucose profile report(2025 edition)[J].Chinese Journal of Diabetes,2025,17(3):311-321.
[4] CHENG M,CHEN Y,LIU Q,et al.InstructTime:Advancing Time Series Classification with Multimodal Language Modeling[C]//Proceedings of the 18th ACM International Conference on Web Search and Data Mining.New York:ACM,2025:792-800.
[5] MARÍN N,SÁNCHEZ D.On generating linguistic descriptions of time series[J].Fuzzy Sets and Systems,2016,285(15):6-30.
[6] ZHANG L,CHEN C,LIANG Y.Hallucinations Proactive Relief in Diabetes Q&A LLM[J].Computer Science,2025,52(S1):55-64.
[7] KACPRZYK J,YAGER R R,ZADROZNY S.Fuzzy Linguistic Summaries of Databases for an Efficient Business Data Analysis and Decision Support[M]//Knowledge Discovery for Business Information Systems.Springer,2002:129-152.
[8] WILBIK A,BARRETO D,BACKUS G.On Relevance of Lin-guistic Summaries-A Case Study from the Agro-Food Domain[C]//Proceedings of Information Processing and Management of Uncertainty in Knowledge-Based Systems.Springer,2020:289-300.
[9] MARTINEZ-CRUZ C,RUEDA A J,POPESCU M,et al.New Linguistic Description Approach for Time Series and its Application to Bed Restlessness Monitoring for Eldercare[J].IEEE Transactions on Fuzzy Systems,2022,30(4):1048-1059.
[10] GENC S,AKAY D,BORAN F E,et al.Yager,Linguistic summarization of fuzzy social and economic networks:An application on the international trade network[J].Soft Computing,2020,24:1511-1527.
[11] CASCALLAR-FUENTES A,GALLEGO-FERNÁNDEZ J,RAMOS-SOTO A,et al.Automatic generation of textual descriptions in data-to-text systems using a fuzzy temporal onto-logy:Application in air quality index data series[J].Applied Soft Computing.2022,119:108612.
[12] NISHIMURA R,OKADA Y,KURODA A,et al.Consensusstatement on novel glucose-related metrics obtained through advanced medical devices:English version[J].Diabetology International,2025,16:595-613.
[13] GAITÁN-GUERRERO J F,RUIZ J L L.Toward an Interpretable Continuous Glucose Monitoring Data Modeling[J].IEEE Internet of Things Journal,2024,11(19):31080-31094.
[14] HEALEY E,TAN A L M T,KRISTEN L F,et al.A case study on using a large language model to analyze continuous glucose monitoring data[J].Science Reports,2025,15:1143.
[15] MARTINEZ-CRUZ C,GUERRERO J F G,RUIZ J L L,et al.A First Approach to the Generation of Linguistic Summaries from Glucose Sensors Using GPT-4[C]//Proceedings of the 15th International Conference on Ubiquitous Computing & Ambient Intelligence.Springer,2023:33-43.
[16] HEALEY E,KOHANE I.LLM-CGM:A Benchmark for Large Language Model-Enabled Querying of Continuous Glucose Monitoring Data for Conversational Diabetes Management[J].Pacific Symposium on Biocomputing,2025,30:82-93.
[17] LU Y,LIU D,LIANG Z,et al.A pretrained transformer model for decoding individual glucose dynamics from continuous glucose monitoring data[J].National Science Review,2025,12(5):nwaf039.
[18] TRABELSI M,BOYD A,CAO J,et al.Time Series LanguageModel for Descriptive Caption Generation[J].Engineering Applications of Artificial Intelligence,2025,162:112673.
[19] CHOW W,GARDINER L,HALLGRÍMSSON H T,et al.Towards Time Series Reasoning with LLMs[C]//NeurIPS 2024.2024.
[20] RAFFEL C,SHAZEER N,ROBERTS A,et al.Exploring the limits of transfer learning with a unified text-to-text transformer[J].The Journal of Machine Learning Research,2020,21(1):5485-5551.
[21] RODRIGUEZ-LEON C,AVILES-PEREZ M D,BANOS O,et al.T1DiabetesGranada:a longitudinal multi-modal dataset of type 1 diabetes mellitus[J].Science Data,2023,10:916.
[22] LIN C Y.ROUGE:A Package for Automatic Evaluation ofSummaries[C]//Text Summarization Branches Out,Association for Computational Linguistics.2004:74-81.
[23] CHEN J,XIAO S,ZHANG P,et al.M3-Embedding:Multi-Linguality,Multi-Functionality,Multi-Granularity Text Embeddings Through Self-Knowledge Distillation[C]//Findings of the 2024 Association for Computational Linguistics.Stroudsburg:ACL,2024:2318-2335.
[24] HARSH J,TAYLOR B K.Truth-Conditional Captions for TimeSeries Data[C]//Proceedings of the 2021 Conference on Empi-rical Methods in Natural Language Processing.Stroudsburg:ACL,2021:719-733.
[25] WANG Y Q,ZHANG Y S,WANG P.High school Englishreading comprehension model integrating knowledge augmentation and contrastive learning[J/OL].Journal of Computer Applications,1-10.[2026-03-03].https://link.cnki.net/urlid/51.1307.tp.20251014.1829.004.
[26] TOUVRON H,MARTIN L,STONE K,et al.Llama 2:OpenFoundation and Fine-Tuned Chat Models[J].arXiv:2307.09288,2023.
[27] QWEN.Qwen2.5 Technical Report[J].arXiv:2412.15115,2024.
[28] LEWIS M,LIU Y,GOYAL N,et al.BART:Denoising Se-quence-to-Sequence Pre-training for Natural Language Generation,Translation,and Comprehension[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.Stroudsburg:ACL,2020:7871-7880.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!