Computer Science ›› 2026, Vol. 53 ›› Issue (9): 375-384.doi: 10.11896/jsjkx.250700002

• Computer Network • Previous Articles     Next Articles

DTD3-AQM:Satellite Network Queue Management Algorithm Based on Information Timeliness

WEI Debin1, WANG Xinrui1, YANG Li2, PAN Chengsheng3   

  1. 1 School of Information Engineering,Dalian University,Dalian,Liaoning 116622,China
    2 School of Automation,Nanjing University of Science and Technology,Nanjing 210094,China
    3 National Engineering Research Center for Communication and Network Technology,Nanjing University of Posts and Telecommunications,Nanjing 210003,China
  • Received:2025-07-01 Revised:2025-10-09 Online:2026-09-15 Published:2026-09-10
  • About author:WEI Debin,born in 1978,Ph.D,asso-ciate professor,is a member of CCF(No.22419M).His main research interests include space-ground integrated network transmission technology,traffic engineering and network optimization.
    PAN Chengsheng,born in 1962,Ph.D,professor,is a member of CCF(No.09154F).His main research interests include intelligent network traffic theory and key technologies.
  • Supported by:
    National Natural Science Foundation of China (U21B2003,61931004).

Abstract: Low Earth Orbit(LEO) satellite communication networks offer advantages such as broad coverage and low propagation delay.However,their dynamically changing topologies and bandwidth constraints may significantly impact the quality of service for time- sensitive applications.To address this issue,this paper proposes an intelligent active queue management algorithm,DTD3-AQM,which integrates the twin delayed deep deterministic policy gradient(TD3) framework with the dueling network architecture.This algorithm introduces time-varying topology matrices and multi-dimensional queue performance metrics to design a multi-objective reward mechanism,achieving congestion control while simultaneously considering data freshness.Simulation experiments conducted on a MATLAB-NS3 joint platform under a Linux environment demonstrate that DTD3-AQM achieves an optimal balance across multiple performance metrics.Under conditions with 200 TCP sources,compared to the system-perfor-mance-oriented QueuePilot,DTD3-AQM reduces queuing delay by 11.3 ms and average age of information(AoI) by 87.2 ms.Although it incurs a slight drop in throughput,it enhances congestion control capability and improves information timeliness.Compared to the freshness-focused DeepAAQM,DTD3-AQM increases throughput by 3.1 Mbps and reduces packet loss rate by 0.29%,though DeepAAQM still holds a significant advantage in reducing AoI.

Key words: Satellite network, Active queue management, Age of information, Reinforcement learning, Dueling network

CLC Number: 

  • TP393
[1] SONG T,KYUNG Y.Deep-reinforcement-learning-based age-of-information-aware low-power active queue management for IoT sensor networks[J].IEEE Internet of Things Journal,2024,11(9):16700-16709.
[2] PAPPAS N,GUNNARSSON J,KRATZ L,et al.Age of information of multiple sources with queue management[C] //2015 IEEE International Conference on Communications(ICC).IEEE,2015:5935-5940.
[3] FANG Z,WANG J,JIANG C,et al.Average peak age of information in underwater information collection with sleep-scheduling[J].IEEE Transactions on Vehicular Technology,2022,71(9):10132-10136.
[4] XU C L,CHEN Z Y,TAO M X,et al.Wireless Multi-User Interactive Virtual Reality in Metaverse with Edge-Device Colla-borative Computing[J].IEEE Transactions on Wireless Communications,2025,27(7):6135-6150.
[5] SU Y,HUANG L,FENG C.QRED:A Q-learning-based active queue management scheme[J].Journal of Internet Technology,2018,19(4):1169-1178.
[6] BYUN D J,BARAS J S.Adaptive virtual queue random earlydetection in satellite networks[M] // Wireless Technology:Applications,Management,and Security.2009:63-82.
[7] YOUSIF A,HASSAN H,MUTTASHER G.Applying rein-forcement learning for random early detection algorithm in adaptive queue management systems[J].Indonesian Journal of Electrical Engineering and Computer Science,2022,26(3):1684.
[8] LIU J Y,WEI D B.Active queue management based on Q-lear-ning traffic predictor[C] //2022 International Conference on Cyber-Physical Social Intelligence(ICCSI).IEEE,2022:399-404.
[9] KIM M,JASEEMUDDIN M,ANPALAGAN A.Deep reinforce-ment learning based active queue management for IoT networks[J].Journal of Network and Systems Management,2021,29(3):34.
[10] FORERO P A,ZHANG P,RADOSEVIC D.Active queue-management policies for undersea networking via deep reinforcement learning[C] //OCEANS 2021:San Diego-Porto.IEEE,2021:1-8.
[11] DERY M,KRUPNIK O,KESLASSY I.Queuepilot:Revivingsmall buffers with a learned aqm policy[C] //IEEE INFOCOM 2023-IEEE Conference on Computer Communications.IEEE,2023:1-10.
[12] HUI T F,ZHAI S H,GONG F K,et al.Research status and development trends of satellite communication multiple access technology[J].Space Electronic Technology,2026,23(1):116-134.
[13] LIU Y,SHI Y,LI J G,et al.Reconnaissance and jamming methods for low earth orbit constellations[J].Space Electronic Technology,2026,23(2):1-8.
[14] YAN Y,WANG L H,TIAN Z,et al.Research on probabilistic routing broadcast protocol for low-orbit satellite networks[J].Space Electronic Technology,2026,23(1):135-144.
[15] GRAZIA C A,PATRICIELLO N,KLAPEZ M,et al.Mitigating congestion and bufferbloat on satellite networks through a rate-based AQM[C] //2017 IEEE International Conference on Communications(ICC).IEEE,2017:1-6.
[16] MA Y,ZHANG T,ZHANG J.An AQM algorithm for LEO sa-tellite network[C] //2011 4th IEEE International Symposium on Microwave,Antenna,Propagation and EMC Technologies for Wireless Communications.IEEE,2011:615-618.
[17] YATES R D,SUN Y,BROWN D R,et al.Age of information:An introduction and survey[J].IEEE Journal on Selected Areas in Communications,2021,39(5):1183-1210.
[18] SUTTON R S,BARTO A G.Reinforcement learning[J].Journal of Cognitive Neuroscience,1999,11(1):126-134.
[19] LILLICRAP T P,HUNT J J,PRITZEL A,et al.Continuouscontrol with deep reinforcement learning[J].arXiv:1509.02971,2015.
[20] FUJIMOTO S,HOOF H,MEGER D.Addressing function approximation error in actor-critic methods[C] //International Conference on Machine Learning.PMLR,2018:1587-1596.
[21] WANG Z,SCHAUL T,HESSEL M,et al.Dueling network architectures for deep reinforcement learning[C] //International Conference on Machine Learning.PMLR,2016:1995-2003.
[22] SHUKUR HH A,AHMAD Y A,YUNUS M S F M,et al.Latency Performance Evaluation of LEO Starlink and SES-12 GEO HTS Network Under Tropical Rainfall Conditions [J].IIUM Engineering Journal,2025,26(2):204-219.
[23] LAI Z,LIU C,WANG X,et al.SpaceRTC:Unleashing the low-latency potential of mega-constellations for real-time communications [C] //IEEE INFOCOM 2022-IEEE Conference on Computer Communications.IEEE,2022:1339-1348.
[24] ZHANG T.Toward automated vehicle teleoperation:Vision,opportunities,and challenges [J].IEEE Internet of Things Journal,2020,7(12):11347-11354.
[25] MA H,XU D,DAI Y,et al.An intelligent scheme for congestion control:When active queue management meets deep reinforcement learning[J].Computer Networks,2021,200:108515.
[1] XIA Bin, SU Jinya, CAO Junmin, XIAO Junwen, SUN Guozi. Research on Log Anomaly Prediction Method for Log Sequence Graph Construction Guided byReinforcement Learning [J]. Computer Science, 2026, 53(9): 405-413.
[2] HE Yulin, XIAO Youqi, YANG Zhenyu, HUANG Zhexue, CUI Laizhong. Adaptive Frequency Tuning Approach for Spark Clusters [J]. Computer Science, 2026, 53(8): 71-84.
[3] LIAN Zhaoyang, SI Bailu. Optimization in Cross-field of Manufacturing and Transportation by Combining Reinforcement Learning and Artificial Hummingbird Algorithm [J]. Computer Science, 2026, 53(8): 307-315.
[4] WANG Hongguang, JIANG Yiming, LIU Xiajun, BAI Luxin. Self-adaptive Load Balancing Strategy Based on Reinforcement Learning for SDSN [J]. Computer Science, 2026, 53(7): 336-342.
[5] SHANG Kefeng, ZHANG Dan, ZHUAN Sunying, LI Dandan, LIU Yan, ZHU Kaige. Multi-party Inter-satellite Collaborative Computing Offloading Algorithm for Time-varying Topologies and Dynamic Heterogeneous Resources [J]. Computer Science, 2026, 53(7): 354-362.
[6] GONG Jing, YANG Yufa, ZHENG Yifan, SUN Zhixin. Multi-objective Intelligent Warehousing Path Planning Based on Conflict Free Path Algorithm [J]. Computer Science, 2026, 53(4): 88-100.
[7] LIU Jiaqi, WANG Yujie, XIANG Guodu, YU Kui, CAO Fuyuan. Long-term Causal Effect Estimation Based on Deep Reinforcement Learning [J]. Computer Science, 2026, 53(4): 235-244.
[8] PAN Jiahao, FENG Xiang, YU Huiqun. SM-PHT:Robust,Scalable,and Efficient Method for Multi-task Reinforcement Learning [J]. Computer Science, 2026, 53(4): 366-376.
[9] ZHENG Cheng, BAN Qingqing. Knowledge-assisted and Reinforced Syntax-driven for Aspect-based Sentiment Analysis [J]. Computer Science, 2026, 53(4): 406-414.
[10] ZHAI Jie, LI Yanhao, CHEN Lexuan, GUO Weibin. Dynamic Recommendation of Personalized Hands-on Learning Materials Based on LightweightEducational LLMs [J]. Computer Science, 2026, 53(2): 48-56.
[11] LI Fang, YUAN Baochun, SHEN Hang, WANG Tianjing, BAI Guangwei. Deep Reinforcement Learning-based Aircraft Task Offloading in Low Earth Orbit Satellite Networks [J]. Computer Science, 2026, 53(2): 406-415.
[12] WAN Shenghua, XU Xingye, GAN Le, ZHAN Dechuan. Pre-training World Models from Videos with Generated Actions by Multi-modal Large Models [J]. Computer Science, 2026, 53(1): 51-57.
[13] WANG Haoyan, LI Chongshou, LI Tianrui. Reinforcement Learning Method for Solving Flexible Job Shop Scheduling Problem Based onDouble Layer Attention Network [J]. Computer Science, 2026, 53(1): 231-240.
[14] DUAN Pengting, WEN Chao, WANG Baoping, WANG Zhenni. Collaborative Semantics Fusion for Multi-agent Behavior Decision-making [J]. Computer Science, 2026, 53(1): 252-261.
[15] ZHU Shihao, PENG Kexing, MA Tinghuai. Graph Attention-based Grouped Multi-agent Reinforcement Learning Method [J]. Computer Science, 2025, 52(9): 330-336.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!