计算机科学 ›› 2026, Vol. 53 ›› Issue (7): 336-342.doi: 10.11896/jsjkx.250500004

• 计算机网络 • 上一篇    下一篇

基于强化学习的软件定义卫星互联网自适应负载均衡策略

汪红光1,3, 江逸茗1,2, 刘侠君3, 白禄鑫1   

  1. 1 信息工程大学信息技术研究所 郑州 450008
    2 网络空间安全教育部重点实验室 郑州 450008
    3 中国人民解放军 66389部队 天津 300100
  • 收稿日期:2025-05-06 修回日期:2025-12-23 出版日期:2026-07-15 发布日期:2026-07-10
  • 通讯作者: 江逸茗(j8403@163.com)
  • 作者简介:(nudtwanghongguang@126.com)
  • 基金资助:
    雄安新区科技创新专项(2022XAGG0111);河南省重大专项课题(22110021090003)

Self-adaptive Load Balancing Strategy Based on Reinforcement Learning for SDSN

WANG Hongguang1,3, JIANG Yiming1,2, LIU Xiajun3, BAI Luxin1   

  1. 1 Institute of Information Technology,Information Engineering University,Zhengzhou 450008,China
    2 Key Laboratory of Cyberspace Security,Ministry of Education of China,Zhengzhou 450008,China
    3 Unit 66389 of PLA,Tianjin 300100,China
  • Received:2025-05-06 Revised:2025-12-23 Published:2026-07-15 Online:2026-07-10
  • About author:WANG Hongguang,born in 1992,postgraduate.His main research interests include satellite Internet and software-defined networking.
    JIANG Yiming,born in 1984,Ph.D,associate researcher.His main research interests include satellite Internet,vir-tualization,and so on.
  • Supported by:
    Xiong'an New Area Science and Technology Innovation Special Project(2022XAGG0111)and Henan Province Major Special Project Topic(22110021090003)

摘要: 针对卫星互联网计算资源有限、链路动态时变、流量分布不均等特点,提出一种基于强化学习的自适应负载均衡策略。借助软件定义卫星互联网控制与转发分离的结构优势,设计2跳范围的软件定义卫星网络区域划分算法,针对同轨/异轨链路的通信质量差异,引入链路状态值(μ)和权重值(w)评估链路性能,确保优先选择同轨低时延链路。基于Actor-Critic深度强化学习框架,设计关键流选择模型SALB-RL,采用多智能体异步训练策略对模型进行训练,通过线性规划计算流量重分布比例,在最小化最大链路利用率的同时降低端到端时延。使用STK(Systems Tool Kit)工具构建低轨Walker星座,并根据卫星网络拓扑结构生成卫星网络流量数据,用于模型训练和实验验证。实验结果表明,SALB-RL仅需对10%的关键流进行重分布即可达到全网流量重分布95%以上的负载均衡性能,与典型卫星互联网强化学习模型和传统地面网络负载均衡优化算法相比,平均负载均衡性能提升了约3%,链路时延性能更加稳定。总之,SALB-RL算法能够兼顾网络负载均衡和路由计算开销,为动态卫星网络的智能管理提供了高效的解决方案。

关键词: 软件定义卫星网络, 负载均衡策略, 强化学习, 关键流

Abstract: Given the challenges of satellite Internet,including limited computational resources,dynamically time-varying links,and uneven traffic distribution,this paper proposes an self-adaptive load balancing strategy based on reinforcement learning.Leveraging the control-data plane separation architecture of software-defined satellite networks(SDSN),a two-hop regional partitioning algorithm for SDSN is designed.To address the communication quality disparities between intra-orbit and inter-orbit links,link state values(μ) and weight values(w) are introduced to quantify link performance,prioritizing intra-orbit low-latency links.Built upon the Actor-Critic deep reinforcement learning framework,the SALB-RL model employs multi-agent asynchronous training to optimize key flow selection.Traffic redistribution ratios are computed via linear programming to minimize maximum link utilization while reducing end-to-end delay.A low earth orbit(LEO) Walker constellation is constructed using STK(Systems Tool Kit),a leading system engineering software,and traffic datasets derived from dynamic network topologies are used for training and validation.Experimental results demonstrate that SALB-RL achieves over 95% network-wide load balancing performance by redistributing only 10% of critical flows.Compared with state-of-the-art satellite Internet DRL models and traditional terrestrial load balancing algorithms,SALB-RL improves average load balancing performance by about 3% while ensuring more stable delay characteristics.This work highlights that SALB-RL effectively balances load balancing efficiency and routing overhead,offering an optimal solution for intelligent management of dynamic satellite networks.

Key words: Software-defined satellite network, Load balancing strategy, Reinforcement learning, Critical flow

中图分类号: 

  • TN927
[1]RAY P P.A review on 6G for space-air-ground integrated network:Key enablers,open challenges,and future direction[J].Journal of King Saud University-Computer and Information Sciences,2022,34(9):6949-6976.
[2]KREUTZ D,RAMOS F M V,ESTEVES V P,et al.Software-Defined Networking:A Comprehensive Survey[J].Proceedings of the IEEE,2015,103(1):14-76.
[3]GOPAL R,RAVISHANKAR C.Software Defined Satellite Networks[C]//32nd AIAA International Communications Satellite Systems Conference.2014:2014-4480.
[4]ZHU Y,QIAN L,DING L,et al.Software defined routing algorithm in LEO satellite networks[C]//2017 International Conference on Electrical Engineering and Informatics(ICELTICs).IEEE,2017:257-262.
[5]CAO X,LI Y,XIONG X,et al.Dynamic Routings in Satellite Networks:An Overview[J].Sensors,2022,22(12):4552.
[6]WEI L H,LIU G W,LIU Y,et al.Research on Routing Optimization in Satellite Internet Based on Deep Reinforcement Lear-ning[J].Space-integrated-ground Information Networks,2022,3(3):65-71.
[7]WU Y,ZHOU S,WEI Y,et al.Deep Reinforcement Learning for Controller Placement in Software Defined Network[C]//IEEE INFOCOM 2020-IEEE Conference on Computer Communications Workshops(INFOCOM WKSHPS).IEEE,2020:1254-1259.
[8]ZHANG S,LIU A,HAN C,et al.Graph Neural Network and Reinforcement Learning Based Routing for Mega LEO Satellite Constellations[C]//2023 9th International Conference on Computer and Communications(ICCC).IEEE,2023:1-6.
[9]PAPA A,DE COLA T,VIZARRETA P,et al.Dynamic SDNController Placement in a LEO Constellation Satellite Network[C]//2018 IEEE Global Communications Conference(GLOBECOM).IEEE,2018.
[10]MAO H,ALIZADEH M,MENACHE I,et al.Resource Management with Deep Reinforcement Learning[C]//Proceedings of the 15th ACM Workshop on Hot Topics in Networks.ACM,2016:50-56.
[11]DEBOPAM B,ANKIT S.Network topology design at 27 000km/hour[C]//Proceedings of the 15th International Conference on Emerging Networking Experiments And Technologies.ACM,2019:341-354.
[12]YANG M Q,DONG X R,HU M,et al.Design and Simulation for Hybrid LEO Communication and Navigation Constellation[C]//2016 IEEE Chinese Guidance,Navigation and Control Conference.Nanjing,China:IEEE,2016:1665-1669.
[13]CHEN Q,GIAMBENE G,YANG L,et al.Analysis of Inter-Satellite Link Paths for LEO Mega-Constellation Networks[J].IEEE Transactions on Vehicular Technology,2021,70(3):2743-2755.
[14]ZHAO P,LIU J,ZHANG R,et al.Self-Healing Motif-Based Distributed Routing Algorithm for Mega-Constellation[C]//2022 5th International Conference on Hot Information-Centric Networking(HotICN).IEEE,2022:90-98.
[15]WANG C,WANG H,WANG W.A Two-Hops State-Aware Routing Strategy Based on Deep Reinforcement Learning for LEO Satellite Networks[J].Electronics,2019,8(9):920.
[16]CHAUDHRY A U,YANIKOMEROGLU H.Laser Intersatellite Links in a Starlink Constellation:A Classification and Ana-lysis[J].IEEE Vehicular Technology Magazine,2021,16(2):48-56.
[17]YUAN S ZHANG H,CAI A L,et al.A Specific Flow Routing Selection Algorithm Based on Self-Attention Deep Reinforcement Learning [J].Optical Communication Technology,2024,48(3):7-12.
[18]ZHANG J,YE M,GUO Z,et al.CFR-RL:Traffic Engineering With Reinforcement Learning in SDN[J].IEEE Journal on Selected Areas in Communications,2020,38(10):2249-2259.
[19]LIU X,ZHOU H,ZHANG Z,et al.Multipath Cooperative Routing in Ultradense LEO Satellite Networks:A Deep-Reinforcement-Learning-Based Approach[J].IEEE Internet of Things Journal,2025,12(2):1789-1804.
[20]HUANG Y,JIANG X,CHEN S,et al.Pheromone Incentivized Intelligent Multipath Traffic Scheduling Approach for LEO Satellite Networks[J].IEEE Transactions on Wireless Communications,2022,21(8):5889-5902.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!