计算机科学 ›› 2026, Vol. 53 ›› Issue (6): 171-184.doi: 10.11896/jsjkx.250800064
吴璨, 肖海力, 王小宁, 赵一宁, 卢莎莎, 和荣
WU Can, XIAO Haili, WANG Xiaoning, ZHAO Yining, LU Shasha, HE Rong
摘要: 在高性能计算(HPC)领域,机器学习驱动的调度算法正逐渐成为研究的热点,这类算法的性能优化对工作负载数据的质量要求较高。因此,开展面向HPC环境的工作负载分析与建模方法研究,对于提升调度算法的效率和适应性具有重要意义。基于超级计算中心公开的运行日志,对HPC工作负载进行了深入分析,并形成了一套分析与建模方法。首先采用DBSCAN算法清洗异常数据,并系统分析了作业到达模式、处理器核数分布特征、应用分布特征、运行时间分布特征以及属性间的相关性。然后基于这些分析结果,开发了一个灵活的高性能计算工作负载生成器,不仅支持默认参数配置,还能自动分析SWF格式的工作负载特征,以满足不同场景的仿真需求。经过性能实验验证,所构建的生成器产生的数据相比于随机生成的数据,更贴近真实数据的特征分布,可以为新型调度算法的研发提供高质量的训练数据,对提升超级计算中心的资源使用效率具有重要的意义。
中图分类号:
| [1]XIANG Y,YANG X M,SUN Y,et al.A Fault-tolerant andCost-efficient Workflow Scheduling Approach Based on Deep Reinforcement Learning for IT Operation and Maintenance[C]//International Conference on Computer Supported Cooperative Work in Design.2023:411-416. [2]WANG B Y,LI H F,LIN Z W,et al.Temporal Fusion Pointer network-based Reinforcement Learning algorithm for Multi-Objective Workflow Scheduling in the cloud[C]//International Joint Conference on Neural Networks.2020:1-8. [3]ASGHARI A,SOHRABI M K,YAGHMAEE F.Online scheduling of dependent tasks of cloud's workflows to enhance resource utilization and reduce the makespan using multiple reinforcement learning-based agents[J].Soft Computing,2020,24:16177-16199. [4]WEI Y,KUDENKO D,LIU S,et al.A Reinforcement Learning Based Workflow Application Scheduling Approach in Dynamic Cloud Environment[C]//International Conference on Collaborative Computing:Networking,Applications and Worksharing.CollaborateCom,2017:120-131. [5]DONG T T,XUE F,TANG H L,et al.Deep reinforcementlearning for fault-tolerant workflow scheduling in cloud environment[J].Applied Intelligence,2022,53(9):9916-9932. [6]LUBLIN U,FEITELSON D.The workload on parallel super-computers:modeling the characteristics of rigid jobs[J].Parallel Distributed Computing,2003,63:1105-1122. [7]CIRNE W,BERMAN F.A comprehensive model of the supercomputer workload[C]//Proceedings of the Fourth Annual IEEE International Workshop on Workload Characterization.2001:140-148. [8]PATEL T,LIU Z C,KETTIMUTHU R,et al.Job Characteris-tics on Large-Scale Systems:Long-Term Analysis,Quantification,and Implications[C]//International Conference for High Performance Computing,Networking,Storage and Analysis.2020:1-17. [9]中国高性能计算工作负载库[EB/OL].https://git.ustc.edu.cn/shenyu/CSWA.git. [10]WANG Q Q,LI J,WANG S,et al.A Novel Two-Step Job Run-time Estimation Method Based on Input Parameters in HPC System[C]//International Conference on Cloud Computing and Big Data Analysis.2019:311-316. [11]WANG Q Q,SHEN Y,LI J.User-level Workload Analysis for Supercomputers[C]//Conference on Software Engineering and Information Management.2021:16-18. [12]FEITELSON D,TSAFIRI D,KRAKOV D.Experience withusing the Parallel Workloads Archive[J].Parallel and Distributed Computing,2014,74(10):2967-2982. [13]LOSUP A,EPEMA D.Grid Computing Workloads[J].IEEEInternet Computing,2011,15(2):19-26. [14]LOSUP A,SONMEZ O,ANOEP S,et al.The performance of bags-of-tasks in large-scale distributed systems[C]//High Performance Distributed Computing.2008:97-108. [15]CARVALHO M,BRASILEIRO F.A User-Based Model of Grid Computing Workloads[C]//International Conference on Grid Computing.2012:40-48. [16]LOSUP A,JAN M,SONMEZ O,et al.The Characteristics and Performance of Groups of Jobs in Grids[C]//International Euro-Par Conference on Parallel Processing.2007:382-393. [17]SCHLAGKAMP S,SILVA R F D,ALLCOCK W,et al.Conse-cutive Job Submission Behavior at Mira Supercomputer[C]//International Symposium on High-Performance Parallel and Distributed Computing.2016:93-96. [18]RODRIGO G P,OSTBERG P O,ELMROTH E,et al.Towards understanding HPC users and systems:A NERSC case study[J].Parallel and Distributed Computing,2018,111:206-221. [19]FAN Y P,LAN Z L,CHILDERS T,et al.Deep Reinforcement Agent for Scheduling in HPC[C]//International Parallel and Distributed Processing Symposium.2021:807-816. |
|
||