计算机科学 ›› 2026, Vol. 53 ›› Issue (6): 153-162.doi: 10.11896/jsjkx.251000113
李金优1, 张文帅2, 沈瑜2, 张运动2, 李会民2, 李京1
LI Jinyou1, ZHANG Wenshuai2, SHEN Yu2, ZHANG Yundong2, LI Huimin2, LI Jing1
摘要: 在高性能计算领域,超级计算机平台利用大规模计算资源池实现多任务的并行处理,显著提升了作业执行速度。其资源调度机制要求用户提交作业时显式指定相应请求资源及其数目,而当前实践仍高度依赖人工经验的并行参数配置,往往导致硬件资源利用率难以逼近理论最优值,形成显著的资源利用效率瓶颈。对于结构复杂的超算软件,其输入参数与实际硬件资源紧密耦合,仅调整个别参数难以构建简明有效的框架以充分探索参数空间,而遍历整个参数搜索空间的应用配置代价极为昂贵。基于机器学习的并行计算应用运行时优化方法,通过多个高效模型分析处理大量运行数据来提供更好的参数配置建议,无需用户指定资源量即可自动化确定作业申请的硬件参数配置,从而降低集群软件使用门槛,提高工作效率,有效解决现有技术中超算集群用户配置并行参数效率低下以及特定任务高质量作业运行数据集匮乏的问题。实验结果表明,相比默认或普遍通用的参数配置,基于机器学习方法给出的参数配置在多例典型VASP实例上实现了普遍的加速效果和较低的机时开销,具有显著的推广实践价值。
中图分类号:
| [1]DIAZ J,MUNOZ-CARO C,NINO A.A survey of parallel programming models and tools in the multi and many-core era[J].IEEE Transactions onParallel and Distributed Systems,2012,23(8):1369-1386. [2]HACKENBERG D,JUCKELAND G,BRUNST H.Performanceanalysis of multi-level parallelism:inter-node,intra-node and hardware accelerators[J].Concurrency and Computation:Practice and Experience,2012,24(1):62-72. [3]TU B B,ZOU M,ZHAN J F,et al.Research on a hierarchical parallel computing model in the memory layer of multicore processor clusters [J].Journal of Computer Research and Development,2008,31(11):1948-1955. [4]JINNOUCHI R,LAHNSTEINER J,KARSAI F,et al.Phasetransitions of hybrid perovskites simulated by machine-learning force fields trained on the fly withbayesian inference[J].Physical Review Letters,2019,122(22):225701. [5]JINNOUCHI R,KARSAI F,KRESSE G.On-the-fly machine learning force field generation:Application to melting points[J].Physical Review B,2019,100(1):014105. [6]THOMPSON A P,AKTULGA H M,BERGER R,et al.Lammps-a flexible simulation tool for particle-based materials modeling at the atomic,meso,and continuum scales[J].Computer Physics Communications,2022,271:108171. [7]KRESSE G,FURTHMÜLLER J.Efficient iterative schemes for ab initio total-energy calculations using a plane-wave basis set[J].PhysicalReview B,1996,54(16):11169. [8]STEGAILOV V,SMIRNOV G,VECHER V.Vasp hits thememory wall:Processors efficiency comparison[J].Concurrency and Computation:Practice and Experience,2019,31(19):e5136. [9]CORSETTI F.Performance analysis of electronic structurecodes onhpc systems:a case study of siesta[J].PLoS One,2014,9(4):e95390. [10]KRESSE G,HAFNER J.Ab initio molecular-dynamics simulation of the liquid-metal-amorphous-semiconductor transition in germanium[J].Physical Review B,1994,49(20):14251. [11]BOWLER D R,MIYAZAKI T.methods in electronic structure calculations[J].Reports on Progress in Physics,2012,75(3):036503. [12]HOU Z,SHEN H,ZHOU X,et al.Prediction of job characteristics for intelligent resource allocation inhpc systems:a survey and future directions [J].Frontiers of Computer Science,2022,16(5):165107. [13]MARIANI G,ANGHEL A,JONGERIUS R,et al.Predictingcloud performance forhpc applications before deployment[J].Future Generation Computer Systems,2018,87:618-628. [14]GARCIA-GASULLA M,BANCHELLI F,PEIRO K,et al.A generic performance analysis technique applied to differentcfd methods for hpc[J].International Journal of Computational Fluid Dynamics,2020,34(7/8):508-528. [15]GAVIN J P.Optimisation of VASP density functional theory(DFT)/hartree-fock(HF)hybrid functional code using mixed-mode parallelism[R/OL].HECTOR Academy,2012[2023-11-20].http://www.hect or.ac.uk/cse/distributedcse/reports/VASP02/VASP02.pdf. [16]CATLOW R,WOODLEY S,DE LEEUW N,et al.Optimising the performance of the VASP 5.2.2 code on HECToR[R/OL].University College London and EPCC,University of Edinburgh(2010-11-09).https://www.hector.ac.uk/cse/distributedcse/reports/VASP/VASP_final.pdf. [17]PRABHAKARAN S.Dynamic resource management and jobscheduling for high performancecomputing[EB/OL].[PDF]Dynamic Resource Management and Job Scheduling for High Performance Computing=Dynamisches Ressourcenmanagement und Job-Scheduling für das Hochleistungsrechnen | Semantic Scholar. [18]GE R,CAMERON K W.Power-aware speedup[C]//2007 IEEEInternational Parallel and Distributed Processing Symposium.IEEE,2007:1-10. [19]KELECHI A H,ALSHARIF M H,BAMEYI O J,et al.Artificial intelligence:An energy efficiency tool for enhancedhigh performance computing[J].Symmetry,2020,12(6):1029. [20]YOO A B,JETTE M A,GRONDONA M.Slurm:Simple linux utility for resource management[C]//Workshop on Job Scheduling Strategies for Parallel Processing.Springer,2003:44-60. [21]MATSUNAGA A,FORTES J A B.On the use of machinelearning to predict the time and resources consumed by applications[C]//CCGRID '10:Proceedings of the 2010 10th IEEE/ACM International Conference on Cluster,Cloud and Grid Computing.IEEE Computer Society,2010:495-504. [22]DHIMAN G,MIHIC K,ROSING T.A system for online power prediction in virtualized environments using gaussian mixture models[C]//Proceedings of the 47th Design Automation Conference.2010:807-812. [23]GUI B W,YU S,WEN S Z,et al.Runtime prediction of jobs for backfilling optimization[J].Journal of Chinese Computer Systems,2019,40(1):6-12. [24]ZHANG W S,LI H M,LI J,et al.An application parallel parameter optimization method integrated into a supercomputing job scheduling system [J].Computer Engineering,2025,51(7):59-67. [25]WANG Y,DU Z,JIANG J,et al.Modeling the parallel efficiency of density functional theory based jobs on sunway taihulight[C]//2019 IEEE International Conference on Computational Science and Engineering(CSE)and IEEE International Conference on Embedded and Ubiquitous Computing(EUC).IEEE,2019:199-204. [26]FEITELSON D G,TSAFRIR D,KRAKOV D.Experience with using the parallel workloads archive[J].Journal of Parallel and Distributed Computing,2014,74(10):2967-2982. [27]WANG Z,O'BOYLE M F.Mapping parallelism to multi-cores:a machine learning based approach[C]//Proceedings of the 14th ACM SIGPLANSymposium on Principles and Practice of Parallel Programming.2009:75-84. [28]GOMATHEESHWARI B,SELVAKUMAR J.Appropriate allocation of workloads on performance asymmetric multicore architectures via deep learning algorithms[J].Microprocessors and Microsystems,2020,73:102996. [29]RAGHU H V,SAURAV S K,BAPU B S.PAAS:power aware algorithm for scheduling in high performance computing[C]//2013 IEEE/ACM 6th International Conference on Utility and Cloud Computing.Dresden,Germany:IEEE,2013:327-332. [30]BLAGODUROV S,FEDOROVA A,VINNIK E,et al.Multi-objective job placement in clusters[C]//Proceedings of the International Conference for High Performance Computing,Networking,Storage and Analysis.2015:1-12. [31]YU L,ZHOU Z,FAN Y,et al.System-wide trade-off modeling of performance,power,and resilience onpetascale systems[J].The Journal of Supercomputing,2018,74:3168-3192. [32]BUDACH L,FEUERPFEIL M,IHDE N,et al.The effects of data quality on ml-model performance[J].CoRR:2207.14529,2022. [33]NIEVES-PÍREZ I,MUÑOZ A,ALMEIDA F,et al.Energy efficiency and performance analysis of a legacy atomic scale materials modeling simulator(vasp)[J].The Journal of Supercomputing,2024,80(11):16679-16702. [34]WENDE F,MARSMAN M,KIM J,et al.Openmp in vasp:Threading and simd[J].International Journal of Quantum Chemistry,2019,119(12):e25851. [35]ZHAO Z,MARSMAN M,WENDE F,et al.Performance of hybridmpi/openmp vasp on cray xc40 based on intel knights landing many integrated core architecture[EB/OL].OPUS 4|Performance of Hybrid MPI/OpenMP VASP on Cray XC40 Based on Intel Knights Landing Many Integrated Core Architecture. |
|
||