计算机科学 ›› 2026, Vol. 53 ›› Issue (6): 153-162.doi: 10.11896/jsjkx.251000113

• 高性能计算 • 上一篇    下一篇

基于机器学习的超算应用并行参数优化

李金优1, 张文帅2, 沈瑜2, 张运动2, 李会民2, 李京1   

  1. 1 中国科学技术大学计算机科学与技术学院 合肥 230026
    2 中国科学技术大学网络信息中心超级计算中心 合肥 230026
  • 收稿日期:2025-10-27 修回日期:2025-12-16 出版日期:2026-06-15 发布日期:2026-06-09
  • 通讯作者: 张文帅(wszhang@ustc.edu.cn)
  • 作者简介:(ljinyou@mail.ustc.edu.cn)
  • 基金资助:
    中国科学院网络安全和信息化专项(CAS-WX2022SDC-SJ13)

Machine Learning-based Parallel Parameter Optimization in High-performance ComputingApplications

LI Jinyou1, ZHANG Wenshuai2, SHEN Yu2, ZHANG Yundong2, LI Huimin2, LI Jing1   

  1. 1 School of Computer Science and Technology,University of Science and Technology of China,Hefei 230026,China
    2 Network Information Center,University of Science and Technology of China,Hefei 230026,China
  • Received:2025-10-27 Revised:2025-12-16 Published:2026-06-15 Online:2026-06-09
  • About author:LI Jinyou,born in 2001,master,is a member of CCF(No.A02025G).Her main research interests include parallel optimization and machine learning.
    ZHANG Wenshuai,born in 1987,Ph.D,senior engineer. His main research in-terests include operation and optimization of supercomputing clusters,high-performance computing and computational physics software,and the construction and rights management of scientific data.
  • Supported by:
    CAS Special Fund for Network Security and Informatization(CAS-WX2022SDC-SJ13).

摘要: 在高性能计算领域,超级计算机平台利用大规模计算资源池实现多任务的并行处理,显著提升了作业执行速度。其资源调度机制要求用户提交作业时显式指定相应请求资源及其数目,而当前实践仍高度依赖人工经验的并行参数配置,往往导致硬件资源利用率难以逼近理论最优值,形成显著的资源利用效率瓶颈。对于结构复杂的超算软件,其输入参数与实际硬件资源紧密耦合,仅调整个别参数难以构建简明有效的框架以充分探索参数空间,而遍历整个参数搜索空间的应用配置代价极为昂贵。基于机器学习的并行计算应用运行时优化方法,通过多个高效模型分析处理大量运行数据来提供更好的参数配置建议,无需用户指定资源量即可自动化确定作业申请的硬件参数配置,从而降低集群软件使用门槛,提高工作效率,有效解决现有技术中超算集群用户配置并行参数效率低下以及特定任务高质量作业运行数据集匮乏的问题。实验结果表明,相比默认或普遍通用的参数配置,基于机器学习方法给出的参数配置在多例典型VASP实例上实现了普遍的加速效果和较低的机时开销,具有显著的推广实践价值。

关键词: 并行优化, 高性能计算, 机器学习, 参数搜索, VASP

Abstract: In the field of high-performance computing(HPC),supercomputing platforms leverage large-scale computational resource pools to enable parallel processing of multiple tasks,significantly accelerating job execution.Their resource scheduling mechanisms require users to explicitly specify requested resources and their quantities when submitting jobs.However,current practice still heavily relies on manual experience-based parallel parameter configuration,often resulting in hardware resource utilization failing to approach theoretical optimal values,creating significant resource efficiency bottlenecks.For structurally complex HPC software,individual input parameters are tightly coupled with actual hardware resources;adjusting isolated parameters alone makes it difficult to construct a concise yet effective framework for thorough exploration,while exhaustive traversal of the entire parameter search space incurs prohibitively expensive configuration costs.Machine learning(ML)-based runtime optimization methods for parallel computing applications employ multiple efficient models to analyze and process extensive execution data,providing superior parameter configuration recommendations.These methods autonomously determine hardware parameter configurations for job requests without requiring users to specify resource amounts,thereby lowering the barrier to cluster software usage and improving work efficiency.This effectively addresses the problems of suboptimal parallel parameter configuration efficiency for supercomputing cluster users and the lack of high-quality job execution datasets for specific tasks in existing technologies.Experimental results demonstrate that compared to default or commonly used parameter configurations,ML-derived parameter configurations achieve consistent speedup and significantly reduced core-hour consumption across multiple typical VASP instances,exhibiting substantial practical value for widespread adoption.

Key words: Parallel optimization, High performance computing, Machine learning, Parameter search, VASP

中图分类号: 

  • TP391
[1]DIAZ J,MUNOZ-CARO C,NINO A.A survey of parallel programming models and tools in the multi and many-core era[J].IEEE Transactions onParallel and Distributed Systems,2012,23(8):1369-1386.
[2]HACKENBERG D,JUCKELAND G,BRUNST H.Performanceanalysis of multi-level parallelism:inter-node,intra-node and hardware accelerators[J].Concurrency and Computation:Practice and Experience,2012,24(1):62-72.
[3]TU B B,ZOU M,ZHAN J F,et al.Research on a hierarchical parallel computing model in the memory layer of multicore processor clusters [J].Journal of Computer Research and Development,2008,31(11):1948-1955.
[4]JINNOUCHI R,LAHNSTEINER J,KARSAI F,et al.Phasetransitions of hybrid perovskites simulated by machine-learning force fields trained on the fly withbayesian inference[J].Physical Review Letters,2019,122(22):225701.
[5]JINNOUCHI R,KARSAI F,KRESSE G.On-the-fly machine learning force field generation:Application to melting points[J].Physical Review B,2019,100(1):014105.
[6]THOMPSON A P,AKTULGA H M,BERGER R,et al.Lammps-a flexible simulation tool for particle-based materials modeling at the atomic,meso,and continuum scales[J].Computer Physics Communications,2022,271:108171.
[7]KRESSE G,FURTHMÜLLER J.Efficient iterative schemes for ab initio total-energy calculations using a plane-wave basis set[J].PhysicalReview B,1996,54(16):11169.
[8]STEGAILOV V,SMIRNOV G,VECHER V.Vasp hits thememory wall:Processors efficiency comparison[J].Concurrency and Computation:Practice and Experience,2019,31(19):e5136.
[9]CORSETTI F.Performance analysis of electronic structurecodes onhpc systems:a case study of siesta[J].PLoS One,2014,9(4):e95390.
[10]KRESSE G,HAFNER J.Ab initio molecular-dynamics simulation of the liquid-metal-amorphous-semiconductor transition in germanium[J].Physical Review B,1994,49(20):14251.
[11]BOWLER D R,MIYAZAKI T.methods in electronic structure calculations[J].Reports on Progress in Physics,2012,75(3):036503.
[12]HOU Z,SHEN H,ZHOU X,et al.Prediction of job characteristics for intelligent resource allocation inhpc systems:a survey and future directions [J].Frontiers of Computer Science,2022,16(5):165107.
[13]MARIANI G,ANGHEL A,JONGERIUS R,et al.Predictingcloud performance forhpc applications before deployment[J].Future Generation Computer Systems,2018,87:618-628.
[14]GARCIA-GASULLA M,BANCHELLI F,PEIRO K,et al.A generic performance analysis technique applied to differentcfd methods for hpc[J].International Journal of Computational Fluid Dynamics,2020,34(7/8):508-528.
[15]GAVIN J P.Optimisation of VASP density functional theory(DFT)/hartree-fock(HF)hybrid functional code using mixed-mode parallelism[R/OL].HECTOR Academy,2012[2023-11-20].http://www.hect or.ac.uk/cse/distributedcse/reports/VASP02/VASP02.pdf.
[16]CATLOW R,WOODLEY S,DE LEEUW N,et al.Optimising the performance of the VASP 5.2.2 code on HECToR[R/OL].University College London and EPCC,University of Edinburgh(2010-11-09).https://www.hector.ac.uk/cse/distributedcse/reports/VASP/VASP_final.pdf.
[17]PRABHAKARAN S.Dynamic resource management and jobscheduling for high performancecomputing[EB/OL].[PDF]Dynamic Resource Management and Job Scheduling for High Performance Computing=Dynamisches Ressourcenmanagement und Job-Scheduling für das Hochleistungsrechnen | Semantic Scholar.
[18]GE R,CAMERON K W.Power-aware speedup[C]//2007 IEEEInternational Parallel and Distributed Processing Symposium.IEEE,2007:1-10.
[19]KELECHI A H,ALSHARIF M H,BAMEYI O J,et al.Artificial intelligence:An energy efficiency tool for enhancedhigh performance computing[J].Symmetry,2020,12(6):1029.
[20]YOO A B,JETTE M A,GRONDONA M.Slurm:Simple linux utility for resource management[C]//Workshop on Job Scheduling Strategies for Parallel Processing.Springer,2003:44-60.
[21]MATSUNAGA A,FORTES J A B.On the use of machinelearning to predict the time and resources consumed by applications[C]//CCGRID '10:Proceedings of the 2010 10th IEEE/ACM International Conference on Cluster,Cloud and Grid Computing.IEEE Computer Society,2010:495-504.
[22]DHIMAN G,MIHIC K,ROSING T.A system for online power prediction in virtualized environments using gaussian mixture models[C]//Proceedings of the 47th Design Automation Conference.2010:807-812.
[23]GUI B W,YU S,WEN S Z,et al.Runtime prediction of jobs for backfilling optimization[J].Journal of Chinese Computer Systems,2019,40(1):6-12.
[24]ZHANG W S,LI H M,LI J,et al.An application parallel parameter optimization method integrated into a supercomputing job scheduling system [J].Computer Engineering,2025,51(7):59-67.
[25]WANG Y,DU Z,JIANG J,et al.Modeling the parallel efficiency of density functional theory based jobs on sunway taihulight[C]//2019 IEEE International Conference on Computational Science and Engineering(CSE)and IEEE International Conference on Embedded and Ubiquitous Computing(EUC).IEEE,2019:199-204.
[26]FEITELSON D G,TSAFRIR D,KRAKOV D.Experience with using the parallel workloads archive[J].Journal of Parallel and Distributed Computing,2014,74(10):2967-2982.
[27]WANG Z,O'BOYLE M F.Mapping parallelism to multi-cores:a machine learning based approach[C]//Proceedings of the 14th ACM SIGPLANSymposium on Principles and Practice of Parallel Programming.2009:75-84.
[28]GOMATHEESHWARI B,SELVAKUMAR J.Appropriate allocation of workloads on performance asymmetric multicore architectures via deep learning algorithms[J].Microprocessors and Microsystems,2020,73:102996.
[29]RAGHU H V,SAURAV S K,BAPU B S.PAAS:power aware algorithm for scheduling in high performance computing[C]//2013 IEEE/ACM 6th International Conference on Utility and Cloud Computing.Dresden,Germany:IEEE,2013:327-332.
[30]BLAGODUROV S,FEDOROVA A,VINNIK E,et al.Multi-objective job placement in clusters[C]//Proceedings of the International Conference for High Performance Computing,Networking,Storage and Analysis.2015:1-12.
[31]YU L,ZHOU Z,FAN Y,et al.System-wide trade-off modeling of performance,power,and resilience onpetascale systems[J].The Journal of Supercomputing,2018,74:3168-3192.
[32]BUDACH L,FEUERPFEIL M,IHDE N,et al.The effects of data quality on ml-model performance[J].CoRR:2207.14529,2022.
[33]NIEVES-PÍREZ I,MUÑOZ A,ALMEIDA F,et al.Energy efficiency and performance analysis of a legacy atomic scale materials modeling simulator(vasp)[J].The Journal of Supercomputing,2024,80(11):16679-16702.
[34]WENDE F,MARSMAN M,KIM J,et al.Openmp in vasp:Threading and simd[J].International Journal of Quantum Chemistry,2019,119(12):e25851.
[35]ZHAO Z,MARSMAN M,WENDE F,et al.Performance of hybridmpi/openmp vasp on cray xc40 based on intel knights landing many integrated core architecture[EB/OL].OPUS 4|Performance of Hybrid MPI/OpenMP VASP on Cray XC40 Based on Intel Knights Landing Many Integrated Core Architecture.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!