计算机科学 ›› 2026, Vol. 53 ›› Issue (6): 185-192.doi: 10.11896/jsjkx.251200067
任荣耀1,2,5, 马百威3,4,5, 邓光华2, 杜琦6, 王跃礼4, 李诗言3,4,5
REN Rongyao1,2,5, MA Baiwei3,4,5, DENG Guanghua2, DU Qi6, WANG Yueli4, LI Shiyan3,4,5
摘要: 针对国产飞腾多核NUMA架构的服务器平台在高性能计算场景中的应用瓶颈,围绕SP基准应用程序开展多层次优化研究,提出并实现了从编译、内存分配、NUMA拓扑和向量化规约4个层次的优化策略。实验分析了1个MPI进程和8个MPI进程下,不同规模数据集在不同并行核数的运行时间、并行效率和加速比。实验分析显示,优化后的运算时间显著缩短,在8个MPI进程128并行核数下,中大规模数据集的性能提升了3~5倍,并缓解了高并发下的性能退化问题。优化后的并行效率更加线性,8个MPI进程的小规模数据集并行效率实现了倍数级提升;中大规模数据集在32,64和128并行核数下并行效率分别保持在90%以上、85%左右和64%~71%,延缓了性能饱和点的出现。加速比方面,8个MPI进程比1个MPI进程在高并发情况下表现出更好的性能增益。实验结果表明,所提多层次优化策略能够有效提升SP应用在目标架构服务器上的计算性能。特别是在核心数增多、NUMA效应凸显时,所提优化方案展现出强大的可扩展性优势,为国产高性能计算平台上的数值仿真与科学计算提供了优化路径。
中图分类号:
| [1]ALMEIDA F,OKON E.Assessing the impact of high-perfor-mance computing on digital transformation:benefits,challenges,and size-dependent differences[J].The Journal of Supercompu-ting,2025,81(6):795. [2]TAN S,JIANG Q,AN H.Uncovering the performance bottleneck of modern HPC processor with static code analyzer:a case study on Kunpeng 920[J].CCF Transactions on High Perfor-mance Computing,2024,6(3):343-364. [3]JIN H,VAN D W R F.Performance characteristics of the multi-zone NAS parallel benchmarks[J].Journal of Parallel and Distributed Computing,2006,66(5):674-685. [4]STONE C P,ELTON B H.Accelerating the multi-zone scalar pentadiagonal CFD algorithm with OpenACC[C]//Proceedings of the Second Workshop on Accelerator Programming Using Directives.2015:1-7. [5]LASZLO E,GILES M,APPLEYARD J.Manycore algorithms for batch scalar and block tridiagonal solvers[J].ACM Transactions on Mathematical Software,2016,42(4):1-36. [6]JESUD R,WEILAND M.Evaluating and optimising compilercode generation for NVIDIA Grace[C]//Proceedings of the 53rd International Conference on Parallel Processing.2024:691-700. [7]CEDRON F,ALVAREZ-GONZALEZ S,RIBAS-RODRIGUEZ A,et al.Efficient Implementation of Multilayer Perceptrons:Reducing Execution Time and Memory Consumption[J].Applied Sciences,2024,14(17):8020. [8]LICKER N.Low-level cross-language post-link optimisation[D].Cambridge:University of Cambridge,2022. [9]DURNER D,LEIS V,NEUMANN T.On the impact of memory allocation on high-performance query processing[C]//Procee-dings of the 15th International Workshop on Data Management on New Hardware.2019:1-3. [10]EVANS J.A scalable concurrent malloc(3) implementation for FreeBSD[C]//Proceedings of the BSDCan Conference.2006. [11]LAMETER C.NUMA(Non-Uniform Memory Access):AnOverview:NUMA becomes more common because memory controllers get close to execution units on microprocessors[J].Queue,2013,11(7):40-51. [12]PAN X,MUELLER F.NUMA-aware memory coloring for multicore real-time systems[J].Journal of Systems Architecture,2021,118:102188. [13]TIKIR M M,HOLLINGSWORTH J K.Hardware monitors for dynamic page migration[J].Journal of Parallel and Distributed Computing,2008,68(9):1186-1200. [14]KANDIAH V,LUSTIG D,VILLA O,et al.Parsimony:Ena-bling SIMD/Vector Programming in Standard Compiler Flows[C]//Proceedings of the 21st ACM/IEEE International Symposium on Code Generation and Optimization.2023:186-198. |
|
||