计算机科学 ›› 2026, Vol. 53 ›› Issue (6): 185-192.doi: 10.11896/jsjkx.251200067

• 高性能计算 • 上一篇    下一篇

面向国产飞腾多核NUMA架构服务器的SP应用多层次优化研究

任荣耀1,2,5, 马百威3,4,5, 邓光华2, 杜琦6, 王跃礼4, 李诗言3,4,5   

  1. 1 国防科技大学计算机科学与技术学院 长沙 410073
    2 天津先进技术研究院 天津 300450
    3 天津七所精密机电技术有限公司 天津 300131
    4 天津航海仪器研究所 天津 300131
    5 天津市抗恶劣环境特种计算机重点实验室 天津 300450
    6 飞腾信息技术有限公司 天津 300450
  • 收稿日期:2025-12-10 修回日期:2026-03-20 出版日期:2026-06-15 发布日期:2026-06-09
  • 通讯作者: 邓光华(16675347519@163.com)
  • 作者简介:(renry9710@163.com)
  • 基金资助:
    天津市抗恶劣特种计算机重点实验室开放基金(202503)

Research on Multi-level Optimization of SP Applications for Domestic Phytium Multi-core NUMAArchitecture Servers

REN Rongyao1,2,5, MA Baiwei3,4,5, DENG Guanghua2, DU Qi6, WANG Yueli4, LI Shiyan3,4,5   

  1. 1 College of Computer Science and Technology,National University of Defense Technology,Changsha 410073,China
    2 Tianjin Advanced Technology Research Institute,Tianjin 300450,China
    3 Tianjin Qisuo Precision Electromechanical Technology Co.,Ltd.,Tianjin 300131,China
    4 Tianjin Navigation Instruments Research Institute,Tianjin 300131,China
    5 Tianjin Key Laboratory of Special Severe Environment Computer,Tianjin 300450,China
    6 Phytium Technology Co.,Ltd.,Tianjin 300450,China
  • Received:2025-12-10 Revised:2026-03-20 Published:2026-06-15 Online:2026-06-09
  • About author:REN Rongyao,born in 1983,Ph.D candidate,researcher.His main research interests include computer architecture,sonar signal processing and machine learning.
    DENG Guanghua,born in 1994,master,assistant engineer.His main research interests include high performance computing and software development.
  • Supported by:
    Tianjin Key Laboratory of Special Severe Environment Computer Open Foundation(202503).

摘要: 针对国产飞腾多核NUMA架构的服务器平台在高性能计算场景中的应用瓶颈,围绕SP基准应用程序开展多层次优化研究,提出并实现了从编译、内存分配、NUMA拓扑和向量化规约4个层次的优化策略。实验分析了1个MPI进程和8个MPI进程下,不同规模数据集在不同并行核数的运行时间、并行效率和加速比。实验分析显示,优化后的运算时间显著缩短,在8个MPI进程128并行核数下,中大规模数据集的性能提升了3~5倍,并缓解了高并发下的性能退化问题。优化后的并行效率更加线性,8个MPI进程的小规模数据集并行效率实现了倍数级提升;中大规模数据集在32,64和128并行核数下并行效率分别保持在90%以上、85%左右和64%~71%,延缓了性能饱和点的出现。加速比方面,8个MPI进程比1个MPI进程在高并发情况下表现出更好的性能增益。实验结果表明,所提多层次优化策略能够有效提升SP应用在目标架构服务器上的计算性能。特别是在核心数增多、NUMA效应凸显时,所提优化方案展现出强大的可扩展性优势,为国产高性能计算平台上的数值仿真与科学计算提供了优化路径。

关键词: SP应用, 非一致内存访问, MPI, OpenMP, NEON

Abstract: This paper addresses the application bottlenecks of domestic Phytium multi-core NUMA architecture server platforms in high-performance computing scenarios,conducting multi-level optimization research around the SP benchmark application.The paper proposes and implements optimization strategies at four levels:compilation,memory allocation,NUMA topology,and vectorized reduction.Experiments analyze the execution time,parallel efficiency,and speedup for different dataset sizes running on 1 MPI process and 8 MPI processes with varying numbers of parallel cores.The analysis shows that the optimized computation time is significantly reduced.With 8 MPI processes and 128 parallel cores,the performance of medium to large datasets improves by 3 to 5 times,and performance degradation under high concurrency is alleviated.The optimized parallel efficiency is more linear,achieving multiple-fold improvements for small datasets with 8 MPI processes,while medium to large datasets maintain approximately 90%,85%,and 64%~71% efficiency at 32,64,and 128 parallel cores respectively,delaying the onset of performance saturation.In terms of speedup,8 MPI processes show better performance gains than 1 MPI process under high concurrency.Experimental results demonstrate that the proposed multi-level optimization strategies can effectively enhance the computational performance of SP applications on the target server architecture.Especially as the number of cores increases and NUMA effects become more pronounced,the optimization scheme exhibits strong scalability advantages,providing an optimization path for numerical simulation and scientific computing on domestic high-performance computing platforms.

Key words: SP applications, Non-uniform memory access, MPI, OpenMP, NEON

中图分类号: 

  • TP311
[1]ALMEIDA F,OKON E.Assessing the impact of high-perfor-mance computing on digital transformation:benefits,challenges,and size-dependent differences[J].The Journal of Supercompu-ting,2025,81(6):795.
[2]TAN S,JIANG Q,AN H.Uncovering the performance bottleneck of modern HPC processor with static code analyzer:a case study on Kunpeng 920[J].CCF Transactions on High Perfor-mance Computing,2024,6(3):343-364.
[3]JIN H,VAN D W R F.Performance characteristics of the multi-zone NAS parallel benchmarks[J].Journal of Parallel and Distributed Computing,2006,66(5):674-685.
[4]STONE C P,ELTON B H.Accelerating the multi-zone scalar pentadiagonal CFD algorithm with OpenACC[C]//Proceedings of the Second Workshop on Accelerator Programming Using Directives.2015:1-7.
[5]LASZLO E,GILES M,APPLEYARD J.Manycore algorithms for batch scalar and block tridiagonal solvers[J].ACM Transactions on Mathematical Software,2016,42(4):1-36.
[6]JESUD R,WEILAND M.Evaluating and optimising compilercode generation for NVIDIA Grace[C]//Proceedings of the 53rd International Conference on Parallel Processing.2024:691-700.
[7]CEDRON F,ALVAREZ-GONZALEZ S,RIBAS-RODRIGUEZ A,et al.Efficient Implementation of Multilayer Perceptrons:Reducing Execution Time and Memory Consumption[J].Applied Sciences,2024,14(17):8020.
[8]LICKER N.Low-level cross-language post-link optimisation[D].Cambridge:University of Cambridge,2022.
[9]DURNER D,LEIS V,NEUMANN T.On the impact of memory allocation on high-performance query processing[C]//Procee-dings of the 15th International Workshop on Data Management on New Hardware.2019:1-3.
[10]EVANS J.A scalable concurrent malloc(3) implementation for FreeBSD[C]//Proceedings of the BSDCan Conference.2006.
[11]LAMETER C.NUMA(Non-Uniform Memory Access):AnOverview:NUMA becomes more common because memory controllers get close to execution units on microprocessors[J].Queue,2013,11(7):40-51.
[12]PAN X,MUELLER F.NUMA-aware memory coloring for multicore real-time systems[J].Journal of Systems Architecture,2021,118:102188.
[13]TIKIR M M,HOLLINGSWORTH J K.Hardware monitors for dynamic page migration[J].Journal of Parallel and Distributed Computing,2008,68(9):1186-1200.
[14]KANDIAH V,LUSTIG D,VILLA O,et al.Parsimony:Ena-bling SIMD/Vector Programming in Standard Compiler Flows[C]//Proceedings of the 21st ACM/IEEE International Symposium on Code Generation and Optimization.2023:186-198.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!