计算机科学 ›› 2026, Vol. 53 ›› Issue (8): 71-84.doi: 10.11896/jsjkx.250800093
何玉林1, 肖又旗1,2, 杨振宇1,2, 黄哲学1, 崔来中2
HE Yulin1, XIAO Youqi1,2, YANG Zhenyu1,2, HUANG Zhexue1, CUI Laizhong2
摘要: 数据中心能效优化是绿色计算的核心研究领域之一。 当前以Spark为代表的大数据分布式计算框架,由于其特有的内存计算模式,能耗持续偏高,这已成为该领域亟待解决的关键问题。 当前数据中心面临资源供给与动态需求的结构性失衡,即昼夜负载波动致使低负载时段硬件高频空转,引发资源闲置浪费;同时,集群计算节点间的负载分布不均会造成低负载节点的资源碎片化损耗。现有的集群能效提升方法存在调控粒度不足、过度依赖历史数据、难以适应负载异构等问题,导致最终能耗优化效果不显著且实现成本高昂。 为此,设计了一种基于深度强化学习的集群能效优化方法。 该方法通过近端策略优化算法构建了基于Actor-Critic的智能体PPOAgent,结合策略熵实现稳定收敛,并通过动态电压频率调节指令,实现硬件资源与作业负载需求的精准匹配,从而达到更有效地降低集群能耗的目的。 实验表明,在Spark-on-YARN集群的HiBench基准测试中,PPOAgent在满足服务水平协议的前提下,相比FAESS-DVFS和GRMSSY,能耗优化幅度最高分别提升9.7%和12.9%,执行时间缩短幅度最高达30.3%和49.7%,并在高负载场景下显著提高了任务的准时完成率。 结果证实,该方法能够在多样化作业与动态负载环境下实现能耗与性能的协同优化,具有重要的实际应用价值。
中图分类号:
| [1] VAVILAPALLI V K,MURTHY A C,DOUGLAS C,et al.Apache hadoop yarn:Yet another resource negotiator[C]//Proceedings of the 4th Annual Symposium on Cloud Computing.2013:1-16. [2] ZAHARIA M,XIN R S,WENDELL P,et al.Apache Spark:a unified engine for big data processing[J].Communications of the ACM,2016,59(11):56-65. [3] TOSHNIWAL A,TANEJA S,SHUKLA A,et al.Storm@twitter[C]//Proceedings of the 2014 ACM SIGMOD International Conference on Management of Data.2014:147-156. [4] MASANET E,SHEHABI A,LEI N,et al.Recalibrating global data center energy-use estimates[J].Science,2020,367(6481):984-986. [5] ZAHARIA M,CHOWDHURY M,DAS T,et al.Resilient distributed datasets:A {Fault-Tolerant} abstraction for {In-Memory} cluster computing[C]//9th USENIX Symposium on Networked Systems Design and Implementation(NSDI 12).2012:15-28. [6] YOUSEFI M H N,GOUDARZI M.A task-based greedy sche-duling algorithm for minimizing energy of mapreduce jobs[J].Journal of Grid Computing,2018,16:535-551. [7] DEAN J,GHEMAWAT S.MapReduce:simplified data proces-sing on large clusters[J].Communications of the ACM,2008,51(1):107-113. [8] VEIGA J,EXPÓSITO R R,PARDO X C,et al.Performanceevaluation of big data frameworks for large-scale data analytics[C]//2016 IEEE International Conference on Big Data(Big Data).IEEE,2016:424-431. [9] SHI W,LI H,GUAN J,et al.Energy-efficient scheduling algo-rithms based on task clustering in heterogeneous Spark clusters[J].Parallel Computing,2022,112:102947. [10] DAYARATHNA M,WEN Y,FAN R.Data center energy consumption modeling:A survey[J].IEEE Communications Surveys & Tutorials,2015,18(1):732-794. [11] GOIRI Í,LE K,NGUYEN T D,et al.Greenhadoop:leveraginggreen energy in data-processing frameworks[C]//Proceedings of the 7th ACM European Conference on Computer Systems.2012:57-70. [12] YING Y,BIRKE R,WANG C,et al.Optimizing energy,locality and prI/Ority in a mapreduce cluster[C]//2015 IEEE International Conference on Autonomic Computing.IEEE,2015:21-30. [13] ISLAM M T,WU H,KARUNASEKERA S,et al.SLA-based scheduling of Spark jobs in hybrid cloud computing environments[J].IEEE Transactions on Computers,2021,71(5):1117-1132. [14] HANIF M,KIM E,HELAL S,et al.SLA-based adaptationschemes in distributed stream processing engines[J].Applied Sciences,2019,9(6):1045. [15] SHABESTARI F,NAVIMIPOUR N J.An energy-aware re-source management strategy based on Spark and yarn in heterogeneous environments[J].IEEE Transactions on Green Communications and Networking,2023,8(2):635-644. [16] MASHAYEKHY L,NEJAD M M,GROSU D,et al.Energy-aware scheduling of mapreduce jobs for big data Applications[J].IEEE Transactions on Parallel and Distributed Systems,2014,26(10):2720-2733. [17] LI H,WANG H,FANG S,et al.An energy-aware scheduling algorithm for big data Applications in Spark[J].Cluster Computing,2020,23:593-609. [18] MIRZA N M,ALI A,MUSA N S,et al.Enhancing Task Management in Apache Spark Through Energy-Efficient Data Segregation and Time-Based Scheduling[J].IEEE Access,2024,12:105080-105095. [19] JALALI KHALIL ABADI Z,MANSOURI N,JAVIDI M M.Deep reinforcement learning-based scheduling in distributed systems:a critical review[J].Knowledge and Information Systems,2024,66(10):5709-5782. [20] GU Y,LIU Z,DAI S,et al.Deep reinforcement learning for job scheduling and resource management in cloud computing:An algorithm-level review[J].arXiv:2501.01007,2025. [21] GRINSZTAJN N,BEAUMONT O,JEANNOT E,et al.Readys:A reinforcement learning based strategy for heterogeneous dynamic scheduling[C]//2021 IEEE International Conference on Cluster Computing(CLUSTER).IEEE,2021:70-81. [22] LIU N,LI Z,XU J,et al.A hierarchical framework of cloud resource allocation and power management using deep reinforcement learning[C]//2017 IEEE 37th International Conference on Distributed Computing Systems(ICDCS).IEEE,2017:372-382. [23] LILHORE U K,SIMAIYA S,SHARMA Y K,et al.Cloud-edge hybrid deep learning framework for scalable IoT resource optimization[J].Journal of Cloud Computing,2025,14(1):5. [24] LUO F,YUAN Y,DING W,et al.An improved particle swarm Optimization algorithm based on adaptive weight for task scheduling in cloud computing[C]//Proceedings of the 2nd International Conference on Computer Science and Application Engineering.2018:1-5. [25] RJOUB G,BENTAHAR J,WAHAB O A,et al.Deep smartscheduling:A deep learning approach for automated big data scheduling over the cloud[C]//2019 7th International Confe-rence on Future Internet of Things and Cloud(FiCloud).IEEE,2019:189-196. [26] LI H,LU L,SHI W,et al.Energy-aware scheduling for Spark job based on deep reinforcement learning in cloud[J].Computing,2023,105(8):1717-1743. [27] BABAIYAN V,BUSHEHRIAN O.A deep-reinforcement-learning-based strategy selection approach for fault-tolerant offloading of delay-sensitive tasks in vehicular edge-cloud computing[J].The Journal of Supercomputing,2025,81(5):1-37. [28] TANG Z,QI L,CHENG Z,et al.An energy-efficient taskscheduling algorithm in DVFS-enabled cloud environment[J].Journal of Grid Computing,2016,14:55-74. [29] LUO L,WU W J,ZHANG F.Energy modeling based on cloud data center[J].Journal of Software,2014,25(7):1371-1387. [30] LI H,WEI Y,XIONG Y,et al.A frequency-aware and energy-saving strategy based on DVFS for Spark[J].The Journal of Supercomputing,2021,77:11575-11596. [31] LI S,ABDELZAHER T,YUAN M.Tapa:Temperature aware power allocation in data center with map-reduce[C]//2011 International Green Computing Conference and Workshops.IEEE,2011:1-8. [32] IBRAHIM S,PHAN T D,CARPEN-AMARIE A,et al.Governing energy consumption in Hadoop through CPU frequency scaling:An analysis[J].Future Generation Computer Systems,2016,54:219-232. [33] WU C M,CHANG R S,CHAN H Y.A green energy-efficient scheduling algorithm using the DVFS technique for cloud datacenters[J].Future Generation Computer Systems,2014,37:141-147. [34] GE R,FENG X,FENG W,et al.Cpu miser:A performance-directed,run-time system for power-aware clusters[C]//2007 International Conference on Parallel Processing(ICPP 2007).IEEE,2007:18-18. [35] MAROULIS S,ZACHEILAS N,KALOGERAKI V.A framework for efficient energy scheduling of Spark workloads[C]//2017 IEEE 37th International Conference on Distributed Computing Systems(ICDCS).IEEE,2017:2614-2615. [36] SCHULMAN J,WOLSKI F,DHARIWAL P,et al.Proximalpolicy Optimization algorithms[J].arXiv:1707.06347,2017. [37] HUANG S,HUANG J,DAI J,et al.The HiBench benchmark suite:Characterization of the MapReduce-based data analysis[C]//2010 IEEE 26th International Conference on Data Engineering Workshops(ICDEW 2010).IEEE,2010:41-51. [38] ISMAIL L,MATERWALA H.Computing server power mo-deling in a data center:Survey,taxonomy,and performance eva-luation[J].ACM Computing Surveys,2020,53(3):1-34. [39] SHI H K,WANG Z S,ZHANG S Z,et al.A survey on general CPU performance benchmark research[J].Acta Electronica Sinica,2023,51(1):246-256. [40] HE Y L,WU D T,PHILIPPE F V,et al.Spark data balanced partitioning method based on priority filling strategy[J].Acta Electronica Sinica,2024,52(10):3322-3335. |
|
||