计算机科学 ›› 2026, Vol. 53 ›› Issue (8): 71-84.doi: 10.11896/jsjkx.250800093

• 数据库 & 大数据 & 数据科学 • 上一篇    下一篇

面向Spark集群的自适应频率调节方法

何玉林1, 肖又旗1,2, 杨振宇1,2, 黄哲学1, 崔来中2   

  1. 1 人工智能与数字经济广东省实验室(深圳) 广东 深圳 518107
    2 深圳大学计算机与软件学院 广东 深圳 518060
  • 收稿日期:2025-08-20 修回日期:2025-12-03 出版日期:2026-08-15 发布日期:2026-08-17
  • 通讯作者: 肖又旗(xiaoyouqi@gml.ac.cn)
  • 作者简介:(csylhe@126.com)
  • 基金资助:
    广东省自然科学基金面上项目(2023B1515120020);深圳市科技重大专项(KJZD20230923114809020)

Adaptive Frequency Tuning Approach for Spark Clusters

HE Yulin1, XIAO Youqi1,2, YANG Zhenyu1,2, HUANG Zhexue1, CUI Laizhong2   

  1. 1 Guangdong Laboratory of Artificial Intelligence and Digital Economy(Shenzhen), Shenzhen, Guangdong 518107, China
    2 College of Computer Science and Software Engineering, Shenzhen University, Shenzhen, Guangdong 518060, China
  • Received:2025-08-20 Revised:2025-12-03 Published:2026-08-15 Online:2026-08-17
  • About author:HE Yulin,born in 1982,Ph.D,research fellow,doctoral supervisor.His main research interests include novel theoretical models for big data,supporting technologies for big data-oriented software systems,machine learning algorithms for big data,complex big data application systems,etc.
    XIAO Youqi,born in 2001,postgra-duate.His main research interests include distributed computing for big data,Spark energy consumption optimization,and high-performance data mi-ning.
  • Supported by:
    Guangdong Basic and Applied Basic Research Foundation(2023B1515120020) and Science and Technology Major Project of Shenzhen(KJZD20230923114809020).

摘要: 数据中心能效优化是绿色计算的核心研究领域之一。 当前以Spark为代表的大数据分布式计算框架,由于其特有的内存计算模式,能耗持续偏高,这已成为该领域亟待解决的关键问题。 当前数据中心面临资源供给与动态需求的结构性失衡,即昼夜负载波动致使低负载时段硬件高频空转,引发资源闲置浪费;同时,集群计算节点间的负载分布不均会造成低负载节点的资源碎片化损耗。现有的集群能效提升方法存在调控粒度不足、过度依赖历史数据、难以适应负载异构等问题,导致最终能耗优化效果不显著且实现成本高昂。 为此,设计了一种基于深度强化学习的集群能效优化方法。 该方法通过近端策略优化算法构建了基于Actor-Critic的智能体PPOAgent,结合策略熵实现稳定收敛,并通过动态电压频率调节指令,实现硬件资源与作业负载需求的精准匹配,从而达到更有效地降低集群能耗的目的。 实验表明,在Spark-on-YARN集群的HiBench基准测试中,PPOAgent在满足服务水平协议的前提下,相比FAESS-DVFS和GRMSSY,能耗优化幅度最高分别提升9.7%和12.9%,执行时间缩短幅度最高达30.3%和49.7%,并在高负载场景下显著提高了任务的准时完成率。 结果证实,该方法能够在多样化作业与动态负载环境下实现能耗与性能的协同优化,具有重要的实际应用价值。

关键词: 大数据计算, Apache Spark, 深度强化学习, DVFS, 能耗模型

Abstract: Data center energy efficiency represents a critical challenge in green computing.Distributed big data frameworks such as Spark consistently exhibit high energy consumption due to their in-memory computation architecture.Current data centers face a structural imbalance between resource supply and dynamic demand:daily workload fluctuations frequently lead to hardware idling during low-activity periods,resulting in substantial resource wastage.Moreover,uneven load distribution across cluster nodes contributes to resource fragmentation and underutilization.Existing energy optimization methods for clusters often suffer from coarse control granularity,heavy reliance on historical data,and limited adaptability to heterogeneous workloads,leading to constrained effectiveness and high operational costs.To address these limitations,this paper introduces a deep reinforcement lear-ning-based approach for cluster energy efficiency optimization.The proposed method constructs an Actor-Critic agent(PPO-Agent) using the proximal policy optimization algorithm,achieves stable convergence through policy entropy regularization,and leverages dynamic voltage and frequency scaling(DVFS) commands to precisely align hardware resources with real-time job requirements.This enables more effective reduction of overall cluster energy consumption.Experimental evaluations under the HiBench benchmark on Spark-on-YARN clusters demonstrate that PPOAgent achieves energy savings of up to 9.7% and 12.9% compared to FAESS-DVFS and GRMSSY,respectively,while consistently meeting service level agreements.Furthermore,PPOAgent reduces execution time by up to 30.3% and 49.7%,respectively,and significantly improves task completion rates under high-load conditions.These results confirm that the proposed method successfully achieves synergistic optimization of energy efficiency and performance across diverse workloads and dynamic environments,demonstrating considerable practical value.

Key words: Big data computing, Apache Spark, Deep reinforcement learning, DVFS, Energy consumption modeling

中图分类号: 

  • TP302.1
[1] VAVILAPALLI V K,MURTHY A C,DOUGLAS C,et al.Apache hadoop yarn:Yet another resource negotiator[C]//Proceedings of the 4th Annual Symposium on Cloud Computing.2013:1-16.
[2] ZAHARIA M,XIN R S,WENDELL P,et al.Apache Spark:a unified engine for big data processing[J].Communications of the ACM,2016,59(11):56-65.
[3] TOSHNIWAL A,TANEJA S,SHUKLA A,et al.Storm@twitter[C]//Proceedings of the 2014 ACM SIGMOD International Conference on Management of Data.2014:147-156.
[4] MASANET E,SHEHABI A,LEI N,et al.Recalibrating global data center energy-use estimates[J].Science,2020,367(6481):984-986.
[5] ZAHARIA M,CHOWDHURY M,DAS T,et al.Resilient distributed datasets:A {Fault-Tolerant} abstraction for {In-Memory} cluster computing[C]//9th USENIX Symposium on Networked Systems Design and Implementation(NSDI 12).2012:15-28.
[6] YOUSEFI M H N,GOUDARZI M.A task-based greedy sche-duling algorithm for minimizing energy of mapreduce jobs[J].Journal of Grid Computing,2018,16:535-551.
[7] DEAN J,GHEMAWAT S.MapReduce:simplified data proces-sing on large clusters[J].Communications of the ACM,2008,51(1):107-113.
[8] VEIGA J,EXPÓSITO R R,PARDO X C,et al.Performanceevaluation of big data frameworks for large-scale data analytics[C]//2016 IEEE International Conference on Big Data(Big Data).IEEE,2016:424-431.
[9] SHI W,LI H,GUAN J,et al.Energy-efficient scheduling algo-rithms based on task clustering in heterogeneous Spark clusters[J].Parallel Computing,2022,112:102947.
[10] DAYARATHNA M,WEN Y,FAN R.Data center energy consumption modeling:A survey[J].IEEE Communications Surveys & Tutorials,2015,18(1):732-794.
[11] GOIRI Í,LE K,NGUYEN T D,et al.Greenhadoop:leveraginggreen energy in data-processing frameworks[C]//Proceedings of the 7th ACM European Conference on Computer Systems.2012:57-70.
[12] YING Y,BIRKE R,WANG C,et al.Optimizing energy,locality and prI/Ority in a mapreduce cluster[C]//2015 IEEE International Conference on Autonomic Computing.IEEE,2015:21-30.
[13] ISLAM M T,WU H,KARUNASEKERA S,et al.SLA-based scheduling of Spark jobs in hybrid cloud computing environments[J].IEEE Transactions on Computers,2021,71(5):1117-1132.
[14] HANIF M,KIM E,HELAL S,et al.SLA-based adaptationschemes in distributed stream processing engines[J].Applied Sciences,2019,9(6):1045.
[15] SHABESTARI F,NAVIMIPOUR N J.An energy-aware re-source management strategy based on Spark and yarn in heterogeneous environments[J].IEEE Transactions on Green Communications and Networking,2023,8(2):635-644.
[16] MASHAYEKHY L,NEJAD M M,GROSU D,et al.Energy-aware scheduling of mapreduce jobs for big data Applications[J].IEEE Transactions on Parallel and Distributed Systems,2014,26(10):2720-2733.
[17] LI H,WANG H,FANG S,et al.An energy-aware scheduling algorithm for big data Applications in Spark[J].Cluster Computing,2020,23:593-609.
[18] MIRZA N M,ALI A,MUSA N S,et al.Enhancing Task Management in Apache Spark Through Energy-Efficient Data Segregation and Time-Based Scheduling[J].IEEE Access,2024,12:105080-105095.
[19] JALALI KHALIL ABADI Z,MANSOURI N,JAVIDI M M.Deep reinforcement learning-based scheduling in distributed systems:a critical review[J].Knowledge and Information Systems,2024,66(10):5709-5782.
[20] GU Y,LIU Z,DAI S,et al.Deep reinforcement learning for job scheduling and resource management in cloud computing:An algorithm-level review[J].arXiv:2501.01007,2025.
[21] GRINSZTAJN N,BEAUMONT O,JEANNOT E,et al.Readys:A reinforcement learning based strategy for heterogeneous dynamic scheduling[C]//2021 IEEE International Conference on Cluster Computing(CLUSTER).IEEE,2021:70-81.
[22] LIU N,LI Z,XU J,et al.A hierarchical framework of cloud resource allocation and power management using deep reinforcement learning[C]//2017 IEEE 37th International Conference on Distributed Computing Systems(ICDCS).IEEE,2017:372-382.
[23] LILHORE U K,SIMAIYA S,SHARMA Y K,et al.Cloud-edge hybrid deep learning framework for scalable IoT resource optimization[J].Journal of Cloud Computing,2025,14(1):5.
[24] LUO F,YUAN Y,DING W,et al.An improved particle swarm Optimization algorithm based on adaptive weight for task scheduling in cloud computing[C]//Proceedings of the 2nd International Conference on Computer Science and Application Engineering.2018:1-5.
[25] RJOUB G,BENTAHAR J,WAHAB O A,et al.Deep smartscheduling:A deep learning approach for automated big data scheduling over the cloud[C]//2019 7th International Confe-rence on Future Internet of Things and Cloud(FiCloud).IEEE,2019:189-196.
[26] LI H,LU L,SHI W,et al.Energy-aware scheduling for Spark job based on deep reinforcement learning in cloud[J].Computing,2023,105(8):1717-1743.
[27] BABAIYAN V,BUSHEHRIAN O.A deep-reinforcement-learning-based strategy selection approach for fault-tolerant offloading of delay-sensitive tasks in vehicular edge-cloud computing[J].The Journal of Supercomputing,2025,81(5):1-37.
[28] TANG Z,QI L,CHENG Z,et al.An energy-efficient taskscheduling algorithm in DVFS-enabled cloud environment[J].Journal of Grid Computing,2016,14:55-74.
[29] LUO L,WU W J,ZHANG F.Energy modeling based on cloud data center[J].Journal of Software,2014,25(7):1371-1387.
[30] LI H,WEI Y,XIONG Y,et al.A frequency-aware and energy-saving strategy based on DVFS for Spark[J].The Journal of Supercomputing,2021,77:11575-11596.
[31] LI S,ABDELZAHER T,YUAN M.Tapa:Temperature aware power allocation in data center with map-reduce[C]//2011 International Green Computing Conference and Workshops.IEEE,2011:1-8.
[32] IBRAHIM S,PHAN T D,CARPEN-AMARIE A,et al.Governing energy consumption in Hadoop through CPU frequency scaling:An analysis[J].Future Generation Computer Systems,2016,54:219-232.
[33] WU C M,CHANG R S,CHAN H Y.A green energy-efficient scheduling algorithm using the DVFS technique for cloud datacenters[J].Future Generation Computer Systems,2014,37:141-147.
[34] GE R,FENG X,FENG W,et al.Cpu miser:A performance-directed,run-time system for power-aware clusters[C]//2007 International Conference on Parallel Processing(ICPP 2007).IEEE,2007:18-18.
[35] MAROULIS S,ZACHEILAS N,KALOGERAKI V.A framework for efficient energy scheduling of Spark workloads[C]//2017 IEEE 37th International Conference on Distributed Computing Systems(ICDCS).IEEE,2017:2614-2615.
[36] SCHULMAN J,WOLSKI F,DHARIWAL P,et al.Proximalpolicy Optimization algorithms[J].arXiv:1707.06347,2017.
[37] HUANG S,HUANG J,DAI J,et al.The HiBench benchmark suite:Characterization of the MapReduce-based data analysis[C]//2010 IEEE 26th International Conference on Data Engineering Workshops(ICDEW 2010).IEEE,2010:41-51.
[38] ISMAIL L,MATERWALA H.Computing server power mo-deling in a data center:Survey,taxonomy,and performance eva-luation[J].ACM Computing Surveys,2020,53(3):1-34.
[39] SHI H K,WANG Z S,ZHANG S Z,et al.A survey on general CPU performance benchmark research[J].Acta Electronica Sinica,2023,51(1):246-256.
[40] HE Y L,WU D T,PHILIPPE F V,et al.Spark data balanced partitioning method based on priority filling strategy[J].Acta Electronica Sinica,2024,52(10):3322-3335.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!