计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250600199-8.doi: 10.11896/jsjkx.250600199

• 大数据&数据科学 • 上一篇    下一篇

基于属性值类重叠度的不平衡数据学习方法

孙博1, 王志军1, 周筑南1, 李清杰2, 王韵1, 耿霞1, 张艳1, 孙晨轩1   

  1. 1 山东农业大学信息科学与工程学院 山东 泰安 271018
    2 山东省泰安英雄山中学 山东 泰安 271000
  • 出版日期:2026-06-16 发布日期:2026-06-12
  • 通讯作者: 王志军(chaisident@sina.com)
  • 作者简介:(sunbo87@126.com)
  • 基金资助:
    山东省自然科学基金(ZR2023MF098,ZR2023MA011,ZR2021MC168,ZR2018QF002);果业智能化生产关键技术研发与应用示范项目(2022TZXD0011)

Imbalanced Data Learning Approach Utilizing Feature Value Based Class Overlap Degree

SUN Bo1, WANG Zhijun1, ZHOU Zhunan1, LI Qingjie2, WANG Yun1, GENG Xia1, ZHANG Yan1 , SUN Chenxuan1   

  1. 1 College of Information Science and Engineering,Shandong Agricultural University,Taian,Shandong 271018,China
    2 Shandong Taian Yingxiongshan Middle School,Taian,Shandong 271000,China
  • Published:2026-06-16 Online:2026-06-12
  • About author:SUN Bo,born in 1987,Ph.D, associate professor,master's supervisor.His main research interests include machine learning,artificial neural network and security artificial intelligence.
    WANG Zhijun,born in 1974,Ph.D,professor. His main research interests include agricultural informationization,information security,and machine lear-ning.
  • Supported by:
    Natural Science Foundation of Shandong Province(ZR2023MF098,ZR2023MA011,ZR2021MC168,ZR2018QF002) and R&D and Application Demonstration of Key Technologies for Intelligent Production in Fruit Industry(2022TZXD0011).

摘要: 类不平衡问题是有监督机器学习领域中的一项重要挑战。在一个不平衡训练集中,虽然少数类的规模显著小于多数类,但少数类往往是人们更为关注的并且比多数类具有更高的误分类代价。大多数分类器算法通常以总体分类精度作为优化目标,容易误分类对分类精度贡献较小的少数类样例。现有不平衡学习方法往往将训练集的类不平衡比例IR作为分类复杂性度量,并将其作为优化目标。然而,最近研究表明,与IR相比,类重叠更能客观度量不平衡数据的学习难度。鉴于类重叠这一数据复杂性指标的重要性,研究从类重叠视角解决不平衡问题,并提出一种基于类重叠信息的不平衡数据学习方法FO-RBU。具体地,利用各属性上类重叠样例的比例分布来衡量不平衡数据的学习难度,并将其作为确定径向欠采样方法RBU合适欠采样程度的理论依据。实验结果表明,基于属性值的类重叠信息能较好地指导合适欠采样比例的确定,并且提出的类不平衡学习方法FO-RBU是有效的。

关键词: 分类, 类不平衡数据, 类不平衡比例IR, 类重叠, 欠采样, 径向欠采样, 欠采样比例, 机器学习

Abstract: Class imbalance problem is an important challenge in supervised machine learning field.In an imbalanced training set,although the minority class is significantly outnumbered by the majority class,it usually attracts more attention from the practitioners and has higher misclassification cost than the latter one.Most classifier learning algorithms usually employ the overall classification accuracy as the optimization goal,and thus easily misclassify the minority class examples that make less contribution to overall classification accuracy.Existing imbalance learning approaches often utilize the class imbalance ratio(IR) of a training set as the classification complexity measure as well as the optimization goal.However,it has recently been indicated that,compared with IR,class overlap can more objectively measure the learning difficulty of an imbalanced dataset.Considering the importance of class overlap in evaluating the data complexity,the imbalance problem is solved from the class overlap perspective,and an imbalanced dataset learning approach FO-RBU utilizing the class overlap information of a training set is proposed.Specifically,the distribution concerning the ratios of feature based class overlap examples is employed to evaluate the learning difficulty of an imbalanced dataset,and further utilized as a theoretical guideline in determining the proper undersampling extent of Radial-Based Undersampling approach.Experimental results show that the feature values based class overlap information is a good indicator in the proper undersampling ratio determination process,and the proposed class imbalance learning approach FO-RBU is effective.

Key words: Classification, Class imbalanced data, Class imbalance ratio IR, Class overlap, Undersampling, Radial-based undersampling, Undersampling ratio, Machine learning

中图分类号: 

  • TP181
[1] DONG J,JIANG Z,PAN D,et al.A survey on confidence calibration of deep learning-based classification models under class imbalance data[J].IEEE Transactions on Neural Networks and Learning Systems,2025,3(1):1-21.
[2] SUN T H,ZHAO G,GUO M Q.Long-tail Distributed Medical Image Classification Based on Large Selective Nuclear Bilateral-branch Networks[J].Computer Science,2025,52(4):231-239.
[3] XIA T,DANG T,HAN J,et al.Uncertainty-Aware Health Diagnostics via Class-Balanced Evidential Deep Learning[J].IEEE Journal of Biomedical and Health Informatics,2024,28(11):6417-6428.
[4] DING H,SUN Y,HUANG N,et al.TMG-GAN:GenerativeAdversarial Networks-Based Imbalanced Learning for Network Intrusion Detection[J].IEEE Transactions on Information Forensics and Security,2023,19(1):1156-1167.
[5] VUTTIPITTAYAMONGKOL P,ELYAN E,PETROVSKI A.On the class overlap problem in imbalanced data classification[J].Knowledge-Based Systems,2021,212(1):1-17.
[6] LU Y,CHEUNG Y M,TANG Y Y.Bayes imbalance impact index:a measure of class imbalanced data set for classification problem[J].IEEE Transactions on Neural Networks and Learning Systems,2019,31(9):3525-3539.
[7] ZHANG R,ZHANG Z,WANG D.RFCL:A new under-sam-pling method of reducing the degree of imbalance and overlap[J].Pattern Analysis and Applications,2021,24(2):641-654.
[8] KOZIARSKI M.Radial-based undersampling for imbalanced data classification[J].Pattern Recognition,2020,102(1):1-11.
[9] MALDONADO S,VAIRETTI C,FERNANDEZ A,et al.FW-SMOTE:a feature-weighted oversampling approach for imbalanced classification[J].Pattern Recognition,2022,124(4):1-13.
[10] NG W W Y,XU S,ZHANG J,et al..Hashing-based undersampling ensemble for imbalanced pattern classification problems[J].IEEE Transactions on Cybernetics,2022,52(2):1269-1279.
[11] WANG A X,LE V T,TRUNG H N,et al.Addressing imbalance in health data:synthetic minority oversampling using deep learning[J].Computers in Biology and Medicine,2025,188(1):109830-109840.
[12] ZHENG J H,LI X M,LIU S Y,et al.Improved Random Forest Imbalance Data Classification Algorithm Combining Cascaded Up-sampling and Down-sampling[J].Computer Science,2021,48(7):145-154.
[13] HUANG Z,SANG Y,SUN Y,et al.Neural Networks Learn Specified Information for Imbalanced Data Classification[J].IEEE Transactions on Knowledge and Data Engineering,2024,36(11):6719-6730.
[14] HOU,R,CHANG H,MA B,et al.Dual compensation residual networks for class imbalanced learning[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2023,45(10):11733-11752.
[15] CHEN C,SHEN W,YANG C,et al.A New Safe-Level Enabled Borderline-SMOTE for Condition Recognition of Imbalanced Dataset[J].IEEE Transactions on Instrumentation and Mea-surement,2023,72(1):1-10.
[16] LIU C L,CHANG Y H.Learning From Imbalanced Data With Deep Density Hybrid Sampling[J].IEEE Transactions on Systems,Man,and Cybernetics:Systems,2022,52(11):7065-7077.
[17] YAN S,ZHAO Z,LIU S,et al.BO-SMOTE:A Novel Bayesian-Optimization-Based Synthetic Minority Oversampling Technique[J].IEEE Transactions on Systems,Man,and Cybernetics:Systems,2023,12(1):1-13.
[18] XU Z,SHEN D,KOU Y,et al.A Synthetic Minority Oversampling Technique Based on Gaussian Mixture Model Filtering for Imbalanced Data Classification[J].IEEE Transactions on Neural Networks Learning Systems,2022,8(1):1-14.
[19] TAHIR MA,KITTLER J,YAN F.Inverse random under sampling for class imbalance problem and its application to multi-label classification[J].Pattern Recognition,2012,45(10):3738-3750.
[20] BACH M,WERNET A,PALT M.The proposal of undersampling method for learning from imbalanced datasets[J].Procedia Computer Science,2019,159(1):125-134.
[21] OFEK N,ROKACH L,STERN R,et al.Fast-CBUS:A fastclustering-based undersampling method for addressing the class imbalance problem[J].Neurocomputing,2017,243(1):88-102.
[22] TSAI C F,LIN W C,HU Y H,et al.Under-sampling class imbalanced datasets by combining clustering analysis and instance selection[J].Information Sciences,2019,477(1):47-54.
[23] DAI Q,WANG L H,XU K L,et al.Class-overlap detection based on heterogeneous clustering ensemble for multi-class imbalance problem[J].Expert Systems with Applications,2024,255(1):1-17.
[24] FERNANDES E,DE CARVALHO A C.Evolutionary inversion of class distribution in overlapping areas for multi-class imbalanced learning[J].Information Sciences,2019,494(8):141-154.
[25] XU Y,YU Z,CHEN C.Classifier Ensemble Based on Multiview Optimization for High-Dimensional Imbalanced Data Classification[J].IEEE Transactions on Neural Networks and Learning Systems,2024,35(1):870-883.
[26] SANTOS M S,ABREU P H,JAPKOWICZ N,et al.On thejoint-effect of class imbalance and overlap:a critical review[J].Artificial Intelligence Review,2022,55(8):6207-6275.
[27] LORENA AC,GARCIA LPF,LEHMANN J,et al.How complex is your classification problem? a survey on measuring classification complexity[J].ACM Computing Surveys(CSUR),2019,52(5):1-34.
[28] DUA D,GRAFF C.UCI Machine Learning Repository [EB/OL].http://archive.ics.uci.edu/ml.
[29] SUN Z,WANG G,LI P,et al.An improved random forest based on the classification accuracy and correlation measurement of decision trees[J].Expert Systems with Applications,2024,237(1):121549.
[30] XU Z Z,SHEN D R,KOU Y,et al.Clinical prediction of C4.5 decision tree classification algorithm with embedded resampling technique[J].Control and Decision,2021,36(6):1342-1350.
[31] LUQUE A,CARRASCO A,MARTÍN A,et al.The impact of class imbalance in classification performance metrics based on the binary confusion matrix[J].Pattern Recognition,2019,91(1):216-231.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!