计算机科学 ›› 2026, Vol. 53 ›› Issue (6): 332-338.doi: 10.11896/jsjkx.250600056

• 数据库&大数据&数据科学 • 上一篇    下一篇

基于均值漂移的模糊三支聚类

倪永婷1, 钱进1,2, 闫少伟1, 吴越洋1   

  1. 1 华东交通大学信息与软件工程学院 南昌 330013
    2 江苏科技大学计算机学院 江苏 镇江 212003
  • 收稿日期:2025-06-09 修回日期:2025-08-21 出版日期:2026-06-15 发布日期:2026-06-09
  • 通讯作者: 钱进(qjqjlqyf@163.com)
  • 作者简介:(Niyongtingting@gmail.com)
  • 基金资助:
    国家自然科学基金(62466017,62066014);江西省自然科学基金(20232ACB202013)

Fuzzy Three-way Clustering Based on Mean Shift

NI Yongting1, QIAN Jin1,2, YAN Shaowei1, WU Yueyang1   

  1. 1 School of Information and Software Engineering,East China Jiaotong University,Nanchang 330013,China
    2 School of Computer Science and Engineering,Jiangsu University of Science and Technology,Zhenjiang,Jiangsu 212003,China
  • Received:2025-06-09 Revised:2025-08-21 Published:2026-06-15 Online:2026-06-09
  • About author:NI Yongting,born in 2001,postgra-duate.Her main research interests include clustering analysis and granular computing.
    QIAN Jin,born in 1975,Ph.D,professor,is a member of CCF(No.06332S).His main research interests include big data mining,intelligent decision and privacy protection.
  • Supported by:
    National Natural Science Foundation of China(62466017,62066014) and Natural Science Foundation of Jiangxi Province,China(20232ACB202013).

摘要: 均值漂移是一种被广泛使用的基于密度的聚类方法,但其受带宽选择的影响较大,且基于访问次数的划分特性可能导致聚类精度不高。为此,提出一种基于均值漂移的模糊三支聚类算法FTWMS,通过引入动态漂移点选择策略和基于模糊隶属度的划分机制,优化了传统均值漂移算法的聚类过程。首先,算法通过动态选择漂移点来优化漂移过程,以此选择更可靠的聚类中心。然后,结合数据点到不同漂移点的距离、不同簇的访问频率和簇的局部密度,计算模糊隶属度,分别划分簇的核心域和边界域,得到最终的三支聚类结果。最后,在6个人工数据集和6个UCI数据集上进行实验测试,结果表明,FTWMS算法在ACC,NMI和ARI这3个聚类评价指标中表现优异,相较于传统均值漂移方法、k-means和基于三支证据理论的密度峰值聚类(3W-PEDP)算法,能够更好地刻画簇的边界域,综合性能更优。

关键词: 数据挖掘, 均值漂移, 三支聚类, 模糊隶属度, 局部密度

Abstract: Mean Shift is a widely used density-based clustering method,but it is greatly affected by bandwidth selection,and the partitioning characteristics based on the number of visits may lead to low clustering accuracy.To address these issues,this paper proposes a fuzzy three-way clustering algorithm based on Mean Shift(FTWMS).By introducing a dynamic drift point selection strategy and a fuzzy membership-based division mechanism,the proposed method optimizes the clustering process of the tradi-tional Mean Shift algorithm.Firstly,the algorithm enhances the drift process by dynamically selecting drift points,which helps identify more reliable cluster centers.In addition,the fuzzy membership is calculated by integrating the distances from data points to different drift points,the visit frequencies of clusters,and the local densities of cluster.Based on this,each cluster is respectively divided into core and boundary regions,yielding the final three-way clustering structure.Finally,experiments are conducted on 6 synthetic datasets and 6 UCI datasets.The results show that the FTWMS algorithm outperforms traditional Mean Shift,K-means and three-way evidence theory-based density peak clustering(3W-PEDP) algorithms in clustering evaluation metrics,including ACC,NMI,and ARI.The algorithm also demonstrates superior ability in delineating cluster boundaries,offering better overall performance than the three comparison algorithms.

Key words: Data mining, Mean shift, Three-way clustering, Fuzzy membership, Local density

中图分类号: 

  • TP311
[1]ALI H H,KADHUM L E.K-means clustering algorithm applications in data mining and pattern recognition[J].International Journal of Science and Research,2017,6(8):1577-1584.
[2]ZHANG Q H,ZHOU J P,DAI Y Y,et al.Density Peaks Clustering Algorithm Based on Representative Points and Knearest Neighbors.[J] Ruan Jian Xue Bao/Journal of Software,2023,34(12):5629-5648
[3]IKOTUN A M,EZUGWU A E,ABUALIGAH L,et al.K-means clustering algorithms:A comprehensive review,variants analysis,and advances in the era of big data[J].Information Sciences,2023(622):178-210.
[4]BEZDEK J C,EHRLICH ROBERT,FULL W.FCM:The fuzzy c-means clustering algorithm[J].Computers & Geosciences,1984,10(2/3):191-203.
[5]KARYPIS G,HAN E H,KUMAR V.Chameleon:Hierarchical clustering using dynamic modeling[J].Computer,1999,32(8):68-75.
[6]LI B J,WU H M.Deep embedded clustering based on dirichlet variational autoencoder[J].Journal of Chongqing Technology and Business University(Natural Science Edition),2025,42(3):52-62.
[7]ESTER M,KRIGEL H P,SANDER J,et al.A Density-BasedAlgorithm for Discovering Clusters in Large Spatial Databases with Noise[C]//Proceedings of the 2nd International Confe-rence on Knowledge Discovery and Data Mining.Association for Computing Machinery,1996:226-231.
[8]RODRIGUEZ A,LAIO A.Clustering by fast search and find of density peaks[J],Science,2014,344(6191):1492-1496.
[9]FUKUNAGA K,HOSTETLER L.The estimation of the gradient of a density function,with applications in pattern recognition[J].IEEE Transactions on Information Theory,1975,21(1):32-40.
[10]CHENG Y Z.Mean shift,mode seeking,and clustering[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,1995,17(8):790-799.
[11]CARREIRA-PERPINAN M A.Gaussian mean-shift is an EMalgorithm[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2007,29(5):767-776.
[12]WEN L Y,PANG K.Adaptive mean shift clustering algorithm based on cover tree[J].Computer Engineering and Design,2024,45(2):452-458.
[13]QIAN J TANG D W,HONG C X.Research on multi-granularity hierarchical sequential three-branch decision model[J].Journal of Shandong University(Science Edition),2022,57(9):33-45.
[14]YU H.A Framework of Three-way Cluster Analysis [C]//International Joint Conference on Rough Sets.Springer Nature,2017:300-312.
[15]WANG P X,YAO Y Y.CE3:A three-way clustering methodbased on mathematical morphology[J].Knowledge-based Systems,2018(155):54-65.
[16]JU H R,LU Y,DING W P,et al.Three-way evidence theory-based density peak clustering with the principle of justifiable granularity[J].Applied Soft Computing,2024(152):111217.
[17]YU H,CHEN Y,LINGRAS P,et al.A three-way cluster en-semble approach for large-scale data[J].International Journal of Approximate Reasoning,2019(11):32-49.
[18]XIONG J,YU H.An adaptive three-way clustering algorithmfor mixed-type data[C] // Foundations of Intelligent Systems:24th International Symposium(ISMIS).Springer Nature,2018:379-388.
[19]JIANG C M,ZHAO S B.Multi-granularity three-branch clustering ensemble based on shadow set[J].Journal of Electronics,2021,49(8):1524-1532.
[20]DU M J,ZHAO J Q,SUN J R,et al.M3W:Multistep three-way clustering[J].IEEE Transactions on Neural Networks and Learning Systems,2022,35(4): 5627-5640.
[21]WANG P X,YANG X B,DING W P,et al.Three-way clustering:Foundations,survey and challenges[J].Applied Soft Computing,2024(151):111131.
[22]CHENG D D,LI Y,XIA S Y,et al.A fast granular-ball-based density peaks clustering algorithm for large-scale data[J].IEEE Transactions on Neural Networks and Learning Systems,2023(35):22.
[23]WANG Y Z,QIAN J X,HASSAN M,et al.Density peak clustering algorithms:A review on the decade 2014-2023[J].Expert Systems with Applications,2024(238):121860.
[24]WU M,SCHÖLKOPF B.A local learning approach for clustering[J].https://proceedings.neurips.cc/paper_files/paper/2006/file/366f0bc7bd1d4bf414073cabbadfdfcd-Paper.pdf.
[25]VINH N X,EPPS J,BAILEY J.Information Theoretic Measures for Clusterings Comparison:Variants,Properties,Normalization and Correction for Chance[J].Journal of Machine Lear-ning Research,2010(11):2837-2854.
[26]FRÄNTI P,REZAEI M,ZHAO Q P.Centroid index:Clusterlevel similarity measure[J].Pattern Recognition,2014,47(9):3034-3045.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!