计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250600127-8.doi: 10.11896/jsjkx.250600127

• 图像处理&多媒体技术 • 上一篇    下一篇

LitchiNet:融合多尺度门控注意力模块与类别不平衡感知的荔枝品种轻量级识别模型

苏烨1,2, 徐鑫3, 赵龙龙1, 李晓丽1, 陈盼1, 陈劲松1   

  1. 1 中国科学院深圳先进技术研究院 广东 深圳 518055
    2 中国科学院大学人工智能学院 北京 101407
    3 华中师范大学计算机学院 武汉 430000
  • 出版日期:2026-06-16 发布日期:2026-06-12
  • 通讯作者: 赵龙龙(ll.zhao@siat.ac.cn)
  • 基金资助:
    国家自然科学基金(42171323);深圳市科技重大专项(可持续发展专项)(KCXFZ20240903093800002)

LitchiNet:Lightweight Litchi Variety Recognition Network with Fused Multi-scale Gated Attention and Class Imbalance Awareness

SU Ye1,2, XU Xin3, ZHAO Longlong1, LI Xiaoli1, CHEN Pan1, CHEN Jinsong1   

  1. 1 Shenzhen Institutes of Advanced Technology,Chinese Academy of Sciences,Shenzhen,Guangdong 518055,China
    2 School of Artificial Intelligence,University of Chinese Academy of Sciences,Beijing 101407,China
    3 School of Computer Science,Central China Normal University,Hubei 430000,China
  • Published:2026-06-16 Online:2026-06-12
  • About author:SU Ye,born in 2001,postgraduate,is a member of CCF(No.K5190G).His main research interests include AutoML,feature selection,ensemble lear-ning,smart agriculture,and intelligent remote sensing interpretation and applications.
    ZHAO Longlong,born in 1988,Ph.D,associate researcher.Her main research interests include artificial intelligence,machine learning,smart agriculture,and intelligent remote sensing interpretation and applications.
  • Supported by:
    National Natural Science Foundation of China(42171323) and Research Foundation of Shenzhen Science and Technology Innovation Bureau(KCXFZ20240903093800002).

摘要: 精准高效的荔枝品种识别,是实现采后荔枝品质智能化检测的关键环节。当前,深度学习模型在该任务中仍面临细粒度特征难以区分、样本数量有限、类别分布不均以及部署资源受限等多重挑战。为应对上述挑战,提出了一种轻量化荔枝品种识别模型(LitchiNet),以预训练的SqueezeNet1.0作为实验验证的主干网络,嵌入一种新颖的多尺度门控注意力模块(Multi-scale Gated Attention,MSGA),通过多尺度卷积融合、通道注意力与轻量门控机制协同作用,在残差连接的基础上强化关键特征区域响应,以增强模型对荔枝品种细粒度差异的识别能力。在LitchiNet的最后阶段,设计了一种计算资源友好、推理效率较高的分类器结构。为了应对类别不平衡问题,提出一种类别不平衡感知损失函数(Class Imbalance Awareness Loss,CIA Loss),通过引入类别权重因子与难易样本调节机制,有效缓解训练样本分布不均带来的偏差。实验基于Github公开荔枝品种识别数据集开展研究。结果显示,LitchiNet在准确率、精确率、召回率、F1分数上均表现优异,其中召回率达到99.40%,明显优于4种主流轻量深度学习模型。同时,模型参数仅为 3.210×106,具备良好的边缘部署适应性。对比4种常见注意力模块的实验结果显示,嵌入 MSGA 的模型具备更快的收敛速度、更低的最终损失及更高的识别精度。此外,LitchiNet的模块化设计支持与任意神经网络主干结构兼容,具备良好的通用性和可拓展性。LitchiNet不仅为荔枝品种智能化识别提供了切实可行的解决方案,也为水果品种细粒度识别研究及农业人工智能技术发展提供了新的思路与技术范式。

关键词: 荔枝, 品种识别, 图像处理, 深度学习, 类别不平衡

Abstract: Accurate and efficient recognition of litchi varieties is essential for intelligent postharvest quality assessment.How-ever,existing deep learning models face several challenges in this task,including fine-grained feature discrimination,limited sample numbers,class imbalance,and constraints on deployment resources.To tackle these challenges,a lightweight litchi variety recognition model,named LitchiNet is proposed.LitchiNet adopts a pretrained SqueezeNet1.0 as its backbone and integrates a novel Multi-Scale Gated Attention(MSGA) module.By combining multi-scale convolutional branches,channel attention,and a lightweight gating mechanism within a residual framework,MSGA enhances the model's ability to capture subtle inter-class diffe-rences and emphasize key feature regions.In the final stage of LitchiNet,a computationally efficient classifier structure is designed to ensure high inference speed and deployment friendliness.To further tackle class imbalance,LitchiNet introduces a Class Imba-lance Awareness Loss(CIA Loss) that incorporates both class weighting and a difficulty-aware modulation term,enabling more robust learning from minority classes.Experiments on a public litchi variety dataset demonstrate that LitchiNet achieves excellent performance,reaching a recall of 99.40% and outperforming four state-of-the-art lightweight models across all metrics.With only 3.210×106 parameters,the model is well-suited for edge deployment.Comparative experiments with four state-of-the-art attention modules further reveal that the inclusion of MSGA leads to faster convergence,lower final loss,and better recognition accuracy.Moreover,the modular design of LitchiNet ensures compatibility with various backbone networks,offering strong generalizability and scalability.LitchiNet provides a practical and effective solution for fine-grained litchi variety recognition,and contri-butes a novel approach to lightweight agricultural AI applications.

Key words: Litchi, Variety recognition, Image processing, Deep learning, Class imbalance

中图分类号: 

  • TP391
[1] ZHUANG L J,QIU Z H.Development characteristics and policy suggestions of China's litchi industry in 2019 [J].China Sou-thern Fruit,2021,50(4):184-188.
[2] FANG Z D,FAN Q,PENG Y X.Analysis of characteristics of ten major litchi varieties in Guangdong Province [J].China Fruit News,2024,41(9):96-101.
[3] MO Y D,ZOU X J,YE M,et al.Eye-in-hand calibration method of litchi picking robot based on Sylvester equation deformation [J].Transactions of the Chinese Society of Agricultural Engineering,2017,33(4):47-54.
[4] GAO F,FU L,ZHANG X,et al.Multi-class fruit-on-plant detection for apple in SNAP system using Faster R-CNN [J].Computers and Electronics in Agriculture,2020,176:105634.
[5] ZHUANG J J,LUO S M,HOU C J,et al.Detection of orchard citrus fruits using a monocular machine vision-based method for automatic fruit picking applications [J].Computers and Electronics in Agriculture,2018,152:64-73.
[6] JIMÉNEZ A R,JAIN A K,CERES R,et al.Automatic fruit re-cognition:a survey and new results using Range/Attenuation images [J].Pattern Recognition,1999,32(10):1719-1736.
[7] BULANON D M,BURKS T F,ALCHANATIS V.Image fusion of visible and thermal images for fruit detection [J].Biosystems Engineering,2009,103(1):12-22.
[8] DENARDA A R,CROCETTI F,COSTANTE G,et al.MangoDetNet:A novel label-efficient weakly supervised fruit detection framework [J].Precision Agriculture,2024,25(6).
[9] SU B F,SHEN L,CHEN S,et al.Multi-feature classification method of grape varieties based on attention mechanism [J].Transactions of the Chinese Society for Agricultural Machinery,2021,52(11):226-233,252.
[10] GENG L,HUANG Y L,GUO Y M.Apple variety classification method based on fused attention mechanism [J].Transactions of the Chinese Society for Agricultural Machinery,2022,53(6):304-310,369.
[11] LIU J,LI Y,XIAO L M,et al.Citrus fruit recognition and localization method based on improved YOLOv4 model [J].Transactions of the Chinese Society of Agricultural Engineering,2022,38(12):173-182.
[12] ZHANG X Y,HUANG G Y,YANG Y T,et al.Strawberry maturity classification method based on improved CNN [J].Food and Machinery,2023,39(10):130-137.
[13] YU L J,XU Z.A litchi fruit recognition method in a natural environment using RGB-D images [J].Biosystems Engineering,2021,204:50-63.
[14] LIU D,WANG L,SUN D W,et al.Lychee variety discrimination by hyperspectral imaging coupled with multivariate classification [J].Food Analytical Methods,2014,7(9):1848-1857.
[15] XIAO Y,WANG J,XIONG H,et al.Lychee cultivar fine-grained image classification method based on improved ResNet-34 residual network [J].Journal of Agricultural Engineering,2024,55(3):84-97.
[16] HUANG M J,CAI W Q,ZHANG Z J,et al.Real-time and accurate recognition algorithm for litchi fruit varieties based on improved YOLOv5 [J].Transactions of the Chinese Society of Agricultural Engineering,2025,41(11):156-164.
[17] HOWARD A G,ZHU M,CHEN B,et al.MobileNets:Efficient convolutional neural networks for mobile vision applications [J].arXiv:1704.04861,2017.
[18] HAN S,MAO H,DALLY W J.Deep compression:Compressing deep neural networks with pruning,trained quantization and Huffman coding [J].arXiv:1510.00149,2015.
[19] SIMONYAN K,ZISSERMAN A.Very deep convolutional networks for large-scale image recognition [J].arXiv:1409.1556,2014.
[20] IANDOLA F N,HAN S,MOSKEWICZ M W,et al.Sque-ezeNet:AlexNet-level accuracy with 50× fewer parameters and <0.5 MB model size [C]//Proceedings of the International Conference on Learning Representations(ICLR).OpenReview.net,2017:1-13.
[21] HOWARD A,SANDLER M,CHU G,et al.Searching for MobileNetV3 [C]//Proceedings of the IEEE/CVF International Conference on Computer Vision(ICCV).IEEE Computer Society,2019:1314-1324.
[22] MA N,ZHANG X,ZHENG H T,et al.ShuffleNet V2:Practical guidelines for efficient CNN architecture design [C]//Procee-dings of the European Conference on Computer Vision(ECCV).Springer,2018:122-138.
[23] TAN M,CHEN B,PANG R,et al.MnasNet:Platform-awareneural architecture search for mobile [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).IEEE Computer Society,2019:2820-2828.
[24] KRIZHEVSKY A,SUTSKEVER I,HINTON G E.ImageNet classification with deep convolutional neural networks [C]//Advances in Neural Information Processing Systems(NeurIPS).2012:1097-1105.
[25] NIU Z,ZHONG G,YU H.A review on the attention mechanism of deep learning [J].Neurocomputing.2021,452:48-62.
[26] BRAUWERS G,FRASINCAR F.A general survey on attention mechanisms indeep learning [J].IEEE Transactions on Know-ledge and Data Engineering,2021,35(4):3279-3298.
[27] GUO M H,XU T X,LIU J J,et al.Attention mechanisms incomputer vision:A survey [J].Computational Visual Media,2022,8(3):331-368.
[28] LI H,WU X J.CrossFuse:A novel cross attention mechanism based infrared and visible image fusion approach [J].Information Fusion,2024,103:102147.
[29] CHEN Y,XIA R,YANG K,et al.DNNAM:Image inpainting algorithm via deep neural networks and attention mechanism [J].Applied Soft Computing,2024,154:111392.
[30] HU J,SHEN L,SUN G.Squeeze-and-excitation networks [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).IEEE Computer Society,2018:7132-7141.
[31] WOO S,PARK J,LEE J Y,et al.CBAM:Convolutional block attention module [C]//Proceedings of the European Conference on Computer Vision(ECCV).Springer,2018:3-19.
[32] WANG Q,WU B,ZHU P,et al.ECA-Net:Efficient channel attention for deep convolutional neural networks [C]//Procee-dings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).IEEE Computer Society,2020:11531-11539.
[33] PARK J,WOO S,LEE J Y,et al.BAM:Bottleneck attention module [J].arXiv:1807.06514,2018.
[34] LI X,WANG W,HU X,et al.Selective kernel networks [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).IEEE Computer Society,2019:510-519.
[35] WANG Y,ZHOU Q,WEI Q,et al.Pyramid attention network for semantic segmentation [C]//Proceedings of the European Conference on Computer Vision(ECCV).Springer,2018:633-648.
[36] DONG Y,JIANG Y,XU D,et al.CSWin Transformer:A generalvision transformer backbone with cross-shaped windows [C]//Advances in Neural Information Processing Systems(NeurIPS).Curran Associates,Inc.,2021,34:12154-12167.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!