Computer Science ›› 2026, Vol. 53 ›› Issue (9): 271-282.doi: 10.11896/jsjkx.250700154

• Artificial Intelligence • Previous Articles     Next Articles

DHMoE:Multimodal Feature-decoupling and Heterogeneous Mixture-of-Experts Model for Alzheimer’s Disease Diagnosis

ZAN Peng, WANG Bin   

  1. College of Computer Science and Technology(College of Data Science),Taiyuan University of Technology,Jinzhong,Shanxi 030600,China
  • Received:2025-07-23 Revised:2025-12-10 Online:2026-09-15 Published:2026-09-10
  • About author:ZAN Peng,born in 1996,postgraduate,is a member of CCF(No.A02491G).His main research interests include deep learning and medical imaging science.
    WANG Bin,born in 1986,Ph.D,professor,Ph.D supervisor.His main research interests include deep learning,brain imaging research,medical images and bioinformatics.
  • Supported by:
    National Natural Science Foundation of China(62176177) and Science and Technology Cooperation and Exchange Special Projects of Shanxi(202304041101034).

Abstract: Alzheimer’s disease(AD) is a neurodegenerative disorder with an extremely complex pathogenesis and is nearly incurable.Due to its irreversible progression,early diagnosis is crucial for patients.Studies show that the joint analysis of multimodal medical imaging data can help reveal the pathological characteristics of AD at different stages and provide a more comprehensive perspective,thereby facilitating early diagnosis and timely intervention.However,existing multimodal methods often fuse information from different modalities directly,neglecting the underlying shared and modality-specific features,which limits the model’s ability to deeply extract AD-related representations.To address this issue,this paper proposes DHMoE,a multimodal deep lear-ning model that integrates feature decoupling learning and a heterogeneous mixture-of-experts(MoE) network,aiming for accurate AD diagnosis.This method constructs modality-specific deep encoding networks to decouple the features of each modality into shared and modality-specific components,which are then mapped into public and private subspaces,respectively.This enables deep modeling of both commonalities and differences within multimodal data,enhancing the model’s representation and discriminative capabilities.Furthermore,a heterogeneous MoE structure is introduced,where the decoupled features are fed into corresponding expert networks.Through a dynamic weight control mechanism,the model achieves adaptive fusion and discriminative modeling of multi-granularity features.Finally,experiments on the ADNI dataset demonstrate that the proposed method outperforms several mainstream approaches in multiple evaluation metrics such as accuracy(ACC) and area under the curve(AUC),achieving more accurate classification results.

Key words: Alzheimer’s disease, sMRI, PET, Multimodal, Feature decoupling, Mixture of experts model

CLC Number: 

  • TP391
[1] SCHELTENS P,DE STROOPER B,KIVIPELTO M,et al.Alzheimer’s disease [J].The Lancet,2021,397(10284):1577-1590.
[2] BROOKMEYER R,JOHNSON E,ZIEGLER-GRAHAM K,et al.Forecasting the global burden of Alzheimers disease [J].Alzheimer’s & Dementia,2007,3(3):186-191.
[3] MULUMBA J,DUAN R,LUO B,et al.The role of neuroima-ging in Alzheimer’s disease:implications for the diagnosis,monitoring disease progression,and treatment[J].Exploration of Neuroscience,2025,4:100675.
[4] YEN C,LIN C L,CHIANG M C.Exploring the frontiers of neuroimaging:a review of recent advances in understandingbrain functioning and disorders [J].Life,2023,13(7):1472.
[5] DU L,LIU F,LIU K,et al.Associating multi-modal brain imaging phenotypes and genetic risk factors via a dirty multi-task learning method [J].IEEE Transactions on Medical Imaging,2020,39(11):3416-3428.
[6] WANG M,SHAO W,HAO X,et al.Identify connectome be-tween genotypes and brain network phenotypes via deep self-reconstruction sparse canonical correlation analysis [J].Bioinformatics,2022,38(8):2323-2332.
[7] ZHU Q,YUAN N,HUANG J,et al.Multi-modal AD classification via self-paced latent correlation analysis [J].Neurocompu-ting,2019,355:143-154.
[8] LITJENS G,KOOI T,BEJNORDI B E,et al.A survey on deep learning in medical image analysis [J].Medical Image Analysis,2017,42:60-88.
[9] TANG H,HUANG Z,LI W,et al.Automatic brain segmentation for PET/MR dual-modal images through a cross-fusion mechanism [J].IEEE Journal of Biomedical and Health Informatics,2025,29(3):1982-1994..
[10] MOHSEN S.Alzheimer’s disease detection using deep learning and machine learning:a review [J].Artificial Intelligence Review,2025,58(9):262.
[11] HAN K,LI G,FANG Z,et al.Multi-template meta-information regularized network for Alzheimer’s disease diagnosis using structural MRI[J].IEEE Transactions on Medical Imaging,2024,43(5):1664-1676.
[12] WEN J,THIBEAU-SUTRE E,DIAZ-MELO M,et al.Convolutional neural networks for classification of Alzheimer’s disease:overview and reproducible evaluation[J].Medical Image Analysis,2020,63:101694.
[13] LIU J,XU Y,LIU Y.Attention-guided 3D CNN with lesion feature selection for early Alzheimer’s disease prediction using longitudinal sMRI[J].IEEE Journal of Biomedical and Health Informatics,2025,29(1):324-332.
[14] JANG J,HWANG D.M3T:Three-dimensional medical image classifierusing multi-plane and multi-slice transformer[C] //Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).2022:20718-20729.
[15] HUANG F,QIU A.Ensemble vision transformer for dementia diagnosis[J].IEEE Journal of Biomedical and Health Informa-tics,2024,28(9):5551-5561.
[16] EL-SAPPAGH S,ALONSO J M,ISLAM S M R,et al.A multilayer multimodal detection and prediction model based on explainable artificial intelligence for Alzheimer’s disease [J].Scientific Reports,2021,11(1):2660.
[17] ZHAO K,LIN J,DYRBA M,et al.Coupling of the spatial distributions between sMRI and PET reveals the progression of Alzheimer’s disease [J].Network Neuroscience,2023,7(1):86-101.
[18] ZHANG T,SHI M.Multi-modal neuroimaging feature fusionfor diagnosis of Alzheimer’s disease [J].Journal of Neuroscience Methods,2020,341:108795.
[19] GOEL T,SHARMA R,TANVEER M,et al.Multimodal neuroimaging based Alzheimer’s disease diagnosis using evolutionary RVFL classifier [J].IEEE Journal of Biomedical and Health Informatics,2025,29(6):3833- 3841.
[20] WU T R,JIAO C N,CUI X,et al.Deep self-reconstruction fusion similarity hashing for the diagnosis of Alzheimer’s disease on multi-modal data [J].IEEE Journal of Biomedical and Health Informatics,2024,28(6):3513-3522.
[21] ZHANG Y,SUN K,LIU Y,et al.A modality-flexible framework for Alzheimer’s disease diagnosis following clinical routine [J].IEEE Journal of Biomedical and Health Informatics,2025,29(1):535-546.
[22] QIANG Y R,ZHANG S W,LI J N,et al.Diagnosis of Alzheimer’s disease by joining dual attention CNN and MLP based on structural MRIs,clinical and genetic data [J].Artificial Intelligence in Medicine,2023,145:102678.
[23] NING Z,XIAO Q,FENG Q,et al.Relation-induced multi-modal shared representation learning for Alzheimer’s disease diagnosis [J].IEEE Transactions on Medical Imaging,2021,40(6):1632-1645.
[24] ZHOU T,THUNG K H,LIU M,et al.Multi-modal latent space inducing ensemble SVM classifier for early dementia diagnosis with neuroimaging data [J].Medical Image Analysis,2020,60:101630.
[25] LIU M,CHENG D,WANG K,et al.Multi-modality cascaded convolutional neural networks for Alzheimer’s disease diagnosis [J].Neuroinformatics,2018,16(3):295-308.
[26] QIU Z,YANG P,XIAO C,et al.3D multimodal fusion network with disease-induced joint learning for early Alzheimer’s disease diagnosis [J].IEEE Transactions on Medical Imaging,2024,43(9):3161-3175.
[27] LYU W,ASHRAFINIA S,MA J,et al.Multi-level multi-moda-lity fusion radiomics:application to PET and CT imaging for prognostication of head and neck cancer [J].IEEE Journal of Biomedical and Health Informatics,2019,24(8):2268-2277.
[28] JIA X,JING X Y,ZHU X,et al.Semi-supervised multi-view deep discriminant representation learning [J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2020,43(7):2496-2509.
[29] HU D,ZHANG H,WU Z,et al.Disentangled-multimodal adversarial autoencoder:application to infant age prediction with incomplete multimodal neuroimages [J].IEEE Transactions on Medical Imaging,2020,39(12):4137-4149.
[30] CHENG J,GAO M,LIU J,et al.Multimodal disentangled variational autoencoder with game theoretic interpretability for glioma grading [J].IEEE Journal of Biomedical and Health Informatics,2021,26(2):673-684.
[31] LI Y,WANG Y,CUI Z.Decoupled multimodal distilling foremotion recognition [C] //Proceedings of the IEEE/CVF Confe-rence on Computer Vision and Pattern Recognition.2023:6631-6640.
[32] JACOBS R A,JORDAN M I,NOWLAN S J,et al.Adaptive mixtures of local experts [J].Neural Computation,1991,3(1):79-87.
[33] GOYAL A,KUMAR N,GUHA T,et al.A multimodal mixture-of-experts model for dynamic emotion prediction in movies [C] //2016 IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP).2016:2822-2826.
[34] CAO B,SUN Y,ZHU P,et al.Multi-modal gated mixture of local-to-global experts for dynamic image fusion [C] //Procee-dings of the IEEE/CVF International Conference on Computer Vision.2023:23555-23564.
[35] DING R,LU H,LIU M.DenseFormer-MoE:a dense transfor-mer foundation model with mixture of experts for multi-task brain image analysis [J].IEEE Transactions on Medical Imaging,2010,44(10):12.
[36] ZHANG X,OU N,BASARAN B D,et al.A foundation model for brain lesion segmentation with mixture of modality experts [C] //International Conference on Medical Image Computing and Computer-Assisted Intervention.2024:379-389.
[37] HE K,ZHANG X,REN S,et al.Deep residual learning forimage recognition [C] //Proceedings of the IEEE Conference on Computer Vision and PatternRecognition(CVPR).2016:770-778.
[38] DOSOVITSKIY A,BEYER L,KOLESNIKOV A,et al.Animage is worth 16×16 words:Transformers for image recognition at scale [J].arXiv:2010.11929,2020.
[39] ZHANG J,HE X,LIU Y,et al.Multi-modal cross-attention network for Alzheimer’s disease diagnosis with multi-modality data[J].Computers in Biology and Medicine,2023,162:107050.
[40] LENG Y,CUI W,PENG Y,et al.Multimodal cross enhanced fusion network for diagnosis of Alzheimer’s disease and subjective memory complaints[J].Computers in Biology and Medicine,2023,157:106788.
[41] ISLAM M A,HASAN MAJUMDER M Z,HUSSEIN M A,et al.A review of machine learning and deep learning algorithms for Parkinson’s disease detection using handwriting and voice datasets[J].Heliyon,2024,10(3):e25469.
[1] QIAO Kexiang, LI Yang, WANG Suge. ITRHG:“Image-Text Rationalization” Fake News Detection Method Based on Hypergraph [J]. Computer Science, 2026, 53(9): 324-332.
[2] CHEN Yarui, HONG Lehan, SHAO Jianlin, YANG Jianning, LIAO Yun, SHI Yancui. Multimodal Probabilistic Generative Model Based on Cross-modal Consistency Constraints [J]. Computer Science, 2026, 53(9): 291-298.
[3] FENG Guang, SUN Xiangli, LIN Yibao, LIU Xinting, CAO Yuqiao, HUANG Junhui, LIAO Beirong. Multimodal Sentiment Analysis Based on Prompt Learning and Guided Gated Fusion Mechanism [J]. Computer Science, 2026, 53(9): 299-308.
[4] REN Yanzhang, GAO Tai, LI Ying, WANG Bin. Gated Bidirectional Mamba Multimodal Feature Fusion Framework for Drug-Target InteractionPrediction [J]. Computer Science, 2026, 53(8): 326-335.
[5] HE Yunong, DING Zhijun. Semantics-aware Fine-grained Parallel Structural Reduction Framework for Petri-net LTL Model Checking [J]. Computer Science, 2026, 53(8): 357-364.
[6] WANG Lihua, WANG Xinyu, YAN Weidan, ZHANG Dengyin. Review of Research on Face Deepfake Detection Technology [J]. Computer Science, 2026, 53(8): 375-387.
[7] FU Le, HUANG Xiaofang, LIAO Min, SONG Luhua. Dynamic Adversarial Detection Framework Based on Multimodal Uncertainty Fusion [J]. Computer Science, 2026, 53(8): 469-477.
[8] WANG Yuqi, ZHANG Yangsen, GUO Yalong, KANG Jing, WANG Yalun. Time Series Language Model for Continuous Glucose Monitoring Interpretation [J]. Computer Science, 2026, 53(8): 29-39.
[9] DENG Jiayan, TIAN Shirui, LIU Hou, ZHU Ningbo, DUAN Mingxing. Zero-shot Pedestrian Trajectory Prediction Method Based on Compositional Motion [J]. Computer Science, 2026, 53(8): 117-126.
[10] KE Xianxin, LI Xuan, SONG Junqi. Research on Facial Emotion Expression Technologies for Humanoid Robots [J]. Computer Science, 2026, 53(7): 1-8.
[11] WANG Hongbiao, ZHAN Qiankun, GAO Ge, LEI Ming. Accurate Prediction of Electric Vehicle Charging Loads Approach Based on Multi-branch Fusionand Multi-head Attention Residual Network [J]. Computer Science, 2026, 53(6A): 250300074-5.
[12] SUN Andong, ZHANG Qingyi. Design of Trend-aware Branch Predictor Based on RISC-V Processor [J]. Computer Science, 2026, 53(6A): 250300124-7.
[13] ZHANG Xinliang, LIU Lilong, CHEN Shangheng, CHEN Ziyang, QIAN Shengsheng. Dual-stream Heterogeneous Social Graph for Micro-video Popularity Prediction [J]. Computer Science, 2026, 53(6A): 250800073-8.
[14] WEI Wei, LI Bicheng, ZHU Zhenshui, ZUO Jun. Semantic Modeling and Co-attention Mechanism for Multimodal Sarcasm Detection Method [J]. Computer Science, 2026, 53(6A): 250400127-6.
[15] ZHENG Haibin, LIN Xiuhao, HAN Ye, CHEN Jinyin, LI Beibei. Black-box Physical Adversarial Attack Against Multimodal Object Detector [J]. Computer Science, 2026, 53(6A): 250700023-10.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!