Computer Science ›› 2026, Vol. 53 ›› Issue (9): 291-298.doi: 10.11896/jsjkx.250900090

• Artificial Intelligence • Previous Articles     Next Articles

Multimodal Probabilistic Generative Model Based on Cross-modal Consistency Constraints

CHEN Yarui1, HONG Lehan1, SHAO Jianlin1, YANG Jianning2, LIAO Yun1, SHI Yancui1   

  1. 1 College of Artificial Intelligence,Tianjin University of Science and Technology,Tianjin 300457,China
    2 School of Communications and Information Engineering,Xi'an University of Posts and Telecommunications,Xi'an 710121,China
  • Received:2025-09-15 Revised:2025-11-28 Online:2026-09-15 Published:2026-09-10
  • About author:CHEN Yarui,born in 1982,Ph.D,associate professor,is a member of CCF(No.47164S).Her main research in-terests include deep learning,image processing and video processing.
    SHI Yancui,born in 1982,Ph.D,asso-ciate professor,is a member of CCF(No.39889M).Her main research interests include Internet of Things applications,recommendation system and social network analysis.
  • Supported by:
    National Natural Science Foundation of China(61976156).

Abstract: With the continuous development of cross-modal generation and understanding tasks,the issue of inter modal consistency has become a key challenge in the research of multimodal deep probability generation models.To address this issue,this paper proposes a multimodal probability generation model(CCVAE) based on cross-modal consistency constraints.This model decouples the hidden space into modal sharing and private hidden space,and uses the expert product mechanism to fuse the shared information of each modality.In the shared latent space,unsupervised alignment constraints are used to guide the alignment of shared representations among different modalities through contrastive learning,reducing the distribution differences between different modalities.Dimensional variance constraints are designed to suppress modal differences in the mean values of various dimensions in the shared space,achieve feature structure alignment between modalities,and improve the stability and consistency of cross-modal shared representations.Comparative experimental results on multiple multimodal datasets show that CCVAE performs superior in cross-modal data cross generation and transformation generation tasks,significantly improving the quality of generated results.The visualization analysis of latent vectors further validates the effectiveness of this model in shared representation alignment and cross-modal discriminative ability.

Key words: Multimodal representation learning, Variational autoencoder, Decoupling, Generate models, Unsupervised learning

CLC Number: 

  • TP183
[1] BALTRUŠAITIS T,AHUJA C,MORENCY L P.Multimodal machine learning:a survey and taxonomy[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2018,41(2):423-443.
[2] SUZUKI M,MATSUO Y.A survey of multimodal deep generative models[J].Advanced Robotics,2022,36(5/6):261-278.
[3] LIU H,YU S,HOU Q,et al.Cross-Modality Image Transformation Using Generative Adversarial Network[C] //International Conference on Man-Machine-Environment System Engineering.Singapore:Springer,2024:463-468.
[4] ZHU Z,LI Y,LYU W,et al.Consistent multimodal generation via a unified GAN framework[C] //Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.2024:5048-5057.
[5] MENG D,TZELEPIS C,PATRAS I,et al.MM2Latent:Text-to-facial image generation and editing in GANs with multimodal assistance[C] //European Conference on Computer Vision.Cham:Springer,2025:88-106.
[6] ZHANG H,KOH J Y,BALDRIDGE J,et al.Cross-modal con-trastive learning for text-to-image generation[C] //Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2021:833-842.
[7] ZHOU D,ZHANG H,MA J,et al.BC-GAN:A Generative Adversarial Network for Synthesizing a Batch of Collocated Clo-thing[J].IEEE Transactions on Circuits and Systems for Video Technology,2023,34(5):3245-3259.
[8] SUZUKI M,NAKAYAMA K,MATSUO Y.Joint multimodal learning with deep generative models[J].arXiv:1611.01891,2016.
[9] WU M,GOODMAN N.Multimodal generative models for scal-able weakly-supervised learning[C] //Advances in Neural Information Processing Systems.2018.
[10] SHI Y,PAIGE B,TORR P.Variational mixture-of-experts autoencoders for multi-modal deep generative models[C] //Advances in Neural Information Processing Systems.2019.
[11] LEE M,PAVLOVIC V.Private-shared disentangled multimodal vae for learning of latent representations[C] //Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2021:1692-1700.
[12] SUTTER T M,MENG Y,AGOSTINI A,et al.Unity by diversity:improved representation learning in multimodal vaes[J].arXiv:2403.05300,2024.
[13] DAUNHAWER I,SUTTER T M,MARCINKEVIČS R,et al.Self-supervised disentanglement of modality-specific and shared factors improves multimodal generative models[C] //DAGM German Conference on Pattern Recognition.Cham:Springer,2020:459-473.
[14] VAN DEN OORD A,LI Y,VINYALS O.Representation lear-ning with contrastive predictive coding[J].arXiv:1807.03748,2018.
[15] PALUMBO E,DAUNHAWER I,VOGT J E.MMVAE+:Enhancing the generative quality of multimodal VAEs without compromises[C] //The Eleventh International Conference on Learning Representations.OpenReview,2023.
[16] MONDAL A K,SAILOPAL A,SINGLA P,et al.SSDMM-VAE:variational multi-modal disentangled representation learning[J].Applied Intelligence,2023,53(7):8467-8481.
[17] MÄRTENS K,YAU C.Disentangling shared and private latent factors in multimodal Variational Autoencoders[C] //Machine Learning in Computational Biology.PMLR,2024:60-75.
[18] ZHOU X,MIAO C.Disentangled graph variational auto-encoder for multimodal recommendation with interpretability[J].IEEE Transactions on Multimedia,2024,26:7543-7554.
[19] MAHAJAN S,BOTSCHEN T,GUREVYCH I,et al.Joint wasserstein autoencoders for aligning multimodal embeddings[C] //Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops.2019.
[20] JIANG Q,CHEN C,ZHAO H,et al.Understanding and constructing latent modality structures in multi-modal representation learning[C] //Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:7661-7671.
[21] MOLINO D,DI FEOLA F,SHEN L,et al.Any-to-Any Vision-Language Model for Multimodal X-ray Imaging and Radiological Report Generation[J].arXiv:2505.01091,2025.
[22] RADFORD A,KIM J W,HALLACY C,et al.Learning transferable visual models from natural language supervision[C] //International Conference on Machine Learning.PMLR,2021:8748-8763.
[23] JIA C,YANG Y,XIA Y,et al.Scaling up visual and vision-language representation learning with noisy text supervision[C] //International Conference on Machine Learning.PMLR,2021:4904-4916.
[24] CHEN Y C,LI L,YU L,et al.Uniter:Universal image-text representation learning[C] //European Conference on Computer Vision.Cham:Springer,2020:104-120.
[25] GAN Z,CHEN Y C,LI L,et al.Large-scale adversarial training for vision-and-language representation learning[J].Advances in Neural Information Processing Systems,2020,33:6616-6628.
[26] FAGHRI F,FLEET D J,KIROS J R,et al.Vse++:Improving visual-semantic embeddings with hard negatives[J].arXiv:1707.05612,2017.
[27] LI J,SELVARAJU R,GOTMARE A,et al.Align before fuse:Vision and language representation learning with momentum distillation[J].Advances in Neural Information Processing Systems,2021,34:9694-9705.
[28] GONZALEZ-GARCIA A,VAN DE WEIJER J,BENGIO Y.Image-to-image translation for cross-domain disentanglement[J].arXiv:1805.09730,2018.
[29] SUTTER T M,DAUNHAWER I,VOGT J E.Generalized multimodal ELBO[J].arXiv:2105.02470,2021.
[30] NETZER Y,WANG T,COATES A,et al.Reading digits in na-tural images with unsupervised feature learning[C] //NIPS Workshop on Deep Lear-ning and Unsupervised Feature Lear-ning.2011:4.
[31] QIU P,ZHU W,KUMAR S,et al.Multimodal Variational Autoencoder:A Barycentric View[C] //Proceedings of the AAAI Conference on Artificial Intelligence.2025:20060-20068.
[32] VAN DER MAATEN L,HINTON G.Visualizing data using t-SNE[J].Journal of Machine Learning Research,2008,9(11):2579-2605.
[1] ZAN Peng, WANG Bin. DHMoE:Multimodal Feature-decoupling and Heterogeneous Mixture-of-Experts Model for Alzheimer’s Disease Diagnosis [J]. Computer Science, 2026, 53(9): 271-282.
[2] CHEN Yifan, DING Cong, CAO Min. Image Anomaly Detection Based on Masked Convolutional Kernel [J]. Computer Science, 2026, 53(7): 45-53.
[3] JI Wendi, WANG Yongquan. Autoregressive Sequence Reconstruction for Unsupervised Anomaly Detection in Medical Insurance [J]. Computer Science, 2026, 53(7): 298-307.
[4] ZHU Yifei, LIU Tianpeng, SUN Tengzhong, LI Yanchen, CHEN Zhihong, FANG Pengfei. Survey of Hyperbolic Geometry in Computer Vision [J]. Computer Science, 2026, 53(7): 9-23.
[5] WANG Hongbiao, ZHAN Qiankun, GAO Ge, LEI Ming. Accurate Prediction of Electric Vehicle Charging Loads Approach Based on Multi-branch Fusionand Multi-head Attention Residual Network [J]. Computer Science, 2026, 53(6A): 250300074-5.
[6] CHANG Yanan, SUN Yi, CUI Jianqun, YAN Xianglong, ZONG Chenglu. Fuzzy Clustering-based DTN Routing Algorithm for IoT [J]. Computer Science, 2026, 53(6): 358-366.
[7] LI Tengjia, MA Chun’ai. Multi-scale Transformer Oil Price Prediction Framework with AEMD and Trend Cross-attention [J]. Computer Science, 2026, 53(5): 157-163.
[8] HUANG Siyang, YAO Ye, ZHU Yian, HAI Duo, XIONG Zhihai. Anomaly Detection and Localization Technology for Gravity Wave Spectral Images Based onPre-trained Networks [J]. Computer Science, 2026, 53(5): 193-206.
[9] LIU Huashuai, TAO Houguo, YUE Kun, DUAN Liang. Bayesian Network Based Fault Root Cause Analysis [J]. Computer Science, 2026, 53(3): 143-150.
[10] WU Jiagao, YI Jing, ZHOU Zehui, LIU Linfeng. Personalized Federated Learning Framework for Long-tailed Heterogeneous Data [J]. Computer Science, 2025, 52(9): 232-240.
[11] JIANG Rui, FAN Shuwen, WANG Xiaoming, XU Youyun. Clustering Algorithm Based on Improved SOM Model [J]. Computer Science, 2025, 52(8): 162-170.
[12] DING Zhengze, NIE Rencan, LI Jintao, SU Huaping, XU Hang. MTFuse:An Infrared and Visible Image Fusion Network Based on Mamba and Transformer [J]. Computer Science, 2025, 52(8): 188-194.
[13] GUO Husheng, ZHANG Xufei, SUN Yujie, WANG Wenjian. Continuously Evolution Streaming Graph Neural Network [J]. Computer Science, 2025, 52(8): 118-126.
[14] ZHENG Chuangrui, DENG Xiuqin, CHEN Lei. Traffic Prediction Model Based on Decoupled Adaptive Dynamic Graph Convolution [J]. Computer Science, 2025, 52(6A): 240400149-8.
[15] AN Rui, LU Jin, YANG Jingjing. Deep Clustering Method Based on Dual-branch Wavelet Convolutional Autoencoder and DataAugmentation [J]. Computer Science, 2025, 52(4): 129-137.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!