计算机科学 ›› 2026, Vol. 53 ›› Issue (7): 34-44.doi: 10.11896/jsjkx.250400046

• 计算机图形学 & 多媒体 • 上一篇    下一篇

FFiT:基于FasterViT的快联合帧插值去模糊方法

程之荣, 徐杨   

  1. 贵州大学大数据与信息工程学院 贵阳 550025
  • 收稿日期:2025-04-10 修回日期:2025-07-14 出版日期:2026-07-15 发布日期:2026-07-10
  • 通讯作者: 徐杨(xuy@gzu.edu.cn)
  • 作者简介:(gs.zrcheng23@gzu.edu.cn)
  • 基金资助:
    贵州省科技计划(黔科合成果(2024)重大004)

FFiT:Faster Frame Interpolation Transformer Based on FasterViT

CHENG Zhirong, XU Yang   

  1. College of Big Data and Information Engineering,Guizhou University,Guiyang 550025,China
  • Received:2025-04-10 Revised:2025-07-14 Published:2026-07-15 Online:2026-07-10
  • About author:CHENG Zhirong,born in 2000,postgraduate.His main research interests include CNN+Transformer and motion video deblur.
    XU Yang,born in 1980,associate professor,is a member of CCF(No.D7127M).His main research interests include deep learning and action gesture recognition.
  • Supported by:
    Guizhou Province Science and Technology Plan(Qiankehe Achievements(2024)Major 004).

摘要: 在动态场景图像采集领域,消费级成像设备通常受限于硬件架构缺陷与曝光参数优化不足,长期存在帧率受限与运动模糊共存的复合型退化问题。现有方法在真实模糊场景下存在重建精度不足、模型泛化能力受限,以及计算效率低下等技术瓶颈。对此,提出一种基于FasterViT架构的联合帧插值与运动解耦优化框架FFiT(Faster Frame Interpolation Transformer),旨在实现动态模糊序列的高效时空联合建模。该框架通过构建CNN+Transformer混合编码架构,集成了以下关键模块:1)改进型多尺度残差Transformer模块(Multi-scale Residual Transformer Block,MRTB),旨在通过时空域自注意力机制解决重建精度问题,并显式构建模糊帧间的事件相关性;2)高质量特征传输网络(High-Quality Feature Transmitter,HQFT),采用跨尺度特征蒸馏机制增强模糊-清晰域转换的语义一致性,以提升模型的泛化能力;3)轻量化动态上采样渲染模块(Dynamic Upsampling and Rendering Module,DYRM),通过可微分动态卷积实现分辨率重建与计算复杂度的解耦优化,解决计算效率低下并实现灵活的分辨率恢复。在Adobe240和RBI数据集上进行实验,结果表明,FFiT在解决动态模糊序列的时空联合建模难题方面展现了卓越性能:不仅将模型参数量降低了70%,还将峰值信噪比(PSNR)较基线模型提升了2.73 dB,在提升图像质量的同时平衡了计算效率,为解决此类复合退化问题提供了有价值的参考。

关键词: 运动模糊, 联合插值去模糊, CNN+Transformer, 轻量化

Abstract: In the realm of dynamic scene image acquisition,consumer-grade imaging devices commonly suffer from a compound degradation problem characterized by coexisting limited frame rates and motion blur.These issues primarily stem from inherent hardware architecture deficiencies and suboptimal exposure parameter optimization.Prevailing methodologies aimed at mitigating these degradations encounter significant technical bottlenecks in real-world blurry scenarios,including insufficient reconstruction accuracy,restricted model generalization capabilities,and low computational efficiency.To address these limitations,this paper introduces FFiT(Faster Frame Interpolation Transformer),a novel joint frame interpolation and motion decoupling optimization framework based on the FasterViT architecture.FFiT is designed to achieve efficient spatio-temporal joint modeling of dynamic blurry sequences.The framework integrates a CNN+Transformer hybrid encoding architecture,incorporating several key mo-dules:1)An improved multi-scale residual transformer block(MRTB),which leverages spatio-temporal self-attention mechanisms to enhance reconstruction accuracy and explicitly model event correlations between blurry frames;2)A high-quality feature transmitter(HQFT) module,employing a cross-scale feature distillation mechanism to bolster semantic consistency during the blurry-to-sharp domain conversion,thereby addressing the challenge of model generalization;3)A lightweight dynamic upsampling and rendering module(DYRM) that utilizes differentiable dynamic convolution to decouple resolution reconstruction from computational complexity,thus tackling computational inefficiency and enabling flexible resolution recovery.Experimental evaluations on the Adobe240 and RBI datasets demonstrate that FFiT exhibits exceptional performance in addressing the intricate spatio-temporal joint modeling of dynamic blurry sequences.Notably,FFiT reduces model parameters by 70% to accommodate computational resource constraints while achieving a 2.73 dB improvement in PSNR(Peak Signal-to-Noise Ratio) compared to baseline models.This research offers a valuable reference for resolving such compound degradation problems by effectively balancing image quality enhancement with computational efficiency.

Key words: Motion blur, Joint interpolation deblurring, CNN+Transfomer, Lightweight

中图分类号: 

  • TP751
[1]ZHONG Z,CAO M,JI X,et al.Blur interpolation transformer for real-world motion from blur[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:5713-5723.
[2]KIM T,LEE J,WANG L,et al.Event-guided deblurring of unknown exposure time videos[C]//European Conference on Computer Vision.Cham:Springer Nature Switzerland,2022:519-538.
[3]JIN M,HU Z,FAVARO P.Learning to extract flawless slow motion from blurry videos[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2019:8112-8121.
[4]SHEN W,BAO W,ZHAI G,et al.Video frame interpolation and enhancement via pyramid recurrent framework[J].IEEE Transactions on Image Processing,2020,30:277-292.
[5]OH J,KIM M.Demfi:deep joint deblurring and multi-frame interpolation with flow-guided attentive correlation and recursive boosting[C]//European Conference on Computer Vision.Cham:Springer Nature Switzerland,2022:198-215.
[6]ROZUMNYI D,OSWALD M R,FERRARI V,et al.Motion-from-blur:3d shape and motion estimation of motion-blurred objects in videos[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:15990-15999.
[7]GUPTA A,AICH A,ROY-CHOWDHURY A K.Alanet:Adaptive latent attention network for joint video deblurring and interpolation[C]//Proceedings of the 28th ACM International Conference on Multimedia.2020:256-264.
[8]杨雪松,何亮田.基于即插即用的盲二值图像去模糊算法[J].重庆工商大学学报(自然科学版),2025,42(4):36-43.
[9]SHANG W,REN D,YANG Y,et al.Joint Video Multi-Frame Interpolation and Deblurring under Unknown Exposure Time[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:13935-13944.
[10]CHO S J,JI S W,HONG J P,et al.Rethinking coarse-to-fine approach in single image deblurring[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2021:4641-4650.
[11]PARK D,KANG D U,KIM J,et al.Multi-temporal recurrent neural networks for progressive non-uniform single image deblurring with incremental temporal training[C]//European Conference on Computer Vision.Cham:Springer International Publishing,2020:327-343.
[12]REN D,ZHANG K,WANG Q,et al.Neural blind deconvolution using deep priors[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2020:3341-3350.
[13]ZHONG Z,GAO Y,ZHENG Y,et al.Efficient spatio-temporal recurrent neural network for video deblurring[C]//Computer Vision-ECCV 2020:16th European Conference.Springer International Publishing,2020:191-207.
[14]PAN J,BAI H,TANG J.Cascaded deep video deblurring using temporal sharpness prior[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2020:3043-3051.
[15]SHANG W,REN D,ZOU D,et al.Bringing events into video deblurring with non-consecutively blurry frames[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2021:4531-4540.
[16]WANG X,CHAN K C K,YU K,et al.Edvr:Video restoration with enhanced deformable convolutional networks[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops.2019:1954-1963.
[17]ZHAO Q,ZHOU D M,YANG H,et al.Image deblurring based on residual attention and multi-feature fusion[J].Computer Science,2023,50(1):147-155.
[18]ZAMIR S W,ARORA A,KHAN S,et al.Multi-stage progressive image restoration[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2021:14821-14831.
[19]ZAMIR S W,ARORA A,KHAN S,et al.Restormer:Efficient transformer for high-resolution image restoration[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:5728-5739.
[20]LIANG J,FAN Y,XIANG X,et al.Recurrent video restoration transformer with guided deformable attention[J].Advances in Neural Information Processing Systems,2022,35:378-393.
[21]CAO M,FAN Y,ZHANG Y,et al.VDTR:Video deblurring with transformer[J].IEEE Transactions on Circuits and Systems for Video Technology,2022,33(1):160-171.
[22]TOUVRON H,CORD M,DOUZE M,et al.Training data-efficient image transformers & distillation through attention[C]//International Conference on Machine Learning.PMLR,2021:10347-10357.
[23]GRAHAM B,EL-NOUBY A,TOUVRON H,et al.Levit:a vision transformer in convnet's clothing for faster inference[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2021:12259-12269.
[24]WANG W,XIE E,LI X,et al.Pyramid vision transformer:A versatile backbone for dense prediction without convolutions[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2021:568-578.
[25]YUAN L,CHEN Y,WANG T,et al.Tokens-to-token vit:Training vision transformers from scratch on imagenet[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2021:558-567.
[26]HATAMIZADEH A,YIN H,HEINRICH G,et al.Global context vision transformers[C]//International Conference on Machine Learning.PMLR,2023:12633-12646.
[27]ZHANG P,DAI X,YANG J,et al.Multi-scale vision longformer:A new vision transformer for high-resolution image encoding[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2021:2998-3008.
[28]YUAN L,HOU Q,JIANG Z,et al.Volo:Vision outlooker for visual recognition[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2022,45(5):6575-6586.
[29]LIU Z,HU H,LIN Y,et al.Swin transformer v2:Scaling up capacity and resolution[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:12009-12019.
[30]CHEN Z,XIE L,NIU J,et al.Visformer:The vision-friendly transformer[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2021:589-598.
[31]CHEN C F R,FAN Q,PANDA R.Crossvit:Cross-attention multi-scale vision transformer for image classification[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2021:357-366.
[32]LIU X F,JIANG M R,HUANG Y Q,et al.Using image stabilization and VLBE algorithm to extract foreground target from jitter video[J].Computer Science,2022,49(S2):404-411.
[33]BAO W,LAI W S,MA C,et al.Depth-aware video frame interpolation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2019:3703-3712.
[34]LEE H,KIM T,CHUNG T,et al.Adacof:Adaptive collaboration of flows for video frame interpolation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2020:5316-5325.
[35]XU X,SIYAO L,SUN W,et al.Quadratic video interpolation[C]//NeurIPS 2019.2020:1636-1645.
[36]CHI Z,MOHAMMADI NASIRI R,LIU Z,et al.All at once:Temporally adaptive multi-frame interpolation with advanced motion modeling[C]//Computer Vision-ECCV 2020:16th European Conference.2020:107-123.
[37]PUROHIT K,SHAH A,RAJAGOPALAN A N.Bringing alive blurred moments[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2019:6830-6839.
[38]ARGAW D M,KIM J,RAMEAU F,et al.Motion-blurred video interpolation and extrapolation[C]//Proceedings of the AAAI Conference on Artificial Intelligence.2021:901-910.
[39]HATAMIZADEH A,HEINRICH G,YIN H,et al.Fastervit:Fast vision transformers with hierarchical attention[J].arXiv:2306.06189,2023.
[40]HAN Q,FAN Z,DAI Q,et al.On the connection between local attention and dynamic depth-wise convolution[J].arXiv:2106.04263,2021.
[41]SU S,DELBRACIO M,WANG J,et al.Deep video deblurring for hand-held cameras[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.2017:1279-1288.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!