[1] Selva J, Johansen AS, Escalera S, Nasrollahi K, Moeslund TB, Clapés A. Video Transformers: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2023;45(11):12922-43. 10.1109/TPAMI.2023.3243465
[2] Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, et al. An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. In: International Conference on Learning Representations; 2021. Available from: https://arxiv.org/abs/2010.11929. 10.48550/arXiv.2010.11929
[3] Mashru D, Vora K. Comparative Analysis of CNN, RNN, LSTM, and Transformer Architectures in Deep Learning.Educational Administration: Theory and Practice. 2023;29(4):5439-43.10.53555/kuey.v29i4.10364
[4] MISSING.
[5] Arnab A, Dehghani M, Heigold G, Sun C, Lučić M, Schmid C. ViViT: A Video Vision Transformer. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. IEEE; 2021. p. 6836-46. 10.1109/ICCV48922.2021.00676
[6] Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention Is All You Need. In: Advances in Neural Information Processing Systems. vol. 30; 2017. p. 5998-6008.10.48550/arXiv.1706.03762
[7] Peruzzo E, Sangineto E, Liu Y, De Nadai M, Bi W, Lepri B, et al.Spatial Entropy as an Inductive Bias for Vision Transformers.Machine Learning. 2024;113:6945-75.10.1007/s10994-024-06570-7
[8] Le DPC, Wang D, Le VT. A Comprehensive Survey of Recent Transformers in Image, Video and Diffusion Models. Computers, Materials & Continua. 2024;80(1):37-60. 10.32604/cmc.2024.050790
[9] Singh S, Dewangan S, Krishna GS, Tyagi V, Reddy S, Medi PR. Video Vision Transformers for Violence Detection; 2022. Preprint. Available from: https://arxiv.org/abs/2209.03561. 10.48550/arXiv.2209.03561
[10] Mulyanto E, Yuniarno EM, Putra OV, Hafidz I, Priyadi A, Purnomo MH. Improvement of Tradition Dance Classification Process Using Video Vision Transformer Based on Tubelet Embedding. International Journal of Intelligent Engineering and Systems. 2024;17(4):530-45. 10.22266/IJIES2024.0831.41
[11] Li L, Zhuang L, Gao S, Wang S.HaViT: Hybrid-Attention Based Vision Transformer for Video Classification. In: Wang L, Gall J, Chin TJ, Sato I, Chellappa R, editors. Computer Vision – ACCV 2022. vol. 13844 of Lecture Notes in Computer Science. Cham: Springer Nature Switzerland; 2023. p. 502-17.10.1007/978-3-031-26316-330
[12] Zhang B, Yu J, Fifty C, Han W, Dai AM, Pang R, et al.. Co-Training Transformer with Videos and Images Improves Action Recognition; 2021. Preprint. Available from: https://arxiv.org/abs/2112.07175. 10.48550/arXiv.2112.07175
[13] Gritsenko AA, Xiong X, Djolonga J, Dehghani M, Sun C, Lučić M, et al. End-to-End Spatio-Temporal Action Localisation with Video Transformers. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; 2024. p. 18373-83.10.1109/CVPR52733.2024.01739
[14] Gowda SN, Arnab A, Huang J.Optimizing Factorized Encoder Models: Time and Memory Reduction for Scalable and Efficient Action Recognition. In: Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T, Varol G, editors. Computer Vision – ECCV 2024. vol. 15068 of Lecture Notes in Computer Science. Cham: Springer Nature Switzerland; 2025. p. 457-74. 10.1007/978-3-031-72684-226
[15] Huang Z, Qing Z, Wang X, Feng Y, Zhang S, Jiang J, et al.. Towards Training Stronger Video Vision Transformers for EPIC-KITCHENS-100 Action Recognition; 2021.Preprint.Available from: https://arxiv.org/abs/2106.05058. 10.48550/arXiv.2106.05058
[16] Koot R, Hennerbichler M, Lu H. Evaluating Transformers for Lightweight Action Recognition; 2021. Preprint. Available from: https://arxiv.org/abs/2111.09641. 10.48550/arXiv.2111.09641
[17] Yang J, Dong X, Liu L, Zhang C, Shen J, Yu D. Recurring the Transformer for Video Action Recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; 2022. p. 14043-53.10.1109/CVPR52688.2022.01367
[18] Li D, Lasenby J. A Revised Video Vision Transformer for Traffic Estimation with Fleet Trajectories. IEEE Sensors Journal. 2022;22(17):17103-12. 10.1109/JSEN.2022.3193663
[19] Kitškerkin J. In Search of Video Transformer Models for Action Recognition on Limited Data. Tilburg University; 2022. The author’s full given name and a stable repository record could not be independently verified from the submitted citation
[20] Truong TD, Bui QH, Duong CN, Seo HS, Phung SL, Li X, et al.DirecFormer: A Directed Attention in Transformer Approach to Robust Action Recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; 2022. p. 19998-20008.10.1109/CVPR52688.2022.01940
[21] Denize J, Liashuha M, Rabarisoa J, Orcesi A, Hérault R. COMEDIAN: Self-Supervised Learning and Knowledge Distillation for Action Spotting Using Transformers. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops. IEEE; 2024. p. 518-28. 10.1109/WACVW60836.2024.00060
[22] Vu NT, Huynh VT, Nguyen TN, Kim SH. Ensemble Spatial and Temporal Vision Transformer for Action Units Detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. IEEE; 2023. p. 5770-6. 10.1109/CVPRW59228.2023.00612
[23] Shaikh MB, Chai D, Shamsul Islam SM, Akhtar N.MAiVAR-T: Multimodal Audio-Image and Video Action Recognizer Using Transformers. In: 2023 11th European Workshop on Visual Information Processing. IEEE; 2023. p. 1-6. 10.1109/EUVIP58404.2023.10323051
[24] Lu H, Jian H, Poppe R, Salah AA. Enhancing Video Transformers for Action Understanding with VLM-Aided Training; 2024. Preprint. Available from: https://arxiv.org/abs/2403.16128. 10.48550/arXiv.2403.16128
[25] Deng F, Yang C, Guo H, Wang Y, Xu L. DA-ViViT: Fatigue Detection Framework Using Joint and Facial Keypoint Features with Dynamic Distributed Attention Video Vision Transformer. Research Square; 2024.Preprint, version 1. Available from: https://doi.org/10.21203/rs.3.rs-4546491/v1. 10.21203/rs.3.rs-4546491/v1
[26] Chen J, Ho CM. MM-ViT: Multi-Modal Video Transformer for Compressed Video Action Recognition. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. IEEE; 2022. p. 786-97. 10.1109/WACV51458.2022.00086
[27] Ma Y, Wang R, Zong M, Ji W, Wang Y, Ye B. Convolutional Transformer Network for Fine-Grained Action Recognition. Neurocomputing. 2024;569:127027. 10.1016/j.neucom.2023.127027
[28] Kuznietsova N, Smirnov S. Application of Vision Transformers and 3D Convolutional Neural Networks for Sign Language Cluster Recognition. In: Computer Modeling and Intelligent Systems. vol. 3392 of CEUR Workshop Proceedings; 2023. p. 151-63. 10.32782/cmis/3392-13
[29] Liang Y, Zhou P, Zimmermann R, Yan S.DualFormer: Local-Global Stratified Transformer for Efficient Video Recognition. In: Avidan S, Brostow G, Cissé M, Farinella GM, Hassner T, editors. Computer Vision – ECCV 2022. vol. 13694 of Lecture Notes in Computer Science. Cham: Springer Nature Switzerland; 2022. p. 577-95.10.1007/978-3-031-19830-433
[30] Liu Z, Ning J, Cao Y, Wei Y, Zhang Z, Lin S, et al.Video Swin Transformer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; 2022. p. 3192-201.10.1109/CVPR52688.2022.00320
[31] Yan S, Xiong X, Arnab A, Lu Z, Zhang M, Sun C, et al. Multiview Transformers for Video Recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; 2022. p. 3323-33.10.1109/CVPR52688.2022.00333
[32] Fan H, Xiong B, Mangalam K, Li Y, Yan Z, Malik J, et al.Multiscale Vision Transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. IEEE; 2021. p. 6804-15.10.1109/ICCV48922.2021.00675
[33] Rahman F, Mubarek Ö, Kira Z. On the Surprising Effectiveness of Transformers in Low-Labeled Video Recognition; 2022. Preprint. Available from: https://arxiv.org/abs/2209.07474. 10.48550/arXiv.2209.07474
[34] Marais M, Brown D, Connan J, Boby A. Spatiotemporal Convolutions and Video Vision Transformers for Signer-Independent Sign Language Recognition. In: 2023 International Conference on Artificial Intelligence, Big Data, Computing and Data Communication Systems. IEEE; 2023. p. 1-6. 10.1109/icABCD59051.2023.10220534
[35] Garg M, Ghosh D, Pradhan PM.GestFormer: Multiscale Wavelet Pooling Transformer Network for Dynamic Hand Gesture Recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. IEEE; 2024. p. 2473-83. 10.1109/CVPRW63382.2024.00254
[36] Hashmi A, Shahzad SA, Lin CW, Tsao Y, Wang HM. AVTENet: A Human-Cognition-Inspired Audio-Visual Transformer-Based Ensemble Network for Video Deepfake Detection. IEEE Transactions on Cognitive and Developmental Systems. 2025;17(6):1360-76. 10.1109/TCDS.2025.3554477
[37] Yuan H, Cai Z, Zhou H, Wang Y, Chen X. TransAnomaly: Video Anomaly Detection Using Video Vision Transformer. IEEE Access. 2021;9:123977-86.10.1109/ACCESS.2021.3109102
[38] Ramadhani KN, Munir R, Utama NS. Improving Video Vision Transformer for Deepfake Video Detection Using Facial Landmark, Depthwise Separable Convolution and Self Attention. IEEE Access. 2024;12:8932-9. 10.1109/ACCESS.2024.3352890
[39] Poh JQ, See J, El Gayar N, Wong LK. Are You Paying Attention? Multimodal Linear Attention Transformers for Affect Prediction in Video Conversations. In: Proceedings of the 2nd InternationalWorkshop on Multimodal and Responsible Affective Computing. Association for Computing Machinery; 2024. p. 15-23. 10.1145/3689092.3689409
[40] Bajgoti A, Gupta R, Balaji P, Dwivedi R, Siwach M, Gupta D. SwinAnomaly: Real-Time Video Anomaly Detection Using Video Swin Transformer and SORT. IEEE Access. 2023;11:111093-105. 10.1109/ACCESS.2023.3321801
[41] Sun J, Dodge HH, Mahoor MH. MC-ViViT: Multi-Branch Classifier-ViViT to Detect Mild Cognitive Impairment in Older Adults Using Facial Videos. Expert Systems with Applications. 2024;238:121929. 10.1016/j.eswa.2023.121929
[42] Pătrăucean V, He XO, Heyward J, Zhang C, Sajjadi MSM, Muraru GC, et al. TRecViT: A Recurrent Video Transformer.Transactions on Machine Learning Research. 2026.Version of record; Google DeepMind publication page dated 9 January 2026. Available from: https://openreview.net/forum?id=Mmi46Ytb1H
[43] MISSING.
[44] Rao A, Jiang X, Wang S, Guo Y, Liu Z, Dai B, et al.. Temporal and Contextual Transformer for Multi-Camera Editing of TV Shows; 2022. Preprint.Available from: https://arxiv.org/abs/2210.08737. 10.48550/arXiv.2210.08737
[45] Naumenka J, Giuffrida V. Evaluating the ConvNeXt-ViViT Hybrid Architecture in Predicting Vehicular Crashes from Video Data. Open Science Framework; 2024. Repository preprint.Available from: https://osf.io/bnkw5/. 10.17605/OSF.IO/BNKW5
[46] Mohamed MAR, Srinath NS, Thomas J, Monica KM. Lane Change Prediction of Surrounding Vehicles Using Video Vision Transformers. International Journal of Computing and Digital Systems. 2025;17(1):1-9. 10.12785/ijcds/1571024445
[47] Jarry R, Chaumont M, Berti-Equille L, Subsol G. Predicting Socio-Economic Indicator Variations with Satellite Image Time Series and Transformer. In: Machine Vision for Earth Observation and Environment Monitoring. Glasgow, United Kingdom; 2024. MVEO 2024 workshop paper. Available from: https://hal-lirmm.ccsd.cnrs.fr/lirmm-04895134v2
[48] Wang J, Torresani L.Deformable Video Transformer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; 2022. p. 14033-42. 10.1109/CVPR52688.2022.01366
[49] Fiorentini G, Önal Ertuğrul I, Salah AA.Fully-Attentive and Interpretable: Vision and Video Vision Transformers for Pain Detection. In: NeurIPS 2022 Workshop on Vision Transformers: Theory and Applications; 2022. p. 1-12. Available from: https://arxiv.org/abs/2210.15769. 10.48550/arXiv.2210.15769
[50] Bargshady G, Joseph C, Hirachan N, Goecke R, Rojas RF.Acute Pain Recognition from Facial Expression Videos Using Vision Transformers. In: 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society. IEEE; 2024. p. 1-4.10.1109/EMBC53108.2024.1078161
[51] Wu Y, Guan R, Liang X, Zhang W, Jiang Y, Cui Y, et al.. Prediction of Immune Checkpoint Inhibitors Treatment Response of Non-Small Cell Lung Cancer Patients from Serial Computed Tomography Scans Based on Global Self-Attention Mechanism. Preprints.org; 2024. Preprint, version 1. Available from: https://www.preprints.org/manuscript/202407.1061/v1. 10.20944/preprints202407.1061.v1
[52] Płotka S, Grzeszczyk MK, Brawura-Biskupski-Samaha R, Gutaj P, Lipa M, Trzciński T. BabyNet: Residual Transformer Module for Birth Weight Prediction on Fetal Ultrasound Video.In: Wang L, Dou Q, Fletcher PT, Speidel S, Li S, editors. Medical Image Computing and Computer Assisted Intervention – MICCAI 2022. Lecture Notes in Computer Science. Cham: Springer Nature Switzerland; 2022. p. 350-9.10.1007/978-3-031-16440-834
[53] Emre T, Oghbaie M, Chakravarty A, Rivail A, Riedl S, Mai J, et al.Pretrained Deep 2.5D Models for Efficient Predictive Modeling from Retinal OCT: A PINNACLE Study Report. In: Antony B, Chen H, Fang H, Fu H, Lee CS, Zheng Y, editors. Ophthalmic Medical Image Analysis. vol. 14096 of Lecture Notes in Computer Science. Cham: Springer Nature Switzerland; 2023. p. 132-41.10.1007/978-3-031-44013-714
[54] Zheng Y. Heart Rate and Oxygen Level Estimation from Facial Videos Using a Hybrid Deep Learning Model. In: Agaian SS, DelMarco SP, Asari VK, editors. Multimodal Image Exploitation and Learning 2024. vol. 13033 of Proceedings of SPIE. National Harbor, Maryland, USA: SPIE; 2024. Eight-page conference paper.10.1117/12.3013956
[55] Hong J, Shin J, Choi J, Ko M. Robust Eye Blink Detection Using Dual Embedding Video Vision Transformer. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. IEEE; 2024. p. 6362-72. 10.1109/WACV57701.2024.00625
[56] Akan T, Alp S, Bhuiyan MS, Disbrow EA, Conrad SA, Vanchiere JA, et al.. Leveraging Video Vision Transformer for Alzheimer’s Disease Diagnosis from 3D Brain MRI; 2025. Preprint. Available from: https://arxiv.org/abs/2501.15733. 10.48550/arXiv.2501.15733
[57] Zheng Y.Hybrid Neural Network Models to Estimate Vital Signs from Facial Videos.BioMedInformatics. 2025;5(1):6.10.3390/biomedinformatics5010006
[58] Yang Y, Xing Z, Yu L, Huang C, Fu H, Zhu L. Vivim: A Video Vision Mamba for Medical Video Segmentation; 2024. Preprint. Available from: https://arxiv.org/abs/2401.14168. 10.48550/arXiv.2401.14168
[59] Yu L, Lezama J, Gundavarapu NB, Versari L, Sohn K, Minnen D, et al. Language Model Beats Diffusion: Tokenizer Is Key to Visual Generation. In: International Conference on Learning Representations; 2024. Available from: https : / / proceedings.iclr.cc/paper_files/paper/2024/hash/036912a83bdbb1fd792baf6532f102d8-Abstract-Conference.html
[60] Koner R, Jain G, Jain P, Tresp V, Paul S. LookupViT: Compressing Visual Information to a Limited Number of Tokens. In: Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T, Varol G, editors. Computer Vision – ECCV 2024. vol. 15144 of Lecture Notes in Computer Science. Cham: Springer Nature Switzerland; 2025. p. 322-37. 10.1007/978-3-031-73016-019
[61] Zhuang Y.Human Action Classification Task Based on Spatial and Temporal Transformer in Videos.Waseda University; 2022.Available from: http://hdl.handle.net/2065/00092760
[62] Mercea OB, Gritsenko A, Schmid C, Arnab A. Time-, Memory- and Parameter-Efficient Visual Adaptation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; 2024. p. 5536-45.10.1109/CVPR52733.2024.00529
[63] Dutson M, Li Y, Gupta M. Eventful Transformers: Leveraging Temporal Redundancy in Vision Transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. IEEE; 2023. p. 16865-77. 10.1109/ICCV51070.2023.01551
[64] Souza JS, Bedin E, Higa GTH, Loebens N, Pistori H. Pig Aggression Classification Using CNN, Transformers and Recurrent Networks; 2024. Preprint. Available from: https://arxiv.org/abs/2403.08528. 10.48550/arXiv.2403.08528
[65] Chowdhury S, Tisha SN, McGarrity ME, Hamerly G.Video-Based Recognition of Aquatic Invasive Species Larvae Using Attention-LSTM Transformer. In: Bebis G, Ghiasi G, Fang Y, Sharf A, Dong Y, Weaver C, et al., editors. Advances in Visual Computing. vol. 14361 of Lecture Notes in Computer Science. Cham: Springer Nature Switzerland; 2023. p. 224-35.10.1007/978-3-031-47969-418
[66] Yokota H, Bozkurtlar M, Yen B, Itoyama K, Nishida K, Nakadai K. A Video Vision Transformer for Sound Source Localization. In: 2024 32nd European Signal Processing Conference. IEEE; 2024. p. 106-10.10.23919/EUSIPCO63174.2024.10715427
[67] Jang HD, Kwon S, Nam H, Chang DE. Chemical Gas Source Localization with Synthetic Time Series Diffusion Data Using Video Vision Transformer. Applied Sciences. 2024;14(11):4451.10.3390/app14114451
[68] Constantin MG, Ionescu B.AIMultimediaLab at MediaEval 2022: Predicting Media Memorability Using Video Vision Transformers and Augmented Memorable Moments.In: Working Notes Proceedings of the MediaEval 2022 Workshop. vol. 3583 of CEUR Workshop Proceedings. Bergen, Norway and Online; 2023. Workshop held 12–13 January 2023.Available from: https://ceur-ws.org/Vol-3583/