Apvit: ViT with adaptive patches for scene text recognition
Abstract Scene texts in nature exhibit varied colors, which serve as a significant distinguishing feature that effectively suppresses background interference. In this study, color clustering is utilized as a prior guide to group patches, enhancing their spatial relationships. Additionally, patch siz...
| Published in: | Discover Applied Sciences |
|---|---|
| Main Authors: | Ning Zhang, Ce Li, Zongshun Wang, Jialin Ma, Zhiqiang Feng |
| Format: | Article |
| Language: | English |
| Published: |
Springer
2025-03-01
|
| Subjects: | |
| Online Access: | https://doi.org/10.1007/s42452-025-06570-9 |
Similar Items
ViTSTR-Transducer: Cross-Attention-Free Vision Transformer Transducer for Scene Text Recognition
by: Rina Buoy, et al.
Published: (2023-12-01)
by: Rina Buoy, et al.
Published: (2023-12-01)
Patch and Model Size Characterization for On-Device Efficient-ViTs on Small Datasets Using 12 Quantitative Metrics
by: Jurn-Gyu Park, et al.
Published: (2025-01-01)
by: Jurn-Gyu Park, et al.
Published: (2025-01-01)
TDA-ViT: A Transformer-Based Framework for Unified Urdu Text Recognition via Topological and Visual Feature Fusion
by: Shahbaz Hassan, et al.
Published: (2025-01-01)
by: Shahbaz Hassan, et al.
Published: (2025-01-01)
Identifying Malignant Breast Ultrasound Images Using ViT-Patch
by: Hao Feng, et al.
Published: (2023-03-01)
by: Hao Feng, et al.
Published: (2023-03-01)
Permeability Prediction Using Vision Transformers
by: Cenk Temizel, et al.
Published: (2025-07-01)
by: Cenk Temizel, et al.
Published: (2025-07-01)
HGR-ViT: Hand Gesture Recognition with Vision Transformer
by: Chun Keat Tan, et al.
Published: (2023-06-01)
by: Chun Keat Tan, et al.
Published: (2023-06-01)
Interpretable Deep Learning for Diabetic Retinopathy: A Comparative Study of CNN, ViT, and Hybrid Architectures
by: Weijie Zhang, et al.
Published: (2025-05-01)
by: Weijie Zhang, et al.
Published: (2025-05-01)
A New Pes Planus Automatic Diagnosis Method: ViT-OELM Hybrid Modeling
by: Derya Avcı
Published: (2025-03-01)
by: Derya Avcı
Published: (2025-03-01)
Tumor ViT-GRU-XAI: Advanced Brain Tumor Diagnosis Framework: Vision Transformer and GRU Integration for Improved MRI Analysis: A Case Study of Egypt
by: Mohammed Aly, et al.
Published: (2024-01-01)
by: Mohammed Aly, et al.
Published: (2024-01-01)
Cross-Parallel Transformer: Parallel ViT for Medical Image Segmentation
by: Dong Wang, et al.
Published: (2023-11-01)
by: Dong Wang, et al.
Published: (2023-11-01)
ViT-RoT: Vision Transformer-Based Robust Framework for Tomato Leaf Disease Recognition
by: Sathiyamohan Nishankar, et al.
Published: (2025-06-01)
by: Sathiyamohan Nishankar, et al.
Published: (2025-06-01)
ViT-based Terrain Recognition System for wearable soft exosuit
by: Fangliang Yang, et al.
Published: (2023-03-01)
by: Fangliang Yang, et al.
Published: (2023-03-01)
Turkish scene text recognition: Introducing extensive real and synthetic datasets and a novel recognition model
by: Serdar Yıldız
Published: (2024-12-01)
by: Serdar Yıldız
Published: (2024-12-01)
A Finger Vein Liveness Detection System Based on Multi-Scale Spatial-Temporal Map and Light-ViT Model
by: Liukui Chen, et al.
Published: (2023-12-01)
by: Liukui Chen, et al.
Published: (2023-12-01)
ViT-PSO-SVM: Cervical Cancer Predication Based on Integrating Vision Transformer with Particle Swarm Optimization and Support Vector Machine
by: Abdulaziz AlMohimeed, et al.
Published: (2024-07-01)
by: Abdulaziz AlMohimeed, et al.
Published: (2024-07-01)
An Experiment of Cone-Sleeve Recognition in UAV Aerial Refueling Using Multi-Modal ViT Technology
by: Zhangjun Sun, et al.
Published: (2026-05-01)
by: Zhangjun Sun, et al.
Published: (2026-05-01)
An Investigation on Prediction of Infrastructure Asset Defect with CNN and ViT Algorithms
by: Nam Lethanh, et al.
Published: (2025-05-01)
by: Nam Lethanh, et al.
Published: (2025-05-01)
Context Transformer and Adaptive Method with Visual Transformer for Robust Facial Expression Recognition
by: Lingxin Xiong, et al.
Published: (2024-02-01)
by: Lingxin Xiong, et al.
Published: (2024-02-01)
An Attention-Based Framework for Detecting Face Forgeries: Integrating Efficient-ViT and Wavelet Transform
by: Yinfei Xiao, et al.
Published: (2025-08-01)
by: Yinfei Xiao, et al.
Published: (2025-08-01)
ViT-Based Face Diagnosis Images Analysis for Schizophrenia Detection
by: Huilin Liu, et al.
Published: (2024-12-01)
by: Huilin Liu, et al.
Published: (2024-12-01)
Efficient Spatiotemporal Vegetation Pixel Classification With Vision Transformers
by: Alan Gomes, et al.
Published: (2026-01-01)
by: Alan Gomes, et al.
Published: (2026-01-01)
Cascaded Dual-Inpainting Network for Scene Text
by: Chunmei Liu
Published: (2025-07-01)
by: Chunmei Liu
Published: (2025-07-01)
ViT-FRD: A Vision Transformer Model for Cardiac MRI Image Segmentation Based on Feature Recombination Distillation
by: Chunyu Fan, et al.
Published: (2023-01-01)
by: Chunyu Fan, et al.
Published: (2023-01-01)
Remote Sensing Crop Water Stress Determination Using CNN-ViT Architecture
by: Kawtar Lehouel, et al.
Published: (2024-05-01)
by: Kawtar Lehouel, et al.
Published: (2024-05-01)
LMS-ViT: a multi-scale vision transformer approach for real-time smartphone-based skin cancer detection
by: A. Anny Leema, et al.
Published: (2025-09-01)
by: A. Anny Leema, et al.
Published: (2025-09-01)
Unlocking the Potential of XAI for Improved Alzheimer’s Disease Detection and Classification Using a ViT-GRU Model
by: S. M. Mahim, et al.
Published: (2024-01-01)
by: S. M. Mahim, et al.
Published: (2024-01-01)
Visual place recognition from end-to-end semantic scene text features
by: Zobeir Raisi, et al.
Published: (2024-09-01)
by: Zobeir Raisi, et al.
Published: (2024-09-01)
A Robust Hybrid CNN+ViT Framework for Breast Cancer Classification Using Mammogram Images
by: Vasudha Rani Patheda, et al.
Published: (2025-01-01)
by: Vasudha Rani Patheda, et al.
Published: (2025-01-01)
ViT-HHO: Optimized vision transformer for diabetic retinopathy detection using Harris Hawk optimization
by: Vishal Awasthi, et al.
Published: (2024-12-01)
by: Vishal Awasthi, et al.
Published: (2024-12-01)
Multi-Scale and Multi-Factor ViT Attention Model for Classification and Detection of Pest and Disease in Agriculture
by: Mingyao Xie, et al.
Published: (2024-07-01)
by: Mingyao Xie, et al.
Published: (2024-07-01)
Gait-ViT: Gait Recognition with Vision Transformer
by: Jashila Nair Mogan, et al.
Published: (2022-09-01)
by: Jashila Nair Mogan, et al.
Published: (2022-09-01)
A Fault Diagnosis Method for Rolling Bearing Based on 1D-ViT Model
by: Pinghu Xu, et al.
Published: (2023-01-01)
by: Pinghu Xu, et al.
Published: (2023-01-01)
ViT-UNet: A Vision Transformer Based UNet Model for Coastal Wetland Classification Based on High Spatial Resolution Imagery
by: Nan Zhou, et al.
Published: (2024-01-01)
by: Nan Zhou, et al.
Published: (2024-01-01)
ViT-BiLSTM Multimodal Learning for Paediatric ADHD Recognition: Integrating Wearable Sensor Data with Clinical Profiles
by: Lin Wang, et al.
Published: (2025-10-01)
by: Lin Wang, et al.
Published: (2025-10-01)
Class-Aware Self-Distillation for Remote Sensing Image Scene Classification
by: Bin Wu, et al.
Published: (2024-01-01)
by: Bin Wu, et al.
Published: (2024-01-01)
A Lightweight Dual-Branch Swin Transformer for Remote Sensing Scene Classification
by: Fujian Zheng, et al.
Published: (2023-05-01)
by: Fujian Zheng, et al.
Published: (2023-05-01)
Enhancing Cervical Pre-Cancerous Classification Using Advanced Vision Transformer
by: Manal Darwish, et al.
Published: (2023-09-01)
by: Manal Darwish, et al.
Published: (2023-09-01)
Benchmarking MedViT and hybrid CNN–ViT architectures for multi-label thoracic disease classification
by: Victor Agbo, et al.
Published: (2026-04-01)
by: Victor Agbo, et al.
Published: (2026-04-01)
A Robot Object Recognition Method Based on Scene Text Reading in Home Environments
by: Shuhua Liu, et al.
Published: (2021-03-01)
by: Shuhua Liu, et al.
Published: (2021-03-01)
Binary and Multi-Class Classification of Colorectal Polyps Using CRP-ViT: A Comparative Study Between CNNs and QNNs
by: Jothiraj Selvaraj, et al.
Published: (2025-07-01)
by: Jothiraj Selvaraj, et al.
Published: (2025-07-01)
Similar Items
-
ViTSTR-Transducer: Cross-Attention-Free Vision Transformer Transducer for Scene Text Recognition
by: Rina Buoy, et al.
Published: (2023-12-01) -
Patch and Model Size Characterization for On-Device Efficient-ViTs on Small Datasets Using 12 Quantitative Metrics
by: Jurn-Gyu Park, et al.
Published: (2025-01-01) -
TDA-ViT: A Transformer-Based Framework for Unified Urdu Text Recognition via Topological and Visual Feature Fusion
by: Shahbaz Hassan, et al.
Published: (2025-01-01) -
Identifying Malignant Breast Ultrasound Images Using ViT-Patch
by: Hao Feng, et al.
Published: (2023-03-01) -
Permeability Prediction Using Vision Transformers
by: Cenk Temizel, et al.
Published: (2025-07-01)
