FastViT-Based Lightweight Batak Script Recognition for Cultural Digital Preservation
DOI:
https://doi.org/10.55123/storage.v5i3.8850Keywords:
Batak Script Recognitin, deep learning, computer visionAbstract
Batak script is a cultural heritage of the Batak people in North Sumatra that requires preservation through digitization. This study proposes the FastViT-SA12 architecture for handwritten Batak script recognition using the Batak Char 20 dataset, which consists of 117 character classes. To address data imbalance, data augmentation was performed using ImgAug, resulting in 650 images per class. The dataset was divided into training, validation, and testing sets with a ratio of 70:20:10, and all images were resized to 64 × 64 pixels to reduce computational overhead. The model was trained for 30 epochs using the Adam optimizer with a learning rate of 0.001 and a batch size of 64. The proposed model achieved a training accuracy of 99.97%, a validation accuracy of 97.96%, and a test accuracy of 89.31%, with weighted precision, recall, and F1-score of 91.84%, 89.31%, and 88.99%, respectively. FastViT-SA12 demonstrates promising potential for recognizing 117 Batak script classes while maintaining computational efficiency with a compact parameter footprint, making it a viable candidate for mobile-first Optical Character Recognition (OCR) systems in manuscript preservation.
Downloads
References
Akallouch, O., Akallouch, M. and Fardousse, K. (2026) “TifinNet: CNN–Transformer Hybrid Architecture for Tifinagh Handwritten Recognition,” IEEE Access, 14, pp. 34830–34844. Available at: https://doi.org/10.1109/ACCESS.2026.3669937.
Alghyaline, S. (2024) “Optimised CNN Architectures for Handwritten Arabic Character Recognition,” Computers, Materials and Continua, 79(3), pp. 4905–4924. Available at: https://doi.org/https://doi.org/10.32604/cmc.2024.052016.
Ashimgaliyev, M. et al. (2026) “Development of Deep Learning Methods for Visual Document Classification Using Hybrid Vision Transformer–EfficientNet Architecture,” IEEE Access, 14, pp. 28041–28053. Available at: https://doi.org/10.1109/ACCESS.2025.3646481.
Ataman, F. (2025) “Data Augmentation Techniques in Deep Image Processing,” in Advances in Medical Diagnosis, Treatment, and Care (AMDTC) Book Series, pp. 231–262. Available at: https://doi.org/10.4018/979-8-3693-9816-6.ch010.
Babita and Nayak, D.R. (2024) “LiCT-Net: Lightweight Convolutional Transformer Network for Multiclass Breast Cancer Classification,” in TENCON 2024 - 2024 IEEE Region 10 Conference (TENCON), pp. 1359–1363. Available at: https://doi.org/10.1109/TENCON61640.2024.10902690.
Barillaro, L. (2025) “Deep Learning Platforms: PyTorch,” in Encyclopedia of Bioinformatics and Computational Biology (Second Edition). Oxford: Elsevier, pp. 167–170. Available at: https://doi.org/https://doi.org/10.1016/B978-0-323-95502-7.00093-2.
Carneiro, T. et al. (2018) “Performance Analysis of Google Colaboratory as a Tool for Accelerating Deep Learning Applications,” IEEE Access, 6, pp. 61677–61685. Available at: https://doi.org/10.1109/ACCESS.2018.2874767.
Gabriel, J. and Zhu, J. (2023) “FastViT : A Fast Hybrid Vision Transformer using Structural Reparameterization,” Computer Vision and Pattern Recognition [Preprint]. Available at: https://doi.org/https://doi.org/10.48550/arXiv.2303.14189.
Hasan, M.M. et al. (2020) “Preprocessing of Continuous Bengali Speech for Feature Extraction,” 2020 11th International Conference on Computing, Communication and Networking Technologies, ICCCNT 2020, pp. 1–4. Available at: https://doi.org/10.1109/ICCCNT49239.2020.9225469.
Hasugian, J.H. (2025) “PUTUSNYA PEWARISAN PENGETAHUAN PUSTAHA LAKLAK BATAK: ANALISIS HISTORIS DAN SOSIO-KULTURAL,” JURNAL NAGUR, 3(2), pp. 52–62.
Hatamizadeh, A. et al. (2024) “FASTERVIT: FAST VISION TRANSFORMERS WITH HIERARCHICAL ATTENTION,” ICLR, pp. 1–24.
Hurtik, P., Molek, V. and Hula, J. (2020) “Data Preprocessing Technique for Neural Networks Based on Image Represented by a Fuzzy Function,” IEEE Transactions on Fuzzy Systems, 28(7), pp. 1195–1204. Available at: https://doi.org/10.1109/TFUZZ.2019.2911494.
Ilić, A.S. (2020) “Preprocessing Image Data for Deep Learning,” in Sinteza, pp. 312–317. Available at: https://doi.org/10.15308/SINTEZA-2020-312-317.
Jayachandran, S., Selvakumar, Y. and Pearline, S.A. (2025) “Multi-Attention Based Convolutional Neural Network for Tamil Handwritten Character Recognition,” IEEE Access, 13, pp. 184360–184375. Available at: https://doi.org/10.1109/ACCESS.2025.3625588.
Jha, R.G. and Samlodia, A. (2024) “GPU-acceleration of tensor renormalization with PyTorch using CUDA,” Computer Physics Communications, 294, p. 108941. Available at: https://doi.org/https://doi.org/10.1016/j.cpc.2023.108941.
Jia, Y., Zheng, S. and Gu, L. (2026) “Enhancing adversarial transferability via hybrid data and model augmentation,” Neurocomputing, 697, p. 134144. Available at: https://doi.org/https://doi.org/10.1016/j.neucom.2026.134144.
Long, F. et al. (2025) “A diabetic retinopathy classification method based on image-text contrastive learning,” Pattern Recognition Letters, 193, pp. 142–148. Available at: https://doi.org/https://doi.org/10.1016/j.patrec.2025.04.017.
Ran, Z. et al. (2025) “FastViT: Real-Time Linear Attention Accelerator for Dense Predictions of Vision Transformer (ViT),” in 2025 IEEE International Symposium on Circuits and Systems (ISCAS), pp. 1–5. Available at: https://doi.org/10.1109/ISCAS56072.2025.11043624.
Siregar, A.M. (2017) “Pelestarian Aksara Batak Toba sebagai Warisan Budaya Melalui Media Edukasi Interaktif,” Jurnal Bahasa dan Budaya, 3(2), pp. 88–95.
Sohail, M. et al. (2024) “Deep Learning Based Multi Pose Human Face Matching System,” IEEE Access, 12, pp. 26046–26061. Available at: https://doi.org/10.1109/ACCESS.2024.3366451.
Stephen, R.K. et al. (2025) “Fast Vision Transformer Framework for Proactive Diabetic Retinopathy Diagnosis in Fundus Images,” in International Conference on Intelligent Systems and Digital Transformation (ICISD 2025). Atlantis Press, pp. 774–786. Available at: https://doi.org/10.2991/978-94-6463-866-0_63.
Sudewo, E.D.B. et al. (2025) “Evaluating the Impact of Optimizer Hyperparameters on ResNet in Hanacaraka Character Recognition,” Preservation, Digital Technology & Culture, pp. 1–11. Available at: https://doi.org/10.1515/pdtc-2024-0061.
Wang, Z. et al. (2026) “A Comprehensive Survey on Data Augmentation,” IEEE Transactions on Knowledge and Data Engineering, 38(1), pp. 47–66. Available at: https://doi.org/10.1109/TKDE.2025.3622600.
Willian, S., Rochadiani, T.H. and Sofian, T. (2023) “Design of Batak Toba Script Recognition System Using Convolutional Neural Network Algorithm,” 8(3), pp. 1609–1618.
Xu, D. et al. (2025) “DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based image captioning,” Pattern Recognition, 164, p. 111522. Available at: https://doi.org/https://doi.org/10.1016/j.patcog.2025.111522.
Ye, H. et al. (2025) “Lightweight Deep Learning Model for Lung Ultrasound Image Scoring,” in 2025 IEEE International Ultrasonics Symposium (IUS), pp. 1–3. Available at: https://doi.org/10.1109/IUS62464.2025.11201808.
Zhang, H. et al. (2021) “A Data Preprocessing Method for Automatic Modulation Classification Based on CNN,” IEEE Communications Letters, 25(4), pp. 1206–1210. Available at: https://doi.org/10.1109/LCOMM.2020.3044755.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Egi Dio Bagus Sudewo, Hainur Rasyid, Muhammad Kunta Biddinika, Kariyamin, Dini Khairunnisa Tanjung

This work is licensed under a Creative Commons Attribution 4.0 International License.



















