RTL Design of a CNN-Based FPGA Accelerator for Handwritten Digit Recognition

Authors

  • Vivek Rathi SCTR'S Pune Institute of Computer Technology
  • Vedant Kulkarni SCTR'S Pune Institute of Computer Technology
  • Harshal Ingle SCTR'S Pune Institute of Computer Technology
  • Mrs. P.S. Agnihotri SCTR'S Pune Institute of Computer Technology

DOI:

https://doi.org/10.18486/ijcsnt/15.1.005

Keywords:

Handwritten Digit Recognition, LeNet-5, FPGA Accelerator, Convolutional Neural Network, RTL Design, Field Programmable Gate Array, Kintex-7

Abstract

Handwritten digit recognition plays a key role in areas like mail sorting, banking, and document processing. Software-based Convolutional Neural Networks (CNNs) achieve high recognition accuracy, but their computing power and energy needs limit their use in embedded systems. This paper presents the design and RTL-level simulation of a CNN-based hardware accelerator for handwritten digit recognition targeting a Xilinx Kintex-7 FPGA. The accelerator implements a modified LeNet-5 CNN architecture entirely in Verilog using fixed-point arithmetic. It consists of convolution layers followed by ReLU activation and max pooling, along with a fully connected output layer. The CNN model is trained offline using Python, and the trained and quantized weights and biases are converted to fixed-point representation and mapped directly into the RTL design. A streaming, line-buffer-based architecture is employed to efficiently generate convolution windows and enable continuous data flow through the network. The design is verified at the RTL level using cycle-accurate simulation with MNIST test images, ensuring functional correctness of all network layers. The proposed accelerator achieves a classification accuracy of 91% on 1,000 MNIST test images while utilizing 12,709 Slice LUTs and 238 DSP slices, with only 796 trainable parameters. These results demonstrate that the RTL-based CNN accelerator provides an effective trade-off between recognition accuracy and hardware resource utilization, making it suitable for embedded FPGA-based inference. Hardware implementation and on-board validation are planned as future work.

References

O. Choudhari, S. Dabhadkar, M. Chopade, V. Ingale, and S. Chopde, “Hardware accelerator: Implementation of CNN on FPGA for digit recognition,” in Proc. IEEE International Conference on Emerging Trends in Information Technology and Engineering (IC-ETITE), Vellore, India, Feb. 2020, pp. 1–6, doi: 10.1109/ICETITE47903.2020.194. DOI: https://doi.org/10.1109/VDAT50263.2020.9190274

R. Leveugle, A. Cogney, A. B. Gah El Hilal, T. Lailler, and M. Pieau, “Hardware acceleration and approximation of CNN computations: Case study on an integer version of LeNet,” Electronics, vol. 13, no. 13, p. 2709, 2024. DOI: https://doi.org/10.3390/electronics13142709

F. Yasir and M. Kazmi, “Acceleration of optical character recognition on Zynq UltraScale+ MPSoC using deep convolutional neural networks,” IEEE Access, vol. 13, pp. 135538–135555, 2025, doi: 10.1109/ACCESS.2025.3593294. DOI: https://doi.org/10.1109/ACCESS.2025.3593294

R. Tanno and K. Yanai, “Caffe2C: A framework for easy implementation of CNN-based mobile applications,” in Adjunct Proceedings of the 13th International Conference on Mobile and Ubiquitous Systems: Computing, Networking and Services, ACM, 2016, pp. 159–164. DOI: https://doi.org/10.1145/3004010.3004025

R. Zhao, W. Luk, X. Niu, H. Shi, and H. Wang, “Hardware acceleration for machine learning,” in IEEE Computer Society Annual Symposium on VLSI, 2017. DOI: https://doi.org/10.1109/ISVLSI.2017.127

Y. Cheng, D. Wang, P. Zhou, and T. Zhang, “A survey of model compression and acceleration for deep neural networks,” IEEE Signal Processing Magazine, Special Issue on Deep Learning for Image Understanding, Sep. 2019.

Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, Nov. 1998. DOI: https://doi.org/10.1109/5.726791

D. Rongshi and T. Yongming, “Accelerator implementation of LeNet-5 convolution neural network based on FPGA with HLS,” in Proceedings of the 3rd International Conference on Circuits, System and Simulation (ICCSS), Nanjing, China, 2019, pp. 64–67, doi: 10.1109/CIRSYSSIM.2019.8935599. DOI: https://doi.org/10.1109/CIRSYSSIM.2019.8935599

H. Madadum and Y. Becerikli, “FPGA-based optimized convolutional neural network framework for handwritten digit recognition,” in Proceedings of the 1st International Informatics and Software Engineering Conference (UBMYK), Ankara, Turkey, 2019, pp. 1–6, doi: 10.1109/UBMYK48245.2019.8965628. DOI: https://doi.org/10.1109/UBMYK48245.2019.8965628

Y. Zhou and J. Jiang, “An FPGA-based accelerator implementation for deep convolutional neural networks,” in Proceedings of the 4th International Conference on Computer Science and Network Technology (ICCSNT), Harbin, China, 2015, pp. 829–832, doi: 10.1109/ICCSNT.2015.7490869. DOI: https://doi.org/10.1109/ICCSNT.2015.7490869

J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller, “Striving for simplicity: The all convolutional net,” in Proceedings of the International Conference on Learning Representations (ICLR), 2015.

K. Abdelouahab, M. Pelcat, J. Serot, and F. Berry, “Accelerating CNNs on FPGA: A survey,” arXiv preprint arXiv:1806.01683, May 2018.

C. Zhang, P. Li, G. Sun, Y. Guan, B. Xiao, and J. Cong, “Optimizing FPGA-based accelerator design for deep convolutional neural networks,” in Proceedings of the ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, Feb. 2015. DOI: https://doi.org/10.1145/2684746.2689060

A. Shawahna, S. M. Sait, and A. El-Maleh, “FPGA-based accelerators of deep learning networks for learning and classification: A review,” IEEE Access, vol. 6, pp. 7823–7859, 2018. DOI: https://doi.org/10.1109/ACCESS.2018.2890150

M. Bettoni, G. Urgese, Y. Kobayashi, E. Macii, and A. Aquaviva, “A convolution neural network fully implemented on FPGA for embedded platforms,” in Proceedings of the IEEE International Conference on Circuits and Systems (CAS), 2017. DOI: https://doi.org/10.1109/NGCAS.2017.16

M. Bojarski et al., “End to end learning for self-driving cars,” arXiv preprint arXiv:1604.07316, 2016.

M. Motamedi, F. Portillo, M. Saffarpour, D. Fong, and S. Ghiasi, “Resource scalable CNN synthesis for IoT applications,” arXiv preprint arXiv:1901.00738, Dec. 2018.

F. Farrukh, M. A. Khan, and S. Rehman, “FPGA-based acceleration of convolutional neural networks for image classification,” in Proceedings of the International Conference on Computing, Mathematics and Engineering Technologies, Sukkur, Pakistan, 2018.

H. Hareth, A. Hassan, and M. Ali, “Efficient FPGA implementation of convolutional neural networks,” Microprocessors and Microsystems, vol. 67, pp. 1–10, 2019.

A. Mujawar and S. Patil, “Design and implementation of convolutional neural network accelerator on FPGA,” in Proceedings of the International Conference on Advances in Computing, Communication and Control, Mumbai, India, 2018.

S. Shan, X. Zhang, and Y. Liu, “High-performance FPGA-based convolutional neural network accelerator,” IEEE Embedded Systems Letters, vol. 12, no. 3, pp. 85–88, 2020.

Y. Wang, L. Chen, and J. Xu, “Optimized FPGA architecture for deep convolutional neural networks,” IEEE Transactions on Circuits and Systems I, vol. 69, no. 4, pp. 1567–1579, 2022.

M. Arshad, A. Khan, and S. A. Malik, “Hardware implementation of deep learning accelerators on FPGA platforms,” in Proceedings of the International Conference on Artificial Intelligence and Data Analytics, Lahore, Pakistan, 2020.

G. Dinelli, M. Poncino, and E. Macii, “Energy-efficient convolutional neural network inference on FPGA,” IEEE Access, vol. 8, pp. 112345–112356, 2020.

Y. Jiang, Q. Liu, and J. Li, “Design of a scalable convolutional neural network accelerator on FPGA,” in Proceedings of the IEEE International Conference on Electronics and Information Engineering, Nanjing, China, 2019.

R. Pisharody and S. K. Nandy, “A high-throughput FPGA accelerator for convolutional neural networks,” ACM Transactions on Embedded Computing Systems, vol. 20, no. 4, pp. 1–24, 2021.

X. Si and H. Zhang, “FPGA-based convolutional neural network accelerator with optimized memory access,” in Proceedings of the International Conference on Computer Engineering and Technology, Chengdu, China, 2018. DOI: https://doi.org/10.1109/ICSICT.2018.8565671

X. Si, H. Zhang, and Y. Chen, “Improved FPGA implementation of convolutional neural networks,” Journal of Systems Architecture, vol. 96, pp. 45–56, 2019.

F. Natale and M. Martina, “Low-power FPGA implementation of convolutional neural networks,” in Proceedings of the IEEE International Conference on Design and Technology of Integrated Systems (DTIS), Palermo, Italy, 2017.

Downloads

Published

2026-04-30

How to Cite

RTL Design of a CNN-Based FPGA Accelerator for Handwritten Digit Recognition. (2026). International Journal of Communication Systems and Network Technologies, 15(1), 69-82. https://doi.org/10.18486/ijcsnt/15.1.005