Low-Complexity lassification Technique and Hardware-Efficient Classify-Unit Architecture for CNN Accelerator

Publications

Low-Complexity lassification Technique and Hardware-Efficient Classify-Unit Architecture for CNN Accelerator

Year : 2024

Publisher : IEEE Computer Society

Source Title : Proceedings of the IEEE International Conference on VLSI Design

Document Type :

Abstract

This paper proposes simplified classification technique to reduce the complexity of softmax-based classification in the convolutional neural network (CNN) inference engine/accelerator. It primarily allows the CNN accelerator to directly classify the object from the activation of fully connected (FC) layer that avoids complex exponential and divisive operations. Corresponding to the suggested technique, this work also presents a hardware-efficient VLSI architecture of classify unit for CNN accelerator. Furthermore, the proposed classify-unit architecture has been ASIC synthesized and post-layout simulated in 28 nm-FDSOI technology node. As a result, our design delivers a peak throughput of 2.5 GIPS with a hardware efficiency of 5.05× 103 GIPS/mW/mm2. Comparison of these results with the relevant reported works indicates that the proposed classify unit manifests 24.1 × lesser area and 12.5× better hardware efficiency than the state-of-the-art implementations. Finally, complete CNN accelerator that is integrated with the proposed classify unit has been functionally validated with the aid of Zynq UltraScale+ ZCU102 FPGA-board in real-world scenario, using the MobileNet-V2 CNN model.