Energy-Efficient FPGA-Based Accelerator Architecture for Real-Time Edge AI Applications in Embedded Systems
Keywords:
FPGA Accelerator; Edge Artificial Intelligence; Embedded Systems; Energy-Efficient Computing; Real-Time Inference; Hardware–Software Co-DesignAbstract
The expanding trend in intelligent edge devices uses in areas of smart healthcare, industrial surveillance, autonomous system, and environmental sensing has generated a significant need to perform real-time artificial intelligence (AI) on embedded systems. Edge AI systems have to handle data streams continuously with hard limitations on the level of latencies, power, and memory consumption in addition to predictable inference performance. The use of conventional CPU and GPU solutions are usually not enough, because of high power requirements, instruction overheads, and low memory access efficiency, to execute deep learning problems on resource-constrained devices. The present paper provides a high-speed edge AI accelerator architecture built on FPGA that meets real-time requirements based on the energy-efficient product development of embedded systems. The architecture proposed is a parallel and pipelined processing architecture that is optimised to perform neural network inferences, which is effective at performing computationally consuming functions like multiply accumulate (MAC) computations. A lean on-chip hierarchy of memory and custom compute units are combined together in order to reduce the amount of data movement and maximise off-chip memory access, enhancing energy efficiency and processing speed. The experimental findings indicate that the suggested FPGA accelerator lowers the inference latency and energy usage largely relative to the traditional processor-based inference models and still incurs the highest inference accuracy levels. The architecture offers an efficient and scalable platform that can be used to deploy real-time AI inference in next-generation embedded and edge computing platforms.
