Reinforcement Learning-Enabled Design Space Exploration for Energy-Efficient AI Accelerator Architectures
Keywords:
Reinforcement learning, Design space exploration, AI accelerators, Energy efficiency, VLSI architecture, Hardware–software co-design.Abstract
The development of AI deep-learning workloads has accelerated and workloads on edge
and cloud computing systems have increased significantly posing a fundamental requirement
on special AI accelerator architectures capable of providing high levels of computational
performance at operation with strict energy requirements. The construction of such
accelerators is a computationally infeasible and scalability issue due to the sheer size and
complexity of the architectural design space that such traditional design space exploration
(DSE) methods as exhaustive search or heuristic optimization methods can explore with ease.
To help manage these issues, this article presents a reinforcement learning (RL)-based DSE
architecture to the automated and energy-efficient design of AI accelerator architectures.
The suggested scheme defines accelerator configuration optimisation as a series of
decision-making, where an RL agent engages with an environment of performance and
energy analysis, and performs successive exploration of architectural parameters, i.e., the
organization of processing elements, memory hierarchy configuration, dataflow strategies,
and arithmetic precision. Multi-objective reward function is used to jointly maximize energy,
latency, and the use of hardware resources. With RL agent can detect the existence of
Pareto-optimal accelerator designs without necessarily searching the entire design space,
the RL agent reaches such optimal designs through a series of trial and error executions, and
feedbacks. Through experimental assessment based on representative representative deep
neural network benchmarks, it has been shown that the RL-based DSE framework is able
to reliably determine architectures with dramatically lower energy use, and with superior
energy performance trade-offs, than traditional heuristic architecture exploration tools. The
findings also indicate a stable convergence pattern and good exploration of high dimensions
design space. In summary, the paper presents reinforcement as a scalable, intelligent and
effective optimization paradigm of next generation energy-efficient design of AI accelerator,
which shows great promise in automated hardware-software co-design of future AI systems.
