Implementation of energy-based methods for on-chip training of Magnetic Tunnel junction based neural networks
Implementing Organization
Indian Institute Of Technology, Gandhinagar
Principal Investigator
Dr. Naveen Sisodia
Indian Institute Of Technology, Gandhinagar, Gujarat
naveen.sisodia@iitgn.ac.in
CO-Principal Investigator
Nil
Project Overview
With the advancements in the field of Artificial Intelligence (AI), there is a strong demand for energy-efficient hardware platforms suitable for emulating software-based Deep Neural Networks (DNNs). Physical Neural Networks (PNNs), which encompass different hardware (memristive, photonic, magnetic, etc.) that can perform calculations required in neural networks using a physical property of the system, have gained attention for use as AI hardware due to their potential energy-efficiency when scaling for large networks. In addition, parallel efforts are also underway to propose algorithms for faster, energy-efficient training of neural networks to minimise the otherwise massive energy requirements during training. A simultaneous effort is thus needed to find optimal energy-efficient hardware along with a compatible training algorithm that can minimize energy costs and accommodate variabilities in real devices. In this project, we aim to demonstrate an energy-based protocol, Equilibrium Propagation (EP), for ultrafast, low-power and parallelized training of Magnetic Tunnel Junction (MTJ) based neural networks. These energy-based protocols are an alternative to conventionally used BackPropagation (BP) and provide a gradient-free method to update the network weights. In addition, they inherently account for device variabilities during the training process making them an excellent choice for training PNNs. The MTJ hardware platform was chosen for our project due to their low-energy, ultrafast operation speed and tunable microwave properties that can be utilized for developing radio frequency neural networks. Such MTJ networks have previously been shown to have performance equivalent to that of software-based neural networks. In this project, we will focus our efforts on implementing Equilibrium Propagation (EP) for training of interconnected MTJ network and propose a novel architecture with bi-directional synaptic connections ensuring compatibility with EP protocol. We account for the device variability and the asymmetry in the forward and backward passes during training by implementing Direct Feedback Alignment (DFA) technique. To perform accurate simulations accounting for potential latency between each layer of the network, we will build a large-scale simulation model integrating the exact physics of each MTJ device (Landau Lifshitz Gilbert equation) in our network with DNN models developed in Pytorch. The developed models will be benchmarked using standard datasets and compared with state-of-the-art software-based networks. We will also exploit the dynamic nature of our network to demonstrate prediction/forecasting tasks with time-series inputs. The developed protocols can be rapidly deployed on MTJ devices with potential applications in IoT devices, smart sensors and edge computing.