×

img Accessibility Controls

Research Projects Banner

Research Projects

Exploring Compute-Memory Integration Strategies for Accelerator-Driven High-Performance and Energy-Efficient Systems

Implementing Organization

Indian Institute Of Technology Madras
Principal Investigator
Prof. Madhu Mutyam
Indian Institute Of Technology Madras
madhu@cse.iitm.ac.in
CO-Principal Investigator
Dr. Debiprasanna Sahoo
Indian Institute Of Technology Bhubaneswar, Argul - Jatni Road, Kansapada,Odisha,Khordha-752050

Project Overview

The widespread adoption of artificial intelligence (AI) and machine learning (ML) has fuelled the evolution of modern computing architectures and led to a paradigm shift toward accelerator-driven architectures. Due to heavy data movement between compute units and memory subsystems, the performance of existing heterogeneous systems is increasingly limited. The classical von Neumann bottleneck and inefficiencies in existing memory hierarchies hinder both performance and energy efficiency. Emerging technologies like high-bandwidth memory (HBM), processing-in-memory (PIM), and chiplet-based architectures offer new opportunities for integration but require systematic exploration and evaluation. Tightly integrated compute and memory systems can significantly reduce energy consumption and latency. However, designing and evaluating such systems requires a new class of tools and workloads. Although there exist various architectural simulators modeling CPUs, accelerators, HBMs, PIMs, and NoCs, a single simulator connecting all these modules is not available. Similarly, while there are various benchmark suites for AI/ML workloads, building a benchmark suite for next-generation workloads really helps in proposing efficient heterogeneous systems. However, although many innovative techniques have been proposed in the literature, as mentioned above, none of them has conducted a systematic study on system-level efficiency. This project addresses these needs by building an ecosystem comprising a modular simulator framework, a benchmark suite that, together with the simulator, enables rigorous analysis and evaluation of emerging scenarios, and innovative architectural techniques for exploring compute-memory integration strategies tailored to accelerator-centric, next-generation computing systems. We propose to design a modular architectural simulator that models CPUs, NPUs, HBMs and PIMs and connects them using a suitable NoC model. We develop a simulation framework for open-source community and use RISC-V-based ISA (an open-source ISA) for the development of NPU components of the simulator. We believe that our efforts will foster open source hardware development, which is one of the major goals of the Indian Semiconductor Mission. We embed existing power models proposed in the literature into our framework to approximate and project the power usage of the overall system. We would like to provide the performance and power simulation infrastructure along with the modern workloads for the swift performance and power tuning of such systems. The simulation infrastructure will be thoroughly verified and validated. We then develop benchmark suites that represent the essential features of next-generation ML workloads and federated learning models. We explore techniques that enhance system-level efficiency in terms of compute, memory, and communication efficiency. From a compute perspective, we limit the scope of this proposal to interactions between CPUs and NPUs. By understanding workload characteristics, we identify the portions of computations that can be scheduled on CPUs and NPUs. We explore data compression techniques to minimise the amount of data transmitted between various compute and memory elements. We explore processing-in-memory mechanisms to address bandwidth/energy issues. If all the objectives are reached, we would have a publicly downloadable modular architectural simulator connecting compute (both CPUs and NPUs) and memory elements (HMB and PIM) with NoC, along with benchmark suite for next-generation ML workloads. We believe our workload analysis provides insights that can help coming up with efficient microarchitectual techniques for heterogenous systems.
Funding Organization
Quick Information
Area of Research
Engineering Sciences
Focus Area
Computer Science And Engineering
Start Date
26 Mar 2026
End Date
25 Mar 2029
Status
ongoing
Output
No. of Research Paper
00
Technologies (If Any)
00
No. of PhD Produced
00
Publications
00
No. of Patents
Filed : 00
Grant : 00
arrowtop
Latest Updates
Loading…