×

img Accessibility Controls

Research Projects Banner

Research Projects

Enhancing Hardware Security and Throughput in Vision Tasks

Implementing Organization

Principal Investigator
Dr. Palash Das
Indian Institute Of Technology Jodhpur, Rajasthan
palashdas@iitj.ac.in
CO-Principal Investigator
Nil

Project Overview

Transformers, known for their self-attention mechanism, have proven highly effective in natural language processing (NLP) tasks such as machine translation, language modeling, and semantic analysis. This success has extended beyond NLP, with transformer-based architectures being adapted for computer vision, giving rise to Vision Transformers (ViTs). What makes ViTs particularly powerful is their ability to learn relationships between image elements, regardless of their spatial positions, enabling them to capture global dependencies across features. ViTs have demonstrated broad applicability across various domains, including image recognition, object detection, segmentation, image super-resolution, video understanding, image processing, text-image synthesis, visual question answering, and many other tasks. Although ViTs are highly effective for vision tasks, they encounter two major challenges: (1) susceptibility to adversaria and bit-flip attacks, and (2) efficiency concerns, especially in resource-constrained environments, due to the quadratic increase in computational complexity as the number of tokens grows. ViTs are vulnerable to adversarial and bit-flip attacks, which can exploit weaknesses in the model’s perception. These attacks subtly alter input images/ models, causing the network to misinterpret and make inaccurate predictions. Recent studies have shown that even small, carefully designed perturbations can lead ViTs to confidently make incorrect classifications. Additionally, ViT models can require hundreds of GFLOPs during inference, making them highly computationally and memory-intensive. Researchers have attempted to address the issue of attacks by designing detection and rectification algorithms. However, these algorithms are typically executed on traditional CPU/GPU-based systems, while the high computational demands of ViTs make them better suited for execution on specialized accelerators. When attack-handling routines and the main ViT algorithm are processed on different units, such as a CPU/GPU and a ViT accelerator, it leads to increased data traffic within the system. This issue helps us originate the idea behind this proposal: that attack-handling routines and ViT execution should be performed in the same location, specifically on accelerators, due to the high GFLOP requirements for real-time execution.
Funding Organization
Quick Information
Area of Research
Engineering Sciences
Focus Area
Electrical, Electronics & Computer Engineering
Start Date
27 Mar 2025
End Date
26 Mar 2028
Status
ongoing
Output
No. of Research Paper
00
Technologies (If Any)
00
No. of PhD Produced
00
Publications
00
No. of Patents
Filed : 00
Grant : 00
arrowtop
Latest Updates
Loading…