×

img Accessibility Controls

Research Projects Banner

Research Projects

AI-Driven Integrative Prediction of Protein Interaction Sites Using Language Models and Structural Data

Implementing Organization

Principal Investigator
Dr. Nagamma patil
National Institute Of Technology Karnataka, Surathkal
nagammapatil@nitk.edu.in

Project Overview

Protein–protein interaction (PPI) sites are central to numerous cellular functions and serve as vital targets for therapeutic development. Accurate prediction of these sites is crucial for advancing drug discovery, protein design, and functional annotation. Traditional computational approaches heavily depend on evolutionary profiles derived from multiple sequence alignments (MSAs) and a limited set of experimentally resolved protein structures. These methods, however, falter in scenarios where homologous sequences or high-quality structural data are absent, particularly for novel, uncharacterized, or synthetic proteins. To address these limitations, the proposed research introduces a novel, MSA-independent framework that leverages state-of-the-art protein language models (PLMs) and predicted 3D structures to enhance PPI site prediction. PLMs trained on large-scale sequence data provide residue-level embeddings that inherently capture evolutionary and functional signals without requiring explicit homologs. This work systematically integrates these embeddings with detailed structural descriptors extracted from both experimentally resolved structures (PDB) and high-confidence AlphaFold-predicted models, which help mitigate missing residue information common in PDB files. The research will test the hypothesis that deep contextual representations from PLMs, when combined with fine-grained structural features or structure-aware PLMs (e.g., SaProt embeddings using 3Di tokens), can significantly improve interaction site prediction accuracy. A key objective is to evaluate whether implicit structural encodings within PLMs can match or outperform explicitly computed structure-based features. To this end, the project will implement and benchmark deep learning models trained using: (i) PLM embeddings alone, (ii) PLM + PDB structural features, (iii) PLM + AlphaFold structural features, and (iv) structure-aware PLM embeddings. The core experiments involve computing and integrating structural features, generating embeddings using PLMs and structure-aware models, training graph-based deep learning models, and conducting both quantitative and qualitative evaluations, including ablation studies, to assess model effectiveness. Ultimately, we aim to develop a continuously accessible online platform that enables annotation of protein–protein interaction sites from user-submitted amino acid sequences, thereby facilitating the creation of a high-quality, customized database of proteins with detailed interaction site information. If successful, the proposed framework will establish a generalizable and scalable approach for PPI site prediction that remains robust even in the absence of evolutionary or experimental structural information. The scientific significance lies in its ability to improve fundamental understanding of protein interaction mechanisms while offering practical utility for high-throughput protein annotation, rational drug design, and synthetic biology applications. The outcomes are poised to contribute meaningfully to the field by addressing a longstanding bottleneck in computational structural biology and by enhancing predictive capabilities in real-world, post-genomic settings.
Funding Organization
Quick Information
Area of Research
Engineering Sciences
Focus Area
Computer Science And Engineering
Start Date
26 Mar 2026
End Date
25 Sep 2029
Status
ongoing
Output
No. of Research Paper
00
Technologies (If Any)
00
No. of PhD Produced
00
Publications
00
No. of Patents
Filed : 00
Grant : 00
arrowtop
Latest Updates
Loading…