This project aims to develop a trustworthy and interpretable AI framework by integrating
advanced generative modeling with risk-controlled Uncertainty Quantification (UQ), tailored
for biomedical applications. The focus lies on addressing a fundamental limitation in
current AI systems i.e., reliable decision-making under uncertainty, especially when applied
to high-dimensional, multimodal clinical inputs such as biomedical images, textual data, and
omics. While powerful generative models like diffusion and transformer-based architectures
are increasingly used to synthesize or impute biomedical information, they often operate
as black-box systems and lack calibrated uncertainty estimates critical for high-stakes
applications such as cancer diagnosis, prognosis, and treatment planning.
To tackle these challenges, this proposal aims to leverage recent advancements
of Conformal Prediction (CP) that enables distribution-free, loss-aware uncertainty
quantification. For example, Risk-Controlling Prediction Sets (RCPS), a principled
extension of CP that incorporates domain-specific loss control while maintaining rigorous
statistical guarantees. The methodology further emphasizes the need for individualized
clinical reasoning by designing biologically-informed nonconformity scores and subgroup- or
instance-level calibration strategies, thereby going beyond the population-level guarantees
typically offered by CP. Drawing upon my prior research experience in uncertainty-aware
modeling, including Bayesian ensemble learning and feature relevance estimation using
ARD-based Gaussian process regression, the proposed approach aims to build personalized
prediction sets that are statistically valid and interpretable.
The framework will be evaluated using public biomedical datasets e.g., pan-cancer
datasets TCGA, CPTAC, that provide multimodal records, including imaging and omics
data. Generative components will be trained to infer missing omics information from visual
or contextual inputs, with a special emphasis on interpretability and robustness under
incomplete data. Explainable AI (XAI) methods will be integrated to clarify not just the
top prediction but also why other plausible alternatives appear in the model’s output set.
Importantly, this research holds direct relevance for resource-constrained clinical
environments, such as rural and semi-urban healthcare centers in India and other
1developing countries, where access to expensive molecular diagnostics is limited. By offering
cost-effective surrogate predictions with quantified reliability, the proposed system can
contribute toward democratizing precision medicine, bridging the gap between high-end AI
innovation and practical, equitable healthcare delivery.