Novel Bayesian statistical methods and software tools for transcriptome-wide association study while calibrating the uncertainty of predicted expression
Implementing Organization
Indian Institute Of Technology Hyderabad
Principal Investigator
Dr. Arunabha Majumdar
Indian Institute Of Technology Hyderabad
biostat.arun@gmail.com
Project Overview
Genome-wide association studies (GWAS) have discovered numerous single-nucleotide polymorphism (SNP) loci associated with many complex human diseases and traits, e.g., cancers, cholesterol levels, intraocular pressure, etc. However, most GWAS signals reside in the non-coding regions of the genome, lacking a comprehensive understanding of the molecular mechanisms underlying the genetic associations. SNPs spotted by GWAS are more likely to be expression quantitative trait loci (eQTLs) and can influence the trait by modulating the expression level of a nearby gene. This critical insight is the primary motivation for developing the transcriptome-wide association study (TWAS). TWAS is a gene-prioritization approach for a complex trait. It aggregates the regulatory effects of multiple eQTLs and tests for association between a gene and a trait. First, it builds a prediction model for the genetically regulated component of expression (GREx) for a gene based on reference transcriptome data. Next, it predicts the GREx in separate GWAS data employing the prediction model. Then, an outcome trait is regressed on the predicted GREx in the GWAS data to test for a gene-trait association. TWAS has become popular and has been applied to a broad variety of traits. A significant criticism of the traditional TWAS is that it disregards the uncertainty of the imputed expression in the second step. Recent studies demonstrated that TWAS inference on gene-trait association can be subject to an uncontrolled rate of false positives due to the non-adjustment of the uncertainty in predicted expression. To contribute to this crucial gap in TWAS methods, we propose to develop integrative Bayesian approaches and related software tools. We will jointly model the transcriptome and GWAS data and circumvent the ad hoc two-step genre of the standard TWAS approaches. We emphasize modelling the unknown sparsity of the eQTL effects simultaneously in the Bayesian framework. A multi-ancestry approach to genomic analysis has a better potential of uncovering the genetic architecture of the trait than an ancestry-specific analysis. We propose to extend the joint Bayesian TWAS approaches for multiple ancestries. We will design joint Bayesian models of the reference transcriptome and GWAS data. We first aim to build a unified Bayesian approach to single-ancestry TWAS. We implement the horseshoe prior to model the unknown sparsity in transcriptome data. We then combine a Dirac spike and slab prior to test for an association between the gene and trait. Here, the focus is to improve the discovery of associated genes and the accuracy of effect size estimation. Next, we aim to generalize the methodology for multiple ancestries to enhance the diversity of genetic ancestry and improve the accuracy of TWAS inference. In another objective, we will build a Bayesian TWAS approach under a causal inference framework to account for the horizontal pleiotropic effects of the local SNPs of the gene on the outcome trait. We aim to distinguish the horizontal pleiotropic effect of the gene’s local SNPs on the trait from the gene’s causal effect on the trait. Next, we will extend the causal inference setup for multiple ancestries. We aim to design these approaches based on summary-level eQTL and GWAS association data without individual-level data. We aim to build efficient, user-friendly open-source software packages implementing the novel Bayesian methodologies. We will apply the methods to analyse real data for various complex traits for different populations/ancestries. We will investigate the biological significance of the discovered genes for a complex trait, concerning the genes' enrichment in relevant biological pathways, and their effects on drug response for related diseases. We will apply the methods to genomic data for Indian populations when available. Thus, we aim to develop novel Bayesian methods and software tools for TWAS, focusing on both single and multi-ancestry approaches.