Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Comparison of methods for transcriptome imputation through application to two common complex diseases.

James J Fryett | Jamie Inshaw | Andrew P Morris | Heather J Cordell
European journal of human genetics : EJHG | 2018

Transcriptome imputation has become a popular method for integrating genotype data with publicly available expression data to investigate the potentially causal role of genes in complex traits. Here, we compare three approaches (PrediXcan, MetaXcan and FUSION) via application to genome-wide association study (GWAS) data for Crohn's disease and type 1 diabetes from the Wellcome Trust Case Control Consortium. We investigate: (i) how the results of each approach compare with each other and with those of standard GWAS analysis; and (ii) how variants in the models used by the prediction tools compare with variants previously reported as eQTLs. We find that all approaches produce highly correlated results when applied to the same GWAS data, although for a subset of genes, mostly in the major histocompatibility complex, the approaches strongly disagree. We also observe that most associations detected by these methods occur near known GWAS risk loci. PrediXcan and MetaXcan's models for predicting expression more consistently recapitulate known effects of genotype on expression, suggesting they are more robust than FUSION. Application of these transcriptome imputation approaches to summary statistics from meta-analyses in Crohn's disease and type 1 diabetes detects 53 significant expression-Crohn's disease associations and 154 significant expression-type 1 diabetes associations, providing insight into biology underlying these diseases. We conclude that while current implementations of transcriptome imputation typically detect fewer associations than GWAS, they nonetheless provide an interesting way of interpreting association signals to identify potentially causal genes, and that PrediXcan and MetaXcan generally produce more reliable results than FUSION.

Pubmed ID: 29976976

Research resources used in this publication

None found

Additional research tools detected in this publication

Antibodies used in this publication

None found

Associated grants

  • Agency: Wellcome Trust, United Kingdom
    Id: 098017
  • Agency: Wellcome Trust, United Kingdom
    Id: 102858/Z/13/Z
  • Agency: Wellcome Trust, United Kingdom
    Id: 076113
  • Agency: Wellcome Trust, United Kingdom
    Id: 102858
  • Agency: Wellcome Trust, United Kingdom
    Id: 107212/Z/15/Z
  • Agency: Biotechnology and Biological Sciences Research Council, United Kingdom
    Id: BB/M011186/1
  • Agency: Wellcome Trust, United Kingdom

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


1000 Genomes Project and AWS (tool)

RRID:SCR_008801

A dataset containing the full genomic sequence of 1,700 individuals, freely available for research use. The 1000 Genomes Project is an international research effort coordinated by a consortium of 75 companies and organizations to establish the most detailed catalogue of human genetic variation. The project has grown to 200 terabytes of genomic data including DNA sequenced from more than 1,700 individuals that researchers can now access on AWS for use in disease research free of charge. The dataset containing the full genomic sequence of 1,700 individuals is now available to all via Amazon S3. The data can be found at: http://s3.amazonaws.com/1000genomes The 1000 Genomes Project aims to include the genomes of more than 2,662 individuals from 26 populations around the world, and the NIH will continue to add the remaining genome samples to the data collection this year. Public Data Sets on AWS provide a centralized repository of public data hosted on Amazon Simple Storage Service (Amazon S3). The data can be seamlessly accessed from AWS services such Amazon Elastic Compute Cloud (Amazon EC2) and Amazon Elastic MapReduce (Amazon EMR), which provide organizations with the highly scalable compute resources needed to take advantage of these large data collections. AWS is storing the public data sets at no charge to the community. Researchers pay only for the additional AWS resources they need for further processing or analysis of the data. All 200 TB of the latest 1000 Genomes Project data is available in a publicly available Amazon S3 bucket. You can access the data via simple HTTP requests, or take advantage of the AWS SDKs in languages such as Ruby, Java, Python, .NET and PHP. Researchers can use the Amazon EC2 utility computing service to dive into this data without the usual capital investment required to work with data at this scale. AWS also provides a number of orchestration and automation services to help teams make their research available to others to remix and reuse. Making the data available via a bucket in Amazon S3 also means that customers can crunch the information using Hadoop via Amazon Elastic MapReduce, and take advantage of the growing collection of tools for running bioinformatics job flows, such as CloudBurst and Crossbow.

View all literature mentions

PEER (tool)

RRID:SCR_009326

Software collection of Bayesian approaches to infer hidden determinants and their effects from gene expression profiles using factor analysis methods. Applications of PEER have * detected batch effects and experimental confounders * increased the number of expression QTL findings by threefold * allowed inference of intermediate cellular traits, such as transcription factor or pathway activations This project offers an efficient and versatile C++ implementation of the underlying algorithms with user-friendly interfaces to R and python.

View all literature mentions