Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Empirical design of a variant quality control pipeline for whole genome sequencing data using replicate discordance.

Robert P Adelson | Alan E Renton | Wentian Li | Nir Barzilai | Gil Atzmon | Alison M Goate | Peter Davies | Yun Freudenberg-Hua
Scientific reports | 2019

The success of next-generation sequencing depends on the accuracy of variant calls. Few objective protocols exist for QC following variant calling from whole genome sequencing (WGS) data. After applying QC filtering based on Genome Analysis Tool Kit (GATK) best practices, we used genotype discordance of eight samples that were sequenced twice each to evaluate the proportion of potentially inaccurate variant calls. We designed a QC pipeline involving hard filters to improve replicate genotype concordance, which indicates improved accuracy of genotype calls. Our pipeline analyzes the efficacy of each filtering step. We initially applied this strategy to well-characterized variants from the ClinVar database, and subsequently to the full WGS dataset. The genome-wide biallelic pipeline removed 82.11% of discordant and 14.89% of concordant genotypes, and improved the concordance rate from 98.53% to 99.69%. The variant-level read depth filter most improved the genome-wide biallelic concordance rate. We also adapted this pipeline for triallelic sites, given the increasing proportion of multiallelic sites as sample sizes increase. For triallelic sites containing only SNVs, the concordance rate improved from 97.68% to 99.80%. Our QC pipeline removes many potentially false positive calls that pass in GATK, and may inform future WGS studies prior to variant effect analysis.

Pubmed ID: 31695094

Research resources used in this publication

None found

Antibodies used in this publication

None found

Associated grants

  • Agency: NIA NIH HHS, United States
    Id: R01 AG057909
  • Agency: U.S. Department of Health & Human Services | NIH | National Institute on Aging (U.S. National Institute on Aging), International
    Id: R01 AG 042188
  • Agency: U.S. Department of Health & Human Services | NIH | National Institute on Aging (U.S. National Institute on Aging), International
    Id: K08AG054727
  • Agency: U.S. Department of Health & Human Services | NIH | National Institute on Aging (U.S. National Institute on Aging), International
    Id: R01 AG 046949
  • Agency: U.S. Department of Health & Human Services | NIH | National Institute on Aging (U.S. National Institute on Aging), International
    Id: R01 AG 618381
  • Agency: U.S. Department of Health & Human Services | NIH | National Institute on Aging (U.S. National Institute on Aging), International
    Id: P01 AG 021654
  • Agency: NIA NIH HHS, United States
    Id: K08 AG054727
  • Agency: NIA NIH HHS, United States
    Id: P30 AG038072

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


GATK (tool)

RRID:SCR_001876

A software package to analyze next-generation resequencing data. The toolkit offers a wide variety of tools, with a primary focus on variant discovery and genotyping as well as strong emphasis on data quality assurance. Its robust architecture, powerful processing engine and high-performance computing features make it capable of taking on projects of any size. This software library makes writing efficient analysis tools using next-generation sequencing data very easy, and second it's a suite of tools for working with human medical resequencing projects such as 1000 Genomes and The Cancer Genome Atlas. These tools include things like a depth of coverage analyzers, a quality score recalibrator, a SNP/indel caller and a local realigner. (entry from Genetic Analysis Software)

View all literature mentions

dbSNP (tool)

RRID:SCR_002338

Database as central repository for both single base nucleotide substitutions and short deletion and insertion polymorphisms. Distinguishes report of how to assay SNP from use of that SNP with individuals and populations. This separation simplifies some issues of data representation. However, these initial reports describing how to assay SNP will often be accompanied by SNP experiments measuring allele occurrence in individuals and populations. Community can contribute to this resource.

View all literature mentions

ClinVar (tool)

RRID:SCR_006169

Archive of aggregated information about sequence variation and its relationship to human health. Provides reports of relationships among human variations and phenotypes along with supporting evidence. Submissions from clinical testing labs, research labs, locus-specific databases, expert panels and professional societies are welcome. Collects reports of variants found in patient samples, assertions made regarding their clinical significance, information about submitter, and other supporting data. Alleles described in submissions are mapped to reference sequences, and reported according to HGVS standard.

View all literature mentions

seqMINER (tool)

RRID:SCR_013020

Software for a genome wide mapping data interpretation platform for NGS (ChIPSeq).

View all literature mentions