Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Accurate and complete genomes from metagenomes.

Lin-Xing Chen | Karthik Anantharaman | Alon Shaiber | A Murat Eren | Jillian F Banfield
Genome research | 2020

Genomes are an integral component of the biological information about an organism; thus, the more complete the genome, the more informative it is. Historically, bacterial and archaeal genomes were reconstructed from pure (monoclonal) cultures, and the first reported sequences were manually curated to completion. However, the bottleneck imposed by the requirement for isolates precluded genomic insights for the vast majority of microbial life. Shotgun sequencing of microbial communities, referred to initially as community genomics and subsequently as genome-resolved metagenomics, can circumvent this limitation by obtaining metagenome-assembled genomes (MAGs); but gaps, local assembly errors, chimeras, and contamination by fragments from other genomes limit the value of these genomes. Here, we discuss genome curation to improve and, in some cases, achieve complete (circularized, no gaps) MAGs (CMAGs). To date, few CMAGs have been generated, although notably some are from very complex systems such as soil and sediment. Through analysis of about 7000 published complete bacterial isolate genomes, we verify the value of cumulative GC skew in combination with other metrics to establish bacterial genome sequence accuracy. The analysis of cumulative GC skew identified potential misassemblies in some reference genomes of isolated bacteria and the repeat sequences that likely gave rise to them. We discuss methods that could be implemented in bioinformatic approaches for curation to ensure that metabolic and evolutionary analyses can be based on very high-quality genomes.

Pubmed ID: 32188701

Research resources used in this publication

None found

Antibodies used in this publication

None found

Associated grants

  • Agency: NIAID NIH HHS, United States
    Id: R01 AI092531
  • Agency: NIGMS NIH HHS, United States
    Id: R01 GM109454

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


Hmmer (tool)

RRID:SCR_005305

Tool for searching sequence databases for homologs of protein sequences, and for making protein sequence alignments. It implements methods using probabilistic models called profile hidden Markov models (profile HMMs). Compared to BLAST, FASTA, and other sequence alignment and database search tools based on older scoring methodology, HMMER aims to be significantly more accurate and more able to detect remote homologs because of the strength of its underlying mathematical models. In the past, this strength came at significant computational expense, but in the new HMMER3 project, HMMER is now essentially as fast as BLAST.

View all literature mentions

Bowtie (tool)

RRID:SCR_005476

Software ultrafast memory efficient tool for aligning sequencing reads. Bowtie is short read aligner.

View all literature mentions

Sickle (tool)

RRID:SCR_006800

Software tool for windowed adaptive trimming for fastq files using quality. Supports quality values like Illumina, Solexa, and Sanger. Takes the quality values and slides a window across them whose length is 0.1 times the length of the read.

View all literature mentions

tRNAscan-SE (tool)

RRID:SCR_010835

Web server to search for tRNA genes in genomic sequence. If you would like to run tRNAscan-SE locally, you can get the UNIX source code (gzip''d tar file).

View all literature mentions

IDBA-UD (tool)

RRID:SCR_011912

Software for an iterative De Bruijn Graph De Novo Assembler for Short Reads Sequencing data with Highly Uneven Sequencing Depth.

View all literature mentions

Prodigal (tool)

RRID:SCR_011936

Software tool for protein coding gene prediction for prokaryotic genomes.

View all literature mentions

MetaBAT (tool)

RRID:SCR_019134

Software statistical framework for reconstructing genomes from metagenome data. Open source software tool for accurately reconstructing single genomes from complex microbial communities.

View all literature mentions

tRNAscan-SE (tool)

RRID:SCR_008637

Web server to search for tRNA genes in genomic sequence. If you would like to run tRNAscan-SE locally, you can get the UNIX source code (gzip''d tar file).

View all literature mentions