Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
Genomes are an integral component of the biological information about an organism; thus, the more complete the genome, the more informative it is. Historically, bacterial and archaeal genomes were reconstructed from pure (monoclonal) cultures, and the first reported sequences were manually curated to completion. However, the bottleneck imposed by the requirement for isolates precluded genomic insights for the vast majority of microbial life. Shotgun sequencing of microbial communities, referred to initially as community genomics and subsequently as genome-resolved metagenomics, can circumvent this limitation by obtaining metagenome-assembled genomes (MAGs); but gaps, local assembly errors, chimeras, and contamination by fragments from other genomes limit the value of these genomes. Here, we discuss genome curation to improve and, in some cases, achieve complete (circularized, no gaps) MAGs (CMAGs). To date, few CMAGs have been generated, although notably some are from very complex systems such as soil and sediment. Through analysis of about 7000 published complete bacterial isolate genomes, we verify the value of cumulative GC skew in combination with other metrics to establish bacterial genome sequence accuracy. The analysis of cumulative GC skew identified potential misassemblies in some reference genomes of isolated bacteria and the repeat sequences that likely gave rise to them. We discuss methods that could be implemented in bioinformatic approaches for curation to ensure that metabolic and evolutionary analyses can be based on very high-quality genomes.
Pubmed ID: 32188701
Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.
Tool for searching sequence databases for homologs of protein sequences, and for making protein sequence alignments. It implements methods using probabilistic models called profile hidden Markov models (profile HMMs). Compared to BLAST, FASTA, and other sequence alignment and database search tools based on older scoring methodology, HMMER aims to be significantly more accurate and more able to detect remote homologs because of the strength of its underlying mathematical models. In the past, this strength came at significant computational expense, but in the new HMMER3 project, HMMER is now essentially as fast as BLAST.
View all literature mentionsSoftware ultrafast memory efficient tool for aligning sequencing reads. Bowtie is short read aligner.
View all literature mentionsSoftware tool for windowed adaptive trimming for fastq files using quality. Supports quality values like Illumina, Solexa, and Sanger. Takes the quality values and slides a window across them whose length is 0.1 times the length of the read.
View all literature mentionsWeb server to search for tRNA genes in genomic sequence. If you would like to run tRNAscan-SE locally, you can get the UNIX source code (gzip''d tar file).
View all literature mentionsSoftware for an iterative De Bruijn Graph De Novo Assembler for Short Reads Sequencing data with Highly Uneven Sequencing Depth.
View all literature mentionsSoftware tool for protein coding gene prediction for prokaryotic genomes.
View all literature mentionsSoftware statistical framework for reconstructing genomes from metagenome data. Open source software tool for accurately reconstructing single genomes from complex microbial communities.
View all literature mentionsWeb server to search for tRNA genes in genomic sequence. If you would like to run tRNAscan-SE locally, you can get the UNIX source code (gzip''d tar file).
View all literature mentions