Whole genome amplification by the multiple displacement amplification (MDA) method allows sequencing of DNA from single cells of bacteria that cannot be cultured. Assembling a genome is challenging, however, because MDA generates highly nonuniform coverage of the genome. Here we describe an algorithm tailored for short-read data from single cells that improves assembly through the use of a progressively increasing coverage cutoff. Assembly of reads from single Escherichia coli and Staphylococcus aureus cells captures >91% of genes within contigs, approaching the 95% captured from an assembly based on many E. coli cells. We apply this method to assemble a genome from a single cell of an uncultivated SAR324 clade of Deltaproteobacteria, a cosmopolitan bacterial lineage in the global ocean. Metabolic reconstruction suggests that SAR324 is aerobic, motile and chemotaxic. Our approach enables acquisition of genome assemblies for individual uncultivated bacteria using only short reads, providing cell-specific genetic information absent from metagenomic studies.
Pubmed ID: 21926975 RIS Download
Mesh terms: Algorithms | Bacteria | Base Sequence | Contig Mapping | Databases, Nucleic Acid | Deltaproteobacteria | Escherichia coli | Genome, Bacterial | Likelihood Functions | Sequence Analysis, DNA | Single-Cell Analysis | Staphylococcus aureus
Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.
Database that contains updated information about the Escherichia coli K-12 genome and proteome sequences, including extensive gene bibliographies. Users are able to download customized tables, perform Boolean query comparisons, generate sets of paired DNA sequences, and download any E. coli K-12 genomic DNA sub-sequence. BLAST functions, microarray data, an alphabetical index of genes, and gene overlap queries are also available. The Database Table Downloads Page provides a full list of EG numbers cross-referenced to the new cross-database ECK numbers and other common accession numbers, as well as gene names and synonyms. Monthly release archival downloads are available, but the live, daily updated version of EcoGene is the default mysql database for download queries.
toolView all literature mentions