Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Ultrafast search of all deposited bacterial and viral genomic data.

Phelim Bradley | Henk C den Bakker | Eduardo P C Rocha | Gil McVean | Zamin Iqbal
Nature biotechnology | 2019

Exponentially increasing amounts of unprocessed bacterial and viral genomic sequence data are stored in the global archives. The ability to query these data for sequence search terms would facilitate both basic research and applications such as real-time genomic epidemiology and surveillance. However, this is not possible with current methods. To solve this problem, we combine knowledge of microbial population genomics with computational methods devised for web search to produce a searchable data structure named BItsliced Genomic Signature Index (BIGSI). We indexed the entire global corpus of 447,833 bacterial and viral whole-genome sequence datasets using four orders of magnitude less storage than previous methods. We applied our BIGSI search function to rapidly find resistance genes MCR-1, MCR-2, and MCR-3, determine the host-range of 2,827 plasmids, and quantify antibiotic resistance in archived datasets. Our index can grow incrementally as new (unprocessed or assembled) sequence datasets are deposited and can scale to millions of datasets.

Pubmed ID: 30718882

Research resources used in this publication

None found

Additional research tools detected in this publication

Antibodies used in this publication

None found

Associated grants

  • Agency: Wellcome Trust, United Kingdom

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


Kraken (tool)

RRID:SCR_005484

A set of software tools ( Reaper, Tally and Sequence Imp) designed to streamline the analysis of next-generation sequencing data. Although designed with small RNA sequence analysis in mind the tools can be used to address issues facing next-generation sequencing in general.

View all literature mentions

Plasmid Genome Database (tool)

RRID:SCR_008228

The Plasmid Genome Database aims to collate biological and genomic data for all bacterial plasmids in the hopes of enabling rapid, interrogation of both meta- and genomic data. Data maintained includes access to all plasmid genomes and information on core genomic features obtained from parsing the original EMBL/DDBJ/NCBI submission. In addition a suite of third party analyses has been performed for each genome to supplement the original annotation. This site also links to Genome Atlases provided by the Centre for Biological Sequence Analysis (CBS). The motivation behind the construction of this site derived from observations from genome sequencing projects: the abundance and inferred importance of the horizontal gene pool (HGP) in bacterial adaptation and evolution. In so far as plasmids are autonomously replicating, extrachromosomal elements they are a readily identifiable and accessible component of the HGP. Also plasmids have been identified in almost all bacterial divisions, ranging in size from less than 2 kbp to > 1.5 Mbp and as such represent a defined, yet diverse and complex sample of genes in the HGP.

View all literature mentions