Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
Horizontal gene transfer (HGT) is an essential force in microbial evolution. Despite detailed studies on a variety of systems, a global picture of HGT in the microbial world is still missing. Here, we exploit that HGT creates long identical DNA sequences in the genomes of distant species, which can be found efficiently using alignment-free methods. Our pairwise analysis of 93,481 bacterial genomes identified 138,273 HGT events. We developed a model to explain their statistical properties as well as estimate the transfer rate between pairs of taxa. This reveals that long-distance HGT is frequent: our results indicate that HGT between species from different phyla has occurred in at least 8% of the species. Finally, our results confirm that the function of sequences strongly impacts their transfer rate, which varies by more than three orders of magnitude between different functional categories. Overall, we provide a comprehensive view of HGT, illuminating a fundamental process driving bacterial evolution.
Pubmed ID: 34121661
Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.
Collection of curated, non-redundant genomic DNA, transcript RNA, and protein sequences produced by NCBI. Provides a reference for genome annotation, gene identification and characterization, mutation and polymorphism analysis, expression studies, and comparative analyses. Accessed through the Nucleotide and Protein databases.
View all literature mentionsA Bioinformatics Resource Center bacterial bioinformatics database and analysis resource that provides researchers with an online resource that stores and integrates a variety of data types (e.g. genomics, transcriptomics, protein-protein interactions (PPIs), three-dimensional protein structures and sequence typing data) and associated metadata. Datatypes are summarized for individual genomes and across taxonomic levels. All genomes, currently more than 10 000, are consistently annotated using RAST, the Rapid Annotations using Subsystems Technology. Summaries of different data types are also provided for individual genes, where comparisons of different annotations are available, and also include available transcriptomic data. PATRIC provides a variety of ways for researchers to find data of interest and a private workspace where they can store both genomic and gene associations, and their own private data. Both private and public data can be analyzed together using a suite of tools to perform comparative genomic or transcriptomic analysis. PATRIC also includes integrated information related to disease and PPIs. The PATRIC project includes three primary collaborators: the University of Chicago, the University of Manchester, and New City Media. The University of Chicago is providing genome annotations and a PATRIC end-user genome annotation service using their Rapid Annotation using Subsystem Technology (RAST) system. The National Centre for Text Mining (NaCTeM) at the University of Manchester is providing literature-based text mining capability and service. New City Media is providing assistance in website interface development. An FTP server and download tool are available.
View all literature mentionsA repository for bacterial small regulatory RNA. They welcome you to submit new experimental validated sRNA targets.
View all literature mentionsAn automated analysis platform for metagenomes providing quantitative insights into microbial populations based on sequence data. The server primarily provides upload, quality control, automated annotation and analysis for prokaryotic metagenomic shotgun samples.
View all literature mentionsSoftware package for functional analysis of sequences by classifying them into families and predicting presence of domains and sites. Scans sequences against InterPro's signatures. Characterizes nucleotide or protein function by matching it with models from several different databases. Used in large scale analysis of whole proteomes, genomes and metagenomes. Available as Web based version and standalone Perl version and SOAP Web Service.
View all literature mentionsAn information resource for peptidases (also termed proteases, proteinases and proteolytic enzymes) and the proteins that inhibit them. The MEROPS database uses an hierarchical, structure-based classification of the peptidases. In this, each peptidase is assigned to a Family on the basis of statistically significant similarities in amino acid sequence, and families that are thought to be homologous are grouped together in a Clan. There is a Summary page for each family and clan, and these have indexes. Each of the Summary pages offers links to supplementary pages. About 3000 individual peptidases and inhibitors are included in the database, and there is a Summary page describing each one. You can navigate to this by any of several routes. There are indexes of Name, MEROPS Identifier and source Organism on the menu bar. Each Summary page describes the classification and nomenclature of the peptidase or inhibitor, and provides links to supplementary pages showing sequence identifiers, the structure if known, literature references and more.
View all literature mentionsDatabase of information about restriction enzymes and related proteins containing published and unpublished references, recognition and cleavage sites, isoschizomers, commercial availability, methylation sensitivity, crystal, genome, and sequence data. DNA methyltransferases, homing endonucleases, nicking enzymes, specificity subunits and control proteins are also included. Several tools are available including REBsites, BLAST against REBASE, NEBcutter and REBpredictor. Putative DNA methyltransferases and restriction enzymes, as predicted from analysis of genomic sequences, are also listed. REBASE is updated daily and is constantly expanding. Users may submit new enzyme and/or sequence information, recommend references, or send them corrections to existing data. The contents of REBASE may be browsed from the web and selected compilations can be downloaded by ftp (ftp.neb.com). Additionally, monthly updates can be requested via email.
View all literature mentionsSoftware tool for protein coding gene prediction for prokaryotic genomes.
View all literature mentions