Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Prediction and characterization of enzymatic activities guided by sequence similarity and genome neighborhood networks.

Suwen Zhao | Ayano Sakai | Xinshuai Zhang | Matthew W Vetting | Ritesh Kumar | Brandan Hillerich | Brian San Francisco | Jose Solbiati | Adam Steves | Shoshana Brown | Eyal Akiva | Alan Barber | Ronald D Seidel | Patricia C Babbitt | Steven C Almo | John A Gerlt | Matthew P Jacobson
eLife | 2014

Metabolic pathways in eubacteria and archaea often are encoded by operons and/or gene clusters (genome neighborhoods) that provide important clues for assignment of both enzyme functions and metabolic pathways. We describe a bioinformatic approach (genome neighborhood network; GNN) that enables large scale prediction of the in vitro enzymatic activities and in vivo physiological functions (metabolic pathways) of uncharacterized enzymes in protein families. We demonstrate the utility of the GNN approach by predicting in vitro activities and in vivo functions in the proline racemase superfamily (PRS; InterPro IPR008794). The predictions were verified by measuring in vitro activities for 51 proteins in 12 families in the PRS that represent ∼85% of the sequences; in vitro activities of pathway enzymes, carbon/nitrogen source phenotypes, and/or transcriptomic studies confirmed the predicted pathways. The synergistic use of sequence similarity networks3 and GNNs will facilitate the discovery of the components of novel, uncharacterized metabolic pathways in sequenced genomes.

Pubmed ID: 24980702

Research resources used in this publication

None found

Antibodies used in this publication

None found

Associated grants

  • Agency: NIGMS NIH HHS, United States
    Id: U54GM093342
  • Agency: NIGMS NIH HHS, United States
    Id: P41 GM103311
  • Agency: NIGMS NIH HHS, United States
    Id: P01 GM071790
  • Agency: NIGMS NIH HHS, United States
    Id: U54 GM074945
  • Agency: NIGMS NIH HHS, United States
    Id: U54GM094662
  • Agency: NIGMS NIH HHS, United States
    Id: U54GM074945
  • Agency: NIGMS NIH HHS, United States
    Id: U54 GM094662
  • Agency: NIGMS NIH HHS, United States
    Id: P01GM071790
  • Agency: NCI NIH HHS, United States
    Id: P30 CA013330
  • Agency: NIGMS NIH HHS, United States
    Id: U54 GM093342
  • Agency: NIGMS NIH HHS, United States
    Id: P41-GM103311

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


FundRef (tool)

RRID:SCR_003218

Funder Registry and associated funding metadata allows everyone to have transparency into research funding and its outcomes. Open registry of persistent identifiers for grant-giving organizations around the world.

View all literature mentions

InterPro (tool)

RRID:SCR_006695

Service providing functional analysis of proteins by classifying them into families and predicting domains and important sites. They combine protein signatures from a number of member databases into a single searchable resource, capitalizing on their individual strengths to produce a powerful integrated database and diagnostic tool. This integrated database of predictive protein signatures is used for the classification and automatic annotation of proteins and genomes. InterPro classifies sequences at superfamily, family and subfamily levels, predicting the occurrence of functional domains, repeats and important sites. InterPro adds in-depth annotation, including GO terms, to the protein signatures. You can access the data programmatically, via Web Services. The member databases use a number of approaches: # ProDom: provider of sequence-clusters built from UniProtKB using PSI-BLAST. # PROSITE patterns: provider of simple regular expressions. # PROSITE and HAMAP profiles: provide sequence matrices. # PRINTS provider of fingerprints, which are groups of aligned, un-weighted Position Specific Sequence Matrices (PSSMs). # PANTHER, PIRSF, Pfam, SMART, TIGRFAMs, Gene3D and SUPERFAMILY: are providers of hidden Markov models (HMMs). Your contributions are welcome. You are encouraged to use the ''''Add your annotation'''' button on InterPro entry pages to suggest updated or improved annotation for individual InterPro entries.

View all literature mentions

Research Collaboratory for Structural Bioinformatics Protein Data Bank (RCSB PDB) (tool)

RRID:SCR_012820

Collection of structural data of biological macromolecules. Database of information about 3D structures of large biological molecules, including proteins and nucleic acids. Users can perform queries on data and analyze and visualize results.

View all literature mentions

UniProt (tool)

RRID:SCR_002380

Collection of data of protein sequence and functional information. Resource for protein sequence and annotation data. Consortium for preservation of the UniProt databases: UniProt Knowledgebase (UniProtKB), UniProt Reference Clusters (UniRef), and UniProt Archive (UniParc), UniProt Proteomes. Collaboration between European Bioinformatics Institute (EMBL-EBI), SIB Swiss Institute of Bioinformatics and Protein Information Resource. Swiss-Prot is a curated subset of UniProtKB.

View all literature mentions