Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Prediction of enzymatic pathways by integrative pathway mapping.

Sara Calhoun | Magdalena Korczynska | Daniel J Wichelecki | Brian San Francisco | Suwen Zhao | Dmitry A Rodionov | Matthew W Vetting | Nawar F Al-Obaidi | Henry Lin | Matthew J O'Meara | David A Scott | John H Morris | Daniel Russel | Steven C Almo | Andrei L Osterman | John A Gerlt | Matthew P Jacobson | Brian K Shoichet | Andrej Sali
eLife | 2018

The functions of most proteins are yet to be determined. The function of an enzyme is often defined by its interacting partners, including its substrate and product, and its role in larger metabolic networks. Here, we describe a computational method that predicts the functions of orphan enzymes by organizing them into a linear metabolic pathway. Given candidate enzyme and metabolite pathway members, this aim is achieved by finding those pathways that satisfy structural and network restraints implied by varied input information, including that from virtual screening, chemoinformatics, genomic context analysis, and ligand -binding experiments. We demonstrate this integrative pathway mapping method by predicting the L-gulonate catabolic pathway in Haemophilus influenzae Rd KW20. The prediction was subsequently validated experimentally by enzymology, crystallography, and metabolomics. Integrative pathway mapping by satisfaction of structural and network restraints is extensible to molecular networks in general and thus formally bridges the gap between structural biology and systems biology.

Pubmed ID: 29377793

Associated grants

  • Agency: NIGMS NIH HHS, United States
    Id: P01 GM118303
  • Agency: NIGMS NIH HHS, United States
    Id: P41 GM103311
  • Agency: NIGMS NIH HHS, United States
    Id: T32 GM008284
  • Agency: NCI NIH HHS, United States
    Id: P30 CA030199
  • Agency: NIDDK NIH HHS, United States
    Id: P30 DK020541
  • Agency: NIH HHS, United States
    Id: U54 GM093342
  • Agency: NIGMS NIH HHS, United States
    Id: P41-GM103311
  • Agency: NIGMS NIH HHS, United States
    Id: U54 GM093342
  • Agency: NIGMS NIH HHS, United States
    Id: P41 GM109824
  • Agency: NIH HHS, United States
    Id: S10 OD021596

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


KEGG (tool)

RRID:SCR_012773

Integrated database resource consisting of 16 main databases, broadly categorized into systems information, genomic information, and chemical information. In particular, gene catalogs in completely sequenced genomes are linked to higher-level systemic functions of cell, organism, and ecosystem. Analysis tools are also available. KEGG may be used as reference knowledge base for biological interpretation of large-scale datasets generated by sequencing and other high-throughput experimental technologies.

View all literature mentions

DOCK (tool)

RRID:SCR_000128

An algorithm used to predict and analyse binding modes of docking molecules. Users can search ligand databases for compounds that inhibit enzymatic activity and bind to particular molecules and nucleic acid targets. Molecular docking is used to predict a predominant binding mode(s) of a ligand in three-dimensional structure. This method can be used for molecular biology and computer-assisted drug design.

View all literature mentions

Integrative Modeling Platform (tool)

RRID:SCR_002982

An open source C++ and Python toolbox for solving complex modeling problems, and a number of applications for tackling some common problems in a user-friendly way. Its broad goal is to contribute to a comprehensive structural characterization of biomolecules ranging in size and complexity from small peptides to large macromolecular assemblies, by integrating data from diverse biochemical and biophysical experiments. It can also be used from the Chimera molecular modeling system, or via one of several web applications.

View all literature mentions

Cytoscape (tool)

RRID:SCR_003032

Software platform for complex network analysis and visualization. Used for visualization of molecular interaction networks and biological pathways and integrating these networks with annotations, gene expression profiles and other state data.

View all literature mentions

scikit-learn (tool)

RRID:SCR_002577

scikit-learn: machine learning in Python

View all literature mentions

Pfam (tool)

RRID:SCR_004726

A database of protein families, each represented by multiple sequence alignments and hidden Markov models (HMMs). Users can analyze protein sequences for Pfam matches, view Pfam family annotation and alignments, see groups of related families, look at the domain organization of a protein sequence, find the domains on a PDB structure, and query Pfam by keywords. There are two components to Pfam: Pfam-A and Pfam-B. Pfam-A entries are high quality, manually curated families that may automatically generate a supplement using the ADDA database. These automatically generated entries are called Pfam-B. Although of lower quality, Pfam-B families can be useful for identifying functionally conserved regions when no Pfam-A entries are found. Pfam also generates higher-level groupings of related families, known as clans (collections of Pfam-A entries which are related by similarity of sequence, structure or profile-HMM).

View all literature mentions

MODELLER (tool)

RRID:SCR_008395

Software tool as Program for Comparative Protein Structure Modelling by Satisfaction of Spatial Restraints. Used for homology or comparative modeling of protein three dimensional structures. User provides alignment of sequence to be modeled with known related structures and MODELLER automatically calculates model containing all non hydrogen atoms.

View all literature mentions

RDKit: Open-Source Cheminformatics Software (tool)

RRID:SCR_014274

An open-source cheminformatics and machine-learning toolkit that is useable from Java or Python. It includes a collection of standard cheminformatics functionality for molecule I/O, substructure searching, chemical reactions, coordinate generation (2D or 3D), fingerprinting, etc., as well as a high-performance database cartridge for working with molecules using the PostgreSQL database. Documentation is available on the main website.

View all literature mentions

OpenEye (tool)

RRID:SCR_014880

Commercial organization which provides molecular modelling and cheminformatics software to the pharmaceutical industry.

View all literature mentions