We support boolean queries, use +,-,<,>,~,* to alter the weighting of terms
The Plasmid Genome Database aims to collate biological and genomic data for all bacterial plasmids in the hopes of enabling rapid, interrogation of both meta- and genomic data. Data maintained includes access to all plasmid genomes and information on core genomic features obtained from parsing the original EMBL/DDBJ/NCBI submission. In addition a suite of third party analyses has been performed for each genome to supplement the original annotation. This site also links to Genome Atlases provided by the Centre for Biological Sequence Analysis (CBS). The motivation behind the construction of this site derived from observations from genome sequencing projects: the abundance and inferred importance of the horizontal gene pool (HGP) in bacterial adaptation and evolution. In so far as plasmids are autonomously replicating, extrachromosomal elements they are a readily identifiable and accessible component of the HGP. Also plasmids have been identified in almost all bacterial divisions, ranging in size from less than 2 kbp to > 1.5 Mbp and as such represent a defined, yet diverse and complex sample of genes in the HGP.
Protein Data Bank (PDB) contains data on the spatial protein structures and their biologically active sites (i.e., ligand binding regions, enzyme catalytic centers, regions subjected to biochemical modifications, etc.). However, neither of the well known systems searching PDB does not provide the user with possibility to make the queries related with the active sites. A database PDBSITE storing the data on biologically active sites contained in the PDB database has been developed. PDBSITE accumulates amino acid content, structure features calculated by spatial protein structures, and physicochemical properties of sites and their spatial surroundings.
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 23, 2016. PDBfun is a web server for structural and functional analysis of proteins at the residue level. pdbFun gives fast access to the whole Protein Data Bank (PDB) organized as a database of annotated residues. The available data (features) range from solvent exposure to ligand binding ability, location in a protein cavity, secondary structure, residue type, sequence functional pattern, protein domain and catalytic activity. PDBfun is an integrated web tool for querying the PDB at the residue level and for local structural comparison. It integrates knowledge on single residues in protein structures coming from other databases or calculated with available or in-house developed instruments for structural analysis. Each set of different annotations represents a feature. Features are listed in PDBfun main page in orange. Features can be used for building residues selections.
PRAGMA was formed to establish sustained collaborations and advance the use of grid technologies in applications among a community of investigators working with leading institutions around the Pacific Rim. Currently there are 35 institutions in PRAGMA, who meet twice a year at PRAGMA Workshops. In PRAGMA, applications are the key, integrating focus that bring together the necessary infrastructure and middleware to advance the applications goals. Sponsors: This resource is made possible by a grant from National Science Foundation Grant No. INT-0314015 and OCI-0627026.
A database of binding affinities for the protein-ligand complexes in the Protein Data Bank (PDB). The PDBbind database is a collection of the experimentally measured binding affinities exclusively for the protein-ligand complexes available in the Protein Data Bank (PDB). It thus provides a link between energetic and structural information of those complexes and may be of great value to various molecular recognition studies. This site was last updated in 2007. The updated version of the resource is maintained by the Shanghai Institute of Organic Chemistry (http://www.pdbbind.org.cn).
THIS RESOURCE IS NO LONGER IN SERVICE, documented on August 20,2019.The COG-database has become a powerful tool in the field of comparative genomics. The construction of this data-base is based on sequence homologies of proteins from different completely sequenced genomes. Highly homologous proteins are assigned to clusters of orthologous groups. The updated collection of orthologous protein sets for prokaryotes and eukaryotes is expected to be a useful platform for functional annotation of newly sequenced genomes, including those of complex eukaryotes, and genome-wide evolutionary studies. The availability of multiple, essentially complete genome sequences of prokaryotes and eukaryotes spurred both the demand and the opportunity for the construction of an evolutionary classification of genes from these genomes. Such a classification system based on orthologous relationships between genes appears to be a natural framework for comparative genomics and should facilitate both functional annotation of genomes and large-scale evolutionary studies. Here is a major update of the previously developed system for delineation of Clusters of Orthologous Groups of proteins (COGs) from the sequenced genomes of prokaryotes and unicellular eukaryotes and the construction of clusters of predicted orthologs for 7 eukaryotic genomes, which we named KOGs after eukaryotic orthologous groups. The COG collection currently consists of 138,458 proteins, which form 4873 COGs and comprise 75% of the 185,505 (predicted) proteins encoded in 66 genomes of unicellular organisms. The eukaryotic orthologous groups (KOGs) include proteins from 7 eukaryotic genomes: three animals (the nematode Caenorhabditis elegans, the fruit fly Drosophila melanogaster and Homo sapiens), one plant, Arabidopsis thaliana, two fungi (Saccharomyces cerevisiae and Schizosaccharomyces pombe), and the intracellular microsporidian parasite Encephalitozoon cuniculi. The current KOG set consists of 4852 clusters of orthologs, which include 59,838 proteins, or approximately 54% of the analyzed eukaryotic 110,655 gene products. Compared to the coverage of the prokaryotic genomes with COGs, a considerably smaller fraction of eukaryotic genes could be included into the KOGs; addition of new eukaryotic genomes is expected to result in substantial increase in the coverage of eukaryotic genomes with KOGs. Examination of the phyletic patterns of KOGs reveals a conserved core represented in all analyzed species and consisting of approximately 20% of the KOG set. This conserved portion of the KOG set is much greater than the ubiquitous portion of the COG set (approximately 1% of the COGs). In part, this difference is probably due to the small number of included eukaryotic genomes, but it could also reflect the relative compactness of eukaryotes as a clade and the greater evolutionary stability of eukaryotic genomes.
THIS RESOURCE IS NO LONGER IN SERVICE, documented on July 15, 2013. This is the official database of the environmental chlamydia genome project. This resource provides access to finished sequence for Parachlamydia-related symbiont UWE25 and to a wide range of manual annotations, automatical analyses and derived datasets. Functional classification and description has been manually annotated according to the Annotation guidelines. Chlamydiae are the major cause of preventable blindness and sexually transmitted disease. Genome analysis of a chlamydia-related symbiont of free-living amoebae revealed that it is twice as large as any of the pathogenic chlamydiae and had few signs of recent lateral gene acquisition. We showed that about 700 million years ago the last common ancestor of pathogenic and symbiotic chlamydiae was already adapted to intracellular survival in early eukaryotes and contained many virulence factors found in modern pathogenic chlamydiae, including a type III secretion system. Ancient chlamydiae appear to be the originators of mechanisms for the exploitation of eukaryotic cells. Environmental chlamydiae have recently been recognized as obligate endosymbionts of free-living amoebae and have been implicated as potential human pathogens. Environmental chlamydiae form a deep branching evolutionary lineage within the medically important order Chlamydiales. Despite their high diversity and ubiquitous distribution in clinical and environmental samples only limited information about genetics and ecology of these microorganisms is available. The Parachlamydia-related Acanthamoeba symbiont UWE25 was therefore selected as representative environmental chlamydia strain for whole genome sequencing. Comparative genome analysis was performed using PEDANT and simap. Sponsors: The environmental chlamydia genome project was funded by the bmb+f (German Federal Ministry of Education and Research) and is part of the Competence Network PathoGenoMiK.
THIS RESOURCE IS NO LONGER IN SERVICE.Documented on September 23,2022. OrthoDisease is a comprehensive database of model organism genes that are orthologous to human disease genes. It was constructed by applying the Inparanoid ortholog detection algorithm to disease genes derived from the Online Mendelian Inheritance in Man database (OMIM). Over 1000 human genes have been associated with a disease phenotype to date and this number continues to rise on a daily basis. Studying their orthologous counterparts in model organisms assists in the understanding of the role these genes both in normal and pathological situations. A confounding factor in this process is that cross-species comparisons often identify genes which, although highly similar, do not represent a true ortholog and may in fact be functionally dissimilar. In order to resolve this we plan to identify truly orthologous genes in several model organisms, initially in the worm C. elegans, the fly D. melanogaster, the mustard weed A. thaliana and the yeast S. cerevisiae. They will collect and curate a set of genes that are associated with diseases by extracting and filtering excerpts from the Online Mendelian Inheritance in Man. The sequences of the resulting genes will then be analyzed and scored for orthology using methods previously developed at the Center for Genomics and Bioinformatics, primarily the Inparanoid algorithm. This will yield assignments of orthologs in model organisms, at different levels of confidence, forming the database OrthoDisease. Orthodisease is constructed primarily using Inparanoid analysis. Inparanoid is a program that automatically detects orthologs (or groups of orthologs) from 2 species. Pairwise comparisons between several model organisms are shown here. The species are derived mainly from the Ensembl resource as well as several submitted organism genomes. The program itself and accessory programs can be downloaded from the Inparanoid website.The algorithm is based on pairwise similarity scores which are by default calculated with NCBI BLAST program. Inparanoid detects best-best hits between sequences from 2 different species. These are the two seed ORTHOLOGS that form an orthologous group. Other sequences are added to this group if they are closely related to one of the main orthologs. These members of the orthologous group are called IN-PARALOGS. A confidence value is provided for each in-paralog that shows how closely related it is to the main ortholog. Genes/proteins are derived from the Online Inheritence In Man (OMIM) database, both from the Morbid map, which contains genes directly linked to specific diseases,and from the Genemap which contains genes with a know cytogenetic location that have been mentioned in OMIM (but not necessarily in morbidmap). Keyword: model, organism, gene, orthologous, human, disease, gene, Mendelian, man, phenotype, pathological, cross-species, C. elegans, D. melanogaster, mustard, weed, A. thaliana, yeast, S cerevisiae, genomic, bioinformatic, specie, Model Organisms and Comparative Genomics Databases
The Online Macromolecular Museum (OMM) is a site for the display and study of macromolecules. Macromolecular structures, as discovered by crystallographic or NMR methods, are scientific objects in much the same sense as fossil bones or dried specimens: they can be archived, studied, and displayed in aesthetically pleasing, educational exhibits. Hence, a museum seems an appropriate designation for the collection of displays that we are assembling. The OMM''s exhibits are interactive tutorials on individual molecules in which hypertextual explanations of important biochemical features are linked to illustrative renderings of the molecule at hand. Why devote a site to detailed visualizations of different macromolecules? In learning about the intricacies of life processes at the molecular level, it is important to understand how natural selection has fashioned the structure and chemistry of macromolecular machines to suit them for particular functions. This understanding is greatly facilitated by the visualization of 3-dimensional structure, when known. So, if static views of molecules (even in stereo) are worth a thousand words, then interactive animations of molecules should be worth much more. Indeed, we have found the types of displays represented here invaluable in gaining an appreciation for the details of key biochemical processes. Sponsors: This resource is federally-funded by the Protein Data Bank
THIS RESOURCE IS NO LONGER IN SERVICE, documented on March 18, 2013. 15 annotated electron micrographs of different parts of the nervous system. Different nerve tissues are depicted.
This website contains a list of elements from the periodic table. It provides information about the atomic mass, excess mass, binding energy, beta decay energy, half life, mode of decay, and possible parent nuclides for each isotope of the elements listed. Aside from the elements listed, the information for a neutron is also listed.
Statistical test for model organisms association mapping correcting for the confounding from population structure and genetic relatedness. EMMA takes advantage of the specific nature of the optimization problem in applying mixed models for association mapping, which substantially increases the computational speed and the reliability of the results. The current implementation of EMMA is available in an R package. The documentation is included in the installation package.
Web application that uses a combination of computational methods to identify those changes most likely to be cancer-associated.
NSDL is a digital library of exemplary resource collections and services, organized in support of science education at all levels. Starting with a partnership of NSDL-funded projects, NSDL is emerging as a center of innovation in digital libraries as applied to education, and a community center for groups focused on digital-library-enabled science education. The National Science Digital Library (NSDL) was created by the National Science Foundation to provide organized access to high quality resources and tools that support innovations in teaching and learning at all levels of science, technology, engineering, and mathematics (STEM) education. As a national network of learning environments, resources, and partnerships, NSDL seeks to serve a vital role as STEM educational cyberlearning for the nation, meeting the informational and technological needs of educators and learners at all levels. Educators need efficient and reliable methods to discover and use science and math materials that help them meet the demands of instruction, assessment, and professional development in an increasingly complex technology-based world. NSDL provides an organized point of access to: -High-quality STEM content aggregated from a variety of other digital libraries, NSF-funded projects, and NSDL-reviewed web sites. -Services and tools that enhance the use of this content in a variety of contexts. NSDL is designed primarily for K-16 educators, but anyone can access NSDL.org and search the library at no cost. Access to most resources discovered through NSDL is free; however, some content providers may require a login, or a nominal fee or subscription to retrieve their specific resources. NSDL serves as a nexus for educators, researchers, policy makers and the public by building bridges: -Between private sector and public interests by providing access to resources such as publisher'' journal articles, teacher-created lesson plans and real-time data sets from scientists -Between the scientific, research and educational communities by applying advanced technologies to stimulate new ways for educators and learners to access and use scientific information -Between teachers and learners at all levels, in all locations by supplying content and tools in open-access, non-proprietary formats in an easily accessible online environment. Sponsors: This work supported by the National Science Foundation under Grant No. 0733600, Grant No. 0424671, Grant No. 0227648, Grant No. 0227656, and Grant No. 0227888.
Concise, clinically useful information for diagnosing, managing and treating infectious diseases in adults; however it does cover some pediatric topics including vaccines. It is designed for primary care providers and other non-infectious disease specialists as a tool that can be used at the point of care to assist in prescribing antibiotics.
Generates data for use in developing and refining computational tools for comparing genomic sequence from multiple species. The NISC Comparative Sequencing Program's goal is to establish a data resource consisting of sequences for the same set of targeted genomic regions derived from multiple animal species. The broader program includes plans for a diverse set of analytical studies using the generated sequence and the publication of a series of papers describing the results of those analysis in peer-reviewed journals in a timely fashion. Experimentally, this project involves the shotgun sequencing of mapped BAC clones. For each BAC, an assembly is first performed when a sufficient number of sequence reads have been generated to provide full shotgun coverage of the clone. At that time, the assembled sequence is submitted to the HTGS division of GenBank. Subsequent refinements of the sequence, including the generation of higher-accuracy finished sequence, results in the updating of the sequence record in GenBank. By immediately submitting our BAC-derived sequences to GenBank, it makes their data available as a public service to allow colleagues to speed up their research, consistent with the now well-established routine of sequencing centers participating in the Human Genome Project. However, at the same time, it has made considerable investment in acquiring these mapping and sequence data, including sizable efforts of graduate students, postdoctoral fellows, and other trainees. Furthermore, in most cases, large data sets involving multiple BAC sequences from multiple species must first be generated, often taking many months to accumulate, before the planned analysis can be performed and the resulting papers written and submitted for publication.
THIS RESOURCE IS NO LONGER IN SERVICE, documented on July 17, 2013. It is freely available as a reference dataset for the statistical analysis of sequence and structure features of proteins in the PDB. It is a dataset of structurally dissimilar proteins. This dataset has been compiled by selecting well resolved representatives from the Topology level of the CATH database which hierarchically classifies all protein structures. These have been been pruned to remove: i) domains that may contain homologous elements (by pairwise sequence comparison and structural superposition of aligned residues) ii) internal duplications (by repeat detection) iii) regions with high B-Factor The statistical analysis of protein structures requires datasets in which structural features can be considered independently distributed, i.e. not related through common ancestry, and that fulfill minimal requirements regarding the experimental quality of the structures it contains. However, non-redundant datasets based on sequence similarity invariably contain distantly related homologues. Here a reference dataset of non-homologous protein domains is provided, assuming that structural dissimilarity at the topology level is incompatible with recognizable common ancestry. It contains the best refined representatives of each Topology level, validates structural dissimilarity and removes internally duplicated fragments. The compilation of Nh3D is fully scripted. The current Nh3D list contains 570 domains with a total of 90780 residues. It covers more than 70% of folds at the Topology level of the CATH database and represents more than 90% of the structures in the PDB that have been classified by CATH. Even though all protein pairs are structurally dissimilar, some pairwise sequence identities after global alignment are greater than 30%. Nh3D is freely available as a reference dataset for the statistical analysis of sequence and structure features of proteins in the PDB.
The NCI DIS 3D database is a collection of 3D structures for over 400,000 drugs. The database is an extension of the NCI Drug Information System. The structural information stored in the DIS is only the connection table for each drug. The connection table is just a list of which atoms are connected and how they are connected. It is essentially a searcheable database of three-dimensional structures has been developed from the chemistry database of the NCI Drug Information System (DIS), a file of about 450,000 primarily organic compounds which have been tested by NCI for anticancer activity. The DIS database is very similar in size and content to the proprietary databases used in the pharmaceutical industry; its development began in the 1950s; and this history led to a number of problems in the generation of 3D structures. This information can be searched to find drugs that share similar patterns of connections, which can correlate with similar biological activity. But the cellular targets for drug action, as well as the drugs themselves, are 3 dimensional objects and advances in computer hardware and software have reached the point where they can be represented as such. In many cases the important points of interaction between a drug and its target can be represented by a 3D arrangement of a small number of atoms. Such a group of atoms is called a pharmacophore. The pharmacophore can be used to search 3D databases and drugs that match the pharmacophore could have similar biological activity, but have very different patterns of atomic connections. Having a diverse set of lead compounds increases the chances of finding an active compound with acceptable properties for clinical development. Sponsor: The ICBG are supported by the Cooperative Agreement mechanism, with funds from nine components of the NIH, the National Science Foundation, and the Foreign Agricultural Service of the USDA.
THIS RESOURCE IS NO LONGER IN SERVICE, documented on 6/24/13. A repository of information on commercially available phospho-specific antibodies to human phosphorylation sites. It provides a BLAST search for phosphorylation sites using as query the amino acid sequence surrounding the site. It also provides direct links to the relevant antibodies from many companies including BD Pharmingen, Biosource International, Cell Signaling Technology (CST), Santa Cruz Biotechnologies, Upstate Biotechnology.
A database of manually annotated mammalian protein complexes. To obtain a high-quality dataset, information was extracted from individual experiments described in the scientific literature. Data from high-throughput experiments was not included.