We support boolean queries, use +,-,<,>,~,* to alter the weighting of terms
POINT is a protein-protein interaction database. It includes annotation of interologs and protein phsophorylation. This work analyzes the applicability of orthologs-based PPI prediction and provide the theoretical upper-bound of this approach.
A relational database that integrates data from rice, maize, and Arabidopsis by placing the complete Arabidopsis and rice proteomes, and the available maize sequences into "putative orthologous groups" (POGS). Annotation efforts are now beginning and will focus on predicted RNA binding proteins (e.g. those with known RNA binding domains or known to influence RNA function). Putative Orthologous Groups (POGs) form the heart of the database, and were assigned using a mutual best hit strategy after performing BLAST comparisons of the predicted Arabidopsis and rice proteomes. Each POG entry includes cross-referenced orthologs and paralogs in Arabidopsis and rice, annotated with domain organization, gene models, phylogenetic trees showing closely-related proteins, and intracellular targeting predictions. The database can be queried to identify POGs with specific domain combinations and predicted intracellular locations.
A plastid protein database. It integrates data from large scale proteome analyses of different plastid types.These include etioplasts, chloroplasts, chromoplasts and the undifferentiated proplastid-like organelles of tobacco BY2 cells. This comparison allows establishing a core proteome that is common to all plastid types and provides furthermore information about plastid type-specific functions.
It is an objective classification system for plan proteins based on cluster analyses of the inferred proteomes of the sequenced angiospermsArabidopsis thaliana v Columbia, Oryza sativa v. japonica (Rice), and Populus trichocarpa (poplar). Sequence data for Carica papaya and Medicago papaya are also included in the current version of Tribes. Results for these species are currently masked from view, but will be available when the genomes are publicly released. In addition to the genome-based tribe scaffold, unigenes from more than 200 plant and algal species TIGR Transcript Assemblies have been associated with each tribe (see documentation), resulting in a global classification of about 4 million putative plant protein sequences. PlantTribes 1.0 incorporates an extensive collection of microarray expression data from Arabidopsis microarray experiments. Expression data is linked to the individual genes in PlantTribes, and can be accessed through any result including Arabidopsis gene sequences. PlantTribes is based on the similarity-based clustering procedure TribeMCL (Enright et al, 2002,2003) to classify protein-coding genes into putative gene families. MCL classifications have been constructed using three clustering stringencies , allowing the user to explore the stability of the protein classification. A second round of MCL clustering identifies SuperTribes that approximate objective superfamilies. PlantTribes also includes information about domains, traditional gene family names, and a unified nomenclature based on common terms.
Software tool that allows to display very long data vectors in a space-efficient manner, allowing the user to visually judge the large scale structure and distribution of features simultaneously with the rough shape and intensity of individual features.
It provides access to a veriety of databases of plant genomics databases. These include: *Arabidopsis PARE Database *Arabidopsis SBS Database *Arabidopsis MPSS Plus Database *Rice MPSS Database *Legume SBS Database *Maize SBS Database *Gallus Gallus SBS Database *Magnaporthe MPSS Database *Grape MPSS Database These data have been generated by a series of different projects, mainly based at the University of Delaware and funded primarily by grants from the National Science Foundation and the United States Department of Agriculture. The database and the web pages that you see were produced by the Meyers lab, with data generated within our lab, by collaborating laboratories like the Green or Wang labs. Some data we've gathered from Genbank or other public sources. Much of the data was produced under contract by Illumina, Inc (Hayward, California), but some SBS libraries are now being sequenced at other sites.
A small ontology for anatomical spatial references, such as dorsal, ventral, axis, and so forth.
A web-based database of protein interaction sites. PiSITE provides not only information of interaction sites of a protein from single PDB entry, but also information of interaction sites of a protein from multiple PDB entries including similar proteins. PiSite also provides a list of sociable proteins, proteins with multiple binding states and multiple binding partners.In PiSITE, the identification of the binding sites of protein chains is performed by searching the same proteins with different binding states in PDB at first, and then mapping those binding sites onto the query proteins. The database PiSITE provides real interaction sites of proteins using the complex structures in PDB. According to the progress of several structural genomic projects, we have a large amount of structural data in PDB. Consequently, we can observe different binding states of proteins in atomic resolutions, and can analyze actual interaction sites of proteins. It will lead better understandings of protein interaction sites in near future. Usual practice to identify the interaction site has been done using a representative complex in PDB. However, for the proteins with multiple partners, non-interaction sites identified by using a single complex structure is not enough, because some part of the non-binding sites may be involved in the interaction sites with another proteins. Therefore, the real interaction sites should be obtained by using all of the binding states in PDB. For the purpose, the identifications of the binding site in PiSITE are done by searching the same proteins with different binding sites in PDB at first, and then mapping the binding sites onto the query proteins. PiSITE also provides the lists of transient hub proteins, which we call sociable proteins to clarify the different of so-called hub proteins. The sociable proteins are identified as the proteins with multiple binding states and multiple binding partners. On the other hand, so-called hub proteins have been identified as the proteins at the hub position in protein-protein interaction networks obtained by large-scale experiments, but the definition of the hub proteins cannot differentiate transient hub proteins from stable ones, although the differentiation is critically important for the better understanding of protein interaction networks. In addition, the usual definition of hub proteins can contain supermolecules as hub proteins. The supermolecules can be identified as the proteins with a single binding state and multiple binding partners, which we call stable hub proteins.
A web analysis system and resource, which provides comprehensive information on piRNAs in the widely studied mammals. It compiles all the possible clusters of piRNAs and also depicts piRNAs along with the associated genomic elements like genes and repeats on a genome wide map. piRNABank mainly provides data onnamely Human, Mouse, Rat, Zebrafish, Platypus and a fruit fly, Drosophila.Search options have been designed to query and obtain useful data from this online resource. It also facilitates abstraction of sequences and structural features from piRNA data. piRNABank provides the following features: * Simple search * Search piRNA clusters * Search homologous piRNAs * piRNA visualization map * Analysis tools, THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
A database of predicted human protein-protein interactions. The predictions have been made using a na&iuml;ve Bayesian classifier to calculate a Score of interaction. There are 37606 interactions with a Score &ge;1 indicating that the interaction is more likely to occur than not to occur.
A protein-protein interactions thermodynamic database which contains data of several thermodynamic parameters along with sequence and structural information experimental conditions and literature information. Each entry contains numerical data for features of the interacting proteins such as the free energy change, dissociation constant, association constant, enthalpy change, and heat capacity change. PINT includes: the name and source of the proteins involved in binding, SWISS-PROT and Protein Data Bank (PDB) codes, secondary structure and solvent accessibility of residues at mutant positions, measuring methods, and experimental conditions such as buffers, ions and additives, and literature information. PINT is cross-linked with other related databases such as PIR, SWISS-PROT, PDB and the NCBI PUBMED literature database.
A database dedicated to the study of host-pathogen protein-protein interactions (PPIs). PIG provides a number of user interfaces for searching available data and tools for predicting interactions between host and pathogen proteins.
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 19, 2016. A database for the study of protein inter-atomic distance distribution. Currently, the distances are extracted from the protein structures determined through X-ray Crystallography, but they could also be obtained from NMR structural models. The known structures with the resolution higher than 2A and less than 70% sequence similarities are selected. Each type of distances is specified in terms of the types of the atoms it involves, the types of the residues containing the atoms, and the types of the residues in between the two end residues in sequence. An automated system is built to generate and process the data dynamically. The system consists of two levels of databases. The first one stores the sequence and structure information for a large set of high-resolution protein structures, with a similar data structure as the structural data represented in the PDB Data Bank. The second one stores the information for the distance distributions, with each record corresponding to a distribution function. The second database is built dynamically from the first one. The database can provide structural information in terms of distance distributions to structural biologists. Such information can be valuable for the study of many fundamental biological problems including protein structure prediction and determination, protein dynamics simulation, molecular design, protein structural analysis and classification, etc.
An online comparative genomics resource that is built upon publicly available sequence and map information from a diverse set of plant species, with a focus on the angiosperms, or flowering plants. It provides an interface to the results from a variety of phylogenomic analyses. Phytome is designed to facilitate functional genomics, molecular breeding and evolutionary studies in model and non-model plant species. Currently, Phytome contains phylogenetic and functional information for predicted protein sequences ("Unipeptides"). Future development will incorporate data and tools for analysis of sequence-based comparative maps.
A database of phylogenetic patterns of evolution between 46 different species. PhyloPat uses the latest release of EnsMart (release 52), and their one-to-one, one-to-many and many-to-many orthologies. First, we stored all of the Ensembl IDs within the 46 species, and the orthologies between them. Second, we determined the evolutionary order of the studied species using the NCBI Taxonomy database. The phylogenetic tree of these species can be viewed here. Third, we used this phylogenetic tree as a starting point for building our phylogenetic lineages. For each gene in the first species (S. cerevisiae), we looked for orthologs in the other species. All orthologs were added to the phylogenetic lineage, and in the next round were checked for orthologs themselves, until no more orthologies were found for any of the genes. This process was repeated for all genes in all species that were not connected to any phylogenetic lineage yet. The complete phylogenetic lineage determination generated 329,998 phylogenetic lineages, consisting of 973,821 genes. These lineages can be queried here by phylogenetic patterns, MySQL regular expressions or simply a list of Ensembl/EMBL/EntrezGene/HGNC IDs. Output can be given in HTML, Excel or plain text format.
Database for phylomes, that is, complete collections of phylogenetic trees for all proteins encoded in a given genome. It aims at providing a repository of high-quality phylogenies and alignments for proteins encoded in model species. To derive a phylome, each protein encoded in a given genome is used as a seed to retrieve its homologs in other complete genomes. These sequences are aligned and processed to derive reliable phylogenies using several phylogenetic methods. Besides providing the evolutionary history of the gene families, phylomeDB includes phylogeny based predictions of orthology and paralogy relationships., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
It contains pre-calculated structural and phylogenomic analyses of over 57,000 protein families and domains. The PhyloFacts resource includes "books" for protein families across the Tree of Life. Each book includes a multiple sequence alignment, one or more phylogenetic trees, predicted subfamilies, predicted 3D protein structures, active sites and other key residues, cellular localization, and Gene Ontology (GO) annotations and evidence codes. PhyloFacts includes hidden Markov models for classification of user-submitted (DNA or protein) sequences to protein families and subfamilies across the tree of life. Our primary current focus is on covering all the gene families represented in the human genome and all structural domains, but plan to expand the resource to include all proteins in all species. The protein families in this resource typically contain homologs from many species. The phylogenetic distribution of a protein family can vary from highly restricted (e.g., to hominidae or mammals) to throughout the tree of life. Gathering homologs from many divergent species enables us to take advantage of experimental investigations in different systems, and allows powerful inferences of function and structure that might not otherwise be possible.
A publicly available database resource containing the assembled partial genomes for ~700 eukaryotic organisms. Partial genomes are generated from expressed sequence tag datasets containing more than 1000 sequences. PartiGeneDB allows users to view sets of genes and identify genes of interest in organisms for which a full genome is not currently available. PartiGeneDB is automatically updated to include new organism datasets as they are generated. PartiGeneDB provides four portals of entry into the database. It is hosted and supported by the Hospital for Sick Children, Toronto. In addition to providing a comprehensive resource facilitating comparative analyses, PartiGeneDB allows researchers to access the partial genomes of organisms that may not be available elsewhere. However, we recommend and encourage users interested in exploring datasets from a single organism in more depth, that you visit the specific web sites associated with the sequencing effort associated with that organism .
Growth, trait and development ontology for soybean