We support boolean queries, use +,-,<,>,~,* to alter the weighting of terms
A database of comprehensive information about mutant phenotypes, reporter-gene expression patterns, flanking sequences of T-DNA insertional sites, seed availability, and others are collected in the database. RMD can be searched by keywords, nucleotide sequence or protein sequence. This database provides three classes of functions: (1) identifying novel genes, (2) identifying regulatory elements, and (3) identifying pattern lines for ectopic expression (misexpression) of target gene at specific tissue or at specific growth stage.
The Saccharomyces Cerevisiae Morphological Database(SCMD) is a collection of micrographs of budding yeast mutants. Micorgraphs of mutants with altered cell morphology from a set of the haploid MATa deleted strains obtained from EUROSCARF. From the micrographs, disruptant cells are automatically extracted by our novel cell-image processing software.
THIS RESOURCE IS NO LONGER IN SERVICE, documented on July 17, 2013. It provides information about primers and probes that can be used to quantitate human and mouse mRNA by reverse transcription polymerase chain reaction (RTx96PCR) assays. Users can search the QPPD to find: * Primer sets and probes for a given gene * Primer location * Amplicon size * Assay type * Positions of single nucleotide polymorphisms (SNPs) * Literature references * Available I.M.A.G.E. cDNA clones * The Primer Viewer, a graphical representation of the gene and primer sets, which includes hyperlinks to Gene Info from the Cancer Gene Anatomy Project (CGAP) and the CGAP SNP viewer.
It provides information on natural and artificial mutants, including random and site-directed ones, for all proteins except members of the globin and immunoglobulin families. The PMD is based on literature, and each entry in the database corresponds to one article which may describe one, several or a number of protein mutants. Each database entry is identified by a serial number and is defined as either natural or artificial, depending on the type of the mutation. For each entry the following are recorded : JOURNAL, TITLE, CROSS-REFERENCE, PROTEIN, N-TERMINAL, CHANGE, FUNCTION, STRUCTURE, STABILITY, etc. CROSS-REFERENCE indicates the code names of the protein given in other databases such as Protein Identification Resources (2). N-TERMINAL shows the N-terminal sequence of five amino acids which may help to show the unambiguous numbering of th e sequence. CHANGE indicates the position and kind of mutations, such as amino acid substitution, insertion and deletion, denoted with a specific notation. Any functional or structural features (FUNCTION, STRUCTURE, STABILITY,etc) observed in the mutant are described immediately after ''CHANGE''. Relative differences in activity and/or stability, in comparison with the wild-type protein, are indicated with symbols (- -),(-),(=),(+) or (+ +). Complete loss of activity is denoted as (0). Data Submission A data submission system was newly prepared in the PMD. We welcome the authors of articles published in academic journals to submit their own mutant data to the PMD. After checking the contents, we will register the data with a unique accession number.
THIS RESOURCE IS NO LONGER IN SERVICE, documented on July 16, 2013. A database of protein-protein homo- and hetero-complexes as well as domain-domain structures. This issue of the database contains 17.024 entries (as of October 2007) of which 1350 are two-chain protein hetero-complexes, 7773 homodimers and 1589 are one-chain proteins parsed into two domains (domain structures). The rest of entries are constructed from PDB files for multi-chain protein complexes by leaving only two interacting chains, (in the database nomenclature designators of those chains are appended to the name of the original PDB file). The homo-complexes in this database are the complexes with two monomers having more than 95 % sequence identity. Each entry consists of the X-ray structure taken from the PDB data bank (in the PDB/ folder), sequence file (in the SEQ/ folder) and the data file (in the INF/ folder) where all relevant information is stored. File names consist of PDB ID and extensions .pdb (X-ray structures), .seq (sequence files or .inf (data files). Current implementation of the database accommodates the following information for each entry, separately for larger (denoted as chain A*) and smaller (chain B) components: * Full sequence; * Number of residues; * Number of residues on the interface; * List of interfacial residues; * Number of helices and strands; * Absolute (in �2) and relative interface areas. The database is searchable with respect to majority of the above parameters. Additional search is possible with respect to the PDB ID and protein names. Downloads are possible for both individual entries (as plain text files) and for the whole content or search (checked) result list (as one gzipped file). Supplementary search and download is available for the subset of the database at 40% sequence identity level (the DPPC40 database).
This database provides a unified resource to analyze the effects of alternative splicing events on the structure of the resulting protein isoforms. ProSAS comprehensively annotates protein structures for several Ensembl genomes and alternative transcripts can be analyzed on the protein structure and protein function level using the intuitive user interface of the database. Users can search based on Ensembl gene or Ensembl transcript ids, Gene descriptions, Uniprot gene names, Genes matching patterns, Swissprot/Uniprot identifiers or Affymetrix probeset ids.
A sister database to ProSite, is constituted of manually created rules that increase the discriminatory power of PROSITE motifs (generally profiles) by providing additional information about functionally and/or structurally critical amino acids and can automatically generate annotation based on PROSITE motifs in the UniProtKB/Swiss-Prot format. Each ProRule is defined in the UniRule format. In addition to these rules corresponding to a unique PROSITE motif, there are also rules triggered by a specific combination of PROSITE motifs called metamotifs. Metamotifs allow the definition of arrangements of domains separated by spacers of variable size, as well as the anchoring to the N- and/or C-termini and the exclusion of a PROSITE motif. ProRule uses the UniRule format that is common to all types of rules created to annotate UniProtKB/Swiss-Prot, including the HAMAP family rules. Each rule contains information used to provide template based annotation associated with the domain or family detected by the PROSITE motif. ProRule is used to create UniProtKB/Swiss-Prot lines with basic and complex annotation derived from the presence of the domain and of biologically critical amino acids: domain name and boundaries, EC number, function, keywords, associated PROSITE patterns, PTMs, active sites, disulfide bonds, etc.). ProRule contains notably the position of structurally and/or functionally critical amino acid(s), as well as the condition(s) they must fulfil to play their biological role(s). Part of these supplementary data are used by ScanProsite that not only provides the protein sequence matched by a profile, but also information about the relevance of biologically meaningful residues, like active sites, binding sites, post-translational modification sites or disulfide bonds, to help function determination
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on May 12,2023. Database of interactions between amino acid residues of enzyme and its ligands. Provides summary of interactions between amino acid residues of enzyme and its various ligands including substrate and transition state analogues, cofactors, inhibitors, and products., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
PPNEMA is abioinformatic database of rRNA genes from plant-parasitic nematodes. It consists of a database of ribosomal cistron sequences from various species grouped according to nematode genera, and a search system allowing data to be extracted according to both text and pattern searching. PPNEMA offers to the scientific community a preprocessed archive of plant parasitic nematode sequences useful for nematologists. It is a tool to retrieve plant nematode multialigned sequences for phylogenetic studies or to recognize a nematode by comparing its rDNA sequence with the PPNEMA available genus specific multialignments.
A Plant Proteome DataBase for Arabidopsis thaliana and maize (Zea mays). The PPDB stores experimental data from in-house proteome and mass spectrometry analysis, curated information about protein function, protein properties and subcellular localization. Importantly, proteins are particularly curated for possible (intra) plastid location and their plastid function. Protein accessions identified in published Arabidopsis (and other Brassicacea) proteomics papers are cross-referenced to rapidly determine previous experimental identification by mass spectrometry. All protein-encoding gene models in the Arabidopsis nuclear and organellar genomes, as assembled by TAIR, as well as all maize EST assemblies (ZmGI) as assembled by DFCI Maize Gene Index project. These are all uploaded in PPDB and are linked to each other via a BLAST alignment. Thus every predicted protein in both species can be searched for experimental and other information (even if not experimentally identified).
A searchable database of pKa values of amino acid residues within proteins. These values have been measured experimentally using a variety of methods, including NMR and the linearized Poisson-Boltzmann equation. This database of protein ionization constants was sourced from the primary literature and contains in excess of 1400 entries. The database contains pKa values for amino acid side-chains, as well as the N and C termini, over 75% of which focus on Glutamate, Lysine, Histidine and Aspartate. These four residues are all key ionizable residues, and therefore the apparent bias is not driven by our selection, but by the available experimental data. Very little data is currently available for Arginine: its pKa value (~12) essentially precludes measurement by titration as proteins will denature at high basic pH.
A database of information on pox viruses. Goals of this project are to acquire and annotate data on poxviruses, and to develop and utilize new tools to facilitate the study of this group of organisms. This basic research is being undertaken with an eye toward the development of novel antiviral therapies, vaccines against human orthopoxvirus infections, new approaches for the environmental detection of virions, and methods to accomplish more rapid diagnosis of disease.
An integrated database of human coding single nucleotide polymorphisms (SNPs) and their annotations. Unlike other databases of similar nature, apart from integrating several coding SNPs (cSNPs) and protein-related information resources, we predict the implications of the non-synonymous SNPs (nsSNPs) using two well known algorithms (SIFT and PolyPhen). The results are presented in an intuitive visualization that depicts the cSNPs mapped onto protein domains and highlights those nsSNPs that are potentially damaging/deleterious or have been reported as disease allelic variants (based on OMIM). The query interface also supports searching for a list of proteins associated with any gene ontology term, pathway, disease term or gene family. Results can also be downloaded as a spreadsheet. The visualization page also provides links to several other related sources and dynamic links to literature references.
A database of quadruplex motifs. It is composed of two parts (EuQuad and ProQuad). EuQuad gives information on quadruplex motifs present in human, chimpanzee, rat and mouse genes. ProQuad contains quadruplex information of 146 prokaryotes. Apart from gene-specific searches QuadBase has a number of other modules. &lsquo;Orthologs Analysis' queries for conserved motifs across species in a user-defined manner; &lsquo;Pattern Search' can be used to fetch specific motifs of interest and the &lsquo;Pattern Finder' tool can search for motifs in any given sequence
A database of mRNA polyadenylation sites. PolyA_DB version 1 contains human and mouse poly(A) sites that are mapped by cDNA/EST sequences. PolyA_DB version 2 contains poly(A) sites in human, mouse, rat, chicken and zebrafish that are mapped by cDNA/EST and Trace sequences. Sequence alignments between orthologous sites are available. PolyA_SVM predicts poly(A) sites using 15 cis elements identified for human poly(A) sites.
POINT is a protein-protein interaction database. It includes annotation of interologs and protein phsophorylation. This work analyzes the applicability of orthologs-based PPI prediction and provide the theoretical upper-bound of this approach.
A relational database that integrates data from rice, maize, and Arabidopsis by placing the complete Arabidopsis and rice proteomes, and the available maize sequences into "putative orthologous groups" (POGS). Annotation efforts are now beginning and will focus on predicted RNA binding proteins (e.g. those with known RNA binding domains or known to influence RNA function). Putative Orthologous Groups (POGs) form the heart of the database, and were assigned using a mutual best hit strategy after performing BLAST comparisons of the predicted Arabidopsis and rice proteomes. Each POG entry includes cross-referenced orthologs and paralogs in Arabidopsis and rice, annotated with domain organization, gene models, phylogenetic trees showing closely-related proteins, and intracellular targeting predictions. The database can be queried to identify POGs with specific domain combinations and predicted intracellular locations.
A plastid protein database. It integrates data from large scale proteome analyses of different plastid types.These include etioplasts, chloroplasts, chromoplasts and the undifferentiated proplastid-like organelles of tobacco BY2 cells. This comparison allows establishing a core proteome that is common to all plastid types and provides furthermore information about plastid type-specific functions.
It is an objective classification system for plan proteins based on cluster analyses of the inferred proteomes of the sequenced angiospermsArabidopsis thaliana v Columbia, Oryza sativa v. japonica (Rice), and Populus trichocarpa (poplar). Sequence data for Carica papaya and Medicago papaya are also included in the current version of Tribes. Results for these species are currently masked from view, but will be available when the genomes are publicly released. In addition to the genome-based tribe scaffold, unigenes from more than 200 plant and algal species TIGR Transcript Assemblies have been associated with each tribe (see documentation), resulting in a global classification of about 4 million putative plant protein sequences. PlantTribes 1.0 incorporates an extensive collection of microarray expression data from Arabidopsis microarray experiments. Expression data is linked to the individual genes in PlantTribes, and can be accessed through any result including Arabidopsis gene sequences. PlantTribes is based on the similarity-based clustering procedure TribeMCL (Enright et al, 2002,2003) to classify protein-coding genes into putative gene families. MCL classifications have been constructed using three clustering stringencies , allowing the user to explore the stability of the protein classification. A second round of MCL clustering identifies SuperTribes that approximate objective superfamilies. PlantTribes also includes information about domains, traditional gene family names, and a unified nomenclature based on common terms.
Software tool that allows to display very long data vectors in a space-efficient manner, allowing the user to visually judge the large scale structure and distribution of features simultaneously with the rough shape and intensity of individual features.