We support boolean queries, use +,-,<,>,~,* to alter the weighting of terms
A database of human, chimpanzee, mouse, and rat proteases and protease inhibitors, as well as as the growing number of hereditary diseases caused by mutations in protease genes. Analysis of the human and mouse genomes has allowed us to annotate 581 human, 580 chimpanzee, 667 mouse, and 655 rat protease genes. Proteases are classified in five different classes according to their mechanism of catalysis. Proteases are a diverse and important group of enzymes representing >2% of the human, chimpanzee, mouse and rat genomes. This group of enzymes is implicated in numerous physiological processes. The importance of proteases is illustrated by the existence of 99 different hereditary diseases due to mutations in protease genes. Furthermore, proteases have been implicated in multiple human pathologies, including vascular diseases, rheumatoid arthritis, neurodegenerative processes, and cancer. During the last ten years, our laboratory has identified and characterized more than 60 human protease genes. Due to the importance of proteolytic enzymes in human physiology and pathology, we have recently introduced the concept of Degradome, as the complete repertoire of proteases expressed by a tissue or organism. Thanks to the recent completion of the human, chimpanzee, mouse, and rat genome sequencing projects, we were able to analyze and compare for the first time the complete protease repertoire in those mammalian organisms, as well as the complement of protease inhibitor genes. This webpage also contains the Supplementary Material of Human and mouse proteases: a comparative genomic approach Nat Rev Genet (2003) 4: 544-558, Genome sequence of the brown Norway rat yields insights into mammalian evolution Nature (2004) 428: 493-521, A genomic analysis of rat proteases and protease inhibitors Genome Res. (2004) 14: 609-622, and Comparative genomic analysis of human and chimpanzee proteases Genomics (2005) 86: 638-647.
The defensins knowledgebase is a manually curated database and information source devoted to the defensin family of antimicrobial peptides. The current version of the database holds a comprehensive collection of 363 defensin records each containing sequence, structure and activity information. A web-based interface provides access to the information and allows for text-based searching on the data fields. With the rapidly increasing interest in defensins, we hope that the knowledgebase will prove to be a valuable resource in the field of antimicrobial peptide research.
Decoys-R-Us is a database of decoys, computer-generated conformations of protein sequences that possess some characteristics of native proteins, but are not biologically real. The primary use of decoys is to test scoring, or energy, functions. All the decoys in the Decoys ''R'' Us database can be downloaded.
:DDOC provides a comprehensive compilation of the published research related to the genes associated with ovarian cancer. DDOC provides details of the cell line, tissue or cell type, expression status, disease stage, tumor grade, OC type and laboratory method provided in the literature. The links to the relevant sources of data used to extract information related to genes are also included. Many aspects of the information provided in the DDOC were curated by biologists, which increases its accuracy. DDOC is freely accessible for academic and non-profit users.
DBTGR provides information on tunicate gene regulation, such as the location of expression, or the identified regulatory elements present in promoter sequences. The database also contains the promoters of homologous genes in multiple species to allow identification of conserved cis elements.
dbPTM is a database that compiles information on protein post-translational modifications (PTM) such as the modified sites, solvent accessibility of surrounding amino acids, protein secondary and tertiary structures, protein domains, and protein variations. The version 2.0 of dbPTM integrates the experimentally validated PTM sites with referable literatures from Swiss-Prot, Phospho.ELM, O-GLYCBASE, and UbiProt. In all of the collected PTM information, about 25 types of PTM with enough experimentally validated sites are trained the profile hidden Markov models (HMMs) to detect the potential PTM sites with 100% specificity against Swiss-Prot proteins. To help users investigating more detail in each type of PTM, the substrate peptide specificity such as positional amino acid frequency, solvent accessibility and secondary structure surrounding the modified sites are also provided. Moreover, the information of orthologous protein clusters is provided to users for analyzing whether the PTM sites located in the evolutionary conserved regions or not., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
The Liver Expression Profile database aims to be an information center of liver protein expression profile. dbLEP contains two datasets, with plans to provide more datasets in the future. For each dataset, dbLEP provides all identification results including none-redundant identified protein, all possible identified proteins, peptides and their spectrums. The detailed annotation is also provided for each identified protein. Benefit from large number of intact data resources, abundant links and flexible search functions, researchers may get all the information for the proteins they are interested in by text query or similarity comparison. Besides of judging the quality of the identified proteins according to their identified peptides and spectrums, researchers could analysis these data using the annotation information. We hope dbLEP could help you step from data to knowledge finally.
The Cypriot National Genetic Database is an online repository of information about inherited disorders in the Cypriot population. The Cypriot National Genetic Database results from the fruitful collaboration among several investigators from Erasmus Medical Center (The Netherlands) and Kypriako Idryma Erevnon Gia Ti Myiki Distrofia (Cyprus) encouraged by the Human Genome Variation Society and financially supported in part by the European FP6 INCO grant MedGeNet and by Asclepion Genetics in Switzerland.
Cyclonet database is a database on cell cycle regulation in eukaryotes. The database contains information about cell cycle specific genes, proteins, protein complexes and their interactions, diagrams of cell cycle regulation for vertebrates, models of cell cycle and results of their analyses, microarray data, literature references and other related resources. The data are compiled from different databases and from literature annotation. Known cell-cycle models are imported from SBML and CellML model repositories and developed by ourselves.
CyanoBase provides an easy way of accessing the sequences and all-inclusive annotation data on the structures of the cyanobacterial genomes. Users can view by data type, search using BLAST2, or search by species.
A database of information about each Cancer-Testis (CT) gene, its gene products and the immune response induced in cancer patients by these proteins. CT antigens are proteins normally expressed only in the human germ line but that are also present in a significant subset of malignant tumors. The practical importance of these proteins is that due to their restricted expression pattern they are frequently recognized by the immune system of cancer patients. Moreover, this antigenicity has raised the possibility of their being used as vaccines to actively stimulate immune responses in order to combat tumor growth. As a result worldwide research into many aspects of CT antigens is rapidly growing prompting the construction of this database as a resource for investigators involved in this area.
CREMOFAC is a database for chromatin remodeling factors has been developed. The database harbors 64 types of remodeling factors from 49 different organisms reported in literature and facilitates a comprehensive search for them. In addition, it also provides in-depth information for the factors reported in the three widely studied mammals namely, human, mouse and rat. Further, information on literature, pathways, and phylogenetic relationships has also been covered.
The Crop EST Database (CR-EST) is a public available online resource providing access to sequence, classification, clustering, and annotation data of crop EST projects at the IPK. Summarized numbers about genomic data of species are listed in tables. The main database content is original sequence data and cDNA library information from different organisms as well as results from BlastX searches against major protein sequence databases contained in NRPEP. Additionally sequence alignments of stackPACK clustering projects are available. This web application allows to BLAST against CR-EST ESTs and to query and retrieve data from Gene Ontology and metabolic pathway annotations as well as sequence similarities from stored results of BLASTX searches against the NRPEP database. CR-EST also features interactive JAVA-based tools, such as open reading frame visualization and explorative analysis of Gene Ontology mappings to ESTs.
The FlyTF database contains information on the manual curation of FlyBase identifiers based on FlyBase/Gene Ontology annotation or the DBD Transcription Factor Database. FlyBase identifiers are putative site-specific transcription factors. There are currently1052 of them in this database.
THIS RESOURCE IS NO LONGER IN SERVICE, documented on July 15, 2013. Non-coding DNA segments that are conserved across multiple homologous genomic sequences are good indicators of putative regulatory elements. We use a systematic approach to delineate such conserved non-coding blocks from a collection of vertebrate species. Upstream regions of homologous gene pairs from man, rhesus monkey, mouse, rat, dog, cow, chicken, tetraodon, zebrafish and xenopus are considered for this purpose. Pairwise as well as Multiple alignments based on the pairwise ones are available. Sequence conservation in non-coding, upstream regions of orthologous genes from man and mouse is likely to reflect common regulatory DNA sites. Motivated by this assumption we have delineated a catalogue of conserved non-coding sequence blocks and provide the CORG-''COmparative Regulatory Genomics''-database. The data were computed based on statistically significant local suboptimal alignments of 15 kb regions upstream of the translation start sites of, currently, 10 793 pairs of orthologous genes. The resulting conserved non-coding blocks were annotated with EST matches for easier detection of non-coding mRNA and with hits to known transcription factor binding sites. CORG data are accessible from the ENSEMBL web site via a DAS service as well as a specially developed web service for query and interactive visualization of the conserved blocks and their annotation.
Comparasite is an integrated database of our original full-length cDNA sequence data. It consists of seven sub-databases of apicomplexa protozoa, Plasmodium falciparum, Plasmodium yoelii, Plasmodium vivax, Toxoplasma gondii, Cryptosporidium parvum, Echinococcus multilocularis. Homologous gene groups are clustered and comparative analysis of any combination of these seven species is implemented, such as interspecies comparisons as to cellular localization, motifs or transmembrane regions and so on. For submitted keywords and other search conditions, Comparasite retrieves orthologous gene groups containing a given protein motif/GO term etc in common or in a species-specific manner. By enabling multi-faceted comparative analyses of genes of apicomplexa protozoa, monophyletic organisms that have evolved to diversify to parasitize various hosts by adopting complex life cycles, Comparasite should help elucidate the mechanism behind parasitism.
COMe is an attempt to classify metalloproteins and some other complex proteins using the concept of bioinorganic motif. COMe consists of three types of entities: Molecule (MOL), Bioinorganic Motif (BIM), and Bioinorganic Proteins (PRX). Both Molecule (MOL) and Bioinorganic Motif (BIM) entities consist of substructure elements. Substructure literally means that these elements can form parts of a bigger structure, i.e. complete protein. MOL is an entity representing small molecule (as opposed to macromolecule) which, in complex with polypeptide, forms a functional protein. Users can query by COMe ID or using a case sensitive or insensitive text search, use predefined queries, or look at the paths via the online ontology provided by COMe.
Database dedicated to the analysis of the genome of Escherichia coli. Its purpose is to collate and integrate various aspects of the genomic information from E. coli, the paradigm of Gram-negative bacteria. Colibri provides a complete dataset of DNA and protein sequences derived from the paradigm strain E. coli K-12, linked to the relevant annotations and functional assignments. It allows one to easily browse through these data and retrieve information, using various criteria (gene names, location, keywords, etc.). The data contained in Colibri originates from two major sources of information, the reference genomic DNA sequence from the E. coli Genome Project and the feature annotations from the EcoGene data collection., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
coliBASE is a database for comparative genome analysis of Enterobactericaiae, e.g. Escherichia, Shigella etc. coliBASE covers a greater sequence diversity than most other online E. coli resources (which tend to be limited to the K12 genome sequence), and provides novel tools such as the alignment viewer and whole genome viewer that are not available elsewhere. It is supported by a 5 year BBSRC grant until Feburary 2012.
COGEME is an ongoing BBSRC-funded study to construct a relational database of genomic information from phytopathogenic fungi. This site also hosts microarray data for Blumeria graminis. Expressed sequence tags (ESTs) obtained from eighteen species of plant pathogenic fungi, two species of phytopathogenic oomycete and three species of saprophytic fungi are included here. Hierarchical clustering software was used to classify together ESTs representing the same gene and produce a single contig, or consensus sequence. The unisequence set for each pathogen therefore represents a set of unique gene sequences, each one consisting of either a single EST or a contig sequence made from a group of ESTs. Unisequences were annotated based on top hits against the NCBI non-redundant protein database using blastx.