We support boolean queries, use +,-,<,>,~,* to alter the weighting of terms
BioThesaurus is a web-based system designed to map a comprehensive collection of protein and gene names to UniProt Knowledgebase protein entries. It covers all UniProtKB protein entries, and consists of several millions of names extracted from multiple resources based on database cross-references in iProClass. The web site allows the retrieval of synonymous names of given protein entries and the identification of ambiguous names shared by multiple proteins. Searches can be done on protein/gene name, organism, or unique identifier.
Bionemo stores manually curated information about proteins and genes directly implicated in biodegradation metabolism. The database includes information on sequence, domains and structures for proteins; and sequence, regulatory elements and transcription units for genes. Bionemo complements other biodegradation databases such as the University of Minessota Biocatalysis/Biodegradation Database, or Metarouter, which focus on the biochemical aspects of biodegradation. Bionemo has been built by manually associating sequences databases entries to biodegradation reactions, using the information extracted from published articles. Information on transcription units and their regulation was also extracted from the literature for biodegradation genes, and linked to the underlying biochemical network.
Biodefense Proteomics Resource Center presents information on Class A-C biodefense organisms. :This list includes Bacillus anthracis, Brucella abortus, Francisella tularensis, salmonella typhi, salmonella typhimurium, Virbio cholerae, Yersinia pestis, Cryptosporidium parvum, Toxoplasma gondii, Avian influenza, SARS, Monkeypox, Vaccinia, and Variola. For each organism, the page provides a general overview of the organism and the diseases it causes, protein (and protein interaction) data, reagents, and data from experiments performed with this organism. Users may also find links to the NCBI Taxonomy center.
The MIPS Ustilago maydis Genome Database aims to present information on the molecular structure and functional network of the entirely sequenced, filamentous fungus Ustilago maydis. The underlying sequence is the initial release of the high quality draft sequence of the Broad Institute. The goal of the MIPS database is to provide a comprehensive genome database in the Genome Research Environment in parallel with other fungal genomes to enable in depth fungal comparative analysis. The specific aims are to: 1. Generate and assemble Whole Genome Shotgun sequence reads yielding 10X coverage of the U. maydis genome 2. Integrate the genomic sequence assembly with physical maps generated by Bayer CropScience 3. Perform automated annotation of the sequence assembly 4. Align the strain 521 assembly with the FB1 assembly provided by Exelixis 5. Release the sequence assembly and results of our annotation and analysis to public Ustilago maydis is a basidiomycete fungal pathogen of maize and teosinte. The genome size is approximately 20 Mb. The fungus induces tumors on host plants and forms masses of diploid teliospores. These spores germinate and form haploid meiotic products that can be propagated in culture as yeast-like cells. Haploid strains of opposite mating type fuse and form a filamentous, dikaryotic cell type that invades plant tissue to reinitiate infection. Ustilago maydis is an important model system for studying pathogen-host interactions and has been studied for more than 100 years by plant pathologists. Molecular genetic research with U. maydis focuses on recombination, the role of mating in pathogenesis, and signaling pathways that influence virulence. Recently, the fungus has emerged as an excellent experimental model for the molecular genetic analysis of phytopathogenesis, particularly in the characterization of infection-specific morphogenesis in response to signals from host plants. Ustilago maydis also serves as an important model for other basidiomycete plant pathogens that are more difficult to work with in the laboratory, such as the rust and bunt fungi. Genomic sequence of U. maydis will also be valuable for comparative analysis of other fungal genomes, especially with respect to understanding the host range of fungal phytopathogens. The analysis of U. maydis would provide a framework for studying the hundreds of other Ustilago species that attack important crops, such as barley, wheat, sorghum, and sugarcane. Comparisons would also be possible with other basidiomycete fungi, such as the important human pathogen C. neoformans. Commercially, U. maydis is an excellent model for the discovery of antifungal drugs. In addition, maize tumors caused by U. maydis are prized in Hispanic cuisine and there is interest in improving commercial production. The complete putative gene set of the Broad Institute''s second release is loaded into the database and in addition all deviating putative genes from a putative gene set produced by MIPS with different gene prediction parameters are also loaded. The complete dataset will then be analysed, gene predictions will be manually corrected due to combined information derived from different gene prediction algorithms and, more important, protein and EST comparisons. Gene prediction will be restricted to ORFs larger than 50 codons; smaller ORFs will be included only if similarities to other proteins or EST matches confirm their existence or if a coding region was postulated by all prediction programs used. The resulting proteins will be annotated. They will be classified according to the MIPS classification catalogue receiving appropriate descriptions. All proteins with a known, characterized homolog will be automatically assigned to functional categories using the MIPS functional catalog. All extracted proteins are in addition automatically analysed and annotated by the PEDANT suite.
Software package to perform phylogeny based association and localization analysis.Used for association detection and localization of susceptibility sites using haplotype phylogenetic trees. Performs these two phylogeny-based analysis: tests association between candidate gene and disease; pinpoints markers (SNPs) that are putative disease susceptibility loci.
It was created in order to create standard datasets on which the performance of machine learning methods can be compared. The collection contains datasets of sequences and structures, each subdivided into positive/negative training/test sets. Such a subdivision is called a classification task. Typical tasks include the classification of structural domains in the SCOP and CATH databases based on their sequences, as fell as various functional and taxonomic classification tasks. Running a performance evaluation test on an entire database can include many different classification tasks. These ensembles of classification tasks are encoded in a simple matrix format - called the cast matrix or membership table - that specifies the role of each sequence (or structure) in the different calculations. Each column of this matrix is a subdivision of the objects (rows) into positive/negative training/test sets. Typically, a database record contains such an ensemble of classification tasks, encoded in a single cast matrix. In addition, there is a collection of distance matrices that contain an all vs. all comparison of the datasets using methods as BLAST, Smith-Waterman, 3D-comparisons etc. Evaluation of a method on a given database consists of calculating a performance measure such as a receiver operating curve (ROC) AUC value. Results of evaluation are deposited along with the data, each dataset is evaluated at least by one classification method, such as 1NN (nearest neighbour) or SVM (support vector machines), ANN (artificial neural networks), RF (random forests) etc.. There are small datasets meant for program developers, as well as downloadable programs for various classification algorithms.
The Bacterial Carbohydrate Structure DataBase is aimed at provision of structural, bibliographic, taxonomic and related information on bacterial carbohydrate structures. It currently requires Internet Explorer in order to function properly. Two key points of this service are: :* covering - is above 95% in the scope of bacterial carbohydrates. This means the negative search answer remains the valuable information too. :* consistence - we manually check the data, and aim at hight quality error-free content The source of data are Carbbank database (University of Georgia, Athens; structures published before 1995, approx. 4000 records) and manual data posting (structures published after 1995, approx. 3000 records). The scope is bacterial carbohydrates and covers nearly all structures of this class published before 2006. Bacterial means that a structure has been found in bacteria or obtained by modification of those found in bacteria. Carbohydrate means a structure composed of any residues linked by glycosidic, ester, amidic, ketal, phospho- or sulpho-diester bonds, in which at least one residue is a sugar or its derivative. Besides the structure itself, each record includes bibliography, abstract, keywords, biological source, methods used to elucidate the structure, bioactivity, NMR assignment tables and a lot of other information. More details, including a format of records, are available at data submission page. You can search the database by IDs, bibliographic data and keywords, biological source, the fragment of structure and NMR data. The substructure search implies either a query language (expert form) or a structure wizard. The database is cross-linked with GlycoSCIENCES DB, which includes all the data from Carbbank (not only bacterial). This means you can search the substructure you entered in GlycoSCIENCES DB, and each record, which contains a structure also present in Carbbank, has a cross-link to data from GlycoSCIENCES DB, including NMR spectra.
Bcipep is collection of the peptides having the role in humoral immunity. The peptides in the database have varying measure of immunogenicity. This database can assist in the development of methods for predicting B cell epitopes, designing synthetic vaccines, and in disease diagnosis. These peptides lead to the generation of antibodies which combine with antigens and are responsible for the host defense, and can be very useful for subunit vaccine designing. The database has 3031 peptide entries. For each peptide, the user can find a plethora of information, including entry number, peptide sequence, pathogen group, protein source, antigen structure, antibody, etc.
THIS RESOURCE IS NO LONGER IN SERVICE, documented on June 10, 2011. BANMOKI is a collection of models of 3D structures of all Bacterial Nucleoside Monophosphate Kinases (NMPK), their pre-computed properties and other relevant information. The web server provides information in regard to sequence identity, 3D structure, Enzyme Commission (EC) number, pKa, desolvation penalty, interaction energy with permanent dipoles and the net charge of folded and unfolded states as a function of pH. The database and datasets are searchable by: seq ID name of protein Enzyme commission number pKa desolvation penalty(in pk units) interaction energy with permanent dipoles(in pk units) No. of basic residues No. of acidic residues No. of titratable residues Total no. of residues The web server development and the corresponding scientific work was supported by NATO collaborative grant CBP.EAP.CLG 981749
THIS RESOURCE IS NO LONGER IN SERVICE, documented on June 04, 2014. Curated database on selected from randomized pools proteins and peptides designed for accumulation of experimental data on protein functionality obtained by in vitro directed evolution methods (phage display, ribosome display, SIP etc.) ASPD is integrated by means of hyperlinks with different databases (SWISS-PROT, PDB, PROSITE, etc). The database also contains modules for pairwise correlation analysis and BLAST search.
Software application for non-parametric linkage analysis using allele sharing in sib pairs (entry from Genetic Analysis Software)
It has been established with the intention of assembling in a central, publicly accessible site information about alternatively spliced genes, their products and expression patterns. Version 2.1 of ASDB consists of two divisions, ASDB(proteins) , which contains amino acid sequences, and ASDB(nucleotides) with genomic sequences.<BR/> SWISS-PROT uses two formats for description of alternative splicing Thus the protein sequences were selected from SWISS-PROT using full text search for both the words alternative splicing (usually in the CC lines) and varsplic (in the FT lines). In order to group proteins that could arise by alternative splicing of the same gene, we developed the clustering procedure. Two proteins were linked if they had a common fragment of at least 20 amino acids, and clusters were initially defined as maximum connected groups of linked proteins. It turned out that some clusters were chimeric, in the sense that they contained members of multi-gene families, but not alternatively spliced variants of one gene. Therefore the multiple alignments were subject to additional analysis aimed at detection of chimeric clusters.<BR/> Each cluster is represented by multiple alignment of its members constructed using CLUSTALW. The distribution of cluster size, representation of species and other relevant statistics of ASDB(proteins) can be accessed through the links below.<BR/> This processing covers the cases when alternatively spliced variants are described in separate SWISS-PROT entries. The other kinds of ASDB records, originating from the SWISS-PROT entries with the varsplic field in the feature table, usually describe the proteins that are not part of any cluster. In these cases, the information on the variable fragments of the several proteins which result from the alternative splicing of a single gene is contained in the entry itself. ASDB(proteins) entries are marked with different symbols to allow for easy differentiation among the three types: those proteins which are part of the ASDB clusters and the corresponding multialignments, those which have the information on different variants in the associated SWISS-PROT entries, and those for which the information on the variants is not available at the present time. ASDB contains internal links between entries and/or clusters, as well as external links to Medline, GenBank and SWISS-PROT entries.<BR/> The ASDB(nucleotides) division was generated by collecting all GenBank entries containing the words alternative splicing and further selection of those entries that contain complete gene sequences (all CDS fields are complete, i.e. they do not have continuation signs).<BR/> Sponsors: This work was supported by the Director, Office of Energy Research, Office of Biological and Environmental Research, of the US Department of Energy under Contract No. DE-ACO3-76SF00098. Additional support came from grants from the Russian Fund of Basic Research (99-04-48347), the Russian State Scientific Program Human Genome (65/99), and the Merck Genome Research Institute (244).<BR/>
This database, AS-ALPS (Alternative Splicing-induced ALteration of Protein Structure), is aimed at providing useful information to analyze effect of AS on protein interaction and network through alteration of protein structure. In AS-ALPS, regions of amino acid sequences changed by AS (AS regions) which are detected in human and mouse transcript sequences in H-InvDB, FANTOM and RefSeq, are linked to information extracted from PDB about residues forming hydrophobic cores and inter-molecular interaction sites. This makes it possible to directly infer whether protein structure and/or interaction are affected by each AS event. In addition, AS-ALPS provides links to a protein network database KEGG, making it easy to know which network and which node in the network can be influenced by AS. :Sponsors: This database was supported by a grant of the Genome Network Project from Ministry of Education, Culture, Sports, Science and Technology of Japan. :
A database is a of mammalian miRNAs and their known or predicted regulatory targets. It provides information on origin of miRNAs, tissue specificity of their expressions and their known or proposed functions, their potential target genes as well as data on miRNA families based on their co-expression and proteins known to be involved in miRNA processing. This database also contains three other navigation tools that can be used to find information relating to miRNA: 1.) Gene Annotations is an information retrieval system for miRNA target genes. It provides comprehensive information from sequence databases and allows to simultaneously search PubMed with all synonyms of a given gene. 2.) miRNA Motif Finder - Argonaute predicts miRNA motifs binding to the gene sequence of the user. The miRNA mature sequences are taken from Agronaute 2 database. miRNA Motif Finder - Custom predicts miRNA motifs binding to the gene sequence, both the gene sequence and miRNA mature sequences provided by the user. 3.) miRNA Statistics provides statistics for the mature miRNA sequences from Argonaute 2 as well as for the miRNA sequences uploaded by the user. It provides statitics on the individual nucleotide as well as pattern of nucleotides apperaing in the sequence.
A database of putative membrane proteins of Thale Cress (Arabidopsis thaliana), Rice (Oryza sativa) and about some 6700 putative membrane proteins of ~300 other seed plants. The database stores data about: * protein, cDNA and genomic sequences * exon predictions (A.thaliana and O.sativa) * different cDNA/protein models of genes (A.thaliana and O.sativa) * ontology terms according to the Gene Ontology (GO) Consortium * protein sequence motifs as predictable by using the PFAM database * transporter classification as predictable by using the TC-system * bibliographic references * predictions for transmembrane spanning proteins (transmembrane alpha helices, beta barrels) * predictions for membrane-anchored proteins (GPI-attachment, prenylation, myristoylation) * prediction of the subcellular location * consensus predictions (transmembrane alpha helices, subcellular location) * isospecic homologs (''paralogs'') * heterospecic homologs (''orthologs'')
Comprehensive catalogue of animal genome size data. Haploid DNA contents (C-values, in picograms) are available for 4972 species (3231 vertebrates and 1741 non-vertebrates) based on 6518 records from 669 published sources. Data may be submitted directly to the database or reprints and notifications of new papers may be sent to database curation staff.
Software application to describe population structure using biomarker data ( typically SNPs, CNVs etc.) available in a population sample. The main features different from PCA are: (1) geometrically motivated and graphic model based; (2)robustness of outliers. (entry from Genetic Analysis Software)
The Annozilla project was designed to view and create annotations associated with a web page, as defined by the W3C Annotea project. Annotations are stored as RDF on a server, using XPointer (or at least XPointer-like constructs) to identify the region of the document being annotated. Additionally, the Annozilla source code is available. The intention of Annozilla is to use Mozilla''''s native facilities to manipulate annotation data - its built-in RDF handling to parse the annotations, and nsIXmlHttpRequest to submit data when creating annotations. To use Annozilla, you will need to install the packages, get set up with a user account with an annotation server (e.g., the W3C test server), and then you should be ready to start using the extension.
AGD is a genome/transcriptome database containing gene annotation and high-density oligonucleotide microarray expression data for protein-coding genes from Ashbya gossypii and the model organism Saccharomyces cerevisiae. It also provides access to comparative genomics data from those two fungi as well as Schizosaccharomyces pombe and Neurospora crassa. Comparative DNA and protein-level data is available, including synteny information in fungi. Additionally, AGD now displays microarray expression data from A.gossypii and S.cerevisiae.
A curated, open-source, web-accessible resource for functional analysis of agricultural plant and animal gene products. Our long-term goal is to serve the needs of the agricultural research communities by facilitating post-genome biology for agriculture researchers and for those researchers primarily using agricultural species as biomedical models. AgBase provides tools designed to assist with the analysis of proteomics data and tools to evaluate experimental datasets using the GO. Additional tools for sequence analysis are also provided. We use controlled vocabularies developed by the Gene Ontology (GO) Consortium to describe molecular function, biological process, and cellular component for genes and gene products in agricultural species. AgBase will also accept annotations from any interested party in the research communities. AgBase develops freely available tools for functional analysis, including tools for using GO. We appreciate any and all questions, comments, and suggestions. AgBase uses the NCBI Blast program for searches for similar sequences. And the Taxonomy Browser allows users to find the NCBI defined taxon ID for or taxon name for different organisms.