We support boolean queries, use +,-,<,>,~,* to alter the weighting of terms
Data repository to preserve and provide access to agricultural and environmental data produced during research projects undertaken at the University of Guelph including datasets on topics such as crop yield, soil moisture, weather and agroforestry. A special emphasis is placed on research funded by Ontario Ministry of Agriculture and Food (OMAF) and MRA.
Service providing functional analysis of proteins by classifying them into families and predicting domains and important sites. They combine protein signatures from a number of member databases into a single searchable resource, capitalizing on their individual strengths to produce a powerful integrated database and diagnostic tool. This integrated database of predictive protein signatures is used for the classification and automatic annotation of proteins and genomes. InterPro classifies sequences at superfamily, family and subfamily levels, predicting the occurrence of functional domains, repeats and important sites. InterPro adds in-depth annotation, including GO terms, to the protein signatures. You can access the data programmatically, via Web Services. The member databases use a number of approaches: # ProDom: provider of sequence-clusters built from UniProtKB using PSI-BLAST. # PROSITE patterns: provider of simple regular expressions. # PROSITE and HAMAP profiles: provide sequence matrices. # PRINTS provider of fingerprints, which are groups of aligned, un-weighted Position Specific Sequence Matrices (PSSMs). # PANTHER, PIRSF, Pfam, SMART, TIGRFAMs, Gene3D and SUPERFAMILY: are providers of hidden Markov models (HMMs). Your contributions are welcome. You are encouraged to use the ''''Add your annotation'''' button on InterPro entry pages to suggest updated or improved annotation for individual InterPro entries.
Curated protein-protein and genetic interaction repository of raw protein and genetic interactions from major model organism species, with data compiled through comprehensive curation efforts.
Non profit research organization for genome sequences to advance understanding of biology of humans and pathogens in order to improve human health globally. Provides data which can be translated for diagnostics, treatments or therapies including over 100 finished genomes, which can be downloaded. Data are publicly available on limited basis, and provided more extensively upon request.
International, curated, digital repository that makes the data underlying scientific publications discoverable, freely reusable, and citable. Particularly data for which no specialized repository exists. Provides the infrastructure for, and promotes the re-use of, data underlying the scholarly literature. Governed by a nonprofit membership organization. Membership is open to any stakeholder organization, including but not limited to journals, scientific societies, publishers, research institutions, libraries, and funding organizations. Most data are associated with peer-reviewed articles, although data associated with non-peer reviewed publications from reputable academic sources, such as dissertations, are also accepted. Used to validate published findings, explore new analysis methodologies, repurpose data for research questions unanticipated by the original authors, and perform synthetic studies.UC system is member organization of Dryad general subject data repository.
Freely available database focused on interactions established by extracellular proteins and polysaccharides, taking into account the multimeric nature of the extracellular proteins (e.g. collagens, laminins and thrombospondins are multimers). MatrixDB is an active member of the International Molecular Exchange (IMEx) consortium and has adopted the PSI-MI standards for annotating and exchanging interaction data. It includes interaction data extracted from the literature by manual curation, and offers access to relevant data involving extracellular proteins provided by the IMEx partner databases through the PSICQUIC webservice, as well as data from the Human Protein Reference Database. The database reports mammalian protein-protein and protein-carbohydrate interactions involving extracellular molecules. Interactions with lipids and cations are also reported. MatrixDB is focused on mammalian interactions, but aims to integrate interaction datasets of model organisms when available. MatrixDB provides direct links to databases recapitulating mutations in genes encoding extracellular proteins, to UniGene and to the Human Protein Atlas that shows expression and localization of proteins in a large variety of normal human tissues and cells. MatrixDB allows researchers to perform customized queries and to build tissue- and disease-specific interaction networks that can be visualized and analyzed with Cytoscape or Medusa. Statistics (2013): 2283 extracellular matrix interactions including 2095 protein-protein and 169 protein-glycosaminoglycan interactions.
General database of genetic variations maintained by the NCBI. Database as central repository for both single base nucleotide substitutions and short deletion and insertion polymorphisms. Distinguishes report of how to assay SNP from use of that SNP with individuals and populations. This separation simplifies some issues of data representation. However, these initial reports describing how to assay SNP will often be accompanied by SNP experiments measuring allele occurrence in individuals and populations. Community can contribute to this resource.
NIH genetic sequence database that provides annotated collection of all publicly available DNA sequences for almost 280 000 formally described species (Jan 2014) .These sequences are obtained primarily through submissions from individual laboratories and batch submissions from large-scale sequencing projects, including whole-genome shotgun (WGS) and environmental sampling projects. Most submissions are made using web-based BankIt or standalone Sequin programs, and GenBank staff assigns accession numbers upon data receipt. It is part of International Nucleotide Sequence Database Collaboration and daily data exchange with European Nucleotide Archive (ENA) and DNA Data Bank of Japan (DDBJ) ensures worldwide coverage. GenBank is accessible through NCBI Entrez retrieval system, which integrates data from major DNA and protein sequence databases along with taxonomy, genome, mapping, protein structure and domain information, and biomedical journal literature via PubMed. BLAST provides sequence similarity searches of GenBank and other sequence databases. Complete bimonthly releases and daily updates of GenBank database are available by FTP.
Database that summarizes diverse information about metabolic pathways in crop plants and allows automatic export of information for the creation of detailed metabolic models. It contains manually curated, highly detailed information about metabolic pathways in crop plants, including pathway diagrams, reactions, locations, transport processes, reaction kinetics, taxonomy and literature. It contains information about seven major crop plants with high agronomical importance and two model plants.
Database that catalogs experimentally verified pathogenicity, virulence and effector genes from fungal, Oomycete and bacterial pathogens, which infect animal, plant, fungal and insect hosts. It is an invaluable resource in the discovery of genes in medically and agronomically important pathogens, which may be potential targets for chemical intervention. In collaboration with the FRAC team, it also includes antifungal compounds and their target genes. Each entry is curated by domain experts and is supported by strong experimental evidence (gene disruption experiments, STM etc), as well as literature references in which the original experiments are described. Each gene is presented with its nucleotide and deduced amino acid sequence, as well as a detailed description of the predicted protein's function during the host infection process. To facilitate data interoperability, genes have been annotated using controlled vocabularies and links to external sources (Gene Ontology terms, EC Numbers, NCBI taxonomy, EMBL, PubMed and FRAC).
Collection of information about chemical structures and biological properties of small molecules and siRNA reagents hosted by the National Center for Biotechnology Information (NCBI).
Public repository that accepts direct submissions and provides archiving, accessioning and distribution of publicly available genomic structural variants, in all species. Variants are accessioned at the study and sample level, granting stable identifiers that can be used in publications. DGVa data is integrated with other EBI resources, including comprehensive EBI search and Ensembl genome browser. Exchanges data with companion database, dbVar, at National Center for Biotechnology Information.NOTE: since 2019 DGVa doesn't accept submissions. Please send the data for submission to European Variation Archive (EVA).
A multidisciplinary repository of public data sets such as the Human Genome and US Census data that can be seamlessly integrated into AWS cloud-based applications. AWS is hosting the public data sets at no charge for the community. Anyone can access these data sets from their Amazon Elastic Compute Cloud (Amazon EC2) instances and start computing on the data within minutes. Users can also leverage the entire AWS ecosystem and easily collaborate with other AWS users. If you have a public domain or non-proprietary data set that you think is useful and interesting to the AWS community, please submit a request and the AWS team will review your submission and get back to you. Typically the data sets in the repository are between 1 GB to 1 TB in size (based on the Amazon EBS volume limit), but they can work with you to host larger data sets as well. You must have the right to make the data freely available.
A non-profit university-governed consortium that facilitates geoscience research and education using geodesy. It rovides access to and submission of Geodetic GPS / GNSS Data, Geodetic Imaging Data, Strain and Seismic Borehole Data, and Meteorological Data. Data access web services/API provides the ability to use a command line interface to query metadata and obtain URLs to data and products. UNAVCO also provides a variety of software, including web applications, and desktop utilities for scientists, instructors, students, and others. Web-based data visualization and mapping tools provide users with the ability to view postprocessed data while web-based geodetic utilities provide ancillary information. Downloadable stand-alone software utilities include applications for configuring instruments, managing data collection, download and transfer, and performing computations on the raw data, e.g., data pre-processing or processing. The UNAVCO Facility in Boulder, Colorado is the primary operational activity of UNAVCO and exists to support university and other research investigators in their use of geophysical sensor technology for Earth sciences research. The Facility performs this task in part by archiving GNSS/GPS data and data products for current and future applications. Other data types that scientists use for Earth deformation studies are also held in the UNAVCO Archive collections. UNAVCO operates a community Archive, which provides long-term secure storage and easy retrieval of GNSS data, strain data, various derived products and related metadata. The Archive primarily stores high-precision geodetic data used for research purposes, collected under National Science Foundation and NASA sponsored projects. UNAVCO provides many learning opportunities including: Short Courses and Workshops, Educational Resources, RESESS Research Student Internships, and Technical Training.
Registry for Familial Mediterranean Fever (FMF) and hereditary inflammatory disorders mutations. As of 2014, it includes twenty genes including: MEFV, MVK, TNFRSF1A, NLRP3, NOD2, PSTPIP1, LPIN2 and NLRP7, and contains over 1338 sequence variants. Confidential data, simple and complex alleles are accepted. For each gene, a menu offers: 1) a tabular list of the variants that can be sorted by several parameters; 2) a gene graph providing a schematic representation of the variants along the gene; 3) statistical analysis of the data according to the phenotype, alteration type, and location of the mutation in the gene; 4) the cDNA and gDNA sequences of each gene, showing the nucleotide changes along the sequence, with a color-based code highlighting the gene domains, the first ATG, and the termination codon; and 5) a download menu making all tables and figures available for the users, which, except for the gene graphs, are all automatically generated and updated upon submission of the variants. The entire database was curated to comply with the HUGO Gene Nomenclature Committee (HGNC) and HGVS nomenclature guidelines, and wherever necessary, an informative note was provided.
Collection of structural data of biological macromolecules. Database of information about 3D structures of large biological molecules, including proteins and nucleic acids. Users can perform queries on data and analyze and visualize results.
Database that gathers, generates, and shares taxa, images, videos, and sounds to freely provide knowledge about life on earth to increase awareness and understanding of living nature. Free EOL memberships are ranked so members have greater authority and editorial abilities based on their level of expertise.
Database of trait mapping data, i.e. QTL (phenotype / expression, eQTL), candidate gene and association data (GWAS) and copy number variations (CNV) mapped to livestock animal genomes, to facilitate locating and comparing discoveries within and between species. New data and database tools are continually developed to align various trait mapping data to map-based genome features, such as annotated genes. QTLdb is open to house QTL/association date from other animal species where feasible. Most scientific journals require that any original QTL/association data be deposited into public databases before paper may be accepted for publication. User curator accounts are provided for direct data deposit. Users can download QTLdb data from each species or individual chromosome.
Collection of genome databases for vertebrates and other eukaryotic species with DNA and protein sequence search capabilities. Used to automatically annotate genome, integrate this annotation with other available biological data and make data publicly available via web. Ensembl tools include BLAST, BLAT, BioMart and the Variant Effect Predictor (VEP) for all supported species.
Cross-species microarray expression database focusing on high-throughput expression data relevant for germline development, meiosis and gametogenesis as well as the mitotic cell cycle. The database contains a unique combination of information: 1) High-throughput expression data obtained with whole-genome high-density oligonucleotide microarrays (GeneChips). 2) Sample annotation (mouse over the sample name and click on it) using the Multiomics Information Management and Annotation System (MIMAS 3.0). 3) In vivo protein-DNA binding data and protein-protein interaction data (available for selected species). 4) Genome annotation information from Ensembl version 50. 5) Orthologs are identified using data from Ensembl and OMA and linked to each other via a section in the report pages. The portal provides access to the Saccharomyces Genomics Viewer (SGV) which facilitates online interpretation of complex data from experiments with high-density oligonucleotide tiling microarrays that cover the entire yeast genome. The database displays only expression data obtained with high-density oligonucleotide microarrays (GeneChips)., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on January 15,2026.