Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://biobases.ibch.poznan.pl/5SData/
A database on nucleotide sequences of 5S rRNAs and their genes. The database contains 1985 primary structures of 5S rRNA and 5S rDNA, and was last updated in 2002, according to the website. They include 60 archaebacterial, 470 eubacterial, 63 plastid, nine mitochondrial and 1383 eukaryotic sequences. The nucleotide sequences of the 5S rRNAs or 5S rDNAs are divided according to the taxonomic position of the source organisms. The sequences for particular organisms can be retrieved as single files using a taxonomic browser or in multiple sequence structural alignments. The multiple sequence alignments of 5S ribosomal RNAs can be downloaded in TAB-delimited and FASTA formats.
Proper citation: 5S Ribosomal RNA Database (RRID:SCR_007545) Copy
http://mips.gsf.de/genre/proj/ustilago/
The MIPS Ustilago maydis Genome Database aims to present information on the molecular structure and functional network of the entirely sequenced, filamentous fungus Ustilago maydis. The underlying sequence is the initial release of the high quality draft sequence of the Broad Institute. The goal of the MIPS database is to provide a comprehensive genome database in the Genome Research Environment in parallel with other fungal genomes to enable in depth fungal comparative analysis. The specific aims are to: 1. Generate and assemble Whole Genome Shotgun sequence reads yielding 10X coverage of the U. maydis genome 2. Integrate the genomic sequence assembly with physical maps generated by Bayer CropScience 3. Perform automated annotation of the sequence assembly 4. Align the strain 521 assembly with the FB1 assembly provided by Exelixis 5. Release the sequence assembly and results of our annotation and analysis to public Ustilago maydis is a basidiomycete fungal pathogen of maize and teosinte. The genome size is approximately 20 Mb. The fungus induces tumors on host plants and forms masses of diploid teliospores. These spores germinate and form haploid meiotic products that can be propagated in culture as yeast-like cells. Haploid strains of opposite mating type fuse and form a filamentous, dikaryotic cell type that invades plant tissue to reinitiate infection. Ustilago maydis is an important model system for studying pathogen-host interactions and has been studied for more than 100 years by plant pathologists. Molecular genetic research with U. maydis focuses on recombination, the role of mating in pathogenesis, and signaling pathways that influence virulence. Recently, the fungus has emerged as an excellent experimental model for the molecular genetic analysis of phytopathogenesis, particularly in the characterization of infection-specific morphogenesis in response to signals from host plants. Ustilago maydis also serves as an important model for other basidiomycete plant pathogens that are more difficult to work with in the laboratory, such as the rust and bunt fungi. Genomic sequence of U. maydis will also be valuable for comparative analysis of other fungal genomes, especially with respect to understanding the host range of fungal phytopathogens. The analysis of U. maydis would provide a framework for studying the hundreds of other Ustilago species that attack important crops, such as barley, wheat, sorghum, and sugarcane. Comparisons would also be possible with other basidiomycete fungi, such as the important human pathogen C. neoformans. Commercially, U. maydis is an excellent model for the discovery of antifungal drugs. In addition, maize tumors caused by U. maydis are prized in Hispanic cuisine and there is interest in improving commercial production. The complete putative gene set of the Broad Institute''s second release is loaded into the database and in addition all deviating putative genes from a putative gene set produced by MIPS with different gene prediction parameters are also loaded. The complete dataset will then be analysed, gene predictions will be manually corrected due to combined information derived from different gene prediction algorithms and, more important, protein and EST comparisons. Gene prediction will be restricted to ORFs larger than 50 codons; smaller ORFs will be included only if similarities to other proteins or EST matches confirm their existence or if a coding region was postulated by all prediction programs used. The resulting proteins will be annotated. They will be classified according to the MIPS classification catalogue receiving appropriate descriptions. All proteins with a known, characterized homolog will be automatically assigned to functional categories using the MIPS functional catalog. All extracted proteins are in addition automatically analysed and annotated by the PEDANT suite.
Proper citation: MIPS Ustilago maydis Database (RRID:SCR_007563) Copy
Collection of data of protein sequence and functional information. Resource for protein sequence and annotation data. Consortium for preservation of the UniProt databases: UniProt Knowledgebase (UniProtKB), UniProt Reference Clusters (UniRef), and UniProt Archive (UniParc), UniProt Proteomes. Collaboration between European Bioinformatics Institute (EMBL-EBI), SIB Swiss Institute of Bioinformatics and Protein Information Resource. Swiss-Prot is a curated subset of UniProtKB.
Proper citation: UniProt (RRID:SCR_002380) Copy
http://www.ebi.ac.uk/swissprot/hpi/hpi.html
THIS RESOURCE IS NO LONGER IN SERVICE, documented on August 03, 2011. IT HAS BEEN REPLACED BY A NEW UniProtKB/Swiss-Prot ANNOTATION PROGRAM CALLED UniProt Chordata protein annotation program. The Human Proteome Initiative (HPI) aims to annotate all known human protein sequences, as well as their orthologous sequences in other mammals, according to the quality standards of UniProtKB/Swiss-Prot. In addition to accurate sequences, we strive to provide, for each protein, a wealth of information that includes the description of its function, domain structure, subcellular location, similarities to other proteins, etc. Although as complete as currently possible, the human protein set they provide is still imperfect, it will have to be reviewed and updated with future research results. They will also create entries for newly discovered human proteins, increase the number of splice variants, explore the full range of post-translational modifications (PTMs) and continue to build a comprehensive view of protein variation in the human population. The availability of the human genome sequence has enabled the exploration and exploitation of the human genome and proteome to begin. Research has now focused on the annotation of the genome and in particular of the proteome. With expert annotation extracted from the literature by biologists as the foundation, it has been possible to expand into the areas of data mining and automatic annotation. With further development and integration of pattern recognition methods and the application of alignments clustering, proteome analysis can now be provided in a meaningful way. These various approaches have been integrated to attach, extract and combine as much relevant information as possible to the proteome. This resource should be valuable to users from both research and industry. We maintain a file containing all human UniProtKB/Swiss-Prot entries. This file is updated at every biweekly release of UniProt and can be downloaded by FTP download, HTTP download or by using a mirroring program which automatically retrieves the file at regular intervals.
Proper citation: Human Proteomics Initiative (RRID:SCR_002373) Copy
http://www.ncbi.nlm.nih.gov/ieb/research/acembly/
THIS RESOURCE IS NO LONGER IN SERVICE, documented May 10, 2017. A pilot effort that has developed a centralized, web-based biospecimen locator that presents biospecimens collected and stored at participating Arizona hospitals and biospecimen banks, which are available for acquisition and use by researchers. Researchers may use this site to browse, search and request biospecimens to use in qualified studies. The development of the ABL was guided by the Arizona Biospecimen Consortium (ABC), a consortium of hospitals and medical centers in the Phoenix area, and is now being piloted by this Consortium under the direction of ABRC. You may browse by type (cells, fluid, molecular, tissue) or disease. Common data elements decided by the ABC Standards Committee, based on data elements on the National Cancer Institute''s (NCI''s) Common Biorepository Model (CBM), are displayed. These describe the minimum set of data elements that the NCI determined were most important for a researcher to see about a biospecimen. The ABL currently does not display information on whether or not clinical data is available to accompany the biospecimens. However, a requester has the ability to solicit clinical data in the request. Once a request is approved, the biospecimen provider will contact the requester to discuss the request (and the requester''s questions) before finalizing the invoice and shipment. The ABL is available to the public to browse. In order to request biospecimens from the ABL, the researcher will be required to submit the requested required information. Upon submission of the information, shipment of the requested biospecimen(s) will be dependent on the scientific and institutional review approval. Account required. Registration is open to everyone., documented August 29, 2016. AceView offers an integrated view of the human, nematode and Arabidopsis genes reconstructed by co-alignment of all publicly available mRNAs and ESTs on the genome sequence. Our goals are to offer a reliable up-to-date resource on the genes and their functions and to stimulate further validating experiments at the bench. AceView provides a curated, comprehensive and non-redundant sequence representation of all public mRNA sequences (mRNAs from GenBank or RefSeq, and single pass cDNA sequences from dbEST and Trace). These experimental cDNA sequences are first co-aligned on the genome then clustered into a minimal number of alternative transcript variants and grouped into genes. Using exhaustively and with high quality standards the available cDNA sequences evidences the beauty and complexity of mammals' transcriptome, and the relative simplicity of the nematode and plant transcriptomes. Genes are classified according to their inferred coding potential; many presumably non-coding genes are discovered. Genes are named by Entrez Gene names when available, else by AceView gene names, stable from release to release. Alternative features (promoters, introns and exons, polyadenylation signals) and coding potential, including motifs, domains, and homologies are annotated in depth; tissues where expression has been observed are listed in order of representation; diseases, phenotypes, pathways, functions, localization or interactions are annotated by mining selected sources, in particular PubMed, GAD and Entrez Gene, and also by performing manual annotation, especially in the worm. In this way, both the anatomy and physiology of the experimentally cDNA supported human, mouse and nematode genes are thoroughly annotated. Our goals are to offer an up-to-date resource on the genes, in the hope to stimulate further experiments at the bench, or to help medical research. AceView can be queried by meaningful words or groups of words as well as by most standard identifiers, such as gene names, Entrez Gene ID, UniGene ID, GenBank accessions.
Proper citation: AceView (RRID:SCR_002277) Copy
http://fullmal.hgc.jp/index_ajax.html
FULL-malaria is a database for a full-length-enriched cDNA library from the human malaria parasite Plasmodium falciparum. Because of its medical importance, this organism is the first target for genome sequencing of a eukaryotic pathogen; the sequences of two of its 14 chromosomes have already been determined. However, for the full exploitation of this rapidly accumulating information, correct identification of the genes and study of their expression are essential. Using the oligo-capping method, this database has produced a full-length-enriched cDNA library from erythrocytic stage parasites and performed one-pass reading. The database consists of nucleotide sequences of 2490 random clones that include 390 (16%) known malaria genes according to BLASTN analysis of the nr-nt database in GenBank; these represent 98 genes, and the clones for 48 of these genes contain the complete protein-coding sequence (49%). On the other hand, comparisons with the complete chromosome 2 sequence revealed that 35 of 210 predicted genes are expressed, and in addition led to detection of three new gene candidates that were not previously known. In total, 19 of these 38 clones (50%) were full-length. From these observations, it is expected that the database contains approximately 1000 genes, including 500 full-length clones. It should be an invaluable resource for the development of vaccines and novel drugs. Full-malaria has been updated in at least three points. (i) 8934 sequences generated from the addition of new libraries added so that the database collection of 11,424 full-length cDNAs covers 1375 (25%) of the estimated number of the entire 5409 parasite genes. (ii) All of its full-length cDNAs and GenBank EST sequences were mapped to genomic sequences together with publicly available annotated genes and other predictions. This precisely determined the gene structures and positions of the transcriptional start sites, which are indispensable for the identification of the promoter regions. (iii) A total of 4257 cDNA sequences were newly generated from murine malaria parasites, Plasmodium yoelii yoelii. The genome/cDNA sequences were compared at both nucleotide and amino acid levels, with those of P.falciparum, and the sequence alignment for each gene is presented graphically. This part of the database serves as a versatile platform to elucidate the function(s) of malaria genes by a comparative genomic approach. It should also be noted that all of the cDNAs represented in this database are supported by physical cDNA clones, which are publicly and freely available, and should serve as indispensable resources to explore functional analyses of malaria genomes. Sponsors: This database has been constructed and maintained by a Grant-in-Aid for Publication of Scientific Research Results from the Japan Society for the Promotion of Science (JSPS). This work was also supported by a Special Coordination Funds for Promoting Science and Technology from the Science and Technology Agency of Japan (STA) and a Grant-in-Aid for Scientific Research on Priority Areas from the Ministry of Education, Science, Sports and Culture of Japan.
Proper citation: Full-Malaria: Malaria Full-Length cDNA Database (RRID:SCR_002348) Copy
http://gladyshevlab.org/SelenoproteinPredictionServer/
Web server to predict eukaryotic selenoproteins and SECIS (SElenoCysteine Insertion Sequences) elements along nucleotide sequences. SECISearch3 replaces its predecessor SECISearch as a tool for prediction of eukaryotic SECIS elements. Seblastian is a method for selenoprotein gene detection that uses SECISearch3 and then predicts selenoprotein sequences encoded upstream of SECIS elements. Seblastian is able to both identify known selenoproteins and predict new selenoproteins.
Proper citation: SECISearch3 and Seblastian (RRID:SCR_003186) Copy
A database of three-dimensional structural information about nucleic acids and their complexes. In addition to primary data, it contains derived geometric data, classifications of structures and motifs, standards for describing nucleic acid features, as well as tools and software for the analysis of nucleic acids. A variety of search capabilities are available, as are many different types of reports. NDB maintains the macromolecular Crystallographic Information File (mmCIF).
Proper citation: Nucleic Acid Database (RRID:SCR_003255) Copy
http://www.ncbi.nlm.nih.gov/taxonomy/
Database for a curated classification and nomenclature that contains the names of all organisms that are represented in the public sequence databases with at least one nucleotide or protein sequence. Data provided encompasses archaea, bacteria, eukaryota, viroids and viruses. The NCBI taxonomy database is not a primary source for taxonomic or phylogenetic information. Furthermore, the database does not follow a single taxonomic treatise but rather attempts to incorporate phylogenetic and taxonomic knowledge from a variety of sources, including the published literature, web-based databases, and the advice of sequence submitters and outside taxonomy experts. Consequently, the NCBI taxonomy database is not a phylogenetic or taxonomic authority and should not be cited as such.
Proper citation: NCBI Taxonomy (RRID:SCR_003256) Copy
A database of human mitochondrial genomes containing mtDNA sequences, polymorphic sites, and the ability to search for specific variants. It contains 1865 complete sequences and 839 coding region sequences.
Proper citation: mtDB - Human Mitochondrial Genome Database (RRID:SCR_002945) Copy
http://bioinfo.mbi.ucla.edu/ASAP/
THIS RESOURCE IS NO LONGER IN SERVICE, documented on 8/12/13. Database to access and mine alternative splicing information coming from genomics and proteomics based on genome-wide analyses of alternative splicing in human (30 793 alternative splice relationships found) from detailed alignment of expressed sequences onto the genomic sequence. ASAP provides precise gene exon-intron structure, alternative splicing, tissue specificity of alternative splice forms, and protein isoform sequences resulting from alternative splicing. They developed an automated method for discovering human tissue-specific regulation of alternative splicing through a genome-wide analysis of expressed sequence tags (ESTs), which involves classifying human EST libraries according to tissue categories and Bayesian statistical analysis. They use the UniGene clusters of human Expressed Sequence Tags (ESTs) to identify splices. The UniGene EST's are clustered so that a single cluster roughly corresponds to a gene (or at least a part of a gene). A single EST represents a portion of a processed (already spliced) mRNA. A given cluster contains many ESTs, each representing an outcome of a series of splicing events. The ESTs in UniGene contain the different mRNA isoforms transcribed from an alternatively spliced gene. They are not predicting alternative splicing, but locating it based on EST analysis. The discovered splices are further analyzed to determine alternative splicing events. They have identified 6201 alternative splice relationships in human genes, through a genome-wide analysis of expressed sequence tags (ESTs). Starting with 2.1 million human mRNA and EST sequences, they mapped expressed sequences onto the draft human genome sequence and only accepted splices that obeyed the standard splice site consensus. After constructing a tissue list of 46 human tissues with 2 million human ESTs, they generated a database of novel human alternative splices that is four times larger than our previous report, and used Bayesian statistics to compare the relative abundance of every pair of alternative splices in these tissues. Using several statistical criteria for tissue specificity, they have identified 667 tissue-specific alternative splicing relationships and analyzed their distribution in human tissues. They have validated our results by comparison with independent studies. This genome-wide analysis of tissue specificity of alternative splicing will provide a useful resource to study the tissue-specific functions of transcripts and the association of tissue-specific variants with human diseases.
Proper citation: ASAP: the Alternative Splicing Annotation Project (RRID:SCR_003415) Copy
Database of polymorphisms and mutations of the human mitochondrial DNA. It reports published and unpublished data on human mitochondrial DNA variation. All data is curated by hand. If you would like to submit published articles to be included in mitomap, please send them the citation and a pdf.
Proper citation: MITOMAP - A human mitochondrial genome database (RRID:SCR_002996) Copy
http://caps.ncbs.res.in/3dswap/index.html
Curated knowledegbase of protein structures that are reported to be involved in 3-dimensional domain swapping. 3DSwap provides literature curated information and structure related information about 3D domain swapping in proteins. Information about swapping, hinge region, swapped region, extent of swapping, etc. are extracted from original research publications after extensive literature curation.
Proper citation: 3DSwap (RRID:SCR_004133) Copy
http://www.hgsc.bcm.tmc.edu/content/hapmap-3-and-encode-3
Draft release 3 for genome-wide SNP genotyping and targeted sequencing in DNA samples from a variety of human populations (sometimes referred to as the HapMap 3 samples). This release contains the following data: * SNP genotype data generated from 1184 samples, collected using two platforms: the Illumina Human1M (by the Wellcome Trust Sanger Institute) and the Affymetrix SNP 6.0 (by the Broad Institute). Data from the two platforms have been merged for this release. * PCR-based resequencing data (by Baylor College of Medicine Human Genome Sequencing Center) across ten 100-kb regions (collectively referred to as ENCODE 3) in 712 samples. Since this is a draft release, please check this site regularly for updates and new releases. The HapMap 3 sample collection comprises 1,301 samples (including the original 270 samples used in Phase I and II of the International HapMap Project) from 11 populations, listed below alphabetically by their 3-letter labels. Five of the ten ENCODE 3 regions overlap with the HapMap-ENCODE regions; the other five are regions selected at random from the ENCODE target regions (excluding the 10 HapMap-ENCODE regions). All ENCODE 3 regions are 100-kb in size, and are centered within each respective ENCODE region. The HapMap 3 and ENCORE 3 data are downloadable from the ftp site.
Proper citation: HapMap 3 and ENCODE 3 (RRID:SCR_004563) Copy
A database of protein families, each represented by multiple sequence alignments and hidden Markov models (HMMs). Users can analyze protein sequences for Pfam matches, view Pfam family annotation and alignments, see groups of related families, look at the domain organization of a protein sequence, find the domains on a PDB structure, and query Pfam by keywords. There are two components to Pfam: Pfam-A and Pfam-B. Pfam-A entries are high quality, manually curated families that may automatically generate a supplement using the ADDA database. These automatically generated entries are called Pfam-B. Although of lower quality, Pfam-B families can be useful for identifying functionally conserved regions when no Pfam-A entries are found. Pfam also generates higher-level groupings of related families, known as clans (collections of Pfam-A entries which are related by similarity of sequence, structure or profile-HMM).
Proper citation: Pfam (RRID:SCR_004726) Copy
http://www.ncbi.nlm.nih.gov/mapview/map_search.cgi?taxid=7165
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on January 11, 2023. A database for the Anopheles gambiae str. PEST genome that was sequenced using a whole genome shotgun approach. The database aims to contribute to the understanding of mosquito genome structure and organization and will assist the development of malaria control strategies and improved anti-malarial drugs and vaccines. Sequences were generated and assembled into contigs for submission to GenBank.
Proper citation: Anopheles gambiae (African malaria mosquito) genome view (RRID:SCR_004402) Copy
http://www.uniprot.org/taxonomy/
NEWT is the taxonomy database maintained by the UniProt group. It integrates taxonomy data compiled in the NCBI database and data specific to the UniProt Knowledgebase. Browse by hierarchy, List all, or Complete proteomes. Organisms are classified in a hierarchical tree structure. Our taxonomy database contains every node (taxon) of the tree. UniProtKB taxonomy data is manually curated: next to manually verified organism names, we provide a selection of external links, organism strains and viral host information. Species with protein sequences stored in the UniProt Knowledgebase are named according to UniProt nomenclature. We endeavour to maintain a list of manually curated species names for which protein sequence data is available. In particular, we have adopted a systematic convention for naming viral and bacterial strains and isolates. Links to external sites are chosen by the UniProt taxonomy team and show pictures and various scientific data of interest (taxonomy, biology, physiology,...).
Proper citation: NEWT (RRID:SCR_004477) Copy
http://amphoranet.pitgroup.org/
Webserver implementation of the AMPHORA2 workflow for phylogenetic analysis of metagenomic shotgun sequencing data. It is capable of assigning a probability-weighted taxonomic group for each phylogenetic marker gene found in the input metagenomic sample.
Proper citation: AmphoraNet (RRID:SCR_005009) Copy
http://www.ebi.ac.uk/biosamples/
Database that aggregates sample information for reference samples (e.g. Coriell Cell lines) and samples for which data exist in one of the EBI''''s assay databases such as ArrayExpress, the European Nucleotide Archive or PRoteomics Identificates DatabasE. It provides links to assays for specific samples, and accepts direct submissions of sample information. The goals of the BioSample Database include: # recording and linking of sample information consistently within EBI databases such as ENA, ArrayExpress and PRIDE; # minimizing data entry efforts for EBI database submitters by enabling submitting sample descriptions once and referencing them later in data submissions to assay databases and # supporting cross database queries by sample characteristics. The database includes a growing set of reference samples, such as cell lines, which are repeatedly used in experiments and can be easily referenced from any database by their accession numbers. Accession numbers for the reference samples will be exchanged with a similar database at NCBI. The samples in the database can be queried by their attributes, such as sample types, disease names or sample providers. A simple tab-delimited format facilitates submissions of sample information to the database, initially via email to biosamples (at) ebi.ac.uk. Current data sources: * European Nucleotide Archive (424,811 samples) * PRIDE (17,001 samples) * ArrayExpress (1,187,884 samples) * ENCODE cell lines (119 samples) * CORIELL cell lines (27,002 samples) * Thousand Genome (2,628 samples) * HapMap (1,417 samples) * IMSR (248,660 samples)
Proper citation: BioSample Database at EBI (RRID:SCR_004856) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on January 11,2023. SuperCAT hosts typing databases for the Bacillus cereus group of bacteria. The databases contain MultiLocus Sequence Typing (MLST), MultiLocus Enzyme Electrophoresis (MLEE), and Amplified Fragment Length Polymorphism (AFLP) phylogenetic data. multilocus, sequence, Bacillus cereus, bacteria, Genomics, non-vertebrate, taxonomy, identification
Proper citation: SuperCAT (RRID:SCR_004882) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the Kravitz Resources search. From here you can search through a compilation of resources used by Kravitz and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that Kravitz has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on Kravitz then you can log in from here to get additional features in Kravitz such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into Kravitz you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within Kravitz that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.