Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://biobases.ibch.poznan.pl/5SData/
A database on nucleotide sequences of 5S rRNAs and their genes. The database contains 1985 primary structures of 5S rRNA and 5S rDNA, and was last updated in 2002, according to the website. They include 60 archaebacterial, 470 eubacterial, 63 plastid, nine mitochondrial and 1383 eukaryotic sequences. The nucleotide sequences of the 5S rRNAs or 5S rDNAs are divided according to the taxonomic position of the source organisms. The sequences for particular organisms can be retrieved as single files using a taxonomic browser or in multiple sequence structural alignments. The multiple sequence alignments of 5S ribosomal RNAs can be downloaded in TAB-delimited and FASTA formats.
Proper citation: 5S Ribosomal RNA Database (RRID:SCR_007545) Copy
http://mips.gsf.de/genre/proj/ustilago/
The MIPS Ustilago maydis Genome Database aims to present information on the molecular structure and functional network of the entirely sequenced, filamentous fungus Ustilago maydis. The underlying sequence is the initial release of the high quality draft sequence of the Broad Institute. The goal of the MIPS database is to provide a comprehensive genome database in the Genome Research Environment in parallel with other fungal genomes to enable in depth fungal comparative analysis. The specific aims are to: 1. Generate and assemble Whole Genome Shotgun sequence reads yielding 10X coverage of the U. maydis genome 2. Integrate the genomic sequence assembly with physical maps generated by Bayer CropScience 3. Perform automated annotation of the sequence assembly 4. Align the strain 521 assembly with the FB1 assembly provided by Exelixis 5. Release the sequence assembly and results of our annotation and analysis to public Ustilago maydis is a basidiomycete fungal pathogen of maize and teosinte. The genome size is approximately 20 Mb. The fungus induces tumors on host plants and forms masses of diploid teliospores. These spores germinate and form haploid meiotic products that can be propagated in culture as yeast-like cells. Haploid strains of opposite mating type fuse and form a filamentous, dikaryotic cell type that invades plant tissue to reinitiate infection. Ustilago maydis is an important model system for studying pathogen-host interactions and has been studied for more than 100 years by plant pathologists. Molecular genetic research with U. maydis focuses on recombination, the role of mating in pathogenesis, and signaling pathways that influence virulence. Recently, the fungus has emerged as an excellent experimental model for the molecular genetic analysis of phytopathogenesis, particularly in the characterization of infection-specific morphogenesis in response to signals from host plants. Ustilago maydis also serves as an important model for other basidiomycete plant pathogens that are more difficult to work with in the laboratory, such as the rust and bunt fungi. Genomic sequence of U. maydis will also be valuable for comparative analysis of other fungal genomes, especially with respect to understanding the host range of fungal phytopathogens. The analysis of U. maydis would provide a framework for studying the hundreds of other Ustilago species that attack important crops, such as barley, wheat, sorghum, and sugarcane. Comparisons would also be possible with other basidiomycete fungi, such as the important human pathogen C. neoformans. Commercially, U. maydis is an excellent model for the discovery of antifungal drugs. In addition, maize tumors caused by U. maydis are prized in Hispanic cuisine and there is interest in improving commercial production. The complete putative gene set of the Broad Institute''s second release is loaded into the database and in addition all deviating putative genes from a putative gene set produced by MIPS with different gene prediction parameters are also loaded. The complete dataset will then be analysed, gene predictions will be manually corrected due to combined information derived from different gene prediction algorithms and, more important, protein and EST comparisons. Gene prediction will be restricted to ORFs larger than 50 codons; smaller ORFs will be included only if similarities to other proteins or EST matches confirm their existence or if a coding region was postulated by all prediction programs used. The resulting proteins will be annotated. They will be classified according to the MIPS classification catalogue receiving appropriate descriptions. All proteins with a known, characterized homolog will be automatically assigned to functional categories using the MIPS functional catalog. All extracted proteins are in addition automatically analysed and annotated by the PEDANT suite.
Proper citation: MIPS Ustilago maydis Database (RRID:SCR_007563) Copy
Collection of transmembrane protein datasets containing experimentally derived topology information from the literature and from public databases. Web interface of TOPDB includes tools for searching, relational querying and data browsing, visualisation tools for topology data.
Proper citation: Topology Data Bank of Transmembrane Proteins (RRID:SCR_007964) Copy
It provides a database based on a pre-computed similarity matrix covering the similarity space formed by >4 million amino acid sequences from public databases and completely sequenced genomes. The database is capable of handling very large datasets and is updated incrementally. For sequence similarity searches and pairwise alignments, we implemented a grid-enabled software system, which is based on FASTA heuristics and the Smith Waterman algorithm. SimpleSIMAP and AdvancedSIMAP retrieve homologs for given protein sequences that need to be contained in the SIMAP database. While SimpleSIMAP provides only selected parameters and preconfigured search spaces, the AdvancedSIMAP allows the user to specify search space, filtering and sorting parameters in a flexible manner. Both types of queries result in lists of homologs that are linked in turn to their homologs. So the web interfaces allow users to explore quickly and interactively the protein world by homology. Sponsors: SIMAP is supported by the Department of Genome Oriented Bioinformatics of the Technische Universitt Mnchen and the Institute for Bioinformatics of the GSF-National Research Center for Environment and Health.
Proper citation: SIMAP (RRID:SCR_007927) Copy
http://www.grt.kyushu-u.ac.jp/spad/
It is divided to four categories based on extracellular signal molecules (Growth factor, Cytokine, and Hormone) and stress, that initiate the intracellular signaling pathway. SPAD is compiled in order to describe information on interaction between protein and protein, protein and DNA as well as information on sequences of DNA and proteins. There are multiple signal transduction pathways: cascade of information from plasma membrane to nucleus in response to an extracellular stimulus in living organisms. Extracellular signal molecule binds specific intracellular receptor, and initiates the signaling pathway. Now, there is a large amount of information about the signaling pathway which controls the gene expression and cellular proliferation. We have developed an integrated database SPAD to understand the overview of signaling transduction.
Proper citation: Signaling Pathway Database (RRID:SCR_008243) Copy
ITFP is an integrated transcription factor (TF) platform, which included abundant TFs and targets message of mammalian. Support vector machine (SVM) algorithm combined with error-correcting output coding (ECOC) algorithm was utilized to identify and classify transcription factor from protein sequence of Human, Mouse and Rat. For transcription factor targets, a reverse engineering method named ARACNE was used to derive potential interaction pairs between transcription factor and downstream regulated gene from Human, Mouse and Rat gene expression profile data. Detailed information of gene expression profile data can be found in help page. Moreover, all data provided by the platform is free for non-commercial users and can be downloaded through links on help page.
Proper citation: Intergrated Transcription Factor Platform (RRID:SCR_008119) Copy
http://pbil.univ-lyon1.fr/databases/homolens.php
Database of homologous genes from Ensembl organisms, structured under ACNUC sequence database management system. It allows to select sets of homologous genes among species, and to visualize multiple alignments and phylogenetic trees. It is possible to search for orthologous genes in a wide range of taxons. HOMOLENS is particularly useful for comparative sequence analysis, phylogeny and molecular evolution studies. More generally, HOMOLENS gives an overall view of what is known about a peculiar gene family. Note that HOMOLENS is split into two databases on this server: HOMOLENS contains the protein sequences while HOMOLENSDNA contains the nucleotide sequences. Protein sequences of HOMOLENS have been generated by translating the CDS of HOMOLENSDNA and using associated cross-references to generate the annotations.
Proper citation: Homologous Sequences in Ensembl Animal Genomes (RRID:SCR_008356) Copy
http://www.primervfx.com/#welcome
PrimerParadise is an online PCR primer database for genomics studies. The database contains predesigned PCR primers for amplification of exons, genes and SNPs of almost all sequenced genomes. Primers can be used for genome-wide projects (resequencing, mutation analysis, SNP detection etc). The primers for eukaryotic genomes have been tested with e-PCR to make sure that no alternative products will be generated. Also, all eukaryotic primers have been filtered to exclude primers that bind excessively throughout the genome. Genes are amplified as amplicons. Amplicons are defined as only one genes exons containing maximaly 3000 bp long dna segments. If gene is longer than 3000 bp then it is split into the segments at length 3000 bp. So for example gene at length 5000 bp is split into two segment and for both segments there were designed a separate primerpair. If genes exons length is over 3000 bp then it is split into amplicons as well. Every SNP has one primerpair. In addition of considering repetitive sequences and mono-dinucleotide repeats, we avoid designing primers to genome regions which contain other SNPs. -There are two ways to search for primers: you can use features IDs ( for SNP primers Reference ID, for gene/exon primers different IDs (Ensembl gene IDs, HUGO IDs for human genes, LocusLink IDs, RefSeq IDs, MIM IDs, NCBI gene names, SWISSPROT IDs for bacterial genes, VEGA gene IDs for human and mouse, Sanger S.pombe systematic gene names and common gene names, S.cerevisiae GeneBanks Locus, AccNo, GI IDs and common gene names) -you can use genome regions (chromosome coordinates, chromosome bands if exists) -Currently we provide 3 primers collections: proPCR for prokaryotic organisms genes primers -euPCR for eukaryotic organisms genes/exons primers -snpPCR for eukaryotic organisms SNP primers Sponsors: PrimerStudio is funded by the University of Tartu.
Proper citation: PrimerStudio (RRID:SCR_008232) Copy
http://www.thearkdb.org/arkdb/
This website contains the mapping sequence of poultry. The ArkDB database system aims to provide a comprehensive public repository for genome mapping data from farmed and other animal species. In doing so, it aims to provide a route in to genomic and other sequence from the initial viewpoint of linkage mapping, RH mapping, physical mapping or - possibly more importantly - QTL mapping data. It's supported, in part, by the USDA-CSREES National Animal Genome Research Program in order to serve the poultry genome mapping community. This system represents a complete rewrite of the original version with the code migrated to java and the underlying database targeted at postgres (although any standards-compliant database engine should suffice). The initial release records details of maps and the markers that they contain. There are alternative entry points that target either a chromosome or a specific mapping analysis as the starting point. Limited relationships between markers are recorded and displayed. As with the previous version, all maps are drawn using data extracted from the database on the fly.
Proper citation: ChickBase (RRID:SCR_008147) Copy
http://locustdb.genomics.org.cn/
The migratory locust (Locusta migratoria) is an orthopteran pest and a representative member of hemimetabolous insects. Its transcriptomic data provide invaluable information for molecular entomology study of the insect and pave a way for comparative studies of other medically, agronomically, and ecologically relevant insects. This first transcriptomic database of the locust (LocustDB) has been developed, building necessary infrastructures to integrate, organize, and retrieve data that are either currently available or to be acquired in the future. It currently hosts 45,474 high quality EST sequences from the locust, which were assembled into 12,161 unigenes. This database contains original sequence data, including homologous/orthologous sequences, functional annotations, pathway analysis, and codon usage, based on conserved orthologous groups (COG), gene ontology (GO), protein domain (InterPro), and functional pathways (KEGG). It also provides information from comparative analysis based on data from the migratory locust and five other invertebrate species, such as the silkworm, the honeybee, the fruitfly, the mosquito and the nematode. LocustDB also provides information from comparative analysis based on data from the migratory locust and five other invertebrate species, such as the silkworm, the honeybee, the fruitfly, the mosquito and the nematode. It starts with the first transcriptome information for an orthopteran and hemimetabolous insect and will be extended to provide a framework for incorporation of in-coming genomic data of relevant insect groups and a workbench for cross-species comparative studies.
Proper citation: Migratory Locust EST Database (RRID:SCR_008201) Copy
http://www.sanger.ac.uk/Projects/C_elegans/index.shtml
The Sanger Institute and the Genome Sequencing Center at the Washington University School of Medicine, St. Louis have collaborated to sequence the genomes of both C. elegans and C. briggsae. The completed C. elegans genome sequence is represented by over 3,000 individual clone sequences which can be accessed through this site (or through WormBase). These sequences are submitted to EMBL whenever the sequence or annotation changes (e.g. modification to gene structures) and these submissions are then mirrored to GenBank and DDBJ. These sequences (along with ESTs and proteins) can be searched on our C. elegans BLAST server. WormBase is the repository of mapping, sequencing and phenotypic information for C. elegans. The worm informatics group at the Sanger Institute play a key role in assembling the whole database. They also curate and develop some of the constituent databases that comprise WormBase.
Proper citation: Caenorhabditis Genome Sequencing Projects (RRID:SCR_008155) Copy
http://www.ebi.ac.uk/asd/aedb/index.html
THIS RESOURCE IS NO LONGER IN SERVICE, documented on March 27, 2013. A manual generated database for alternative exons and their properties from numerous species - the data is gathered from literature where these exons have been experimentally verified. Most alternative exons are cassette exons and are expressed in more than two tissues. Of all exons whose expression was reported to be specific for a certain tissue, the majority were expressed in the brain. At the moment, AEdb products that are available are sequence (a database of alternative exons), function (a database of functions attributed to constitutive and alternative exon), regulatory sequence (a database of transcript regulatory motifs), minigenes (a table of minigenes and their associations to splicing events), and diseases (a table of diseases associated with splicing and their associations to AltSplice). Alternative splicing is an important regulatory mechanism of mammalian gene expression. The alternative splicing database (ASD) consortium is systematically collecting and annotating data on alternative splicing. The continuation and upgrade of the ASD consists of computationally and manually generated data. Its largest parts are AltSplice, a value-added database of computationally delineated alternative splicing events. Its data include alternatively spliced introns/exons, events, isoform splicing patterns and isoform peptide sequences. AltSplice data are generated by examining gene-transcript alignments. The data are annotated for various biological features including splicing signals, expression states, (SNP)-mediated splicing and cross-species conservation. AEdb forms the manually curated component of ASD. It is a literature-based data set containing sequence and properties of alternatively spliced exons, functional enumeration of observed splicing events, characterization of observed splicing regulatory elements, and a collection of experimentally clarified minigene constructs.
Proper citation: Alternative Exon Database (RRID:SCR_008157) Copy
http://mips.gsf.de/services/genomes/uwe25/
THIS RESOURCE IS NO LONGER IN SERVICE, documented on July 15, 2013. This is the official database of the environmental chlamydia genome project. This resource provides access to finished sequence for Parachlamydia-related symbiont UWE25 and to a wide range of manual annotations, automatical analyses and derived datasets. Functional classification and description has been manually annotated according to the Annotation guidelines. Chlamydiae are the major cause of preventable blindness and sexually transmitted disease. Genome analysis of a chlamydia-related symbiont of free-living amoebae revealed that it is twice as large as any of the pathogenic chlamydiae and had few signs of recent lateral gene acquisition. We showed that about 700 million years ago the last common ancestor of pathogenic and symbiotic chlamydiae was already adapted to intracellular survival in early eukaryotes and contained many virulence factors found in modern pathogenic chlamydiae, including a type III secretion system. Ancient chlamydiae appear to be the originators of mechanisms for the exploitation of eukaryotic cells. Environmental chlamydiae have recently been recognized as obligate endosymbionts of free-living amoebae and have been implicated as potential human pathogens. Environmental chlamydiae form a deep branching evolutionary lineage within the medically important order Chlamydiales. Despite their high diversity and ubiquitous distribution in clinical and environmental samples only limited information about genetics and ecology of these microorganisms is available. The Parachlamydia-related Acanthamoeba symbiont UWE25 was therefore selected as representative environmental chlamydia strain for whole genome sequencing. Comparative genome analysis was performed using PEDANT and simap. Sponsors: The environmental chlamydia genome project was funded by the bmb+f (German Federal Ministry of Education and Research) and is part of the Competence Network PathoGenoMiK.
Proper citation: Protochlamydia amoebophila UWE25 (RRID:SCR_008222) Copy
http://www.bioinf.mdc-berlin.de/splice/db/
THIS RESOURCE IS NO LONGER IN SERVICE, documented on July 15, 2013. An online available compendium of alternative splice forms for several organisms (Arabidopsis thaliana, Bos taurus, Caenorhabditis elegans, Drosophila melanogaster, Danio rerio, Homo sapiens, Mus musculus, Rattus norvegicus, Xenopus laevis). Alternative splice forms are defined by comparing high-scoring ESTs to mRNA sequences (both from GenBank) with known exon-intron information (from ENSEMBL database) using BLAST. Repetitive sequences of all mRNAs have beforehand been masked by MaskerAid. Filtering programs with defined parameters compare the ends of each aligned sequence pair for deletions or insertions in the EST sequence, which suggest the existence of alternative splice forms. The database is accessible by typing in accession numbers (ACC) or keywords like description, gene names, organism or other keywords. (If more than one hit was found a list of all results is given.) And the result page is divided into 4 major parts. The first part (General Information About The Entry) summarizes the most important information as database ids, organism, and description. The so called alternative splice profile (ASP) of each human sequence is shown in the second part (Alternative Splice Frequency). The ASP indicates the number of alternatively spliced ESTs (NAE), the number of constitutively spliced ESTs (NCE) as well as the number of alternative splice sites (NSS) per mRNA. NAE and NCE corresponds to the EST coverage and can be used as a quality value for the predicted alternative splice variants. The NSS value specifies the splice propensity of a gene. Moreover the number of ESTs from cancerous tissues is shown. The histological source and the developmental stages are illustrated with several colors to enables the user to get an overview of the origins of the matching ESTs. Also, the Splice Site View shows graphically all alternative splice sites for the whole transcript.
Proper citation: Extended Alternatively Spliced EST Database (RRID:SCR_008186) Copy
http://mpr.nci.nih.gov/MPR/BrowseProteins.aspx
THIS RESOURCE IS NO LONGER IN SERVICE, documented on 6/24/13. A repository of information on commercially available phospho-specific antibodies to human phosphorylation sites. It provides a BLAST search for phosphorylation sites using as query the amino acid sequence surrounding the site. It also provides direct links to the relevant antibodies from many companies including BD Pharmingen, Biosource International, Cell Signaling Technology (CST), Santa Cruz Biotechnologies, Upstate Biotechnology.
Proper citation: Mammalian Phosphorylation Resource (RRID:SCR_008210) Copy
http://www.schematikon.org/Nh3D.html
THIS RESOURCE IS NO LONGER IN SERVICE, documented on July 17, 2013. It is freely available as a reference dataset for the statistical analysis of sequence and structure features of proteins in the PDB. It is a dataset of structurally dissimilar proteins. This dataset has been compiled by selecting well resolved representatives from the Topology level of the CATH database which hierarchically classifies all protein structures. These have been been pruned to remove: i) domains that may contain homologous elements (by pairwise sequence comparison and structural superposition of aligned residues) ii) internal duplications (by repeat detection) iii) regions with high B-Factor The statistical analysis of protein structures requires datasets in which structural features can be considered independently distributed, i.e. not related through common ancestry, and that fulfill minimal requirements regarding the experimental quality of the structures it contains. However, non-redundant datasets based on sequence similarity invariably contain distantly related homologues. Here a reference dataset of non-homologous protein domains is provided, assuming that structural dissimilarity at the topology level is incompatible with recognizable common ancestry. It contains the best refined representatives of each Topology level, validates structural dissimilarity and removes internally duplicated fragments. The compilation of Nh3D is fully scripted. The current Nh3D list contains 570 domains with a total of 90780 residues. It covers more than 70% of folds at the Topology level of the CATH database and represents more than 90% of the structures in the PDB that have been classified by CATH. Even though all protein pairs are structurally dissimilar, some pairwise sequence identities after global alignment are greater than 30%. Nh3D is freely available as a reference dataset for the statistical analysis of sequence and structure features of proteins in the PDB.
Proper citation: Nh3D: A Reference Dataset of Structures of Non-homologous Proteins (RRID:SCR_008212) Copy
The European Bioinformatics Institute (EBI) toolbox area provides a comprehensive range of tools for the field of bioinformatics. These are subdivided into categories in the left menu for convenience. EBI has developed a large number of very useful bioinformatics tools. A few examples include: - Similarity & Homology - the BLAST or FASTA programs can be used to look for sequence similarity and infer homology. - Protein Functional Analysis - InterProScan can be used to search for motifs in your protein sequence. - Proteomic Services NEW - UniProt DAS server allows researchers to show their research results in the context of UniProtKB/Swiss-Prot annotation. - Sequence Analysis - ClustalW2 a sequence alignment tool. - Structural Analysis - MSDfold can be used to query your protein structure and compare it to those in the Protein Data Bank (PDB). - Web Services - provide programmatic access to the various databases and retrieval/analysis services EBI provides. - Tools Miscellaneous - Expression Profiler a set of tools for clustering, analysis and visualization of gene expression and other genomic data. Sponsors: This resource is sponsored by EBI.
Proper citation: Toolbox at the European Bioinformatics Institute (RRID:SCR_002872) Copy
https://github.com/sanger-pathogens/Fastaq
Software application for diverse collection of scripts that perform useful and common FASTA/FASTQ manipulation tasks, such as filtering, merging, splitting, sorting, trimming, search/replace, etc. Input and output files can be gzipped (format is automatically detected) and individual Fastaq commands can be piped together.
Proper citation: Fastaq (RRID:SCR_016091) Copy
http://www.sanger.ac.uk/science/tools/ssaha2-0
A program designed for the efficient mapping of sequence reads onto genomic references. The software is capable of reading most sequencing platforms and giving a range of outputs are supported.
Proper citation: Sequence Search and Alignment by Hashing Algorithm (RRID:SCR_000544) Copy
http://depts.washington.edu/yeastrc/
Biomedical technology research center that (1) exploits the budding yeast Saccharomyces cerevisiae to develop novel technologies for investigating and characterizing protein function and protein structure (2) facilitates research and extension of new technologies through collaboration, and (3) actively disseminates data and technology to the research community. Through collaboration, the YRC freely provides resources and expertise in six core technology areas: Protein Tandem Mass Spectrometry, Protein Sequence-Function Relationships, Quantitative Phenotyping, Protein Structure Prediction and Design, Fluorescence Microscopy, Computational Biology.
Proper citation: Yeast Resource Center (RRID:SCR_007942) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the SPARC SAWG Resources search. From here you can search through a compilation of resources used by SPARC SAWG and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that SPARC SAWG has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on SPARC SAWG then you can log in from here to get additional features in SPARC SAWG such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into SPARC SAWG you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within SPARC SAWG that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.