Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
A clustering and visualization tool that enables the interactive exploration of genome-wide data, with a specialization in epigenomics data. Spark is also available as a service within the Epigenome toolset of the Genboree Workbench. The approach utilizes data clusters as a high-level visual guide and supports interactive inspection of individual regions within each cluster. The cluster view links to gene ontology analysis tools and the detailed region view connects to existing genome browser displays taking advantage of their wealth of annotation and functionality.
Proper citation: Spark (RRID:SCR_006207) Copy
An open-membership International community to promote mechanisms that standardize the description of genomes and the exchange and integration of genomic data. Community-driven standards have the best chance of success if developed within the auspices of international working groups. Participants in the GSC include biologists, computer scientists, those building genomic databases and conducting large-scale comparative genomic analyses, and those with experience of building community-based standards. The mission of the GSC is to work with the wider community towards: * the implementation of new genomic standards * methods of capturing and exchanging metadata * harmonization of metadata collection and analysis efforts across the wider genomics community
Proper citation: Genomic Standards Consortium (RRID:SCR_006273) Copy
http://bioinformatics.ubc.ca/ermineJ/
Data analysis software for gene sets in expression microarray data or other genome-wide data that results in rankings of genes. A typical goal is to determine whether particular biological pathways are doing something interesting in the data. The software is designed to be used by biologists with little or no informatics background. A command-line interface is available for users who wish to script the use of ermineJ. Major features include: * Implementation of multiple methods for gene set analysis: ** Over-representation analysis ** A resampling-based method that uses gene scores ** A rank-based method that uses gene scores ** A resampling-based method that uses correlation between gene expression profiles (a type of cluster-enrichment analysis). * Gene sets receive statistical scores (p-values), and multiple test correction is supported. * Support of the Gene Ontology terminology; users can choose which aspects to analyze. * User files use simple text formats. * Users can modify gene sets or create new ones. * The results can be visualized within the software. * It is simple to compare multiple analyses of the same data set with different settings. * User-definable hyperlinks are provided to external sites to allow more efficient browsing of the results. * For programmers, there is a command line interface as well as a simple application programming interface that can be used to plug ermineJ functionality into your own code Platform: Online tool, Windows compatible, Mac OS X compatible, Linux compatible, Unix compatible
Proper citation: ErmineJ (RRID:SCR_006450) Copy
Public archive providing a comprehensive record of the world''''s nucleotide sequencing information, covering raw sequencing data, sequence assembly information and functional annotation. All submitted data, once public, will be exchanged with the NCBI and DDBJ as part of the INSDC data exchange agreement. The European Nucleotide Archive (ENA) captures and presents information relating to experimental workflows that are based around nucleotide sequencing. A typical workflow includes the isolation and preparation of material for sequencing, a run of a sequencing machine in which sequencing data are produced and a subsequent bioinformatic analysis pipeline. ENA records this information in a data model that covers input information (sample, experimental setup, machine configuration), output machine data (sequence traces, reads and quality scores) and interpreted information (assembly, mapping, functional annotation). Data arrive at ENA from a variety of sources including submissions of raw data, assembled sequences and annotation from small-scale sequencing efforts, data provision from the major European sequencing centers and routine and comprehensive exchange with their partners in the International Nucleotide Sequence Database Collaboration (INSDC). Provision of nucleotide sequence data to ENA or its INSDC partners has become a central and mandatory step in the dissemination of research findings to the scientific community. ENA works with publishers of scientific literature and funding bodies to ensure compliance with these principles and to provide optimal submission systems and data access tools that work seamlessly with the published literature. ENA is made up of a number of distinct databases that includes the EMBL Nucleotide Sequence Database (Embl-Bank), the newly established Sequence Read Archive (SRA) and the Trace Archive. The main tool for downloading ENA data is the ENA Browser, which is available through REST URLs for easy programmatic use. All ENA data are available through the ENA Browser. Note: EMBL Nucleotide Sequence Database (EMBL-Bank) is entirely included within this resource.
Proper citation: European Nucleotide Archive (ENA) (RRID:SCR_006515) Copy
Set of measures intended for use in large-scale genomic studies. Facilitate replication and validation across studies. Includes links to standards and resources in effort to facilitate data harmonization to legacy data. Measurement protocols that address wide range of research domains. Information about each protocol to ensure consistent data collection.Collections of protocols that add depth to Toolkit in specific areas.Tools to help investigators implement measurement protocols.
Proper citation: Phenotypes and eXposures Toolkit (RRID:SCR_006532) Copy
Database for genetic, genomic, phenotype, and disease data generated from rat research. Centralized database that collects, manages, and distributes data generated from rat genetic and genomic research and makes these data available to scientific community. Curation of mapped positions for quantitative trait loci, known mutations and other phenotypic data is provided. Facilitates investigators research efforts by providing tools to search, mine, and analyze this data. Strain reports include description of strain origin, disease, phenotype, genetics, immunology, behavior with links to related genes, QTLs, sub-strains, and strain sources.
Proper citation: Rat Genome Database (RGD) (RRID:SCR_006444) Copy
http://www.gigasciencejournal.com/
An online open-access open-data journal, publishing ''big-data'' studies from the entire spectrum of life and biomedical sciences whose publication format links standard manuscript publication with its affiliated database, GigaDB, that hosts all associated data, provides data analysis tools, cloud-computing resources, and a DOI assignment to every dataset. GigaScience covers not just ''omic'' type data and the fields of high-throughput biology currently serviced by large public repositories, but also the growing range of more difficult-to-access data, such as imaging, neuroscience, ecology, cohort data, systems biology and other new types of large-scale sharable data. Supporting the open-data movement, they require that all supporting data and source code be publicly available in a suitable public repository and/or under a public domain CC0 license in the BGI GigaScience database. Using the BGI cloud as a test environment, they also consider open-source software tools / methods for the analysis or handling of large-scale data. When submitting a manuscript, please contact them if you have datasets or cloud applications you would like them to host. To maximize data usability submitters are encouraged to follow best practice for metadata reporting and are given the opportunity to submit in ISA-Tab format.
Proper citation: GigaScience (RRID:SCR_006565) Copy
Database of Drosophila genetic and genomic information with information about stock collections and fly genetic tools. Gene Ontology (GO) terms are used to describe three attributes of wild-type gene products: their molecular function, the biological processes in which they play a role, and their subcellular location. Additionally, FlyBase accepts data submissions. FlyBase can be searched for genes, alleles, aberrations and other genetic objects, phenotypes, sequences, stocks, images and movies, controlled terms, and Drosophila researchers using the tools available from the "Tools" drop-down menu in the Navigation bar.
Proper citation: FlyBase (RRID:SCR_006549) Copy
Model organism database that provides organization of and access to scientific data for the fission yeast Schizosaccharomyces pombe. PomBase supports genomic sequence and features, genome-wide datasets and manual literature curation. PomBase also provides a community hub for researchers, providing genome statistics, a community curation interface, news, events, documentation, mailing lists, and welcomes data submissions.
Proper citation: PomBase (RRID:SCR_006586) Copy
Service providing functional analysis of proteins by classifying them into families and predicting domains and important sites. They combine protein signatures from a number of member databases into a single searchable resource, capitalizing on their individual strengths to produce a powerful integrated database and diagnostic tool. This integrated database of predictive protein signatures is used for the classification and automatic annotation of proteins and genomes. InterPro classifies sequences at superfamily, family and subfamily levels, predicting the occurrence of functional domains, repeats and important sites. InterPro adds in-depth annotation, including GO terms, to the protein signatures. You can access the data programmatically, via Web Services. The member databases use a number of approaches: # ProDom: provider of sequence-clusters built from UniProtKB using PSI-BLAST. # PROSITE patterns: provider of simple regular expressions. # PROSITE and HAMAP profiles: provide sequence matrices. # PRINTS provider of fingerprints, which are groups of aligned, un-weighted Position Specific Sequence Matrices (PSSMs). # PANTHER, PIRSF, Pfam, SMART, TIGRFAMs, Gene3D and SUPERFAMILY: are providers of hidden Markov models (HMMs). Your contributions are welcome. You are encouraged to use the ''''Add your annotation'''' button on InterPro entry pages to suggest updated or improved annotation for individual InterPro entries.
Proper citation: InterPro (RRID:SCR_006695) Copy
http://pre.ensembl.org/index.html
Database of genomes that are in the process of being annotated are provided as an early access site for users. Genomes are here when the initial BLAST analysis on a new assembly has been done but the gene build has not been completed. Owing to the preliminary nature of the data, Pre-Ensembl provides views of the assembly, BLAST against the assembly and download of portions of the assembly - and little else. A number of ready-made tools for processing your data are also available. In general a full Ensembl release takes months depending on how complex the data are and the time constraints of people in the team. Occasionally a more complete gene build will be released on this site, but without any comparative genomics, variation or other additional data. Many other species with fully annotated genomic data, more website features and documentation are available at www.ensembl.org
Proper citation: Pre Ensembl (RRID:SCR_006766) Copy
Data resource that includes a large collection of genome-wide ChIP-Seq experiments performed on transcription factors (TFs), histone modifications, RNA polymerases and others. Enriched peak regions from the ChIP-Seq experiments are crossed with the genomic coordinates of a set of input genes, to identify which of the experiments present a statistically significant number of peaks within the input genes' loci. The input can be a cluster of co-expressed genes, or any other set of genes sharing a common regulatory profile. Users can thus single out which TFs are likely to be common regulators of the genes, and their respective correlations. Also, by examining results on promoter activation, transcription, histone modifications, polymerase binding and so on, users can investigate the effect of the TFs (activation or repression of transcription) as well as of the cell or tissue specificity of the genes' regulation and expression.
Proper citation: Cscan (RRID:SCR_006756) Copy
Database of peer-reviewed, continually updated annotation for the Pseudomonas aeruginosa PAO1 reference strain genome expanded to include all Pseudomonas species to facilitate cross-strain and cross-species genome comparisons with high quality comparative genomics. The database contains robust assessment of orthologs, a novel ortholog clustering method, and incorporates five views of the data at the sequence and annotation levels (Gbrowse, Mauve and custom views) to facilitate genome comparisons. Other features include more accurate protein subcellular localization predictions and a user-friendly, Boolean searchable log file of updates for the reference strain PAO1. The current annotation is updated using recent research literature and peer-reviewed submissions by a worldwide community of PseudoCAP (Pseudomonas aeruginosa Community Annotation Project) participating researchers. If you are interested in participating, you are invited to get involved. Many annotations, DNA sequences, Orthologs, Intergenic DNA, and Protein sequences are available for download.
Proper citation: Pseudomonas Genome Database (RRID:SCR_006590) Copy
Collection of data related to crop plant and model organism Zea mays. Used to synthesize, display, and provide access to maize genomics and genetics data, prioritizing mutant and phenotype data and tools, structural and genetic map sets, and gene models and to provide support services to the community of maize researchers. Data stored at MaizeGDB was inherited from the MaizeDB and ZmDB projects. Sequence data are from GenBank. Data are searchable by phenotype, traits, Pests, Gel Pattern, and Mutant Images.
Proper citation: MaizeGDB (RRID:SCR_006600) Copy
Encyclopedia of DNA elements consisting of list of functional elements in human genome, including elements that act at protein and RNA levels, and regulatory elements that control cells and circumstances in which gene is active. Enables scientific and medical communities to interpret role of human genome in biology and disease. Provides identification of common cell types to facilitate integrative analysis and new experimental technologies based on high-throughput sequencing. Genome Browser containing ENCODE and Epigenomics Roadmap data. Data are available for entire human genome.
Proper citation: ENCODE (RRID:SCR_006793) Copy
Functional Analysis of Transcriptional Networks (FunNet) is designed as an integrative tool for analyzing gene co-expression networks built from microarray expression data. The analytical model implemented in this tool involves two abstraction layers: transcriptional (i.e. gene expression profiles) and functional (i.e. biological themes indicating the roles of the analyzed transcripts). A functional analysis technique, which relies on Gene Ontology and KEGG annotations, is applied to extract a list of relevant biological themes from microarray gene expression data. Afterwards multiple-instance representations are built to relate relevant biological themes to their annotated transcripts. An original non-linear dynamical model is used to quantify the contextual proximity of relevant genomic themes based on their patterns of propagation in the gene co-expression network (i.e. capturing the similarity of the expression profiles of the transcriptional instances of annotating themes). In the end an unsupervised multiple-instance spectral clustering procedure is used to explore the modular architecture of the co-expression network by grouping together biological themes demonstrating a significant relationship in the co-expression network. Functional and transcriptional representations of the co-expression network are provided, together with detailed information on the contextual centrality of related transcripts and genomic themes. FunNet is provided both as a web-based tool and as a standalone R package. The standalone R implementation can be run on any operating system for which an R environment implementation is available (Windows, Mac OS, various flavors of Linux and Unix) and can be downloaded from the FunNet website, or from the worldwide mirrors of CRAN. Both implementations of the FunNet tool are provided freely under the GNU General Public License 2.0. Platform: Online tool, Windows compatible, Mac OS X compatible, Linux compatible, Unix compatible
Proper citation: FunNet - Transcriptional Networks Analysis (RRID:SCR_006968) Copy
The HumanCyc database describes human metabolic pathways and the human genome. By presenting metabolic pathways as an organizing framework for the human genome, HumanCyc provides the user with an extended dimension for functional analysis of Homo sapiens at the genomic level. A computational pathway analysis of the human genome assigned human enzymes to predicted metabolic pathways. Pathway assignments place genes in their larger biological context, and are a necessary step toward quantitative modeling of metabolism. HumanCyc contains the complete genome sequence of Homo sapiens, as presented in Build 31. Data on the human genome from Ensembl, LocusLink and GenBank were carefully merged to create a minimally redundant human gene set to serve as an input to SRI''s PathoLogic software, which generated the database and predicted Homo sapiens metabolic pathways from functional information contained in the genome''s annotation. SRI did not re-annotate the genome, but worked with the gene function assignments in Ensembl, LocusLink, and GenBank. The resulting pathway/genome database (PGDB) includes information on 28,783 genes, their products and the metabolic reactions and pathways they catalyze. Also included are many links to other databases and publications. The Pathway Tools software/database bundle includes HumanCyc and the Pathway Tools software suite and is available under license. This form of HumanCyc is faster and more powerful than the Web version.
Proper citation: HumanCyc: Encyclopedia of Homo sapiens Genes and Metabolism (RRID:SCR_007050) Copy
http://tritrypdb.org/tritrypdb/
An integrated genomic and functional genomic database providing access to genome-scale datasets for kinetoplastid parasites, and supporting a variety of complex queries driven by research and development needs. Currently, TriTrypDB integrates datasets from Leishmania braziliensis, L. infantum, L. major, L. tarentolae, Trypanosoma brucei and T. cruzi. Users may examine individual genes or chromosomal spans in their genomic context, including syntenic alignments with other kinetoplastid organisms. Data within TriTrypDB can be interrogated utilizing a sophisticated search strategy system that enables a user to construct complex queries combining multiple data types. All search strategies are stored, allowing future access and integrated searches. ''''User Comments'''' may be added to any gene page, enhancing available annotation; such comments become immediately searchable via the text search, and are forwarded to curators for incorporation into the reference annotation when appropriate. TriTrypDB provides programmatic access to its searches, via REST Web Services. The result of a web service request is a list of records (genes, ESTs, etc) in either XML or JSON format. REST services can be executed in a browser by typing a specific URL. TriTrypDB and its continued development are possible through the collaborative efforts between EuPathDB, GeneDB and colleagues at the Seattle Biomedical Research Institute (SBRI).
Proper citation: TriTrypDB (RRID:SCR_007043) Copy
http://www.transcriptionfactor.org/index.cgi?Home
Database of predicted transcription factors in completely sequenced genomes. The predicted transcription factors all contain assignments to sequence specific DNA-binding domain families. The predictions are based on domain assignments from the SUPERFAMILY and Pfam hidden Markov model libraries. Benchmarks of the transcription factor predictions show they are accurate and have wide coverage on a genomic scale. The DBD consists of predicted transcription factor repertoires for 930 completely sequenced genomes.
Proper citation: DBD: Transcription factor prediction database (RRID:SCR_002300) Copy
http://mirna.imbb.forth.gr/SSCprofiler.html
Tool which can be used to identify novel miRNA gene candidates in the human genome.
Proper citation: SSCprofiler (RRID:SCR_001282) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the T1D Resources search. From here you can search through a compilation of resources used by T1D and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that T1D has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on T1D then you can log in from here to get additional features in T1D such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into T1D you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within T1D that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.