Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
Software toolkit for biological sequence analysis and -presentation combined into a single binary. It is used for genome analysis, efficient processing of structured genome annotations and contains binaries for sequence and annotation handling, sequence compression, index structure generation and access, annotation visualization.
Proper citation: GenomeTools (RRID:SCR_016120) Copy
https://cell-innovation.nig.ac.jp/maser/Tools/visualization_top_en.html
One stop platform for NGS big data from analysis to visualization. There are about 400 analysis pipelines integrated on Maser. List of all analysis pipelines, including descriptions and approximate execution times, can be found on page for ‘All pipelines’ in the User Guide. loadGtfToGe_db software loads GTF files to a database for Genome Explorer. It allows the user to browse the results through the GE.
Proper citation: loadGtfToGe_db (RRID:SCR_015998) Copy
http://sanger-pathogens.github.io/circlator/
Software that automates assembly circularization and produces accurate linear representations of circular sequences. It is used for assembling of DNA sequence data of complete bacterial and small eukaryotic genomes.
Proper citation: Circlator (RRID:SCR_016058) Copy
https://github.com/genomeannotation/GAG
Command line program to read, modify, annotate and generate genomic data. Can write files to .gff3 or to the NCBI's .tbl format.
Proper citation: Genome Annotation Generator (RRID:SCR_016053) Copy
https://gemma.msl.ubc.ca/phenotypes.html
Database that consolidates information on genes and phenotypes across multiple resources and allows tracking and exploring of the associations. Part of Gemma, a web site, database and a set of tools for the meta-analysis, re-use and sharing of genomics data.
Proper citation: Phenocarta (RRID:SCR_016273) Copy
Collects and provides data on the human genome and epigenome to facilitate genetic studies of type 2 diabetes and its complications. A component of the AMP T2D consortium, which includes the National Institute for Diabetes and Digestive and Kidney Diseases (NIDDK) and an international collaboration of researchers.
Proper citation: Diabetes Epigenome Atlas (RRID:SCR_016441) Copy
Database that describes the families of structurally-related catalytic and carbohydrate-binding modules (or functional domains) of enzymes that degrade, modify, or create glycosidic bonds. This specialist database is dedicated to the display and analysis of genomic, structural and biochemical information on Carbohydrate-Active Enzymes (CAZymes). CAZy data are accessible either by browsing sequence-based families or by browsing the content of genomes in carbohydrate-active enzymes. New genomes are added regularly shortly after they appear in the daily releases of GenBank. New families are created based on published evidence for the activity of at least one member of the family and all families are regularly updated, both in content and in description. An original aspect of the CAZy database is its attempt to cover all carbohydrate-active enzymes across organisms and across subfields of glycosciences. One can search for CAZY Family pages using the Protein Accession (Genpept Accession, Uniprot Accession or PDB ID), Cazy family name or EC number. In addition, genomes can be searched using the NCBI TaxID. This search can be complemented by Google-based searches on the CAZy site.
Proper citation: CAZy- Carbohydrate Active Enzyme (RRID:SCR_012909) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on August 18,2025. A web-based plant genome assembly simulation platform whose resources include out of the box scripts for analyzing assembly data, an on-demand web graphing tool to model your experiment, and a downloadable database with metrics and parameters from over 3,000 simulated genome assemblies.
Proper citation: Plantagora (RRID:SCR_001227) Copy
https://www.sanger.ac.uk/collaboration/sequencing-idd-regions-nod-mouse-genome/
Genetic variations associated with type 1 diabetes identified by sequencing regions of the non-obese diabetic (NOD) mouse genome and comparing them with the same areas of a diabetes-resistant C57BL/6J reference mouse allowing identification of single nucleotide polymorphisms (SNPs) or other genomic variations putatively associated with diabetes in mice. Finished clones from the targeted insulin-dependent diabetes (Idd) candidate regions are displayed in the NOD clone sequence section of the website, where they can be downloaded either as individual clone sequences or larger contigs that make up the accession golden path (AGP). All sequences are publicly available via the International Nucleotide Sequence Database Collaboration. Two NOD mouse BAC libraries were constructed and the BAC ends sequenced. Clones from the DIL NOD BAC library constructed by RIKEN Genomic Sciences Centre (Japan) in conjunction with the Diabetes and Inflammation Laboratory (DIL) (University of Cambridge) from the NOD/MrkTac mouse strain are designated DIL. Clones from the CHORI-29 NOD BAC library constructed by Pieter de Jong (Children's Hospital, Oakland, California, USA) from the NOD/ShiLtJ mouse strain are designated CHORI-29. All NOD mouse BAC end-sequences have been submitted to the International Nucleotide Sequence Database Consortium (INSDC), deposited in the NCBI trace archive. They have generated a clone map from these two libraries by mapping the BAC end-sequences to the latest assembly of the C57BL/6J mouse reference genome sequence. These BAC end-sequence alignments can then be visualized in the Ensembl mouse genome browser where the alignments of both NOD BAC libraries can be accessed through the Distributed Annotation System (DAS). The Mouse Genomes Project has used the Illumina platform to sequence the entire NOD/ShiLtJ genome and this should help to position unaligned BAC end-sequences to novel non-reference regions of the NOD genome. Further information about the BAC end-sequences, such as their alignment, variation data and Ensembl gene coverage, can be obtained from the NOD mouse ftp site.
Proper citation: Sequencing of Idd regions in the NOD mouse genome (RRID:SCR_001483) Copy
Functional Analysis of Transcriptional Networks (FunNet) is designed as an integrative tool for analyzing gene co-expression networks built from microarray expression data. The analytical model implemented in this tool involves two abstraction layers: transcriptional (i.e. gene expression profiles) and functional (i.e. biological themes indicating the roles of the analyzed transcripts). A functional analysis technique, which relies on Gene Ontology and KEGG annotations, is applied to extract a list of relevant biological themes from microarray gene expression data. Afterwards multiple-instance representations are built to relate relevant biological themes to their annotated transcripts. An original non-linear dynamical model is used to quantify the contextual proximity of relevant genomic themes based on their patterns of propagation in the gene co-expression network (i.e. capturing the similarity of the expression profiles of the transcriptional instances of annotating themes). In the end an unsupervised multiple-instance spectral clustering procedure is used to explore the modular architecture of the co-expression network by grouping together biological themes demonstrating a significant relationship in the co-expression network. Functional and transcriptional representations of the co-expression network are provided, together with detailed information on the contextual centrality of related transcripts and genomic themes. FunNet is provided both as a web-based tool and as a standalone R package. The standalone R implementation can be run on any operating system for which an R environment implementation is available (Windows, Mac OS, various flavors of Linux and Unix) and can be downloaded from the FunNet website, or from the worldwide mirrors of CRAN. Both implementations of the FunNet tool are provided freely under the GNU General Public License 2.0. Platform: Online tool, Windows compatible, Mac OS X compatible, Linux compatible, Unix compatible
Proper citation: FunNet - Transcriptional Networks Analysis (RRID:SCR_006968) Copy
The HumanCyc database describes human metabolic pathways and the human genome. By presenting metabolic pathways as an organizing framework for the human genome, HumanCyc provides the user with an extended dimension for functional analysis of Homo sapiens at the genomic level. A computational pathway analysis of the human genome assigned human enzymes to predicted metabolic pathways. Pathway assignments place genes in their larger biological context, and are a necessary step toward quantitative modeling of metabolism. HumanCyc contains the complete genome sequence of Homo sapiens, as presented in Build 31. Data on the human genome from Ensembl, LocusLink and GenBank were carefully merged to create a minimally redundant human gene set to serve as an input to SRI''s PathoLogic software, which generated the database and predicted Homo sapiens metabolic pathways from functional information contained in the genome''s annotation. SRI did not re-annotate the genome, but worked with the gene function assignments in Ensembl, LocusLink, and GenBank. The resulting pathway/genome database (PGDB) includes information on 28,783 genes, their products and the metabolic reactions and pathways they catalyze. Also included are many links to other databases and publications. The Pathway Tools software/database bundle includes HumanCyc and the Pathway Tools software suite and is available under license. This form of HumanCyc is faster and more powerful than the Web version.
Proper citation: HumanCyc: Encyclopedia of Homo sapiens Genes and Metabolism (RRID:SCR_007050) Copy
http://tritrypdb.org/tritrypdb/
An integrated genomic and functional genomic database providing access to genome-scale datasets for kinetoplastid parasites, and supporting a variety of complex queries driven by research and development needs. Currently, TriTrypDB integrates datasets from Leishmania braziliensis, L. infantum, L. major, L. tarentolae, Trypanosoma brucei and T. cruzi. Users may examine individual genes or chromosomal spans in their genomic context, including syntenic alignments with other kinetoplastid organisms. Data within TriTrypDB can be interrogated utilizing a sophisticated search strategy system that enables a user to construct complex queries combining multiple data types. All search strategies are stored, allowing future access and integrated searches. ''''User Comments'''' may be added to any gene page, enhancing available annotation; such comments become immediately searchable via the text search, and are forwarded to curators for incorporation into the reference annotation when appropriate. TriTrypDB provides programmatic access to its searches, via REST Web Services. The result of a web service request is a list of records (genes, ESTs, etc) in either XML or JSON format. REST services can be executed in a browser by typing a specific URL. TriTrypDB and its continued development are possible through the collaborative efforts between EuPathDB, GeneDB and colleagues at the Seattle Biomedical Research Institute (SBRI).
Proper citation: TriTrypDB (RRID:SCR_007043) Copy
http://yetfasco.ccbr.utoronto.ca/
Collection of all available transcription factor (TF) specificities for the yeast Saccharomyces cerevisiae in Position Frequency Matrix (PFM) or Position Weight Matrix (PWM) formats. The specificities are evaluated for quality using several metrics. With this website, you can scan sequences with the motifs to find where potential binding sites lie, inspect precomputed genome-wide binding sites, find which TFs have similar motifs to one you have found, and download the collection of motifs. Submissions are welcome.
Proper citation: YeTFaSCo (RRID:SCR_006893) Copy
http://www.broad.mit.edu/node/305
The Connectivity Map aims to generate a detailed map that links gene patterns associated with disease to corresponding patterns produced by drug candidates and a variety of genetic manipulations. The Connectivity Map is the most comprehensive effort yet for using genomics in a drug-discovery framework. It allows researchers to screen compounds against genome-wide disease signatures, rather than a pre-selected set of target genes. Drugs are paired with diseases using sophisticated pattern-matching methods with a high level of resolution and specificity. To build a Connectivity Map, the Broad Institute brings together molecular biologists, genomics specialists, computational scientists, pharmacologists, chemists and chemical biologists, as well as expertise from across the breadth and depth of medicine.Connectivity map is a large public database of signatures of drugs and genes, and pattern-matching tools to detect similarities among these signatures.The parent site for the Broad Institute at MIT has a software library of software applications developed for use in genetic analysis.
Proper citation: National Institute of Mental Health (NIMH) Human Genetics Initiative (RRID:SCR_007436) Copy
http://www.geisha.arizona.edu/geisha/
Online repository for chicken in situ hybridization information. This site presents whole mount in situ hybridization images and corresponding probe and genomic information for genes expressed in chicken embryos in Hamburger Hamilton stages 1-25 (0.5-5 days). The GEISHA project began in 1998 to investigate using high throughput whole mount in situ hybridization to identify novel, differentially expressed genes in chicken embryos. An initial expression screen of approximately 900 genes demonstrated feasibility of the approach, and also highlighted the need for a centralized repository of in situ hybridization expression data. Objectives: The goals of the GEISHA project are to obtain whole mount in situ hybridization expression information for all differentially expressed genes in the chicken embryo between HH stages 1-25, to integrate expression data with the chicken genome browsers, and to offer this information through a user-friendly graphical user interface. In situ hybridization images are obtained from three sources: 1. In house high throughput in situ hybridization screening: cDNAs obtained from several embryonic cDNA libraries or from EST repositories are screened for expression using high throughput in situ hybridization approaches. 2. Literature curation: Agreements with journals permit posting of published in situ hybridization images and related information on the GEISHA site. 3. Unpublished in situ hybridization information from other laboratories: laboratories generally publish only a small fraction of their in situ hybridization data. High quality images for which probe identity can be verified are welcome additions to GEISHA.
Proper citation: GEISHA - Gallus Expression in Situ Hybridization Analysis: A Chicken Embryo Gene Expression Database (RRID:SCR_007440) Copy
Comprehensive catalogue of animal genome size data. Haploid DNA contents (C-values, in picograms) are available for 4972 species (3231 vertebrates and 1741 non-vertebrates) based on 6518 records from 669 published sources. Data may be submitted directly to the database or reprints and notifications of new papers may be sent to database curation staff.
Proper citation: Animal Genome Size Database (RRID:SCR_007551) Copy
http://gene3d.biochem.ucl.ac.uk/Gene3D/
A large database of CATH protein domain assignments for ENSEMBL genomes and Uniprot sequences. Gene3D is a resource of form studying proteins and the component domains. Gene3D takes CATH domains from Protein Databank (PDB) structures and assigns them to the millions of protein sequences with no PDB structures using Hidden Markov models. Assigning a CATH superfamily to a region of a protein sequence gives information on the gross 3D structure of that region of the protein. CATH superfamilies have a limited set of functions and so the domain assignment provides some functional insights. Furthermore most proteins have several different domains in a specific order, so looking for proteins with a similar domain organization provides further functional insights. Strict confidence cut-offs are used to ensure the reliability of the domain assignments. Gene3D imports functional information from sources such as UNIPROT, and KEGG. They also import experimental datasets on request to help researchers integrate there data with the corpus of the literature. The website allows users to view descriptions for both single proteins and genes and large protein sets, such as superfamilies or genomes. Subsets can then be selected for detailed investigation or associated functions and interactions can be used to expand explorations to new proteins. The Gene3D web services provide programmatic access to the CATH-Gene3D annotation resources and in-house software tools. These services include Gene3DScan for identifying structural domains within protein sequences, access to pre-calculated annotations for the major sequence databases, and linked functional annotation from UniProt, GO and KEGG., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: Gene3D (RRID:SCR_007672) Copy
http://genomics.senescence.info/
Collection of databases and tools designed to help researchers study the genetics of human ageing using modern approaches such as functional genomics, network analyses, systems biology and evolutionary analyses. A major resource in HAGR is GenAge, which includes a curated database of genes related to human aging and a database of ageing- and longevity-associated genes in model organisms. Another major database in HAGR is AnAge. Featuring over 4,000 species, AnAge provides a compilation of data on aging, longevity, and life history that is ideal for the comparative biology of aging. GenDR is a database of genes associated with dietary restriction based on genetic manipulation experiments and gene expression profiling. Other projects include evolutionary studies, genome sequencing, cancer genomics, and gene expression analyses. The latter allowed them to identify a set of genes commonly altered during mammalian aging which represents a conserved molecular signature of aging. Software, namely in the form of scripts for Perl and SPSS, is made available for users to perform a variety of bioinformatic analyses potentially relevant for studying aging. The Perl toolkit, entitled the Ageing Research Computational Tools (ARCT), provides modules for parsing files, data-mining, searching and downloading data from the Internet, etc. Also available is an SPSS script that can be used to determine the demographic rate of aging for a given population. An extensive list of links regarding computational biology, genomics, gerontology, and comparative biology is also available.
Proper citation: Human Ageing Genomic Resources (RRID:SCR_007700) Copy
Database of information about restriction enzymes and related proteins containing published and unpublished references, recognition and cleavage sites, isoschizomers, commercial availability, methylation sensitivity, crystal, genome, and sequence data. DNA methyltransferases, homing endonucleases, nicking enzymes, specificity subunits and control proteins are also included. Several tools are available including REBsites, BLAST against REBASE, NEBcutter and REBpredictor. Putative DNA methyltransferases and restriction enzymes, as predicted from analysis of genomic sequences, are also listed. REBASE is updated daily and is constantly expanding. Users may submit new enzyme and/or sequence information, recommend references, or send them corrections to existing data. The contents of REBASE may be browsed from the web and selected compilations can be downloaded by ftp (ftp.neb.com). Additionally, monthly updates can be requested via email.,
Proper citation: REBASE (RRID:SCR_007886) Copy
https://github.com/vetscience/Assemblosis
Software tool as a Common Workflow Language (CWL) based automated bioinformatics workflow to assemble haploid/diploid eukaryote genomes of non-model organisms using PacBio long-reads and Illumina short-reads.
Proper citation: Assemblosis (RRID:SCR_016571) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the T1D Resources search. From here you can search through a compilation of resources used by T1D and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that T1D has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on T1D then you can log in from here to get additional features in T1D such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into T1D you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within T1D that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.