Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://hcv.lanl.gov/content/sequence/HCV/ToolsOutline.html
The HCV sequence database collects and annotates sequence data and provides them to the public via a website that contains a user-friendly search interface and a large number of sequence analysis tools, based on the model of the highly regarded Los Alamos HIV database. The hepatitis C virus (HCV) is a significant threat to public health worldwide. The virus is highly variable and evolves rapidly, making it an elusive target for the immune system and for vaccine and drug design. At present, some 30 000 HCV sequences have been published. This central website provides annotated sequences and analysis tools that will be helpful to HCV scientists worldwide. Things you can do: * Find sequences in the database * Download sequences from the database * Retrieve data about the sequences * Analyze sequences * Work with the sequences using our tools * Download ready-made alignments The HCV sequence database was officially launched in September 2003. Since then, its usage has steadily increased and is now at an average of approximately 280 visits per day from distinct IP addresses.
Proper citation: HCV Sequence Database (RRID:SCR_006019) Copy
http://athina.biol.uoa.gr/bioinformatics/PRED-GPCR/
A prediction tool for GPCR Family Classification from sequence alone based on a probabilistic method that uses family-specific profile Hidden Markov Models. The PRED-GPCR system is based on a probabilistic method that uses family specific profile HMMs in order to determine to which GPCR family a query sequence belongs or resembles. The approach proposed in this method exploits the descriptive power of profile HMMs along with an exhaustive discrimination assessment method to select only highly selective and sensitive profiles, for each family. The collection of these profiles constitutes a signature library, which is scanned, for significant matches with a given query sequence. The output report for a query sequence consists of two sections: * A ranked list of the profile HMM matches, below the selected individual motif E-value cutoff, along with their corresponding family. * A ranked list of the Combined P-values, E-values as well as the number of profiles matched for each family. To cross-evaluate your results you can browse through Swiss-Prot, Trembl, Pfam and Prosite family related entries.
Proper citation: PRED-GPCR (RRID:SCR_006196) Copy
Database of peer-reviewed, continually updated annotation for the Pseudomonas aeruginosa PAO1 reference strain genome expanded to include all Pseudomonas species to facilitate cross-strain and cross-species genome comparisons with high quality comparative genomics. The database contains robust assessment of orthologs, a novel ortholog clustering method, and incorporates five views of the data at the sequence and annotation levels (Gbrowse, Mauve and custom views) to facilitate genome comparisons. Other features include more accurate protein subcellular localization predictions and a user-friendly, Boolean searchable log file of updates for the reference strain PAO1. The current annotation is updated using recent research literature and peer-reviewed submissions by a worldwide community of PseudoCAP (Pseudomonas aeruginosa Community Annotation Project) participating researchers. If you are interested in participating, you are invited to get involved. Many annotations, DNA sequences, Orthologs, Intergenic DNA, and Protein sequences are available for download.
Proper citation: Pseudomonas Genome Database (RRID:SCR_006590) Copy
https://docs.python.org/2/library/random.html
This module implements pseudo-random number generators for various distributions. For integers, uniform selection from a range. For sequences, uniform selection of a random element, a function to generate a random permutation of a list in-place, and a function for random sampling without replacement. On the real line, there are functions to compute uniform, normal (Gaussian), lognormal, negative exponential, gamma, and beta distributions. For generating distributions of angles, the von Mises distribution is available. Sponsors: This resource is supported by ASTi logo Advanced Simulation Technology Inc. (ASTi); Array BioPharma Inc.; BizRate.com; Canonical Ltd.; CCP Games; cPacket Networks; EarnMyDegree.com; Enthought Inc.; Exoweb Ltd.; Google; HitMeister Inc.; IronPort Systems; KNMP; Lucasfilm; Madison Tyler LLC.; Merfin, LLC.; Microsoft; OpenEye Scientific Software; Opsware, Inc.; O''Reilly & Associates, Inc.; PropertySold.ca; Rogue Wave; SEO Moves; Strakt Holdings, Inc.; Sun Microsystems; Tabblo; ZeOmega, LLC., and Zope Corporation.
Proper citation: Generate Pseudo-Random Numbers (RRID:SCR_006535) Copy
Collection of data related to crop plant and model organism Zea mays. Used to synthesize, display, and provide access to maize genomics and genetics data, prioritizing mutant and phenotype data and tools, structural and genetic map sets, and gene models and to provide support services to the community of maize researchers. Data stored at MaizeGDB was inherited from the MaizeDB and ZmDB projects. Sequence data are from GenBank. Data are searchable by phenotype, traits, Pests, Gel Pattern, and Mutant Images.
Proper citation: MaizeGDB (RRID:SCR_006600) Copy
Database for genetic, genomic, phenotype, and disease data generated from rat research. Centralized database that collects, manages, and distributes data generated from rat genetic and genomic research and makes these data available to scientific community. Curation of mapped positions for quantitative trait loci, known mutations and other phenotypic data is provided. Facilitates investigators research efforts by providing tools to search, mine, and analyze this data. Strain reports include description of strain origin, disease, phenotype, genetics, immunology, behavior with links to related genes, QTLs, sub-strains, and strain sources.
Proper citation: Rat Genome Database (RGD) (RRID:SCR_006444) Copy
Database of Drosophila genetic and genomic information with information about stock collections and fly genetic tools. Gene Ontology (GO) terms are used to describe three attributes of wild-type gene products: their molecular function, the biological processes in which they play a role, and their subcellular location. Additionally, FlyBase accepts data submissions. FlyBase can be searched for genes, alleles, aberrations and other genetic objects, phenotypes, sequences, stocks, images and movies, controlled terms, and Drosophila researchers using the tools available from the "Tools" drop-down menu in the Navigation bar.
Proper citation: FlyBase (RRID:SCR_006549) Copy
https://github.com/uclinfectionimmunity/Decombinator
Software suite for analysis of T cell receptor repertoire data. Used for fast, efficient analysis of T cell receptor (TcR) repertoire samples, designed to be accessible to those with no previous programming experience.
Proper citation: Decombinator (RRID:SCR_006732) Copy
International collaboration producing an extensive public catalog of human genetic variation, including SNPs and structural variants, and their haplotype contexts, in an effort to provide a foundation for investigating the relationship between genotype and phenotype. The genomes of about 2500 unidentified people from about 25 populations around the world were sequenced using next-generation sequencing technologies. Redundant sequencing on various platforms and by different groups of scientists of the same samples can be compared. The results of the study are freely and publicly accessible to researchers worldwide. The consortium identified the following populations whose DNA will be sequenced: Yoruba in Ibadan, Nigeria; Japanese in Tokyo; Chinese in Beijing; Utah residents with ancestry from northern and western Europe; Luhya in Webuye, Kenya; Maasai in Kinyawa, Kenya; Toscani in Italy; Gujarati Indians in Houston; Chinese in metropolitan Denver; people of Mexican ancestry in Los Angeles; and people of African ancestry in the southwestern United States. The goal Project is to find most genetic variants that have frequencies of at least 1% in the populations studied. Sequencing is still too expensive to deeply sequence the many samples being studied for this project. However, any particular region of the genome generally contains a limited number of haplotypes. Data can be combined across many samples to allow efficient detection of most of the variants in a region. The Project currently plans to sequence each sample to about 4X coverage; at this depth sequencing cannot provide the complete genotype of each sample, but should allow the detection of most variants with frequencies as low as 1%. Combining the data from 2500 samples should allow highly accurate estimation (imputation) of the variants and genotypes for each sample that were not seen directly by the light sequencing. All samples from the 1000 genomes are available as lymphoblastoid cell lines (LCLs) and LCL derived DNA from the Coriell Cell Repository as part of the NHGRI Catalog. The sequence and alignment data generated by the 1000genomes project is made available as quickly as possible via their mirrored ftp sites. ftp://ftp.1000genomes.ebi.ac.uk ftp://ftp-trace.ncbi.nlm.nih.gov/1000genomes
Proper citation: 1000 Genomes: A Deep Catalog of Human Genetic Variation (RRID:SCR_006828) Copy
http://ecoliwiki.net/colipedia/index.php/T4-like_genome_database
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 22, 2016. A database of information on bacterial phages. It contains multiple phage genomes, which users can BLAST and MegaBLAST, and also hosts a Phage Forum in which users can discuss phage data. Interactive browsing of completed phage genomes is available using the program. The browser allows users to scan the genome for particular features and to download sequence information plus analyses of those features. Views of the genome are generated showing named genes BLAST similarities to other phages predicted tRNAs and other sequence features.
Proper citation: T4-like genome database (RRID:SCR_005367) Copy
http://genome.jgi.doe.gov/programs/plants/index.jsf
The goal of the DOE JGI Plant Genome Program is to shed light on the fundamental biology of photosynthesis and transduction of solar to chemical energy. Other areas of interest include characterizing: * Ecosystems and the role of terrestrial plants and oceanic phytoplankton-in carbon sequestration. * The role of plants in coping with toxic pollutants in soils by hyper-accumulation and detoxification. * Feedstocks for biofuels, e.g., biodiesel from soybean; cellulosic ethanol from perennial grasses. * The ability to respond to environmental change (e.g., loss of diversity from monoculture produces vulnerabilities; nitrogen fixing nodules in legumes reduce fertilizer need). * The generation of useful secondary metabolites (produced largely for disease resistance)- for positive/negative control in agriculture, with attendant influence on global carbon cycle. The Plant Genome Program accomplishes the above through the following activities: # Sequence. Produce genome sequences of key plant (and algal) species to accelerate biofuel development and understand response to climate change. # Function. Develop datasets (and synthetic biology tools) to elucidate functional elements in plant genomes, with special focus on handful of flagship genomes. # Variation. Characterize natural genomic variation in plants (and their associated microbiomes), and relate to biofuel sustainability and adaptation to climate change. # Integration. Provide a centralized hub for the retrieval and deep integrated analysis of plant genome datasets.
Proper citation: Plant Genome Resource at JGI (RRID:SCR_005315) Copy
Database of known and predicted protein interactions. The interactions include direct (physical) and indirect (functional) associations and are derived from four sources: Genomic Context, High-throughput experiments, (Conserved) Coexpression, and previous knowledge. STRING quantitatively integrates interaction data from these sources for a large number of organisms, and transfers information between these organisms where applicable. The database currently covers 5''214''234 proteins from 1133 organisms. (2013)
Proper citation: STRING (RRID:SCR_005223) Copy
http://bioinfo.iitk.ac.in/MIPModDB/
This is a database of comparative protein structure models of MIP (Major Intrinsic Protein) family of proteins. The nearly completed sets of MIPs have been identified from the completed genome sequence of organisms available at NCBI. The structural models of MIP proteins were created by defined protocol. The database aims to provide key information of MIPs in particular based on sequence as well as structures. This will further help to decipher the function of uncharacterized MIPs. For each MIP entry, this database contains information about the source, gene structure, sequence features, substitutions in the conserved NPA motifs, structural model, the residues forming the selectivity filter and channel radius profile. For selected set of MIPs, it is possible to derive structure-based sequence alignment and evolutionary relationship. Sequences and structures of selected MIPs can be downloaded from MIPModDB database.
Proper citation: MIPModDB (RRID:SCR_006058) Copy
http://prorepeat.bioinformatics.nl/
ProRepeat is an integrated curated repository and analysis platform for in-depth research on the biological characteristics of amino acid tandem repeats. ProRepeat collects repeats from all proteins included in the UniProt knowledgebase, together with 85 completely sequenced eukaryotic proteomes contained within the RefSeq collection. It contains non-redundant perfect tandem repeats, approximate tandem repeats and simple, low-complexity sequences, covering the majority of the amino acid tandem repeat patterns found in proteins. The ProRepeat web interface allows querying the repeat database using repeat characteristics like repeat unit and length, number of repetitions of the repeat unit and position of the repeat in the protein. Users can also search for repeats by the characteristics of repeat containing proteins, such as entry ID, protein description, sequence length, gene name and taxon. ProRepeat offers powerful analysis tools for finding biological interesting properties of repeats, such as the strong position bias of leucine repeats in the N-terminus of eukaryotic protein sequences, the differences of repeat abundance among proteomes, the functional classification of repeat containing proteins and GC content constrains of repeats' corresponding codons.
Proper citation: ProRepeat (RRID:SCR_006113) Copy
http://operons.ibt.unam.mx/OperonPredictor/
The Prokaryotic Operon DataBase (ProOpDB) constitutes one of the most precise and complete repository of operon predictions in our days. Using our novel and highly accurate operon algorithm, we have predicted the operon structures of more than 1,200 prokaryotic genomes. ProOpDB offers diverse alternatives by which a set of operon predictions can be retrieved including: i) organism name, ii) metabolic pathways, as defined by the KEGG database, iii) gene orthology, as defined by the COG database, iv) conserved protein motifs, as defined by the Pfam database, v) reference gene, vi) reference operon, among others. In order to limit the operon output to non-redundant organisms, ProOpDB offers an efficient protocol to select the more representative organisms based on a precompiled phylogenetic distances matrix. In addition, the ProOpDB operon predictions are used directly as the input data of our Gene Context Tool (GeConT) to visualize their genomic context and retrieve the sequence of their corresponding 5�� regulatory regions, as well as the nucleotide or amino acid sequences of their genes. The prediction algorithm The algorithm is a multilayer perceptron neural network (MLP) classifier, that used as input the intergenic distances of contiguous genes and the functional relationship scores of the STRING database between the different groups of orthologous proteins, as defined in the COG database. Nevertheless, the operon prediction of our method is not restricted to only those genes with a COG assignation, since we successfully defined new groups of orthologous genes and obtained, by extrapolation, a set of equivalent STRING-like scores based on conserved gene pairs on different genomes. Since the STRING functional relationships scores are determined in an un-bias manner and efficiently integrates a large amount of information coming from different sources and kind of evidences, the prediction made by our MLP are considerably less influenced by the bias imposed in the training procedure using one specific organism.
Proper citation: ProOpDB (RRID:SCR_006111) Copy
High quality ribosomal RNA databases providing comprehensive, quality checked and regularly updated datasets of aligned small (16S/18S, SSU) and large subunit (23S/28S, LSU) ribosomal RNA (rRNA) sequences for all three domains of life (Bacteria, Archaea and Eukarya). Supplementary services include a rRNA gene aligner, online tools for probe and primer evaluation and optimized browsing, searching and downloading on the website. The extensively curated SILVA taxonomy and the new non-redundant SILVA datasets provide an ideal reference for high-throughput classification of data from next-generation sequencing approaches. Alignment tool, SINA, is available for download as well as available for use online.
Proper citation: SILVA (RRID:SCR_006423) Copy
ViralZone is a SIB Swiss Institute of Bioinformatics web-resource for all viral genus and families, providing general molecular and epidemiological information, along with virion and genome figures. Each virus or family page gives an easy access to UniProtKB/Swiss-Prot viral protein entries. ViralZone project is handled by the virus program of SwissProt group. Proteins popups were developed in collaboration with Prof. Christian von Mering and Andrea Franceschini, Bioinformatics Group , Institute of Molecular Life Sciences, University of Zurich, Winterthurerstrasse 190, CH-8057 Zurich, Switzerland, funded in part by the SIB Swiss Institute of bioinformatics. All pictures in ViralZone are copyright of the SIB Swiss Institute of Bioinformatics.
Proper citation: ViralZone (RRID:SCR_006563) Copy
http://compbio.soe.ucsc.edu/yeast_introns.html
Database of information about the spliceosomal introns of the yeast Saccharomyces cerevisiae. Listed are known spliceosomal introns in the yeast genome and the splice sites actually used are documented. Through the use of microarrays designed to monitor splicing, they are beginning to identify and analyze splice site context in terms of the nature and activities of the trans-acting factors that mediate splice site recognition. In version 3.0, expression data that relates to the efficiency of splicing relative to other processes in strains of yeast lacking nonessential splicing factors is included. These data are displayed on each intron page for browsing and can be downloaded for other types of analysis.
Proper citation: Yeast Intron Database (RRID:SCR_007144) Copy
http://www.broadinstitute.org/annotation/tetraodon/
This database have been funded by the National Human Genome Research Institute (NHGRI) to produce shotgun sequence of the Tetraodon nigriviridis genome. The strategy involves Whole Genome Shotgun (WGS) sequencing, in which sequence from the entire genome is generated. Whole genome shotgun libraries were prepared from Tetraodon genomic DNA obtained from the laboratory of Jean Weissenbach at Genoscope. Additional sequence data of approximately 2.5X coverage of Tetraodon has also been generated by Genoscope in plasmid and BAC end reads. Broad and Genoscope intend to pool their data and generate whole genome assemblies. Tetraodon nigroviridis is a freshwater pufferfish of the order Tetraodontiformes and lives in the rivers and estuaries of Indonesia, Malaysia and India. This species is 20-30 million years distant from Fugu rubripes, a marine pufferfish from the same family. The gene repertoire of T. nigroviridis is very similar to that of other vertebrates. However, its relatively small genome of 385 Mb is eight times more compact than that of human, mostly because intergenic and intronic sequences are reduced in size compared to other vertebrate genomes. These genome characteristics along with the large evolutionary distance between bony fish and mammals make Tetraodon a compact vertebrate reference genome - a powerful tool for comparative genetics and for quick and reliable identification of human genes.
Proper citation: Tetraodon nigroviridis Database (RRID:SCR_007123) Copy
http://www-personal.umich.edu/~jianghui/rseq/
A software toolkit for RNA sequence data analysis. It contains programs that cover several aspects of RNA-Seq data analysis such as read quality assessment, reference sequence generation, sequence mapping, and gene and isoform expressions estimations.
Proper citation: rSeq (RRID:SCR_000562) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the SPARC SAWG Resources search. From here you can search through a compilation of resources used by SPARC SAWG and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that SPARC SAWG has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on SPARC SAWG then you can log in from here to get additional features in SPARC SAWG such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into SPARC SAWG you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within SPARC SAWG that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.