We support boolean queries, use +,-,<,>,~,* to alter the weighting of terms
To test a sample population of genes for overrepresentation of GO terms, the R/BioC function GOHyperGAll computes for all GO nodes a hypergeometric distribution test and returns the corresponding p-values. A subsequent filter function performs a GO Slim analysis using default or custom GO Slim categories. Basic knowledge about R and BioConductor is required for using this tool. Platform: Windows compatible, Mac OS X compatible, Linux compatible, Unix compatible, THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
A portal for hormone-related health information for the public, physicians, allied health professionals and the media. It serves as a resource for the public by promoting the prevention, treatment and cure of hormone-related conditions through outreach and education. It provides free educational materials, public forums, physician referral service, and media education campaigns. It offers a library of educational materials and programs covering a wide range of endocrine topics, including adrenal disorders, breast cancer, diabetes, osteoporosis, stress, thyroid disease and cancer.
The Peptide Sequence Database contains putative peptide sequences from human, mouse, rat, and zebrafish. Compressed to eliminate redundancy, these are about 40 fold smaller than a brute force enumeration. Current and old releases are available for download. Each species'' peptide sequence database comprises peptide sequence data from releveant species specific UniGene and IPI clusters, plus all sequences from their consituent EST, mRNA and protein sequence databases, namely RefSeq proteins and mRNAs, UniProt''s SwissProt and TrEMBL, GenBank mRNA, ESTs, and high-throughput cDNAs, HInv-DB, VEGA, EMBL, IPI protein sequences, plus the enumeration of all combinations of UniProt sequence variants, Met loss PTM, and signal peptide cleavages. The README file contains some information about the non amino-acid symbols O (digest site corresponding to a protein N- or C-terminus) and J (no digest sequence join) used in these peptide sequence databases and information about how to configure various search engines to use them. Some search engines handle (very) long sequences badly and in some cases must be patched to use these peptide sequence databases. All search engines supported by the PepArML meta-search engine can (or can be patched to) successfully search these peptide sequence databases.
The PeptideMapper Web-Service provides alignments of peptide sequence alignments to proteins, mRNA, EST, and HTC sequences from Genbank, RefSeq, UniProt, IPI, VEGA, EMBL, and HInvDb. This mapping infrastructure is supported, in part, by the compressed peptide sequence database infrastructure (Edwards, 2007) which enables a fast, suffix-tree based mapping of peptide sequences to gene identifiers and a gene-focused detailed mapping of peptide sequences to source sequence evidence. The PeptideMapper Web-Service can be used interactively or as a web-service using either HTTP or SOAP requests. Results of HTTP requests can be returned in a variety of formats, including XML, JSON, CSV, TSV, or XLS, and in some cases, GFF or BED; results of SOAP requests are returned as SOAP responses. The PeptideMapper Web-Service maps at most 20 peptides with length between 5 and 30 amino-acids in each request. The number of alignments returned, per peptide, gene, and sequence type, is set to 10 by default. The default can be changed on the interactive alignments search form or by using the max web-service parameter.
A web server that predicts the functional impact of amino-acid substitutions in proteins, such as mutations discovered in cancer or nonsynonymous polymorphisms. The functional impact is assessed based on evolutionary conservation of the affected amino acid in protein homologs. The method has been validated on a large set (51k) of disease associated (OMIM) and polymorphic variants., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
ALCHEMY is a genotype calling algorithm for Affymetrix and Illumina products which is not based on clustering methods. Features include explicit handling of reduced heterozygosity due to inbreeding and accurate results with small sample sizes. ALCHEMY is a method for automated calling of diploid genotypes from raw intensity data produced by various high-throughput multiplexed SNP genotyping methods. It has been developed for and tested on Affymetrix GeneChip Arrays, Illumina GoldenGate, and Illumina Infinium based assays. Primary motivations for ALCHEMY''s development was the lack of available genotype calling methods which can perform well in the absence of heterozygous samples (due to panels of inbred lines being genotyped) or provide accurate calls with small sample batches. ALCHEMY differs from other genotype calling methods in that genotype inference is based on a parametric Bayesian model of the raw intensity data rather than a generalized clustering approach and the model incorporates population genetic principles such as Hardy-Weinberg equilibrium adjusted for inbreeding levels. ALCHEMY can simultaneously estimate individual sample inbreeding coefficients from the data and use them to improve statistical inference of diploid genotypes at individual SNPs. The main documentation for ALCHEMY is maintained on the sourceforge-hosted MediaWiki system. Features * Population genetic model based SNP genotype calling * Simultaneous estimation of per-sample inbreeding coefficients, allele frequencies, and genotypes * Bayesian model provides posterior probabilities of genotype correctness as quality measures * Growing number of scripts and supporting programs for validation of genotypes against control data and output reformating needs * Multithreaded program for parallel execution on multi-CPU/core systems * Non-clustering based methods can handle small sample sets for empirical optimization of sample preparation techniques and accurate calling of SNPs missing genotype classes ALCHEMY is written in C and developed on the GNU/Linux platform. It should compile on any current GNU/Linux distribution with the development packages for the GNU Scientific Library (gsl) and other development packages for standard system libraries. It may also compile and run on Mac OS X if gsl is installed.
Database of gene expression in the marmoset brain.Comparative anatomy of marmoset and mouse cortex from genomic expression. Atlas comparing brain of neonatal marmoset with mouse using in situ hybridization.
A ChIP-Seq peak calling or differential binding analysis tool that is primarily designed for data with biological replicates. It uses a negative binomial distribution to model the read counts among the samples in the same group, and look for consistent differences between ChIP and control group or two ChIP groups run under different conditions.
Collect, share, and distribute information about protein three-dimensional structures. It serves as a portal for the scientific community to learn about protein structures solved by SG centers, and also to contribute their expertise in annotating protein function. The premise of the TOPSAN project is that, no matter how much any individual knows about a particular protein, there are other members of the scientific community who know more about certain aspects of the same protein, and that the collective analyses from experts will be far more informative than any local group, let alone individual, could contribute. They believe that, if the members of the biological community are given the opportunity, authorship incentives, and an easy way to contribute their knowledge to the structure annotation, they would do so. Therefore, borrowing elements from successful, distributed, collaborative projects, such as Wikipedia (the free encyclopedia anyone can edit) and from other open source software development projects, TOPSAN will be a broad, collaborative effort to annotate protein structures, initially, those determined at the JCSG. They believe that the annotation of proteins solved by structural genomics consortia offers a unique opportunity to challenge the extant paradigm of how biological data is collected and distributed, and to connect structural genomics and structural biology to the entire biological research community. TOPSAN is designed to be scalable, modular and extensible. Furthermore, it is intended to be immediately useful in a simplistic way and will accommodate incremental improvements to functionality as usage becomes more sophisticated. Their annotation pages will offer the end user a combination of automatically generated as well as expert-curated annotations of protein structures. They will use available technology to increase the speed and granularity of the exchange of scientific ideas, and use incentive mechanisms that will encourage collaborative participation.
Software that utilizes a multiobjective evolutionary algorithm for genetic mapping. It is based on a the ECJ evolutionary software package written by Sean Luke and includes the Strength Pareto Evoluationary Algorithm Version 2 changes for multiobjective analysis. The code runs on any platform with Java Version 2. A genetic mapping project, typically implemented during a search for genes responsible for a disease, requires the acquisition of a set of data from each of a large number of individuals. This data set includes the values of multiple genetic markers. These genetic markers occur at discrete positions along the genome, which is a collection of one or more linear chromosomes. Typing the value of a marker in an individual carries a cost; one seeks to minimize the number of markers typed without excessively jeopardizing the probability of detecting an association between a marker and a disease phenotype. MAGMA is a project which employ''s a multiobjective evolutionary algorithm to solve this problem.
The American Cancer Society is the nationwide, community-based, voluntary health organization dedicated to eliminating cancer as a major health problem by preventing cancer, saving lives, and diminishing suffering from cancer, through research, education, advocacy, and service. Together with our millions of supporters, the American Cancer Society (ACS) saves lives and creates a world with less cancer and more birthdays by helping people stay well, helping people get well, by finding cures, and by fighting back. Headquartered in Atlanta, Georgia, the ACS has 12 chartered Divisions, more than 900 local offices nationwide, and a presence in more than 5,100 communities.
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on August 20,2019.Database and analysis environment for experimentally determined binding sites of RNA-binding proteins. It supports the automatic functional annotation of short reads resulting primarily from crosslinking and immunoprecipitation experiments (CLIP) performed with RNA-binding proteins in order to identify the binding sites of these proteins. The functional annotation could be also applied to short reads resulting from other types of experiments such as mRNA-Seq, Digital Gene Expression, small RNA cloning, etc. The platform enables visualization and mining of individual data sets as well as analysis involving multiple experimental data sets. The platform can support collaborative projects involving multiple users and groups of users as well as public and private datasets.
A database of oncogenes and tumor suppressor genes. Users can search by genes, chromosomes, and keywords. The coAnsensus domain analysis tool functions to identify conserved protein domains and GO terms among selected TAG genes, while the &ldquo;oncogenic domain analysis&rdquo; can analyze oncogenic potential of any user-provided protein based on a weighed term frequency table calculated from the TAG proteins. The completion of human genome sequences allows one to rapidly identify and analyze genes of interest through the use of computational approach. The available annotations including physical characterization and functional domains of known tumor-related genes thus can be used to study the role of genes involved in carcinogenesis. The tumor-associated gene (TAG) database was designed to utilize information from well-characterized oncogenes and tumor suppressor genes to facilitate cancer research. All target genes were identified through text-mining approach from the PubMed database. A semi-automatic information retrieving engine was built to collect specific information of these target genes from various resources and store in the TAG database. At current stage, 519 TAGs including 198 oncogenes, 170 tumor suppressor genes, and 151 genes related to oncogenesis were collected. Information collected in TAG database can be browsed through user-friendly web interfaces that provide searching genes by chromosome or by keywords. The &ldquo;consensus domain analysis&rdquo; tool functions to identify conserved protein domains and GO terms among selected TAG genes. In addition, the &ldquo;oncogenic domain analysis&rdquo; can analyze oncogenic potential of any user-provided protein based on a weighed term frequency table calculated from the TAG proteins. This study was supported by grant from National research program for genomic medicine (NRPGM) and personnel from Bioinformatics Center of Center for Biotechnology and Biosciences in the National Cheng Kung University, Taiwan.
The Human Adenovirus Type Classification coordinates the naming of candidate new types, prior to manuscript submission for peer review. This resource contains a method of submitting candidate HAdV, criteria for a new HAdV type, and a Serotyping tool, which displays all potential types corresponding to the query serotype entered by a user. The criteria are based on discussions at the International Adenovirus Meeting (Dobog��k, Hungary; 26-30 April, 2009) and the NIH Human Adenovirus Working Group Workshop (Bethesda, MD. USA; 3 February 2011), which are summarized in a Letter to the Editor.
THIS RESOURCE IS NO LONGER IN SERVICE, documented on July 10, 2012. Cluster Assignment for Biological Inference (CLASSIFI) is a data-mining tool that can be used to identify significant co-clustering of genes with similar functional properties (e.g. cellular response to DNA damage). Briefly, CLASSIFI uses the Gene Ontology gene annotation scheme to define the functional properties of all genes/probes in a microarray data set, and then applies a cumulative hypergeometric distribution analysis to determine if any statistically significant gene ontology co-clustering has occurred. Platform: Online tool
Opasnet is a wiki-based website and workspace for helping societal decision making. The website collects, synthesizes, and distributes people''s values and scientific information. Opasnet welcomes anyone who wants to promote science-based decision-making in any field. The specialty is that the information is structured for both scientific scrutiny and for policy use at the same time. In practice, you can do original research, store data, make models, and perform policy assessments and discuss all of that work in one workspace. Originally, the developers of Opasnet came from the environmental health, i.e. a research field that studies the impacts of environment on human health. We are actively working, among other things, on climate change and air pollution, but you can also start a new assessment about a decision of your own interest, or participate in an existing assessment. Opasnet is a website that has basically two parts. One part is a wiki site (called Opasnet wiki or simply Opasnet) that has descriptive pages with text, figures, and tables; it also contains files. The other part is a database called Opasnet Base that contains quantitative estimates about anything that is described in Opasnet. The majority of information is openly available. However, both Opasnet wiki and Opasnet Base have a protected area for working with material that is non-public for some reason.
omniBiomarker is a web-application for analysis of high-throughput -omic data. Its primary function is to identify differentially expressed biomarkers that may be used for diagnostic or prognostic clinical prediction. Currently, omniBiomarker allows users to analyze their data with many different ranking methods simultaneously using a high-performance compute cluster. The next release of omniBiomarker will automatically select the most biologically relevant ranking method based on user input regarding prior knowledge. The omniBiomarker workflow * Data: Gene Expression * Algorithms: Knowledge-Driven Gene Ranking * Differentially expressed Genes * Clinical / Biological Validation * Knowledge: NCI Thesaurus of Cancer, Cancer Gene Index * back to Algorithms
Public research university in Louisville, Kentucky. It is part of the Kentucky state university system.
A Cytoscape plug-in that visualizes the non-redundant biological terms for large clusters of genes in a functionally grouped network. It can be used in combination with GOlorize. The identifiers can be uploaded from a text file or interactively from a network of Cytoscape. The type of identifiers supported can be easily extended by the user. ClueGO performs single cluster analysis and comparison of clusters. From the ontology sources used, the terms are selected by different filter criteria. The related terms which share similar associated genes can be combined to reduce redundancy. The ClueGO network is created with kappa statistics and reflects the relationships between the terms based on the similarity of their associated genes. On the network, the node colour can be switched between functional groups and clusters distribution. ClueGO charts are underlying the specificity and the common aspects of the biological role. The significance of the terms and groups is automatically calculated. ClueGO is easy updatable with the newest files from Gene Ontology and KEGG. Platform: Windows compatible, Mac OS X compatible, Linux compatible, Unix compatible, THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.