We support boolean queries, use +,-,<,>,~,* to alter the weighting of terms
An algorithmic software tool for detecting known viruses and their integration sites using next-generation sequencing of human cancer tissue. VirusSeq takes FASTQ files (paired-end reads) as input.
Software tool for efficient and accurate detection of viruses and their integration sites in host genomes through next generation sequencing data. Specifically, it detects virus infection, co-infection with multiple viruses, virus integration sites in host genomes, as well as mutations in the virus genomes. It also facilitates virus discovery by reporting novel contigs, long sequences assembled from short reads that map neither to the host genome nor to the genomes of known viruses. VirusFinder 2 works with both paired-end and single-end data, unlike the previous 1.x versions that accepted only paired-end reads. The types of NGS data that VirusFinder 2 can deal with include whole genome sequencing (WGS), whole transcriptome sequencing (RNA-Seq), targeted sequencing data such as whole exome sequencing (WES) and ultra-deep amplicon sequencing.
A highly scalable parallel software program to identify non-host sequences (of potential pathogen origin) and estimate their genome relative abundance in high-throughput sequence datasets.
A computational tool for the identification and analysis of microbial sequences in high-throughput human sequencing data that is designed to work with large numbers of sequencing reads in a scalable manner. This process is composed of a subtractive phase in which input reads are subtracted by alignment to human reference sequences, and an analytic phase in which the remaining reads are aligned to microbial reference sequences (viral, fungal, bacterial, archaeal) and de novo assembled. PathSeq is currently available in a cloud computing environment via Amazon Web Services The typical approach one would take to pathogen discovery with PathSeq: RNA or DNA is extracted from the tissue of interest and sequencing libraries are constructed to be run on the next-generation DNA sequencing platform of choice. The resulting sequence data is run through the PathSeq pipeline in a cloud computing environment. PathSeq reports potential microbes in the sequence data as well as the complete set of reads that could not be identified as human or microbial sequences.
A blog that explores local science, nature, and environment issues & experiences in Northern California. A collaborative effort, our many writers come from local museums, zoos, science centers and research institutions, as well as KQED''s TV and Radio producers covering stories in the field.
Informatics software tool to identify patient sequences that are too similar to happen by chance alone. Highly similar sequences are likely to occur from contamination or other situations like geographic linkage.
ScienceBlogs posts about Jobs.
Software that characterizes coexisting subpopulations (SPs) in a tumor using copy number and allele frequencies derived from exome- or whole genome sequencing input data. The model amplifies the statistical power to detect coexisting genotypes, by fully exploiting run-specific tradeoffs between depth of coverage and breadth of coverage. ExPANdS predicts the number of clonal expansions, the size of the resulting SPs in the tumor bulk, the mutations specific to each SP and tumor purity. The main function runExPANdS provides the complete functionality needed to predict coexisting SPs from single nucleotide variations (SNVs) and associated copy numbers. The robustness of the subpopulation predictions by ExPANdS increases with the number of mutations provided. It is recommended that at least 200 mutations are used as an input to obtain stable results.
Software to estimate purity / ploidy, and from that compute absolute copy-number and mutation multiplicities. When DNA is extracted from an admixed population of cancer and normal cells, the information on absolute copy number per cancer cell is lost in the mixing. The purpose of ABSOLUTE is to re-extract these data from the mixed DNA population. This process begins by generation of segmented copy number data, which is input to the ABSOLUTE algorithm together with pre-computed models of recurrent cancer karyotypes and, optionally, allelic fraction values for somatic point mutations. The output of ABSOLUTE then provides re-extracted information on the absolute cellular copy number of local DNA segments and, for point mutations, the number of mutated alleles.
From climate change to intelligent design, HIV/AIDS to stem cells, science education to space exploration, science is figuring prominently in our discussions of politics, religion, philosophy, business and the arts. New insights and discoveries in neuroscience, theoretical physics and genetics are revolutionizing our understanding of who are are, where we come from and where we''re heading. Launched in January 2006, ScienceBlogs is a portal to this global dialogue, a digital science salon featuring the leading bloggers from a wide array of scientific disciplines. Today, ScienceBlogs is the largest online community dedicated to science. We believe in providing our bloggers with the freedom to exercise their own editorial and creative instincts. We do not edit their work and we do not tell them what to write about. We have selected our 80+ bloggers based on their originality, insight, talent, and dedication and how we think they would contribute to the discussion at ScienceBlogs. Our role, as we see it, is to create and continue to improve this forum for discussion, and to ensure that the rich dialogue that takes place at ScienceBlogs resonates outside the blogosphere. ScienceBlogs is always interested in bringing new contributors into our community. If you''re interested in blogging with us, please fill out our application, and we''ll be in touch.
A database for comparative genomics of group A and group B streptococci. It is based on OGeR (Open Genome Resource for comparative analysis of prokaryotic genomes) and includes all sequenced GAS and GBS strains and serovars available as EMBL genome review or NCBI GenBank files. Strepto-DB identifies the homologous proteins deduced from the genomes of interest. It allows for the elucidation of the GAS and GBS core- and pan-genomes via genome-wide comparisons. Moreover, an intergenic region analysis tool provides alignments and predictions for transcription factor binding sites in the non-coding sequences. An interactive genome browser visualizes functional annotations. Strepto-DB (http://oger.tu-bs.de/strepto_db) was created by the use of OGeR, the Open Genome Resource for comparative analysis of prokaryotic genomes. OGeR is a newly developed open source database and tool platform for the web-based storage, distribution, visualization and comparison of prokaryotic genome data. The system automatically creates the dedicated relational database and web interface and imports an arbitrary number of genomes derived from standardized genome files.
From the editors and reporters of Scientific American, this blog delivers commentary, opinion and analysis on the latest developments in science and technology and their influence on society and policy. From reasoned arguments and cultural critiques to personal and skeptical takes on interesting science news, you''ll find a wide range of scientifically relevant insights here.
Analysis tool that can report the functional properties of any variant in all the human, mouse or rat genes (and soon new model organisms will be added) and the corresponding neighborhoods. Also other non-coding extra-genic regions, such as miRNAs are included in the analysis. It not only reports the obvious functional effects in the coding regions but also analyzes noncoding SNVs situated both within the gene and in the neighborhood that could affect different regulatory motifs, splicing signals, and other structural elements. These include: Jaspar regulatory motifs, miRNA targets, splice sites, exonic splicing silencers, calculations of selective pressures on the particular polymorphic positions, etc. Software analysis pipelines used in the analysis of NGS data are highly modular, heterogeneous, and rapidly evolving. VARIANT can easily be incorporated into a NGS resequencing pipeline either as a CLI or invoked a webservice. It inputs data directly from the most widely used programs for SNV detection., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
A web-based tool for using biological databases to prioritize single nucleotide polymorphisms (SNPs) after a genome-wide association study (GWAS). The site allows users to upload a list of SNPs and GWAS P-values and returns a prioritized list of SNPs using the GIN method. Users can specify candidate genes or genomic regions with custom levels of prioritization. The results can be downloaded or viewed in the browser where users can interactively explore the details of each SNP, including graphical representations of the genomic information network (GIN) method. For investigators interested in incorporating biological databases into a post-GWAS SNP selection strategy, the SPOT web tool is an easily implemented and flexible solution.
A web server for functional annotation of novel and publicly known genetic variants that was developed to assess the potential significance of known and novel SNPs on the major transcriptome, proteome, regulatory and structural variation models in order to identify the phenotypically important variants. A broader range of variations have been incorporated such as insertions / deletions, block substitutions, IUPAC codes submission and region-based analysis, expanding the query size limit, and most importantly including additional categories for the assessment of functional impact. SNPnexus provides a comprehensive set of annotations for genomic variation data by characterizing related functional consequences at the transcriptome/proteome levels of seven major annotation systems with in-depth analysis of potential deleterious effects, inferring physical and cytogenetic mapping, reporting information on HapMap genotype/allele data, finding overlaps with potential regulatory elements, structural variations and conserved elements, and retrieving links with previously reported genetic disease studies.
Genetic variant annotation and effect prediction software toolbox that annotates and predicts effects of variants on genes (such as amino acid changes). By using standards, such as VCF, SnpEff makes it easy to integrate with other programs.
A database to fill the annotation gap left by the high cost of experimental testing for functional significance of protein variants. It joins related bits of knowledge, currently distributed throughout various databases, into a consistent, easily accessible, and updatable resource. It currently covers over 155,000 protein sequences which come from more than 2,600 organisms. Overall more than one million single amino acid substitutions (SAASs) are referenced consisting of natural variants, SAASs from mutagenesis experiments and sequencing conflicts. SNPdbe offers the following pieces of information (if available) on each SAAS: * Experimentally derived functional and structural impact * Predicted functional effect * Associated disease * Average heterozygosity * Experimental evidence of the nsSNP * Evolutionary conservation of wildtype and mutant amino acid * Link-outs to external databases A convenient webinterface to query SAASs on the following levels is offered: * Protein and gene identifiers and keywords * Disease keywords * Protein sequence on different sequence identity thresholds * Variant identifier (dbSNP rs, SwissVar, PMD) or specific mutant like XposY and specified sequence They offer the possibility to submit protein sequences along with experimentally substantiated mutations in order to predict their functional effect and inclusion into our database.
Software package containing tools to process bam files in order to evaluate and analyze de novo assembly / assemblers and identify Structural Variations suspicious genomics regions. The tools have been already successfully applied in several de novo and resequencing projects. This package contains two tools: # FRCbam: tool to compute Feature Response Curves in order to validate and rank assemblies and assemblers # FindTranslocations: tool to identify chromosomal rearrangements using Mate Pairs
A software tool for resolving multi-mappings within an RNA-Seq SAM file.
A simple and easy to use high through-put analysis tool which can provide comprehensive annotation of both novel and known single nucleotide polymorphisms (SNPs) for any organism with a draft sequence and annotation. SNPdat makes possible analyses involving non-model organisms that are not supported by the vast majority of SNP annotation tools currently available. It is especially intended for use by researchers with limited bioinformatic experience.