Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

GeneWalk identifies relevant gene functions for a biological context using network representation learning.

Robert Ietswaart | Benjamin M Gyori | John A Bachman | Peter K Sorger | L Stirling Churchman
Genome biology | 2021

A bottleneck in high-throughput functional genomics experiments is identifying the most important genes and their relevant functions from a list of gene hits. Gene Ontology (GO) enrichment methods provide insight at the gene set level. Here, we introduce GeneWalk ( github.com/churchmanlab/genewalk ) that identifies individual genes and their relevant functions critical for the experimental setting under examination. After the automatic assembly of an experiment-specific gene regulatory network, GeneWalk uses representation learning to quantify the similarity between vector representations of each gene and its GO annotations, yielding annotation significance scores that reflect the experimental context. By performing gene- and condition-specific functional analysis, GeneWalk converts a list of genes into data-driven hypotheses.

Pubmed ID: 33526072

Associated grants

  • Agency: NHGRI NIH HHS, United States
    Id: R01 HG007173

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


cPath (tool)

RRID:SCR_001749

Data management software that runs the Pathway Commons web service. It makes it easy to aggregate custom pathway data sets available in standard exchange formats from multiple databases, present pathway data to biologists via a customizable web interface, and export pathway data via a web service to third-party software, such as Cytoscape, for visualization and analysis. cPath is software only, and does not include new pathway information. Main features: * Import pipeline capable of aggregating pathway and interaction data sets from multiple sources, including: MINT, IntAct, HPRD, DIP, BioCyc, KEGG, PUMA2 and Reactome. * Import/Export support for the Proteomics Standards Initiative Molecular Interaction (PSI-MI) and the Biological Pathways Exchange (BioPAX) XML formats. * Data visualization and analysis via Cytoscape. * Simple HTTP URL based XML web service. * Complete software is freely available for local install. Easy to install and administer. * Partly funded by the U.S. National Cancer Institute, via the Cancer Biomedical Informatics Grid (caBIG) and aims to meet silver-level requirements for software interoperability and data exchange.

View all literature mentions

MEDLINE (tool)

RRID:SCR_002185

A premier bibliographic database that contains over 18 million references to journal articles in life sciences with a concentration on biomedicine. A distinctive feature is that the records are indexed with NLM Medical Subject Headings (MeSH). PubMed provides free access to MEDLINE and links to full text articles when possible. The great majority of journals are selected for MEDLINE based on the recommendation of the Literature Selection Technical Review Committee (LSTRC), an NIH-chartered advisory committee of external experts analogous to the committees that review NIH grant applications. Some additional journals and newsletters are selected based on NLM-initiated reviews, e.g., history of medicine, health services research, AIDS, toxicology and environmental health, molecular biology, and complementary medicine, that are special priorities for NLM or other NIH components. These reviews generally also involve consultation with an array of NIH and outside experts or, in some cases, external organizations with which NLM has special collaborative arrangements. MEDLINE is the primary component of PubMed, part of the Entrez series of databases provided by the NLM National Center for Biotechnology Information (NCBI). MEDLINE may also be searched via the NLM Gateway. Time coverage: generally 1946 to the present, with some older material. Source: Currently, citations from approximately 5,516 worldwide journals in 39 languages; 60 languages for older journals. Citations for MEDLINE are created by the NLM, international partners, and collaborating organizations.

View all literature mentions

HTSeq (tool)

RRID:SCR_005514

THIS RESOURCE IS NO LONGER IN SERVICE. Documented on February 28,2023. Software Python package that provides infrastructure to process data from high-throughput sequencing assays. While the main purpose of HTSeq is to allow you to write your own analysis scripts, customized to your needs, there are also a couple of stand-alone scripts for common tasks that can be used without any Python knowledge.

View all literature mentions

Systems Transcriptional Activity Reconstruction (tool)

RRID:SCR_005622

A next-generation web-based application that aims to provide an integrated solution for both visualization and analysis of deep-sequencing data, along with simple access to public datasets.

View all literature mentions

SciPy (tool)

RRID:SCR_008058

A Python-based environment of open-source software for mathematics, science, and engineering. The core packages of SciPy include: NumPy, a base N-dimensional array package; SciPy Library, a fundamental library for scientific computing; and IPython, an enhanced interactive console.

View all literature mentions

word2vec (tool)

RRID:SCR_014776

Software tool which provides implementation of the continuous bag-of-words and skip-gram architectures for computing vector representations of words. These representations can be used in many natural language processing applications and for further research. It takes a text corpus as input and produces the word vectors as output. It first constructs a vocabulary from the training text data and then learns vector representation of words. The resulting word vector file can be used as features in natural language processing and machine learning applications.

View all literature mentions

DESeq2 (tool)

RRID:SCR_015687

Software package for differential gene expression analysis based on the negative binomial distribution. Used for analyzing RNA-seq data for differential analysis of count data, using shrinkage estimation for dispersions and fold changes to improve stability and interpretability of estimates.

View all literature mentions

biomaRt (tool)

RRID:SCR_019214

Software package that integrates BioMart data resources with data analysis software in Bioconductor. Can annotate range of gene or gene product identifiers including Entrez Gene and Affymetrix probe identifiers with information such as gene symbol, chromosomal coordinates, Gene Ontology and OMIM annotation. Enables retrieval of genomic sequences and single nucleotide polymorphism information, which can be used in data analysis.

View all literature mentions

GeneWalk (tool)

RRID:SCR_023787

Software for individual genes functions determination that are relevant in particular biological context and experimental condition. Quantifies similarity between vector representations of gene and annotated GO terms through representation learning with random walks on condition specific gene regulatory network. Similarity significance is determined through comparison with node similarities from randomized networks.

View all literature mentions

goatools (tool)

RRID:SCR_025305

Software Python library for Gene Ontology analysis. Performs gene ontology enrichment analyses to determine over- and under-represented terms.

View all literature mentions

HeLa S3 (tool)

RRID:CVCL_0058

Cell line HeLa S3 is a Cancer cell line with a species of origin Homo sapiens (Human)

View all literature mentions