Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
Integration of public genome-wide gene expression data together with Cox regression analysis is a powerful weapon to identify new prognostic gene signatures for cancer diagnosis and prognosis. Hepatitis B virus (HBV) is a major cause of hepatocellular carcinoma (HCC), however, it remains largely unknown about the specific gene prognostic signature of HBV-associated HCC. Using Robust Rank Aggreg (RRA) method to integrate seven whole genome expression datasets, we identified 82 up-regulated genes and 577 down-regulated genes in HBV-associated HCC patients. Combination of several enrichment analysis, univariate and multivariate Cox proportional hazards regression analysis, we revealed that a three-gene (SPP2, CDC37L1, and ECHDC2) prognostic signature could act as an independent prognostic indicator for HBV-associated HCC in both the discovery cohort and the internal testing cohort. Gene set enrichment analysis showed that the high-risk group with lower expression levels of the three genes was enriched in bladder cancer and cell cycle pathway, whereas the low-risk group with higher expression levels of the three genes was enriched in drug metabolism-cytochrome P450, PPAR signaling pathway, fatty acid and histidine metabolisms. This indicates that patients of HBV-associated HCC with higher expression of these three genes may preserve relatively good hepatic cellular metabolism and function, which may also protect HCC patients from persistent drug toxicity in response to various medication. Our findings suggest a three-gene prognostic model that serves as a specific prognostic signature for HBV-associated HCC.
Pubmed ID: 29896284
Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.
Integrated database resource consisting of 16 main databases, broadly categorized into systems information, genomic information, and chemical information. In particular, gene catalogs in completely sequenced genomes are linked to higher-level systemic functions of cell, organism, and ecosystem. Analysis tools are also available. KEGG may be used as reference knowledge base for biological interpretation of large-scale datasets generated by sequencing and other high-throughput experimental technologies.
View all literature mentionsFunctional genomics data repository supporting MIAME-compliant data submissions. Includes microarray-based experiments measuring the abundance of mRNA, genomic DNA, and protein molecules, as well as non-array-based technologies such as serial analysis of gene expression (SAGE) and mass spectrometry proteomic technology. Array- and sequence-based data are accepted. Collection of curated gene expression DataSets, as well as original Series and Platform records. The database can be searched using keywords, organism, DataSet type and authors. DataSet records contain additional resources including cluster tools and differential expression queries.
View all literature mentionsSoftware platform for complex network analysis and visualization. Used for visualization of molecular interaction networks and biological pathways and integrating these networks with annotations, gene expression profiles and other state data.
View all literature mentionsSoftware package for interpreting gene expression data. Used for interpretation of a large-scale experiment by identifying pathways and processes.
View all literature mentionsDatabase of known and predicted protein interactions. The interactions include direct (physical) and indirect (functional) associations and are derived from four sources: Genomic Context, High-throughput experiments, (Conserved) Coexpression, and previous knowledge. STRING quantitatively integrates interaction data from these sources for a large number of organisms, and transfers information between these organisms where applicable. The database currently covers 5''214''234 proteins from 1133 organisms. (2013)
View all literature mentionsSoftware package for the analysis of gene expression microarray data, especially the use of linear models for analyzing designed experiments and the assessment of differential expression.
View all literature mentionsBioconductor software package for Empirical analysis of Digital Gene Expression data in R. Used for differential expression analysis of RNA-seq and digital gene expression data with biological replication.
View all literature mentionsFunctional genomics data repository supporting MIAME-compliant data submissions. Includes microarray-based experiments measuring the abundance of mRNA, genomic DNA, and protein molecules, as well as non-array-based technologies such as serial analysis of gene expression (SAGE) and mass spectrometry proteomic technology. Array- and sequence-based data are accepted. Collection of curated gene expression DataSets, as well as original Series and Platform records. The database can be searched using keywords, organism, DataSet type and authors. DataSet records contain additional resources including cluster tools and differential expression queries.
View all literature mentions