Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Proteogenomic landscape of squamous cell lung cancer.

Paul A Stewart | Eric A Welsh | Robbert J C Slebos | Bin Fang | Victoria Izumi | Matthew Chambers | Guolin Zhang | Ling Cen | Fredrik Pettersson | Yonghong Zhang | Zhihua Chen | Chia-Ho Cheng | Ram Thapa | Zachary Thompson | Katherine M Fellows | Jewel M Francis | James J Saller | Tania Mesa | Chaomei Zhang | Sean Yoder | Gina M DeNicola | Amer A Beg | Theresa A Boyle | Jamie K Teer | Yian Ann Chen | John M Koomen | Steven A Eschrich | Eric B Haura
Nature communications | 2019

How genomic and transcriptomic alterations affect the functional proteome in lung cancer is not fully understood. Here, we integrate DNA copy number, somatic mutations, RNA-sequencing, and expression proteomics in a cohort of 108 squamous cell lung cancer (SCC) patients. We identify three proteomic subtypes, two of which (Inflamed, Redox) comprise 87% of tumors. The Inflamed subtype is enriched with neutrophils, B-cells, and monocytes and expresses more PD-1. Redox tumours are enriched for oxidation-reduction and glutathione pathways and harbor more NFE2L2/KEAP1 alterations and copy gain in the 3q2 locus. Proteomic subtypes are not associated with patient survival. However, B-cell-rich tertiary lymph node structures, more common in Inflamed, are associated with better survival. We identify metabolic vulnerabilities (TP63, PSAT1, and TFRC) in Redox. Our work provides a powerful resource for lung SCC biology and suggests therapeutic opportunities based on redox metabolism and immune cell infiltrates.

Pubmed ID: 31395880

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


GATK (tool)

RRID:SCR_001876

A software package to analyze next-generation resequencing data. The toolkit offers a wide variety of tools, with a primary focus on variant discovery and genotyping as well as strong emphasis on data quality assurance. Its robust architecture, powerful processing engine and high-performance computing features make it capable of taking on projects of any size. This software library makes writing efficient analysis tools using next-generation sequencing data very easy, and second it's a suite of tools for working with human medical resequencing projects such as 1000 Genomes and The Cancer Genome Atlas. These tools include things like a depth of coverage analyzers, a quality score recalibrator, a SNP/indel caller and a local realigner. (entry from Genetic Analysis Software)

View all literature mentions

RefSeq (tool)

RRID:SCR_003496

Collection of curated, non-redundant genomic DNA, transcript RNA, and protein sequences produced by NCBI. Provides a reference for genome annotation, gene identification and characterization, mutation and polymorphism analysis, expression studies, and comparative analyses. Accessed through the Nucleotide and Protein databases.

View all literature mentions

RSeQC (tool)

RRID:SCR_005275

Software package to comprehensively evaluate different aspects of RNA-seq experiments, such as sequence quality, GC bias, polymerase chain reaction bias, nucleotide composition bias, sequencing depth, strand specificity, coverage uniformity and read distribution over the genome structure. RSeQC takes both SAM and BAM files as input, which can be produced by most RNA-seq mapping tools as well as BED files, which are widely used for gene models.

View all literature mentions

HTSeq (tool)

RRID:SCR_005514

THIS RESOURCE IS NO LONGER IN SERVICE. Documented on February 28,2023. Software Python package that provides infrastructure to process data from high-throughput sequencing assays. While the main purpose of HTSeq is to allow you to write your own analysis scripts, customized to your needs, there are also a couple of stand-alone scripts for common tasks that can be used without any Python knowledge.

View all literature mentions

Picard (tool)

RRID:SCR_006525

Java toolset for working with next generation sequencing data in the BAM format.

View all literature mentions

1000 Genomes Project and AWS (tool)

RRID:SCR_008801

A dataset containing the full genomic sequence of 1,700 individuals, freely available for research use. The 1000 Genomes Project is an international research effort coordinated by a consortium of 75 companies and organizations to establish the most detailed catalogue of human genetic variation. The project has grown to 200 terabytes of genomic data including DNA sequenced from more than 1,700 individuals that researchers can now access on AWS for use in disease research free of charge. The dataset containing the full genomic sequence of 1,700 individuals is now available to all via Amazon S3. The data can be found at: http://s3.amazonaws.com/1000genomes The 1000 Genomes Project aims to include the genomes of more than 2,662 individuals from 26 populations around the world, and the NIH will continue to add the remaining genome samples to the data collection this year. Public Data Sets on AWS provide a centralized repository of public data hosted on Amazon Simple Storage Service (Amazon S3). The data can be seamlessly accessed from AWS services such Amazon Elastic Compute Cloud (Amazon EC2) and Amazon Elastic MapReduce (Amazon EMR), which provide organizations with the highly scalable compute resources needed to take advantage of these large data collections. AWS is storing the public data sets at no charge to the community. Researchers pay only for the additional AWS resources they need for further processing or analysis of the data. All 200 TB of the latest 1000 Genomes Project data is available in a publicly available Amazon S3 bucket. You can access the data via simple HTTP requests, or take advantage of the AWS SDKs in languages such as Ruby, Java, Python, .NET and PHP. Researchers can use the Amazon EC2 utility computing service to dive into this data without the usual capital investment required to work with data at this scale. AWS also provides a number of orchestration and automation services to help teams make their research available to others to remix and reuse. Making the data available via a bucket in Amazon S3 also means that customers can crunch the information using Hadoop via Amazon Elastic MapReduce, and take advantage of the growing collection of tools for running bioinformatics job flows, such as CloudBurst and Crossbow.

View all literature mentions

ComBat (tool)

RRID:SCR_010974

Adjusting batch effects in microarray expression data using Empirical Bayes methods.

View all literature mentions

CoMet (tool)

RRID:SCR_011925

A web-server for fast comparative functional profiling of metagenomes.

View all literature mentions

ProteoWizard (tool)

RRID:SCR_012056

Software that enables rapid tool creation by providing a robust, pluggable development framework that simplifies and unifies data file access, and performs standard proteomics and LCMS dataset computations.

View all literature mentions

cBioPortal (tool)

RRID:SCR_014555

A portal that provides visualization, analysis and download of large-scale cancer genomics data sets.

View all literature mentions

Pfam (tool)

RRID:SCR_004726

A database of protein families, each represented by multiple sequence alignments and hidden Markov models (HMMs). Users can analyze protein sequences for Pfam matches, view Pfam family annotation and alignments, see groups of related families, look at the domain organization of a protein sequence, find the domains on a PDB structure, and query Pfam by keywords. There are two components to Pfam: Pfam-A and Pfam-B. Pfam-A entries are high quality, manually curated families that may automatically generate a supplement using the ADDA database. These automatically generated entries are called Pfam-B. Although of lower quality, Pfam-B families can be useful for identifying functionally conserved regions when no Pfam-A entries are found. Pfam also generates higher-level groupings of related families, known as clans (collections of Pfam-A entries which are related by similarity of sequence, structure or profile-HMM).

View all literature mentions

Agilent TapeStation Laptop (tool)

RRID:SCR_019547

TapeStation Laptop is used to standardize data acquisition and analysis.

View all literature mentions

BioAnalyzer 2100 (tool)

RRID:SCR_019715

2100 Bioanalyzer system is an established automated electrophoresis tool for the sample quality control of biomolecules. The 2100 Bioanalyzer instrument, together with the 2100 Expert Software and Bioanalyzer assays, provide highly precise analytical evaluation of various samples types in many workflows, including next generation sequencing (NGS), gene expression, biopharmaceutical, and gene editing research. Digital data is provided in a timely manner and delivers objective assessment of sizing, quantitation, integrity and purity from DNA, RNA, and proteins. Minimal sample volumes are required for an accurate result, and the data may be exported in a many different formats for ease-of-use.

View all literature mentions