Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Genome-wide association study of breast cancer in Latinas identifies novel protective variants on 6q25.

Laura Fejerman | Nasim Ahmadiyeh | Donglei Hu | Scott Huntsman | Kenneth B Beckman | Jennifer L Caswell | Karen Tsung | Esther M John | Gabriela Torres-Mejia | Luis Carvajal-Carmona | María Magdalena Echeverry | Anna Marie D Tuazon | Carolina Ramirez | COLUMBUS Consortium | Christopher R Gignoux | Celeste Eng | Esteban Gonzalez-Burchard | Brian Henderson | Loic Le Marchand | Charles Kooperberg | Lifang Hou | Ilir Agalliu | Peter Kraft | Sara Lindström | Eliseo J Perez-Stable | Christopher A Haiman | Elad Ziv
Nature communications | 2014

The genetic contributions to breast cancer development among Latinas are not well understood. Here we carry out a genome-wide association study of breast cancer in Latinas and identify a genome-wide significant risk variant, located 5' of the Estrogen Receptor 1 gene (ESR1; 6q25 region). The minor allele for this variant is strongly protective (rs140068132: odds ratio (OR) 0.60, 95% confidence interval (CI) 0.53-0.67, P=9 × 10(-18)), originates from Indigenous Americans and is uncorrelated with previously reported risk variants at 6q25. The association is stronger for oestrogen receptor-negative disease (OR 0.34, 95% CI 0.21-0.54) than oestrogen receptor-positive disease (OR 0.63, 95% CI 0.49-0.80; P heterogeneity=0.01) and is also associated with mammographic breast density, a strong risk factor for breast cancer (P=0.001). rs140068132 is located within several transcription factor-binding sites and electrophoretic mobility shift assays with MCF-7 nuclear protein demonstrate differential binding of the G/A alleles at this locus. These results highlight the importance of conducting research in diverse populations.

Pubmed ID: 25327703

Associated grants

  • Agency: NHLBI NIH HHS, United States
    Id: HHSN268201100001I
  • Agency: NCI NIH HHS, United States
    Id: UM1 CA164920
  • Agency: NHLBI NIH HHS, United States
    Id: HHSN268201100046C
  • Agency: Howard Hughes Medical Institute, United States
  • Agency: NCI NIH HHS, United States
    Id: U01CA86117
  • Agency: NIA NIH HHS, United States
    Id: P30-AG15272
  • Agency: NHLBI NIH HHS, United States
    Id: HL078885
  • Agency: NCI NIH HHS, United States
    Id: R01 CA063446
  • Agency: NCI NIH HHS, United States
    Id: R01 CA132839
  • Agency: NIGMS NIH HHS, United States
    Id: T32 GM007377
  • Agency: NHLBI NIH HHS, United States
    Id: HHSN268201100003I
  • Agency: NCI NIH HHS, United States
    Id: K01 CA160607
  • Agency: NIGMS NIH HHS, United States
    Id: R25 GM056765
  • Agency: NCI NIH HHS, United States
    Id: CA77305
  • Agency: NCI NIH HHS, United States
    Id: UM1 CA164973
  • Agency: NIAID NIH HHS, United States
    Id: AI077439
  • Agency: NCI NIH HHS, United States
    Id: 5UM1CA164973
  • Agency: NCI NIH HHS, United States
    Id: P30 CA015704
  • Agency: NHLBI NIH HHS, United States
    Id: HHSN268201100004I
  • Agency: WHI NIH HHS, United States
    Id: HHSN268201100003C
  • Agency: NHLBI NIH HHS, United States
    Id: R01 HL078885
  • Agency: NCI NIH HHS, United States
    Id: R37 CA54281
  • Agency: NIA NIH HHS, United States
    Id: P30 AG015272
  • Agency: NIAID NIH HHS, United States
    Id: U19 AI077439
  • Agency: NIA NIH HHS, United States
    Id: HHSN271201100004C
  • Agency: NCI NIH HHS, United States
    Id: R01 CA063464
  • Agency: NCI NIH HHS, United States
    Id: U01 CA086117
  • Agency: WHI NIH HHS, United States
    Id: HHSN268201100002C
  • Agency: NCI NIH HHS, United States
    Id: K12 CA138464
  • Agency: NHLBI NIH HHS, United States
    Id: HHSN268201100002I
  • Agency: NCI NIH HHS, United States
    Id: R01 CA63464
  • Agency: NCI NIH HHS, United States
    Id: P30 CA013330
  • Agency: NHLBI NIH HHS, United States
    Id: R01 HL088133
  • Agency: NCI NIH HHS, United States
    Id: R37 CA054281
  • Agency: NCI NIH HHS, United States
    Id: R01 CA120120
  • Agency: WHI NIH HHS, United States
    Id: HHSN268201100001C
  • Agency: NCI NIH HHS, United States
    Id: K24 CA169004
  • Agency: NHLBI NIH HHS, United States
    Id: HL088133
  • Agency: WHI NIH HHS, United States
    Id: HHSN268201100004C

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


1000 Genomes: A Deep Catalog of Human Genetic Variation (tool)

RRID:SCR_006828

International collaboration producing an extensive public catalog of human genetic variation, including SNPs and structural variants, and their haplotype contexts, in an effort to provide a foundation for investigating the relationship between genotype and phenotype. The genomes of about 2500 unidentified people from about 25 populations around the world were sequenced using next-generation sequencing technologies. Redundant sequencing on various platforms and by different groups of scientists of the same samples can be compared. The results of the study are freely and publicly accessible to researchers worldwide. The consortium identified the following populations whose DNA will be sequenced: Yoruba in Ibadan, Nigeria; Japanese in Tokyo; Chinese in Beijing; Utah residents with ancestry from northern and western Europe; Luhya in Webuye, Kenya; Maasai in Kinyawa, Kenya; Toscani in Italy; Gujarati Indians in Houston; Chinese in metropolitan Denver; people of Mexican ancestry in Los Angeles; and people of African ancestry in the southwestern United States. The goal Project is to find most genetic variants that have frequencies of at least 1% in the populations studied. Sequencing is still too expensive to deeply sequence the many samples being studied for this project. However, any particular region of the genome generally contains a limited number of haplotypes. Data can be combined across many samples to allow efficient detection of most of the variants in a region. The Project currently plans to sequence each sample to about 4X coverage; at this depth sequencing cannot provide the complete genotype of each sample, but should allow the detection of most variants with frequencies as low as 1%. Combining the data from 2500 samples should allow highly accurate estimation (imputation) of the variants and genotypes for each sample that were not seen directly by the light sequencing. All samples from the 1000 genomes are available as lymphoblastoid cell lines (LCLs) and LCL derived DNA from the Coriell Cell Repository as part of the NHGRI Catalog. The sequence and alignment data generated by the 1000genomes project is made available as quickly as possible via their mirrored ftp sites. ftp://ftp.1000genomes.ebi.ac.uk ftp://ftp-trace.ncbi.nlm.nih.gov/1000genomes

View all literature mentions

PLINK (tool)

RRID:SCR_001757

Open source whole genome association analysis toolset, designed to perform range of basic, large scale analyses in computationally efficient manner. Used for analysis of genotype/phenotype data. Through integration with gPLINK and Haploview, there is some support for subsequent visualization, annotation and storage of results. PLINK 1.9 is improved and second generation of the software.

View all literature mentions

Framingham Heart Study (tool)

RRID:SCR_008963

A longitudinal, epidemiologic study to identify the common risk factors or characteristics that contribute to cardiovascular disease by following its development over a long period of time in a large group of participants who had not yet developed overt symptoms or suffered a heart attack or stroke. Since that time the FHS has studied three generations of participants resulting in biological specimens and data from nearly 15,000 participants. Since 1994, two groups from minority populations, including related individuals have been added to the FHS. FHS welcomes proposals from outside investigators for data and biospecimens. The researchers recruited 5,209 men and women between the ages of 30 and 62 from the town of Framingham, Massachusetts, and began the first round of extensive physical examinations and lifestyle interviews that they would later analyze for common patterns related to CVD development. Since 1948, the subjects have continued to return to the study every two years for a detailed medical history, physical examination, and laboratory tests, and in 1971, the Study enrolled a second generation - 5,124 of the original participants'''' adult children and their spouses - to participate in similar examinations. In 1994, the need to establish a new study reflecting a more diverse community of Framingham was recognized, and the first Omni cohort of the Framingham Heart Study was enrolled. In April 2002 the Study entered a new phase, the enrollment of a third generation of participants, the grandchildren of the Original Cohort. In 2003, a second group of Omni participants was enrolled. Over the years, careful monitoring of the Framingham Study population has led to the identification of major CVD risk factors, as well as valuable information on the effects of these factors such as blood pressure, blood triglyceride and cholesterol levels, age, gender, and psychosocial issues. Risk factors for other physiological conditions such as dementia have been and continue to be investigated. In addition, the relationships between physical traits and genetic patterns are being studied. FHS clinical and research data is stored in the dbGaP and NHLBI Repository repositories and may be accessed by application. Please check the following repositories before applying for data through FHS. Investigators seeking data that is not available through dbGaP or BioLINCC or seeking biological specimens may submit a proposal through the FHS web-based research application. The FHS data repository may be accessed through this FHS website, under the For Researchers link, then Description of Data, in order to determine if and how the desired data is stored. Proposals may involve the use of existing data, the collection of new data, either directly from participants or from previously collected samples, images, or other materials (e.g., medical records). The FHS Repository also has biological specimens available for genetic and non-genetic research proposals. Specimens include urine, blood and blood products, as well as DNA.

View all literature mentions

METAL (tool)

RRID:SCR_002013

Software application designed to facilitate meta-analysis of large datasets (such as several whole genome scans) in a convenient, rapid and memory efficient manner. (entry from Genetic Analysis Software)

View all literature mentions

STRUCTURE (tool)

RRID:SCR_002151

Software package for using multi locus genotype data to investigate population structure. Used for inferring presence of distinct populations, assigning individuals to populations, studying hybrid zones, identifying migrants and admixed individuals, and estimating population allele frequencies in situations where many individuals are migrants or admixed. Can be applied to most of commonly used genetic markers, including SNPS, microsatellites, RFLPs and Amplified Fragment Length Polymorphisms.

View all literature mentions

International HapMap Project (tool)

RRID:SCR_002846

THIS RESOURCE IS NO LONGER IN SERVICE, documented August 22, 2016. A multi-country collaboration among scientists and funding agencies to develop a public resource where genetic similarities and differences in human beings are identified and catalogued. Using this information, researchers will be able to find genes that affect health, disease, and individual responses to medications and environmental factors. All of the information generated by the Project will be released into the public domain. Their goal is to compare the genetic sequences of different individuals to identify chromosomal regions where genetic variants are shared. Public and private organizations in six countries are participating in the International HapMap Project. Data generated by the Project can be downloaded with minimal constraints. HapMap project related data, software, and documentation include: bulk data on genotypes, frequencies, LD data, phasing data, allocated SNPs, recombination rates and hotspots, SNP assays, Perlegen amplicons, raw data, inferred genotypes, and mitochondrial and chrY haplogroups; Generic Genome Browser software; protocols and information on assay design, genotyping and other protocols used in the project; and documentation of samples/individuals and the XML format used in the project.

View all literature mentions

1000 Genomes Project and AWS (tool)

RRID:SCR_008801

A dataset containing the full genomic sequence of 1,700 individuals, freely available for research use. The 1000 Genomes Project is an international research effort coordinated by a consortium of 75 companies and organizations to establish the most detailed catalogue of human genetic variation. The project has grown to 200 terabytes of genomic data including DNA sequenced from more than 1,700 individuals that researchers can now access on AWS for use in disease research free of charge. The dataset containing the full genomic sequence of 1,700 individuals is now available to all via Amazon S3. The data can be found at: http://s3.amazonaws.com/1000genomes The 1000 Genomes Project aims to include the genomes of more than 2,662 individuals from 26 populations around the world, and the NIH will continue to add the remaining genome samples to the data collection this year. Public Data Sets on AWS provide a centralized repository of public data hosted on Amazon Simple Storage Service (Amazon S3). The data can be seamlessly accessed from AWS services such Amazon Elastic Compute Cloud (Amazon EC2) and Amazon Elastic MapReduce (Amazon EMR), which provide organizations with the highly scalable compute resources needed to take advantage of these large data collections. AWS is storing the public data sets at no charge to the community. Researchers pay only for the additional AWS resources they need for further processing or analysis of the data. All 200 TB of the latest 1000 Genomes Project data is available in a publicly available Amazon S3 bucket. You can access the data via simple HTTP requests, or take advantage of the AWS SDKs in languages such as Ruby, Java, Python, .NET and PHP. Researchers can use the Amazon EC2 utility computing service to dive into this data without the usual capital investment required to work with data at this scale. AWS also provides a number of orchestration and automation services to help teams make their research available to others to remix and reuse. Making the data available via a bucket in Amazon S3 also means that customers can crunch the information using Hadoop via Amazon Elastic MapReduce, and take advantage of the growing collection of tools for running bioinformatics job flows, such as CloudBurst and Crossbow.

View all literature mentions

MCF-7 (tool)

RRID:CVCL_0031

Cell line MCF-7 is a Cancer cell line with a species of origin Homo sapiens (Human)

View all literature mentions