Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

A reference set of curated biomedical data and metadata from clinical case reports.

J Harry Caufield | Yijiang Zhou | Anders O Garlid | Shaun P Setty | David A Liem | Quan Cao | Jessica M Lee | Sanjana Murali | Sarah Spendlove | Wei Wang | Li Zhang | Yizhou Sun | Alex Bui | Henning Hermjakob | Karol E Watson | Peipei Ping
Scientific data | 2018

Clinical case reports (CCRs) provide an important means of sharing clinical experiences about atypical disease phenotypes and new therapies. However, published case reports contain largely unstructured and heterogeneous clinical data, posing a challenge to mining relevant information. Current indexing approaches generally concern document-level features and have not been specifically designed for CCRs. To address this disparity, we developed a standardized metadata template and identified text corresponding to medical concepts within 3,100 curated CCRs spanning 15 disease groups and more than 750 reports of rare diseases. We also prepared a subset of metadata on reports on selected mitochondrial diseases and assigned ICD-10 diagnostic codes to each. The resulting resource, Metadata Acquired from Clinical Case Reports (MACCRs), contains text associated with high-level clinical concepts, including demographics, disease presentation, treatments, and outcomes for each report. Our template and MACCR set render CCRs more findable, accessible, interoperable, and reusable (FAIR) while serving as valuable resources for key user groups, including researchers, physician investigators, clinicians, data scientists, and those shaping government policies for clinical trials.

Pubmed ID: 30457569

Research resources used in this publication

None found

Antibodies used in this publication

None found

Associated grants

  • Agency: NHLBI NIH HHS, United States
    Id: R35 HL135772
  • Agency: NIGMS NIH HHS, United States
    Id: U54 GM114833

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


MEDLINE (tool)

RRID:SCR_002185

A premier bibliographic database that contains over 18 million references to journal articles in life sciences with a concentration on biomedicine. A distinctive feature is that the records are indexed with NLM Medical Subject Headings (MeSH). PubMed provides free access to MEDLINE and links to full text articles when possible. The great majority of journals are selected for MEDLINE based on the recommendation of the Literature Selection Technical Review Committee (LSTRC), an NIH-chartered advisory committee of external experts analogous to the committees that review NIH grant applications. Some additional journals and newsletters are selected based on NLM-initiated reviews, e.g., history of medicine, health services research, AIDS, toxicology and environmental health, molecular biology, and complementary medicine, that are special priorities for NLM or other NIH components. These reviews generally also involve consultation with an array of NIH and outside experts or, in some cases, external organizations with which NLM has special collaborative arrangements. MEDLINE is the primary component of PubMed, part of the Entrez series of databases provided by the NLM National Center for Biotechnology Information (NCBI). MEDLINE may also be searched via the NLM Gateway. Time coverage: generally 1946 to the present, with some older material. Source: Currently, citations from approximately 5,516 worldwide journals in 39 languages; 60 languages for older journals. Citations for MEDLINE are created by the NLM, international partners, and collaborating organizations.

View all literature mentions

UniProt (tool)

RRID:SCR_002380

Collection of data of protein sequence and functional information. Resource for protein sequence and annotation data. Consortium for preservation of the UniProt databases: UniProt Knowledgebase (UniProtKB), UniProt Reference Clusters (UniRef), and UniProt Archive (UniParc), UniProt Proteomes. Collaboration between European Bioinformatics Institute (EMBL-EBI), SIB Swiss Institute of Bioinformatics and Protein Information Resource. Swiss-Prot is a curated subset of UniProtKB.

View all literature mentions

MeSH (tool)

RRID:SCR_004750

A controlled vocabulary thesaurus that consists of sets of terms naming descriptors in a hierarchical structure that permits searching at various levels of specificity. MeSH, in machine-readable form, is provided at no charge via electronic means. MeSH descriptors are arranged in both an alphabetic and a hierarchical structure. At the most general level of the hierarchical structure are very broad headings such as Anatomy or Mental Disorders. More specific headings are found at more narrow levels of the twelve-level hierarchy, such as Ankle and Conduct Disorder. There are 27,149 descriptors in 2014 MeSH. There are also over 218,000 entry terms that assist in finding the most appropriate MeSH Heading, for example, Vitamin C is an entry term to Ascorbic Acid. In addition to these headings, there are more than 219,000 headings called Supplementary Concept Records (formerly Supplementary Chemical Records) within a separate thesaurus. The MeSH thesaurus is used by NLM for indexing articles from 5,400 of the world''''s leading biomedical journals for the MEDLINE/PubMED database. It is also used for the NLM-produced database that includes cataloging of books, documents, and audiovisuals acquired by the Library. Each bibliographic reference is associated with a set of MeSH terms that describe the content of the item. Similarly, search queries use MeSH vocabulary to find items on a desired topic.

View all literature mentions

cTAKES (tool)

RRID:SCR_006379

An open-source natural language processing system for information extraction from electronic medical record clinical free-text. This is a system through which one creates one or more pipelines to process clinical notes and to identify clinical named entities. It processes clinical notes, identifying types of clinical named entities, drugs, diseases/disorders, signs/symptoms, anatomical sites and procedures. Each named entity that is found is given attributes for the text span, the ontology mapping code, the context (family history of, current, unrelated to patient), and negated/not negated. cTAKES is built on the UIMA framework. cTAKES 2.5 does not provide a GUI of its own for installation or processing. The cTAKES documentation shows how to use the GUIs provided by the UIMA framework, and how to run cTAKES from a command line. Before using cTAKES you need to know that cTAKES does not provide any mechanisms of its own to handle patient data securely. It is assumed that cTAKES is installed on a system that can process patient data, or that any data being processed by cTAKES has already been through a deidentification step in order to comply with any applicable laws. The tool has been developed and deployed at Mayo Clinic since early 2000.

View all literature mentions

HMDB (tool)

RRID:SCR_007712

Curated collection of human metabolite and human metabolism data which contains records for endogenous metabolites, with each metabolite entry containing detailed chemical, physical, biochemical, concentration, and disease information. This is further supplemented with thousands of NMR and MS spectra collected on purified reference metabolites.

View all literature mentions