Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
The carcinogenicity of drugs can have a serious impact on human health, so carcinogenicity testing of new compounds is very necessary before being put on the market. Currently, many methods have been used to predict the carcinogenicity of compounds. However, most methods have limited predictive power and there is still much room for improvement. In this study, we construct a deep learning model based on capsule network and attention mechanism named DCAMCP to discriminate between carcinogenic and non-carcinogenic compounds. We train the DCAMCP on a dataset containing 1564 different compounds through their molecular fingerprints and molecular graph features. The trained model is validated by fivefold cross-validation and external validation. DCAMCP achieves an average accuracy (ACC) of 0.718 ± 0.009, sensitivity (SE) of 0.721 ± 0.006, specificity (SP) of 0.715 ± 0.014 and area under the receiver-operating characteristic curve (AUC) of 0.793 ± 0.012. Meanwhile, comparable results can be achieved on an external validation dataset containing 100 compounds, with an ACC of 0.750, SE of 0.778, SP of 0.727 and AUC of 0.811, which demonstrate the reliability of DCAMCP. The results indicate that our model has made progress in cancer risk assessment and could be used as an efficient tool in drug design.
Pubmed ID: 37525507
Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.
Collection of information about chemical structures and biological properties of small molecules and siRNA reagents hosted by the National Center for Biotechnology Information (NCBI).
View all literature mentionsToxicology data file of the National Library of Medicine''s (NLM) Toxicology Data Network (TOXNET). It is a scientifically evaluated and fully referenced data bank, developed and maintained by the National Cancer Institute (NCI). It contains over 9,000 chemical records with carcinogenicity, mutagenicity, tumor promotion, and tumor inhibition test results. Data are derived from studies cited in primary journals, current awareness tools, NCI reports, and other special sources. Test results have been reviewed by experts in carcinogenesis and mutagenesis. CCRIS is easily accessible and free of charge. Users can search by chemical or other name, chemical name fragment, Chemical Abstracts Service Registry Number (RN), and/or subject terms. Search results can easily be viewed, printed or downloaded. Search results are displayed in relevancy ranked order. Users may select to display any combination of data from the following broad groupings: (a) Carcinogenicity Studies (b) Tumor Promotion Studies (c) Mutagenicity Studies (d) Tumor Inhibition Studies Users can easily conduct their CCRIS search strategy against other databases: Hazardous Substances Data Bank, Integrated Risk Information System, GENE-TOX, TOXLINE, and ChemIDplus.
View all literature mentions