Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
Lysine formylation is a newly discovered post-translational modification in histones, which plays a crucial role in epigenetics of chromatin function and DNA binding. In this study, a novel bioinformatics tool named CKSAAP_FormSite is proposed to predict lysine formylation sites. An effective feature extraction method, the composition of k-spaced amino acid pairs, is employed to encode formylation sites. Moreover, a biased support vector machine algorithm is proposed to solve the class imbalance problem in the prediction of formylation sites. As illustrated by 10-fold cross-validation, CKSAAP_FormSite achieves an satisfactory performance with an AUC of 0.8234. Therefore, CKSAAP_FormSite can be a useful bioinformatics tool for the prediction of formylation sites. Feature analysis shows that some amino acid pairs, such as 'KA', 'SxxxxK' and 'SxxxA' around formylation sites may play an important role in the prediction. The results of analysis and prediction could offer useful information for elucidating the molecular mechanisms of formylation.
Pubmed ID: 31175975
Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on February 28,2023. Software program for clustering biological sequences with many applications in various fields such as making non-redundant databases, finding duplicates, identifying protein families, filtering sequence errors and improving sequence assembly etc. It is very fast and can handle extremely large databases. CD-HIT helps to significantly reduce the computational and manual efforts in many sequence analysis tasks and aids in understanding the data structure and correct the bias within a dataset. The CD-HIT package has CD-HIT, CD-HIT-2D, CD-HIT-EST, CD-HIT-EST-2D, CD-HIT-454, CD-HIT-PARA, PSI-CD-HIT, CD-HIT-OTU and over a dozen scripts. * CD-HIT (CD-HIT-EST) clusters similar proteins (DNAs) into clusters that meet a user-defined similarity threshold. * CD-HIT-2D (CD-HIT-EST-2D) compares 2 datasets and identifies the sequences in db2 that are similar to db1 above a threshold. * CD-HIT-454 identifies natural and artificial duplicates from pyrosequencing reads. * CD-HIT-OTU cluster rRNA tags into OTUs The usage of other programs and scripts can be found in CD-HIT user''s guide. CD-HIT was originally developed by Dr. Weizhong Li at Dr. Adam Godzik''s Lab at the Burnham Institute (now Sanford-Burnham Medical Research Institute).
View all literature mentionsAn integrated software for support vector classification, (C-SVC, nu-SVC), regression (epsilon-SVR, nu-SVR) and distribution estimation (one-class SVM) from the laboratory of Chih-Chung Chang and Chih-Jen Lin. It supports multi-class classification.
View all literature mentions