采用高斯过程模拟预测域/肽识别和相互作用
摘要
Many protein-protein interactions involved in cell signaling networks conduct with the manner so-called folding-on-binding, which are mediated by the binding of a globular domain in one protein to a short peptide stretch in another. Thus, systematic analysis and reliable prediction of domain-peptide recognition and interaction are fundamentally important for our understanding of the molecular mechanism and biological implications underlying cell signaling. Herein, we report the use of a new and powerful machine learning technique called Gaussian process (GP) to carry out statistical modeling and structural analysis for four categories of domain-peptide systems, including SH3, PDZ, 14-3-3, and GYF domains. The results of the modeling are compared systematically to those deriving from classical partial least square (PLS) regression and sophisticated supporting vector machine (SVM). We demonstrate that GP is comparable with or even better than nonlinear SVM, and is much well to linear PLS. In addition, GP possesses some more merits as it is capable of handling linearity and nonlinearity-hybrid problems effectively, determining algorithmic parameters automatically, interpreting obtained models straightforwardly, and providing additional evaluation for predictions quantitatively. All of these come together to suggest that GP would be a promising tool not only for exploring the domain-peptide interaction behavior, but also for solving other chemistry and biology-related problems.