The Experts below are selected from a list of 270 Experts worldwide ranked by ideXlab platform

Jijun Tang - One of the best experts on this subject based on the ideXlab platform.

  • ICIC (2) - Identification of DNA-binding Proteins via Fuzzy Multiple Kernel Model and Sequence Information
    Intelligent Computing Theories and Application, 2019
    Co-Authors: Yijie Ding, Jijun Tang
    Abstract:

    DNA-binding Proteins is the molecular basis for understanding the basic processes of life activities. Many diseases are associated with DNA binding Proteins. The methods of detecting DNA-binding Proteins are mainly realized by biochemical experiment, which is time consuming and extremely expensive. A lot of computational methods based on Machine Learning (ML) algorithm have been developed to detect DNA-binding Proteins. In this study, we propose a novel DNA-binding Proteins model via a Fuzzy Multiple Kernel Support Vector Machine. The multiple features of sequence and evolutionary are extracted and constructed as multiple kernels, respectively. Next, these corresponding kernels are integrated by Multiple Kernel Learning (MKL) algorithm. At last, Fuzzy Support Vector Machine (FSVM) is employed to build an effective DNA-binding protein predictor. Comparing with other outstanding methods, our proposed approach achieves good results. The accuracy of our model are 82.98% and 81.70% on PDB1075 (benchmark data set of DNA-binding Proteins) and PDB186 (independent test set), respectively. Our approach is comparable to previous methods.

  • Computational Methods for Predicting DNA Binding Proteins
    Current Proteomics, 2019
    Co-Authors: Jiandong Wang, Liang Zhao, William Hoskins, Jijun Tang
    Abstract:

    Background: DNA-binding Proteins are very important to many biomolecular functions. The traditional experimental methods are expensive and time consuming. So computational methods that can predict whether a protein is a DNA-binding protein or not are very helpful to researchers. Machine learning has been widely used in many research areas. Many researchers proposed machine learning methods to do DNA-binding protein prediction. To know their advantage and disadvantage is meaningful. Objective: There are many computational methods that can predict DNA-binding Proteins. Every method uses different features and different classifier algorithms. We want to take a review over those methods to find out some common procedures that can help researchers to develop more accurate methods. Method: Firstly, we talked about the information stored in the protein sequence and gene sequence. That information is the basement to find out the patterns leading to bind. Then, feature extraction methods and classifier algorithms are discussed. At last, give out some commonly used benchmark dataset and evaluate several methods. Conclusion: In this review, we analyze some popular computational methods which can do DNA-binding protein. From those methods, we find out many useful skills to build up an accurate DNA-binding protein classifier. Those can help researchers to build up more useful computational tools. Currently, there are some machine learning methods have good performance on predicting DNA-binding Proteins. The performance can be improved by using different kind of features and classifiers

  • Identification of DNA-binding Proteins by multiple kernel support vector machine and sequence information
    Current Proteomics, 2019
    Co-Authors: Yijie Ding, Jijun Tang, Feng Chen, Hongjie Wu
    Abstract:

    Background: The DNA-binding Proteins is an important process in multiple biomolecular functions. However, the tradition experimental methods for DNA-binding Proteins identification are still time consuming and extremely expensive. Objective: In past several years, various computational methods have been developed to detect DNA-binding Proteins. However, most of them do not integrate multiple information. Method: In this study, we propose a novel computational method to predict DNA-binding Proteins by two steps Multiple Kernel Support Vector Machine (MK-SVM) and sequence information. Firstly, we extract several feature and construct multiple kernels. Then, multiple kernels are linear combined by Multiple Kernel Learning (MKL). At last, a final SVM model, constructed by combined kernel, is built to predict DNA-binding Proteins. Results: The proposed method is tested on two benchmark data sets. Compared with other existing method, our approach is comparable, even better than other methods on some data sets. Conclusion: We can conclude that MK-SVM is more suitable than common SVM, as the classifier for DNA-binding Proteins identification.

  • improved detection of dna binding Proteins via compression technology on pssm information
    PLOS ONE, 2017
    Co-Authors: Yubo Wang, Jijun Tang, Yijie Ding
    Abstract:

    Since the importance of DNA-binding Proteins in multiple biomolecular functions has been recognized, an increasing number of researchers are attempting to identify DNA-binding Proteins. In recent years, the machine learning methods have become more and more compelling in the case of protein sequence data soaring, because of their favorable speed and accuracy. In this paper, we extract three features from the protein sequence, namely NMBAC (Normalized Moreau-Broto Autocorrelation), PSSM-DWT (Position-specific scoring matrix—Discrete Wavelet Transform), and PSSM-DCT (Position-specific scoring matrix—Discrete Cosine Transform). We also employ feature selection algorithm on these feature vectors. Then, these features are fed into the training SVM (support vector machine) model as classifier to predict DNA-binding Proteins. Our method applys three datasets, namely PDB1075, PDB594 and PDB186, to evaluate the performance of our approach. The PDB1075 and PDB594 datasets are employed for Jackknife test and the PDB186 dataset is used for the independent test. Our method achieves the best accuracy in the Jacknife test, from 79.20% to 86.23% and 80.5% to 86.20% on PDB1075 and PDB594 datasets, respectively. In the independent test, the accuracy of our method comes to 76.3%. The performance of independent test also shows that our method has a certain ability to be effectively used for DNA-binding protein prediction. The data and source code are at https://doi.org/10.6084/m9.figshare.5104084.

Fenfei Leng - One of the best experts on this subject based on the ideXlab platform.

  • Protein-induced DNA linking number change by sequence-specific DNA binding Proteins and its biological effects
    Biophysical Reviews, 2016
    Co-Authors: Fenfei Leng
    Abstract:

    Sequence-specific DNA-binding Proteins play essential roles in many fundamental biological events such as DNA replication, recombination, and transcription. One common feature of sequence-specific DNA-binding Proteins is to introduce structural changes to their DNA recognition sites including DNA-bending and DNA linking number change (ΔLk). In this article, I review recent progress in studying protein-induced ΔLk by several sequence-specific DNA-binding Proteins, such as E. coli cAMP receptor protein (CRP) and lactose repressor (LacI). It was demonstrated recently that protein-induced ΔLk is an intrinsic property for sequence-specific DNA-binding Proteins and does not correlate to protein-induced other structural changes, such as DNA bending. For instance, although CRP bends its DNA recognition site by 90°, it was not able to introduce a ΔLk to it. However, LacI was able to simultaneously bend and introduce a ΔLk to its DNA binding sites. Intriguingly, LacI also constrained superhelicity within LacI– lac O1 complexes if (−) supercoiled DNA templates were provided. I also discuss how protein-induced ΔLk help sequence-specific DNA-binding Proteins regulate their biological functions. For example, it was shown recently that LacI utilizes the constrained superhelicity (ΔLk) in LacI- lac O1 complexes and serves as a topological barrier to constrain free, unconstrained (−) supercoils within the 401-bp DNA loop. These constrained (−) supercoils enhance LacI’s binding affinity and therefore the repression of the lac promoter. Other biological functions include how DNA replication initiators λ O and DnaA use the induced ΔLk to open/melt bacterial DNA replication origins.

  • DNA linking number change induced by sequence-specific DNA-binding Proteins
    Nucleic Acids Research, 2010
    Co-Authors: Bo Chen, Yazhong Xiao, Chen-zhong Li, Fenfei Leng
    Abstract:

    Sequence-specific DNA-binding Proteins play a key role in many fundamental biological processes, such as transcription, DNA replication and recombination. Very often, these DNA-binding Proteins introduce structural changes to the target DNA-binding sites including DNA bending, twisting or untwisting and wrapping, which in many cases induce a linking number change (ΔLk) to the DNA-binding site. Due to the lack of a feasible approach, ΔLk induced by sequence-specific DNA-binding Proteins has not been fully explored. In this paper we successfully constructed a series of DNA plasmids that carry many tandem copies of a DNA-binding site for one sequence-specific DNA-binding protein, such as λ O, LacI, GalR, CRP and AraC. In this case, the protein-induced ΔLk was greatly amplified and can be measured experimentally. Indeed, not only were we able to simultaneously determine the protein-induced ΔLk and the DNA-binding constant for λ O and GalR, but also we demonstrated that the protein-induced ΔLk is an intrinsic property for these sequence-specific DNA-binding Proteins. Our results also showed that protein-mediated DNA looping by AraC and LacI can induce a ΔLk to the plasmid DNA templates. Furthermore, we demonstrated that the protein-induced ΔLk does not correlate with the protein-induced DNA bending by the DNA-binding Proteins.

Wei Wang - One of the best experts on this subject based on the ideXlab platform.

  • Analysis and prediction of single-stranded and double-stranded DNA binding Proteins based on protein sequences
    BMC Bioinformatics, 2017
    Co-Authors: Wei Wang, Shiguang Zhang, Hongjun Zhang, Tianhe Xu, Keliang Li
    Abstract:

    DNA-binding Proteins perform important functions in a great number of biological activities. DNA-binding Proteins can interact with ssDNA (single-stranded DNA) or dsDNA (double-stranded DNA), and DNA-binding Proteins can be categorized as single-stranded DNA-binding Proteins (SSBs) and double-stranded DNA-binding Proteins (DSBs). The identification of DNA-binding Proteins from amino acid sequences can help to annotate protein functions and understand the binding specificity. In this study, we systematically consider a variety of schemes to represent protein sequences: OAAC (overall amino acid composition) features, dipeptide compositions, PSSM (position-specific scoring matrix profiles) and split amino acid composition (SAA), and then we adopt SVM (support vector machine) and RF (random forest) classification model to distinguish SSBs from DSBs. Our results suggest that some sequence features can significantly differentiate DSBs and SSBs. Evaluated by 10 fold cross-validation on the benchmark datasets, our prediction method can achieve the accuracy of 88.7% and AUC (area under the curve) of 0.919. Moreover, our method has good performance in independent testing. Using various sequence-derived features, a novel method is proposed to distinguish DSBs and SSBs accurately. The method also explores novel features, which could be helpful to discover the binding specificity of DNA-binding Proteins.

  • Identification of single-stranded and double-stranded dna binding Proteins based on protein structure
    BMC Bioinformatics, 2014
    Co-Authors: Wei Wang, Xionghui Zhou
    Abstract:

    Background Protein-DNA interactions are essential for many biological processes. However, the structural mechanisms underlying these interactions are not fully understood. DNA binding Proteins can be classified into double-stranded DNA binding Proteins (DSBs) and single-stranded DNA binding Proteins (SSBs), and they take part in different biological functions. DSBs usually act as transcriptional factors to regulate the genes' expressions, while SSBs usually play roles in DNA replication, recombination, and repair, etc. Understanding the binding specificity of a DNA binding protein is helpful for the research of protein functions.

  • Analysis and classification of DNA-binding sites in single-stranded and double-stranded DNA-binding Proteins using protein information
    IET Systems Biology, 2014
    Co-Authors: Wei Wang, Yi Xiong, Xionghui Zhou
    Abstract:

    Single-stranded DNA-binding Proteins (SSBs) and double-stranded DNA-binding Proteins (DSBs) play different roles in biological processes when they bind to single-stranded DNA (ssDNA) or double-stranded DNA (dsDNA). However, the underlying binding mechanisms of SSBs and DSBs have not yet been fully understood. Here, the authors firstly constructed two groups of ssDNA and dsDNA specific binding sites from two non-redundant sets of SSBs and DSBs. They further analysed the relationship between the two classes of binding sites and a newly proposed set of features (residue charge distribution, secondary structure and spatial shape). To assess and utilise the predictive power of these features, they trained a classification model using support vector machine to make predictions about the ssDNA and the dsDNA binding sites. The author's analysis and prediction results indicated that the two classes of binding sites can be distinguishable by the three types of features, and the final classifier using all the features achieved satisfactory performance. In conclusion, the proposed features will deepen their understanding of the specificity of Proteins which bind to ssDNA or dsDNA.

  • distinguishing single stranded and double stranded dna binding Proteins based on structural information
    Bioinformatics and Biomedicine, 2013
    Co-Authors: Wei Wang
    Abstract:

    Protein-DNA interactions play a critical role in many biological processes. However, the structural mechanisms underlying these interactions are not fully understood. DNA binding Proteins can be classified into double-stranded DNA binding Proteins (DSBs) and single-stranded DNA binding Proteins (SSBs). Understanding the binding specificity of a DNA binding protein is helpful for the research of protein functions. Though there are some researches [1] on the SSB and DSB respectively, few attentions have been paid on investigating what makes SSB and DSB have such different ability of the specific binding. With the development of biotechnology, a large amount of Proteins has been sequenced. However, SSBs have shown to have little sequence conservation [2]. Even DSBs involved in similar functions may have conserved subsequences, different kinds of DSBs with different functions seems to show few common subsequences. Therefore, it is hard to recognize SSB sequences from DSB sequences, or vice versa. In fact, up to Jan. 25, 2013, the Protein Data Bank (PDB) [3] contains 3391 structures for DNA binding Proteins, among them only about 30% and 5% are annotated as DSBs and SSBs, respectively, and whether the remainders belong to DSBs or SSBs are still not very clear. Therefore, a computational method is required to annotate the DNA binding protein as DSB or SSB automatically. The surface of a protein is generally irregular, containing many clefts and grooves of varying shapes and sizes. Previous researches have shown that a large cleft can provide an increased opportunity for the protein to form interactions with other molecules, particularly small ligands [4]. Therefore, some researches used a particularly large and deep cleft to characterize the binding active sites of the Proteins [5]. We guess that for DNA binding Proteins, the cleft properties on the surface may also play important roles on the dsDNA/ssDNA binding specificity. In this work, we applied CAVER 3.0 package [6] to detect the clefts and the corresponding indexes of the largest clefts on the protein surfaces, to investigate whether they are possible to be used for distinguishing the potential interfaces between SSBs and DSBs. Concretely, we mainly got three indexes of the detected tunnels: length, curvature and bottleneck radius. Research results have shown that although the sequences of different SSBs are very different, there are well-conserved elements in the structures. That is, most SSBs contain one or more OB (oligonucleotide/oligosaccharide binding) - fold domains [2]. A typical OB-fold has a five-stranded beta-sheet coiled to form a closed beta-barrel. This barrel is capped by an alpha-helix located between the third and fourth strands. The OB-fold plays critical role in binding with ssDNA. Although it is hard to say that the OB-fold is unique for SSBs, we think that it should also be used as an important descriptor to distinguish SSBs from DSBs. Therefore, we use the protein structure alignment package TM-align [7] to compare its structure with each of the six OB-fold protein templates and use the maximal alignment score TM-score as the OB-fold feature of the protein. We aim to investigate the structural differences between collected SSBs and DSBs, and extract the structure-based features related to surface clefts and OB-folds, Based on which, we construct a computational model that can automatically classify the DNA protein as a DSB or SSB by using the widely used support vector machine (SVM), with prediction accuracy of HOLO-set 0.87, APO-set 0.83, and mixed-set 0.83, respectively. The promising performance suggests that our method will be useful in the protein function annotation and refinement. This work is supported by grant from the National Science Foundation of China (61272274); Program for New Century Excellent Talents in Universities (NCET-10-0644), the Open Research Fund of State Key Laboratory of Hybrid Ri ce (Wuhan University) (KF201301) and the Fundamental Research Funds for the Central Universities (No. 2012211020204).

Yijie Ding - One of the best experts on this subject based on the ideXlab platform.

  • ICIC (2) - A Prediction Method of DNA-binding Proteins Based on Evolutionary Information
    Intelligent Computing Theories and Application, 2019
    Co-Authors: Weizhong Lu, Yijie Ding, Hongjie Wu, Zhengwei Song, Hongmei Huang
    Abstract:

    DNA is the carrier of genetic information in organisms, and DNA-binding protein is one type of unwinding enzymes, which plays a key role in various biological molecular functions. That has greatly promoted the research of various methods for identifying DNA-binding Proteins. In recent years, researchers have developed a Machine Learning-based method to predict DNA-binding Proteins quickly and accurately. Although the prediction accuracy of current methods is considerable, the performance of their prediction can be further improved. In this paper, a DNA-binding Proteins prediction model based on PSSM (Position Specific Scoring Matrix) features and Random Forest classifier is proposed. The results of experiments show that the proposed method can achieve great prediction performance on PDB1075 and PDB186 datasets, whose accuracy is 82.14% and 79.0%, respectively. Experiments show that the method can be compared with other methods, and even surpass the previous methods on some datasets.

  • ICIC (2) - Identification of DNA-binding Proteins via Fuzzy Multiple Kernel Model and Sequence Information
    Intelligent Computing Theories and Application, 2019
    Co-Authors: Yijie Ding, Jijun Tang
    Abstract:

    DNA-binding Proteins is the molecular basis for understanding the basic processes of life activities. Many diseases are associated with DNA binding Proteins. The methods of detecting DNA-binding Proteins are mainly realized by biochemical experiment, which is time consuming and extremely expensive. A lot of computational methods based on Machine Learning (ML) algorithm have been developed to detect DNA-binding Proteins. In this study, we propose a novel DNA-binding Proteins model via a Fuzzy Multiple Kernel Support Vector Machine. The multiple features of sequence and evolutionary are extracted and constructed as multiple kernels, respectively. Next, these corresponding kernels are integrated by Multiple Kernel Learning (MKL) algorithm. At last, Fuzzy Support Vector Machine (FSVM) is employed to build an effective DNA-binding protein predictor. Comparing with other outstanding methods, our proposed approach achieves good results. The accuracy of our model are 82.98% and 81.70% on PDB1075 (benchmark data set of DNA-binding Proteins) and PDB186 (independent test set), respectively. Our approach is comparable to previous methods.

  • Identification of DNA-binding Proteins by multiple kernel support vector machine and sequence information
    Current Proteomics, 2019
    Co-Authors: Yijie Ding, Jijun Tang, Feng Chen, Hongjie Wu
    Abstract:

    Background: The DNA-binding Proteins is an important process in multiple biomolecular functions. However, the tradition experimental methods for DNA-binding Proteins identification are still time consuming and extremely expensive. Objective: In past several years, various computational methods have been developed to detect DNA-binding Proteins. However, most of them do not integrate multiple information. Method: In this study, we propose a novel computational method to predict DNA-binding Proteins by two steps Multiple Kernel Support Vector Machine (MK-SVM) and sequence information. Firstly, we extract several feature and construct multiple kernels. Then, multiple kernels are linear combined by Multiple Kernel Learning (MKL). At last, a final SVM model, constructed by combined kernel, is built to predict DNA-binding Proteins. Results: The proposed method is tested on two benchmark data sets. Compared with other existing method, our approach is comparable, even better than other methods on some data sets. Conclusion: We can conclude that MK-SVM is more suitable than common SVM, as the classifier for DNA-binding Proteins identification.

  • improved detection of dna binding Proteins via compression technology on pssm information
    PLOS ONE, 2017
    Co-Authors: Yubo Wang, Jijun Tang, Yijie Ding
    Abstract:

    Since the importance of DNA-binding Proteins in multiple biomolecular functions has been recognized, an increasing number of researchers are attempting to identify DNA-binding Proteins. In recent years, the machine learning methods have become more and more compelling in the case of protein sequence data soaring, because of their favorable speed and accuracy. In this paper, we extract three features from the protein sequence, namely NMBAC (Normalized Moreau-Broto Autocorrelation), PSSM-DWT (Position-specific scoring matrix—Discrete Wavelet Transform), and PSSM-DCT (Position-specific scoring matrix—Discrete Cosine Transform). We also employ feature selection algorithm on these feature vectors. Then, these features are fed into the training SVM (support vector machine) model as classifier to predict DNA-binding Proteins. Our method applys three datasets, namely PDB1075, PDB594 and PDB186, to evaluate the performance of our approach. The PDB1075 and PDB594 datasets are employed for Jackknife test and the PDB186 dataset is used for the independent test. Our method achieves the best accuracy in the Jacknife test, from 79.20% to 86.23% and 80.5% to 86.20% on PDB1075 and PDB594 datasets, respectively. In the independent test, the accuracy of our method comes to 76.3%. The performance of independent test also shows that our method has a certain ability to be effectively used for DNA-binding protein prediction. The data and source code are at https://doi.org/10.6084/m9.figshare.5104084.

Xionghui Zhou - One of the best experts on this subject based on the ideXlab platform.

  • Identification of single-stranded and double-stranded dna binding Proteins based on protein structure
    BMC Bioinformatics, 2014
    Co-Authors: Wei Wang, Xionghui Zhou
    Abstract:

    Background Protein-DNA interactions are essential for many biological processes. However, the structural mechanisms underlying these interactions are not fully understood. DNA binding Proteins can be classified into double-stranded DNA binding Proteins (DSBs) and single-stranded DNA binding Proteins (SSBs), and they take part in different biological functions. DSBs usually act as transcriptional factors to regulate the genes' expressions, while SSBs usually play roles in DNA replication, recombination, and repair, etc. Understanding the binding specificity of a DNA binding protein is helpful for the research of protein functions.

  • Analysis and classification of DNA-binding sites in single-stranded and double-stranded DNA-binding Proteins using protein information
    IET Systems Biology, 2014
    Co-Authors: Wei Wang, Yi Xiong, Xionghui Zhou
    Abstract:

    Single-stranded DNA-binding Proteins (SSBs) and double-stranded DNA-binding Proteins (DSBs) play different roles in biological processes when they bind to single-stranded DNA (ssDNA) or double-stranded DNA (dsDNA). However, the underlying binding mechanisms of SSBs and DSBs have not yet been fully understood. Here, the authors firstly constructed two groups of ssDNA and dsDNA specific binding sites from two non-redundant sets of SSBs and DSBs. They further analysed the relationship between the two classes of binding sites and a newly proposed set of features (residue charge distribution, secondary structure and spatial shape). To assess and utilise the predictive power of these features, they trained a classification model using support vector machine to make predictions about the ssDNA and the dsDNA binding sites. The author's analysis and prediction results indicated that the two classes of binding sites can be distinguishable by the three types of features, and the final classifier using all the features achieved satisfactory performance. In conclusion, the proposed features will deepen their understanding of the specificity of Proteins which bind to ssDNA or dsDNA.