A sequence-based approach for identifying recombination spots in Saccharomyces cerevisiae by using hyper-parameter optimization in FastText and support vector machine

Duyen Thi Do, Nguyen Quoc Khanh Le

研究成果: 雜誌貢獻文章同行評審

4 引文 斯高帕斯(Scopus)

摘要

Meiotic recombination is a biological process which plays a crucial role in genetic evolution. Therefore, the ability of machine learning models in extracting desire information embedded in DNA sequences has drawn a great deal of attention among biologists. Recently, several attempts have been made to address this problem, however, the performance results still need to be improved. The current study aims to investigate the relationship between natural language processing model and supervised learning in classifying DNA sequences. The idea is to treat DNA sequences by FastText model, including sub-word information and then use them as features in a suitable supervised learning algorithm. To the end, this hybrid approach helps us classify DNA recombination spots with achieved sensitivity of 90%, specificity of 94.76%, accuracy of 92.6%, and MCC of 0.851. These results have suggested that our newly proposed method is superior to other methods on the same benchmark dataset. This study, therefore, could shed the light on developing the prediction models for recombination spots in particular, and DNA sequences in general.
原文英語
文章編號103855
期刊Chemometrics and Intelligent Laboratory Systems
194
DOIs
出版狀態已發佈 - 十一月 15 2019

ASJC Scopus subject areas

  • 分析化學
  • 軟體
  • 製程化學與技術
  • 光譜
  • 電腦科學應用

指紋

深入研究「A sequence-based approach for identifying recombination spots in Saccharomyces cerevisiae by using hyper-parameter optimization in FastText and support vector machine」主題。共同形成了獨特的指紋。

引用此