<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.3 20210610//EN" "JATS-journalpublishing1-3.dtd">
<article article-type="research-article" dtd-version="1.3" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xml:lang="en"><front><journal-meta><journal-id journal-id-type="publisher-id">donstu</journal-id><journal-title-group><journal-title xml:lang="en">Advanced Engineering Research (Rostov-on-Don)</journal-title><trans-title-group xml:lang="ru"><trans-title>Advanced Engineering Research (Rostov-on-Don)</trans-title></trans-title-group></journal-title-group><issn pub-type="epub">2687-1653</issn><publisher><publisher-name>Don State Technical University</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.23947/2687-1653-2022-22-2-169-176</article-id><article-id custom-type="elpub" pub-id-type="custom">donstu-1888</article-id><article-categories><subj-group subj-group-type="heading"><subject>Research Article</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="en"><subject>INFORMATION TECHNOLOGY, COMPUTER SCIENCE AND MANAGEMENT</subject></subj-group><subj-group subj-group-type="section-heading" xml:lang="ru"><subject>ИНФОРМАТИКА, ВЫЧИСЛИТЕЛЬНАЯ ТЕХНИКА И УПРАВЛЕНИЕ</subject></subj-group></article-categories><title-group><article-title>Analysis of natural language processing technology: modern problems and approaches</article-title><trans-title-group xml:lang="ru"><trans-title>Анализ технологии обработки естественного языка: современные проблемы и подходы</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-8669-3383</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Казакова</surname><given-names>М. А.</given-names></name><name name-style="western" xml:lang="en"><surname>Kazakova</surname><given-names>M. A.</given-names></name></name-alternatives><bio xml:lang="ru"><p>г. Казань, ул. К. Маркса, 10</p></bio><bio xml:lang="en"><p>Kazakova, Maria A., graduate student of the Applied Mathematics and Computer Science Department</p><p>10, K. Marx St., Kazan, 420111  </p></bio><email xlink:type="simple">kazakovamaria2609@gmail.com</email><xref ref-type="aff" rid="aff-1"/></contrib><contrib contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-2723-4715</contrib-id><name-alternatives><name name-style="eastern" xml:lang="ru"><surname>Султанова</surname><given-names>А. П.</given-names></name><name name-style="western" xml:lang="en"><surname>Sultanova</surname><given-names>A. P.</given-names></name></name-alternatives><bio xml:lang="ru"><p>г. Казань, ул. К. Маркса, 10</p></bio><bio xml:lang="en"><p>Sultanova, Alina P., associate professor of the Department of Foreign Languages, Russian as a Foreign Language</p><p>10, K. Marx St., Kazan, 420111 </p></bio><email xlink:type="simple">alinasultanova@mail.ru</email><xref ref-type="aff" rid="aff-1"/></contrib></contrib-group><aff-alternatives id="aff-1"><aff xml:lang="ru"><institution>Казанский национальный исследовательский технический университет им. А. Н. Туполева-КАИ</institution><country>Россия</country></aff><aff xml:lang="en"><institution>Kazan National Research Technical University named after A.N. Tupolev–KAI</institution><country>Russian Federation</country></aff></aff-alternatives><pub-date pub-type="collection"><year>2022</year></pub-date><pub-date pub-type="epub"><day>11</day><month>07</month><year>2022</year></pub-date><volume>22</volume><issue>2</issue><fpage>169</fpage><lpage>176</lpage><permissions><copyright-statement>Copyright &amp;#x00A9; Kazakova M.A., Sultanova A.P., 2022</copyright-statement><copyright-year>2022</copyright-year><copyright-holder xml:lang="ru">Казакова М.А., Султанова А.П.</copyright-holder><copyright-holder xml:lang="en">Kazakova M.A., Sultanova A.P.</copyright-holder><license xml:lang="ru" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>Данная работа распространяется под лицензией Creative Commons Attribution 4.0.</license-p></license><license xml:lang="en" license-type="creative-commons-attribution" xlink:href="https://creativecommons.org/licenses/by/4.0/" xlink:type="simple"><license-p>This work is licensed under a Creative Commons Attribution 4.0 License.</license-p></license></permissions><self-uri xlink:href="https://www.vestnik-donstu.ru/jour/article/view/1888">https://www.vestnik-donstu.ru/jour/article/view/1888</self-uri><abstract><sec><title>Introduction</title><p>Introduction. The article presents an overview of modern neural network models for natural language processing. Research into natural language processing is of interest as the need to process large amounts of audio and text information accumulated in recent decades has increased. The most discussed in foreign literature are the features of the processing of spoken language. The aim of the work is to present modern models of neural networks in the field of oral speech processing.</p></sec><sec><title>Materials and Methods</title><p>Materials and Methods. Applied research on understanding spoken language is an important and far-reaching topic in the natural language processing. Listening comprehension is central to practice and presents a challenge. This study meets a method of hearing detection based on deep learning. The article briefly outlines the substantive aspects of various neural networks for speech recognition, using the main terms associated with this theory. A brief description of the main points of the transformation of neural networks into a natural language is given.</p></sec><sec><title>Results</title><p>Results. A retrospective analysis of foreign and domestic literary sources was carried out alongside with a description of new methods for oral speech processing, in which neural networks were used. Information about neural networks, methods of speech recognition and synthesis is provided. The work includes the results of diverse experimental works of recent years. The article elucidates the main approaches to natural language processing and their changes over time, as well as the emergence of new technologies. The major problems currently existing in this area are considered.</p><p>Discussion and Conclusions. The analysis of the main aspects of speech recognition systems has shown that there is currently no universal system that would be self-learning, noise-resistant, recognizing continuous speech, capable of working with large dictionaries and at the same time having a low error rate.</p></sec></abstract><trans-abstract xml:lang="ru"><sec><title>Введение</title><p>Введение. В статье представлен обзор современных моделей нейронных сетей для обработки естественного языка. Исследования обработки естественного языка представляют интерес в связи с тем, что в последнее время возросла потребность в обработке больших объемов аудио- и текстовой информации, накопленной за последние десятилетия. Наиболее обсуждаемой в зарубежной литературе является особенности обработки разговорной речи. Цель работы представить современные модели нейронных сетей в области обработки устной речи.</p></sec><sec><title>Материалы и методы</title><p>Материалы и методы. Прикладное исследование понимания устной речи является сложной и далеко идущей темой обработки естественного языка. Понимание на слух занимает центральное место в исследовании и представляет собой проблему. В этой статье предлагается метод понимания на слух, основанный на глубоком обучении. В статье кратко излагаются содержательные аспекты различных этапов создания нейронной сети по распознаванию речи, приводятся основные термины, связанные с этой теорией. Приводится краткая характеристика основных существующих на текущий момент нейронных сетей по обработке естественного языка.</p></sec><sec><title>Результаты исследования</title><p>Результаты исследования. Проведен ретроспективный анализ зарубежных и отечественных литературных источников с описанием новых методов обработки устной речи, в которых использовались нейронные сети. Предоставлена информация о нейронных сетях, методах распознавания и синтеза речи. В работу включены результаты разноплановых экспериментальных работ последних лет. В статье подробно описаны основные подходы к обработке естественного языка и их изменения с течением времени, а также появление новых технологий. Рассмотрены основные проблемы, существующие в настоящее время в этой области.</p></sec><sec><title>Обсуждение и заключения</title><p>Обсуждение и заключения. Анализ основных аспектов систем распознавания речи показал, что в настоящее время не существует универсальной системы, которая была бы самообучающейся, шумоизоляционной, распознающей непрерывную речь, способной работать с большими словарями и в то же время имеющей низкий уровень ошибок.</p></sec></trans-abstract><kwd-group xml:lang="ru"><kwd>обработка естественного языка</kwd><kwd>устная речь</kwd><kwd>нейронные сети</kwd><kwd>автоматическая обработка естественного языка</kwd><kwd>семантическая согласованность</kwd></kwd-group><kwd-group xml:lang="en"><kwd>Natural Language Processing</kwd><kwd>oral speech</kwd><kwd>neural networks</kwd><kwd>automated natural language processing</kwd><kwd>semantic consistency</kwd></kwd-group></article-meta></front><back><ref-list><title>References</title><ref id="cit1"><label>1</label><citation-alternatives><mixed-citation xml:lang="ru">Lee A, Auli M, Ranzato MA. Discriminative reranking for neural machine translation. In: ACL-IJCNLP 2021 – 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Proceedings of the Conference, 2021. P. 7250–7264. http://dx.doi.org/10.18653/v1/2021.acl-long.563</mixed-citation><mixed-citation xml:lang="en">Lee A, Auli M, Ranzato MA. Discriminative reranking for neural machine translation. In: ACL-IJCNLP 2021 – 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Proceedings of the Conference, 2021. P. 7250–7264. http://dx.doi.org/10.18653/v1/2021.acl-long.563</mixed-citation></citation-alternatives></ref><ref id="cit2"><label>2</label><citation-alternatives><mixed-citation xml:lang="ru">Prashant Johri, Sunil Kumar Khatri, Ahmad T Al-Taani, et al. Natural language processing: History, evolution, application, and future work. In: Proc. 3rd International Conference on Computing Informatics and Networks. 2021;167:365-375. http://dx.doi.org/10.1007/978-981-15-9712-1_31</mixed-citation><mixed-citation xml:lang="en">Prashant Johri, Sunil Kumar Khatri, Ahmad T Al-Taani, et al. Natural language processing: History, evolution, application, and future work. In: Proc. 3rd International Conference on Computing Informatics and Networks. 2021;167:365-375. http://dx.doi.org/10.1007/978-981-15-9712-1_31</mixed-citation></citation-alternatives></ref><ref id="cit3"><label>3</label><citation-alternatives><mixed-citation xml:lang="ru">Nitschke R. Restoring the Sister: Reconstructing a Lexicon from Sister Languages using Neural Machine Translation. In: Proc. 1st Workshop on Natural Language Processing for Indigenous Languages of the Americas, AmericasNLP 2021. 2021. P. 122 – 130. http://dx.doi.org/10.18653/v1/2021.americasnlp-1.13</mixed-citation><mixed-citation xml:lang="en">Nitschke R. Restoring the Sister: Reconstructing a Lexicon from Sister Languages using Neural Machine Translation. In: Proc. 1st Workshop on Natural Language Processing for Indigenous Languages of the Americas, AmericasNLP 2021. 2021. P. 122 – 130. http://dx.doi.org/10.18653/v1/2021.americasnlp-1.13</mixed-citation></citation-alternatives></ref><ref id="cit4"><label>4</label><citation-alternatives><mixed-citation xml:lang="ru">Pokrovskii MM. Izbrannye raboty po yazykoznaniyu. Moscow : Izd-vo Akad. nauk SSSR; 1959. 382 p. (In Russ.)</mixed-citation><mixed-citation xml:lang="en">Pokrovskii MM. Izbrannye raboty po yazykoznaniyu. Moscow : Izd-vo Akad. nauk SSSR; 1959. 382 p. (In Russ.)</mixed-citation></citation-alternatives></ref><ref id="cit5"><label>5</label><citation-alternatives><mixed-citation xml:lang="ru">Ryazanov VV. Modeli, metody, algoritmy i arkhitektury sistem raspoznavaniya rechi. Moscow: Vychislitel'nyi tsentr im. A.A. Dorodnitsyna; 2006. 138 p. (In Russ.)</mixed-citation><mixed-citation xml:lang="en">Ryazanov VV. Modeli, metody, algoritmy i arkhitektury sistem raspoznavaniya rechi. Moscow: Vychislitel'nyi tsentr im. A.A. Dorodnitsyna; 2006. 138 p. (In Russ.)</mixed-citation></citation-alternatives></ref><ref id="cit6"><label>6</label><citation-alternatives><mixed-citation xml:lang="ru">Lixian Hou, Yanling Li, Chengcheng Li, et al. Review of research on task-oriented spoken language understanding. Journal of Physics Conference Series. 2019;1267:012023. http://dx.doi.org/10.1088/1742-6596/1267/1/012023</mixed-citation><mixed-citation xml:lang="en">Lixian Hou, Yanling Li, Chengcheng Li, et al. Review of research on task-oriented spoken language understanding. Journal of Physics Conference Series. 2019;1267:012023. http://dx.doi.org/10.1088/1742-6596/1267/1/012023</mixed-citation></citation-alternatives></ref><ref id="cit7"><label>7</label><citation-alternatives><mixed-citation xml:lang="ru">Ashish Vaswani, Noam Shazeer, Niki Parmar, et al. Attention Is All You Need. In: proc. 31st Conf. on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA. https://arxiv.org/abs/1706.03762</mixed-citation><mixed-citation xml:lang="en">Ashish Vaswani, Noam Shazeer, Niki Parmar, et al. Attention Is All You Need. In: proc. 31st Conf. on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA. https://arxiv.org/abs/1706.03762</mixed-citation></citation-alternatives></ref><ref id="cit8"><label>8</label><citation-alternatives><mixed-citation xml:lang="ru">Jacob Devlin, Ming-Wei Chang, Kenton Lee, et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Computing Research Repository. 2018. P. 1-16.</mixed-citation><mixed-citation xml:lang="en">Jacob Devlin, Ming-Wei Chang, Kenton Lee, et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Computing Research Repository. 2018. P. 1-16.</mixed-citation></citation-alternatives></ref><ref id="cit9"><label>9</label><citation-alternatives><mixed-citation xml:lang="ru">Matthew E Peters, Mark Neumann, Mohit Iyyer, et al. Deep contextualized word representations. In: Proc. NAACL-HLT. 2018;1:2227–2237.</mixed-citation><mixed-citation xml:lang="en">Matthew E Peters, Mark Neumann, Mohit Iyyer, et al. Deep contextualized word representations. In: Proc. NAACL-HLT. 2018;1:2227–2237.</mixed-citation></citation-alternatives></ref><ref id="cit10"><label>10</label><citation-alternatives><mixed-citation xml:lang="ru">Alec Radford, Karthik Narasimhan, Tim Salimans, et al. Improving Language Understanding by Generative PreTraining. Preprint. https://pdf4pro.com/amp/view/improving-language-understanding-by-generative-pre-training5b6487.html</mixed-citation><mixed-citation xml:lang="en">Alec Radford, Karthik Narasimhan, Tim Salimans, et al. Improving Language Understanding by Generative PreTraining. Preprint. https://pdf4pro.com/amp/view/improving-language-understanding-by-generative-pre-training5b6487.html</mixed-citation></citation-alternatives></ref><ref id="cit11"><label>11</label><citation-alternatives><mixed-citation xml:lang="ru">Shinji Watanabe, Takaaki Hori, Shigeki Karita, et al. ESPnet: End-to-End Speech Processing Toolkit. 2018. https://arxiv.org/abs/1804.00015</mixed-citation><mixed-citation xml:lang="en">Shinji Watanabe, Takaaki Hori, Shigeki Karita, et al. ESPnet: End-to-End Speech Processing Toolkit. 2018. https://arxiv.org/abs/1804.00015</mixed-citation></citation-alternatives></ref><ref id="cit12"><label>12</label><citation-alternatives><mixed-citation xml:lang="ru">Jason Li, Vitaly Lavrukhin, Boris Ginsburg, et al. Jasper: An End-to-End Convolutional Neural Acoustic Model. 2019. https://arxiv.org/abs/1904.03288</mixed-citation><mixed-citation xml:lang="en">Jason Li, Vitaly Lavrukhin, Boris Ginsburg, et al. Jasper: An End-to-End Convolutional Neural Acoustic Model. 2019. https://arxiv.org/abs/1904.03288</mixed-citation></citation-alternatives></ref><ref id="cit13"><label>13</label><citation-alternatives><mixed-citation xml:lang="ru">Chaitra Hegde, Shrikumar Patil. Unsupervised Paraphrase Generation using Pre-trained Language Models. 2020. https://arxiv.org/abs/2006.05477</mixed-citation><mixed-citation xml:lang="en">Chaitra Hegde, Shrikumar Patil. Unsupervised Paraphrase Generation using Pre-trained Language Models. 2020. https://arxiv.org/abs/2006.05477</mixed-citation></citation-alternatives></ref><ref id="cit14"><label>14</label><citation-alternatives><mixed-citation xml:lang="ru">Vineel Pratap, Awni Hannun, Qiantong Xu, et al. Wav2letter++: The Fastest Open-source Speech Recognition System. In: Proc. 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). https://doi.org/10.1109/ICASSP.2019.8683535</mixed-citation><mixed-citation xml:lang="en">Vineel Pratap, Awni Hannun, Qiantong Xu, et al. Wav2letter++: The Fastest Open-source Speech Recognition System. In: Proc. 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). https://doi.org/10.1109/ICASSP.2019.8683535</mixed-citation></citation-alternatives></ref><ref id="cit15"><label>15</label><citation-alternatives><mixed-citation xml:lang="ru">Schneider S, Baevski A, Collobert R, et al. Wav2vec: Unsupervised Pre-Training for Speech Recognition. In: Proc. Interspeech 2019, 20th Annual Conference of the International Speech Communication Association. P. 3465- 3469. https://doi.org/10.21437/Interspeech.2019-1873</mixed-citation><mixed-citation xml:lang="en">Schneider S, Baevski A, Collobert R, et al. Wav2vec: Unsupervised Pre-Training for Speech Recognition. In: Proc. Interspeech 2019, 20th Annual Conference of the International Speech Communication Association. P. 3465- 3469. https://doi.org/10.21437/Interspeech.2019-1873</mixed-citation></citation-alternatives></ref><ref id="cit16"><label>16</label><citation-alternatives><mixed-citation xml:lang="ru">Alexis Conneau, Guillaume Lample. Cross-lingual Language Model Pretraining. In: Proc. 33rd Conference on Neural Information Processing Systems. 2019. P. 7057-7067.</mixed-citation><mixed-citation xml:lang="en">Alexis Conneau, Guillaume Lample. Cross-lingual Language Model Pretraining. In: Proc. 33rd Conference on Neural Information Processing Systems. 2019. P. 7057-7067.</mixed-citation></citation-alternatives></ref><ref id="cit17"><label>17</label><citation-alternatives><mixed-citation xml:lang="ru">Zhilin Yang, Zihang Dai, Yiming Yang, et al. XLNet: Generalized Autoregressive Pretraining for Language Understanding. 2019. https://arxiv.org/abs/1906.08237</mixed-citation><mixed-citation xml:lang="en">Zhilin Yang, Zihang Dai, Yiming Yang, et al. XLNet: Generalized Autoregressive Pretraining for Language Understanding. 2019. https://arxiv.org/abs/1906.08237</mixed-citation></citation-alternatives></ref><ref id="cit18"><label>18</label><citation-alternatives><mixed-citation xml:lang="ru">Yinhan Liu, Myle Ott, Naman Goyal, et al. RoBERTa: A Robustly Optimized BERT Pretraining Approach. ICLR 2020 Conference Blind Submission. 2019. https://doi.org/10.48550/arXiv.1907.11692</mixed-citation><mixed-citation xml:lang="en">Yinhan Liu, Myle Ott, Naman Goyal, et al. RoBERTa: A Robustly Optimized BERT Pretraining Approach. ICLR 2020 Conference Blind Submission. 2019. https://doi.org/10.48550/arXiv.1907.11692</mixed-citation></citation-alternatives></ref><ref id="cit19"><label>19</label><citation-alternatives><mixed-citation xml:lang="ru">Manning Kevin Clark, Minh-Thang Luong, Quoc V Le, et al. ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. ICLR 2020 Conference Blind Submission. 2020. https://openreview.net/forum?id=r1xMH1BtvB</mixed-citation><mixed-citation xml:lang="en">Manning Kevin Clark, Minh-Thang Luong, Quoc V Le, et al. ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators. ICLR 2020 Conference Blind Submission. 2020. https://openreview.net/forum?id=r1xMH1BtvB</mixed-citation></citation-alternatives></ref><ref id="cit20"><label>20</label><citation-alternatives><mixed-citation xml:lang="ru">Medennikov I, Korenevsky M, Prisyach T, et al. The STC System for the CHiME-6 Challenge. In: Proc. 6th International Workshop on Speech Processing in Everyday Environments (CHiME 2020). 2020. P. 36-41.</mixed-citation><mixed-citation xml:lang="en">Medennikov I, Korenevsky M, Prisyach T, et al. The STC System for the CHiME-6 Challenge. In: Proc. 6th International Workshop on Speech Processing in Everyday Environments (CHiME 2020). 2020. P. 36-41.</mixed-citation></citation-alternatives></ref><ref id="cit21"><label>21</label><citation-alternatives><mixed-citation xml:lang="ru">Greg Brockman, Mira Murati, Peter Welinder. OpenAI API. 2020 : OpenAI Blog. https://openai.com/blog/openai-api/</mixed-citation><mixed-citation xml:lang="en">Greg Brockman, Mira Murati, Peter Welinder. OpenAI API. 2020 : OpenAI Blog. https://openai.com/blog/openai-api/</mixed-citation></citation-alternatives></ref><ref id="cit22"><label>22</label><citation-alternatives><mixed-citation xml:lang="ru">Zhenzhong Lan, Mingda Chen, Sebastian Goodman, et al. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. 2020. https://arxiv.org/abs/1909.11942</mixed-citation><mixed-citation xml:lang="en">Zhenzhong Lan, Mingda Chen, Sebastian Goodman, et al. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. 2020. https://arxiv.org/abs/1909.11942</mixed-citation></citation-alternatives></ref><ref id="cit23"><label>23</label><citation-alternatives><mixed-citation xml:lang="ru">Yiming Cui, Wanziang Che, Ting Liu, et al. Pre-Training With Whole Word Masking for Chinese BERT. In: IEEE/ACM Transactions on Audio, Speech, and Language Processing. 2021;29:3504-3514. https://doi.org/10.1109/TASLP.2021.3124365</mixed-citation><mixed-citation xml:lang="en">Yiming Cui, Wanziang Che, Ting Liu, et al. Pre-Training With Whole Word Masking for Chinese BERT. In: IEEE/ACM Transactions on Audio, Speech, and Language Processing. 2021;29:3504-3514. https://doi.org/10.1109/TASLP.2021.3124365</mixed-citation></citation-alternatives></ref><ref id="cit24"><label>24</label><citation-alternatives><mixed-citation xml:lang="ru">Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, et al. PaLM: Scaling Language Modeling with Pathways. 2022. https://arxiv.org/abs/2204.02311</mixed-citation><mixed-citation xml:lang="en">Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, et al. PaLM: Scaling Language Modeling with Pathways. 2022. https://arxiv.org/abs/2204.02311</mixed-citation></citation-alternatives></ref></ref-list><fn-group><fn fn-type="conflict"><p>The authors declare that there are no conflicts of interest present.</p></fn></fn-group></back></article>
