Публікація:
Research of the text processing methods in organization of electronic storages of information objects

dc.contributor.authorBarkovska, O.
dc.contributor.authorKhomych, V.
dc.contributor.authorNastenko, O.
dc.date.accessioned2024-01-18T18:33:49Z
dc.date.available2024-01-18T18:33:49Z
dc.date.issued2022
dc.description.abstractThe subject matter of the article is electronic storage of information objects (IO) ordered by specified rules at the stage of accumulation of qualification thesis and scientific work of the contributors of the offered knowledge exchange system provided to the system in different formats (text, graphic, audio). Classified works of contributors of the system are the ground for organization of thematic rooms for discussion to spread scientific achievements, to adopt new ideas, to exchange knowledge and to look for employers or mentors in different countries. The goal of the work is to study the libraries of text processing and analysis to speed-up and increase accuracy of the scanned text documents classification in the process of serialized electronic storage of information objects organization. The following tasks are: to study the text processing methods on the basis of the proposed generalized model of the system of classification of scanned documents with the specified location of the block of text processing and analysis; to investigate the statistics of change in the execution time of the developed parallel modification of the methods of the word processing module for the system with shared memory for collections of text documents of different sizes; analyze the results. The methods used are the following: parallel digital sorting methods, methods of mathematical statistics, linguistic methods of text analysis. The following results were obtained: in the course of the research fulfillment the generalized model of the scanned documents classification system that consist of image processing unit and text processing unit that include unit of the scanned image previous processing; text detection unit; previous text processing; compiling of the frequency dictionary; text proximity detection was offered. Conclusions: the proposed parallel modification of the previous text processing unit gives acceleration up to 3,998 times. But, at a very high computational load (collection of 18144 files, about 1100 MB), the resources of an ordinary multiprocessor-based computer with the shared memory obviously is not enough to solve such problems in the mode close to real time.
dc.identifier.citationBarkovska O. Research of the text processing methods in organization of electronic storages of information objects / O. Barkovska, V. Khomych, O. Nastenko // Сучасний стан наукових досліджень та технологій в промисловості. – 2022. – № 1(19). – С. 5–12.
dc.identifier.urihttps://openarchive.nure.ua/handle/document/25339
dc.language.isoen
dc.publisherХНУРЕ
dc.subjectinformation system
dc.subjectparallelism
dc.subjectword processing
dc.subjectlinguistic programming
dc.subjectacceleration
dc.titleResearch of the text processing methods in organization of electronic storages of information objects
dc.typeArticle
dspace.entity.typePublication

Файли

Оригінальний пакет
Зараз показано 1 - 1 з 1
Завантаження...
Зображення мініатюри
Назва:
EOM_SSND_Pr_2022_n1_5-12.pdf
Розмір:
522.58 KB
Формат:
Adobe Portable Document Format
Ліцензійний пакет
Зараз показано 1 - 1 з 1
Немає доступних мініатюр
Назва:
license.txt
Розмір:
9.55 KB
Формат:
Item-specific license agreed upon to submission
Опис: