Публікація:
Methods for creating datasets for small-resource language pairs

dc.contributor.authorLysenko, D. O.
dc.date.accessioned2026-07-14T07:25:09Z
dc.date.issued2026
dc.description.abstractThis research addresses the critical challenge of data scarcity in developing direct neural machine translation (NMT) systems for low-resource language pairs without relying on a high-resource pivot language like English. To overcome extreme data limitations, the research proposes a fully automated pipeline that combines bitext mining via language-agnostic embedding spaces, synthetic data generation using Large Language Models (LLMs), and internal structural data augmentation. By integrating these generative methodologies with rigorous automated noise filtering, the proposed approach successfully scales microscopic seed dictionaries into robust, high-quality parallel corpora suitable for training accurate and culturally authentic translation models
dc.identifier.citationLysenko D. O. Methods for creating datasets for small-resource language pairs // Радіоелектроніка та молодь у XXI столітті : матеріали 30-го Міжнар. молодіж. форуму, 22–24 квітня 2026 р. Харків, 2026. Т. 7. С. 19-21.
dc.identifier.urihttps://openarchive.nure.ua/handle/document/35526
dc.language.isoen
dc.publisherХНУРЕ
dc.subjectneural machine translation
dc.subjectlow-resource languages
dc.subjectparallel corpus construction
dc.titleMethods for creating datasets for small-resource language pairs
dc.typeConference proceedings
dspace.entity.typePublication

Файли

Оригінальний пакунок

Зараз показано 1 - 1 з 1
Завантаження...
Зображення мініатюри
Назва:
MRF_2026_T7_KZSA_19-21.pdf
Розмір:
190.37 KB
Формат:
Adobe Portable Document Format

Пакунок ліцензії

Зараз показано 1 - 1 з 1
Завантаження...
Зображення мініатюри
Назва:
license.txt
Розмір:
10.74 KB
Формат:
Item-specific license agreed upon to submission
Опис: