Публікація: Methods for creating datasets for small-resource language pairs
Завантаження...
Дата
Автори
Назва журналу
ISSN журналу
Назва тому
Видавець
ХНУРЕ
Анотація
This research addresses the critical challenge of data scarcity in developing direct neural machine translation (NMT) systems for low-resource language pairs without relying on a high-resource pivot language like English. To overcome extreme data limitations, the research proposes a fully automated pipeline that combines bitext mining via language-agnostic embedding spaces, synthetic data generation using Large Language Models (LLMs), and internal structural data augmentation. By integrating these generative methodologies with rigorous automated noise filtering, the proposed approach successfully scales microscopic seed dictionaries into robust, high-quality parallel corpora suitable for training accurate and culturally authentic translation models
Опис
Ключові слова
neural machine translation, low-resource languages, parallel corpus construction
Цитування
Lysenko D. O. Methods for creating datasets for small-resource language pairs // Радіоелектроніка та молодь у XXI столітті : матеріали 30-го Міжнар. молодіж. форуму, 22–24 квітня 2026 р. Харків, 2026. Т. 7. С. 19-21.