Публікація:
Methods for creating datasets for small-resource language pairs

Завантаження...
Зображення мініатюри

Дата

Автори

Назва журналу

ISSN журналу

Назва тому

Видавець

ХНУРЕ

Дослідницькі проекти

Організаційні одиниці

Випуск журналу

Анотація

This research addresses the critical challenge of data scarcity in developing direct neural machine translation (NMT) systems for low-resource language pairs without relying on a high-resource pivot language like English. To overcome extreme data limitations, the research proposes a fully automated pipeline that combines bitext mining via language-agnostic embedding spaces, synthetic data generation using Large Language Models (LLMs), and internal structural data augmentation. By integrating these generative methodologies with rigorous automated noise filtering, the proposed approach successfully scales microscopic seed dictionaries into robust, high-quality parallel corpora suitable for training accurate and culturally authentic translation models

Опис

Ключові слова

neural machine translation, low-resource languages, parallel corpus construction

Цитування

Lysenko D. O. Methods for creating datasets for small-resource language pairs // Радіоелектроніка та молодь у XXI столітті : матеріали 30-го Міжнар. молодіж. форуму, 22–24 квітня 2026 р. Харків, 2026. Т. 7. С. 19-21.

DOI

Схвалення

Рецензія

Доповнено

На які посилаються