Публікація: Methods for creating datasets for small-resource language pairs
| dc.contributor.author | Lysenko, D. O. | |
| dc.date.accessioned | 2026-07-14T07:25:09Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | This research addresses the critical challenge of data scarcity in developing direct neural machine translation (NMT) systems for low-resource language pairs without relying on a high-resource pivot language like English. To overcome extreme data limitations, the research proposes a fully automated pipeline that combines bitext mining via language-agnostic embedding spaces, synthetic data generation using Large Language Models (LLMs), and internal structural data augmentation. By integrating these generative methodologies with rigorous automated noise filtering, the proposed approach successfully scales microscopic seed dictionaries into robust, high-quality parallel corpora suitable for training accurate and culturally authentic translation models | |
| dc.identifier.citation | Lysenko D. O. Methods for creating datasets for small-resource language pairs // Радіоелектроніка та молодь у XXI столітті : матеріали 30-го Міжнар. молодіж. форуму, 22–24 квітня 2026 р. Харків, 2026. Т. 7. С. 19-21. | |
| dc.identifier.uri | https://openarchive.nure.ua/handle/document/35526 | |
| dc.language.iso | en | |
| dc.publisher | ХНУРЕ | |
| dc.subject | neural machine translation | |
| dc.subject | low-resource languages | |
| dc.subject | parallel corpus construction | |
| dc.title | Methods for creating datasets for small-resource language pairs | |
| dc.type | Conference proceedings | |
| dspace.entity.type | Publication |
Файли
Оригінальний пакунок
1 - 1 з 1
Завантаження...
- Назва:
- MRF_2026_T7_KZSA_19-21.pdf
- Розмір:
- 190.37 KB
- Формат:
- Adobe Portable Document Format
Пакунок ліцензії
1 - 1 з 1
Завантаження...
- Назва:
- license.txt
- Розмір:
- 10.74 KB
- Формат:
- Item-specific license agreed upon to submission
- Опис: