Публікація:
Аналіз питання перекладу тексту на зображеннях на основі мультимодальних великих мовних моделей

Завантаження...
Зображення мініатюри

Дата

Назва журналу

ISSN журналу

Назва тому

Видавець

ХНУРЕ

Дослідницькі проекти

Організаційні одиниці

Випуск журналу

Анотація

Traditional image translation typically relies on sequential OCR and machine translation, which often fail to account for visual layouts and context. This research evaluates multimodal large language models in simultaneously processing textual and visual data to ensure accurate, context-aware translations while preserving document topology. The study demonstrates that these models automate 80–90% of visual content localisation by maintaining structural integrity across various formats. Despite these advancements, human verification remains necessary for low-quality images to prevent errors. Future work in this field will focus on technologies like visual font transfer to achieve seamless, design-consistent localisation

Опис

Ключові слова

мультимодальні великі мовні моделі, MLLM, переклад тексту на зображеннях

Цитування

Лазарєва Д. Є. Аналіз питання перекладу тексту на зображеннях на основі мультимодальних великих мовних моделей // Радіоелектроніка та молодь у XXI столітті : матеріали 30-го Міжнар. молодіж. форуму, 22–24 квітня 2026 р. Харків, 2026. Т. 7. С. 94-96.

DOI

Схвалення

Рецензія

Доповнено

На які посилаються