Публікація: Гібридна система детекції синтезованого мовлення на основі SSL-ембеддінгів
Завантаження...
Дата
Автори
Назва журналу
ISSN журналу
Назва тому
Видавець
ХНУРЕ
Анотація
This paper presents a hybrid system for synthetic speech detection that integrates spectral, phase-based, and prosodic acoustic features with self-supervised learning (SSL) embeddings extracted from the WavLM model. The proposed multimodal approach combines low-level handcrafted features (LFCC, jitter, shimmer, phase stability) with high-level context-aware representations derived from intermediate transformer layers. Such integration enables the model to capture both fine-grained acoustic artifacts and global structural inconsistencies in synthetic speech. The proposed architecture shows strong potential for applications in biometric security and digital media integrity.
Опис
Ключові слова
SSL-ембеддінг, синтезоване мовлення
Цитування
Ростовцев В. В., Гриньов С. А. Гібридна система детекції синтезованого мовлення на основі SSL-ембеддінгів // Радіоелектроніка та молодь у XXI столітті : матеріали 30-го Міжнар. молодіж. форуму, 22–24 квітня 2026 р. Харків, 2026. Т. 6. С. 105-107.