استنساخ الصوت باستخدام الشبكات العصبية العميقة: التقنيات، التقييم، التطبيقات، والاعتبارات الأخلاقية
DOI:
https://doi.org/10.5281/zenodo.20266741الكلمات المفتاحية:
استنساخ الصوت، تحويل النص إلى كلام، توليد الكلام، التكنولوجيا المساعدة، الذكاء الاصطناعي الأخلاقيالملخص
يُعد استنساخ الصوت أحد التطبيقات التحويلية للتقنيات المعتمدة على الشبكات العصبية العميقة، حيث يُمكن من خلاله توليد أصوات اصطناعية تشبه إلى حد كبير الصوت البشري الحقيقي. تقدم هذه الورقة مراجعة شاملة لتقنيات استنساخ الصوت، مع التركيز على تطور أنظمة تحويل النص إلى كلام (TTS) من النماذج التقليدية إلى النماذج الحديثة المعتمدة على التعلم العميق مثل, Tacotron WaveNet ,VALL-E .
نستعرض في هذه الدراسة مكونات أنظمة تحويل النص إلى كلام، بما في ذلك المشفرات الصوتية، والمولدات، والمركبات الصوتية العصبية، مع التمييز بين أنظمة استنساخ الصوت أحادية المتكلم ومتعددة المتكلمين. كما نناقش التطبيقات الواقعية في مجالات الاتصالات والتعليم ودعم ذوي الاحتياجات الخاصة والترفيه، إلى جانب التحديات الأخلاقية المهمة مثل انتهاك الخصوصية، ونشر المعلومات المضللة، والتلاعب العاطفي.
تختتم الورقة بعرض لأبرز التحديات التقنية الحالية والاتجاهات المستقبلية، بما في ذلك التعلم الموحد (Federated Learning) والمركبات الصوتية المعتمدة على المحولات (Transformer Vocoders) ونماذج الانتشار (Diffusion Models)، والتي تهدف إلى تحسين جودة وكفاءة واستدامة هذه التكنولوجيا مع مراعاة الجوانب الأخلاقية
التنزيلات
المراجع
[1] F. Khanam et al., “Text to speech synthesis: A systematic review, deep learning based architecture and future research direction,” Journal of Advances in Information Technology, vol. 13, no. 5, 2022.
[2] T. Dutoit, “High-quality text-to-speech synthesis: An overview,” Journal of Electrical and Electronics Engineering Australia, vol. 17, no. 1, pp. 25–36, 1997.
[3] D. Sasirekha and E. Chandra, “Text to speech: A simple tutorial,” International Journal of Soft Computing and Engineering (IJSCE), vol. 2, no. 1, pp. 275–278, 2012.
[4] X. Tan et al., “A survey on neural speech synthesis,” arXiv preprint arXiv:2106.15561, 2021.
[5] D. Schwarz, “Concatenative sound synthesis: The early years,” Journal of New Music Research, vol. 35, no. 1, pp. 3–22, 2006.
[6] H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,” Speech Communication, vol. 51, no. 11, pp. 1039–1064, 2009.
[7] Y. Ren et al., “FastSpeech: Fast, robust and controllable text to speech,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 32, 2019.
[8] Z. Qin et al., “OpenVoice: Versatile instant voice cloning,” arXiv preprint arXiv:2312.01479, 2023.
[9] Đ. T. T. Trang, “Overview of voice cloning.”
[10] P. Neekhara et al., “Expressive neural voice cloning,” in Proceedings of the Asian Conference on Machine Learning (ACML), 2021.
[11] A. van den Oord et al., “WaveNet: A generative model for raw audio,” arXiv preprint arXiv:1609.03499, 2016.
[12] D. Rethage, J. Pons, and X. Serra, “A WaveNet for speech denoising,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2018.
[13] C. McGettigan et al., “Voice cloning: Psychological and ethical implications of intentionally synthesising familiar voice identities,” 2024.
[14] B. Wells-Edwards, “What’s in a voice? The legal implications of voice cloning,” Arizona Law Review, vol. 64, p. 1213, 2022.
[15] N. Veerasamy and H. Pieterse, “Rising above misinformation and deepfakes,” in Proceedings of the International Conference on Cyber Warfare and Security, vol. 17, no. 1, 2022.
[16] O. M. Ijiga et al., “Harmonizing the voices of AI: Exploring generative music models, voice cloning, and voice transfer for creative expression,” World Journal of Advanced Engineering and Technology Sciences, vol. 11, 2024.
التنزيلات
منشور
الرخصة
الحقوق الفكرية (c) 2025 مجلة الجامعة الوطنية الخاصة

هذا العمل مرخص بموجب Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.