The Impact of AI in the Field of Sound for Picture. A Historical, Practical, and Ethical Consideration
DOI:
https://doi.org/10.37130/bf780v59Keywords:
AI, sound restoration, generative speech, generative music, spleeter, machine learning, demixing, demucsAbstract
The paper covers the emergence and evolution of AI in the commercial field of audio: from the pioneering German audio software company Prosoniq Products Software, which first used artificial neural networks for commercial audio processing, creating in 1997 the program Pandora Music Decomposition Series that managed to separate music into its components, to the indispensable technologies for sound processing involving machine learning (iZotope and Audionamix ADX Trax Pro 3), to the spleeter technology, ChatGPT from OpenAI that applies AI technology also to the area of sound processing, to Adobe or ElevenLabs that offers online AI services with Text to Speech and Speech to Speech functions. The second part of the paper will review specific audio plug-ins that use Artificial Intelligence, making it possible to identify to what extent AI is an important tool in the field of sound and music, but also what are its current limitations. The last part of the paper refers to the ethics of AI in the field of sound. Even if plugins facilitate the work process in sound design, music production, voice cloning and synthesis, or sound restoration, thus reducing costs and personnel, what are the ethical limits of using these tools?
References
Adobe (2016) #VoCo. Adobe Audio Manipulator Sneak Peak with Jordan Peele [Online]. Available at: https://www.youtube.com/watch?v=I3l4XLZ59iw/ (Accessed: 20 May 2024).
Adobe (2023a) Behind the Tech: Enhance Speech in Adobe Podcast [Online]. Available at: https://research.adobe.com/news/behind-the-tech-enhance-speech-in-adobe-podcast/ (Accessed: 20 May, 2024).
Adobe (2023b) #ProjectDubDubDub | Adobe MAX Sneaks 2023 [Online]. Available at: https://www.youtube.com/watch?v=fZY-Cv1Q8NY/ (Accessed: 20 May 2024).
Agostinelli, A. et al (2023) ‘MusicLM: Generating Music From Text’, arXiv:2301.11325v1 [cs.SD] [Online]. Available at: https://arxiv.org/pdf/2301.11325 (Accessed: 20 May 2024).
Aicrowd (2024) Crowdsourcing AI to Solve Real-World Problems [Online]. https://www.aicrowd.com (Accessed: 20 May 2024).
Andersen, A. (2024a) Huge sound design contest! 7 chances to win fantastic prizes from Mattia Cellotto [Online]. Available at: https://www.asoundeffect.com/sounddesign24/ (Accessed: 20 May 2024).
Andersen, A. (2024b) Here are the winners of the Mattia Cellotto sound design contest! [Online]. Available at: https://www.asoundeffect.com/sounddesign-contest-results/ (Accessed: 20 May 2024).
Arhiva Activă UNATC (2024) Oamenii nu sunt capre [Online]. Available at: https://arhiva.unatc.ro/filme/oamenii-nu-sunt-capre/ (Accessed: 20 May 2024).
ASoundEffect (2024) Sound Design Contest: Create the sound for this video for 7 chances to win wild prizes! [Online]. Available at: https://www.youtube.com/watch?v=X_wsXpUxOaY (Accessed: 20 May 2024).
Collins, E. and Eck, D. (2024) New generative media models and tools, built with and for creators [Online]. Available at: https://blog.google/technology/ai/google-generative-ai-veo-imagen-3/ (Accessed: 20 May, 2024).
Computer Music Journal (1997) ‘Products of Interest’, Computer Music Journal, vol. 21, no. 3, p. 116. [Online]. Available at: http://www.jstor.org/stable/3681029 (Accessed: 20 May 2024).
Défossez, A. et al (2019) Music Source Separation in the Waveform Domain’, arXiv:1911.13254v1 [cs.SD] [Online]. Available at: https://arxiv.org/pdf/1911.13254v1 (Accessed: 20 May 2024).
Descript (2020) Introducing Descript [Online]. Available at: https://www.youtube.com/watch?v=Bl9wqNe5J8U/ (Accessed: 20 May 2024).
Devideconcept (2024) Audio AI Comparisons [Online]. Available at: https://divideconcept.github.io (Accessed: 20 May 2024).
ElevenLabs (2024) Introducing Voice Actor Payouts [Online]. Available at: https://elevenlabs.io/blog/introducing-voice-actor-payouts/ (Accessed: 20 May 2024).
Grant, N. (2024) Google Chatbot’s A.I. Images Put People of Color in Nazi-Era Uniforms [Online]. Available at: https://www.nytimes.com/2024/02/22/technology/google-gemini-german-uniforms.html (Accessed: 20 May 2024).
Hiatt, Brian (2024) A ChatGPT for Music Is Here. Inside Suno, the Startup Changing Everything [Online]. Available at: https://www.rollingstone.com/music/music-features/suno-ai-chatgpt-for-music-1234982307/ (Accessed: 20 May 2024).
Hsu et al (2023) Audiobox: Unified Audio Generation with Natural Language Prompts [Online]. Available at: https://ai.meta.com/research/publications/audiobox-unified-audio-generation-with-natural-language-prompts/ (Accessed: 20 May 2024).
The Beatles (2020) The Beatles: Get Back - A Sneak Peek from Peter Jackson [Online]. Available at: https://www.youtube.com/watch?v=UocEGvQ10OE (Accessed: 20 May 2024).
Hurwitz, M. (2021) The Making of The Beatles’ Let It Be and Peter Jackson’s Get Back Peter Jackson’s The Beatles: Get Back [Online]. Available at: https://www.soundandvision.com/content/making-beatles-let-it-be-and-peter-jacksons-get-back-peter-jacksons-beatles-get-back/ (Accessed: 20 May 2024).
Lehrman, P. D. (1998) ‘Prosoniq Sonicworx Artist. Audio Editing Software.’ Sound on Sound, vol. 13, issue 5, March. [Online]. Available at: https://www.soundonsound.com/reviews/prosoniq-sonicworx-artist/ (Accessed 20 May 2024).
Lin, David Chuan-En et al (2021) Soundify: Matching Sound Effects to Video, arXiv:2112.09726v1 [cs.SD] [Online]. Available at: https://arxiv.org/pdf/2112.09726v1 (Accessed: 20 May 2024).
Lin, David Chuan-En et al (2021) Soundify: Matching Sound Effects to Video, arXiv:2112.09726v3 [cs.SD] [Online]. Available at: https://arxiv.org/pdf/2112.09726v3 (Accessed: 20 May 2024).
Meta (2023a) Introducing Voicebox: The Most Versatile AI for Speech Generation [Online]. Available at: https://about.fb.com/news/2023/06/introducing-voicebox-ai-for-speech-generation/ (Accessed: 20 May 2024).
Meta (2023b) Open sourcing AudioCraft: Generative AI for audio made simple and available to all [Online]. Available at: https://ai.meta.com/blog/audiocraft-musicgen-audiogen-encodec-generative-ai-audio/ (Accessed: 20 May 2024).
Meta (2023c) Introducing Voicebox: The Most Versatile AI for Speech Generation [Online]. Available at: https://about.fb.com/news/2023/06/introducing-voicebox-ai-for-speech-generation/ (Accessed: 20 May 2024).
Microsoft Copilot (2023) Turn your ideas into songs with Suno on Microsoft Copilot [Online]. Available at: https://www.microsoft.com/en-us/microsoft-copilot/blog/2023/12/19/turn-your-ideas-into-songs-with-suno-on-microsoft-copilot/ (Accessed: 20 May 2024).
Morrison, R. (2024) Pika Labs sound effects now available for all users — I tried it and here is how it sounds [Online]. Available at: https://www.tomsguide.com/ai/ai-image-video/pika-labs-sound-effects-now-available-for-all-users-i-tried-it-and-here-is-how-it-sounds/ (Accessed: 20 May 2024).
Moussallam, M. (2019) Releasing Spleeter: Deezer Research source separation engine [Online]. Available at: https://deezer.io/releasing-spleeter-deezer-r-d-source-separation-engine-2b88985e797e/ (Accessed: 20 May 2024).
Nuñez, Michael (2024) Former Google DeepMind researchers launch AI-powered music creation app Udio [Online]. Available at https://venturebeat.com/ai/former-google-deepmind-researchers-launch-ai-powered-music-creation-app-udio/ (Accessed: 20 May 2024).
OpenAI (2024a) Navigating the Challenges and Opportunities of Synthetic Voices [Online]. Available at: https://openai.com/index/navigating-the-challenges-and-opportunities-of-synthetic-voices/ (Accessed: 20 May 2024).
OpenAI (2024b) Video generation models as world simulators [Online]. Available at: https://openai.com/index/video-generation-models-as-world-simulators/ (Accessed: 20 May 2024).
Owens, A. et al (2016) Visually Indicated Sounds, arXiv:1512.08512v2 [cs.CV] [Online]. Available at: https://arxiv.org/pdf/1512.08512 (Accessed: 20 May, 2024).
Paul, L. and Millman, E. (2023) Viral Drake and The Weeknd AI Collaboration Pulled From Apple, Spotify [Online]. Available at: https://www.rollingstone.com/music/music-news/viral-drake-and-the-weeknd-collaboration-is-completely-ai-generated-1234716154/ (Accessed: 20 May 2024).
Prevost, John J. and Ghose, Sanchita (2020) AutoFoley: Artificial Synthesis of Synchronized Sound Tracks for Silent Videos with Deep Learning, arXiv:2002.10981v1 [cs.SD] [Online]. Available at: https://arxiv.org/pdf/2002.10981/ (Accessed: 20 May 2024).
Prosoniq (2010) Prosoniq get the horn. Free World Cup audio plug-in released [Online]. Available at: https://www.soundonsound.com/news/prosoniq-get-horn/ (Accessed: 20 May 2024).
Reid, G. (2003) ‘Hartmann Neuron. Neuronal Resynthesizing Keyboard’ [Online]. Available at: https://www.soundonsound.com/reviews/hartmann-neuron/ (Accessed: 20 May 2024)
Rose, Jey (2018) Neural Networks: A new way to think about processing, CAS Quarterly [Online]. Available online at: https://digital.copcomm.com/i/933783-winter-2018/29? (Accessed: 20 May 2024).
Rucăreanu, D. Ș. (2024a) AI Restored example - student short film from UNATC archives [Online]. Available at: https://www.youtube.com/watch?v=5wc5HyzWbho (Accessed: 20 May 2024).
Rucăreanu, D. Ș. (2024b) AI English Dubbing Example - student short film from UNATC archives [Online]. Available at: https://www.youtube.com/watch?v=sKjcw1ADxlQ (Accessed: 20 May 2024).
Rucăreanu, D. Ș. (2024c) Audio Example 1 - generated with Meta Audiobox (Beta) [Online]. Available at: https://www.youtube.com/watch?v=lpas-3FkCKU (Accessed: 20 May 2024).
Rucăreanu, D. Ș. (2024d) Audio Example 2 - Generated with ElevenLabs Sound Effect (Beta) [Online]. Available at: https://www.youtube.com/watch?v=lmo099JKVhc (Accessed: 20 May 2024).
Rucăreanu, D. Ș. (2024e). AI Sound Design Example - audio elements generated with ElevenLabs Sound effects (Beta) [Online]. Available at: https://www.youtube.com/watch?v=teJ0g2pkmtI (Accessed: 20 May 2024).
Staniszewski, M. (2024) Introducing Dubbing Studio. Localize videos with precise control over transcript, translation, timing, and more [Online]. Available at: https://elevenlabs.io/blog/introducing-dubbing-studio/ (Accessed: 20 May 2024).
Thomas, B. (2017) Audionamix ADX Trax Pro 3 SP. Speech & Melody Separation Software [Online]. Available online at: https://www.soundonsound.com/reviews/audionamix-adx-trax-pro-3-sp/ (Accessed: 20 May 2024).
van den Oord, A. and Dieleman, S. (2016) WaveNet: A generative model for raw audio [Online]. Available at: https://deepmind.google/discover/blog/wavenet-a-generative-model-for-raw-audio/ (Accessed: 20 May 2024).
Wichern, G. (2017) What the Machine Learning in RX 6 Advanced Means for the Future of Audio Repair Technology [Online]. Available at: https://www.izotope.com/en/learn/what-the-machine-learning-in-rx-6-advanced-means-for-the-future-of-audio-repair-technology.html (Accessed: 20 May 2024).
Zhou, Y. et al (2018) Visual to Sound: Generating Natural Sound for Videos in the Wild, arXiv:1712.01393v2 [cs.CV] [Online]. Available at: https://arxiv.org/pdf/1712.01393/ (Accessed: 20 May 2024).
Zynaptiq (2013) Zyn1aptiq Acquires Prosoniq Product Line And Technologies [Online]. Available at: https://www.zynaptiq.com/more/zynaptiq-acquires-prosoniq-product-line-and-technologies/d106b1606afa940345ffb4a939c4588c/ (Accessed: 20 May 2024).
People are not Goats (1954) [DVD] Directed by Dragos Witkowski [Film]. București: Arhiva UNATC.