The Impact of AI in the Field of Sound for Picture. A Historical, Practical, and Ethical Consideration

Authors

  • Dan-Ștefan Rucăreanu I.L. Caragiale UNATC, Bucharest, Romania Author
  • Laura Lăzărescu-Thois I.L. Caragiale UNATC, Bucharest, Romania. University Politehnica of Bucharest Author

DOI:

https://doi.org/10.37130/bf780v59

Keywords:

AI, sound restoration, generative speech, generative music, spleeter, machine learning, demixing, demucs

Abstract

The paper covers the emergence and evolution of AI in the commercial field of audio: from the pioneering German audio software company Prosoniq Products Software, which first used artificial neural networks for commercial audio processing, creating in 1997 the program Pandora Music Decomposition Series that managed to separate music into its components, to the indispensable technologies for sound processing involving machine learning (iZotope and Audionamix ADX Trax Pro 3), to the spleeter technology, ChatGPT from OpenAI that applies AI technology also to the area of sound processing, to Adobe or ElevenLabs that offers online AI services with Text to Speech and Speech to Speech functions. The second part of the paper will review specific audio plug-ins that use Artificial Intelligence, making it possible to identify to what extent AI is an important tool in the field of sound and music, but also what are its current limitations. The last part of the paper refers to the ethics of AI in the field of sound. Even if plugins facilitate the work process in sound design, music production, voice cloning and synthesis, or sound restoration, thus reducing costs and personnel, what are the ethical limits of using these tools?

Author Biographies

  • Dan-Ștefan Rucăreanu, I.L. Caragiale UNATC, Bucharest, Romania

    Dan-Ștefan Rucăreanu is a university lecturer at the I.L. Caragiale National University of Theatre and Film in Bucharest, in the Multimedia Department: Sound and Film Editing, teaching Sound Design for film, animation, new media, video games and film archive restoration. He is also the Study Programme Coordinator of Sound in the Multimedia Department. In parallel, he works as a sound designer and re-recording mixer in the film industry, participating in the making of a large number of films, documentaries and animations, many of which have been internationally awarded. He is a member of the European Film Academy and has been nationally awarded for sound design. With over 16 years of experience in the field of sound for cinematography and media, he is actively anchored in the domain of artificial intelligence for sound, each time looking for new methods involving the use of emerging AI technologies in his practice.

  • Laura Lăzărescu-Thois, I.L. Caragiale UNATC, Bucharest, Romania. University Politehnica of Bucharest

    Editing Department of the Film Faculty of the I.L. Caragiale UNATC in Bucharest. In 2009 she became a student of the Doctoral School of UNATC, with a research topic relating to sound in the American animation film. She held two research grants in Berlin at the Hochschule für Film und Fernsehen “Konrad Wolf” in PotsdamBabelsberg, Germany, and in June 2012 she obtained her PhD degree in the field of Cinematography and Media. In 2018 she finished a second study, BA in marketing at the Academy of Economic Studies in Bucharest. She works in both film and video production (as an editor, sound designer or director), and in the marketing field, being responsible for the strategy, the content or the online marketing for different festivals, agencies, or clients. Laura also published three books, scientific articles and market research articles concerning innovative educational methods and the modernization of the teaching process, which she presented at different international conferences. Currently she is an Associate Professor PhD, teaching at the Sound Department of the Film Faculty of UNATC.

References

Adobe (2016) #VoCo. Adobe Audio Manipulator Sneak Peak with Jordan Peele [Online]. Available at: https://www.youtube.com/watch?v=I3l4XLZ59iw/ (Accessed: 20 May 2024).

Adobe (2023a) Behind the Tech: Enhance Speech in Adobe Podcast [Online]. Available at: https://research.adobe.com/news/behind-the-tech-enhance-speech-in-adobe-podcast/ (Accessed: 20 May, 2024).

Adobe (2023b) #ProjectDubDubDub | Adobe MAX Sneaks 2023 [Online]. Available at: https://www.youtube.com/watch?v=fZY-Cv1Q8NY/ (Accessed: 20 May 2024).

Agostinelli, A. et al (2023) ‘MusicLM: Generating Music From Text’, arXiv:2301.11325v1 [cs.SD] [Online]. Available at: https://arxiv.org/pdf/2301.11325 (Accessed: 20 May 2024).

Aicrowd (2024) Crowdsourcing AI to Solve Real-World Problems [Online]. https://www.aicrowd.com (Accessed: 20 May 2024).

Andersen, A. (2024a) Huge sound design contest! 7 chances to win fantastic prizes from Mattia Cellotto [Online]. Available at: https://www.asoundeffect.com/sounddesign24/ (Accessed: 20 May 2024).

Andersen, A. (2024b) Here are the winners of the Mattia Cellotto sound design contest! [Online]. Available at: https://www.asoundeffect.com/sounddesign-contest-results/ (Accessed: 20 May 2024).

Arhiva Activă UNATC (2024) Oamenii nu sunt capre [Online]. Available at: https://arhiva.unatc.ro/filme/oamenii-nu-sunt-capre/ (Accessed: 20 May 2024).

ASoundEffect (2024) Sound Design Contest: Create the sound for this video for 7 chances to win wild prizes! [Online]. Available at: https://www.youtube.com/watch?v=X_wsXpUxOaY (Accessed: 20 May 2024).

Collins, E. and Eck, D. (2024) New generative media models and tools, built with and for creators [Online]. Available at: https://blog.google/technology/ai/google-generative-ai-veo-imagen-3/ (Accessed: 20 May, 2024).

Computer Music Journal (1997) ‘Products of Interest’, Computer Music Journal, vol. 21, no. 3, p. 116. [Online]. Available at: http://www.jstor.org/stable/3681029 (Accessed: 20 May 2024).

Défossez, A. et al (2019) Music Source Separation in the Waveform Domain’, arXiv:1911.13254v1 [cs.SD] [Online]. Available at: https://arxiv.org/pdf/1911.13254v1 (Accessed: 20 May 2024).

Descript (2020) Introducing Descript [Online]. Available at: https://www.youtube.com/watch?v=Bl9wqNe5J8U/ (Accessed: 20 May 2024).

Devideconcept (2024) Audio AI Comparisons [Online]. Available at: https://divideconcept.github.io (Accessed: 20 May 2024).

ElevenLabs (2024) Introducing Voice Actor Payouts [Online]. Available at: https://elevenlabs.io/blog/introducing-voice-actor-payouts/ (Accessed: 20 May 2024).

Grant, N. (2024) Google Chatbot’s A.I. Images Put People of Color in Nazi-Era Uniforms [Online]. Available at: https://www.nytimes.com/2024/02/22/technology/google-gemini-german-uniforms.html (Accessed: 20 May 2024).

Hiatt, Brian (2024) A ChatGPT for Music Is Here. Inside Suno, the Startup Changing Everything [Online]. Available at: https://www.rollingstone.com/music/music-features/suno-ai-chatgpt-for-music-1234982307/ (Accessed: 20 May 2024).

Hsu et al (2023) Audiobox: Unified Audio Generation with Natural Language Prompts [Online]. Available at: https://ai.meta.com/research/publications/audiobox-unified-audio-generation-with-natural-language-prompts/ (Accessed: 20 May 2024).

The Beatles (2020) The Beatles: Get Back - A Sneak Peek from Peter Jackson [Online]. Available at: https://www.youtube.com/watch?v=UocEGvQ10OE (Accessed: 20 May 2024).

Hurwitz, M. (2021) The Making of The Beatles’ Let It Be and Peter Jackson’s Get Back Peter Jackson’s The Beatles: Get Back [Online]. Available at: https://www.soundandvision.com/content/making-beatles-let-it-be-and-peter-jacksons-get-back-peter-jacksons-beatles-get-back/ (Accessed: 20 May 2024).

Lehrman, P. D. (1998) ‘Prosoniq Sonicworx Artist. Audio Editing Software.’ Sound on Sound, vol. 13, issue 5, March. [Online]. Available at: https://www.soundonsound.com/reviews/prosoniq-sonicworx-artist/ (Accessed 20 May 2024).

Lin, David Chuan-En et al (2021) Soundify: Matching Sound Effects to Video, arXiv:2112.09726v1 [cs.SD] [Online]. Available at: https://arxiv.org/pdf/2112.09726v1 (Accessed: 20 May 2024).

Lin, David Chuan-En et al (2021) Soundify: Matching Sound Effects to Video, arXiv:2112.09726v3 [cs.SD] [Online]. Available at: https://arxiv.org/pdf/2112.09726v3 (Accessed: 20 May 2024).

Meta (2023a) Introducing Voicebox: The Most Versatile AI for Speech Generation [Online]. Available at: https://about.fb.com/news/2023/06/introducing-voicebox-ai-for-speech-generation/ (Accessed: 20 May 2024).

Meta (2023b) Open sourcing AudioCraft: Generative AI for audio made simple and available to all [Online]. Available at: https://ai.meta.com/blog/audiocraft-musicgen-audiogen-encodec-generative-ai-audio/ (Accessed: 20 May 2024).

Meta (2023c) Introducing Voicebox: The Most Versatile AI for Speech Generation [Online]. Available at: https://about.fb.com/news/2023/06/introducing-voicebox-ai-for-speech-generation/ (Accessed: 20 May 2024).

Microsoft Copilot (2023) Turn your ideas into songs with Suno on Microsoft Copilot [Online]. Available at: https://www.microsoft.com/en-us/microsoft-copilot/blog/2023/12/19/turn-your-ideas-into-songs-with-suno-on-microsoft-copilot/ (Accessed: 20 May 2024).

Morrison, R. (2024) Pika Labs sound effects now available for all users — I tried it and here is how it sounds [Online]. Available at: https://www.tomsguide.com/ai/ai-image-video/pika-labs-sound-effects-now-available-for-all-users-i-tried-it-and-here-is-how-it-sounds/ (Accessed: 20 May 2024).

Moussallam, M. (2019) Releasing Spleeter: Deezer Research source separation engine [Online]. Available at: https://deezer.io/releasing-spleeter-deezer-r-d-source-separation-engine-2b88985e797e/ (Accessed: 20 May 2024).

Nuñez, Michael (2024) Former Google DeepMind researchers launch AI-powered music creation app Udio [Online]. Available at https://venturebeat.com/ai/former-google-deepmind-researchers-launch-ai-powered-music-creation-app-udio/ (Accessed: 20 May 2024).

OpenAI (2024a) Navigating the Challenges and Opportunities of Synthetic Voices [Online]. Available at: https://openai.com/index/navigating-the-challenges-and-opportunities-of-synthetic-voices/ (Accessed: 20 May 2024).

OpenAI (2024b) Video generation models as world simulators [Online]. Available at: https://openai.com/index/video-generation-models-as-world-simulators/ (Accessed: 20 May 2024).

Owens, A. et al (2016) Visually Indicated Sounds, arXiv:1512.08512v2 [cs.CV] [Online]. Available at: https://arxiv.org/pdf/1512.08512 (Accessed: 20 May, 2024).

Paul, L. and Millman, E. (2023) Viral Drake and The Weeknd AI Collaboration Pulled From Apple, Spotify [Online]. Available at: https://www.rollingstone.com/music/music-news/viral-drake-and-the-weeknd-collaboration-is-completely-ai-generated-1234716154/ (Accessed: 20 May 2024).

Prevost, John J. and Ghose, Sanchita (2020) AutoFoley: Artificial Synthesis of Synchronized Sound Tracks for Silent Videos with Deep Learning, arXiv:2002.10981v1 [cs.SD] [Online]. Available at: https://arxiv.org/pdf/2002.10981/ (Accessed: 20 May 2024).

Prosoniq (2010) Prosoniq get the horn. Free World Cup audio plug-in released [Online]. Available at: https://www.soundonsound.com/news/prosoniq-get-horn/ (Accessed: 20 May 2024).

Reid, G. (2003) ‘Hartmann Neuron. Neuronal Resynthesizing Keyboard’ [Online]. Available at: https://www.soundonsound.com/reviews/hartmann-neuron/ (Accessed: 20 May 2024)

Rose, Jey (2018) Neural Networks: A new way to think about processing, CAS Quarterly [Online]. Available online at: https://digital.copcomm.com/i/933783-winter-2018/29? (Accessed: 20 May 2024).

Rucăreanu, D. Ș. (2024a) AI Restored example - student short film from UNATC archives [Online]. Available at: https://www.youtube.com/watch?v=5wc5HyzWbho (Accessed: 20 May 2024).

Rucăreanu, D. Ș. (2024b) AI English Dubbing Example - student short film from UNATC archives [Online]. Available at: https://www.youtube.com/watch?v=sKjcw1ADxlQ (Accessed: 20 May 2024).

Rucăreanu, D. Ș. (2024c) Audio Example 1 - generated with Meta Audiobox (Beta) [Online]. Available at: https://www.youtube.com/watch?v=lpas-3FkCKU (Accessed: 20 May 2024).

Rucăreanu, D. Ș. (2024d) Audio Example 2 - Generated with ElevenLabs Sound Effect (Beta) [Online]. Available at: https://www.youtube.com/watch?v=lmo099JKVhc (Accessed: 20 May 2024).

Rucăreanu, D. Ș. (2024e). AI Sound Design Example - audio elements generated with ElevenLabs Sound effects (Beta) [Online]. Available at: https://www.youtube.com/watch?v=teJ0g2pkmtI (Accessed: 20 May 2024).

Staniszewski, M. (2024) Introducing Dubbing Studio. Localize videos with precise control over transcript, translation, timing, and more [Online]. Available at: https://elevenlabs.io/blog/introducing-dubbing-studio/ (Accessed: 20 May 2024).

Thomas, B. (2017) Audionamix ADX Trax Pro 3 SP. Speech & Melody Separation Software [Online]. Available online at: https://www.soundonsound.com/reviews/audionamix-adx-trax-pro-3-sp/ (Accessed: 20 May 2024).

van den Oord, A. and Dieleman, S. (2016) WaveNet: A generative model for raw audio [Online]. Available at: https://deepmind.google/discover/blog/wavenet-a-generative-model-for-raw-audio/ (Accessed: 20 May 2024).

Wichern, G. (2017) What the Machine Learning in RX 6 Advanced Means for the Future of Audio Repair Technology [Online]. Available at: https://www.izotope.com/en/learn/what-the-machine-learning-in-rx-6-advanced-means-for-the-future-of-audio-repair-technology.html (Accessed: 20 May 2024).

Zhou, Y. et al (2018) Visual to Sound: Generating Natural Sound for Videos in the Wild, arXiv:1712.01393v2 [cs.CV] [Online]. Available at: https://arxiv.org/pdf/1712.01393/ (Accessed: 20 May 2024).

Zynaptiq (2013) Zyn1aptiq Acquires Prosoniq Product Line And Technologies [Online]. Available at: https://www.zynaptiq.com/more/zynaptiq-acquires-prosoniq-product-line-and-technologies/d106b1606afa940345ffb4a939c4588c/ (Accessed: 20 May 2024).

People are not Goats (1954) [DVD] Directed by Dragos Witkowski [Film]. București: Arhiva UNATC.

Downloads

Published

2024-12-18

How to Cite

Rucăreanu, D.- Ștefan, & Lăzărescu-Thois, L. (2024). The Impact of AI in the Field of Sound for Picture. A Historical, Practical, and Ethical Consideration. CONCEPT, 29(2), 112-128. https://doi.org/10.37130/bf780v59