This AI Makes "Audio Deepfakes"!

Two Minute Papers · Advanced ·🛡️ AI Safety & Ethics ·6y ago

Key Takeaways

The video discusses advanced AI techniques for creating audio deepfakes, including voice cloning and neural voice puppetry, with tools like Weights & Biases and techniques like Techo Tron 2 and Neural Voice Puppetry.

Full Transcript

your fellow scholars this is two minute papers with this man's name that is impossible to pronounce my name is dr. Khurana if I hear and indeed it seems that pronouncing my name requires some advanced technology so what was this I promise to tell you in a moment but to understand what happened here first let's have a look at this deep fake technique we showcased a few videos ago as you see we are at a point where our mouth head and eye movements are also realistically translated to a chosen target subject and perhaps the most remarkable part of this work was that we don't even need a video of this target person just one photograph however these deep fake techniques mainly help us in transferring video content so what about voice synthesis is it also as advanced as this technique we are looking at well let's have a look at an example and you can decide for yourself this is a recent work that goes by the name techo Tron 2 and it performs aei based voice cloning all this technique requires is a 5 second sound sample of us and is able to synthesize new sentences in our voice as if we utter these words ourselves let's listen to a couple examples the Norseman considered the rainbow is a bridge over which the gods passed from Earth to their home in the sky take a look at these pages 4 Crooked Creek Drive there are several listings for gas station here's the forecast for the next four days Wow these are truly incredible the timbre of the voice is very similar and it is able to synthesize sounds and consonants that have to be inferred because they were not heard in the original voice sample and now let's jump to the next level and use a new technique that takes a sound sample and animates the video footage as if the target subject said it themselves this technique is called neural voice puppetry and even though the voices here are synthesized by this previous taiko Tron 2 method that you heard a moment ago we shouldn't judge this technique by its audio quality but how well the video follows these given sounds let's go the president of the United States is the head of state and head of government of the United States indirectly elected to a four-year term by the people through the Electoral College the office holder leads the executive branch of the federal government and is the commander-in-chief of the United States Armed Forces there are currently 4 living former presidents if you decide to stay until the end of this video there will be another fun video sample waiting for you there now note that this is not the first technique to achieve results like this so I can't wait to look under the hood and see what's new here after processing the incoming audio the gestures are applied to an intermediate 3d model which is specific to each person since each speaker has their own way of expressing themselves you can see this intermediate 3d model here but we are not done yet we feed it through a neural renderer and what this does is apply this motion to the particular face model shown in the video you can imagine the intermediate 3d model as a crude mask that models the gesture as well but does not look like the face of any one where the neural render adapts the mask to our target subject this includes adapting it to the current resolution lighting face position and more all of which is specific to what is seen in the video what is even cooler is that this new rendering part runs in real-time so what do we get from all this well one superior quality but at the same time it also generalizes to multiple targets have a look here you know I think we're in a moment of history we're probably the most important thing we need to do is to bring the country together and one of the skills that I bring to bear and the list of great news is not over yet you can try it yourself the link is available in the video description make sure to leave a comment with your results to sum up by combining multiple existing techniques it is important that everyone knows about the fact that we can both perform joint video and audio synthesis for a target subject this episode has been supported by weights and biases here they show you how to use their tool to perform phase swapping and improve your model that performs it also weights and biases provides tools to track your experiments in your deep learning projects their system is designed to save you a ton of time and money and it is actively used in projects at prestigious labs such as open AI Toyota research github and more and the best part is that if you are an academic or have an open-source project you can use their tools for free it really is as good as it gets make sure to visit them through w and be calm slash papers or just click the link in the video description and you can get a free demo today our thanks to weights and biases for their long-standing support and for helping us make better videos for you thanks for watching and for your generous support and I will see you next time

Original Description

❤️ Check out Weights & Biases and sign up for a free demo here: https://www.wandb.com/papers Their blog post on #deepfakes is available here: https://www.wandb.com/articles/improving-deepfake-performance-with-data 📝 The paper "Neural Voice Puppetry: Audio-driven Facial Reenactment" and its online demo are available here: Paper: https://justusthies.github.io/posts/neural-voice-puppetry/ Demo - **Update: seems to have been disabled in the meantime, apologies!** : http://kaldir.vc.in.tum.de:9000/ ❤️ Watch these videos in early access on our Patreon page or join us here on YouTube: - https://www.patreon.com/TwoMinutePapers - https://www.youtube.com/channel/UCbfYPyITQ-7l4upoX8nvctg/join 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible: Alex Haro, Alex Paden, Andrew Melnychuk, Angelos Evripiotis, Anthony Vdovitchenko, Benji Rabhan, Brian Gilman, Bryan Learn, Daniel Hasegan, Dennis Abts, Eric Haddad, Eric Martel, Evan Breznyik, Geronimo Moralez, James Watt, Javier Bustamante, Kaiesh Vohra, Kasia Hayden, Kjartan Olason, Levente Szabo, Lorin Atzberger, Lukas Biewald, Marcin Dukaczewski, Marten Rauschenberg, Maurits van Mastrigt, Michael Albrecht, Michael Jensen, Nader Shakerin, Owen Campbell-Moore, Owen Skarpness, Raul Araújo da Silva, Rob Rowe, Robin Graham, Ryan Monsurate, Shawn Azman, Steef, Steve Messina, Sunil Kim, Taras Bobrovytsky, Thomas Krcmar, Torsten Reil, Tybie Fitzhugh. https://www.patreon.com/TwoMinutePapers Meet and discuss your ideas with other Fellow Scholars on the Two Minute Papers Discord: https://discordapp.com/invite/hbcTJu2 Károly Zsolnai-Fehér's links: Instagram: https://www.instagram.com/twominutepapers/ Twitter: https://twitter.com/karoly_zsolnai Web: https://cg.tuwien.ac.at/~zsolnai/ #audiodeepfake #voicedeepfake #deepfake
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Two Minute Papers · Two Minute Papers · 0 of 60

← Previous Next →
1 Fluid Simulations with Blender and Wavelet Turbulence | Two Minute Papers #1
Fluid Simulations with Blender and Wavelet Turbulence | Two Minute Papers #1
Two Minute Papers
2 Capturing Waves of Light With Femto-photography | Two Minute Papers #2
Capturing Waves of Light With Femto-photography | Two Minute Papers #2
Two Minute Papers
3 Artificial Neural Networks and Deep Learning | Two Minute Papers #3
Artificial Neural Networks and Deep Learning | Two Minute Papers #3
Two Minute Papers
4 Blender Rendering - Top 7 LuxRender Features
Blender Rendering - Top 7 LuxRender Features
Two Minute Papers
5 Simulating Breaking Glass | Two Minute Papers #4
Simulating Breaking Glass | Two Minute Papers #4
Two Minute Papers
6 Time Lapse Videos From Community Photos | Two Minute Papers #5
Time Lapse Videos From Community Photos | Two Minute Papers #5
Two Minute Papers
7 AI Learns Van Gogh's Art
AI Learns Van Gogh's Art
Two Minute Papers
8 Hydrographic Printing | Two Minute Papers #7
Hydrographic Printing | Two Minute Papers #7
Two Minute Papers
9 Announcing LuxRender 1.5
Announcing LuxRender 1.5
Two Minute Papers
10 Digital Creatures Learn To Walk | Two Minute Papers #8
Digital Creatures Learn To Walk | Two Minute Papers #8
Two Minute Papers
11 Manipulating Photorealistic Renderings | Two Minute Papers #9
Manipulating Photorealistic Renderings | Two Minute Papers #9
Two Minute Papers
12 Adaptive Fluid Simulations | Two Minute Papers #10
Adaptive Fluid Simulations | Two Minute Papers #10
Two Minute Papers
13 Building Bridges With Flying Machines | Two Minute Papers #11
Building Bridges With Flying Machines | Two Minute Papers #11
Two Minute Papers
14 Reconstructing Sound From Vibrations | Two Minute Papers #12
Reconstructing Sound From Vibrations | Two Minute Papers #12
Two Minute Papers
15 Creating Photographs Using Deep Learning | Two Minute Papers #13
Creating Photographs Using Deep Learning | Two Minute Papers #13
Two Minute Papers
16 Adaptive Cloth Simulations | Two Minute Papers #14
Adaptive Cloth Simulations | Two Minute Papers #14
Two Minute Papers
17 Synthesizing Sound From Collisions | Two Minute Papers #15
Synthesizing Sound From Collisions | Two Minute Papers #15
Two Minute Papers
18 Metropolis Light Transport | Two Minute Papers #16
Metropolis Light Transport | Two Minute Papers #16
Two Minute Papers
19 3D Printing a Glockenspiel | Two Minute Papers #17
3D Printing a Glockenspiel | Two Minute Papers #17
Two Minute Papers
20 Modeling Colliding and Merging Fluids | Two Minute Papers #18
Modeling Colliding and Merging Fluids | Two Minute Papers #18
Two Minute Papers
21 Recurrent Neural Network Writes Music and Shakespeare Novels | Two Minute Papers #19
Recurrent Neural Network Writes Music and Shakespeare Novels | Two Minute Papers #19
Two Minute Papers
22 Gradients, Poisson's Equation and Light Transport | Two Minute Papers #20
Gradients, Poisson's Equation and Light Transport | Two Minute Papers #20
Two Minute Papers
23 Real-Time Facial Expression Transfer | Two Minute Papers #21
Real-Time Facial Expression Transfer | Two Minute Papers #21
Two Minute Papers
24 Automatic Lecture Notes From Videos | Two Minute Papers #22
Automatic Lecture Notes From Videos | Two Minute Papers #22
Two Minute Papers
25 Be a Part of Two Minute Papers on Patreon!
Be a Part of Two Minute Papers on Patreon!
Two Minute Papers
26 Recurrent Neural Network Writes Sentences About Images | Two Minute Papers #23
Recurrent Neural Network Writes Sentences About Images | Two Minute Papers #23
Two Minute Papers
27 How Does Deep Learning Work? | Two Minute Papers #24
How Does Deep Learning Work? | Two Minute Papers #24
Two Minute Papers
28 Cryptography, Perfect Secrecy and One Time Pads | Two Minute Papers #25
Cryptography, Perfect Secrecy and One Time Pads | Two Minute Papers #25
Two Minute Papers
29 Terrain Traversal with Reinforcement Learning | Two Minute Papers #26
Terrain Traversal with Reinforcement Learning | Two Minute Papers #26
Two Minute Papers
30 Multiple-Scattering Microfacet BSDFs with the Smith Model
Multiple-Scattering Microfacet BSDFs with the Smith Model
Two Minute Papers
31 Google DeepMind's Deep Q-Learning & Superhuman Atari Gameplays | Two Minute Papers #27
Google DeepMind's Deep Q-Learning & Superhuman Atari Gameplays | Two Minute Papers #27
Two Minute Papers
32 Are We Living In a Computer Simulation? | Two Minute Papers #28
Are We Living In a Computer Simulation? | Two Minute Papers #28
Two Minute Papers
33 Artificial Superintelligence [Audio only] | Two Minute Papers #29
Artificial Superintelligence [Audio only] | Two Minute Papers #29
Two Minute Papers
34 Automatic Parameter Control for Metropolis Light Transport | Two Minute Papers #30
Automatic Parameter Control for Metropolis Light Transport | Two Minute Papers #30
Two Minute Papers
35 Randomness and Bell's Inequality [Audio only] | Two Minute Papers #31
Randomness and Bell's Inequality [Audio only] | Two Minute Papers #31
Two Minute Papers
36 OpenAI - Non-profit AI company by Elon Musk and Sam Altman
OpenAI - Non-profit AI company by Elon Musk and Sam Altman
Two Minute Papers
37 How Do Genetic Algorithms Work? | Two Minute Papers #32
How Do Genetic Algorithms Work? | Two Minute Papers #32
Two Minute Papers
38 Painting with Fluid Simulations | Two Minute Papers #33
Painting with Fluid Simulations | Two Minute Papers #33
Two Minute Papers
39 Peer Review #1 [Audio only] | Two Minute Papers
Peer Review #1 [Audio only] | Two Minute Papers
Two Minute Papers
40 Neural Programmer-Interpreters Learn To Write Programs | Two Minute Papers #34
Neural Programmer-Interpreters Learn To Write Programs | Two Minute Papers #34
Two Minute Papers
41 9 Cool Deep Learning Applications | Two Minute Papers #35
9 Cool Deep Learning Applications | Two Minute Papers #35
Two Minute Papers
42 Designing Cities and Furnitures With Machine Learning | Two Minute Papers #36
Designing Cities and Furnitures With Machine Learning | Two Minute Papers #36
Two Minute Papers
43 Designing 3D Printable Robotic Creatures | Two Minute Papers #37
Designing 3D Printable Robotic Creatures | Two Minute Papers #37
Two Minute Papers
44 3D Printing Objects With Caustics | Two Minute Papers #38
3D Printing Objects With Caustics | Two Minute Papers #38
Two Minute Papers
45 Interactive Editing of Subsurface Scattering | Two Minute Papers #39
Interactive Editing of Subsurface Scattering | Two Minute Papers #39
Two Minute Papers
46 Simulating Viscosity and Melting Fluids | Two Minute Papers #40
Simulating Viscosity and Melting Fluids | Two Minute Papers #40
Two Minute Papers
47 What Do Virtual Objects Sound Like? | Two Minute Papers #41
What Do Virtual Objects Sound Like? | Two Minute Papers #41
Two Minute Papers
48 How DeepMind Conquered Go With Deep Learning (AlphaGo) | Two Minute Papers #42
How DeepMind Conquered Go With Deep Learning (AlphaGo) | Two Minute Papers #42
Two Minute Papers
49 Breaking Deep Learning Systems With Adversarial Examples | Two Minute Papers #43
Breaking Deep Learning Systems With Adversarial Examples | Two Minute Papers #43
Two Minute Papers
50 Extrapolations and Crowdfunded Research (Experiment) | Two Minute Papers #44
Extrapolations and Crowdfunded Research (Experiment) | Two Minute Papers #44
Two Minute Papers
51 Biophysical Skin Aging Simulations | Two Minute Papers #45
Biophysical Skin Aging Simulations | Two Minute Papers #45
Two Minute Papers
52 What is Impostor Syndrome? | Two Minute Papers #46
What is Impostor Syndrome? | Two Minute Papers #46
Two Minute Papers
53 Should You Take the Stairs at Work? (For Weight Loss) | Two Minute Papers #47
Should You Take the Stairs at Work? (For Weight Loss) | Two Minute Papers #47
Two Minute Papers
54 Artistic Manipulation of Caustics | Two Minute Papers #48
Artistic Manipulation of Caustics | Two Minute Papers #48
Two Minute Papers
55 Deep Learning Program Learns to Paint | Two Minute Papers #49
Deep Learning Program Learns to Paint | Two Minute Papers #49
Two Minute Papers
56 Interactive Photo Recoloring | Two Minute Papers #50
Interactive Photo Recoloring | Two Minute Papers #50
Two Minute Papers
57 How To Get Started With Machine Learning? | Two Minute Papers #51
How To Get Started With Machine Learning? | Two Minute Papers #51
Two Minute Papers
58 Awesome Research For Everyone! - Two Minute Papers Channel Trailer
Awesome Research For Everyone! - Two Minute Papers Channel Trailer
Two Minute Papers
59 10 More Cool Deep Learning Applications | Two Minute Papers #52
10 More Cool Deep Learning Applications | Two Minute Papers #52
Two Minute Papers
60 How DeepMind's AlphaGo Defeated Lee Sedol | Two Minute Papers #53
How DeepMind's AlphaGo Defeated Lee Sedol | Two Minute Papers #53
Two Minute Papers

The video showcases advanced AI techniques for creating audio deepfakes, including voice cloning and neural voice puppetry, and discusses the importance of AI safety and responsible use of these technologies.

Key Takeaways
  1. Understand the basics of audio deepfakes and voice cloning
  2. Explore the Techo Tron 2 technique for AI-based voice cloning
  3. Learn about Neural Voice Puppetry and its applications
  4. Use Weights & Biases for experiment tracking and model improvement
  5. Consider the AI safety implications of audio deepfakes
💡 The combination of voice cloning and neural voice puppetry can create highly realistic audio deepfakes, but also raises important AI safety concerns.

Related Reads

📰
OpenAI’s Hugging Face Breach Shows Frontier AI Guardrails Are Failing
OpenAI's Hugging Face breach exposes weaknesses in AI lab guardrails, highlighting risks of autonomous agents
Forbes Innovation
📰
NSF invierte $83 millones en un backbone de datos para IA científica
La NSF invierte $83 millones en un backbone de datos para IA científica para acelerar la investigación y el descubrimiento
Dev.to · lu1tr0n
📰
Nobody Actually Knows What’s Going On, and That Includes the People Who Built It
Understanding the limitations and uncertainties of AI development, even among its creators, is crucial for responsible innovation
Medium · AI
📰
AI image fraud will cost $40 billion next year - can these international standards help?
International standards may help combat AI image fraud, which is projected to cost $40 billion next year, by providing a unified approach to identification and mitigation.
ZDNet
Up next
Your AI Output Is Wrong and You Don't Know It Yet
Kevin Farugia AI Automation
Watch →