This AI Makes "Audio Deepfakes"!
Key Takeaways
The video discusses advanced AI techniques for creating audio deepfakes, including voice cloning and neural voice puppetry, with tools like Weights & Biases and techniques like Techo Tron 2 and Neural Voice Puppetry.
Full Transcript
your fellow scholars this is two minute papers with this man's name that is impossible to pronounce my name is dr. Khurana if I hear and indeed it seems that pronouncing my name requires some advanced technology so what was this I promise to tell you in a moment but to understand what happened here first let's have a look at this deep fake technique we showcased a few videos ago as you see we are at a point where our mouth head and eye movements are also realistically translated to a chosen target subject and perhaps the most remarkable part of this work was that we don't even need a video of this target person just one photograph however these deep fake techniques mainly help us in transferring video content so what about voice synthesis is it also as advanced as this technique we are looking at well let's have a look at an example and you can decide for yourself this is a recent work that goes by the name techo Tron 2 and it performs aei based voice cloning all this technique requires is a 5 second sound sample of us and is able to synthesize new sentences in our voice as if we utter these words ourselves let's listen to a couple examples the Norseman considered the rainbow is a bridge over which the gods passed from Earth to their home in the sky take a look at these pages 4 Crooked Creek Drive there are several listings for gas station here's the forecast for the next four days Wow these are truly incredible the timbre of the voice is very similar and it is able to synthesize sounds and consonants that have to be inferred because they were not heard in the original voice sample and now let's jump to the next level and use a new technique that takes a sound sample and animates the video footage as if the target subject said it themselves this technique is called neural voice puppetry and even though the voices here are synthesized by this previous taiko Tron 2 method that you heard a moment ago we shouldn't judge this technique by its audio quality but how well the video follows these given sounds let's go the president of the United States is the head of state and head of government of the United States indirectly elected to a four-year term by the people through the Electoral College the office holder leads the executive branch of the federal government and is the commander-in-chief of the United States Armed Forces there are currently 4 living former presidents if you decide to stay until the end of this video there will be another fun video sample waiting for you there now note that this is not the first technique to achieve results like this so I can't wait to look under the hood and see what's new here after processing the incoming audio the gestures are applied to an intermediate 3d model which is specific to each person since each speaker has their own way of expressing themselves you can see this intermediate 3d model here but we are not done yet we feed it through a neural renderer and what this does is apply this motion to the particular face model shown in the video you can imagine the intermediate 3d model as a crude mask that models the gesture as well but does not look like the face of any one where the neural render adapts the mask to our target subject this includes adapting it to the current resolution lighting face position and more all of which is specific to what is seen in the video what is even cooler is that this new rendering part runs in real-time so what do we get from all this well one superior quality but at the same time it also generalizes to multiple targets have a look here you know I think we're in a moment of history we're probably the most important thing we need to do is to bring the country together and one of the skills that I bring to bear and the list of great news is not over yet you can try it yourself the link is available in the video description make sure to leave a comment with your results to sum up by combining multiple existing techniques it is important that everyone knows about the fact that we can both perform joint video and audio synthesis for a target subject this episode has been supported by weights and biases here they show you how to use their tool to perform phase swapping and improve your model that performs it also weights and biases provides tools to track your experiments in your deep learning projects their system is designed to save you a ton of time and money and it is actively used in projects at prestigious labs such as open AI Toyota research github and more and the best part is that if you are an academic or have an open-source project you can use their tools for free it really is as good as it gets make sure to visit them through w and be calm slash papers or just click the link in the video description and you can get a free demo today our thanks to weights and biases for their long-standing support and for helping us make better videos for you thanks for watching and for your generous support and I will see you next time
Original Description
❤️ Check out Weights & Biases and sign up for a free demo here: https://www.wandb.com/papers
Their blog post on #deepfakes is available here:
https://www.wandb.com/articles/improving-deepfake-performance-with-data
📝 The paper "Neural Voice Puppetry: Audio-driven Facial Reenactment" and its online demo are available here:
Paper: https://justusthies.github.io/posts/neural-voice-puppetry/
Demo - **Update: seems to have been disabled in the meantime, apologies!** : http://kaldir.vc.in.tum.de:9000/
❤️ Watch these videos in early access on our Patreon page or join us here on YouTube:
- https://www.patreon.com/TwoMinutePapers
- https://www.youtube.com/channel/UCbfYPyITQ-7l4upoX8nvctg/join
🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
Alex Haro, Alex Paden, Andrew Melnychuk, Angelos Evripiotis, Anthony Vdovitchenko, Benji Rabhan, Brian Gilman, Bryan Learn, Daniel Hasegan, Dennis Abts, Eric Haddad, Eric Martel, Evan Breznyik, Geronimo Moralez, James Watt, Javier Bustamante, Kaiesh Vohra, Kasia Hayden, Kjartan Olason, Levente Szabo, Lorin Atzberger, Lukas Biewald, Marcin Dukaczewski, Marten Rauschenberg, Maurits van Mastrigt, Michael Albrecht, Michael Jensen, Nader Shakerin, Owen Campbell-Moore, Owen Skarpness, Raul Araújo da Silva, Rob Rowe, Robin Graham, Ryan Monsurate, Shawn Azman, Steef, Steve Messina, Sunil Kim, Taras Bobrovytsky, Thomas Krcmar, Torsten Reil, Tybie Fitzhugh.
https://www.patreon.com/TwoMinutePapers
Meet and discuss your ideas with other Fellow Scholars on the Two Minute Papers Discord: https://discordapp.com/invite/hbcTJu2
Károly Zsolnai-Fehér's links:
Instagram: https://www.instagram.com/twominutepapers/
Twitter: https://twitter.com/karoly_zsolnai
Web: https://cg.tuwien.ac.at/~zsolnai/
#audiodeepfake #voicedeepfake #deepfake
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Two Minute Papers · Two Minute Papers · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Fluid Simulations with Blender and Wavelet Turbulence | Two Minute Papers #1
Two Minute Papers
Capturing Waves of Light With Femto-photography | Two Minute Papers #2
Two Minute Papers
Artificial Neural Networks and Deep Learning | Two Minute Papers #3
Two Minute Papers
Blender Rendering - Top 7 LuxRender Features
Two Minute Papers
Simulating Breaking Glass | Two Minute Papers #4
Two Minute Papers
Time Lapse Videos From Community Photos | Two Minute Papers #5
Two Minute Papers
AI Learns Van Gogh's Art
Two Minute Papers
Hydrographic Printing | Two Minute Papers #7
Two Minute Papers
Announcing LuxRender 1.5
Two Minute Papers
Digital Creatures Learn To Walk | Two Minute Papers #8
Two Minute Papers
Manipulating Photorealistic Renderings | Two Minute Papers #9
Two Minute Papers
Adaptive Fluid Simulations | Two Minute Papers #10
Two Minute Papers
Building Bridges With Flying Machines | Two Minute Papers #11
Two Minute Papers
Reconstructing Sound From Vibrations | Two Minute Papers #12
Two Minute Papers
Creating Photographs Using Deep Learning | Two Minute Papers #13
Two Minute Papers
Adaptive Cloth Simulations | Two Minute Papers #14
Two Minute Papers
Synthesizing Sound From Collisions | Two Minute Papers #15
Two Minute Papers
Metropolis Light Transport | Two Minute Papers #16
Two Minute Papers
3D Printing a Glockenspiel | Two Minute Papers #17
Two Minute Papers
Modeling Colliding and Merging Fluids | Two Minute Papers #18
Two Minute Papers
Recurrent Neural Network Writes Music and Shakespeare Novels | Two Minute Papers #19
Two Minute Papers
Gradients, Poisson's Equation and Light Transport | Two Minute Papers #20
Two Minute Papers
Real-Time Facial Expression Transfer | Two Minute Papers #21
Two Minute Papers
Automatic Lecture Notes From Videos | Two Minute Papers #22
Two Minute Papers
Be a Part of Two Minute Papers on Patreon!
Two Minute Papers
Recurrent Neural Network Writes Sentences About Images | Two Minute Papers #23
Two Minute Papers
How Does Deep Learning Work? | Two Minute Papers #24
Two Minute Papers
Cryptography, Perfect Secrecy and One Time Pads | Two Minute Papers #25
Two Minute Papers
Terrain Traversal with Reinforcement Learning | Two Minute Papers #26
Two Minute Papers
Multiple-Scattering Microfacet BSDFs with the Smith Model
Two Minute Papers
Google DeepMind's Deep Q-Learning & Superhuman Atari Gameplays | Two Minute Papers #27
Two Minute Papers
Are We Living In a Computer Simulation? | Two Minute Papers #28
Two Minute Papers
Artificial Superintelligence [Audio only] | Two Minute Papers #29
Two Minute Papers
Automatic Parameter Control for Metropolis Light Transport | Two Minute Papers #30
Two Minute Papers
Randomness and Bell's Inequality [Audio only] | Two Minute Papers #31
Two Minute Papers
OpenAI - Non-profit AI company by Elon Musk and Sam Altman
Two Minute Papers
How Do Genetic Algorithms Work? | Two Minute Papers #32
Two Minute Papers
Painting with Fluid Simulations | Two Minute Papers #33
Two Minute Papers
Peer Review #1 [Audio only] | Two Minute Papers
Two Minute Papers
Neural Programmer-Interpreters Learn To Write Programs | Two Minute Papers #34
Two Minute Papers
9 Cool Deep Learning Applications | Two Minute Papers #35
Two Minute Papers
Designing Cities and Furnitures With Machine Learning | Two Minute Papers #36
Two Minute Papers
Designing 3D Printable Robotic Creatures | Two Minute Papers #37
Two Minute Papers
3D Printing Objects With Caustics | Two Minute Papers #38
Two Minute Papers
Interactive Editing of Subsurface Scattering | Two Minute Papers #39
Two Minute Papers
Simulating Viscosity and Melting Fluids | Two Minute Papers #40
Two Minute Papers
What Do Virtual Objects Sound Like? | Two Minute Papers #41
Two Minute Papers
How DeepMind Conquered Go With Deep Learning (AlphaGo) | Two Minute Papers #42
Two Minute Papers
Breaking Deep Learning Systems With Adversarial Examples | Two Minute Papers #43
Two Minute Papers
Extrapolations and Crowdfunded Research (Experiment) | Two Minute Papers #44
Two Minute Papers
Biophysical Skin Aging Simulations | Two Minute Papers #45
Two Minute Papers
What is Impostor Syndrome? | Two Minute Papers #46
Two Minute Papers
Should You Take the Stairs at Work? (For Weight Loss) | Two Minute Papers #47
Two Minute Papers
Artistic Manipulation of Caustics | Two Minute Papers #48
Two Minute Papers
Deep Learning Program Learns to Paint | Two Minute Papers #49
Two Minute Papers
Interactive Photo Recoloring | Two Minute Papers #50
Two Minute Papers
How To Get Started With Machine Learning? | Two Minute Papers #51
Two Minute Papers
Awesome Research For Everyone! - Two Minute Papers Channel Trailer
Two Minute Papers
10 More Cool Deep Learning Applications | Two Minute Papers #52
Two Minute Papers
How DeepMind's AlphaGo Defeated Lee Sedol | Two Minute Papers #53
Two Minute Papers
More on: AI Safety Engineering
View skill →Related Reads
📰
📰
📰
📰
OpenAI’s Hugging Face Breach Shows Frontier AI Guardrails Are Failing
Forbes Innovation
NSF invierte $83 millones en un backbone de datos para IA científica
Dev.to · lu1tr0n
Nobody Actually Knows What’s Going On, and That Includes the People Who Built It
Medium · AI
AI image fraud will cost $40 billion next year - can these international standards help?
ZDNet
🎓
Tutor Explanation
DeepCamp AI