DAViD: Data-efficient and Accurate Vision Models from Synthetic Data
Key Takeaways
The video presents DAViD, a cost-effective and robust human-centric vision model trained solely on high-fidelity synthetic data, achieving highly accurate results on dense prediction tasks such as soft foreground segmentation, depth estimation, and surface normal prediction using a dense prediction transformer (DPT) architecture.
Full Transcript
In this work, we present highly accurate, cost-effective, and robust human- ccentric vision models trained solely on highfidelity synthetic data. Our approach provides accurate, robust human-centric dense prediction models trained on small yet highfidelity synthetic data with low training and inference costs. Particularly we show that we can achieve highly accurate results on human ccentric dense prediction tasks of soft foreground segmentation depth estimation and surface normal prediction. In the literature there are different approaches to train models that generalize to diverse scenarios. Big transfer leverages a massive data set of 300 million images with noisy labels for pre-training, providing a strong backbone for various classification tasks. For generalizable depth estimation, depth anything v2 leverages the giant version of Dino V2. First training on 600,000 synthetic images with perfect depth annotations, then generating pseudo ground truth for 62 million diverse images to train student models. Depth Pro in contrast first trains a custom DPT model on a mix of real and synthetic data sets to enhance fine details. It continues training on highresolution synthetic data with perfect annotations. In another work, Sapiens focuses on human centric vision tasks and pre-trains its backbone on 300 million real images of people using self-supervised learning. The pre-trained model then initializes weights for various downstream tasks such as depth and normal prediction fine-tuned with synthetic data. While all these approaches show promising accuracy and generalization, they typically require very large models and expensive training. In this work, we take a simpler approach of training on a diverse highfidelity synthetic data set that generalizes to real humanentric scenarios. We show that training on a small yet highfidelity data set is enough to achieve very competitive accuracy. Our models require only a fraction of the cost of training and inference when compared with large foundational models of similar accuracy. To train our models, we exclusively use synthetic data generated through an existing pipeline that creates realistic human-centric data sets along with precise ground truth annotations. In this work, we extend the use of this highquality synthetic data to dense prediction tasks where both realism and annotation quality are more critical and for which annotations on real data are often impossible. We train all three models for depth estimation, surface normal estimation, and soft foreground segmentation on a single training data set of 300,000 images. The images cover faces, upper body, and full body scenarios equally. We designed this data set such that it is diverse in terms of poses, environments, lighting, and appearances and not tailored to any specific evaluation set. Our models are built on top of dense prediction transformer or DPT. We first resize the input image and pass it to a VIT encoder for a fixed time encoder inference. Similar to a unit, we use a convolutional decoder with four blocks in which each block is connected to the encoder through skip connections. To enable inference on any input image resolution, we use a simple convolutional model called resizer. It injects image features into decoder blocks, allowing the decoder to reason about image features at input resolution. We use the same architecture for soft foreground segmentation, depth estimation, and surface normal prediction. Since we use the same architecture design and same training data set across all three tasks, it is feasible to extend the model into a multitask one by adding decoder heads. This is computationally advantageous over running three separate models when all three output modalities are needed. Our human centric dense prediction models deliver high quality results that reliably capture a wide range of human characteristics, handle diverse lighting conditions, and preserve fine details like hair strands and subtle facial features. This demonstrates our model's robustness and ability to handle complex real world scenarios with high accuracy. Here we compare our multitask dense prediction results with the existing soda models. Sapiens 2B. We show that not only our model produces highly accurate predictions with fine details, it runs orders of magnitude faster than competing methods. Our relative depth estimation model generates a 2.5D representation from a single image, enabling 3D point cloud creation and novel viewpoint rendering while preserving complex depth relationships and object proportions. Our soft foreground segmentation model enables precise subject extraction, preserving fine details for highquality background replacement in applications like video conferencing. Another application of our model is relighting from surface normals. By predicting accurate surface normal maps, we can rerender images under novel lighting conditions, creating shading and ambient effects that lead to a visually plausible relighting result.
Original Description
The state of the art in human-centric computer vision achieves high accuracy and robustness across a diverse range of tasks. The most effective models in this domain have billions of parameters, thus requiring extremely large datasets, expensive training regimes, and compute-intensive inference. In this paper, we demonstrate that it is possible to train models on much smaller but high-fidelity synthetic datasets, with no loss in accuracy and higher efficiency. Using synthetic training data provides us with excellent levels of detail and perfect labels, while providing strong guarantees for data provenance, usage rights, and user consent. Procedural data synthesis also provides us with explicit control on data diversity, that we can use to address unfairness in the models we train. Extensive quantitative assessment on real input images demonstrates accuracy of our models on three dense prediction tasks: depth estimation, surface normal estimation, and soft foreground segmentation. Our models require only a fraction of the cost of training and inference when compared with foundational models of similar accuracy.
Project page: https://aka.ms/DAViD
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Microsoft Research · Microsoft Research · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Frontiers in ML: Learning from Limited Labeled Data: Challenges and Opportunities for NLP
Microsoft Research
Frontiers in Machine Learning: Climate Impact of Machine Learning
Microsoft Research
Frontiers in Machine Learning: Security and Machine Learning
Microsoft Research
Hope Speech and Help Speech: Surfacing Positivity Amidst Hate
Microsoft Research
Early Indicators of the Effect of the Global Shift to Remote Work on People with Disabilities
Microsoft Research
Remote Work and Well-Being
Microsoft Research
Challenges and Gratitude of Software Developers During COVID-19 Working From Home
Microsoft Research
Towards a Practical Virtual Office for Mobile Knowledge Workers
Microsoft Research
Impact of COVID-19 crisis on the future of work in India
Microsoft Research
Empowering and Supporting Remote Software Development Team Members through a Culture of Allyship
Microsoft Research
How Work From Home Affects Collaboration: Information Workers in a Natural Experiment During COVID19
Microsoft Research
Phong Surface: Efficient 3D Model Fitting using Lifted Optimization
Microsoft Research
Managing Tasks Across the Work-Life Boundary: Opportunities, Challenges, and Directions
Microsoft Research
Microsoft Urban Futures Summer Workshop | Data Driven Urban Transformation [Day 1]
Microsoft Research
Microsoft Urban Futures Summer Workshop | Sensors and Data [Day 2]
Microsoft Research
Microsoft Urban Futures Summer Workshop | Policy and Social Impact [Day 3]
Microsoft Research
Directions in ML: Algorithmic foundations of neural architecture search
Microsoft Research
MineRL Competition 2020
Microsoft Research
Can we make better software by using ML and AI techniques? With Chandra Maddila and Chetan Bansal
Microsoft Research
From Paper to Product
Microsoft Research
SkinnerDB: Regret Bounded Query Evaluation using RL
Microsoft Research
From SqueezeNet to SqueezeBERT: Developing Efficient Deep Neural Networks
Microsoft Research
Programming with Proofs for High-assurance Software
Microsoft Research
Platform for Situated Intelligence Overview
Microsoft Research
Directional Sources & Listeners in Interactive Sound Propagation using Reciprocal Wave Field Coding
Microsoft Research
Galactic Bell Star Music Demo
Microsoft Research
Importing Animations in Microsoft Expressive Pixels (9 of 9)
Microsoft Research
Welcome to Microsoft Expressive Pixels (1 of 9)
Microsoft Research
Getting Started with Microsoft Expressive Pixels (2 of 9)
Microsoft Research
Creating an Image in Microsoft Expressive Pixels (3 of 9)
Microsoft Research
Creating Animations in Microsoft Expressive Pixels (4 of 9)
Microsoft Research
Managing Animation Galleries in Microsoft Expressive Pixels (5 of 9)
Microsoft Research
Creating Fragments in Microsoft Expressive Pixels (6 of 9)
Microsoft Research
Using Layers in Microsoft Expressive Pixels (7 of 9)
Microsoft Research
Exporting Animations with Microsoft Expressive Pixels (8 of 9)
Microsoft Research
What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 2/2)
Microsoft Research
What Kind of Computation is Human Cognition? A Brief History of Thought (Episode 1/2)
Microsoft Research
Planeverb: Interactive sound propagation for dynamic scenes using 2D wave simulation
Microsoft Research
Making cryptography accessible, efficient, and scalable with Dr. Divya Gupta and Dr. Rahul Sharma
Microsoft Research
Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 Talk)
Microsoft Research
Optics for the cloud – Light at the end of the tunnel? (SIGCOMM 2020 Workshop)
Microsoft Research
Beyond the mega-data center: networking multi-data center regions (SIGCOMM 2020 short talk)
Microsoft Research
Sirius: A Flat Datacenter Network with Nanosecond Optical Switching (SIGCOMM 2020 short talk)
Microsoft Research
Novel Image Captioning
Microsoft Research
Forest Sound Scene Simulation and Bird Localization with Distributed Microphone Arrays
Microsoft Research
Decoding Music Attention from “EEG headphones”: a User-friendly Auditory Brain-computer Interface
Microsoft Research
How does holographic storage work?
Microsoft Research
The physics of hologram formation in iron doped lithium niobate
Microsoft Research
Introduction to coax: A Modular RL Package
Microsoft Research
Directions in ML: "Neural architecture search: Coming of age"
Microsoft Research
Microsoft Research AI Breakthroughs 2020: 20 minute research talks + Q&A panel
Microsoft Research
Fireside Chat with Johannes Gehrke during Microsoft Research AI Breakthroughs 2020
Microsoft Research
Fireside Chat with Susan Dumais during Microsoft Research AI Breakthroughs 2020
Microsoft Research
Microsoft Research AI Breakthroughs 2020: 20 minute research talks, Q&A panel, and event wrap-up
Microsoft Research
Clinical Research with FHIR
Microsoft Research
Soundscape Street Preview
Microsoft Research
Tilt-Responsive Techniques for Digital Drawing Boards
Microsoft Research
SurfaceFleet: Exploring Distributed Interactions Unbounded from Device, Application, User, and Time
Microsoft Research
Haptic PIVOT: On-Demand Handhelds in VR
Microsoft Research
SurfaceFleet Supplemental Video Demonstration (UIST 2020)
Microsoft Research
More on: Modern CV Models
View skill →Related Reads
📰
📰
📰
📰
Building Stable Video Portrait Outlines in the Browser with MODNet, SlimSAM, MediaPipe, and Optical Flow
Dev.to AI
RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation
ArXiv cs.AI
Your Memory Of That Face Is Lying To You — Here's The Proof
Dev.to AI
How to Blur Sensitive Data in Images Locally in Your Browser Using Canvas
Dev.to · JSDev Space
🎓
Tutor Explanation
DeepCamp AI