Contrastive Learning for Tabular Data - SCARF

Connor Shorten · Beginner ·📰 AI News & Updates ·5y ago

Key Takeaways

The SCARF algorithm applies self-supervised contrastive learning to tabular data using random feature corruption, leveraging frameworks like Bootstrap Your Own Latent and empirical marginal distributions to form positive pairs and improve predictive performance.

Full Transcript

this video will provide a quick overview of the scarf algorithm bringing self-supervised contrastive learning to tabular data using random feature corruption so this is using frameworks like bootstrap your own layton where you have these two positive pairs one formed with the data augmentation view of the original example and bringing that kind of idea to tabular data so this is tabular data say it's information about a customer and a business things like name age maybe buying preferences and these different features that describe customer data these kind of type of data sets and we'll look through this catalog this openml cc18 is a data set of uh as a collection of tabular datasets if you're interested in just picking one of these to do academic research or getting a better sense of what these data sets look like compared to the usual deep learning data sets of vision language audio and so on even though the type of data are more common in the machine learning real world anyways but so this is a strategy to bring this kind of advancement of self-supervised pre-training contrastive learning into tabular data this quote communicates what i think is the key id of the algorithm which is how we're going to mask out and corrupt the data augmentation to form the positive pair to align the representations of something like a momentum encoding average like the bootstrap your own latent framework so describe we sample we sample some refraction of the features uniformly at random and replace each of those features by a random draw and then particularly this line from that features empirical marginal distribution so each of these features say it's the age of the customers or how many toilet paper rolls they bought if this is like a shopping basket analysis kind of problem you would have that empirical district marginal distribution marginal distribution of each of the random variables uh say you have x1 x2 x3 so and then they each have their own separate uh distributions as well as the joint distribution and so on so you're going to sample how you're going to mask that and corrupt the example from the marginal distribution of each of the features so if it's age and so on it would have a different distribution of how you sample a new age value and different densities like just say it's a different gaussian distribution with different mean variance parameters or whether it's some poisson distribution total other distribution but so that's how you're going to be sampling which features you're going to be using corrupting in order to form the data augmentation view which is the second positive pairing as you align these v v prime representations in the contrastive learning framework this is the openml cc18 collection of these tabular data sets the authors are going to test this scarf pre-training algorithm with we have all these different domains of tabular data we see the number of instances the number of features in each of the tabular records and the number of class labels to classify them in for supervised evaluation and then supervised fine tuning so each of these domains with each of the features has its own empirical marginal distribution which is how you sample the corruption for forming the data augmentation view for doing this self-supervised pre-training representation learning that facilitates with supervised learning for mapping these data sets to the this number of classes in the in this collection of data so here's some high level takeaways from the study they find the scarf pre-training improves predictive performance it improves performance in the presence of label noise and controlled experiments improves performance when labeled data is limited also in controlled experiments and then the other corruption strategies like not sampling from this empirical marginal distribution are less effective and more sensitive to feature scaling scarf isn't sensitive to batch size and is fairly insensitive to corruption rate and temperature tweaks and then the tweaks of the corruption don't work any better and then the alternatives to the info noise contrast of estimation contrastive loss don't really work any better than the tested algorithm thanks for watching this quick overview of the scarf algorithm bringing latest advances in contrast of self-supervised learning to tabular data please stay tuned for the rest of the ai weekly update series on henry ai labs [Music]

Original Description

Notion Link: https://ebony-scissor-725.notion.site/Henry-AI-Labs-Weekly-Update-July-15th-2021-a68f599395e3428c878dc74c5f0e1124 Thanks for watching! Please Subscribe!
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Connor Shorten · Connor Shorten · 0 of 60

← Previous Next →
1 DenseNets
DenseNets
Connor Shorten
2 DeepWalk Explained
DeepWalk Explained
Connor Shorten
3 Inception Network Explained
Inception Network Explained
Connor Shorten
4 StackGAN
StackGAN
Connor Shorten
5 StyleGAN
StyleGAN
Connor Shorten
6 Progressive Growing of GANs Explained
Progressive Growing of GANs Explained
Connor Shorten
7 Improved Techniques for Training GANs
Improved Techniques for Training GANs
Connor Shorten
8 Word2Vec Explained
Word2Vec Explained
Connor Shorten
9 Must Read Papers on GANs
Must Read Papers on GANs
Connor Shorten
10 Unsupervised Feature Learning
Unsupervised Feature Learning
Connor Shorten
11 Self-Supervised GANs
Self-Supervised GANs
Connor Shorten
12 Embedding Graphs with Deep Learning
Embedding Graphs with Deep Learning
Connor Shorten
13 Transfer Learning in GANs
Transfer Learning in GANs
Connor Shorten
14 ReLU Activation Function
ReLU Activation Function
Connor Shorten
15 AC-GAN Explained
AC-GAN Explained
Connor Shorten
16 SimGAN Explained
SimGAN Explained
Connor Shorten
17 DC-GAN Explained!
DC-GAN Explained!
Connor Shorten
18 ResNet Explained!
ResNet Explained!
Connor Shorten
19 Graph Convolutional Networks
Graph Convolutional Networks
Connor Shorten
20 Neural Architecture Search
Neural Architecture Search
Connor Shorten
21 Henry AI Labs
Henry AI Labs
Connor Shorten
22 Video Classification with Deep Learning
Video Classification with Deep Learning
Connor Shorten
23 BigGANs in Data Augmentation
BigGANs in Data Augmentation
Connor Shorten
24 Introduction to Deep Learning
Introduction to Deep Learning
Connor Shorten
25 EfficientNet Explained!
EfficientNet Explained!
Connor Shorten
26 Self-Attention GAN
Self-Attention GAN
Connor Shorten
27 Curriculum Learning in Deep Neural Networks
Curriculum Learning in Deep Neural Networks
Connor Shorten
28 Deep Learning Podcast #1 | Edward Dixon | Stochastic Weight Averaging
Deep Learning Podcast #1 | Edward Dixon | Stochastic Weight Averaging
Connor Shorten
29 Deep Compression
Deep Compression
Connor Shorten
30 Skin Cancer Classification with Deep Learning
Skin Cancer Classification with Deep Learning
Connor Shorten
31 Deep Learning Podcast #2 | Edward Peake | Deep Learning in Medical Imaging
Deep Learning Podcast #2 | Edward Peake | Deep Learning in Medical Imaging
Connor Shorten
32 The Lottery Ticket Hypothesis Explained!
The Lottery Ticket Hypothesis Explained!
Connor Shorten
33 SqueezeNet
SqueezeNet
Connor Shorten
34 GauGAN Explained!
GauGAN Explained!
Connor Shorten
35 AutoML with Hyperband
AutoML with Hyperband
Connor Shorten
36 DL Podcast #3 | Yannic Kilcher | Population-Based Search
DL Podcast #3 | Yannic Kilcher | Population-Based Search
Connor Shorten
37 Weakly Supervised Pretraining
Weakly Supervised Pretraining
Connor Shorten
38 Image Data Augmentation for Deep Learning
Image Data Augmentation for Deep Learning
Connor Shorten
39 Unsupervised Data Augmentation
Unsupervised Data Augmentation
Connor Shorten
40 Wide ResNet Explained!
Wide ResNet Explained!
Connor Shorten
41 RevNet: Backpropagation without Storing Activations
RevNet: Backpropagation without Storing Activations
Connor Shorten
42 GANs with Fewer Labels
GANs with Fewer Labels
Connor Shorten
43 BigBiGAN Unsupervised Learning!
BigBiGAN Unsupervised Learning!
Connor Shorten
44 Self-Supervised Learning
Self-Supervised Learning
Connor Shorten
45 Multi-Task Self-Supervised Learning
Multi-Task Self-Supervised Learning
Connor Shorten
46 Self-Supervised GANs
Self-Supervised GANs
Connor Shorten
47 Population Based Training
Population Based Training
Connor Shorten
48 Show, Attend and Tell
Show, Attend and Tell
Connor Shorten
49 Siamese Neural Networks
Siamese Neural Networks
Connor Shorten
50 WaveGAN Explained!
WaveGAN Explained!
Connor Shorten
51 VAE-GAN Explained!
VAE-GAN Explained!
Connor Shorten
52 Evolution in Neural Architecture Search!
Evolution in Neural Architecture Search!
Connor Shorten
53 AI Research Weekly Update August 18th, 2019
AI Research Weekly Update August 18th, 2019
Connor Shorten
54 Weight Agnostic Neural Networks Explained!
Weight Agnostic Neural Networks Explained!
Connor Shorten
55 AI Research Weekly Update August 25th, 2019
AI Research Weekly Update August 25th, 2019
Connor Shorten
56 Neuroevolution of Augmenting Topologies (NEAT)
Neuroevolution of Augmenting Topologies (NEAT)
Connor Shorten
57 CoDeepNEAT
CoDeepNEAT
Connor Shorten
58 AI Research Weekly Update September 1st, 2019
AI Research Weekly Update September 1st, 2019
Connor Shorten
59 Randomly Wired Neural Networks
Randomly Wired Neural Networks
Connor Shorten
60 Genetic CNN
Genetic CNN
Connor Shorten

The SCARF algorithm brings self-supervised contrastive learning to tabular data, improving predictive performance and robustness to label noise and limited labeled data. This is achieved through random feature corruption and empirical marginal distributions.

Key Takeaways
  1. Sample features uniformly at random
  2. Replace each feature with a random draw from its empirical marginal distribution
  3. Form positive pairs using the corrupted features
  4. Align representations using a momentum encoding average
  5. Evaluate the algorithm on various tabular data sets
💡 The SCARF algorithm's use of empirical marginal distributions for corruption allows it to be more effective and robust than other corruption strategies.

Related Reads

📰
How AI Is Transforming the Newsroom: A 2026 Overview
Learn how AI is transforming the newsroom in 2026, from research and discovery to ethics and transparency, and what it means for readers and journalists.
Medium · Machine Learning
📰
Your AI Job Interview Is Failing You — Here’s When To Walk Away
Learn when to walk away from an AI job interview that's not going well and how to prioritize your own needs
Forbes Innovation
📰
From Ideas to Execution: What I’m Learning as an AI Developer Intern
Learn how to turn ideas into execution as an AI developer intern and improve collaboration in a team
Medium · Startup
📰
How AI Is Changing the Future of Jobs: Opportunity, Adaptation, and the Skills That Matter Most
Learn how AI is changing the future of jobs and the skills that matter most in this new landscape
Medium · AI
Up next
FREE Unlimited AI Videos 🔥 Higgsfield Is Now Completely Free for 24 hours only #ai #aivideo
Raj Photo Editing and Much More
Watch →