The Real Reason Huge AI Models Actually Work [Prof. Andrew Wilson]

Machine Learning Street Talk · Beginner ·📐 ML Fundamentals ·10mo ago

About this lesson

Why can billion-parameter models perform so well without catastrophically overfitting? The answer lies in the mysterious "simplicity bias" that emerges at scale, a core concept of the double descent phenomenon. Professor Andrew Wilson from NYU explains why many common-sense ideas in artificial intelligence might be wrong. For decades, the rule of thumb in machine learning has been to fear complexity. The thinking goes: if your model has too many parameters (is "too complex") for the amount of data you have, it will "overfit" by essentially memorizing the data instead of learning the underlying patterns. This leads to poor performance on new, unseen data. This is known as the classic "bias-variance trade-off" i.e. a balancing act between a model that's too simple and one that's too complex. **SPONSOR MESSAGES** — Tufa AI Labs is an AI research lab based in Zurich. **They are hiring ML research engineers!** This is a once in a lifetime opportunity to work with one of the best labs in Europe Contact Benjamin Crouzier - https://tufalabs.ai/ — Take the Prolific human data survey - https://www.prolific.com/humandatasurvey?utm_source=mlst and be the first to see the results and benchmark their practices against the wider community! — cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy Oct SF conference - https://dagihouse.com/?utm_source=mlst - Joscha Bach keynoting(!) + OAI, Anthropic, NVDA,++ Hiring a SF VC Principal: https://talent.cyber.fund/companies/cyber-fund-2/jobs/57674170-ai-investment-principal#content?utm_source=mlst Submit investment deck: https://cyber.fund/contact?utm_source=mlst — Description Continued: Professor Wilson challenges this fundamental belief (fearing complexity). He makes a few surprising points: **Bigger Can Be Better**: massive models don't just get more flexible; they also develop a stronger "simplicity bias". So, if your model is overfitting, the solution might paradoxi

Original Description

Why can billion-parameter models perform so well without catastrophically overfitting? The answer lies in the mysterious "simplicity bias" that emerges at scale, a core concept of the double descent phenomenon. Professor Andrew Wilson from NYU explains why many common-sense ideas in artificial intelligence might be wrong. For decades, the rule of thumb in machine learning has been to fear complexity. The thinking goes: if your model has too many parameters (is "too complex") for the amount of data you have, it will "overfit" by essentially memorizing the data instead of learning the underlying patterns. This leads to poor performance on new, unseen data. This is known as the classic "bias-variance trade-off" i.e. a balancing act between a model that's too simple and one that's too complex. **SPONSOR MESSAGES** — Tufa AI Labs is an AI research lab based in Zurich. **They are hiring ML research engineers!** This is a once in a lifetime opportunity to work with one of the best labs in Europe Contact Benjamin Crouzier - https://tufalabs.ai/ — Take the Prolific human data survey - https://www.prolific.com/humandatasurvey?utm_source=mlst and be the first to see the results and benchmark their practices against the wider community! — cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy Oct SF conference - https://dagihouse.com/?utm_source=mlst - Joscha Bach keynoting(!) + OAI, Anthropic, NVDA,++ Hiring a SF VC Principal: https://talent.cyber.fund/companies/cyber-fund-2/jobs/57674170-ai-investment-principal#content?utm_source=mlst Submit investment deck: https://cyber.fund/contact?utm_source=mlst — Description Continued: Professor Wilson challenges this fundamental belief (fearing complexity). He makes a few surprising points: **Bigger Can Be Better**: massive models don't just get more flexible; they also develop a stronger "simplicity bias". So, if your model is overfitting, the solution might paradoxi
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

📰
Reproducing OpenAI’s “persistently beneficial models” - GRPO trait install barely moves. Ideas? [P] [R]
Reproduce OpenAI's persistently beneficial models by troubleshooting GRPO trait installation with small-scale RL
Reddit r/MachineLearning
📰
Mixture Density Networks
Learn how Mixture Density Networks can improve prediction models for complex tasks like self-driving cars
Medium · Machine Learning
📰
How to Crack Technical Interviews as a Fresher
Learn strategies to crack technical interviews as a fresher and increase your chances of landing your first job in the tech industry
Dev.to · Subhalaxmi Paikaray
📰
Programming Assignments: A Complete Guide to Solving Coding Problems Faster and Smarter
Learn to solve coding problems faster and smarter with a complete guide to programming assignments
Medium · Programming
Up next
SQLite3 Tutorial - Learn SQL for Python in 17 Minutes
Thomas Janssen
Watch →