How to train a Million Context LLM — with Mark Huang of Gradient.ai
Key Takeaways
This video teaches how to train a large context LLM using tools like Gradient.ai and techniques such as ALiBi and RoPE
Original Description
Scaling Llama3 beyond 1M context window with ~perfect utilization, the difference between ALiBi and RoPE, how to use GPT-4 to create synthetic data for your context extension finetunes, and more!
Full writeup: https://www.latent.space/p/gradient
00:00:00 Introductions
00:01:30 Founding story of Gradient and its mission
00:04:35 "Minimum viable agents"
00:09:19 Differentiating ML and AI, focusing on out-of-domain generalization
00:10:12 Extending Llama3 to 1M tokens
00:14:32 Technical challenges with long context sequences
00:17:45 Data quality and the importance of diverse datasets
00:19:45 What's a theta value?
00:22:42 RoPE vs Ring Attention vs ALiBi vs YaARN
00:25:06 Why RingAttention matters
00:28:01 How to refine datasets for context extension
00:33:34 Multi-stage training data and avoiding overfitting to recent data
00:34:27 The potential of using synthetic data in training
00:38:22 Applying LoRa adapters to extend model capabilities
00:42:25 Benchmarking long context models and evaluating their performance
00:47:20 Pushing to 4M context and output quality degradation
00:50:08 What do you need this context for?
00:52:57 Impact of long context in chat vs Docs Summarization
00:56:25 Future directions for long context models and multimodality
00:59:38 How do you know what research matters?
01:02:47 Routine for staying updated with AI research and industry news
01:05:33 Deciding which AI developments to invest time in
01:10:37 Request for collaboration and data set construction for long context
AI explanation not available for this lesson yet
This lesson is still being prepared for the AI tutor. In the meantime, explore lessons that are ready.
Browse explainer-ready lessons →
Playlist
Uploads from Latent Space · Latent Space · 30 of 60
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
▶
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Ep 18: Petaflops to the People — with George Hotz of tinycorp
Latent Space
FlashAttention-2: Making Transformers 800% faster AND exact
Latent Space
RWKV: Reinventing RNNs for the Transformer Era
Latent Space
Generating your AI Media Empire - with Youssef Rizk of Wondercraft.ai
Latent Space
RAG is a hack - with Jerry Liu of LlamaIndex
Latent Space
The End of Finetuning — with Jeremy Howard of Fast.ai
Latent Space
Why AI Agents Don't Work (yet) - with Kanjun Qiu of Imbue
Latent Space
Powering your Copilot for Data - with Artem Keydunov from Cube.dev
Latent Space
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Latent Space
The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Latent Space
The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
Latent Space
The AI-First Graphics Editor - with Suhail Doshi of Playground AI
Latent Space
The Accidental AI Canvas - with Steve Ruiz of tldraw
Latent Space
The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Latent Space
The Four Wars of the AI Stack - Dec 2023 Recap
Latent Space
The State of AI in production — with David Hsu of Retool
Latent Space
Building an open AI company - with Ce and Vipul of Together AI
Latent Space
Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal
Latent Space
A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
Latent Space
Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Latent Space
Making Transformers Sing - with Mikey Shulman of Suno
Latent Space
A Comprehensive Overview of Large Language Models - Latent Space Paper Club
Latent Space
Why Google failed to make GPT-3 -- with David Luan of Adept
Latent Space
Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
Latent Space
Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
Latent Space
Breaking down the OG GPT Paper by Alec Radford
Latent Space
High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Latent Space
This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Latent Space
LLM Asia Paper Club Survey Round
Latent Space
How to train a Million Context LLM — with Mark Huang of Gradient.ai
Latent Space
How AI is Eating Finance - with Mike Conover of Brightwave
Latent Space
How To Hire AI Engineers (ft. James Brady and Adam Wiggins of Elicit)
Latent Space
State of the Art: Training 70B LLMs on 10,000 H100 clusters
Latent Space
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Latent Space
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Latent Space
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
Latent Space
Synthetic data + tool use for LLM improvements 🦙
Latent Space
RLHF vs SFT to break out of local maxima 📈
Latent Space
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
Latent Space
Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
Latent Space
Answer.ai & AI Magic with Jeremy Howard
Latent Space
Is finetuning GPT4o worth it?
Latent Space
Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
Latent Space
Building AGI with OpenAI's Structured Outputs API
Latent Space
Q* for model distillation 🍓
Latent Space
Finetuning LoRAs on BILLIONS of tokens 🤖
Latent Space
Cursor UX team is CRACKED 💻
Latent Space
Choosing the BEST OpenAI model 🏆
Latent Space
How will OpenAI voice mode change API design?
Latent Space
STEALING OpenAI models data 🥷
Latent Space
[Paper Club] 🍓 On Reasoning: Q-STaR and Friends!
Latent Space
[Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
Latent Space
The Ultimate Guide to Prompting - with Sander Schulhoff from LearnPrompting.org
Latent Space
llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
Latent Space
Prompt Engineer is NOT a job 📝
Latent Space
Prompt Mining LLMs for better prompts ⛏️
Latent Space
The six pillars of few-shot prompting 🔧
Latent Space
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Latent Space
[Paper Club] Who Validates the Validators? Aligning LLM-Judges with Humans (w/ Eugene Yan)
Latent Space
Can you separate intelligence and knowledge?
Latent Space
Related Reads
📰
📰
📰
📰
Jev AI: The Rise of System One Models and a New Interface Between AI and Software
Medium · LLM
Claude Opus 5.5: cheaper, faster, and it won't stop thinking
Medium · Machine Learning
VEKTOR v1.9.5 is live: your choice of over 600 LLMs across 9 providers
Medium · LLM
How Model Context Protocol (MCP) changes SaaS feature rollouts forever
Dev.to AI
Chapters (23)
Introductions
1:30
Founding story of Gradient and its mission
4:35
"Minimum viable agents"
9:19
Differentiating ML and AI, focusing on out-of-domain generalization
10:12
Extending Llama3 to 1M tokens
14:32
Technical challenges with long context sequences
17:45
Data quality and the importance of diverse datasets
19:45
What's a theta value?
22:42
RoPE vs Ring Attention vs ALiBi vs YaARN
25:06
Why RingAttention matters
28:01
How to refine datasets for context extension
33:34
Multi-stage training data and avoiding overfitting to recent data
34:27
The potential of using synthetic data in training
38:22
Applying LoRa adapters to extend model capabilities
42:25
Benchmarking long context models and evaluating their performance
47:20
Pushing to 4M context and output quality degradation
50:08
What do you need this context for?
52:57
Impact of long context in chat vs Docs Summarization
56:25
Future directions for long context models and multimodality
59:38
How do you know what research matters?
1:02:47
Routine for staying updated with AI research and industry news
1:05:33
Deciding which AI developments to invest time in
1:10:37
Request for collaboration and data set construction for long context
🎓
Tutor Explanation
DeepCamp AI