SFT vs. RL: What Changes Inside the Model?
📰 Medium · Deep Learning
Understand how SFT and RL affect model weights and singular values, and why it matters for deep learning
Action Steps
- Read the article on SFT vs. RL to understand the concepts
- Analyze the weight spectra of a model before and after fine-tuning using SFT
- Compare the singular values of a model trained with SFT vs. RL
- Visualize the rotation of frames in the model's weight space using dimensionality reduction techniques
- Implement SFT and RL in a deep learning project to observe the differences in practice
Who Needs to Know This
Researchers and engineers working on deep learning models can benefit from understanding the differences between SFT and RL, and how they impact model performance
Key Insight
💡 SFT rewrites the singular values of a model's weights, while RL keeps them and rotates the frames around
Share This
🤖 SFT vs. RL: what changes inside the model? 🤔
Full Article
An intuitive guide to SFT, RLVR, and weight spectra: fine-tuning rewrites the singular values, RL keeps them and rotates the frames around… Continue reading on Medium »
Related Videos
⚡
You're 1 lesson closer to your goal
Sign in free and we'll turn this lesson into a structured roadmap — starting with ⚡30 free Sparks for your first AI explanation or skill path.
Create free account →No credit card required.
DeepCamp AI