AI Chatbot Training From Pre Training to Human Feedback
Key Takeaways
The video covers AI chatbot training, specifically pre-training and reinforcement learning with human feedback, using techniques such as autocompleting text passages and flagging unhelpful predictions.
Full Transcript
This is only part of the story though. This whole process is called pre-training. The goal of autocompleting a random passage of text from the internet is very different from the goal of being a good AI assistant. To address this, chatbots undergo another type of training just as important called reinforcement learning with human feedback. Workers flag unhelpful or problematic predictions, and their corrections further change the model's parameters, making them more likely to give predictions that users prefer.
Original Description
Pre-training is only part of how chatbots learn. Reinforcement learning with human feedback is also crucial. Workers flag unhelpful ...
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
More on: LLM Foundations
View skill →Related Reads
📰
📰
📰
📰
Help Choosing Neural Network Architecture for Matrix Classification
Reddit r/deeplearning
How to Choose the Best Deep Learning Model for Medical Imaging
Medium · Deep Learning
Another Way to Read Neural Geometry
Medium · Data Science
Another Way to Read Neural Geometry
Medium · Deep Learning
🎓
Tutor Explanation
DeepCamp AI