OpenAI DevDay 2024 | OpenAI Research

OpenAI · Intermediate ·📄 Research Papers Explained ·1y ago

Key Takeaways

The video discusses building with o1, a reasoning model trained with reinforcement learning, and its applications in various domains, including math, coding, and science. It highlights the differences between o1 and previous models, such as GPT-40, and provides guidance on when to use o1 preview and o1 mini.

Full Transcript

hello my name is Yan uh today with Jason we're happy to share our thoughts on how you can build with oan so oan is a reasoning model and we train oan to think with reinforcement learning and during the training phase it learns among other things to refine the thinking strategies and recognize and correct its mistakes so when one one attempts to solve a very difficult problem it may not get to the working strategy in one go but just by trying a strategy even if it's unsuccessful can give some cues as to what to try next and so 01 does this and then eventually it gets to the better strategy and so on so it's very patient it's a very different typees of model and so last month when we released on preview we uh had showed the some of the examples of actual Channel thought going on so in this particular example we have the cipher text the model is trying to decipher a some kind of a text and I'll show you some uh reasoning patterns so in this case the model says H so probably it's like recognizing that this current thinking strategy is not really leading to anywhere and probably trying trying something different in this case the model tries something out and then says wait a minute so the probably realizing slightly better approach and then trying that next and then in this case the model has more conrete idea to try out so it says let's test this Theory and if after a while model gets to a more correct situation correct like a strategy and says it's perfect so the Model Behavior is very very different and in fact it's so different that we believe o1 represents a new paradigm and new paradigm changes so many things that I think we should think about a having a New Perspective so what really changes I think a good starting point to think about this question is where we were and then where we are now and where we're heading towards and so I invite you to think about questions like these what just became possible with o1 that wasn't really possible with the previous gener generation of the model and what will become possible with the future versions of o1 and obviously the answers to these questions will be very different uh depending on the specific domain but just thinking about these questions will put you into a a modo thinking where okay now we're building with the future models uh in mind instead of thinking that the current generations of the model will stay as is so um you might say okay I'm not building o1 myself so how do I know what future generations of the o1 might look like prev unlike the previous uh paradigms o one Paradigm is much simpler in that it's a reasoning model so its reasonings will be just better meaning that it will be able to think better in Prett much anything that requires thinking so with that I think it's useful to think about um when you're building something uh considering questions like this what would you want to build if reasoning is 50% better than what we have now what would you want to do differently and then maybe more importantly what would you not want to build if reasoning is 50% better and we have seen in many cases that as the model gets generally smarter some of the problems that we think are difficult uh in the past just became trably possible so if we believe that the reasoning will just continually get better we should think about what problems to not solve as well so with this new paradigm I have been working for a while but I'm still really uh having some trouble because I'm so used to the previous Paradigm of um and so I think this is be really useful I hope this Sparks some interest in thinking about how to uh build with this new reasoning Paradigm with that I'll pass it on to Jason thanks youngan um I wanted to talk a little more specifically I wanted to talk a little more specifically about a few EV valves that we've shown at the blog post that might guide you for when you might want to use 01 and 01 preview uh compared to GPT 40 so I think one of the best use cases for models in the 01 Paradigm are for extremely hard math and code problems so you'll see in these bar charts uh we have on the left IM which is competition math and on the right code forces um and there's uh these three bars so GPT 40 inal 01 preview uh and then 01 and the thing to note here is that GPT 40 and 01 preview are barely solving a few questions in these benchmarks um and then you could see o1 preview can solve more than half and o1 uh can solve the majority of problems in this data set so the point is there's some subset of tasks where GPT 40 is really struggling um and uh uh 01 can solve the majority of problems um here's a broader evaluation Suite uh that we published uh in the blog post um and I'll highlight a few things here so uh first um if you look at uh the performance on some of the math benchmarks so uh math from Hendrick uh physics college math uh elsat there's a huge performance gain um when you use 01 preview compared to GPT 40 and then conversely you don't get a huge performance gain for every task so there's some tasks like AP English Language literature sat public relations where we actually don't see o1 preview doing a lot better than GPT 40 and so I made this table that sort of summarizes when you might want to use models from the 01 Paradigm uh versus GPT 40 so I would say the pros of using models from 01 preview and 01 is if you're trying to do tasks that are extremely challenging prompts um and in their domains of science math and coding um and the other use case would be if you don't care about any other constraints and you just want the best answer uh then 01 will likely be the most performant model generally the cons here are obviously that because as hangan as hangan mentioned um 01 preview and o1 require time to think it's going to be a lot more expensive and there's going to be a much higher latency when you use these models and then for gb24 I would say that's still great for the majority of use cases that people are currently using the API for obviously it's less expensive and lower latency uh than o1 preview and 01 um and I guess the con would be that it's weaker on than o1 on prompts that require uh reasoning or or strong coding or math um and then there's the question of when should you use o1 preview versus 01 mini um and this plot here shows the inference cost versus performance of a few of these models um so you could see uh the inference cost on the x-axis um and then on the y- AIS is performance on imy which is competition math um and interestingly we see that 01 pre 01 mini is actually strictly better than 01 preview and this is because we really specialized 01 mini to be a fast but performant model on things like math and coding um and I would say you should use o1 mini if you're doing math coding and or if you want the answer uh more quickly or cheaply um but I would say in other cases uh o1 preview is a good choice um finally I wanted to highlight a few use cases uh in the API of o1 preview and o1 mini so I like this one from the uh open a cookbook here uh it's basically a uh medical inaccuracy detection so you given a bunch of information um and diagnosis um and then o1 preview tries to detect whether it's a correct diagnosis or not um obviously coding uh is another great example of uh uh o1 preview shining so for use cases like cursor I think o1 preview would do a great job um hard Sciences research is another great use case that that o1 preview is particularly strong at um and we've also heard that these models have been good uh as a brainstorming partner on math problems or uh on legal domain reasoning um so I'll end here and enjoy using preview and mini thank you

Original Description

Building with o1
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from OpenAI · OpenAI · 0 of 60

← Previous Next →
1 Robots that Learn
Robots that Learn
OpenAI
2 Emergence of Grounded Compositional Language in Multi-Agent Populations
Emergence of Grounded Compositional Language in Multi-Agent Populations
OpenAI
3 OpenAI + Dota 2
OpenAI + Dota 2
OpenAI
4 Dendi vs. OpenAI at The International 2017
Dendi vs. OpenAI at The International 2017
OpenAI
5 Competitive Self-Play
Competitive Self-Play
OpenAI
6 Learning a Hierarchy
Learning a Hierarchy
OpenAI
7 Physical Spam Detection
Physical Spam Detection
OpenAI
8 Ingredients for Robotics Research
Ingredients for Robotics Research
OpenAI
9 OpenAI Five
OpenAI Five
OpenAI
10 OpenAI Five: Dota Gameplay
OpenAI Five: Dota Gameplay
OpenAI
11 Learning Dexterity
Learning Dexterity
OpenAI
12 Learning Dexterity: Uncut
Learning Dexterity: Uncut
OpenAI
13 OpenAI Five Benchmark: Post-Game Analysis
OpenAI Five Benchmark: Post-Game Analysis
OpenAI
14 Investigating Model Based RL for Continuous Control | Alex Botev | 2018 Summer Intern Open House
Investigating Model Based RL for Continuous Control | Alex Botev | 2018 Summer Intern Open House
OpenAI
15 Generative Modelling | Sadhika Malladi | 2018 Summer Intern Open House
Generative Modelling | Sadhika Malladi | 2018 Summer Intern Open House
OpenAI
16 A pathway to more efficient generative models | Will Grathwohl | 2018 Summer Intern Open House
A pathway to more efficient generative models | Will Grathwohl | 2018 Summer Intern Open House
OpenAI
17 Learning Dexterity | Alex Ray | 2018 Summer Intern Open House
Learning Dexterity | Alex Ray | 2018 Summer Intern Open House
OpenAI
18 Robust Vision-Based State Estimation | Hsiao-Yu 'Fish' Tung | 2018 Summer Intern Open House
Robust Vision-Based State Estimation | Hsiao-Yu 'Fish' Tung | 2018 Summer Intern Open House
OpenAI
19 Using Semantic Trees In Place of Sentences | Munashe Shumba | OpenAI Scholars Demo Day 2018
Using Semantic Trees In Place of Sentences | Munashe Shumba | OpenAI Scholars Demo Day 2018
OpenAI
20 Reinforcement Learning with Prediction-Based Rewards
Reinforcement Learning with Prediction-Based Rewards
OpenAI
21 OpenAI Spinning Up in Deep RL Workshop
OpenAI Spinning Up in Deep RL Workshop
OpenAI
22 Arena Announcement and Closing | OpenAI Five Finals (6/6)
Arena Announcement and Closing | OpenAI Five Finals (6/6)
OpenAI
23 Co-Op Match | OpenAI Five Finals (5/6)
Co-Op Match | OpenAI Five Finals (5/6)
OpenAI
24 OpenAI Five vs. OG, Game 2 | OpenAI Five Finals (4/6)
OpenAI Five vs. OG, Game 2 | OpenAI Five Finals (4/6)
OpenAI
25 OpenAI Five vs. OG, Game 1 | OpenAI Five Finals (3/6)
OpenAI Five vs. OG, Game 1 | OpenAI Five Finals (3/6)
OpenAI
26 Pre-Match Panel Discussion | OpenAI Five Finals (2/6)
Pre-Match Panel Discussion | OpenAI Five Finals (2/6)
OpenAI
27 Opening Keynote | OpenAI Five Finals (1/6)
Opening Keynote | OpenAI Five Finals (1/6)
OpenAI
28 OpenAI Robotics Symposium 2019
OpenAI Robotics Symposium 2019
OpenAI
29 OpenAI Scholars Demo Day 2019
OpenAI Scholars Demo Day 2019
OpenAI
30 Multi-Agent Hide and Seek
Multi-Agent Hide and Seek
OpenAI
31 Solving Rubik’s Cube with a Robot Hand: Uncut
Solving Rubik’s Cube with a Robot Hand: Uncut
OpenAI
32 Solving Rubik’s Cube with a Robot Hand: Perturbations
Solving Rubik’s Cube with a Robot Hand: Perturbations
OpenAI
33 Solving Rubik’s Cube with a Robot Hand
Solving Rubik’s Cube with a Robot Hand
OpenAI
34 Music Generation | Christine Payne | OpenAI Scholars Demo Day 2018
Music Generation | Christine Payne | OpenAI Scholars Demo Day 2018
OpenAI
35 Deephypebot | Nadja Rhodes | OpenAI Scholars Demo Day 2018
Deephypebot | Nadja Rhodes | OpenAI Scholars Demo Day 2018
OpenAI
36 Physics Net | Ifu Aniemeka | OpenAI Scholars Demo Day 2018
Physics Net | Ifu Aniemeka | OpenAI Scholars Demo Day 2018
OpenAI
37 Art Composition Attributes + CycleGAN | Holly Grimm | OpenAI Scholars Demo Day 2018
Art Composition Attributes + CycleGAN | Holly Grimm | OpenAI Scholars Demo Day 2018
OpenAI
38 Generating Emotional Landscapes | Hannah Davis | OpenAI Scholars Demo Day 2018
Generating Emotional Landscapes | Hannah Davis | OpenAI Scholars Demo Day 2018
OpenAI
39 Looking For Grammar In All The Right Places | Alethea Power | OpenAI Scholars Demo Day 2020
Looking For Grammar In All The Right Places | Alethea Power | OpenAI Scholars Demo Day 2020
OpenAI
40 Semantic Parsing English to GraphQL | Andre Carerra | OpenAI Scholars Demo Day 2020
Semantic Parsing English to GraphQL | Andre Carerra | OpenAI Scholars Demo Day 2020
OpenAI
41 Long term credit assignment with temporal reward transp… | Cathy Yeh | OpenAI Scholars Demo Day 2020
Long term credit assignment with temporal reward transp… | Cathy Yeh | OpenAI Scholars Demo Day 2020
OpenAI
42 Social learning in independent multi-agent reinfor… | Kamal N’dousse | OpenAI Scholars Demo Day 2020
Social learning in independent multi-agent reinfor… | Kamal N’dousse | OpenAI Scholars Demo Day 2020
OpenAI
43 Quantifying Interpretability of Models Trained on Coi… | Jorge Orbay | OpenAI Scholars Demo Day 2020
Quantifying Interpretability of Models Trained on Coi… | Jorge Orbay | OpenAI Scholars Demo Day 2020
OpenAI
44 Towards Epileptic Seizure Prediction with Deep Network | Kata Slama | OpenAI Scholars Demo Day 2020
Towards Epileptic Seizure Prediction with Deep Network | Kata Slama | OpenAI Scholars Demo Day 2020
OpenAI
45 Universal Adversarial Perturbations and Language M… | Pamela Mishkin | OpenAI Scholars Demo Day 2020
Universal Adversarial Perturbations and Language M… | Pamela Mishkin | OpenAI Scholars Demo Day 2020
OpenAI
46 Introductions by Sam Altman & Greg Brockman | OpenAI Scholars Demo Day 2020
Introductions by Sam Altman & Greg Brockman | OpenAI Scholars Demo Day 2020
OpenAI
47 Introduction by Sam Altman | OpenAI Scholars Demo Day 2021
Introduction by Sam Altman | OpenAI Scholars Demo Day 2021
OpenAI
48 Breaking Contrastive Models with the SET Card Game | Legg Yeung | OpenAI Scholars Demo Day 2021
Breaking Contrastive Models with the SET Card Game | Legg Yeung | OpenAI Scholars Demo Day 2021
OpenAI
49 Large Scale Reward Modeling | Jonathan Ward | OpenAI Scholars Demo Day 2021
Large Scale Reward Modeling | Jonathan Ward | OpenAI Scholars Demo Day 2021
OpenAI
50 Words to Bytes: Exploring Language Tokenizations | Sam Gbafa | OpenAI Scholars Demo Day 2021
Words to Bytes: Exploring Language Tokenizations | Sam Gbafa | OpenAI Scholars Demo Day 2021
OpenAI
51 Learning Multiple Modes of Behavior in a Continuous… | Tyna Eloundou | OpenAI Scholars Demo Day 2021
Learning Multiple Modes of Behavior in a Continuous… | Tyna Eloundou | OpenAI Scholars Demo Day 2021
OpenAI
52 Scaling Laws for Language Transfer Learning | Christina Kim | OpenAI Scholars Demo Day 2021
Scaling Laws for Language Transfer Learning | Christina Kim | OpenAI Scholars Demo Day 2021
OpenAI
53 Contrastive Language Encoding | Ellie Kitanidis | OpenAI Scholars Demo Day 2021
Contrastive Language Encoding | Ellie Kitanidis | OpenAI Scholars Demo Day 2021
OpenAI
54 Characterizing Test Time Compute on Graph Structur… | Kudzo Ahegbebu | OpenAI Scholars Demo Day 2021
Characterizing Test Time Compute on Graph Structur… | Kudzo Ahegbebu | OpenAI Scholars Demo Day 2021
OpenAI
55 Studying Scaling Laws for Transformer Architecture … | Shola Oyedele | OpenAI Scholars Demo Day 2021
Studying Scaling Laws for Transformer Architecture … | Shola Oyedele | OpenAI Scholars Demo Day 2021
OpenAI
56 Feedback Loops in Opinion Modeling | Danielle Ensign | OpenAI Scholars Demo Day 2021
Feedback Loops in Opinion Modeling | Danielle Ensign | OpenAI Scholars Demo Day 2021
OpenAI
57 Creating a Space Game with OpenAI Codex
Creating a Space Game with OpenAI Codex
OpenAI
58 “Hello World” with OpenAI Codex
“Hello World” with OpenAI Codex
OpenAI
59 Talking to Your Computer with OpenAI Codex
Talking to Your Computer with OpenAI Codex
OpenAI
60 Data Science with OpenAI Codex
Data Science with OpenAI Codex
OpenAI

The video discusses the applications and differences of o1, a reasoning model, and provides guidance on when to use o1 preview and o1 mini. It highlights the importance of considering the future of AI models when building with them.

Key Takeaways
  1. Understand the basics of o1 and its applications
  2. Evaluate the differences between o1 and previous models, such as GPT-40
  3. Determine when to use o1 preview and o1 mini based on the specific use case
  4. Consider the future of AI models when building with them
💡 The o1 model represents a new paradigm in AI, with the ability to reason and think differently than previous models, and its applications will continue to grow as the model improves.

Related Reads

Up next
The Government Just Made Investing in India Way Easier for NRIs
marketfeed
Watch →