NEW RL Method: FlowRL (GFlowNets)

Discover AI · Beginner ·🎮 Reinforcement Learning ·10mo ago

Key Takeaways

The video discusses FlowRL, a new reinforcement learning method that uses GFlowNets to enhance LLM reasoning by shifting from traditional reward maximization to reward distribution matching via flow balancing. This method addresses mode collapse, a fundamental weakness in maximization-based methods, and is considered a better paradigm for finding solutions in complex domains.

Full Transcript

Hello community. So great that you are back. We have a new reinforcement learning idea and a brand new paper. I already showed you in my last week's video September 18. We have here flow reinforcement learning matching the reward distribution for LLM reasoning. And if you just look here at the Kback Liber divergence here for a reward maximizing here for PO and GRPO, you see we are really high. And if you look here for flow reinforcement learning you see that our kubek liber divergence here the difference between two probability distribution is almost zero and this is exactly what we want. So what beautiful thing is here right in front of us. Let's have a deep dive. Now you know whenever we fine-tune here our new reasoning model our huge reasoning models you have here the finetuning you go with a PO from 2017 or DPO or whatever you have but you fight always against the mode collapse where your model output becomes repetitive and lack of diversity and it's just collapsing into one dimensionality. Now flow reinforcement learning is not just a cleav hack. It is a complete new idea with an extensive mathematical body from the generative flow networks G flow nets and the energy based model providing you really a solid mathematical justification. Why this system now is the new reinforcement learning paradigm in AI because you know what it does it reframes the policy optimization by data. I know it is fascinating. So we have to learn a little bit. I had to look at this paper here. This is from 2021. This is here from Benju. It's here one of the first works about the flow network based generative models for this idea. And he says you know I want to have a view of a generative process as a flow network making it possible to handle here the tricky case where different trajectories can yield the same final state and he did this here on the molecular graph complexity. Now if you look at this what's the most important idea in this paper? You say hey I do have a mark of decision process and now I view this as a flow network. So this means I leverage here my directed a graph structure of a mark of decision process and learn a flow rather than estimate here myo values as a sum of descendant rewards as elaborated below. So a flow network mock of decision process and you immediately understand if you have seen my video that AI every AI model runs on uncertainty and of course at the core of every AI is a partially observable mock of decision process. Why this is here really shaking here up the base of any AI model that we have today. Now in the study they go here on the flow condition that require for any node the incoming flow equals the outgoing flow and the total flow and everything that is great. But you know what's really important is here they found here to compute here a loss function and now you see the loss function at first here is a square which makes it much easier and we have the logs of our expression. I will explain later why this is so beautiful. So we will use now this complex mathematical apparatus for reinforcement learning. And you said okay but but we were talking about reinforcement learning. So how does it fit? Now if we have a reward distribution, how can we reduce our coolback lava divergence here from a horrible 8.6 to a beautiful 0.1? Well, that's it. So let's have a look. And I do not want to give you right in the front of the pure mathematical explanation. But I would like to explain it in simple terms and I would like to find an example that makes it simple to understand because a lot of my subscriber ask can you explain it in simple terms. So I will try my best here and I invested more time for the simple explanation than for the mathematical explanation. So I would say let's start here. So I want to understand here this complexity from a flow perspective you know. So I say I have a flow whatever the flow is. So let's go with the flow that you can imagine a flow of water. So we have rain. We have clouds and it's raining over the mountains. Yes, I'm sitting here in Austria, beautiful part of Europe. And this rain is here, the source of the water, you know. So we have a flow pouring down. And then we have here a complexity. And now this complexity is a three-dimensional mountain terrain as you can imagine. No, rocky terrain all over. This is here our complexity representation in this Gdankan experiment. And now the water is raining flowing down here this complex terrain on the Rocky Mountains into different valleys. And now in those valleys no that is more or less our stable conditions they have different fatalities colorar indicators for the crops that grow in this valley. So you see the simple idea. Now the flow reinforcement learning is now a mechanism to steer here the right amount of water into each of the values in our mountain given here the particular fatalities color indicator of each and every particular value. So every value should have the right amount of water to make it prosper. The aim is now to have a perfect flow from the rain raining down here on this rocky terrain to provide now each path down the rocky terrain down to the valley. And this path is of course the reasoning path, the chain of sort path if you want from the LLM that is generated with enough water to irrigate now the soil in all the valleys perfectly for their specific scolar fertility coefficient. for their complexity. Now let's have a comparison. Maybe you don't get it immediately. So what is the rain over the mountains? The rain not over the mountains is simply the initial state of our complete system. It's your prompt. So this is the source of all potential solutions. No. So for a given mathematical problem, this is the problem statement itself. Find here a quantum field theoretical explanation of ABC. Great. Then the flow of water. The flow of water is in our statistical model of course the probability mass that we are dealing with and the the amount of water is the probability itself. Either you normally to one or 100% doesn't mind. Flow reinforcement learning decides how to distribute this particular probability of the probability mass or of the flow of the water down over the rocky terrain. And the rocky terrain is of course our gradient landscape. So the mountain terrain is the space of all possible token sequences. Let's change to this perspective. If we say the mountain terrain is the space of all if you want answers given over those token sequences a short token sequences a short sentence whatever a vast complex landscape of all possible chains of sort that leads to the solution of the problem. So some paths are easy and direct. This will be our deep tunnels. Other are convoluted and unlikely. So have a low probability some rocky trickles only. But there's more. Now of course there's the path down to the valley and this is now in the complexity a reasoning trajectory why and if you want this is our chain of sort. So this is like here a specific sequence of step. This is the answer given by our LLM or our coding LLM. If you say hey I want a solution to this sequence of codes or a mathematical LLM where I say hey solve this particular equation it's the answer p down to a valley but the value itself now as I told you are the final solution or if you want to see this from a state representation the trouble states of our system is at the end of the reasoning path we have an answer here the final answer generated by a vision language model and LLM now the fertility of the value is of course remember we are working here in the reinforcement learning representation the reward function and let's go here with a scalar reward that has here I called it here the fertility of the land a high fertility value represents of course a correct elegant or high quality solution so a high reward reply and if you have here born rocky value this is more or less an incorrect solution so I'm working with this very simple analog you see yes of course looking out of the window I almost see this so you understand why The current way to do this, the current algorithms like PO and GRPO are pure reward maximizers. So they behave like an aggressive civil engineer that comes to the mountain one goal. Maximize here the crop yield in the single most fertile valley. So the engineer says, hey, I see seven valleys down here looking down from the mountain, but only one green valley. And now I focus all the water the downstream here to this single value. This is a reward maximizer. But we don't want this. We want to have here new solution. We don't want that every system just collapses here to one particular solution. We don't want a mode collapse happening for our system. But because this is what's happening now in our GPO result, we have one champion valley here. This is perfectly irrigated and incredible productive here and this is the go-to solution that the AI has for a particular problem coding explanation whatever you have and all the other values all the other possible solution paths even those that are maybe a 90% correct yeah but have maybe a path that would be essential for the next problem for a higher complexity problem we would have to explore this value with only 90% fertility It is now dried out becomes arid deserts. So the reasoning ecosystem collapses now into a single mode and we only have a water pipeline down into one valley into one if you want solution for one particular problem. But imagine we modify now the problem then this known path will not provide any solution anymore. And this is what we say is a mode collapse of our reasoning system. Now what is the solution? The solution is by flow reinforcement learning system and this is now a distribution matcher. It is now not looking for the perfect for the perfect singular answer but it is looking at a complete variety of answers. So let's say it acts like a wise ecologist trying to create here a healthy diverse and resilient watershed. And the goal is now to ensure the amount of water reaching now each and every valley is perfectly proportional to its natural fertility. So to solution we found there down in the valley. So we follow the reasoning path. We learn the reasoning path. explore the reasoning path even if they're not 100% correct but also maybe 90% correct because the next task to this AI might be a high complexity level and then it has to find new pathways forward. So you might ask how does flow reinforcement learning achieves this without building now dams and just focusing on channeling everything down to one valley. Well the idea is and this is a little bit that's a little bit of challenge. It reshapes the three-dimensional mountain terrain itself because if we look down at the mathematical representation, it changes here our policy pi theta. Theta are the parameters and the weights and everything. So this is here if you want with the magic of the flow balance objective function comes in. This is now where here this new mathematical the new mathematical apparatus of the flow distribution come in into reinforcement learning. Now let's look at the core objective again through the lens here of this water flow. Now we can rewrite now some of the equation here into this very simple flow equation or to make it here even simpler here in my example now with the rain. Now so we have a term that is more or less the total rate. Now it's the water pouring down from the sky. This is here if we have a flowbased analysis this is the total rain the input. Then we have a lot of pathways with a particular flowliness. No, how fast the water goes down this particular reasoning path down to the valley down to the solution. And of course this has to be proportional somehow. Then at the end of the flow the target water that arrives in all the different valleys. So you see simple idea. Yeah. If you're not familiar here with this notation, a lot of people always ask me what are you talking about? What is a pya? is a probability of the model generated. If you have here a specific prompt X and you want to generate a specific answer Y. So you see this is here simply the strategy or the policy pi theta where theta are the free trainable parameters of the system. Now look at this from a different perspective here. The first term here the yeah it's of course log for simply we have just to sum and not multiplication the total rain it's a model estimate of the total amount of rainfall for the entire mountain range for the specific problem for the specific task and of course now the problem is how to determine this parameter more about this in a minute then I already told you here we have the strategy or the policy pi data it represents how easy it is for the water to flow down here a specific solution path y given the complexity the task complexity x. So a high probability path is like a deep smooth riverbed and that low probability is just a tiny trickle of water down to the valley. And here of course our or you guessed it this is the reward function. The target amount of water that should have should end up in the valley at the end of each path. Why? based on its specific fertility scala fertility efficients and you see it is a very simple equation. No. Yeah. Again this here why we write have here the log again I was asked here hey we can transfer this here instead of a chain of multiplication into a much more stable sum and here you have the sum. So this is just to have a simpler mathematical calculation. But how this complete balance act works now here if we have multiple valleys and multiple rivers downstream multiple augmentation path multiple complexities that we need to explore. Now we have more or less two option the valley is under irerrigated here by this stream of the rain. So what is happening? We have an equation that is imbalance. We have a very high loss function. To fix this, the optimizer will erode now the path leading to this valley coughing here a deeper channel by increasing here the policy probability pi theta more water now flows down there naturally because the fertility of this valley is high it just needs more water it just needs a better reasoning path to this particular solution or you have the exact opposite no value is over irrigated so again the equation is unbalanced the loss is also really high. So the optimizer will now try to let's say reduce here the flow by deposing depositing here some sediments into that particular downstream path making it harder for the water to flow and of course yeah this is a regularization terms and mathematical speech but I want to give you a simple example by decreasing here the policy probability pi theta great so you see okay it's just a control loop now by applying now this balancing if you want pressure all over again over every path everything that it samples. It doesn't just force here on a single unique solution, but it adjusts here. Think about the different valleys here. The entire landscape of the probabilities until a beautiful equilibrium is reached. The flow of all the probability mass naturally distributes itself across all reasoning paths into all the values that are fertile in a perfect proportion to the resulting rewards. This is the fertility scholar. So you see I try to explain a mathematical complexity in a very simple image. Now you know of course so you understand this pie here the policy or the strategy of our agent deciding now what is the next action to take this here is here the target water that is the valley. So more or less now we just have to focus here on the first term here on what is Z what is this set and let's have a look at it what is in my example what it is in a mathematical construct now we already identified set P here as the total amount of rainfall but there is of course a twist nobody knows what the total amount of rainfall will be and no chance at all if you know this in advance so if you want to See now the mountain range in its totality. It is of course the problem space that we have to look at and we have to explore. You remember explorability versus exploitability here in reinforcement learning. And now this problem space is vast and unknown and no idea. So you just can't put a giant measuring cup over it and just collect your old rain or measure this somehow. No. So you hire now and this is the solution here a learnable weather forecaster and this weather forecaster is nothing else here than our neural network zed set of pi here let's start the flow so at the beginning the forecaster has absolutely no idea we have to learn the system we have to train this system from scratch so no idea how much reward potential a problem holds so for a new mountain range a new prompt given by the human it might make Now just a guess it just start here. Now I said hey I predict a light drizzle today. So a small value for zed and then the system lets some water flow naturally down a few path. So the policy the given policy pi data generates here a few reasoning trajectories why and these are like exploratory probes now and these probes flow into different valleys and let's say you can measure then very simple here what is the effect of the water in the valleys you have a certain fertility and this of course is our reward function. So the results are now observed remember monof decision process partially observable monof decision process. So let's in our example those are observed the probes are sent back to the weather forecaster. So if we have an underestimation the forecaster predicted a light drizzle but the probes discover several incredible fertile valleys. No. So there's a high reward coming back says yeah if you give me more water into my valley I will have a a a crop production 100%. So there is now a mismatch of what it could be. So the system knows that the forecast was wrong. So there's far more fertility potential in the valleys than it was expected. So the loss function now the learning of the system is now in a way that the loss function penalizes now the forecaster for its low prediction forcing it to increase the estimate of zed. And guess what? If we have exactly the opposite an overestimation, the system also penalizes you the forecaster for being too optimistic and forcing it to decrease the estimate of zed. So you see we will have a loop that comes to an equilibrium. So the set network isn't just measuring the total rainfall. It is learning to predict like any ice system the total integrated fertality of the entire landscape given we have a particular flow from the rain over the complex mountain terrain into the multitude of valleys. So it learns this here by constantly comparing its global forecast against the local evidence guarded here by the individual water flows and the result of the water flows that make your crop prosper and blah beautiful. So let's come back and have a different perspective set is now the self-correcting estimate of the total reward budget of course which allows you the system now to make sense of the individual flows. So the system is learning a self-correcting self-learning system and a river flow is only meaningful when compared to the total amount of rain that feeds it. So let's have a look now since we have now this understanding here of a flow in and a flow out and a flow balance equation. What does it mean in scientific terms? It is nothing else than a partition function, a learnable partition function for the system. Now is the time that I tell you if you really want to have a deep dive, I would recommend this paper. This is my weekend paper. You have here 75 pages here. This is already from June 2023. Real interesting. The Gflow net foundations. Have a look at this. This is a detailed mathematical. Gives you everything you need to know. But we just want to focus here on the main idea, the main understanding before we have a mathematical deep dive. So what is the goal? The goal is to create a target probability distribution where the probability of a solution y is proportional to it exponentiated reward. So our reward function let's just say reward function forget the beta forget the exponential our reward. And to make this a valid probability distribution to make it to sum over one over all probability distribution you notice we must divide it now of course by normalization constant. And guess what? This normalization constant is our partition function Z. Now, isn't this beautiful? This is it. Yeah, absolutely. So, what is set? How can we calculate set? It is this idea simple. No, you go through every possible single reasoning path Y in the universe of all the solution of the set of Y. calculate ex its reward function here with a hyperparameter and exponentiated and everything and then you add them all up. Now there's a problem. Now if you have all the possible reasoning p this is unlimited almost no it's astronomically no so there's no way you can compute this no fundamental barrier and this was a problem and now the idea that flow reinforcement learning is and it kind of borrowed this from the mathematical framework of gflow nets is to say hey if we can't compute it let's learn to approximate it. So they create now a separate simple simpler neural network Z with your parameters and the sole job of this neural network is now to output a single scala number that is our best guess at the time for the value of the log offset. You see it's approximation wherever you look. Now there is now the next beautiful idea is you say how do we build this loss function and you want to say hey let's train this simultaneously so we train here the policy network our strategy of our AI system and we integrate this with the training here of set of our set neural network and we build this using the same loss function and here you have this loss function and it's a trajectory balance SLOs as you see which is the core objective function that flow RL seeks here to minimize. Look how beautiful it is here our expectation. Now of course this loss and you see here our Z function and our strategy and our reward function here this loss creates a beautiful self-consistent loop. So for every sample Y the loss implicitly tells Z what its value should have been to make the equation balance out. It's a flow equation now given the current policies probability PI theta and the observed reward R. So average now over many many loops and samples our Z is now pushed to converge to a value that is consistent with the reward distribution being explored by the policy by the strategy itself. This is the beauty. So again have a look at this loss function. There's an elegance to it. Now the loss function panalyzes here the model whenever the probability it assigns to a reasoning path is not proportional to the reward that the path achieves. Thereby teaching it to match the entire reward distribution. Remember the very beginning I showed you here that the coolback lava divergence here from flow RL is almost 0.11 extreme small this is the reason this is the main understanding what's happening here if you want to see it here in your initial flow your path probability and your terminal reward this is here what the system learns simultaneously so main insight that is here trackable learnable will stand in for an unsolvable global normaliz computational unsolvable global normalization constant. Yeah. And this if you want trick allows here flow reinforcement learning to frame the reinforcement learning exercise as a distribution matching problem without ever needing to compute here this impossible sum. turning here theoretical elegant but impractical idea into a really working algorithm that we can use and use now for reinforcement learning optimization. So this that is the key technique and if you read the paper yes of course it's in in a relation to the energy based toggle distribution have a look at it it's beautiful. So therefore, if you just want to remember one sentence, this is it. This new methodology of flow reinforcement learning cleverly uses now here from the separate mathematical body of Gflow net trajectory balance objectives as a stable and practical surrogate for what was currently impossible to calculate for the intractable Kubaklavic divergence minimization objective. So we found a way to come close to this. So you want the distribution of your model output your pi theta to be identical to the true reward distribution pi gilder. And of course you know how the mathematical tool to measure the difference is of the coolback lava divergence here. And the perfect objective is to minimize this value to zero. And with flow RL as I showed you at the very beginning we're real close 0.11. Now yeah as I showed you the Kubaklava divergence if you really want to do this is almost computational impossible to do. So therefore we have this beautiful as I showed you trajectory balance loss here with our terms that we went through and this is now the objective function that the fl minimizes here as a mathematical thing. Now look at this you will notice there is now a beautiful simplification that helps us immensely in the computation. This formula doesn't require summing over all possible y's all possible solution path at the same time. You do have a stable squared loss form rather than the complexity of a coolback liberal optimization formula. And and this is one of the most important formulas you will see in a complete um preprint here. Have a look at proposition one. This explains why you allowed to do this. You have here your minimization of the of the Kubak Leler and then you say why suddenly I can go to a trajectory balance function. Now they prove here in a mathematical formula that if you consistently minimize the local easy to calculate trajectory balance error from random samples y this here the gradients so this means the update direction for your model on a complex manifold are the same as the gradients you would have gotten if you would have magically with unlimited computer power calculated here the global impossible to calculate coolback lava divergence They say this is here for the gradients they're the same. This is not easy to understand. It is not easy to have to understand this on a mathematical level. But they provide some beautiful guidance and this is what I will enjoy this weekend. Okay. So here we are now. Flow reinforcement learning replaces here an impossible global comparison with an easy and efficient local balancing act that achieves you the same net result with our flow. Now of course you might say okay so what do we have now for reinforcement learning from human feedback and PO and everything and GRPO and DPO now we have flow reinforcement learning when to use what what is better what happens of course you find here in the paper yeah this here the last one is here the best outperforms 5 to 10% everything else but it is not that simple just let's have a look at this so I would formulate it in a different way the GRPO aims to maximize last year the scalar reward and its goal is to find the policy it produces here the pi data the highest possible score it's a peak climber it looks just for the highest mountain peaks no this is it dpo is different it maximizes here the lock likelihood of the human preferences no it is if you want a preference maximizer and now what is flow reinforcement learning it aims to match your target probability distribution derived from the rewards So its goal is for its output to reflect the entire landscape of all the values of good solution of all the good fertility scholars in proportion to their quality. So if you want it is a complete ecosystem manager. So you have a peak optimizer, a preference optimizer and then an ecosystem optimizer. And given that your task will increase in time and you will have higher complexities to solve an ecosystem manager might be the right way forward. So therefore I would now say if your task requires creativity kind of a robustness from the mathematical and computational side and finding a multiple valid solution for mathematics strategic planning or any scientific discovery where there's not only one path forward there are multiple solution path and you just have to find the right combinatorial multitude of those subpaths you know then flow reinforcement learning represents I think a better paradigm to find a solution Because why it addresses here the fundamental weakness we have in the other systems here mode collapse and it really is here mode collapse something that we have here in GRPO and everything that plagues here the maximization based methods on those complex domains especially if you go to mathematic or strategic planning or planning tool use and so on. So therefore I think this study here by Shanghai Geodong University, Shanghai AI laboratory, Microsoft research, Chinua University, Picking University, Renman University of China, Stanford University and Toyota Technological Institute at Chicago. This is a beautiful paper. You have to have a look at this paper. I think it's really a significant better solution for the next frontier of AI systems, especially for complex causal reasoning. I hope you enjoyed it. If you subscribe, I see you in my next video.

Original Description

NEW FlowRL, a reinforcement learning algorithm for enhancing LLM reasoning by shifting from traditional reward maximization (employed in methods like PPO and GRPO) to reward distribution matching via flow balancing, inspired by GFlowNets from 2021. The main new insight is that minimizing the reverse Kullback-Leibler divergence between the policy distribution π_θ(y|x) and a target reward-induced distribution exp(β r(x,y)) / Z_φ(x), where Z_φ(x) is a learnable partition function, promotes diverse exploration of reasoning trajectories, mitigating mode collapse and improving generalization in chain-of-thought tasks. This is achieved through a trajectory balance objective reformulated with length normalization (scaling log-probabilities by 1/|y| to address gradient explosion in sequences up to 8K tokens) and clipped importance sampling for off-policy stability. Flow RL yielding empirical gains of 10.0% over GRPO and 5.1% over PPO on average across six math benchmarks, alongside consistent improvements on three code benchmarks, as validated by diversity analyses of generated paths. 00:00 FlowRL ArXiv 01:34 GFlowNet explained 03:37 A simple Explanation of FlowRL 08:51 FlowRL compared to DPO, GRPO 11:16 The Solution 14:14 The core Objective 17:54 The Weather Forecaster Z 22:04 The Partition Function Z 26:30 Main Insight FlowRL All rights w/ authors: "FlowRL: Matching Reward Distributions for LLM Reasoning" Xuekai Zhu 1, Daixuan Cheng 6, Dinghuai Zhang 3, Hengli Li 5, Kaiyan Zhang 4, Che Jiang 4, Youbang Sun 4, Ermo Hua 4, Yuxin Zuo 4, Xingtai Lv 4, Qizheng Zhang 7, Lin Chen 1, Fanghao Shao 1, Bo Xue 1, Yunchong Song 1, Zhenjie Yang 1, Ganqu Cui 2, Ning Ding 4,2, Jianfeng Gao 3, Xiaodong Liu 3, Bowen Zhou 4,2‡, Hongyuan Mei 8, Zhouhan Lin 1,2 from 1 Shanghai Jiao Tong University 2 Shanghai AI Laboratory 3 Microsoft Research 4 Tsinghua University 5 Peking University 6 Renmin University of China 7 Stanford University 8 Toyota Technological Institute at Chicago
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Discover AI · Discover AI · 0 of 60

← Previous Next →
1 Step Into the Unknown (by YouChat) - May 2023 be your best year yet
Step Into the Unknown (by YouChat) - May 2023 be your best year yet
Discover AI
2 Wishing you all an amazing 2023 filled with Love, Laughter, and Happiness!
Wishing you all an amazing 2023 filled with Love, Laughter, and Happiness!
Discover AI
3 Create a Smarter Future!
Create a Smarter Future!
Discover AI
4 The Art of Text to Vector Transformation: A Comprehensive Look at AI and NLP Transformers
The Art of Text to Vector Transformation: A Comprehensive Look at AI and NLP Transformers
Discover AI
5 Feature Vectors: The Key to Unlocking the Power of BERT and SBERT Transformer Models
Feature Vectors: The Key to Unlocking the Power of BERT and SBERT Transformer Models
Discover AI
6 Domain-Specific AI Models: How to Create Customized BERT and SBERT Models for Your Business
Domain-Specific AI Models: How to Create Customized BERT and SBERT Models for Your Business
Discover AI
7 Achieve Unimaginable Levels of Domain Knowledge through SBERT Extreme in 3D   (SBERT 48)
Achieve Unimaginable Levels of Domain Knowledge through SBERT Extreme in 3D (SBERT 48)
Discover AI
8 Unlocking Scientific Domain Knowledge w/ BPE Tokenizer: An Amazing Journey!  (SBERT 49)
Unlocking Scientific Domain Knowledge w/ BPE Tokenizer: An Amazing Journey! (SBERT 49)
Discover AI
9 SBERT Extreme 3D: Train a BERT Tokenizer  on your (scientific) Domain Knowledge  (SBERT 50)
SBERT Extreme 3D: Train a BERT Tokenizer on your (scientific) Domain Knowledge (SBERT 50)
Discover AI
10 Discover Vision Transformer (ViT) Tech in 2023
Discover Vision Transformer (ViT) Tech in 2023
Discover AI
11 Pre-Train BERT from scratch: Solution for Company Domain Knowledge Data | PyTorch (SBERT 51)
Pre-Train BERT from scratch: Solution for Company Domain Knowledge Data | PyTorch (SBERT 51)
Discover AI
12 Flan-T5-XL model on a free COLAB | A free LLM - that explains itself w/ reasoning /write essay | AI
Flan-T5-XL model on a free COLAB | A free LLM - that explains itself w/ reasoning /write essay | AI
Discover AI
13 BERT and GPT in Language Models like ChatGPT or BLOOM |  EASY Tutorial on Large Language Models LLM
BERT and GPT in Language Models like ChatGPT or BLOOM | EASY Tutorial on Large Language Models LLM
Discover AI
14 Free Alternative to ChatGPT: Flan-T5-XL GUI (open-source)  #shorts
Free Alternative to ChatGPT: Flan-T5-XL GUI (open-source) #shorts
Discover AI
15 From T5 to T5X: A Game-Changing Evolution with JAX & FLAX
From T5 to T5X: A Game-Changing Evolution with JAX & FLAX
Discover AI
16 How to start with ChatGPT?  | Short Introduction to OpenAI API #shorts
How to start with ChatGPT? | Short Introduction to OpenAI API #shorts
Discover AI
17 The Future of Conversational AI? Google's PaLM w/ RLHF  | LLM ChatGPT Competitor
The Future of Conversational AI? Google's PaLM w/ RLHF | LLM ChatGPT Competitor
Discover AI
18 Microsoft and ChatGPU
Microsoft and ChatGPU
Discover AI
19 From Zero to FLAN-T5 XL Model GUI with Gradio: A Step-by-Step Guide on Free COLAB Notebook PyTorch
From Zero to FLAN-T5 XL Model GUI with Gradio: A Step-by-Step Guide on Free COLAB Notebook PyTorch
Discover AI
20 Google's 2nd Answer to "BING ChatGPT":  Sparrow | after BARD w/ LaMDA | 2nd Gen Conversational AI
Google's 2nd Answer to "BING ChatGPT": Sparrow | after BARD w/ LaMDA | 2nd Gen Conversational AI
Discover AI
21 TF2: Pre-Train BERT from scratch (a Transformer), fine-tune & run inference on text | KERAS NLP
TF2: Pre-Train BERT from scratch (a Transformer), fine-tune & run inference on text | KERAS NLP
Discover AI
22 3D Visualization for BERT: How to Pre-Train with a New Layer & Fine-Tune with Downstream Task Layer
3D Visualization for BERT: How to Pre-Train with a New Layer & Fine-Tune with Downstream Task Layer
Discover AI
23 FLAN-T5-XXL on NVIDIA A100 GPU w/ HF Inference Endpoints, let's explore 11b models!
FLAN-T5-XXL on NVIDIA A100 GPU w/ HF Inference Endpoints, let's explore 11b models!
Discover AI
24 ChatGPT - Can it Lie to you?
ChatGPT - Can it Lie to you?
Discover AI
25 ChatGPT Alternative: Perplexity by Perplexity.AI
ChatGPT Alternative: Perplexity by Perplexity.AI
Discover AI
26 2023 KerasNLP Tutorial: Explore Latest KERAS Toolbox & NLP Processing Library for BERT - TF2
2023 KerasNLP Tutorial: Explore Latest KERAS Toolbox & NLP Processing Library for BERT - TF2
Discover AI
27 Self-aware AI: You.com/chat vs Perplexity.ai | Live Demo, LLMs show Future of ChatGPT w/ BING
Self-aware AI: You.com/chat vs Perplexity.ai | Live Demo, LLMs show Future of ChatGPT w/ BING
Discover AI
28 BLOOM 176B Inference on AWS  | Bigger than GPT-3 for more Power!
BLOOM 176B Inference on AWS | Bigger than GPT-3 for more Power!
Discover AI
29 Fine-tune ChatGPT? Buy Embeddings /OpenAI? What are Embeddings?  My own ChatGPT? | Visual Q+A
Fine-tune ChatGPT? Buy Embeddings /OpenAI? What are Embeddings? My own ChatGPT? | Visual Q+A
Discover AI
30 Unleashing the Power of BLOOM 176B with AWS ml.p4de.24xlarge, DJL & DeepSpeed: The Ultimate Boost!
Unleashing the Power of BLOOM 176B with AWS ml.p4de.24xlarge, DJL & DeepSpeed: The Ultimate Boost!
Discover AI
31 After ChatGPT: NEW BioGPT by Microsoft | Do YOU trust Microsoft for your Medication?
After ChatGPT: NEW BioGPT by Microsoft | Do YOU trust Microsoft for your Medication?
Discover AI
32 Improve ChatGPT: Modular, Adaptive, Smart LLM | Inside ChatGPT
Improve ChatGPT: Modular, Adaptive, Smart LLM | Inside ChatGPT
Discover AI
33 Fine-tune ChatGPT w/  in-context learning ICL - Chain of Thought, AMA, reasoning & acting: ReAct
Fine-tune ChatGPT w/ in-context learning ICL - Chain of Thought, AMA, reasoning & acting: ReAct
Discover AI
34 The Intersection of Copyright Law and Human Faces: Exploring Virtual K-Pop with MAVE
The Intersection of Copyright Law and Human Faces: Exploring Virtual K-Pop with MAVE
Discover AI
35 New TECH: Vision Transformer 2023 on Image Classification | AI
New TECH: Vision Transformer 2023 on Image Classification | AI
Discover AI
36 PyTorch code Vision Transformer: Apply ViT models pre-trained and fine-tuned  | AI  Tech
PyTorch code Vision Transformer: Apply ViT models pre-trained and fine-tuned | AI Tech
Discover AI
37 New BING ChatGPT: Unlock the Power of Emotions in your Search Engine!
New BING ChatGPT: Unlock the Power of Emotions in your Search Engine!
Discover AI
38 New BING ChatGPT loses its mind
New BING ChatGPT loses its mind
Discover AI
39 Self-Attention Heads of last Layer of Vision Transformer (ViT) visualized (pre-trained with DINO)
Self-Attention Heads of last Layer of Vision Transformer (ViT) visualized (pre-trained with DINO)
Discover AI
40 Visualizing the Self-Attention Head of the Last Layer in DINO ViT: A Unique Perspective on Vision AI
Visualizing the Self-Attention Head of the Last Layer in DINO ViT: A Unique Perspective on Vision AI
Discover AI
41 Microsoft strongly restricts access to ChatGPT on new BING - WHY?
Microsoft strongly restricts access to ChatGPT on new BING - WHY?
Discover AI
42 PyTorch ViT: The Ultimate Guide to Fine-Tuning for Object Identification (COLAB)
PyTorch ViT: The Ultimate Guide to Fine-Tuning for Object Identification (COLAB)
Discover AI
43 New BING Chat AGGRESSIVE
New BING Chat AGGRESSIVE
Discover AI
44 Panoptic Image Segmentation: Mask2Former explained | Identify all objects!
Panoptic Image Segmentation: Mask2Former explained | Identify all objects!
Discover AI
45 Code Panoptic Image Segmentation w/ Vision Transformer & Mask2Former - A PyTorch tutorial
Code Panoptic Image Segmentation w/ Vision Transformer & Mask2Former - A PyTorch tutorial
Discover AI
46 Dream Job Alert: AI Prompt Engineer - $335K  |  AI Prompt Design: A Crash Course
Dream Job Alert: AI Prompt Engineer - $335K | AI Prompt Design: A Crash Course
Discover AI
47 Streamlining Similar Image Detection with ViT in PyTorch: A Step-by-Step Guide
Streamlining Similar Image Detection with ViT in PyTorch: A Step-by-Step Guide
Discover AI
48 Microsoft's CEO in Trouble   #shorts
Microsoft's CEO in Trouble #shorts
Discover AI
49 Why wait for KOSMOS-1? Code a VISION - LLM w/ ViT, Flan-T5 LLM and BLIP-2: Multimodal LLMs (MLLM)
Why wait for KOSMOS-1? Code a VISION - LLM w/ ViT, Flan-T5 LLM and BLIP-2: Multimodal LLMs (MLLM)
Discover AI
50 OpenAI's ChatGPT can NOW summarize external Sources on the Internet?
OpenAI's ChatGPT can NOW summarize external Sources on the Internet?
Discover AI
51 ChatGPT polarizes
ChatGPT polarizes
Discover AI
52 Hospital /Clinic AI Decision Models: Performance of 12 AI LLM Systems (incl $$) Radiology, Biomed
Hospital /Clinic AI Decision Models: Performance of 12 AI LLM Systems (incl $$) Radiology, Biomed
Discover AI
53 ChatGPT Prompt Engineering w/ in-context learning (ICL)  - 7 Examples | Tutorial
ChatGPT Prompt Engineering w/ in-context learning (ICL) - 7 Examples | Tutorial
Discover AI
54 Chat with your Image!  BLIP-2 connects Q-Former w/ VISION-LANGUAGE models (ViT & T5 LLM)
Chat with your Image! BLIP-2 connects Q-Former w/ VISION-LANGUAGE models (ViT & T5 LLM)
Discover AI
55 ChatGPT:  Multidimensional Prompts
ChatGPT: Multidimensional Prompts
Discover AI
56 ChatGPT:  In-context Retrieval-Augmented Learning (IC-RALM) | In-context Learning (ICL) Examples
ChatGPT: In-context Retrieval-Augmented Learning (IC-RALM) | In-context Learning (ICL) Examples
Discover AI
57 Code your BLIP-2 APP: VISION Transformer (ViT) + Chat LLM (Flan-T5) = MLLM
Code your BLIP-2 APP: VISION Transformer (ViT) + Chat LLM (Flan-T5) = MLLM
Discover AI
58 Buy Microsoft "Azure OpenAI Service" or buy from OpenAI its API for ChatGPT access & tuning?
Buy Microsoft "Azure OpenAI Service" or buy from OpenAI its API for ChatGPT access & tuning?
Discover AI
59 Pretraining vs Fine-tuning vs In-context Learning of LLM (GPT-x) EXPLAINED | Ultimate Guide ($)
Pretraining vs Fine-tuning vs In-context Learning of LLM (GPT-x) EXPLAINED | Ultimate Guide ($)
Discover AI
60 Reversible Transformer: ReFORMER for GPU Memory Optimization! Reversible Residual Layers?
Reversible Transformer: ReFORMER for GPU Memory Optimization! Reversible Residual Layers?
Discover AI

The video teaches the concept of FlowRL, a new reinforcement learning method that uses GFlowNets to enhance LLM reasoning. It discusses the limitations of traditional reward maximization methods and how FlowRL addresses mode collapse. The video is useful for researchers and practitioners interested in reinforcement learning and complex causal reasoning.

Key Takeaways
  1. Read the research paper on FlowRL
  2. Implement the FlowRL algorithm
  3. Understand the concept of GFlowNets
  4. Apply FlowRL to complex domains
  5. Analyze the effectiveness of FlowRL
  6. Reproduce the results of the FlowRL paper
💡 FlowRL addresses mode collapse, a fundamental weakness in maximization-based methods, by using a flow balance objective function to optimize reinforcement learning.

Related Reads

📰
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
Learn how to improve off-policy reinforcement learning with auxiliary branches, enhancing reasoning in large language models
ArXiv cs.AI
📰
A Practical Guide to Implementing the REINFORCE Algorithm in Python (Part 5)
Implement the REINFORCE algorithm in Python using PyTorch and Gymnasium for reinforcement learning tasks
Medium · Machine Learning
📰
Gimitest: A Comprehensive Tool for Testing Reinforcement Learning Policies
Learn how to test reinforcement learning policies with Gimitest, a comprehensive tool for ensuring reliability and safety
ArXiv cs.AI
📰
RLVP: Penalize the Path, Reward the Outcome
Learn how to implement RLVP, a new reinforcement learning approach that prioritizes outcome over path, and apply it to real-world problems with costly interactions
ArXiv cs.AI

Chapters (9)

FlowRL ArXiv
1:34 GFlowNet explained
3:37 A simple Explanation of FlowRL
8:51 FlowRL compared to DPO, GRPO
11:16 The Solution
14:14 The core Objective
17:54 The Weather Forecaster Z
22:04 The Partition Function Z
26:30 Main Insight FlowRL
Up next
How Netflix Uses Reinforcement Learning to Recommend Movies #ai #coding #machinelearning #netflix
Ascent
Watch →