Mind Evolution: Deeper Thinking at Inference (by Google)

Discover AI · Advanced ·🧠 Large Language Models ·1y ago

Key Takeaways

Presents Google's Mind Evolution research on evolving deeper LLM thinking during Test-Time-Compute

Full Transcript

hello Community let's have a look at the latest research by ghoul called Mind Evolution and yeah it's a snowstorm outside I'm sitting here in the Austrian Alps it's beautiful snow all over it's below zero so just stay inside right next to the fireplace and now I got some complaints that one of my last videos were a little bit too complicated so today is Story Time relax sit right next to your fireplace and here we go I have multiple llms that sege the internet and I get some EI generated summaries and this is one of the llms that I employ and look at this what I got today mind Evolution a groundbreaking innovation in intelligent problem solving with genetic algorithms and refinement through critical conversation it is transforms problem solving unparalleled resulting Benchmark yeah the power of Mind Evolution and but what nonsense is this and you know I realized that a simple command like write the technical summary of the attached PDF can generate quite some nonsense even with the best available llms so let's explore this where we are today and I ask J gbd for Omni here for example here hey tell me in scientific terms what are the genetic operators in the latest I research of Google because for me also working here in biotechnology in genetics a genetic operator for me in my world is something completely different to whatever was chosen here in this paper interestingly even jgpt stays within the paper given my system Proms and all my examples it stays within this notation and the description here and gives me just here not a summary but just here you know some sentences that are just taken out from the PDF and then I was a little bit angry and I said okay now describe this simple genetic operators in non-technical terms for let's say a prompt optimization and finally finally it happened that system came back and said okay selection is imagine you have several versions of a prompt you picked the best one this is selection and I said wow and then crossover takes the best part of two different prompts and combines them to a new prompt and I said amazing mutation introduced small changes to a prompt to see if it gets better result and I was blown away here by the marketing effort that Google put here into its scientific paper so you see our llms more and more even if I say hey scientific terms explain it to me are more and more related here to the marketing slack and even if I ask those systems and I have a lot of roadblocks that I do not get this marketing BS it is more and more getting that those elements are just shining through and there's hardly any scientific explanation here in this okay so science let's talk about what is happening here we here not into training we are here definitely at inference time so Google wants here also like 01 or 03 models by openi to optimize the performance in test time computer of our llm by scaling here the inference time computation the goal is a deeper sying process of our llm when we wait 1 2 3 5 minutes maybe half an hour an hour for the answer and it is now implemented by exploring here an evolutionary marketing search strategy so let's have a look at is the main question was hey how can an llm be guided to think deeper about a real complex problem and leverage inference time compute to improve here its own problem solving abilities and the idea was simple if we have a solution evaluator for our search process the search strategy that we're going to apply have an advantage of being able to rely to improve here the problem solving ability with an increased computer so the more inference time compute we offer to the llm the llm will provide better solution because it has more time for some elaborated search strategies and an inherent solution evaluator that guide its way here to the best available solution at a given time and I've shown you one in my last video here the best of end solution and now Google goes the next step into development so they have now the idea we propos as Google an evolutionary search strategy that combines here the free flowing stochastic exploration and the large scale iterative refinement process and I say of what what are the objects explain this to me give me more details anyway they call this now the Mind evolution in marketing terms and oh great okay what is the idea idea I think is that this mind Evolution methodology is not restricted to searching now in a formal space you know with r we have a vector space or we have some embedding or some other mathematical space we construct here but now we do all of this here simply by optimizing solution in the space of our natural language and you might say that's that's strange no we built agents that have memory that have function calling to our tools so we have calculators we have python environments we have all the computer simulation here whatever we need to have those tools available and now we go back to the space of natural language to have a deep thinking reasoning progress yeps this is the way we are going and here we have the paper the title is evolving deeper llm thinking published today for me recording this 20th of January 2025 and it is worth having a look at this because there are some surprises hidden so they tell us hey it is a genetic search strategy and me and my simple mind I was still locked to the genetic biotechnology because over there we also have genetic search strategy but on complete different topics so whatever you use here in in an EI publication terms that are from a different part of science please don't do this because you know every specific Tech IAL term has its own environment so they say genetic s strategy yes yes yes and feedback from an evaluator great and again Divergent thinking convergent thinking they Hallmark of intelligent problem solving behavior and you see what what is now creeping into these scientific Publications but you know it becomes clear if you look here simple at the task and the task is here you talk to your Al and say Hey you plan to visit five European cities for 16 days in total you only take di flag to community between the cities you spend five days in Madrid from day three to 7 an will show you want to do this and this and you say now they collected here different methodology here the one pass with the best of end methodology and then finally they show here the last line here that their solution is the best that can solve this particular problem of course it is a handpicked problem but okay we end understand for this particular task this is the best methodology now if I look now at this from the official publication by Google of Mind Evolution as a genetic based evolutionary search strategy I'm a little bit confused because this is rather simple now I mean a sample solution beautiful so we just ask the LM to come up with different plans for day one day two and then we have a feedback from an evaluator what whatever this evaluator might be and the feedback is fed back and so then we have now an evaluation some things are better some things are worse we have a feedback Loof of a refinement or improv the previous solution now with the genetic operator of selection cross over and mutation which I call just prompt engineering but never mind and then we have than one maybe the best solution or one of the best solution given we have a restricted maximum compute time budget or financial budget and looking at this I might say tell me where is the innovation in this this is something I'm kind of familiar with this now and I thought maybe it is here in the evaluation function so in principle any function tells us Google that can evaluate the solution quality can be used including a pure llm evaluation by the llm itself of its solution generated and they call it of course in the mind Evolution scheme they call this the fitness function great so scoring solutions by measuring here the optimization objective or verify whether the solution satisfy the given constraints and provide corresponding textual feedback so we have a simple feedback loop and then it dawned on me this whole mind evolution is just here a natural language planning exercise that is all there is yes it is also a little bit about reasoning but we are here focusing in my simple explanation on natural language planning let's have a look at this of course we started the game here with a population initialization in mind Evolution so we have something here from from population Dynamics from genetics we apply this now to an EI term and they say here give them a Target Target problem we dependently sample now in an initial Solution by prompting an llm with a description of the problem any information needed for solving the problem and some relevant instruction so this is simply a prompt that we have here llm and I want just to sample here I don't know five or 10 times a given llm with my problem I don't need a population initialization but okay and then each of these initial Solutions are then evaluated here and refined through additional turns here of the refinement through critical conversation process and I said what the hell is the this process so I just looked at an example here the refinement through critical conversation or RCC and you know in the llm summarization all these Bus words were there without any explanation but you know what it is you have a task you have a first initial solution then you evaluate here the return you have maybe a critique llm that says hey was given that in Tokyo it should be 5 days instead of the three days that are given here by the first initial solution so therefore we have to refine this so just have a feedback loop but now this is called an RCC methodology okay Google we have an RCC methodology but I guess you and I we understand what we're talking about but but then then it happened to me because here looking now in detail look at this particular table I understand that I did not get it because there was something hidden in the terms that was really there let me show you what I mean so we have here engine the maximum number of generation to search for a solution the many independent populations we want to evolve how many conversation per Island and how many terms per conversation and at this time I asked myself what are exactly Islands here for this and I started to search and I found here 25 years ago in the year 2000 there was a survey of parallel genetic algorithms from University of Illinois here the genetic algorithms laboratory and I'm sorry but what happened 25 years in it I I was not familiar with this term so those genetic algorithms were something okay great and then reading this understood here that this island model is more or less a strategy where the overall population is divided into multiple subpopulations and we refer to those subpopulation as Islands so each island evolves independently applying genetic operations like selection Cross of mutation within its own subpopulation and periodically individuals migrate between between the island producing new genetic materials so I think this is from Darin right I know going back to the old Millennium okay and I understood this was a term a technical term that is now used but have given me complete different Vibes a complete different visual environment because we do have genetic algorithms today in 2025 but they are definitely not those so you see G allowing different subpopulations to explore various reason so I would say today so what we have we have different prompts we get different replies by the llm we can cluster this replies that we get from an llm into thematic topics thematic clusters those are the islands and then I can go maybe find here the focus point of a cluster and then maybe have here an edge functionality to another cluster and the cluster allowed to exchange information about the most important fact or the most important methodology that they apply for the solution so you see interesting you can find something complete different with a different kind of language that you use if you use those terms interesting so in the context of mind Evolution whatever this means an island model is now employed here by Google and this creates now multiple groups of prompts those group of prompts are another the island and then the prompt from one Island from one cluster is shared with the other cluster and this is finally the solution to all of this marketing speech because if I just look here what the llm gives me as a summary of this paper it would lead me in a complete different direction so current performance of llms on this topic horrible great coming back to the main topic where I told you hey we have agent we have tools we have everything that we need to have complex computer simulation and now we go back here to real world real natural language and we think we can solve the problems of the world here without any computer modulation without any calculation numerical calculation itical function using the human language is enough to find all solution and then I found that they have here Google itself has here also referring here to the travel planner this is a particular Benchmark for real world planning you see again we are talking only about planning not so much about really executing here the reasoning of language agent this is October 2024 fan University Ohio State University Pennsylvania state univers was in in meta but really then came here Google deepat has its own natural plan Benchmark that it used here from June 2024 benchmarking llms on natural language planning so now this makes it really Crystal Clear if you read a complete paper and you follow here all the references what they doing they are looking here at a real small Spectrum only what is a natural language planning exercise and in my understanding only for this natural language planning exercise they provide here a new methodology with a feedback mechanism and they call this here something from generic operators for genetic operators and yes God knows what so unfortunately it takes quite some time but you see a publication is not simply equal to any other publication and I noticed this in the last weeks and months here all this publication to get attention here from the community they start to invent scientific terms invent here complete notation for I don't know for what reason we do know how to communicate this but now everything has to be brand new everything has to have a new name and everything has to have an unbel believable marketing methodology nomenclature so my goodness so all these papers so all these new research is for natural language planning exercise without any tool use and they have here this particular genetic operators prompt engineering with a feedback loop and this is the content here of this L paper by Google and you see this is the beauty that outside is a snowstorm I can not go anywhere else I'm stuck here in my little cabin here and I just enjoy here the winter if you are living here on the sou hemisphere of our planet oh gee I'm jealous if I think here of the people here in Australia down under maybe at Bondi Beach enjoying here surfing oh wow I'm here in a snowstorm and I'm reading here Google Publications okay so you see even from such a paper there are some insight and I just wanted to share with you be careful do not trust here those automated summarization by llms because they can get things horribly wrong or impressed by the marketing slogan that were maybe generated by another llm and then we end up with this paper which could be completely Rewritten with a clear Focus but yeah I think this is the time where we living so therefore enjoy your snowstorm if you have one enjoy your fireplace and you see today's video was only about storytelling so if you want to subscribe and get notified with my next video we will focus a little bit more on a scientific topic

Original Description

The latest AI research by Google regarding AI deep reasoning (Jan 20, 2025). Called Mind Evolution, evolving deeper LLM Thinking during Test-Time-Compute (TTC). All rights w/ authors: Evolving Deeper LLM Thinking (Mind Evolution) Kuang-Huei Lee, Ian Fischer, Yueh-Hua Wu, Dave Marwood, Shumeet Baluja, Dale Schuurmans and Xinyun Chen from Google DeepMind, UC San Diego, University of Alberta #science #reasoningskills #airesearch #mind
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Discover AI · Discover AI · 0 of 60

← Previous Next →
1 Step Into the Unknown (by YouChat) - May 2023 be your best year yet
Step Into the Unknown (by YouChat) - May 2023 be your best year yet
Discover AI
2 Wishing you all an amazing 2023 filled with Love, Laughter, and Happiness!
Wishing you all an amazing 2023 filled with Love, Laughter, and Happiness!
Discover AI
3 Create a Smarter Future!
Create a Smarter Future!
Discover AI
4 The Art of Text to Vector Transformation: A Comprehensive Look at AI and NLP Transformers
The Art of Text to Vector Transformation: A Comprehensive Look at AI and NLP Transformers
Discover AI
5 Feature Vectors: The Key to Unlocking the Power of BERT and SBERT Transformer Models
Feature Vectors: The Key to Unlocking the Power of BERT and SBERT Transformer Models
Discover AI
6 Domain-Specific AI Models: How to Create Customized BERT and SBERT Models for Your Business
Domain-Specific AI Models: How to Create Customized BERT and SBERT Models for Your Business
Discover AI
7 Achieve Unimaginable Levels of Domain Knowledge through SBERT Extreme in 3D   (SBERT 48)
Achieve Unimaginable Levels of Domain Knowledge through SBERT Extreme in 3D (SBERT 48)
Discover AI
8 Unlocking Scientific Domain Knowledge w/ BPE Tokenizer: An Amazing Journey!  (SBERT 49)
Unlocking Scientific Domain Knowledge w/ BPE Tokenizer: An Amazing Journey! (SBERT 49)
Discover AI
9 SBERT Extreme 3D: Train a BERT Tokenizer  on your (scientific) Domain Knowledge  (SBERT 50)
SBERT Extreme 3D: Train a BERT Tokenizer on your (scientific) Domain Knowledge (SBERT 50)
Discover AI
10 Discover Vision Transformer (ViT) Tech in 2023
Discover Vision Transformer (ViT) Tech in 2023
Discover AI
11 Pre-Train BERT from scratch: Solution for Company Domain Knowledge Data | PyTorch (SBERT 51)
Pre-Train BERT from scratch: Solution for Company Domain Knowledge Data | PyTorch (SBERT 51)
Discover AI
12 Flan-T5-XL model on a free COLAB | A free LLM - that explains itself w/ reasoning /write essay | AI
Flan-T5-XL model on a free COLAB | A free LLM - that explains itself w/ reasoning /write essay | AI
Discover AI
13 BERT and GPT in Language Models like ChatGPT or BLOOM |  EASY Tutorial on Large Language Models LLM
BERT and GPT in Language Models like ChatGPT or BLOOM | EASY Tutorial on Large Language Models LLM
Discover AI
14 Free Alternative to ChatGPT: Flan-T5-XL GUI (open-source)  #shorts
Free Alternative to ChatGPT: Flan-T5-XL GUI (open-source) #shorts
Discover AI
15 From T5 to T5X: A Game-Changing Evolution with JAX & FLAX
From T5 to T5X: A Game-Changing Evolution with JAX & FLAX
Discover AI
16 How to start with ChatGPT?  | Short Introduction to OpenAI API #shorts
How to start with ChatGPT? | Short Introduction to OpenAI API #shorts
Discover AI
17 The Future of Conversational AI? Google's PaLM w/ RLHF  | LLM ChatGPT Competitor
The Future of Conversational AI? Google's PaLM w/ RLHF | LLM ChatGPT Competitor
Discover AI
18 Microsoft and ChatGPU
Microsoft and ChatGPU
Discover AI
19 From Zero to FLAN-T5 XL Model GUI with Gradio: A Step-by-Step Guide on Free COLAB Notebook PyTorch
From Zero to FLAN-T5 XL Model GUI with Gradio: A Step-by-Step Guide on Free COLAB Notebook PyTorch
Discover AI
20 Google's 2nd Answer to "BING ChatGPT":  Sparrow | after BARD w/ LaMDA | 2nd Gen Conversational AI
Google's 2nd Answer to "BING ChatGPT": Sparrow | after BARD w/ LaMDA | 2nd Gen Conversational AI
Discover AI
21 TF2: Pre-Train BERT from scratch (a Transformer), fine-tune & run inference on text | KERAS NLP
TF2: Pre-Train BERT from scratch (a Transformer), fine-tune & run inference on text | KERAS NLP
Discover AI
22 3D Visualization for BERT: How to Pre-Train with a New Layer & Fine-Tune with Downstream Task Layer
3D Visualization for BERT: How to Pre-Train with a New Layer & Fine-Tune with Downstream Task Layer
Discover AI
23 FLAN-T5-XXL on NVIDIA A100 GPU w/ HF Inference Endpoints, let's explore 11b models!
FLAN-T5-XXL on NVIDIA A100 GPU w/ HF Inference Endpoints, let's explore 11b models!
Discover AI
24 ChatGPT - Can it Lie to you?
ChatGPT - Can it Lie to you?
Discover AI
25 ChatGPT Alternative: Perplexity by Perplexity.AI
ChatGPT Alternative: Perplexity by Perplexity.AI
Discover AI
26 2023 KerasNLP Tutorial: Explore Latest KERAS Toolbox & NLP Processing Library for BERT - TF2
2023 KerasNLP Tutorial: Explore Latest KERAS Toolbox & NLP Processing Library for BERT - TF2
Discover AI
27 Self-aware AI: You.com/chat vs Perplexity.ai | Live Demo, LLMs show Future of ChatGPT w/ BING
Self-aware AI: You.com/chat vs Perplexity.ai | Live Demo, LLMs show Future of ChatGPT w/ BING
Discover AI
28 BLOOM 176B Inference on AWS  | Bigger than GPT-3 for more Power!
BLOOM 176B Inference on AWS | Bigger than GPT-3 for more Power!
Discover AI
29 Fine-tune ChatGPT? Buy Embeddings /OpenAI? What are Embeddings?  My own ChatGPT? | Visual Q+A
Fine-tune ChatGPT? Buy Embeddings /OpenAI? What are Embeddings? My own ChatGPT? | Visual Q+A
Discover AI
30 Unleashing the Power of BLOOM 176B with AWS ml.p4de.24xlarge, DJL & DeepSpeed: The Ultimate Boost!
Unleashing the Power of BLOOM 176B with AWS ml.p4de.24xlarge, DJL & DeepSpeed: The Ultimate Boost!
Discover AI
31 After ChatGPT: NEW BioGPT by Microsoft | Do YOU trust Microsoft for your Medication?
After ChatGPT: NEW BioGPT by Microsoft | Do YOU trust Microsoft for your Medication?
Discover AI
32 Improve ChatGPT: Modular, Adaptive, Smart LLM | Inside ChatGPT
Improve ChatGPT: Modular, Adaptive, Smart LLM | Inside ChatGPT
Discover AI
33 Fine-tune ChatGPT w/  in-context learning ICL - Chain of Thought, AMA, reasoning & acting: ReAct
Fine-tune ChatGPT w/ in-context learning ICL - Chain of Thought, AMA, reasoning & acting: ReAct
Discover AI
34 The Intersection of Copyright Law and Human Faces: Exploring Virtual K-Pop with MAVE
The Intersection of Copyright Law and Human Faces: Exploring Virtual K-Pop with MAVE
Discover AI
35 New TECH: Vision Transformer 2023 on Image Classification | AI
New TECH: Vision Transformer 2023 on Image Classification | AI
Discover AI
36 PyTorch code Vision Transformer: Apply ViT models pre-trained and fine-tuned  | AI  Tech
PyTorch code Vision Transformer: Apply ViT models pre-trained and fine-tuned | AI Tech
Discover AI
37 New BING ChatGPT: Unlock the Power of Emotions in your Search Engine!
New BING ChatGPT: Unlock the Power of Emotions in your Search Engine!
Discover AI
38 New BING ChatGPT loses its mind
New BING ChatGPT loses its mind
Discover AI
39 Self-Attention Heads of last Layer of Vision Transformer (ViT) visualized (pre-trained with DINO)
Self-Attention Heads of last Layer of Vision Transformer (ViT) visualized (pre-trained with DINO)
Discover AI
40 Visualizing the Self-Attention Head of the Last Layer in DINO ViT: A Unique Perspective on Vision AI
Visualizing the Self-Attention Head of the Last Layer in DINO ViT: A Unique Perspective on Vision AI
Discover AI
41 Microsoft strongly restricts access to ChatGPT on new BING - WHY?
Microsoft strongly restricts access to ChatGPT on new BING - WHY?
Discover AI
42 PyTorch ViT: The Ultimate Guide to Fine-Tuning for Object Identification (COLAB)
PyTorch ViT: The Ultimate Guide to Fine-Tuning for Object Identification (COLAB)
Discover AI
43 New BING Chat AGGRESSIVE
New BING Chat AGGRESSIVE
Discover AI
44 Panoptic Image Segmentation: Mask2Former explained | Identify all objects!
Panoptic Image Segmentation: Mask2Former explained | Identify all objects!
Discover AI
45 Code Panoptic Image Segmentation w/ Vision Transformer & Mask2Former - A PyTorch tutorial
Code Panoptic Image Segmentation w/ Vision Transformer & Mask2Former - A PyTorch tutorial
Discover AI
46 Dream Job Alert: AI Prompt Engineer - $335K  |  AI Prompt Design: A Crash Course
Dream Job Alert: AI Prompt Engineer - $335K | AI Prompt Design: A Crash Course
Discover AI
47 Streamlining Similar Image Detection with ViT in PyTorch: A Step-by-Step Guide
Streamlining Similar Image Detection with ViT in PyTorch: A Step-by-Step Guide
Discover AI
48 Microsoft's CEO in Trouble   #shorts
Microsoft's CEO in Trouble #shorts
Discover AI
49 Why wait for KOSMOS-1? Code a VISION - LLM w/ ViT, Flan-T5 LLM and BLIP-2: Multimodal LLMs (MLLM)
Why wait for KOSMOS-1? Code a VISION - LLM w/ ViT, Flan-T5 LLM and BLIP-2: Multimodal LLMs (MLLM)
Discover AI
50 OpenAI's ChatGPT can NOW summarize external Sources on the Internet?
OpenAI's ChatGPT can NOW summarize external Sources on the Internet?
Discover AI
51 ChatGPT polarizes
ChatGPT polarizes
Discover AI
52 Hospital /Clinic AI Decision Models: Performance of 12 AI LLM Systems (incl $$) Radiology, Biomed
Hospital /Clinic AI Decision Models: Performance of 12 AI LLM Systems (incl $$) Radiology, Biomed
Discover AI
53 ChatGPT Prompt Engineering w/ in-context learning (ICL)  - 7 Examples | Tutorial
ChatGPT Prompt Engineering w/ in-context learning (ICL) - 7 Examples | Tutorial
Discover AI
54 Chat with your Image!  BLIP-2 connects Q-Former w/ VISION-LANGUAGE models (ViT & T5 LLM)
Chat with your Image! BLIP-2 connects Q-Former w/ VISION-LANGUAGE models (ViT & T5 LLM)
Discover AI
55 ChatGPT:  Multidimensional Prompts
ChatGPT: Multidimensional Prompts
Discover AI
56 ChatGPT:  In-context Retrieval-Augmented Learning (IC-RALM) | In-context Learning (ICL) Examples
ChatGPT: In-context Retrieval-Augmented Learning (IC-RALM) | In-context Learning (ICL) Examples
Discover AI
57 Code your BLIP-2 APP: VISION Transformer (ViT) + Chat LLM (Flan-T5) = MLLM
Code your BLIP-2 APP: VISION Transformer (ViT) + Chat LLM (Flan-T5) = MLLM
Discover AI
58 Buy Microsoft "Azure OpenAI Service" or buy from OpenAI its API for ChatGPT access & tuning?
Buy Microsoft "Azure OpenAI Service" or buy from OpenAI its API for ChatGPT access & tuning?
Discover AI
59 Pretraining vs Fine-tuning vs In-context Learning of LLM (GPT-x) EXPLAINED | Ultimate Guide ($)
Pretraining vs Fine-tuning vs In-context Learning of LLM (GPT-x) EXPLAINED | Ultimate Guide ($)
Discover AI
60 Reversible Transformer: ReFORMER for GPU Memory Optimization! Reversible Residual Layers?
Reversible Transformer: ReFORMER for GPU Memory Optimization! Reversible Residual Layers?
Discover AI

Related Reads

Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →