True AI Reasoning: Graph-Based CPT
Key Takeaways
The video discusses Graph-Based CPT for true AI reasoning, utilizing GraphPile, a large-scale corpus for continue-pretraining, and integrating graph problems into LLM pre-training to enhance general reasoning abilities across multiple domains. It explores the application of graph-based CPT in various areas, including logic, common sense, code, and graph problems, and evaluates the performance of different models, such as Graph Mind and LLaMA, in these tasks.
Full Transcript
Hello community. So great that you are back. Yes, we have another grand idea in the eye research. We use the graph as a universal teacher for LLMs and we optimize LLMs today. Now you know the LLM world is a sub symbolic world. So we have a vast high dimensional space defined by billions of numerical tensor weights and biases and there's no node variable like we have in a graph or in a knowledge graph. Instead all the concepts in a large language model here in a transform architecture are represented as vectors embeddings and knowledge is not stored in a database. It is encoded in the geometric relationship exactly between those vectors. Now you know llams can have beautiful tool calling abilities. Huh? This is a traditional program that is symbolic now and it has variables like node equal 5 and executes hard-coded rules like this one. So you see with our tools, with our code programs, with our Python environments, with our C++, the logic is explicit and transparent. Yes, of course, if you just add a little bit of memory, you have an agent. But let's come back to the tool. We have with the tool a numerical computer simulation for some theoretical physics experiment in C++. So this tool now calculates here a numerical result given here a specific I don't know astrophysical phenomena. Now the way that this calculation was performed in code has an inherent reasoning structure. Somebody thought about how to code it before they started here to calculate the red shift of galaxies. So you see in this code hidden is a flow of logic and patterns. And now when the LLM is pre-trained on a new form of data set that would integrate these inherent reasoning structures from other languages like a code, it is now forced to map the symbolic world of logic of a C++ onto its own sub symbolic world. And this is where we can enjoy something beautiful because we add new reasoning primitives to the reasoning capacity of an LLM. And how we do this simply with a continued pre-training. So you know with continued pre-training we have another problem highly specialized knowledge but it lacks transferability. This means even a mathematical exquisite model is not necessarily a better logical reasoning for causal reasoning and I would have assumed hey Matt and causal reasoning so close no in the model in the AI model they failed to show the performance. So what we are creating is really narrow intelligence. But how do we foster here a more broader a more general reasoning ability? And yes, you guessed it. Today's new research paper is exactly about how to open up to general reasoning abilities. And they say if we could identify a class of problems so rich in foundational logic terms that mastering them would enhance the reasoning capabilities of our large language model across all domains without relying on tools. You see this is a second class of thinking and they say tools are great but we want to integrate this into the reasoning capabilities of an LLM. So what they decided the authors of today's study they said we need a diverse set of reasoning skill that go beyond mathematics. We need skills in logical topological computational and enumerative and I will show you all of this in a minute. So those skills they say are now more foundational more on a deeper level of an understanding than required here some simple traditional mathematical skills or text problem skills for some text bots. So we are going deeper and deeper into the rabbit hole of reasoning and intelligence of a large linguistic language model. Now here this is one of the drawings here of the beautiful paper and you have here the topological reasoning clearly in a cluster with a logical reasoning and Hamiltonian p topological sort maximum click we're going to talk about this in a second and you see if they argue now if we find here the the smallest denominator if we fall here the core of all of this thinking of this reasoning this would now boost here the reasoning performance on all other domains topbological relationships. Yeah, let's talk about it. So, this is a brand new idea. They say integrating now the graph problems that we have on a knowledge graph for example now into the pre-training process of LLMs. And please notice this is not supervised fine-tuning. This is not reinforcement learning. We're going back and we go here to the first step, the pre-training of LLMs. The authors say, "We aim to unlock a powerful tool for enhancing general reasoning abilities across multiple domains." So, CPT continued pre-training here is the tool, the methodology that we're going to have a look at for the green grasshoppers, for all newcomers. What is the difference between a CPT and SFT and on a reinforcement learning? This is it. Continual pre-training. Just think about general knowledge and adapt to new domains. And we do this by large unlabeled text corpora. It's like reading a new set of British anticolas here to expand your world view. So basing is much more specialized for a specific task. We talk about here training for a downstream task to follow instruction. Normally we do this with labeled examples like a prompt response pair but this is here very limited. You learn only one new skill and reinforcement learning. Okay, let's go classical 2017 with open eye reinforcement learning from human feedback. Here this is just here an alignment and a specific orientation to human values and you have a teacher providing you the guidance telling you the model what is good or bad and what is the expected behavior of those models. So we are now going here to the real expansive pre-training but we are not doing the pre-training. We start with some open-source models some small language model and we will do a continual pre-training and this is a lot of operational problems. Let's stick here to the theoretical idea. So continual pre-training expanding the model's core knowledge base much like the initial pre-training phase but we built on something that Microsoft did or Google did or openi did it's absolute in contrast to fine reinforcement learning GRPO DPO whatever you know about shaping here the model's behavior no so we are not only trying to improve the mathematical reasoning but also other form of complex reasoning in multiple domains such as algorithmic problem solving and pure logical reasoning without relying here on specific tools like proofreader like lean for or anything else. So the authors seek here to bridge the gap between the domain specific pre-training and the development of a more universally capable reasoning model by integrating our complete new reasoning traces here. And they say this would be the holy grail of artificial intelligence. And here we have the study. Those are the Hong Kong University of Science Technology. A beautiful study. Read it. Published July 23, 2025. improving LLM's generalized reasoning abilities by graph problems. We bring now here the complete understanding that we have from graph theory from mathematical graph theory that is if you want a symbolic theory into a subsyolic representation of an LLM quite a lot of fun let's start we needed a training data set and they built here a new data set a first data set for continued pre-training using here the graph problem reasoning data they call their training said you remember in the history we had the pile now we have a graph pile and approximately 11 billion tokens and if you want to see this here we have a chain of sword with 2.8 billion token real world graph yes you got it so we have over 2.6 million samples and remember yes pre-training is we need a lot a lot I mean really a lot of training data 2.6 6 million is okay on the lower side. It is simple. It is everything that we know. Chain of sort, program of sort, trace of execution. There's nothing specific to it. But let's have a look. Chain of sword. You know that this is teaching here the LLM the Y. So we synthesize here a stepbystep natural linguistic language explanation for solving now graph problems. And you have your beautiful example here from the literature. So let's test the connectivity. You have here a question and an answer. And now the LLM learns here this graph theoretical problem. And here the if you want the linguistic semantic connectivity here of our tokenized words here in a large language model. It is like learning here another code language. You know that's not so difficult. We know this works great although it has massive limitation. So teaches the LLM to develop you and articulate a systematic reasoning process. Not really deep but it works. The second is here the program of sort. This teaches you the how. This leverages here on LLM to generate precise executable code solution for growth problems. So if you have your code solution, you just learn this more or less. No, it's just a sequency of tokens. The third one is a real world grounding from our abstract theoretical base. We go now to concrete examples that we use here. Knowledge graph you might know or social network if you are into mathematical graph theory. The mold's ability to reason about complex real practical scenarios like here find the shortest path here for if you travel from different airports to your particular um goal in you want to go to Paris. This is all known. This is all there. And finally, this is the new element if you want the trace of execution. This is now something that is a little bit more interesting because it's analyzing here the graph algorithmic code and tries to find here the flow. It tries to find here a linear sequence, a flow in an argumentation why the code was written in a particular way. So the model learns to predict intermediate variable states of the system. So it tries to find your argumentation why the code was written in the way and not in any other different way. So this here they try here to to capture here the dynamics of the reasoning process. A much deeper skill than just a pattern matching. If you go here to GitHub you just copy and say this is the pattern. They want to understand the internal dynamic of this pattern for the reasoning process. This here is really an interesting innovative element. Let's have a look at a little bit more detailed. I know you want more logical reasoning task simple. All graph problems are fundamentally based on reasoning derived from the logical rules. Thereby they inherently task of logical reasoning. So two examples you want to find here. You give them a graph. You check if there it contains a cycle. This is the definition of a cycle or other elements and you apply specific logical rules. No, a cycle exists in a graph if a path from a vertex revisits here the same vertex. So those rules are simple. Topological wisiting task are a little bit more challenging. Exploring here the relationship between the node and the edges in the graph and making inferences based on those relationship. We have everything that we know about mathematical topology and topological reasoning in the simplest case transfer here to a linguistic model and we have the example of topological sorting common neighbors. This is everything that you know you know here topological sorting here reveals the hierarchical relationship indirected as cyclic graphs. Common neighbors highlight local connection between those nodes. More interesting are the enumeration tasks. This is not because we sometimes we have to list all possible solution, all possible configuration, everything that is in a mathematical combinatorical way a possible theoretical solution to my given query. Everything that goes with a hum with a optimization procedure or everything that is here a search a guided search or just a trial and error search. So this is here interesting. If we could analyze this in more detail and have examples, thousands and thousands of examples to find here enough training data for the LLM to learn this reasoning complexity. And this is for example the Hamiltonian P for the maximum click problem. So they did this they found here 2.6 6 million examples and then they said okay now we have the training data set guess what we do the continuous pre-training this is a complexity in itself I will neglect this step because I will preparing here a video here on this particular topic how is the best way to continue pre-training your own model but they went here and yeah pre-training is very expensive with very small models from two billion free trainable parameters to 8 billion free trainable parameters and they did three version of the reasoning model and now they call their new model with the new training data set in a continued pre-training mode graph mind. Here we have it and at first they look at mathematics. Yeah. So three models we have a gemma 2 2 billion really minimal configuration a llama 38b open source and a llama 3.18p open source and you see here for all the mathematical benchmark that you can wish for and here in the end we have an average and you have here the model itself as you get it if you take it here from the shelf or if you do your discontinued pre-training with the data set graph pile And here you have this and you see there's a a huge amount here where really graph pile outperforms here the normal standard vanilla llama 3.1 8b. Notice that in the 2B there are significant jumps sometimes from 39 to 41 but in 8b it's almost the same. So the effect in mathematics is not really massive. three, four, five percentage points. But look at this. Look at logic. Look at common sense. Look at code. Look at graph problems. Now we see something moving from 16 to 62. From 33 to 75. Yeah, there are some let's say outliers from 3 to 50. But you see now now we have an effect that we say okay given the new training data set given the new complex learning patterns that the LLM learns now in a continuous pre-training it shows an effect not really on mathematics but on everything else that is close to mathematics logic common sense code graph so we have a much better performant performance gain in out of domain tasks Now they claim here, yeah, this graph might outperforms your older base, your tiny two billion base models, up to 5% in mathematic reasoning, and then 22% in other domains. Sounds reasonable. They gave you here for graphs here for classical benchmark. You see, hm okay just look at this. No, gamma 2 at 2 billion is 12.4%. graph pile is just a little bit better but let's say four. If we go now to a more modern llama 3.18 be a little bit bigger a little bit warn you see the increase here is not really there no or here 71.7 to 73.0 zero. There's a little bit sometimes missing, but graph risk looks interesting. Look here at this performance jump. So again, you could ask, hey, wait a minute. Is this really a general ability for an improved reasoning across all domains or is it rather here domains that are rather close here to the original domain that the pre-training data encompassed here? No. If you look at the evaluation detail of the graph reasoning data set for graph with and graph instruct, you can find that there are some discontinuities. No, look here at all. Yeah, here this I explained to you graph connectivity flow shortest path topological sorting. We went through this your maximum flow comma niba page rank. Here you have all the benchmarks. You see sometimes graph pile is much better but sometimes yeah the classical model is better which is an effect I do not currently understand how this could happen. So there is a not really absolute consistency in this benchmark data. Although here if you look at this Alama 3.18B yeah you would say yeah okay this looks really good especially if you go here from 7% to 93%. You would say okay here it really shines. No here the graph instruct benchmark really shows a significant improvement in the benchmark data. Quite a lot of my viewers ask hey can you include here the temperature sensitivity? Great. So here we have a llama 38B in orange and in green we have now this super continued pre-trained graph mind 8B which is based on a llama 38B. You see here in general there is say some parallel development. However just look here this cursor clo directly tests you the very algorithmic reasoning skill. This is one of the main core reasoning skills that we want to improve here for a general reasoning improvement the algorithmic part. But if you go a temperature 0.0 zero. Yeah, you see. Wow. Look here. Graph mind heal close to 50% compared to I don't know 2% from the llama sweep. But the moment you increase the temperature, it completely goes away. This effect I cannot explain this. You have a lot of numerical data and test data in this report. Have a look at this report. especially the annex. There are tons and tons of numerical data and I've chosen this one to tell you I think there are still some discontinuities here that I for example I cannot explain where this performance jump is happening. Something seems not to be right yet here for the absolute 100% performance jumps if you're interested is of course here by Google. Great. So what we found reasoning ability is a very narrowly transferable skill. But there is now a shift from a domain specific continuous pre-training as a learning methodology now including here higher complexity reasoning pattern continuous pre-training methods. This is absolutely fascinating. If you think in a transformer layer structure or if you think the tensor weights, if you think about the complexity of the tensor weights and their hyper plane orientation here in I don't know a 10,000 dimensional space, they seem to open up new areas, new new quarters here of this vector space for higher complexity reasoning patterns where now here the vector embedding points to this high complexity objects in a vector space. I will do a specific video on this because I think this is an absolutely fascinating new development. Yeah. And the quest is on to identify other universal curricular to improve yet the general reasoning. Maybe we can include more of the formal logic in a linguistic form for the LLMs or more understanding here in the semantic formulation of physics and the the laws of physics to enable here an LLM a better reasoning performance or you might say hey why we have the tools we have everything that we need from C++ to lean for to whatever we don't need to try to squeeze is more of pattern recognition of higher complexity reasoning pattern into the LLM. We have the tools for this. This is the reason why we have agents. They can connect to the world. They can interact with the world. They can use other comput resources. Yes, but the answer is generated by the LLM. I understand also the other group that says, "Yeah, great to have tools, great that we have access to all the computer programming and whatever to PPR, but imagine how good it would be if we could further increase the reasoning capability of the LLM of the large language model in the semantic complexity for the reasoning process in general when we get back the results from the tools. So this is an absolute fascinating way to go forward and I suppose both has its benefits. Yeah, bonus for you. If you made it to this time in the video, you know, just 5 hours ago here I posted here on my YouTube channel here a simple questioner. Is this statement true? So the intelligence of an LLM is just the ability to configure higher complexity patterns into learn vector space representation that now offer a complex solution path for advanced graph reasoning and of course I did this after I read the paper and I was just interested how you see this. So after 5 hours in we have here you think here either you have read here as a scientific community already the paper and you know the statement is true because you know the latest AI research so 65% of you said yes I know the paper I know that you tried to trick us here into an incorrect statement great but you know what I went to GPT I mean I didn't log in nothing this is here the free version this is here everybody has access to no and I put in the same statement Interestingly, Chad GPD said, "Hey, this is false." No, because Yes. Yes. Yes. And I said, "Read it again. It says that intelligence is the ability to configure high complexity patterns into learned vector space representations." This is a true statement. And if you have seen one of my last videos, you know that small large language models are very sensitive to anchors that are positioned at the end of the statement. And you know what suddenly Jimid said hey you are absolutely right in pointing out that to a significant extent here intelligent in an M can be understood as my things why the statement is true now what a surprise so whatever you do with this great but it came back and said you know there are still some nuances here the second part of my statement and then I explained the second part of my statement and then I asked again so then my statement is true. No. After if you want judge read also this new publication came back and said your refined statement where explain what is now the content here of this video. Your refined statement is supported by current research and can be considered true and at this point I thought chat TPT you have no idea what you're talking about absolutely no idea. And you can modify and polarize any statement you want from a AI system independent of the true, independent of the fact, independent of data. And yeah, and suddenly of course my statement is true and why your statement holds no in hybrid context. And now Jet GPT explains to me why I'm correct. It's kind of funny, but it shows you the current state of AI. And yes, this is another video I would like to make. A lot of my viewers or subscribers hopefully ask me, "Hey, can you explain the complete state of AI in a 10 minutes video?" Yeah, maybe next weekend if I have time. Yeah, it's a challenge. Sounds interesting. If you want to see more, subscribe and I see you in the next video.
Original Description
CPT For Complex Graph Reasoning injected into LLM.
True AI Reasoning: Graph-Based CPT.
New research introduces GraphPile, the first large-scale corpus specifically designed for CPT (continue-pretraining) using GPR (Graph Problem Reasoning) data.
Using GraphPile, the authors train GraphMind on three popular base models-Llama 3&3.1 and Gemma 2-achieving up to 4.9% higher accuracy in mathematical reasoning and up to 21.2% improvement in non-mathematical reasoning tasks, like logical and commonsense reasoning.
All rights w/ authors:
"Improving LLMs’ Generalized Reasoning Abilities by Graph
Problems"
Qifan Zhang, Nuo Chen, Zehua Li, Miao Peng, Jing Tang, Jia Li
from
The Hong Kong University of Science and Technology (Guangzhou)
#airesearch
#reasoning
#knowledgegraph
#knowledge
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from Discover AI · Discover AI · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Step Into the Unknown (by YouChat) - May 2023 be your best year yet
Discover AI
Wishing you all an amazing 2023 filled with Love, Laughter, and Happiness!
Discover AI
Create a Smarter Future!
Discover AI
The Art of Text to Vector Transformation: A Comprehensive Look at AI and NLP Transformers
Discover AI
Feature Vectors: The Key to Unlocking the Power of BERT and SBERT Transformer Models
Discover AI
Domain-Specific AI Models: How to Create Customized BERT and SBERT Models for Your Business
Discover AI
Achieve Unimaginable Levels of Domain Knowledge through SBERT Extreme in 3D (SBERT 48)
Discover AI
Unlocking Scientific Domain Knowledge w/ BPE Tokenizer: An Amazing Journey! (SBERT 49)
Discover AI
SBERT Extreme 3D: Train a BERT Tokenizer on your (scientific) Domain Knowledge (SBERT 50)
Discover AI
Discover Vision Transformer (ViT) Tech in 2023
Discover AI
Pre-Train BERT from scratch: Solution for Company Domain Knowledge Data | PyTorch (SBERT 51)
Discover AI
Flan-T5-XL model on a free COLAB | A free LLM - that explains itself w/ reasoning /write essay | AI
Discover AI
BERT and GPT in Language Models like ChatGPT or BLOOM | EASY Tutorial on Large Language Models LLM
Discover AI
Free Alternative to ChatGPT: Flan-T5-XL GUI (open-source) #shorts
Discover AI
From T5 to T5X: A Game-Changing Evolution with JAX & FLAX
Discover AI
How to start with ChatGPT? | Short Introduction to OpenAI API #shorts
Discover AI
The Future of Conversational AI? Google's PaLM w/ RLHF | LLM ChatGPT Competitor
Discover AI
Microsoft and ChatGPU
Discover AI
From Zero to FLAN-T5 XL Model GUI with Gradio: A Step-by-Step Guide on Free COLAB Notebook PyTorch
Discover AI
Google's 2nd Answer to "BING ChatGPT": Sparrow | after BARD w/ LaMDA | 2nd Gen Conversational AI
Discover AI
TF2: Pre-Train BERT from scratch (a Transformer), fine-tune & run inference on text | KERAS NLP
Discover AI
3D Visualization for BERT: How to Pre-Train with a New Layer & Fine-Tune with Downstream Task Layer
Discover AI
FLAN-T5-XXL on NVIDIA A100 GPU w/ HF Inference Endpoints, let's explore 11b models!
Discover AI
ChatGPT - Can it Lie to you?
Discover AI
ChatGPT Alternative: Perplexity by Perplexity.AI
Discover AI
2023 KerasNLP Tutorial: Explore Latest KERAS Toolbox & NLP Processing Library for BERT - TF2
Discover AI
Self-aware AI: You.com/chat vs Perplexity.ai | Live Demo, LLMs show Future of ChatGPT w/ BING
Discover AI
BLOOM 176B Inference on AWS | Bigger than GPT-3 for more Power!
Discover AI
Fine-tune ChatGPT? Buy Embeddings /OpenAI? What are Embeddings? My own ChatGPT? | Visual Q+A
Discover AI
Unleashing the Power of BLOOM 176B with AWS ml.p4de.24xlarge, DJL & DeepSpeed: The Ultimate Boost!
Discover AI
After ChatGPT: NEW BioGPT by Microsoft | Do YOU trust Microsoft for your Medication?
Discover AI
Improve ChatGPT: Modular, Adaptive, Smart LLM | Inside ChatGPT
Discover AI
Fine-tune ChatGPT w/ in-context learning ICL - Chain of Thought, AMA, reasoning & acting: ReAct
Discover AI
The Intersection of Copyright Law and Human Faces: Exploring Virtual K-Pop with MAVE
Discover AI
New TECH: Vision Transformer 2023 on Image Classification | AI
Discover AI
PyTorch code Vision Transformer: Apply ViT models pre-trained and fine-tuned | AI Tech
Discover AI
New BING ChatGPT: Unlock the Power of Emotions in your Search Engine!
Discover AI
New BING ChatGPT loses its mind
Discover AI
Self-Attention Heads of last Layer of Vision Transformer (ViT) visualized (pre-trained with DINO)
Discover AI
Visualizing the Self-Attention Head of the Last Layer in DINO ViT: A Unique Perspective on Vision AI
Discover AI
Microsoft strongly restricts access to ChatGPT on new BING - WHY?
Discover AI
PyTorch ViT: The Ultimate Guide to Fine-Tuning for Object Identification (COLAB)
Discover AI
New BING Chat AGGRESSIVE
Discover AI
Panoptic Image Segmentation: Mask2Former explained | Identify all objects!
Discover AI
Code Panoptic Image Segmentation w/ Vision Transformer & Mask2Former - A PyTorch tutorial
Discover AI
Dream Job Alert: AI Prompt Engineer - $335K | AI Prompt Design: A Crash Course
Discover AI
Streamlining Similar Image Detection with ViT in PyTorch: A Step-by-Step Guide
Discover AI
Microsoft's CEO in Trouble #shorts
Discover AI
Why wait for KOSMOS-1? Code a VISION - LLM w/ ViT, Flan-T5 LLM and BLIP-2: Multimodal LLMs (MLLM)
Discover AI
OpenAI's ChatGPT can NOW summarize external Sources on the Internet?
Discover AI
ChatGPT polarizes
Discover AI
Hospital /Clinic AI Decision Models: Performance of 12 AI LLM Systems (incl $$) Radiology, Biomed
Discover AI
ChatGPT Prompt Engineering w/ in-context learning (ICL) - 7 Examples | Tutorial
Discover AI
Chat with your Image! BLIP-2 connects Q-Former w/ VISION-LANGUAGE models (ViT & T5 LLM)
Discover AI
ChatGPT: Multidimensional Prompts
Discover AI
ChatGPT: In-context Retrieval-Augmented Learning (IC-RALM) | In-context Learning (ICL) Examples
Discover AI
Code your BLIP-2 APP: VISION Transformer (ViT) + Chat LLM (Flan-T5) = MLLM
Discover AI
Buy Microsoft "Azure OpenAI Service" or buy from OpenAI its API for ChatGPT access & tuning?
Discover AI
Pretraining vs Fine-tuning vs In-context Learning of LLM (GPT-x) EXPLAINED | Ultimate Guide ($)
Discover AI
Reversible Transformer: ReFORMER for GPU Memory Optimization! Reversible Residual Layers?
Discover AI
More on: LLM Foundations
View skill →Related Reads
📰
📰
📰
📰
Building Your First Model Context Protocol (MCP) Server with TypeScript and Zod
Dev.to AI
Beyond Code Generation: How OmniSVG Rethinks Vector Graphics with Vision-Language Models
Dev.to · Shrijith Venkatramana
LangGraph Checkpointing: Three Production Rewrites to Stop Losing State
Dev.to · Elena Revicheva
AI builder essentials: tokens, context windows and RAG 101
Dev.to · Tilde A. Thurium
🎓
Tutor Explanation
DeepCamp AI