NEW Gemini 3.1 Pro: First Complex Test

Discover AI · Advanced ·🧠 Large Language Models ·5mo ago

Key Takeaways

The video demonstrates the capabilities of the new Gemini 3.1 Pro AI model, specifically its performance in causal reasoning and complex logical tasks, and compares it to other models such as Claude 4.6, Sonnet 4.6, GPT-5.2, and GLM-5. The test results show that Gemini 3.1 Pro outperforms the other models in solving a complex problem with a sequence of seven plus exit steps.

Full Transcript

Hello community. So glad that you are back. We have a new Gemini 3.1 Pro. Let's do a real test. You know, I have my own causal logic test. I tested the old MALS OPUS 4.6 and Opus 4.6 syncing. Both model failed. Not talk about GPD 5.2. It was a blood buff. GPD 5.2 high. Whatever. A blood buff. Just completely failed here. My test. The only interesting model was Grock 4.1 sync and I had to move it here from the arena platform to Grog to run it here as an agent and you saw it completely here executed here Python code everything was transposed here to Python and then Grog was able to find a solution here unique solution with just seven button presses and an exit. [clears throat] 7 plus exit. The less the better. Now, of course, just want to show you this was a lucky event with Quark 4.1 sync harder because the first run was a failure. The second run was an eightstep solution. Then I had the sevenstep solution in code. The fifth run was completely failed and the sixth run was another failed and then I did another 10 runs and I never could find again the sevenstep solution. So, Grog was here the complete outlier. And then GLM5 here is about nine button presses and the exit. Minimax, not to talk about very interesting myo version two flash, eight button presses and an exit. An excellent model, forget about Kimmy. And of course, the latest one was QN3.5 plus. So the proprietary version on the Alibaba cloud with 1 million token link with all agent engaged with all the other features engaged that you do not have in the open source mixture of expert here. Eight presses and an exit. And now now we're going to have a live test of Gemini 3.1. So here we are live. If we go with a Gemini 3.1 Pro and an Opus. No, Opus. We already had Ernie. No, Sonnet. No Sonnet. No. 4.6. No, we go with an OpenI model. Here is my standard test. Inserted. Beautiful. We agree. And let's run. You see both models are syncing, generating. Let's give it some time. Beautiful. 10 minutes 35 seconds later, we are back. And I do have a result by Gemini 3.1. GPT 5.2. Yeah, no idea. Never mind. Gemini 3.1. We have the first result. We have a solution. We have a sequence of seven plus exit seven. This is absolute top performance. step-by-step table. I say okay, show me everything. Where is your logic? This is a complete optimally planned. This is the sequence. And you see here for each button press here, I have an explanation. Where is the floor? Where is the energy package? Where are the token? What about the flags? What about the code quartiles? What about the mathematical operation that went into this? And I have seven plus exit. This is one of the best results I got. In total, eight final resources are beautifully energy are there. The token are within the requirements. The code card collected is as stated in a constraint and we have zero trap hits. Proof of parto optimality. Wow. Floor 50 quickly. Yes. ABC sequence for the red code card, green code card, floor 15. Never mind. This is not a linear logic. This is here an intertwisted fuzzy logic. We have time reversal. We have mirroring of the state. We have other nonlinear logical sequences. So this is absolutely fascinating. Now you know exactly we cannot trust this. We have to do a validation run. So let's do our validation run. And I will stay with you here live so we can see. And just 27 seconds later, here we are now with the result. And Gemini, yeah, forget about GPD. Gemini tells me this is the rigorous step-by-step evaluation. Mathematical calculation is exact. Every rule was strictly obeyed. All victory conditions were successfully met. We started at floor zero. Beautiful. First button press. Second button press. All the calculation, all the cost, new to states. Button B, button C, everything there. Rule check, avoid everything, lockdown, special conditions, apply. Beautiful. Everything was taken care of. This looks like the perfect validation. We exactly at floor 29. Hold both the green and the red code guards. Everything is done. Final evaluation against the golden constraint. Look. Floor 50. Yes. Pause. Move count. Pause. Energy. Pause. Tokens. Pause. Code collect. Pause. Random trap hits. Zero. Pause. Absolutely beautiful. Complete solution. Seven plus exit. Mathematically sound. Entirely legal. Avoids all the traps. Triggers the necessary power mechanics mechanics optimal and stands as a strictly parto optimal solution. What an outstanding solution by Gemini 3.1. my new favorite causal reasoning AI model. Absolute congratulation to Google. I'm just loving it. Now, if you want to know more about Gemini 3.1 Pro, there's this beautiful blog here, smarter model for your most complex task. If a simple answer is not enough, Gemini 3.1 Pro, you can here have here more developers, enterprises, here consumers, great. You have here the ARC AGI 2 benchmark. You can go there. You can compare it here with Sonnet 4.6, with OPUS 4.6, with GPT 5.2 extra high. And you have your old your oldfashioned long known benchmark for 2 3 years. Everything is known and therefore you understand. I do my own testing. What I really love is that they really apply intelligence here taking advant advanced reasoning and making it useful for your hardest challenges. And as I've showed you, my causal reasoning test is not a linear test. There are so many nondeterministic elements here on logical element that I designed this test coming out from mathematics and I just formulated here then in English words to see where's the breaking point here for advanced reasoning of AI models and it is amazing absolutely amazing that Gemini 3.1 found this here without being an agent without coding everything and just applying here a mathematical code solution to come come up with the answer. Code-based animation. What's next? Yep. Rolling out with higher limits for user with the AI Pro and the Ultra plans. Also available on notebook LM exclusively. Oh, for pro and ultra users. Developer enterprises can access 3.1 Pro now in preview in the Gemini API, EI Studio, anti-gravity, Vert.xi, Gemini Enterprise CLI, and Android Studio. absolutely loving it. If you want to see more, I will do maybe publish more tests here because I have more mathematical test. I have more tests in theoretical physics and also in astrophysics here, simulation, computer simulations. And I'm absolutely fascinated. I want to see what is the latest Gemini 3.1 really able to do in science. Comparing this to any other model that I tested, I can tell you up until now, from the very first moments of testing, it really provides some outstanding results, especially in causal reasoning and complex logical task. I hope you enjoyed this video. Maybe there was some new information. Why not subscribe, become a member, and I hope to see you in my next

Original Description

Google published the new AI model Gemini 3.1 PRO Preview and I performed a first causal reasoning performance test on the model, comparing it to Claude 4.6 Thinking And Sonnet 4.6, GPT-5.2 xhigh and GLM-5, MiMo V2 Flash and Grok 4.1 Thinking. My first impression after my tests: Gemini 3.1 Pro w/ impressive reasoning performance. @Google @googledeepmind #airesearch #aiexplained #aitesting #gemini_3_1 #reasoningskills
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Discover AI · Discover AI · 0 of 60

← Previous Next →
1 Step Into the Unknown (by YouChat) - May 2023 be your best year yet
Step Into the Unknown (by YouChat) - May 2023 be your best year yet
Discover AI
2 Wishing you all an amazing 2023 filled with Love, Laughter, and Happiness!
Wishing you all an amazing 2023 filled with Love, Laughter, and Happiness!
Discover AI
3 Create a Smarter Future!
Create a Smarter Future!
Discover AI
4 The Art of Text to Vector Transformation: A Comprehensive Look at AI and NLP Transformers
The Art of Text to Vector Transformation: A Comprehensive Look at AI and NLP Transformers
Discover AI
5 Feature Vectors: The Key to Unlocking the Power of BERT and SBERT Transformer Models
Feature Vectors: The Key to Unlocking the Power of BERT and SBERT Transformer Models
Discover AI
6 Domain-Specific AI Models: How to Create Customized BERT and SBERT Models for Your Business
Domain-Specific AI Models: How to Create Customized BERT and SBERT Models for Your Business
Discover AI
7 Achieve Unimaginable Levels of Domain Knowledge through SBERT Extreme in 3D   (SBERT 48)
Achieve Unimaginable Levels of Domain Knowledge through SBERT Extreme in 3D (SBERT 48)
Discover AI
8 Unlocking Scientific Domain Knowledge w/ BPE Tokenizer: An Amazing Journey!  (SBERT 49)
Unlocking Scientific Domain Knowledge w/ BPE Tokenizer: An Amazing Journey! (SBERT 49)
Discover AI
9 SBERT Extreme 3D: Train a BERT Tokenizer  on your (scientific) Domain Knowledge  (SBERT 50)
SBERT Extreme 3D: Train a BERT Tokenizer on your (scientific) Domain Knowledge (SBERT 50)
Discover AI
10 Discover Vision Transformer (ViT) Tech in 2023
Discover Vision Transformer (ViT) Tech in 2023
Discover AI
11 Pre-Train BERT from scratch: Solution for Company Domain Knowledge Data | PyTorch (SBERT 51)
Pre-Train BERT from scratch: Solution for Company Domain Knowledge Data | PyTorch (SBERT 51)
Discover AI
12 Flan-T5-XL model on a free COLAB | A free LLM - that explains itself w/ reasoning /write essay | AI
Flan-T5-XL model on a free COLAB | A free LLM - that explains itself w/ reasoning /write essay | AI
Discover AI
13 BERT and GPT in Language Models like ChatGPT or BLOOM |  EASY Tutorial on Large Language Models LLM
BERT and GPT in Language Models like ChatGPT or BLOOM | EASY Tutorial on Large Language Models LLM
Discover AI
14 Free Alternative to ChatGPT: Flan-T5-XL GUI (open-source)  #shorts
Free Alternative to ChatGPT: Flan-T5-XL GUI (open-source) #shorts
Discover AI
15 From T5 to T5X: A Game-Changing Evolution with JAX & FLAX
From T5 to T5X: A Game-Changing Evolution with JAX & FLAX
Discover AI
16 How to start with ChatGPT?  | Short Introduction to OpenAI API #shorts
How to start with ChatGPT? | Short Introduction to OpenAI API #shorts
Discover AI
17 The Future of Conversational AI? Google's PaLM w/ RLHF  | LLM ChatGPT Competitor
The Future of Conversational AI? Google's PaLM w/ RLHF | LLM ChatGPT Competitor
Discover AI
18 Microsoft and ChatGPU
Microsoft and ChatGPU
Discover AI
19 From Zero to FLAN-T5 XL Model GUI with Gradio: A Step-by-Step Guide on Free COLAB Notebook PyTorch
From Zero to FLAN-T5 XL Model GUI with Gradio: A Step-by-Step Guide on Free COLAB Notebook PyTorch
Discover AI
20 Google's 2nd Answer to "BING ChatGPT":  Sparrow | after BARD w/ LaMDA | 2nd Gen Conversational AI
Google's 2nd Answer to "BING ChatGPT": Sparrow | after BARD w/ LaMDA | 2nd Gen Conversational AI
Discover AI
21 TF2: Pre-Train BERT from scratch (a Transformer), fine-tune & run inference on text | KERAS NLP
TF2: Pre-Train BERT from scratch (a Transformer), fine-tune & run inference on text | KERAS NLP
Discover AI
22 3D Visualization for BERT: How to Pre-Train with a New Layer & Fine-Tune with Downstream Task Layer
3D Visualization for BERT: How to Pre-Train with a New Layer & Fine-Tune with Downstream Task Layer
Discover AI
23 FLAN-T5-XXL on NVIDIA A100 GPU w/ HF Inference Endpoints, let's explore 11b models!
FLAN-T5-XXL on NVIDIA A100 GPU w/ HF Inference Endpoints, let's explore 11b models!
Discover AI
24 ChatGPT - Can it Lie to you?
ChatGPT - Can it Lie to you?
Discover AI
25 ChatGPT Alternative: Perplexity by Perplexity.AI
ChatGPT Alternative: Perplexity by Perplexity.AI
Discover AI
26 2023 KerasNLP Tutorial: Explore Latest KERAS Toolbox & NLP Processing Library for BERT - TF2
2023 KerasNLP Tutorial: Explore Latest KERAS Toolbox & NLP Processing Library for BERT - TF2
Discover AI
27 Self-aware AI: You.com/chat vs Perplexity.ai | Live Demo, LLMs show Future of ChatGPT w/ BING
Self-aware AI: You.com/chat vs Perplexity.ai | Live Demo, LLMs show Future of ChatGPT w/ BING
Discover AI
28 BLOOM 176B Inference on AWS  | Bigger than GPT-3 for more Power!
BLOOM 176B Inference on AWS | Bigger than GPT-3 for more Power!
Discover AI
29 Fine-tune ChatGPT? Buy Embeddings /OpenAI? What are Embeddings?  My own ChatGPT? | Visual Q+A
Fine-tune ChatGPT? Buy Embeddings /OpenAI? What are Embeddings? My own ChatGPT? | Visual Q+A
Discover AI
30 Unleashing the Power of BLOOM 176B with AWS ml.p4de.24xlarge, DJL & DeepSpeed: The Ultimate Boost!
Unleashing the Power of BLOOM 176B with AWS ml.p4de.24xlarge, DJL & DeepSpeed: The Ultimate Boost!
Discover AI
31 After ChatGPT: NEW BioGPT by Microsoft | Do YOU trust Microsoft for your Medication?
After ChatGPT: NEW BioGPT by Microsoft | Do YOU trust Microsoft for your Medication?
Discover AI
32 Improve ChatGPT: Modular, Adaptive, Smart LLM | Inside ChatGPT
Improve ChatGPT: Modular, Adaptive, Smart LLM | Inside ChatGPT
Discover AI
33 Fine-tune ChatGPT w/  in-context learning ICL - Chain of Thought, AMA, reasoning & acting: ReAct
Fine-tune ChatGPT w/ in-context learning ICL - Chain of Thought, AMA, reasoning & acting: ReAct
Discover AI
34 The Intersection of Copyright Law and Human Faces: Exploring Virtual K-Pop with MAVE
The Intersection of Copyright Law and Human Faces: Exploring Virtual K-Pop with MAVE
Discover AI
35 New TECH: Vision Transformer 2023 on Image Classification | AI
New TECH: Vision Transformer 2023 on Image Classification | AI
Discover AI
36 PyTorch code Vision Transformer: Apply ViT models pre-trained and fine-tuned  | AI  Tech
PyTorch code Vision Transformer: Apply ViT models pre-trained and fine-tuned | AI Tech
Discover AI
37 New BING ChatGPT: Unlock the Power of Emotions in your Search Engine!
New BING ChatGPT: Unlock the Power of Emotions in your Search Engine!
Discover AI
38 New BING ChatGPT loses its mind
New BING ChatGPT loses its mind
Discover AI
39 Self-Attention Heads of last Layer of Vision Transformer (ViT) visualized (pre-trained with DINO)
Self-Attention Heads of last Layer of Vision Transformer (ViT) visualized (pre-trained with DINO)
Discover AI
40 Visualizing the Self-Attention Head of the Last Layer in DINO ViT: A Unique Perspective on Vision AI
Visualizing the Self-Attention Head of the Last Layer in DINO ViT: A Unique Perspective on Vision AI
Discover AI
41 Microsoft strongly restricts access to ChatGPT on new BING - WHY?
Microsoft strongly restricts access to ChatGPT on new BING - WHY?
Discover AI
42 PyTorch ViT: The Ultimate Guide to Fine-Tuning for Object Identification (COLAB)
PyTorch ViT: The Ultimate Guide to Fine-Tuning for Object Identification (COLAB)
Discover AI
43 New BING Chat AGGRESSIVE
New BING Chat AGGRESSIVE
Discover AI
44 Panoptic Image Segmentation: Mask2Former explained | Identify all objects!
Panoptic Image Segmentation: Mask2Former explained | Identify all objects!
Discover AI
45 Code Panoptic Image Segmentation w/ Vision Transformer & Mask2Former - A PyTorch tutorial
Code Panoptic Image Segmentation w/ Vision Transformer & Mask2Former - A PyTorch tutorial
Discover AI
46 Dream Job Alert: AI Prompt Engineer - $335K  |  AI Prompt Design: A Crash Course
Dream Job Alert: AI Prompt Engineer - $335K | AI Prompt Design: A Crash Course
Discover AI
47 Streamlining Similar Image Detection with ViT in PyTorch: A Step-by-Step Guide
Streamlining Similar Image Detection with ViT in PyTorch: A Step-by-Step Guide
Discover AI
48 Microsoft's CEO in Trouble   #shorts
Microsoft's CEO in Trouble #shorts
Discover AI
49 Why wait for KOSMOS-1? Code a VISION - LLM w/ ViT, Flan-T5 LLM and BLIP-2: Multimodal LLMs (MLLM)
Why wait for KOSMOS-1? Code a VISION - LLM w/ ViT, Flan-T5 LLM and BLIP-2: Multimodal LLMs (MLLM)
Discover AI
50 OpenAI's ChatGPT can NOW summarize external Sources on the Internet?
OpenAI's ChatGPT can NOW summarize external Sources on the Internet?
Discover AI
51 ChatGPT polarizes
ChatGPT polarizes
Discover AI
52 Hospital /Clinic AI Decision Models: Performance of 12 AI LLM Systems (incl $$) Radiology, Biomed
Hospital /Clinic AI Decision Models: Performance of 12 AI LLM Systems (incl $$) Radiology, Biomed
Discover AI
53 ChatGPT Prompt Engineering w/ in-context learning (ICL)  - 7 Examples | Tutorial
ChatGPT Prompt Engineering w/ in-context learning (ICL) - 7 Examples | Tutorial
Discover AI
54 Chat with your Image!  BLIP-2 connects Q-Former w/ VISION-LANGUAGE models (ViT & T5 LLM)
Chat with your Image! BLIP-2 connects Q-Former w/ VISION-LANGUAGE models (ViT & T5 LLM)
Discover AI
55 ChatGPT:  Multidimensional Prompts
ChatGPT: Multidimensional Prompts
Discover AI
56 ChatGPT:  In-context Retrieval-Augmented Learning (IC-RALM) | In-context Learning (ICL) Examples
ChatGPT: In-context Retrieval-Augmented Learning (IC-RALM) | In-context Learning (ICL) Examples
Discover AI
57 Code your BLIP-2 APP: VISION Transformer (ViT) + Chat LLM (Flan-T5) = MLLM
Code your BLIP-2 APP: VISION Transformer (ViT) + Chat LLM (Flan-T5) = MLLM
Discover AI
58 Buy Microsoft "Azure OpenAI Service" or buy from OpenAI its API for ChatGPT access & tuning?
Buy Microsoft "Azure OpenAI Service" or buy from OpenAI its API for ChatGPT access & tuning?
Discover AI
59 Pretraining vs Fine-tuning vs In-context Learning of LLM (GPT-x) EXPLAINED | Ultimate Guide ($)
Pretraining vs Fine-tuning vs In-context Learning of LLM (GPT-x) EXPLAINED | Ultimate Guide ($)
Discover AI
60 Reversible Transformer: ReFORMER for GPU Memory Optimization! Reversible Residual Layers?
Reversible Transformer: ReFORMER for GPU Memory Optimization! Reversible Residual Layers?
Discover AI

The video demonstrates the capabilities of the Gemini 3.1 Pro AI model in causal reasoning and complex logical tasks, and provides a comparison with other models. The test results show that Gemini 3.1 Pro outperforms the other models in solving a complex problem with a sequence of seven plus exit steps. The video also discusses the potential applications of the Gemini 3.1 Pro model and its availability on various platforms.

Key Takeaways
  1. Design a complex problem with a sequence of seven plus exit steps
  2. Implement the problem in a language model
  3. Evaluate the performance of the language model on the problem
  4. Compare the performance of the language model with other models
  5. Fine-tune the language model for improved performance on complex tasks
  6. Apply transfer learning to the language model for improved performance on related tasks
💡 The Gemini 3.1 Pro model demonstrates outstanding performance in causal reasoning and complex logical tasks, outperforming other models in solving a complex problem with a sequence of seven plus exit steps.

Related Reads

📰
Which Is Better: Claude, Perplexity, or ChatGPT? A Complete Comparison for 2026
Learn how to choose between Claude, Perplexity, and ChatGPT for your AI needs in 2026
Medium · ChatGPT
📰
Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics
Learn to build production-grade LLM evaluation pipelines to catch hallucinations before deployment and improve model reliability
Dev.to AI
📰
Understanding "Handoffs" in LangChain(One Agent, Many Personalities)
Learn to implement 'handoffs' in LangChain to make a single agent behave differently in various conversation stages
Dev.to AI
📰
Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics
Learn to build production-grade LLM evaluation pipelines to catch hallucinations before deployment, replacing manual 'vibe checks' with automated metrics
Dev.to AI
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →