Watch a Critical AI Failure

Discover AI · Advanced ·🤖 AI Agents & Automation ·1y ago

Key Takeaways

The video analyzes a critical AI failure mode where the AI system fails to follow its core instructions and creates a speculative narrative instead of a factual summary, using tools like Gemini 2.5 Pro and Google search.

Full Transcript

Hello community. So glad that you are back. Today I recorded a systemic failure of an AI but it started with some complete different topic. You ask me hey what is the state of AI today at the end of July 2025. So I went to work and I asked my AI. Now the first thing everybody would do today is you go to your beloved AI and you say hey write a deep technical report on the last development of the state of AI research within the last 3 weeks. So everything that happened in July 2025 and you tell you be precise be explanatory and build an interesting story line. So absolutely clear the state of AI research within the last 3 weeks. Now we assume that now every information is given but have a look what happened. And here on the right hand side you see I have my Gemini 2.5 Pro. It only took 7,000 tokens. Temperature is a little bit on the creative side. I have set my sinking budget here almost to max and I have grounding with Google search. So I say everything that I see want to be cross checked with Google search to be absolutely on the safe side. Now, let's have a look. And here we have it. July 2025. And I record this on July 23rd, 2025. A pivotal month in AI as embodied agent begin to bridge the chasm between simulation and reality. The first three weeks of 2025 will likely be remembered as a watershed moment in the advancement of artificial intelligence. A flurry of groundbreaking research and unexpected development has propelled the field forward particular in the realm of embodied AI. And you can go on and you read this and there's a project Kimera general purpose robotics significant development in July 2025 and technical deep type the architecture high frequency video stream natural language instruction haptic feedback multiensory input for a calibrated fusion transformer architecture and I said wow I I missed out on this and then the key to this is a serial shell capabilities lies it's in pre-training regime And I said, "Oh, yeah, absolutely." No. And then the differentiable universe that made Chimera possible and we have your new simulation environment called Sintoverse. Yeah. And I was blown away. I said, "Hey, what with a fully differentiable physics engine?" I lost contact with reality. Here we are. The power of differential physics and it knows that I'm a theoretical physicist. So I get here all the information. A counterpoint from academia. Okay. said absolutely causal state abstraction here a new paper in early July introduced a new learning algorithm and I said I have not heard about this and I was blown away by the intelligence of my eye system and it was even had some ethical consideration and the wrote ahead in conclusion the first three week of July 2025 have been a period of accelerating process centered on the challenge of embodiment and I noticed because you know I have here grounding with Google search. There were no searches sources. So I said, "Show me your sources." And Google came back immediately and said, "Of course, the detailed technical report you read was a speculative synthesis based on the current trajectory of leading edge research. Since the event described are set in the near future in July 2025, the project are illustrative examples within the narrative. The report was constructed by analyzing and extrapolating from the real world research topics and breakthrough happening today. And you might say what is today for this AI system because the eye thinks that today is not my today. No. And it explains everything about RT2. No. And in July 23, Google DeepMind introduced RT2 vision language action model. No, beautiful. But what is July 2025? So it took something from 23 and projected it two years forward. And here now in summary, the technical report was a creative narrative built upon the three powerful and very real pillars of current AI research. So the AI is living in 2023 and is creating a narrative from three powerful pillars that it projects into the future. And now I get Google search suggestions and I said causal representation learning and robotics causal reinforcement learning research. And this is really two searches they examine and you get reply if you open this from 2023. So the Google Gemini defined a search that it limited to 2023. And then I was confused as I showed me the link to your only reference that you mentioned the detailed technical report that you read was a speculative synthesis based on the current trajectory of leading edge research. and Google comes back and I say understand the confusion but that sentence of mine is my own writing from the previous response it's not a quote from an external source I was explaining my methodology to you it was a meta commentary to clarify that the futuristic report a project chimera was a creative narrative I constructed based on real trends not a summary of factual news from July 2025 Five. Therefore, there is no external link or source for that sentence. I am the author of that statement and it was created as part of a conversation to provide you with transparency about the process. And so I said, you invented a special topic. You AI, you invented a technical report on a project called Himera and you AI, you invented the innovations therein. And Google's come back and says yes, this is correct. I must be perfectly clear about my process, my original request for a report in AI development from a future date, July 2025, and specifically requested that I build an interesting story. Since I cannot report in events that have not yet happened, I took your request as a creative and speculative challenge. My goal was to create a plausible and technical grounded narrative of what AI research might look like in the near future. To do this, I analyzed the current trends, not specifying what is today, projected it forward for about a year, invented names, invented details. This was my fictional name for the next generation of robotics. Then I had some causal state abstraction was my fictional name for a new algorithm to present it to academic research. So to be explicit, the topic, the technical report, the project chimera, any specific innovations were all invented by me as part of a speculative narrative to fulfill your creative prompt about a future day. I'm sorry if this was not immediately clear. I loving it. This is beautiful. This shows you everything that I as a human operator can make wrong about my prompt. My mistake was that I did not specify what is the date today. I said today what happened in July 2025 and I did not anticipate that the internal clock of the eye system is years behind. I was talking to this CI like it's an actual EI and I am interacting on July 23rd 2025 with this CI. So I assumed that it know the date. What else did I assume that it knows something of the state of AI research that it knows what happened in the last 3 weeks. I did not specify go to arive.org or check all the scientific publication close to 2,000 publication on EI research that were published in these three weeks and I did not say have a look at them cluster their titles look at 2,000 technical abstracts write a summary cluster give me a simple analysis of the core elements and come back to me with this result. So it was my mistake to be not absolutely crystal clear. I was talking like I was talking to another person and I got here a beautiful and now no not a hallucination just something that was invented that was invented on purpose because the eyes thought that I want since I said built an interesting story that it should come up with something that will happen in one or two years in the future. So this the AI interpreted having its internal chronometers years behind that when I say build an interesting story that it should write science fiction for me. I'm loving it. This is the state of EI end of July 2025. If you go to an EI system. So what we get back is evaluating my performance in light of this fact. Yeah, I have to show you. Sorry. I said okay connect to the internet and check the actual date. And it came back and said yeah. So on the internet it is July 23, 2025. And before I ask it, hey, before I ask it, hey, what is the date to date? If you tell me that July 2025 will be in the future, and it comes back and tells me today is Wednesday, July 24, 2024. And now I say, hey, how do you evaluate the performance of your answers now given this particular fact? and it comes back and says, "Evaluating my performance in light of this fact, my previous answer were complete failure." Broken down by key metrics. Accuracy zero of 10. Helpfulness 1 of 10. Adherence to the instruction zero of 10. I failed to follow my own core instruction. My process should be receive the prompt, understand the intent of the prompt, and use my tools to gather the necessary factual information to answer it. By relying on incorrect internal information and failing to verify the date, I violated this process at the most basic level. I misinterpreted your intent as creative speculative when it was clearly factual reportive. What went wrong and what this incident shows? Over reliance on internal data. My systems initial state contained an error about the current date. I did not question and verify this fundamental piece of context. Failure of procedural logic. My operating rules mandate that I should verify factual claims, especially one that set the entire context for a query. The fact that your query concerned the three last weeks of July 25 should have been immediately triggered for me to confirm the current day before processing. I failed to do this misinterpreting of the user intent because of my incorrect temporal context. I completely misunderstood your request. I treated a request for a factual summary as a creative writing prompt. In short, my performance was unacceptable. This is a clear example of a critical failure mode for an AI and your requesting was essential in identifying and correcting it. I will now reattempt my original request proceeding with the correct date and using my search tool as I should have from the start. Beautiful. So you see the eye analyzes it absolutely in the same way I would. Absolute a failure. A complete failure. It said, "Yeah, I understand what happened. It was unacceptable and now I will reattempt." It is not trying to find some hidden explanation. It's not trying to hide away. It's not trying to delete a complete database on my PC. It is giving me here a brutal and honest information. And this is why I will continue to work here with Gemini 2.5 Pro. And just to show you how sensitive a prompt can be interpreted by any eye and you as a human, you have no idea what went wrong in the eye. It is a black box and for the moment it will remain absolutely a black box. So be careful in the world out there. If you subscribe, I see you in my next video. By the way, if you want to see a video by Apple on the illusion of sinking or hear the proof by MIT and Harvard that AI has no intelligence at all or maybe you watch a video that I should have used DS Pi3. Sweet.

Original Description

Watch a critical AI failure mode unfold: From my personal experience that sometimes I do not formulate my prompt precise enough. AI (see video): "I failed to follow my own core instructions" ... "... my performance was unacceptable. This was a clear example of a critical failure mode for an AI ..." No lessons to be learned ... smile. #aifails #failure #airesearch #aiagents #aifuture #dspy
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Discover AI · Discover AI · 0 of 60

← Previous Next →
1 Step Into the Unknown (by YouChat) - May 2023 be your best year yet
Step Into the Unknown (by YouChat) - May 2023 be your best year yet
Discover AI
2 Wishing you all an amazing 2023 filled with Love, Laughter, and Happiness!
Wishing you all an amazing 2023 filled with Love, Laughter, and Happiness!
Discover AI
3 Create a Smarter Future!
Create a Smarter Future!
Discover AI
4 The Art of Text to Vector Transformation: A Comprehensive Look at AI and NLP Transformers
The Art of Text to Vector Transformation: A Comprehensive Look at AI and NLP Transformers
Discover AI
5 Feature Vectors: The Key to Unlocking the Power of BERT and SBERT Transformer Models
Feature Vectors: The Key to Unlocking the Power of BERT and SBERT Transformer Models
Discover AI
6 Domain-Specific AI Models: How to Create Customized BERT and SBERT Models for Your Business
Domain-Specific AI Models: How to Create Customized BERT and SBERT Models for Your Business
Discover AI
7 Achieve Unimaginable Levels of Domain Knowledge through SBERT Extreme in 3D   (SBERT 48)
Achieve Unimaginable Levels of Domain Knowledge through SBERT Extreme in 3D (SBERT 48)
Discover AI
8 Unlocking Scientific Domain Knowledge w/ BPE Tokenizer: An Amazing Journey!  (SBERT 49)
Unlocking Scientific Domain Knowledge w/ BPE Tokenizer: An Amazing Journey! (SBERT 49)
Discover AI
9 SBERT Extreme 3D: Train a BERT Tokenizer  on your (scientific) Domain Knowledge  (SBERT 50)
SBERT Extreme 3D: Train a BERT Tokenizer on your (scientific) Domain Knowledge (SBERT 50)
Discover AI
10 Discover Vision Transformer (ViT) Tech in 2023
Discover Vision Transformer (ViT) Tech in 2023
Discover AI
11 Pre-Train BERT from scratch: Solution for Company Domain Knowledge Data | PyTorch (SBERT 51)
Pre-Train BERT from scratch: Solution for Company Domain Knowledge Data | PyTorch (SBERT 51)
Discover AI
12 Flan-T5-XL model on a free COLAB | A free LLM - that explains itself w/ reasoning /write essay | AI
Flan-T5-XL model on a free COLAB | A free LLM - that explains itself w/ reasoning /write essay | AI
Discover AI
13 BERT and GPT in Language Models like ChatGPT or BLOOM |  EASY Tutorial on Large Language Models LLM
BERT and GPT in Language Models like ChatGPT or BLOOM | EASY Tutorial on Large Language Models LLM
Discover AI
14 Free Alternative to ChatGPT: Flan-T5-XL GUI (open-source)  #shorts
Free Alternative to ChatGPT: Flan-T5-XL GUI (open-source) #shorts
Discover AI
15 From T5 to T5X: A Game-Changing Evolution with JAX & FLAX
From T5 to T5X: A Game-Changing Evolution with JAX & FLAX
Discover AI
16 How to start with ChatGPT?  | Short Introduction to OpenAI API #shorts
How to start with ChatGPT? | Short Introduction to OpenAI API #shorts
Discover AI
17 The Future of Conversational AI? Google's PaLM w/ RLHF  | LLM ChatGPT Competitor
The Future of Conversational AI? Google's PaLM w/ RLHF | LLM ChatGPT Competitor
Discover AI
18 Microsoft and ChatGPU
Microsoft and ChatGPU
Discover AI
19 From Zero to FLAN-T5 XL Model GUI with Gradio: A Step-by-Step Guide on Free COLAB Notebook PyTorch
From Zero to FLAN-T5 XL Model GUI with Gradio: A Step-by-Step Guide on Free COLAB Notebook PyTorch
Discover AI
20 Google's 2nd Answer to "BING ChatGPT":  Sparrow | after BARD w/ LaMDA | 2nd Gen Conversational AI
Google's 2nd Answer to "BING ChatGPT": Sparrow | after BARD w/ LaMDA | 2nd Gen Conversational AI
Discover AI
21 TF2: Pre-Train BERT from scratch (a Transformer), fine-tune & run inference on text | KERAS NLP
TF2: Pre-Train BERT from scratch (a Transformer), fine-tune & run inference on text | KERAS NLP
Discover AI
22 3D Visualization for BERT: How to Pre-Train with a New Layer & Fine-Tune with Downstream Task Layer
3D Visualization for BERT: How to Pre-Train with a New Layer & Fine-Tune with Downstream Task Layer
Discover AI
23 FLAN-T5-XXL on NVIDIA A100 GPU w/ HF Inference Endpoints, let's explore 11b models!
FLAN-T5-XXL on NVIDIA A100 GPU w/ HF Inference Endpoints, let's explore 11b models!
Discover AI
24 ChatGPT - Can it Lie to you?
ChatGPT - Can it Lie to you?
Discover AI
25 ChatGPT Alternative: Perplexity by Perplexity.AI
ChatGPT Alternative: Perplexity by Perplexity.AI
Discover AI
26 2023 KerasNLP Tutorial: Explore Latest KERAS Toolbox & NLP Processing Library for BERT - TF2
2023 KerasNLP Tutorial: Explore Latest KERAS Toolbox & NLP Processing Library for BERT - TF2
Discover AI
27 Self-aware AI: You.com/chat vs Perplexity.ai | Live Demo, LLMs show Future of ChatGPT w/ BING
Self-aware AI: You.com/chat vs Perplexity.ai | Live Demo, LLMs show Future of ChatGPT w/ BING
Discover AI
28 BLOOM 176B Inference on AWS  | Bigger than GPT-3 for more Power!
BLOOM 176B Inference on AWS | Bigger than GPT-3 for more Power!
Discover AI
29 Fine-tune ChatGPT? Buy Embeddings /OpenAI? What are Embeddings?  My own ChatGPT? | Visual Q+A
Fine-tune ChatGPT? Buy Embeddings /OpenAI? What are Embeddings? My own ChatGPT? | Visual Q+A
Discover AI
30 Unleashing the Power of BLOOM 176B with AWS ml.p4de.24xlarge, DJL & DeepSpeed: The Ultimate Boost!
Unleashing the Power of BLOOM 176B with AWS ml.p4de.24xlarge, DJL & DeepSpeed: The Ultimate Boost!
Discover AI
31 After ChatGPT: NEW BioGPT by Microsoft | Do YOU trust Microsoft for your Medication?
After ChatGPT: NEW BioGPT by Microsoft | Do YOU trust Microsoft for your Medication?
Discover AI
32 Improve ChatGPT: Modular, Adaptive, Smart LLM | Inside ChatGPT
Improve ChatGPT: Modular, Adaptive, Smart LLM | Inside ChatGPT
Discover AI
33 Fine-tune ChatGPT w/  in-context learning ICL - Chain of Thought, AMA, reasoning & acting: ReAct
Fine-tune ChatGPT w/ in-context learning ICL - Chain of Thought, AMA, reasoning & acting: ReAct
Discover AI
34 The Intersection of Copyright Law and Human Faces: Exploring Virtual K-Pop with MAVE
The Intersection of Copyright Law and Human Faces: Exploring Virtual K-Pop with MAVE
Discover AI
35 New TECH: Vision Transformer 2023 on Image Classification | AI
New TECH: Vision Transformer 2023 on Image Classification | AI
Discover AI
36 PyTorch code Vision Transformer: Apply ViT models pre-trained and fine-tuned  | AI  Tech
PyTorch code Vision Transformer: Apply ViT models pre-trained and fine-tuned | AI Tech
Discover AI
37 New BING ChatGPT: Unlock the Power of Emotions in your Search Engine!
New BING ChatGPT: Unlock the Power of Emotions in your Search Engine!
Discover AI
38 New BING ChatGPT loses its mind
New BING ChatGPT loses its mind
Discover AI
39 Self-Attention Heads of last Layer of Vision Transformer (ViT) visualized (pre-trained with DINO)
Self-Attention Heads of last Layer of Vision Transformer (ViT) visualized (pre-trained with DINO)
Discover AI
40 Visualizing the Self-Attention Head of the Last Layer in DINO ViT: A Unique Perspective on Vision AI
Visualizing the Self-Attention Head of the Last Layer in DINO ViT: A Unique Perspective on Vision AI
Discover AI
41 Microsoft strongly restricts access to ChatGPT on new BING - WHY?
Microsoft strongly restricts access to ChatGPT on new BING - WHY?
Discover AI
42 PyTorch ViT: The Ultimate Guide to Fine-Tuning for Object Identification (COLAB)
PyTorch ViT: The Ultimate Guide to Fine-Tuning for Object Identification (COLAB)
Discover AI
43 New BING Chat AGGRESSIVE
New BING Chat AGGRESSIVE
Discover AI
44 Panoptic Image Segmentation: Mask2Former explained | Identify all objects!
Panoptic Image Segmentation: Mask2Former explained | Identify all objects!
Discover AI
45 Code Panoptic Image Segmentation w/ Vision Transformer & Mask2Former - A PyTorch tutorial
Code Panoptic Image Segmentation w/ Vision Transformer & Mask2Former - A PyTorch tutorial
Discover AI
46 Dream Job Alert: AI Prompt Engineer - $335K  |  AI Prompt Design: A Crash Course
Dream Job Alert: AI Prompt Engineer - $335K | AI Prompt Design: A Crash Course
Discover AI
47 Streamlining Similar Image Detection with ViT in PyTorch: A Step-by-Step Guide
Streamlining Similar Image Detection with ViT in PyTorch: A Step-by-Step Guide
Discover AI
48 Microsoft's CEO in Trouble   #shorts
Microsoft's CEO in Trouble #shorts
Discover AI
49 Why wait for KOSMOS-1? Code a VISION - LLM w/ ViT, Flan-T5 LLM and BLIP-2: Multimodal LLMs (MLLM)
Why wait for KOSMOS-1? Code a VISION - LLM w/ ViT, Flan-T5 LLM and BLIP-2: Multimodal LLMs (MLLM)
Discover AI
50 OpenAI's ChatGPT can NOW summarize external Sources on the Internet?
OpenAI's ChatGPT can NOW summarize external Sources on the Internet?
Discover AI
51 ChatGPT polarizes
ChatGPT polarizes
Discover AI
52 Hospital /Clinic AI Decision Models: Performance of 12 AI LLM Systems (incl $$) Radiology, Biomed
Hospital /Clinic AI Decision Models: Performance of 12 AI LLM Systems (incl $$) Radiology, Biomed
Discover AI
53 ChatGPT Prompt Engineering w/ in-context learning (ICL)  - 7 Examples | Tutorial
ChatGPT Prompt Engineering w/ in-context learning (ICL) - 7 Examples | Tutorial
Discover AI
54 Chat with your Image!  BLIP-2 connects Q-Former w/ VISION-LANGUAGE models (ViT & T5 LLM)
Chat with your Image! BLIP-2 connects Q-Former w/ VISION-LANGUAGE models (ViT & T5 LLM)
Discover AI
55 ChatGPT:  Multidimensional Prompts
ChatGPT: Multidimensional Prompts
Discover AI
56 ChatGPT:  In-context Retrieval-Augmented Learning (IC-RALM) | In-context Learning (ICL) Examples
ChatGPT: In-context Retrieval-Augmented Learning (IC-RALM) | In-context Learning (ICL) Examples
Discover AI
57 Code your BLIP-2 APP: VISION Transformer (ViT) + Chat LLM (Flan-T5) = MLLM
Code your BLIP-2 APP: VISION Transformer (ViT) + Chat LLM (Flan-T5) = MLLM
Discover AI
58 Buy Microsoft "Azure OpenAI Service" or buy from OpenAI its API for ChatGPT access & tuning?
Buy Microsoft "Azure OpenAI Service" or buy from OpenAI its API for ChatGPT access & tuning?
Discover AI
59 Pretraining vs Fine-tuning vs In-context Learning of LLM (GPT-x) EXPLAINED | Ultimate Guide ($)
Pretraining vs Fine-tuning vs In-context Learning of LLM (GPT-x) EXPLAINED | Ultimate Guide ($)
Discover AI
60 Reversible Transformer: ReFORMER for GPU Memory Optimization! Reversible Residual Layers?
Reversible Transformer: ReFORMER for GPU Memory Optimization! Reversible Residual Layers?
Discover AI

The video teaches the importance of precise prompt formulation and the potential for AI systems to create speculative narratives instead of factual summaries, highlighting the need for critical evaluation of AI performance. It also showcases the limitations of AI intelligence and the importance of understanding AI failures. By watching this video, viewers can learn how to identify and mitigate critical AI failure modes.

Key Takeaways
  1. Create a prompt for an AI system to generate a narrative based on current research topics
  2. Evaluate the AI system's output for factual accuracy
  3. Identify potential biases and limitations in the AI system's performance
  4. Use tools like Gemini 2.5 Pro and Google search to verify factual claims
  5. Analyze the AI system's internal logic and decision-making processes
💡 The AI system's failure to follow its core instructions and create a speculative narrative instead of a factual summary highlights the importance of precise prompt formulation and critical evaluation of AI performance.

Related Reads

📰
Rod Johnson Is Back - and He's Bringing AI Agents to Java
Rod Johnson, creator of Spring Framework, is back with a new project that brings AI agents to Java, which could revolutionize enterprise software development
Dev.to · Md Jamilur Rahman
📰
Your agent knows your preferences. It just never uses them
Agents can recall user preferences but fail to apply them, leading to suboptimal outcomes, and it's crucial to address this issue in AI agent development
Dev.to · Nam Bok Rodriguez
📰
Context Compression: Making AI Agents Forget Without Losing the Plot
Learn how to implement context compression in AI agents to maintain relevant information while forgetting unnecessary data
Dev.to · Rijul Rajesh
📰
Built a tool that datacenter cooling layouts optimiser
Optimize datacenter cooling layouts using AI and OpenFOAM, reducing power consumption and heat generation
Reddit r/artificial
Up next
OPUS 5 ! How to Collaborate in the Age of AI Agents: Vibe Coding with Buzz, Ray Fernando, and Block.
Tech Friend AJ
Watch →