Self-improving AI, Opus 4.8, Nvidia bangers, game-ready 3D models, juggling robots: AI NEWS
Skills:
Staying Current in AI90%
Key Takeaways
Covers recent AI news, including updates on Claude Opus 4.8, Step 3.7, DeepSWE, and other AI tools and technologies
Full Transcript
AI never sleeps and this week has been absolutely insane. Anthropic releases their best and latest model, Opus 4.8. This AI can reconstruct an entire scene from just a casual phone video. Nvidia dropped some open- source bangers this week, including a super fast and accurate object detector, as well as an incredible image upscaler. Plus, they also released a world simulator of up to four different players at once. We have another AI that can create 3D models that are simulation ready. So, these can be animated or moved around with accurate articulation. This AI can generate images on just your phone. Roblox releases an open-source 3D model generator that can create game ready assets with just a text prompt. We have an open-source agentic system for automating scientific research. This AI can relight an image at any angle, plus change the harshness of the light. We have some ridiculous humanoid robot demos and a lot more. So, let's jump right in. Thanks to HubSpot for sponsoring this video. First up, Nvidia releases a really powerful vision language grounding model called locate anything. In simple terms, you can give it an image or a video, and it can tell you exactly where anything is in the video. So, if you ask it to detect and segment a certain animal or a certain object in an image or video, even if it's a really crowded scene or even if there are multiple instances of that specific object, it's able to accurately output bounding boxes around all of them. The clever part is how it predicts those boxes. A lot of vision language models generate coordinates token by token, almost like spelling out a location one number at a time. This can be slow and it can also make the box geometry less reliable. But what locate anything does is it uses something called parallel box decoding which means it predicts the whole bounding box together in just one step. This makes it way faster and more geometrically consistent. The data set they used to train this is also huge. So this includes 103 million language queries, 785 million bounding boxes across object detection, detecting things in interfaces, OCR, layout understanding, and a ton of other different tasks. The awesome thing is they've released this already. So at the top, if you click on this GitHub button, it takes you to this page and if you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. They also give you the code on how to train this and fine-tune it. Note that this is fairly tiny at only 3 billion parameters and 7.8 GB in size, so this should be able to fit on most consumer GPUs. If you're interested in reading further, I'll link to this main page in the description below. Also, this week we have a really useful AI for image editing. So, this is called control light, and this is a tool for fixing dark images or editing the lighting of an image. It gives you exact control over how the image gets brightened. So, here are some examples. I can drag the slider to adjust the brightness of the light of this room like this. Or here's another example where I can drag this to adjust the brightness. Now, traditional tools like Photoshop also have a brightness slider, of course. But if you drag the brightness too much, then it introduces a lot of artifacts and noise. But by using the AI to generatively add brightness to the scene, it's able to preserve all the details of the original image without introducing a lot of errors. So here, the model understands how real photos should look. So instead of just cranking up exposure or making everything washed out, this is designed to be way more flexible. The model learns different degrees of brightness instead of just one before and after pair. It also focuses on keeping the structure of the image stable as the light changes, so it doesn't introduce any errors. So, this could be useful for like surveillance or restoring photos or night scenes where using a traditional brightness slider is not enough. The awesome thing is they've released everything already. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. Notice that this is based off of Flux 2 client, which should be able to run on most consumer devices. And then here are the instructions on how you can run this. But in addition, they also released the code on how to train this as well as the data set that was used to train this. So, this is fully open source. If you're interested in reading further, I'll link to this main page in the description below. Next up, this AI is really interesting. It's called Triclat, and this is a 3D reconstruction model that can turn a set of images into a 3D scene. And this is simulation ready. So, after creating the 3D scene, for example, you can also input a robot to navigate across the scene. Now, how this works is really interesting. You see, a lot of 3D reconstruction models use Gausian splats. These are kind of like dots that are scattered across the 3D space, but they're usually not usable for clean surfaces in like robotic simulations. If you want physics, collisions, or game engine interaction, you need extra conversion steps to extract a mesh. Well, tripplat completely skips that step by representing the entire splat as these triangle primitives from the start. In other words, you can think of this entire scene as just made of triangles instead of dots. And the result here is that it's not just something that looks 3D, but it's actually much closer to a usable 3D scene. You can see here that this new triplat, which is in red, is able to reconstruct these 3D scenes way faster. Whereas for other methods that use gausian splattings, you also need to convert those splats into meshes first. So it takes much longer. Now, if you scroll up to the top of the page, they have released the code to this. So if you click on this code button, it takes you to this page. And if you scroll down a bit here, it contains all the instructions on how to install and run this locally on your computer. Note that the size of each model is only like 4.4 GB in size, so it's pretty tiny. It should be able to fit on most consumer devices. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Nvidia continues to cook some really cool open- source stuff. So they also released PID and this is a really powerful way to upscale images and turn them into higher resolution with much better detail. So here are some examples for your reference. You can upscale this to like 2K or beyond. And the details are really good. Here are some additional examples for your reference. Notice the insane increase in quality after you plug it through this P step. And then here's another example. Everything is super clean and detailed and consistent. Now, here's the issue it's trying to solve. For most traditional image generation models, if you want to generate highresolution images, then the final decoding step can be a bottleneck. Usually, the system decodes the image first from latent space into pixel space and then it might use a separate upscaler to make the image larger. Well, P completely replaces the step with just one pixel diffusion decoder. In other words, it directly outputs a high resolution in pixel space. This is also incredibly fast. So you can see here that P is able to upscale a 512 x 512 image into 2K in under 1 second which is crazy. This is almost six times faster than seed VR2 which is currently one of the best upscalers out there. And on this chart this is the win rate of P compared to the other models. And as you can see P is able to win the majority of instances. The awesome thing is they've released this already. So at the top if you click on this code button and you scroll down a bit here it contains all the instructions on how to download and run this locally on your computer. Note that this already works in comi plus the current released model supports flux 2 and z image and also sd3 and these other models. They also plans to release support for quen image and sdxl in the future. If you're interested in reading further I'll link to this main page in the description below. Also this week this AI is pretty interesting. It's called instruct AV toAVV. And this is a system for editing video and audio together just from a simple prompt. So for example, we can input a video and then get that person to say something completely different. And it's able to edit not only the audio but her lip sync. Here's the example. >> Yeah, I see what you mean. But I really think we should give it another shot. This is more than just art. It's a statement. >> Or here's another example. I understand the situation and I believe that you can do it. I really think we should try our best to do it. >> Or here's another example where we can change the person's voice to sound more like a man and also change the words. >> I did find an address for a mother up in Hartwell. >> I understand, but I think we need to consider. >> Now, we don't have to change what they say. We can keep the speech, but change the man into a woman. And here's what we get. >> We must act now before it's too late. >> We must act now before it's too late. >> So, a pretty interesting tool. Now, if you scroll up to the top of the page, they have released a code button plus the model weights button, but currently there's nothing on this GitHub page yet. Anyway, if you're interested in reading further, I'll link to this page in the description below. Next up, this AI is great for real estate VR. So, it's called Gen Recon. And as you can see from this video, this is basically a system that can turn casual smartphone videos or even just a set of images of a room into a complete 3D scene which you can edit further. So the input is a set of images or a video of a room and the output is a PBR ready mesh, meaning a 3D model with materials that can be used for post-processing or rendering. For example, you can also relight it at different angles or change the surfaces or colors of stuff in the scene if you want. So, here's how it works. Gen Recon takes a set of images or frames in a video and then it generates chunks of the 3D scene together so that the full scene stays consistent. It also uses Trellis 2 as what they call a generative shape prior. So, think of this as like giving the system a strong sense of what real 3D shapes should look like instead of just forcing it to only use what's directly visible in the images. And this allows it to reconstruct the scene a lot better and more realistically. And so afterwards, you get a 3D scene where you can relight it from different angles or even move stuff around or edit the objects. Now, at the top, it does have a coming soon button, so hopefully they will open source this. For now, they've only released a technical paper, but if you're interested in reading further, I'll link to this main page in the description below. Also this week, we have an AI called scope, and this is designed to make first-person shooter game worlds that are actually playable. So, the input is a starting frame plus controller actions, and the output is basically a video, but it can respond to these actions in real time. So, for example, you can move or aim and fire, reload, switch weapons, interact with the environment, all while the camera is changing. So, the AI model has to be trained to adapt to all these different key presses and respond to different actions. Now, the quality is not perfect. You can see that the gun and the details of the scene do kind of warp over time, but this is one of the first instances of a generative world model that can respond to so many different actions, including firing and reloading and shooting enemies. So, they trained it on a huge data set with almost 70,000 clips from seven firstperson shooter games, including 10 different types of controller signals. So, this lets the model learn the action patterns across games instead of just memorizing one title. and the benchmark results are pretty good. So, in terms of visual quality, motion quality, and consistency, then on average, Scope even outperforms some recent game generators like Matrix Game 3 or Hunyan World 5. If you scroll up to the top of the page, they have released the code to this already. So, if you click on this code button over here, it contains all the instructions on how to download and run this locally on your computer. Note that this is based off of 1 2.2 2. And the total size of this is around 30 GB. So, you'll need a high-end GPU to run this. The awesome thing is they've also released the data set to this. So, this contains like almost 70,000 clips from different games with 10 different controller signals. A super valuable data set if you want to generate your own video game model. If you're interested in reading further, I'll link to this main page in the description below. AI agents are everywhere right now, but honestly, most people still don't really understand what they actually are or which tools are even worth using. Well, this guide called your AI agents cheat sheet by HubSpot is a great way for you to understand the entire AI agent landscape. It breaks down the biggest AI agent tools right now, who each one is for, what they're actually good at, and the best use cases to start with. What I really like is it doesn't just explain the tool in abstract terms. It gives you practical workflows and prompts you can literally copy and paste immediately. For example, it explains the difference between normal AI and actual AI agents. Plus, the guide breaks down the four core things that make something a true AI agent: planning, tool use, autonomy, and selfcorrection. Then, it walks through seven of the biggest AI agent platforms right now. And here's the useful part. For every single tool, it explains who it's best for. It gives you some real-time use cases and starter prompts you can use immediately. It also includes a full comparison section showing the setup time, skill level, and pricing for all seven tools side by side, which makes the whole AI agent space way less confusing. So, if you've been hearing everyone talk about AI agents, but you're not really sure where to begin, this is one of the best beginnerfriendly overviews I've seen. You can access it for free using the link in the description below. This resource was made by HubSpot, the sponsor of this video. Also, this week, we have a really useful AI called Fizz X Omni. This is solving a really important problem in 3D generation. So, this AI can make objects that not only look good, but actually work inside physics simulations. Instead of just generating a car as just one unified 3D model that can't move, this can actually create simulation ready assets. That means the object has geometry, scale, material properties, and even motion understanding. So, in this example, you can see the wheels can move as you drag the car around. Or here's another example where you can see this object has accurate joints at the right place. This can be moved around very realistically. Or here's yet another example. What's interesting is that most 3D generators are either focused on appearance or they only work well with one type of object like rigid objects. But Physex Omni tries to unify all of this into just one framework. And if you compare this new PhysX Omni, which is in red, against other competitor models like Articulate Anything or PhysX Gen, PhysX Anything, which I've gone over on my channel before, you can see that Physex Omni on average performs better across all these different benchmarks. It has the highest surface area in this chart. The awesome thing is they've released everything already. If you scroll up to the top of the page and click on this code button here, it contains all the instructions on how to download and run this locally on your computer. Plus, they also released the training script and the data set for this as well. So, this is completely open source. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have a new benchmark for coding agents called Deepswe. And the core idea is very simple. Current coding benchmarks are getting too easy to game. It's too saturated and too contaminated. But Deep Suite tries to measure whether AI agents can handle actual real software engineering work, not just solve public GitHub issues that the AI might have already seen during training. So these tasks are written from scratch, spread across 91 active open-source repositories and cover different languages from Typescript, Go, Python, JavaScript, and Rust. The prompts are intentionally short and realistic, just like how a developer would actually prompt an agent. So instead of giving the model a giant checklist, it might get a compact request. Then it has to autonomously do the work itself. Explore the repo, figure out where the change belongs, implement it, and make sure it doesn't break anything. What makes this benchmark stand out is the task size. So here you can see that deep sweep prompts are actually shorter than Sweepbench, which are the current standard coding benchmarks. But the reference solutions require much more work. it needs to write way more lines of code and also edit a lot more files. This benchmark also uses handwritten behavioral verifiers, meaning it checks whether the final software behaves correctly, not whether the model just copied one specific implementation. And the results were pretty interesting. GPT 5.5 scores the highest, followed by the Claude models, and then followed by Gemini 3.5 Flash, and then the open- source models like Kimmy K 2.6 6 and GLM actually score way worse which is quite interesting. So in summary, Deep Suite is trying to answer the question developers actually care about. Can this agent survive a messy realistic coding task that requires multiple steps or working with multiple files? Can it actually explore the codebase, make the right decisions, reliably edit the code and not break the rest of the codebase? If you're interested in reading further, I'll link to this main page in the description below. Also this week, Anthropic releases their latest model, Opus 4.8. And you can see across the board, at least according to their self-reported benchmarks, it's not only better than Opus 4.7, but also better than OpenAI's flagship model GPT 5.5 in terms of agentic coding, reasoning, computer use, knowledge, and financial analysis. Although in terms of agentic terminal coding, GPT 5.5 is still the leader. Here it says that one of the most prominent improvements in Opus 4.8 8 is its honesty. This model is more likely to flag uncertainties about its work and less likely to make unsupported claims. So instead of just hallucinating and sounding confident in something it doesn't know, it's more likely to say that it just doesn't know. Opus 4.8 is four times less likely to allow flaws in its code without noticing them. And it will also push back on bad or weak plans and stay reliable during agentic workflows. Now, if you look at this independent leaderboard by artificial analysis, you can see that Opus 4.8 Max is ranked number one, but just one point above OpenAI's GPT 5.5. So, not like an insane lead. And also note that the current top open source model, Kimik 2.6, is not far behind either. Now, if you look at the price of this, then surprisingly, Opus 4.8 is a bit cheaper than GPT 5.5. Now, contrary to their claims of being honest and more reliable, if you look at this omniscience accuracy index, which measures the proportion of correctly answered questions out of all questions, you can see that Opus 4.8 Max isn't actually the most accurate. GPT 5.5 and surprisingly Gemini 3.1 Pro as well as Gemini Flash perform even better. Now, in terms of hallucination rate, then Opus 4.8 8 hallucinates the same as the predecessor 4.7, but still not as low as some open models like GLM miniax as well as Xiaomi's Mimo. And if you look at this leaderboard called LiveBench by Abacus AI, here you can see that Opus 4.8 isn't the top. GPT 5.5 is still number one, followed by Gemini 3.1 Pro. So you can see that Opus 4.8 8 is slightly better in terms of reasoning, but it's not really that good in terms of coding or math, data analysis, language, or instruction following. So, depending on which leaderboard you refer to, it's mixed results. Opus 4.8 isn't like way better than GPT 5.5. Anyways, I'll continue testing this some more, and if it's worth it, I'll make a full review video on it. If you're interested in reading further, I'll link to this main page in the description below. In humanoid robot news, Astrobot unveils their T1 robot for home use, which is actually really cheap. Here you can see it helping in the kitchen. It can also help you place clothes in a washing machine and operate the washing machine and then later take the clothing out and even iron it on an ironing board. It can also act as a bartender if you ever need one or help play with your kids and do stuff in a lab setting. This can also help you work in industrial and warehouse settings as well as you can see here. So, it can use a variety of different tools and manipulate a variety of different objects. What I think is the most impressive is the price. So, this is rumored to cost around only $13,000, which is a pretty low entry point for a humanoid robot that could potentially help you with a ton of chores at home. Although, one thing to note is that this has a wheeled base, so it doesn't walk on two legs, and it'll likely have trouble walking up and down stairs. So, this is limited to only rolling along flat surfaces. But still, it's a pretty cute and capable model to put on your radar. In other humanoid robot news, we also have this demo from Rye Institute of their Athena Zero robot. Here you can see that it has learned to juggle complex patterns and they claim that it learned all of this in less than 10 minutes of realworld interaction. As you can see from the label in the bottom left, it's actually switching between five different juggling styles. This requires the AI to completely rewire its patterns on the fly. Moving directly from one pattern to the next proves that the robot is highly adaptive and flexible rather than just memorizing a single rigid script on how to juggle. Now, this is actually quite difficult for humanoid robots to pull off. Juggling is a classic benchmark because it requires a perfect combo of high-speed software and hardware as well as real-time spatial computing and coordination. All these need to work together in sync in order to well juggle three balls at once. The robot needs to track all three moving targets. And the software needs to predict the parabolic path of each ball in real time, instantly adjusting for minor variations, for example, from imperfect throws. And the fact that it doesn't just do one variation of juggling, but five different styles makes it even more impressive. Also, this week, Stefun releases their latest model, Step 3.7 Flash. This is a highly efficient model built for realworld agents. In other words, the model isn't just trying to answer questions. It's designed to look at images, read interfaces, search the web, work with office style tools, and stay coherent across very long runs. This is multimodal, so the input can be text, images, documents, charts, product interfaces, and other stuff. So, you can get it to analyze images or automate some workflows on your browser. And it's able to handle this very well. What makes it interesting is that this is a flash model. It's supposed to be efficient but strong enough for agentic work. Here are some benchmarks for your reference. And as you can see, this even beats other open and closed flash models, and it's edging very close to even the performance of GPT 5.5 and Opus 4.7 in terms of SweetBench Pro. And then in terms of multimodal handling as well as general agentic tasks, it's also very performant and knowledgeable. The awesome thing is they've open sourced this. So, at the top, if you click on this GitHub repo and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. Now, even though it's just a flash model, it's still pretty huge being a multimodal unified model. So, the total size is around 400 GB. You'll need a DJX Spark or like multiple GPUs to run this. If you're interested in reading further, I'll link to this main page in the description below. We also have another 3D model generator this week called cube part. And this is also really powerful. This can generate an object from a text prompt, but decompose it into multiple parts based on what you specify. For example, a car needs wheels, a body, doors, and maybe a steering wheel. A robot needs arms and legs and a torso. The nice part about QART is it's able to segment everything for you. So the output is a set of separate 3D meshes, one for each part, which can assemble into one coherent object. And this allows you to immediately plug this object into a game or virtual simulation, and everything just works right out of the box. This can move around and be animated realistically. Here's another example. Under the hood, Cube Part uses a two-stage process. First, it generates the overall shape. Then it decomposes that shape into part level meshes using a diffusion transformer with cross part attention which basically lets all the parts communicate so that the final object still fits together well. The nice thing about this is it's super flexible. You can decide the number of parts yourself. For example, you can get it to generate this object with two parts or four parts or eight parts depending on what you want to segment. And if you compare this to other 3D model generators with separate parts, then this new cube part performs a lot better. It produces cleaner boundaries and also stronger geometric fidelity. The awesome thing is they've released this already. So at the top, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this locally on your computer. Note that it's quite tiny. The total size of everything is less than 10 GB in size. So, this should be able to fit on most consumer GPUs. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, Google unveiled something called relightable hollowported characters. This is a method for capturing a moving human and then placing that entire person into a new scene with realistic lighting. So for example, you can capture this woman moving around with multiple cameras and then it would create a full body avatar of her and then you can place her in different environments. You can change the lighting warmth, the brightness and also the lighting angle and the avatar should be able to blend in very realistically with any environment. Here's another example where this is the input. You need four cameras capturing this person moving and then afterwards you can plug this person into any environment with different lighting conditions. Currently, it says the code is coming soon. They haven't released anything yet, so I'm not going to spend too much time on it. If you're interested in learning more, I'll link to this main page in the description below. Also, this week, we have this project called self-improving language models with birectional evolutionary search or bees for short. Now, this self-improving part is a bit misleading, but this evolutionary search mechanism is pretty interesting. So, here's how it works. Instead of asking a model to just keep sampling answers until one works, bees helps it search smarter in both directions. On one side, it does forward search where the model tries to build possible solutions step by step. But instead of only following the most likely path, it can also mix and recombine partial attempts almost like evolution. So it can discover solutions that a normal single rollout would probably miss. Then on the other side, it does backward search as well, where the original goal gets broken down into smaller subgoals. Think of it like solving a hard puzzle by both building possible answers from the bottom up and breaking the final answer into smaller clues from the top down. This matters because a lot of current model improvement methods depend on sparse feedback. In other words, the model only finds out whether it got the final answer right or wrong. But this evolutionary mechanism gives the model more useful guidance along the way. And in experiments, it shows real gains on challenging post-training tasks where other methods failed to improve. It also showed better performance on open problem-solving benchmarks. In the top right corner, they have released the code to this. And here it contains instructions on how you can run this yourself. If you're interested in reading further, I'll link to this main page in the description below. Also, this week we have a new agentic framework for automating scientific research. The funny thing is this week I already published a full video on two other AI agents that have automated scientific discoveries. Both papers were published in nature and both of these came up with some genuine medical breakthroughs such as new treatments for cancer, vision loss, liver fibrosis, and more. In fact, see this video if you want to learn more. But coincidentally, one day after my video was published, another agentic framework came out. So this is a new research system where AI agents don't just run one experiment at a time. They organize themselves into research teams. They explore different scientific ideas in parallel and keep improving over long runs. So this system kind of works like a small decentralized research lab. Every agent can read the current state of the project, see what experiments have already worked or failed, and help test new ideas. And this is actually one of the biggest problems with automating research is that real science is messy. You usually don't know the right direction in advance. Some ideas start strong and then they hit a wall. Other ideas look boring at first, but suddenly become useful. So, the important part is not just generating hypotheses. It's being able to organize all this chaos and keep track of your progress. So, Autoscientist has a shared state that all agents can see, including the current best solution, the experiment log, a discussion forum, and also deadend registries, which are basically records of ideas that did not work. So, the system doesn't keep wasting time on these. Each agent repeatedly does the same simple loop. It reads the shared state, decide on what to do next, and then takes action and writes the results back. Some agents act like analysts, so they would read past experiments and forum discussions, write hypothesis notes, and keep track of all the dead ends. Other agents act like experimenters, so they would actually propose an experiment and then apply the code change or train and evaluate it and then report the results back to the shared state. So think of this as like a group of researchers working together. And the results are pretty strong. So on this bioml bench which includes 24 biomedical machine learning tasks across imaging, drug discovery, protein engineering and more, autoscientist was able to beat other agentic frameworks. If you scroll up to the top, they've already released the code to this. So if you click on this code button and you scroll down a bit here, it contains the instructions on how to set this up and run it on your computer. If you're interested in reading further, I'll link to this main page in the description below. Now, Nvidia is not done cooking yet. They also unveiled another project this week called Gamma World. This can essentially generate simulations of multiple agents playing the same game at once. And this is a big step because most interactive world models are designed around one player or one controllable viewpoint. But Gamma World is built for scenes where multiple players or robots are moving at the same time, all affecting the same shared environment. The hard part is keeping each agent independent while still making everything consistent. Each agent needs its own identity and control and the model can't fall apart when the number of agents changes. So, Gamma World handles this with something called a simplex rotary agent encoding, which is basically a way to give each agent a distinct signal without locking the system into a fixed number of players. And specifically in terms of specs, this can generate real-time videos at 24 frames per second and it can generalize from two to four players. So here are some examples of two player simulations. And here are some examples of the same environment but with four players. Now at the top they have released a GitHub repo and here it says the code to this is coming soon as well as the training scripts and the data set preparation tool. So it looks like they're planning to fully open source this. If you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have a really cool AI called Pantheon 360. This is a 360° video generation model for building digital twins. So, this takes several 360° images plus a camera path, and it's able to output a highquality panoramic video that stays consistent as the camera moves through the scene. The reason this matters is because normal video generators only see a narrow field of view. So, you can't really generate a full 3D panoramic video just from a regular video model. The AI also needs to stitch together lots of partial views and things can drift or become inconsistent. But what Pantheon 360 does is it takes these 360° images and then it reconstructs a 3D point cloud model. So, think of this as like a rough geometric reconstruction of the scene and it uses this to keep the video stable and grounded as the camera moves through the path. So, this is especially important for creating digital twins or training robots or autonomous driving. Basically, any situation where you want the scene to stay consistent the whole time. Now, if you scroll up to the top of the page, it does say the code is coming soon. So, it looks like they are planning to open source this, which is fantastic. For now, if you're interested in reading further, I'll link to this main page in the description below. Also, this week, we have a new image generator that can run locally and offline on just your phone. So this is called Bonsai image and here are some example generations for your reference. Now they released two different variants. There's a one-bit variant and a turnary variant. The one bit model uses two possible values for transformer weights whereas the turnary model uses three possible values which gives it a bit more flexibility and better image quality. Both of these are designed to run locally on just an iPhone or similar consumer devices. Now, it's important to mention that this is actually just Flux 2 Klein. So, it's not like they've created a new image model from scratch. However, they did take Flux 2 Klein, which is originally almost 8 GB in size, and they compressed it down to just around 1 GB. So, they significantly shrunk the memory and compute required to run one of the best open-source image models out there. So, here's an example of Bonsai image in action on an iPhone 17 Pro Max. And here they say that it can generate a 512 x 512 image in just 9.4 seconds. So, if you are interested on running an image generator offline on just your iPhone, this might be a good option for you. On the right sidebar, they've released a GitHub link which contains all the instructions on how to run this. If you're interested in reading further, I'll link to this main page in the description below. Next up, we have a tiny but surprisingly capable model built for local deployment, especially for smaller devices. So, this is called mini CPM51B by OpenBM. And this is just a tiny 1 billion parameter dense model. But if you compare this to similarsized models, note how performant this is. So, across benchmarks on general knowledge, coding, math, logical reasoning, and agentic use, it outperforms other competitors. The ridiculous thing is this is only 2 gigabytes in size. So you can run this on like most laptops or even potentially a mobile phone. Now on this page, it contains all the instructions on how to set this up locally via all these different platforms. If you're interested in reading further, I'll link to this main page in the description below. Also, this week we have a new method to create highresolution images. This is called Sega, and here are some of its sample generations. You can create really high resolution images. So, for example, this one is like over 4K and then this one is a 4K image. Everything is super detailed and sharp. Here are some other examples for your reference. Notice the water droplets and the sharpness of the whiskers. This is really good. And here's another result for your reference. Again, a very detailed image, even if you zoom way in. And it's hard to spot any errors or flaws with this. Very impressive. Now, this doesn't just have to use Flux as the base model. This also supports Quen. So, here's an example using Quen image. And again, if you zoom in on the details of the buildings and everything else in the scene, it's just really sharp and accurate. And check out the resolution of this. This is like 6,144 pixels on each side. A massive image. And if you compare this to other upscalers like DYP, which I've gone over on my channel before, notice that Sega is a lot more consistent and less errorprone. This is currently one of the best methods to use to generate high quality images. If you scroll up to the top, they have released the code to this. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to run this either for Flux one or Quinn image. Note that there's no indication whether they will also roll this out to Flux 2 or Z image. If you're interested in reading further, I'll link to this main page in the description below. Also this week we have Pixel Relights. And as the name implies, this is a new image relighting system that lets you take a single photo and control the lighting in whatever way you want. For example, if this is my original photo, I can just drag my cursor around to change the light from any angle. It's able to detect everything in the scene, including the tables and the chairs, and relight everything realistically. Now, instead of just a flashlight, I can also turn this into a spotlight to make the light harder and more narrow. And here's my result. Or here's another example. Again, a very messy scene with a ton of different objects, but it's able to understand the location of everything, even though this is just a 2D image, and relight everything according to where I place my cursor. Here's another example. So, a very performant tool that is able to understand where things are in the scene and relight everything accordingly. Here's how it works. Basically, it takes your image as input and then it estimates a rough 3D understanding of the scene and then that 3D scene is actually plugged into Blender where the user can relight the scene and then it uses that as reference to create the final reit prediction. Now, at the top of the page, they have released the code to this. So, if you click on this code button and you scroll down a bit here, it contains all the instructions on how to download and run this. If you're interested in reading further, I'll link to this main page in the description below. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this. Which piece of news was your favorite? And which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay uptod date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next
Original Description
HUGE AI NEWS: Claude Opus 4.8, Step 3.7, DeepSWE, MiniCPM5, & more #ai #ainews #aitools #aivideo #agi #singularity
Thanks to our sponsor Hubspot. Access the AI Agent Cheat Sheet for free https://clickhubspot.com/a71cc9
LocateAnything https://research.nvidia.com/labs/lpr/locate-anything/
ControlLight https://yfyang007.github.io/ControlLight/
TriSplat https://lhmd.top/trisplat/
PiD https://research.nvidia.com/labs/sil/projects/pid/
InstructAV2AV https://hjzheng.net/projects/InstructAV2AV/
GenRecon https://kasothaphie.github.io/GenRecon/
Scope https://z2tong.github.io/SCOPE/
PhysX Omni https://physx-omni.github.io/
DeepSWE https://deepswe.datacurve.ai/blog
Opus 4.8 https://www.anthropic.com/news/claude-opus-4-8
Step 3.7 Flash https://static.stepfun.com/blog/step-3.7-flash/
CubePart https://cubepart.github.io/
Relightable chars https://vcai.mpi-inf.mpg.de/projects/RHC/
Self Improving AI https://guoweixu.com/bes/
AI coscientist: https://youtu.be/QvN6Tu6dHYM
AutoScientists https://autoscientists.openscientist.ai/
Gamma World http://research.nvidia.com/labs/sil/projects/gamma-world
Pantheon 360 https://koi953215.github.io/pantheon360_page/
Bonsai Image https://prismml.com/news/bonsai-image-4b
MiniCPM5 1B https://huggingface.co/openbmb/MiniCPM5-1B
Sega https://rajabi2001.github.io/sega/
PixlRelight https://mlfarinha.github.io/pixl-relight/
Timestamps
0:00 AI news intro
1:03 LocateAnything
2:42 ControlLight
4:22 TriSplat
5:58 PiD
7:51 InstructAV2AV
9:05 GenRecon
10:25 Scope
12:12 Hubspot Agent Cheat Sheet
13:35 PhysX Omni
15:08 DeepSWE
17:10 Opus 4.8
19:45 Astribot T1
20:53 Rai AthenaZero
22:06 Step 3.7 Flash
23:33 CubePart
25:18 Relightable chars
26:11 Self Improving AI
27:43 AutoScientists
30:19 Gamma World
31:44 Pantheon 360
33:00 Bonsai Image
34:26 MiniCPM5 1B
35:14 Sega
36:46 PixlRelight
Newsletter: https://aisearch.substack.com/
Find AI tools & jobs: https://ai-search.io/
Support: https://ko-fi.com/aisearch
Here's my equipment, in case you're wondering:
Lenovo Th
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
More on: Staying Current in AI
View skill →Related Reads
📰
📰
📰
📰
What smart people are saying about IBM’s AI warning and SaaSpocalypse fears
Dev.to AI
Dutch company ASML is $300bn from a trillion. AI could close the gap
The Next Web AI
ARR 2026 Meta Review score [D]
Reddit r/MachineLearning
The AI Debate Isn’t New. History Has Heard It Before.
Medium · AI
Chapters (25)
AI news intro
1:03
LocateAnything
2:42
ControlLight
4:22
TriSplat
5:58
PiD
7:51
InstructAV2AV
9:05
GenRecon
10:25
Scope
12:12
Hubspot Agent Cheat Sheet
13:35
PhysX Omni
15:08
DeepSWE
17:10
Opus 4.8
19:45
Astribot T1
20:53
Rai AthenaZero
22:06
Step 3.7 Flash
23:33
CubePart
25:18
Relightable chars
26:11
Self Improving AI
27:43
AutoScientists
30:19
Gamma World
31:44
Pantheon 360
33:00
Bonsai Image
34:26
MiniCPM5 1B
35:14
Sega
36:46
PixlRelight
🎓
Tutor Explanation
DeepCamp AI