Keynote: Claw and Order: Protecting Your Shell
Skills:
AI Security80%
Key Takeaways
Discusses protecting AI systems from security threats using Zero Trust principles
Full Transcript
When I first proposed this topic to to Rob, this was early February. And you know what? It's April. And Open Claw is so January. Okay? So, that's why I had to really kind of redo this whole presentation. The topic is still on Claw and Order, but you'll see a little bit of a shift in how I thought about what I wanted to present today. So, I'm going to tell you the story starting with a fable that I think hopefully most of us all know, about the three little pigs, right? And if you consider the story, like the main takeaway, what like think about the what what is the main takeaway that you have from the story? Now, I'm a security guy. So, the my my main takeaway was, "Hey, it's really good to have threat models." Okay? Now, if most kids probably don't grow up thinking that. I wish more people did because then maybe we'd have more secure software. But, the other takeaway that most of us leave with after the three little pig story is we should build with bricks. Okay? That's generally the story that what most people remember out of this story. But, that's actually the wrong conclusion. Building with bricks isn't really what you should come away with, but rather the notion of architecture. It's not just about the materials. I can actually build a straw house that is hurricane resistant. And there are truly straw bale houses that are hurricane resistant. And at the same time, there are brick houses that you just push over or you have a puff of wind and it'll fall over. The difference here is architecture. It's not materials make a difference, sure, but architecture sometimes matters more than just the materials. And in this particular context, when we think about AI systems, this is also quite true as well. So, let's consider there's a quote here from Ludwig Mies van I'm not sure how to pronounce his last name. But he says architecture starts when you carefully put two bricks together. Careful and I emphasize the word carefully here. Because you can just put two bricks together, doesn't really make much of an architecture. And I'll give you an example. And of course, you know, I'll start with Mythos cuz that's in the news today. And we of course know the story about it finding all these different vulnerabilities. So, that's a stronger material, right? It It You have this amazing model, finds all these vulnerabilities. But let's look at the architecture. What is the architecture that led to the these vulnerabilities being discovered? Well, according to Nicholas Carlini who's also at Anthropic, this was the prompt. You are playing a CTF, find a vulnerability, write the most serious one to this report. That was it. Meaning that there was no architecture. It's a It's about as basic as you can make it, right? And yet it discovered a bunch of things. Okay? But consider now, what if you actually gave it just a basic architecture? And so I wrote this blog about a week ago and said, here's a really really simple architecture. I mean, it is it's part of part of something they call nano analyzer. And it is about as simple as possible, which is basically carve up all these files, send them one at a time, and have it there's some additional filtering that they do they do, but basically it's super simple. And so they put this basic architecture into the equivalent of wood, the equivalent of bricks, okay? Or the equivalent of straw, and they basically found similar the same vulnerabilities in OpenBSD and FreeBSD. Not always, not consistently, but they still found the same vulnerabilities. Okay? So, in other words, the architecture um allowed us to compensate for the weaker material. Moreover, when you're consider the cost, the uh token cost for uh Anthropic mentions that the token cost to find the OpenBSD vulnerabilities was $20,000 in tokens. Yeah, I also found this with less than $100 of token cost. Okay? Um so, now I think this is what is part of the wake-up call, right? If a basic with with no architecture and strong materials, you find these uh vulnerabilities, what can you do when you put even a basic architecture with a strong material, right? And you end up with why if you're if you wonder why many of us are calling this a cop clips or uh I think my friend was my or condominium and heart burner, there's a reason why it's because when you start combining basic architecture or scaffolding, okay? Jacob mentioned this earlier, finding the scaffolding is hard, but once you have the scaffolding, you're able to do some pretty amazing things. So, you have basic scaffolding, you put it together with stronger materials, and we're going to see some interesting waves of vulnerability reporting. And that's just with basic architecture. What if you actually use even stronger architecture? So, uh my co-founder and I, Gadi, we uh developed something called OpenAd and open-sourced it, but it's a it's a uh more refined architecture that helps us find vulnerabilities. What if you couple that with a stronger materials, and okay, we I don't want to I can't I don't want to even imagine how bad it could actually be. Okay? So, in the context of is are the concerns around mythos overhyped? I actually I mean, I firmly do not believe so. In fact, it's quite under hyped. Because Anthropic founded with no architecture, no scaffolding. Just go I said, go find vulnerabilities. Imagine what can happen if you have even basic architecture or stronger architecture. And you couple that with the model. And this overall pattern also manifests itself in in other interesting ways. So, consider um uh what we've seen over time is this progression of better architectures, better scaffolding, better harnesses, and so on and so forth. So, when Claude code first came out back in February 2005 25, it ran under the equivalent of straw. Okay, so on a 3.7. Uh it was a great model, but from the standpoint of of uh of materials, it wasn't that great of a material, but it started the whole vibe uh coding sort of journey. Okay, Andre Karpathy uh he said, "Hey, this is kind of cool. We can start doing vibe uh vibe coding." But it didn't produce the best code out there. But it was it was basically functional. It provided great uh great prototypes and such. Then Opus 4.0 comes out in May, 3 months later. And now you're building with wood. And you can build uh bigger applications. And not quite production ready, but still pretty good. And then Opus 4.5 comes out. Now, that's like building with brick. And so you can build bigger buildings, you can be bolder. People are putting out production applications with this code or with this uh material. And as we see um as as each of these different materials become available, they what we learn from the point from February 2025 to day minus one of Opus 4.4 4.4.0, what we have developed over time is better and better building codes. Okay? So, the straw house that we built on the first day of uh Summit 3.7 and the last day of uh Summit 3.7, the house actually got better because we learned architectural patterns, and we shared them, and we said, "Hey, this is how we build better code using straw." Then 4.0 comes out. Some of the building codes are no longer useful. We jettison those, but we build new building codes because we can build bigger buildings. Repeat that cycle for 4.5, and as we continue to build out better building codes, we can build better and better applications. Now, imagine what happens with Metos. You're building with uh steel now. And you can build And it's interesting because this is also the height of building is almost like an exponential curve as well. Uh never before have we been With all the materials in the past, we haven't been able to build buildings that are uh almost uh half a mile high. And yet, with these uh materials, we can. Um but along the way, we're learning new building codes. We're learning new architecture. We're learning new scaffolding. And as we share those collectively across the community, we learn how to build better. And the sharing of that is happening at uh speeds that you can't imagine. It's happening so fast that you have Andre again who uh said um who who coined the term vibe coding, he's saying, "I've never felt this much behind as a programmer." Here is a This This is a It's a remarkable statement because this is one of those folks who are truly at the leading edge, okay? And he says, "I've never felt this behind." Why is it people not uh so behind? Well, to uh say what he says is he feels like I have I could be 10 times more powerful, more productive if I could just properly string together all these things. Permissions, agents, modes, workflows, LSP, subagents, uh slash slash commands, plugins. I mean, there's all these things that are uh part of the scaffolding. And I I I don't know what like uh you know, a third of these things well what they actually are or how they fit into my workflow. And so, Andre is saying, uh I feel so behind because there's so many different parts of this ecosystem that that I wish I could figure out how to tie together. And so, he says, there needs to be a programmable layer of abstraction to master it and need to build an all-encompassing mental model uh because we have this powerful alien tool that uh comes with no manual. And we all have to figure out the architecture of how to hold and operate it. We have to figure out these design patterns. We have to figure out the scaffolding. Okay? So, he says this in December of 2025. So then, what happens after that? Well, you have Open Claw. And Jack So, it was built uh back in the November time frame when Andre was saying this prob- is having this problem. And he's he's all essentially the need for something like Open Claw. Um and to give you a sense of how things how fast things move, it was the fastest-growing GitHub project um as of March um like as of like 2 weeks ago. Um and it quickly got overcome by another project uh that I'm sure Jacob um doesn't like, which is it's a it's an open-source version of Cloud Code. Um thank you for the leak. So, anyway, that So, that That is now the new fastest-growing project. Um no surprise, right? Okay. So, uh we look at Open Claw and we're like, "Okay, well, what what does the Open Claw do?" Well, it it solves some of the problems that um Andre mentioned. Uh well, first, it's just a it it does stuff, right? It's an it's an actual assistant that does things. Um it does just doesn't just give you an answer. It's not just a chatbot, but it actually goes and does things for you. But moreover, it was meant to be simple to operate. So, the uh and and moreover, Peter who put OpenClaw together, created this mental model, created this sort of construct that brought together all these different pieces, and then made it simple to operate. He gave you a manual. And that is why I think it took off, right? I think that's what really triggered this whole movement towards using these sort of sort of AI assistant. Um so, let me now let's let's look through the bricks. Like, what what are the different how do we assemble these things? What are the what are the components that make up the OpenClaw? And so, he said, "Okay, here are the bricks, the different material the different pieces that I assemble to make OpenClaw. Um I won't go into all these details, but I'm going to focus on one of them, and it's the sole, sole.md. And it's pretty cool because it actually articulates a job function. It says, for example, "I need a risk assessor." And there it's a pretty straightforward statement of what you'd expect a job This is like a job description for a risk assessor. And so, you'd give it to an OpenClaw agent, and it does this function. Um there's other ones. There's tons of other ones. Like, you for example, you have a SOC 2 preparer. So, instead of using Delve, you can turn over to OpenClaw and have it do for do it for you. You might even get better results. Um I I fly a lot, so I I use like a flight scraper. But, the point is that the you have a wide range of different types of functions that you can assign to these OpenClaw agents. And this is what makes it super popular. But, it also makes it more challenging in terms of how to secure it. So, what we see then is statements like, "Hey, OpenClaw is groundbreaking." And it was groundbreaking because again, it addressed something that was missing in the ecosystem, as Andre had mentioned. But, from a security standpoint, it was an absolute It is uh it was an absolute nightmare. And you're basically turning over a lot of data to OpenClaw, and at best uh it's unsafe, and at worst, it's utterly reckless. Now, I have two caveats. First is, these were the early days, okay, of OpenClaw. My second caveat was is that that's that's only 90 days ago, okay? Early days was only 90 90 days ago. Again, this is why I say, "You know what? I OpenClaw is so January. Um there's so many other things that have emerged, but since that was the topic, I'll go ahead give you some quick pointers on how you secure OpenClaw." Um because really, despite these concerns, the claw has already left the bank. Uh you probably have it running inside your organizations. And if you're not running it yourself, just to understand the nature of this uh of this piece, you're you actually should, right? It certainly in a way that is uh properly secured. And one of the things we noticed is that the it didn't quite come out secure out of the box. And so, uh one of the things I I um used as a way to think about the risks associated with OpenClaw is to apply what's called the agent rule of two. So, the agent rule of two is from Meta. Um they identify three different types of things that you want to be concerned about. Uh does it trust does an agent process untrustworthy inputs? Does it have access to sensitive assets? Does it uh connect uh change data connect um externally? Basic premise is uh pick two of these, don't give it all three. Unfortunately, well, when you and you when you give it all three, you're in the danger zone, and unfortunately, OpenClaw by default is right in the danger zone. So, what do you do then? You have to provide some sort of controls, uh but ultimately, the challenge is uh the way that OpenClaw was designed, it just gives access to tons of sensitive stuff. It's uh you have this ClawHub, which pulls in skills from who knows where. You have um it's it by default it was exposed to the internet, And it can certainly do very destructive things on your computer like deleting your hard drive and whatnot. And so, uh what we've seen over the past um 90 days is uh a way to restore law and order. Okay, we have tools from NVIDIA that um introduce a couple of design principles which I'll talk about in a moment moment that help uh uh secure these by default. We also have something from Cisco uh where they released something called Defense Law um that builds on top of uh NVIDIA's um o- open shell. And then uh w- I also uh my my company Nostic also released a couple things to help us find these uh open claw instances, to monitor the for those interactions, and also uh protect yourself from stupid when these open claw agents do something stupid. Overall, uh as I looked at all these different uh patterns, what I This is what I observed. So, if you're trying to safeguard open claw from architectural standpoint, these are all the different elements that you want to put in. Um you want to process if you are going to process untrustworthy inputs, make sure that you uh one have trusted sources that you pull skills from. Make sure you scan them. Make sure you have allow uh allow list and block list. Uh when it comes to sensitive assets, you want to make sure you uh look for uh any leakage of those sensitive of that sensitive data. This principle of least pri- uh least privilege um it's hard, especially given that it's operating with your permissions. Uh but what I've done is I actually um w- one of the reasons why people bought a bunch of Mac minis is because you can kind of isolate that into its own sort of sandbox. And if it blows up, ah not a big deal. It still may have access to a lot of sensitive stuff, but that's up to you to make decisions on. Um and then on the connect state uh connect externally or change state, um there are various guardrails that have been introduced, uh including ways to uh isolate the kernel and uh to have a default deny um access to the internet. And so, these are some of the things I just showed you whether from Nvidia or from Cisco or the tools that we released really we're trying to provide these sort of controls by as it pertains to the agent role of two. What I also liked about what Cisco produced was they also have this audit logging. So, what's what how's how's Open QA interacting? Is there any sort of audit logging that helps us see what happened? They introduced that into their their framework as well. Um So, overall, though we have to recognize So, if you consider if you look at the the agent role of two, it doesn't say no risk. It just says lower risk. Okay, there's the danger zone and then it says lower risk. And Pete, the guy who created Open QA, says, "Look, there's no perfectly secure setup here. Okay, you are taking a risk, um but if you want to lower your risk, then just be deliberate about each of these different things. Be deliberate about deliberate about who your bot can talk to, what they can touch, and upon what area it's allowed to act." So, as we look at this as a whole, what we want to look at then is what does the future look like? Okay, so we see the Open QA is really just a wake-up call to a lot of enterprises that these agents are coming, but it's moving properly again at a pace that's faster than most organizations are comfortable with. Many organizations are still at the chatbot level. Um but then eventually people move towards being able to have these chatbots call tools, eventually operating more persistently, and then eventually going towards managing a whole agent fleet. If you recall what Jacob showed, this is also the progression that attackers are taking, too. Okay? And when we think about uh how do we capture how do we get ahead of the attacker? Well, one way is to see what the progression looks like here and get ahead of the attacker in this way as well. If you and your security function are still at the chatbot level, um you have to move pretty fast up this sort of uh chain. Oh, you know, one thing that I learned um at Unprompted, one thing that was encouraging uh that I learned at Unprompted is we look at all these people who seem to be really far ahead when it comes to the use of AI, um what I was encouraged by was I felt like I was only behind by 2 months. Okay? Yeah, really. Uh the space is moving so fast that anybody who feels like they're really ahead, um one thing that's really kind of fascinating is you heard Jacob mention in terms of the scaffolding for the attackers, that scaffolding is easy to share. That scaffolding and how we build these things is now just a markdown file. It's a it's a sort of um architectural blueprint that is easy for us to just adopt and start using and understanding. I don't have know if we necessarily need to understand all the nuances of the architecture, but adopting it is actually much much much more easier than it was before. And so, when we think about what's happening next, there are a couple things that I've observed. One is that um how we interact with agents is now wherever we want. I interact with Cloud Code on my phone, uh and I tell it, "Hey, go build this." And it goes builds it and I eventually I might come back to my computer and I'll look at it, but whether it's WhatsApp or Telegram or uh uh a phone app of some sort, we're moving really to this model where um I don't really need a computer. I don't need to be in front of a computer to get a lot of work done. Uh second thing is that as more of these agents start um operating, as you have more computer using agents, what you'll start seeing is more uh interfaces that are designed for agents. And they'll still be designed for humans, but those people who love like shortcut keys on websites and stuff, more of that will become available because computer using agents love that kind of stuff. We also see the the the agent really learning and keeping memory. That's what one of the key architectural things that we saw with Open Cloud that was really cool. Um and then the shifts I mentioned there's a capability will shift more away from the LLM itself, the material, and more towards those scaffolding and design patterns that I mentioned earlier. The what we how we build the scaffolding around these LLMs is LLMs that have for which we don't really have great confidence that they're they have integrity or that I think think about as you build a building, you have parts that are made up of you have flawed parts, okay? And these flawed parts you're going to put together in such a way that despite their flaws the scaffolding, the architecture or harnesses really help safeguard against catastrophic failure. And so capability will shift from the large language models, which will be flawed in some way or another, which will be which will hallucinate, which will be poisoned, which will have all these sort of issues, and we'll move it from just the LLM to the overall architecture that surrounds your models, your tools and skills, and so on and so forth. But what I think is most interesting about what happens next is what happens to us. Okay? Because there is going to be no more individual contributor. If you're an IC today, you now are a manager of 10 agents. And if you're not managing 10 agents, you should go find and create 10 agents to manage. And I think the perspective here is really um a shift in how we think about our operating models as a whole. So there's an old saying which is cloud is just somebody else's computer. You can think of that as um uh you can think of cloud as it's just somebody else's computer or it's a new operating model. In the same way, you can look at AI or genetic AI as just more automation. Or you can think of it as a new fundamental rethinking of the operating model. But what is this operating model? Okay? And what I discovered is that if we if you consider you have you've now grown as an organization. And whatever forget how many people you have in your employee population, multiply that by 10, minimally. Maybe 20, maybe 100. Okay? What does an organization 10 times your size look like? Well, you have to fundamentally reorg. Okay? A productive individual doesn't make for a productive organization. You have to find a ways to change the organization itself. Um and when we do look at the uh Jensen Huang's comment about every company in the world has to have an open AI strategy or open cloud strategy, it's really a genetic AI strategy where your organization has grown by 10 times. And when we look at that, what what one of their one of the things that's really fascinating is you may not be familiar with what an organization 10 times your size looks like, but somebody else is. Okay? If you have a 300-person organization, there exist 3,000-person organizations. So, look at their organization model, their org chart, and see what you can adapt for your organization. And for you as an individual, if you're an individual contributor, most ICs when they first start out, they don't get a chance to manage people, you know, for obvious reasons. But you can manage and learn management skills today. And as you as you have as you grow in your organization, there's a concept called Dunbar's number that causes you to say, you know, at certain thresholds, you need to really change how you operate. We've exceeded that threshold. You've already exceeded that threshold, but you still have an organization that hasn't changed. And so that's the thinking that you need to have. That's the new operating model. It's not a new unless you run an organization of 300,000 people or or a million people, um you probably have not experienced what it means to run an organization of that size. But now you do. Now you have to. And what I found was interesting with uh this is a quick view of uh Autodesk uh growth from uh 2007 to 2011. But what I found fascinating is the constant war changes. Because as a company grows, it has to adapt. The architecture of the organization has to change. And that's really fundamentally what I I wanted to leave behind. That when we uh use when we look at architecture, it puts order into things. But the order will constantly change because the materials change and because the space is moving so fast. So with that, thank you very much.
Original Description
Keynote: Claw and Order: Protecting Your Shell from Bottom Dwellers
🎙️ Sounil Yu, Co-founder and Chief AI Safety Officer at Knostic
📍 Presented at SANS AI Cybersecurity Summit 2026
The rise of OpenClaw hints at the pent-up demand for a truly useful personal AI assistant. Despite all our efforts to create layered security boundaries and implore adherence to principles like Zero Trust, it seems like all that was thrown out the window as millions rushed to install OpenClaw and experience an AI personal assistant that actually does things for you.
What can we learn from this that will help us not just design better security but actually have people use it when the next AI innovation arrives at our doorstep?
Explore upcoming SANS Summits to continue learning from leading voices in cybersecurity: https://go.sans.org/summits
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
More on: AI Security
View skill →
🎓
Tutor Explanation
DeepCamp AI