Why AI engineering needs old-school discipline

The New Stack · Beginner ·🤖 AI Agents & Automation ·2mo ago

Key Takeaways

AI engineering needs old-school discipline with a systems-thinking approach

Full Transcript

[music] >> ThoughtWorks is a leading global technology consultancy that [music] delivers extraordinary impact by blending design, engineering, and AI expertise. For over three decades, [music] we've led in technology innovation, and today we're at the forefront of AI-powered [music] software and data engineering. Welcome back to another episode of The New Stack Makers. I'm Frederick Lardinois, the senior editor for AI at The New Stack. And I'm Nimisha Asthagiri. Um at ThoughtWorks, I'm a data and AI advisor. Well, Nimisha, thank you for being here. Thank you, Frederick, for having me. Absolutely, it's a pleasure. Now, one thing I was thinking, ThoughtWorks, I've heard the name, but I'm thinking quite a few folks in our audience may not have heard it yet. They may not be familiar with what you're doing, so maybe give us a short backgrounder on on what it is that ThoughtWorks does. Yes, of course. So, ThoughtWorks, we're a global digital consulting firm. We help our clients around the world with various things, you know, from strategy, design, engineering, and these days really AI. And our phrase is, you know, it's not just AI, but AI that works. And why that is is really taking up taking our decades of experience bringing you know, thought leadership to the industry and bringing the latest and greatest ways of thinking about building software, building it right, and building the right things. So, everything from our product thinking capabilities and and and skills, as well as you know, agile and engineering and platform engineering practices. What is it in your experience right now that these companies are looking for specifically? Like what what are they struggling with still? Well, I mean, these days there is so since 2023 and 2020 November 2022, right? This whole wave of generative AI and then the next, you know, year of doing a lot of proof of concepts and you know, really exercising it trying to now upskill the the their their own employee force. Um companies are looking for understanding how do they actually bring a lot of their proof of concepts to production. They're also looking for how do we get our the development work that we're doing, so it's not just accumulating code, but it's actually viable code and code that can really, you know, execute and run successfully. So, you know, I think Gartner is saying like 40% of agentic projects will be canceled by 2027. So, these are staggering numbers and they just imagine the billions of daughter dollars that are being put into the industry and then therefore waste in some ways. Not a waste in terms of learnings, there's been a lot of learnings and rapid learning with the pace that is there in the industry, but definitely in terms of business ROI, one may not actually achieve that. Mhm. No, absolutely. I And I keep hearing the same numbers kind of. I don't think it's really the needle hasn't moved there all that much, but what is it that these companies are getting wrong? Why are 40% of these projects getting canceled? Why aren't we better at, you know, getting value out of these tools which are, you know, it's really powerful. I I think it's the question that is being asked. The question that we're hearing a lot from executives and others is how do we go faster? How do we go faster? How do we keep relevant? Where I think the question the right question or another alternative better question here might be what do we build given the latest technology that we couldn't build before? So, reimagining and reinventing what we couldn't do before, and that's where we focus our energy rather than everything. And then secondly, how do we build differently? Now that we do have AI and it's not just as an assistance, but actually even a you know, part of the team, you know, with human and machine agents working together, what what does that look like? And that is not therefore just a tool change, but a systemic change a systems change. Rethinking your models of of what to build and how to build. And that that also then implies how do you measure? And and the measurements from the in the past where we might have just measured output and production of of, you know, how quickly are people bringing in pull requests or, you know, code into production. Instead of that, it could be more about, for instance, iteration cycles and interactions with with your with your AI and with others. Those collaboration and interaction metrics that that are also how your first pass acceptance rate, right? How often is the AI that code that's generated, how quickly can it be accepted and with minimal rework? >> [clears throat] >> As you mentioned systems thinking there, talk to me a little bit more about that. How what needs to change there? How you know, it's a both a change that's practical, but also a change kind of in the the culture, I think, of how these projects develop over time. Exactly. Exactly. So, I think it is the, you know, the the typical people process tech coming together, but I think, you know, when we're when we're think rethinking this, I think there's a lot of um uh thinking about the organization at large. And you're thinking about the technical platforms that you have in your organization and what platform capabilities you may need to bring in that become strategic assets. And so therefore that then supports the system at large. And you're using therefore your platform as paved roads that people can use to accelerate their work, but also becomes that the the the conduit to the governance that you also need to apply. Uh and the efficiency gains. So, all of that kind of coming together and for ThoughtWorks, like we've we've been progenitor of of a lot of platform and engineering excellence uh you know, capabilities. And that is now once again coming to the forefront. Where like, yes, it's there's a lot of fundamentals from the past that need to be actually and at this point reinforced if anything. Um and and so then you start thinking about if you have that harness and the system technical system in place, then that's where now you're elevating the people to say, "Hey, use more of your human judgment. Let the other things that we have been jaded in the past of like repetitive work that we may have just that's what that's what I do when I come to work, but kind of rethink that." So, you there's a little bit of unlearning so that you can learn this new way and where your human judgment and the higher order you know, value of your human effort can come into play. Yeah. Yeah. In practical terms, as we're thinking about some of these kind of standards that we've had for many years, like what is it that we need to bring back there or focus on? You know, what are some of the examples that you've seen? >> Yes. Yes. So, for instance, it might be everything from like mutation testing, right? So, test-driven development, right? So, we're really thinking about creating those feedback sensors for AI. So, that the AI when it's auto now with autonomous coding agents. And there's been a drastic change, right? In December with the latest models and so forth. So, and there's a lot of therefore keen interest in generating and designing these autonomous coding agents. But you want to what do you want to bring back? So, I think there are things like mutation testing, also a lot of our testing principles. Um and and there might be things with test-driven development, for instance, designing that tight feedback loop for the AI so that it can continuously learn and evolve. Um so, that being a key component. I think the the other thing would be even the metrics such as even like Dora metrics. Right? So, I think those types of things with deployment frequency, lead time, change failure rate, these are our lagging metrics that will then ensure that, okay, yes, we are moving the needle in the right place. But I think there's also other things like uh zero trust security architecture, right? And really thinking about the uh ensuring that we have proper identity management and security as well. So, when we're seeing the the the propagation of uh agents and now they're now they're they're coexisting along with you when you're working on your laptop and your desktop and, you know, a lot of that um uh uh changes that are happening, there is zero trust architecture is critical. And, you know, being able to know who did what, as well as the authentication and the authorization uh of the work that is happening. So, a lot of a lot of traditional fundamental ways of thinking about engineering discipline, but just really becoming now back into the forefront. Yeah, it's walk about me through a little bit what like what's the best case scenario here that you've seen? Like what company that has done this really well? Kind of what does that look like? Yeah, I I think that this is where we're going back, right? To the systems thinking. I think the companies that we find who aren't who are thinking about their overall strategy, right? And designing that. And so, there's a little bit of let's think this through before we, you know, go ahead and and jump and require top-down um you know, mandates. I think those those are not as successful and those those are finding as anti-patterns. So, it is more about what we're the ones that are successful are doing the due diligence. It's hard. It is hard work. But like, you know, to to to to provide literacy and enablement within your organization for the people and then to really leverage the ROI of your work. It's a lot of strategic thinking as well to think about where do we invest as much as, you know, how um and with what tools. Yeah. I feel to me like in this time of AI FOMO, strategic thinking isn't always at the forefront. Is that something that, you know, really needs to Would you say a lot of these companies need to need to slow down a little bit potentially and just, you know, think over what they're doing and whether maybe AI is even the right tool for them at at this time? Yeah, I mean, definitely. So, yes. And I think um there is a responsible AI perspective here as well as responsible leadership and technology aspect to it. But why why now more than ever is because you you can lend yourself to create generating AI slop or AI is going to produce a lot of you know, what you tell it to produce. And so, and also what you tell it not to produce. You know, right? There there is without the proper feedback loops in place. So, I think um that is why like once again, even even the engineering discipline we talked about, there's strategic disciplines as well. And, you know, bringing back um the the the disciplines that might be in place about what is your competitive advantage, right? I mean, you don't need to do uh what the Joneses are doing just cuz they're that's what they're doing, right? So, um how do you want to differentiate and where do you invest your money? Um so, yes. Definitely, that's also part of that and that requires getting the This is where the human judgment comes into play. Mhm. >> [clears throat] >> Uh for sure. For sure. Now, one we haven't talked about agents specifically yet, but we should talk about that a little bit cuz cuz that's, you know, basically becoming the default now for at least development teams. Um how's that changing how you're talking to your customers and the problems they're facing? And and you mean coding agents or you're talking about >> Yeah, coding agents. Sorry, coding agents. Yeah, specifically coding agents. Yeah. Yeah. So, I I guess um uh seeing that in in especially two different ways. Um but but but um one is the you can think about it from the topology shifts that we're seeing as well as architectural shifts. Um and uh um uh so, I'm part of this uh global team uh the uh that puts together the that works technology radar where we're uh you know, finding trends and we're looking learning from uh on the ground actually experiences from ThoughtWorkers from our ThoughtWorkers globally of 10,000 plus p- employees. Um and one of the things that we're finding is that we're using uh we're putting together um techniques to ensure that the code base is not uh is not shifting or drifting away from architectural principles. Right? So, so those types of things are very important when we're also thinking about coding agents. And so, I think there is a the latest technologies with um agent skills um as well as uh agent side and D was from before, right? So, a lot of these things that help provide the context to the agents. Um we're we're thinking about that as essentially like forward engineering, right? Like uh forward uh constraints for them in terms of the guiding principles. What are those architectural decision records that we would ask the uh agent itself to document so that the humans could review that and could reference that cuz at this point like humans reading all the code, that's going to become, you know, just yeah, un- unsurmountable. Um So, so having those ways to to be able to uh to you know, to validate that. Um and and on the other hand, the team topologies are also changing cuz now we have humans and uh machine agents working together and collaborating. And, you know, so there we're seeing these teams of coding agents where it might be a lot more strategic and intentional uh with deliber- deliberate design of, you know, who might be orchestrating very role specific. Okay, this machine agent for back end versus front end and whatnot. So, that we are finding and we put in our radar as a mechanism as a uh a technique that people can assess. But the thing that we put on the technique as people to maybe caution and watch out for and just don't jump in without some other additional um experimentation and testing are the coding agent swarms that are coming out where it's like hun- hundreds of agents being tasked to do the same thing potentially. And then you have to think about collaboration and, you know, um it they might run into conflicts that they have to resolve. And so, there, you know, right now we're like it's still it's still maturing. So, uh for organizations that are on the forefront of trying these things out, then okay, fine. Go ahead and continue to evolve and provide more practice best practices for their for the industry. But for others who have uh maybe more regulatory compliance um requirements and other things, there might be more cautionary steps towards it. So, overall, I'll just say that really this is still an evolving field and it's so it's dy- dy- dramatically changing. Um and one of the great things about the tech radar is that it just allows you to give gives you a glimpse of what are the things that we're all testing, experimenting versus what are the things that have evolved much further that you can, you know, uh choose to to adopt. Yeah. As you looked at that radar lately, anything else that stood out for you? Any surprises there? Um surprises there. Uh yeah, that's interesting. I guess the the So, this is this is my third time being part of this team um that put this together. One of the biggest things I'll say was uh we we used to call this we used to coin this term too complex to bit blip, which is still the case sometimes cuz our blips are very short and sweet and very quickly get a get a glimpse of um of a new technology or a new technique or or older techniques technologies that we want to reinforce. Um but this time, there were a lot of too young to blip. Okay. >> [laughter] >> Cuz it's moving so quick Everything is moving so quickly. >> quickly and we're and like a ThoughtWorker may have put something out there like cuz everyone's also experimenting on our organization and others. So, there they might put out a a blip that that we're like, oh wow, this might this is a interesting open source project that tackles this white space in the industry, but it just came out two weeks ago. >> [laughter] >> So, we're like, okay, well, we publish our radar maybe two months after we have this meeting. So, um would it be, you know, is it the right time to still anticipate the maturity of this or is it um you know, one of those things that comes out and then >> [laughter] >> may not last. So, I think there is there is definitely um yeah, something about that as well. Yeah, two weeks is a long hype cycle right now. A lot of these projects. Yes, yes. They'll probably get 50,000 stars in that time. Um one [snorts] thing I just wanted to go back to. You said, you know, I think it's another fundamental issue all dealing with is uh a lot of code is being generated now, a lot more than before with all of these tools. And one thing that keeps coming up in my discussions with people is how that creates new bottlenecks all across kind of the life cycle from code reviews on right really after that the code has been written. Like what are you seeing there? How are people dealing with that and what's your advice? Yes, yes. And we blipped this actually in the past radar and also we reinforced it here as well as um First of all, I think a perspective of cognitive load. Cognitive load for the human agents as well as for the machine agents. And and this is where once again, good architectural principles like um thinking about the boundaries of your code. Um and you know, what from before we've talked about modularity and, you know, from 1970s, uh you know, Parnas's paper, like once again, it becomes very evident here as well. So, why that that as a technique is important is that for the machine agents themselves, I think for their own uh you know, what we feed into their context uh windows, um but also for human uh cognitive load to be able to review and understand the code, right? A lot of that modularity um helps and and comes into bear. Um and I think uh uh once you have that, then you can start also thinking about the the harness that you are developing. And that harness including, you know, those guardrails or architectural guardrails, what are those feedback sensors, right? So, um in addition to, right, your the the the the feed forward of your context that you provide your agents with, the feedback with the sensors and the um you know, the tests and linters and a lot of those common practices that come in. >> Sure. Sure. Those are not going away. They're not going away and if anything, it's like, yes, how might we do better? Mhm. I I do want to share though that I think the other thing is to maybe rethink a little bit. This is maybe a little bit more advanced uh advanced in the industry right now, but rethink of how we even think about code itself. because yes, um the the quantity of code is is is is just going to, you know, dramatically increase with how quickly AI's able to produce it, and then humans become the bottleneck in that case. But I I think uh there's an opportunity to rethink where does code matter? And the volatility and durability of that code. Mhm. What I mean [clears throat] to say is, first of all, there's a lot of strategic thinking is important about what's actually valuable for the organization. Where do you want to build versus buy? How do you want to actually differentiate your organization, right? Should we even bother building it? So that is very important because it's so easy to build these days that it's Yeah. it's going to get even easier that cuz code is going to going to become a commodity to generate. Um that you don't necessarily need to spit it becomes there's going to be a lot of dark code, right? Like we have a lot of dark data. We've been collecting a lot of data through our big data initiatives, but now, you know, there's going to be a lot of dark code more more so than even before. Mhm. Um but uh not that there wasn't dark code code. >> [laughter] >> But the second thing is uh so viewing, you know, should this code even exist? I think the other thing is what is the volatility of the code? So right because [clears throat] it's so quickly now to also create POCs, how might you architecturally or from your system standpoint think about uh documenting that code as having a a um a a life cycle that would then eventually be you know, get deleted. Mhm. You know, so [clears throat] the the um retainment of that code um being very explicit about it. But secondly, um uh the um thinking about code that could be dynamically and more ephemerally generated. Yeah. Yeah. Right? Because like someone like a user wants to be able to access a particular uh interface, let's say, or an API. Well, if I don't have the agent skill for it or if I don't have that built already, and it's uh not a necessarily reusable, you know, um feature, then why not go ahead and just dynamically generate it for that particular single of it for purpose use, and then you're done. So So I think there's going to be that perspective of how we really think about code differently. Yeah. >> Yeah. Yeah, it's interesting. It came up, I think, first in discussions a year ago or so, and the models weren't quite there yet, I think, at the time. But people were thinking about it already, like maybe it's the spec is more important than the code cuz I'll just regenerate the code as I need it, and the model gets better, and then just have better code at the end, and just don't worry about all of the Yes. kind of But then we're going to still shift the needle into the spec. And the spec becomes a massive load, which we're also seeing as well. Um and for that, I think we do have a technique, by the way, that our ThoughtWorkers proposed um we were calling about progressive context disclosure. Okay. Um where, you know, progressively uh cuz remember like said, the machine's cognitive load is much as important as our human ones. Um and uh and with the progressive dis- context disclosure, you know, we are being explicit and intentional, I'd say, about um what matters for this particular request. Right. Right. Cuz that super spec is also going to become difficult, and then you want to think about that modularly, as well. >> [laughter] >> You got to get the spec a mono repo of specs. >> Yes, exactly. It'll keep us It'll keep us busy for a while. Um for those who want to learn more about the the radar, is there a place they can go to? Is that public or is that only for your clients? Or what does that look like? Oh, no, no, no, definitely it's it's public. I mean, our um our goal is to take a lot of the learnings that we're having and then the and then share it with the industry. Um so if you go to thoughtworks.com/radar, you'll be able to see our latest one. Uh we do have our uh the the the next version, the next edition is coming out next week, actually, April 15th. So volume 34 will be out. So um yeah, if you just wait a few days and and and and and see that one, then you'll get the latest and greatest. But yeah. Awesome. Perfect. Well, Naeemishah, it's been a pleasure. Yes, thank you so much, Fredrik. It was great to to be here. Same here. Same here. Thank you so much. Okay.

Original Description

In this episode of The New Stack Makers, Nimisha Asthagiri of Thoughtworks explores why many AI initiatives stall between proof of concept and production. A key issue is that organizations focus on speed—asking how to move faster—rather than rethinking what new capabilities AI actually enables. Successful companies take a systems-thinking approach, investing in organizational literacy and aligning teams around meaningful use cases instead of retrofitting AI into existing workflows. Asthagiri highlights that core engineering practices are returning o prominence. As AI-generated code increases, so does the risk of “cognitive debt,” where developers lose understanding of their own systems. To counter this, teams are reviving fundamentals like test-driven development, mutation testing, observability, and zero-trust security, especially as autonomous agents contribute to production code. She also introduces the concept of “dark code”—AI-generated code that may never be used—and argues for more intentional lifecycle management, including ephemeral code. Ultimately, the focus shifts from code itself to specifications, context management, and disciplined engineering practices. Here's the full article to go along with the video: https://thenewstack.io/thoughtworks-radar-agentic-ai/ Learn more from The New Stack around the latest about system-thinking approaches: System Two AI: The Dawn of Reasoning Agents in Business https://thenewstack.io/system-two-ai-the-dawn-of-reasoning-agents-in-business/ A practical systems engineering guide: Architecting AI-ready infrastructure for the agentic era https://thenewstack.io/ai-ready-infrastructure/ Join our community of newsletter subscribers to stay on top of the news and at the top of your game. https://thenewstack.io/newsletter
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

📰
Built a tool that datacenter cooling layouts optimiser
Optimize datacenter cooling layouts using AI and OpenFOAM, reducing power consumption and heat generation
Reddit r/artificial
📰
Beyond Logs: Building a Real-Time AI Observability Dashboard That Surfaces Database Rows, Not Just Latency Percentiles
Learn to build a real-time AI observability dashboard that provides detailed insights beyond latency percentiles, including database rows, to improve system performance and debugging
Dev.to · Robert Pelloni
📰
BizNode runs entirely on your machine — no cloud, no subscriptions, no monthly fees. Your AI business operator that works 24/7
Run an autonomous AI business operator on your local machine with BizNode, eliminating monthly fees and data leakage risks
Dev.to AI
📰
How to Deploy an AI Crypto Trading Bot on Your VPS
Deploy an AI crypto trading bot on your VPS with this step-by-step guide, covering prerequisites, API keys, ML training, and costs
Dev.to · Dinesh Wijethunga
Up next
OPUS 5 ! How to Collaborate in the Age of AI Agents: Vibe Coding with Buzz, Ray Fernando, and Block.
Tech Friend AJ
Watch →