End-to-End Interoperable Data Platform: How Bosch Leverages Databricks Supply Chain Consolidation

Databricks · Beginner ·🔄 Data Engineering ·1y ago

Key Takeaways

Bosch's journey in consolidating supply chain information using Databricks platform, integrating with dbt and Large Language Models (LLMs) for enhanced insights and operational excellence

Full Transcript

So, hi all. First of all, to get you a bit warmed up, it's been a long day. I guess it's like 400 p.m. 4:10. Maybe you're a bit tired. Maybe raise your right arm if you've been to the keynote earlier today. Yeah, you liked it. It's good. So you can leave the room now again. Uh because what we saw when we were earlier this morning in the keynote, we saw that like there were some kind of similarities to what we were supposed to show today. So let's see how it goes. So we're going to talk about like our end to end interoperable data platform. That title is definitely not from ChatVt. Um and we'll be showcasing a bit of like how we leverage data bricks in order to make data more robust in our end. Before we introduce ourselves, quick rational quick uh kickoff. So, one really cool invention of modern supply chain and trade are container ships. I guess I'm not I'm not a ship person myself, but I think container ships in general are cool. They bring you stuff. They can do that quite reliably, quite cheap, and they can go from basically any harbor in the world to any other harbor. So, they are quite nostalgic for global supply chain and trade. Yet another cool invention, canals. It's going to be a bit more interesting than that, but canals are cool because they connect two points. Usually, they are kind of a shortcut. So, they can bring down your roundtrip time for container buses and ships quite tremendously. combining them two like iconic duo. So let's imagine there's a canal in northern Africa. Let's say this canal passes through Egypt between the cine peninsula and the rest of Egypt. Maybe some of you remember maybe some of you can already tell where this goes. Um let's say there is a very large container vessel passing through one of 60 actually. There are usually roughly 60 ships passing each day. And let's imagine there is the ship. It's catching wind. It's slowly tilting eventually blocking the whole canal. It's quite of a screwed up situation. And if you're working logistics operation, that's really not fun. Like you're going to be on fire for a long time. Egyp Egyptian authorities were like looking heavily into resolving this because they would be getting paid based on the number of vessels that would pass through each day. So they want to resolve this as soon as possible. You have those contractual obligations. So they wanted to fix it as soon as possible. They brought up heavy machinery. They were like looking to dig this out. They were looking into like towing it away. It didn't work. Didn't work. So, took some days, one, two days, three days. Nothing worked. And this is when the carriers are getting nervous obviously because they would now need to decide whether they want to have their ships wait maybe for another day or maybe for another week or a month, you don't know, or if they want to have their ships going all the way around Africa, which will increase costs and increase the time for the trip. the colleagues working in logistics operation in our company, they were like having a hard time and um so they were thinking about okay what can we do about it? What can what can we do in order to make sure that our production lines are not coming to a hold because we lack supply. So the first thing that comes to mind is going all air freight. If you have the money, you can do that. But obviously capacities were quite tight. And not only the container prices on ships were going through the roof, but I think roughly 60% plus in very short amount of time, but also everyone who wanted to make sure their supply is fine, they were going to all in. So the point is this wouldn't work quite easily. They would figure this out. So, I'm not kidding if I tell you that in order to make sure we were still getting chips from China to Europe, people were actually considering, hey, could we maybe not send trucks from China to Europe, like shipping those? And I'm not kidding. This really happened. This were really like legit discussions. It would take like, I don't know, a few weeks or whatever, and maybe the the truck would break down. You would have to ship them every once in a while. And it's ridiculous. So luckily after six days um there was full moon and high tide in the sewish canal and the ship got unblocked and everyone was happy. It was really cool. What was not so cool is that during these six days the estimated amount of um goods that were delayed or the value of goods delayed was roughly $50 billion US and the am the number of ships that were uh stuck are roughly equivalent to the number of chairs we have in the room. So my colleague counted them in the very beginning. They're roughly 360. So roughly 360 ships were delayed. Now you may ask why am I telling you this? Uh answer is simple. Uh those basically any kind of event that you read about in the news that you hear in some podcasts or anything that you just get fed in will have an impact on supply chain. I guess most of you will know and uh maybe I'm not telling you anything new. And how do you mitigate these issues? How do you risk management those uh events? You do this through data, right? So I guess it's also obvious. So and while we have those many events going on, there are wars, there are like chip shortages which we also see on the right hand side, there are many things happening each and every day. You basically open your news app and there is some stuff going on quite for sure. All these events will have an impact and you will need to resolve them. And what we are making sure is we are making sure that at least the data is reliable that you use in order to make these decisions and risk mitigations. Hi, my name is Mark. I'm a project lead at Bosch responsible for the global supply analytics platform and I'm here with my dear colleague Satish today to explain a bit of how we make data more robust um for supply chain operations and use cases. Thanks Mark. So complex business needs demands datadriven approach. I think everybody would agree this. But what it really makes more sense uh to make this approach really work is to have a reliable platform. So let me ask you this. How many people think a reliable platform is definitely needed and you're really investing a lot of time on it. So show your hands probably can get some perfect. I really know that with whom I'm talking I mean this is perfect. This is precisely what we really wanted to showcase today. We just want to give you an output of what we really wanted to what we have achieved so far. And I'm Satish uh the development and technical lead for the global supply chain platform at Robert Porch. And I'm here to talk about the ecosystem what we have set up and how we are making the data reliable. Before that I hand over to Mark uh to take you on a journey on how we arrived at this ecosystem. Over to you Mark. Thanks. So maybe a quick thing, I hope you trust me. I hope you are fine with me asking you for a favor. Close your eyes. Close your eyes everyone. Let's do a small cognitive trick so that you remember the session afterwards as well. Let's imagine you're sitting in your kitchen and you're having like maybe on a coffee table. You have some cool hot drink like drink of choice, coffee, tea, whatever you want. And we'll come back to this later. you can open your eyes again. So to give you some key figures um of the complexity of supply chain uh that Bosch is dealing with um we have roughly 400,000 employees worldwide and almost 500 different subsidiaries and local entities and we're serving four different major business sectors. Um the one that is the highest in turnover by roughly 70% is mobility. Maybe not so many people know about it, but Bosch is the um major global um automotive supplier. We also serve industrial technology. Think uh yeah machinery um consumer goods, you know, I guess the best from home appliances and power tools and last but not least, energy and building technology. So there are lots of different goods, complex supply chain, lots of coverage. To make that happen, we have roughly 35,000 people working in supply chain. Uh we have a supply base of 35,000 suppliers. Um we own 240 production sites and operate roughly 800 warehouses. With all of that that's just some figures. Um with all of that there comes two main challenges. The first one is we have distributed functions. So functions are usually quite closely located to the actual value chain. So means in warehouse means in a production plant. And secondly, distributed data, which means that data is not just scaled all over the place. You have like those Oracle virtual machines or those databases in some some plant which you may not even know about. You have like different systems that come into play to make supply chain operations smooth. Like 10 years ago, there were some smart people that thought, okay, let's look into fixing that partially at least. So the idea was, let's move all our data to a so-called data lake. uh I'm not lying like 98% of data is tabula so we heavily relying on SAP um and the point here is that this would help a lot in solving those two challenges so data was colloccated you could easily query the data um you would have some kind of um yeah single point of of data or information just it turned out quite early on that um this wouldn't help to write like answer all the business questions Because data still is hard to understand. Just dumping raw data into a lake like data swamp is a term that is frequently used wouldn't help. You would need to figure out okay how does this data work? How to what's the relationships how to use it? And secondly um the workloads are still quite heterogeneous. So we have workloads that are uh quite operational where we basically compute stuff each and every hour. We have stuff that is just executed once a week once a week like uh computing the coverage for our supply chains and so on. So this wouldn't quite well work on an on-remise infrastructure. So we've had a lake and we started to build a house on top of it. So what did we build? Um a clear lake house based on data bricks platform. Uh we did the ETL fetch the data from the data lake. U move it into the uh data bricks. I think it's more or less a clear reference architecture we used. uh and it's could be more familiar to all the ones who are using data bricks as a medallion architecture and so on and so so forth but what makes it really interesting is uh we use dbt for the data transformation that makes it the cornerstone of the data transformation uh and what what exactly uh made us use uh dbt historically we were using dbt for the azure data warehouse one of the own azure data warehouse platform and when we thought about migrating the data bricks We found it really fascinating. We brought it with us and we started to use this. And what DBT actually does, it directly integrates with data bricks and data bricks. If you have a workflow or something like that, uh we can directly integrate uh DBT without having to create any kind of new ecosystem for it and it makes the migration also seamless. And um it it actually made our journey the data transformation platform agnostic. And this is uh more or less what uh what we achieved uh in in getting the getting the uh lakehouse in place. So modeling the data actually making it analytics ready helped a lot. Obviously this is like a very tedious task as you all know I guess. So semantics help a lot and also like um being able to scale quite easily uh quite well also helped in fixing our hogenous workloads or accommodating them. So Lakers was really cool. Always going good. Days past data processed like you know you check those pipelines you see everything is running smooth you're happy life's good and month go by. Perfect. So Mark mentioned that everything is going good steady and stable scalable. And what comes obviously is what how do we leverage the data and we start looking at okay LLM applications chat ports a lot of fancy stuff and that start that started our journey to start using lots and lots of value out of the data this obviously happens when you have 12 terabytes of data in your fingertips and that's precisely what we do but Mark what do you think what did the management really think about it and how did they read it they loved it right I mean you know if you can sell geni AI aentic AI whatever kind of buzzword you throw and like they read it in some papers and they they think that's like the coolest thing ever. So, they were really on fire. They were bought and they wanted this to happen yesterday, you know. I I guess I'm not telling you anything either. Yeah. So, we've had an incident. Um I hope I'm not the only one now saying that this has happened to us. I guess maybe you have seen similar things happening. So, one sunny day we were in the office. Everything was good. Like at one point we were just getting spammed on teams, getting bug requests, getting emails, getting calls. People were like crazy because they were like the planning guys. They were complaining that the downstream jobs wouldn't work. Data was outdated. They they need to fall back to some kind of manual procedure. Um as said, we are also supporting operational activities. Instead of using some fancy dashboard during the activities with a nice overview, they had to crunch SAP transactions again. And this is not fun at all. Right? So, you know, you wouldn't be getting an overview of all your parts. You would need to go into your transactions, crunch each part number and and you know, get some kind of overview of what's your stock, what's your demand, etc. Not fun. People were really upset and they also let us know about this. So, this was not a happy day for us. We looked into it. Unexpected schema change from production in our data m uh easy thing. Column got dropped. You may say, "Oh, why did you guys not see this?" Yeah, we have code reviews. Obviously, we do, but code reviews are done by people, right? So, you do some mistakes, you miss something, maybe it's been a long day, you're context switching. So, code reviews are cool in general, but any every code review you don't need to do, right? That's a good code review, I would say. Hope you agree partially at least. So, schema stability was not enforced. So, um we would need to see in the code review if a column got dropped or renamed or whatever. So it went unnoticed a set and the people had to fall back to manual operations and as said this was uh like a pivotal point for us right so this this gave us a clear clear indication something is really not the way we had to think about so we had to step back and start thinking about what do we do uh now it's like shift left and start re reflecting on what we have done. Um it's really good, really interesting, really fancy to have a lot of applications, but the most important thing is data. Um your insights is as good as how reliable your data is. It it comes from how fresh is your data and how correct the data and so on and so forth starting from the source to the consumption. So as you can see we can reach the wisdom but when the data itself uh is really not taken care of um that that then comes to a big question mark. Um this is what made us think about uh use of model contracts. So we I think a lot of people would agree uh they do a lot of testing and and verify the information and so on and so forth but I don't think a lot of people enforce it and because the cost of enforcing the data quality and the model contracts is pretty high and the efforts needed is pretty pretty complicated but if we don't do this we really cannot reach the point and that's when we thought we definitely need a model contract uh to achieve this and and to go on from there we need to uh look into it. Before I go to the next point, how many of you use DBT for data transformation in this room? Few few of them for the folks uh who who are not aware of DBT. DBT is is it lets you write SQL Python based data transformation models. It also enforces data validation. uh it's I I would call it a data transformation as a code and it enforce a data validation on that and then we can also uh make sure that the automatic lineages are generated out of it. So uh like I mentioned in the previous uh one of the previous slides, DBT directly integrates with data bricks and if you really wanted to get uh started we can start up you know data transformation with DBT on a small post or something like that and when we really know we wanted to use it on on data bricks we can just directly port it. Of course there is a small level of migration which needs to happen but most of the heavy lifting uh can be done done without uh major issue. As you can see on the left hand side uh the DBT uh data transformation model looks like this. It's a simple SQL transformation and on the right hand side you can see the model contracts where you can see the the code clearly says it's enforcing the data type verifications referential integrity checks and so on and so forth and that's the part we were missing uh because we did do a data data test and all the stuff but if you don't enforce it and if you just move the model to the last stage anything can happen. So we wanted to enforce it and this precisely the missing bit which you wanted to uh build on and um what it really means to uh build is not that easy because you're talking about an environment with 12 terabytes of data more than,300 data models and the developers need to adopt they need to annotate their their their work and they're really as you as you all know the developers every developers have their own way of writing. Standardization is a challenge and we need to have a common standard and common design patterns and so on and so forth. constraints verification needs to happen and everything needs to happen on a manually it needs to be man manually done which needs who wa satish question like this is not planned we we don't have it in time timeline right so there is no capacity there is no time and I can promise you the stakeholders won't be really amused if we now come and say okay we need to make your data more robust more stable so let's spend another three months not delivering anything that is on the road map not delivering anything that we committed on and that's that's precisely what I heard That's precisely what I heard. I mean, do we really want to invest on it and rather let's wait for something to break and let's fix it if the cost is low and we didn't do the manual stuff. Uh rather we started to brainstorm on several levels of uh ways of handling it LLM based modeling or some kind of anomaly detection in the flow and and so on and so forth. Um I mentioned uh when when you spoke about uh the amount of data 12 terabytes and the the platform is stable and a lot of people think about using LLM and agenti to get fancy stuff chatbots and so on and so forth but what we really failed to think about is can't we use the same technology same geni stuff to validate the data to generate these model contracts because modeling as a code can can simply be enable able during this and that's precisely the approach which we took and for the ones who are interested in these areas this is something which I'm going to take it from from from now on um so what is it all about I spoke about the overall platform you can see um the data gets transformed in different stages I really don't want to revisit it but the part which you can see exactly in the middle is a model contract which needs every time the model is generated it goes to the next stage it needs to be filtered out of the model contracts otherwise we'll not be able to generate it and that's something which we which we are achieved and any any raw data going into the loss layer be it medallion architecture or anything needs to get the validation done from the model contracts and how did we actually achieve it um we we thought about it a lot because 1,300 plus models it's still growing and the complexity of the data lots and lots of objects around com lineages, hierarchical models, everything is there. What do we need to do? which means um we need to we had to think about a clear strategy for it and that's when we thought about mosaici framework and what the the there are a lot of things which we could have used but we started with the vector store um and um how many of you use mosaici framework just just so I know okay a couple of them basically the mosaic kaii framework actually offers you uh the vector search index directly sitting on the on the delta tables which means you really don't have to think about bringing in a specialized vector store for this purpose rather we can actually leverage the existing mosaic framework to build your vector store I've been hearing a lot about since since couple of weeks that the mosaici framework itself or the vector store itself is going to get a lot faster and they have a lot of enhanced features but it's something which I've not uh tried it out but it's quite promising uh what we use for the ones who really have not been using vector store. Vector store is basically the data warehouse of your geni applications and we used the vector store uh it generally uses uh the so-called hnssw uh pretty pretty complicated word hierarchical navigable small world algorithm to get the nearest neighbor because we had to get the similarity searches and we needed a vector store for this and um if you have to get the cosine similarities we need to normalize the data and that's precisely What we did we did not just feed all the models directly but we normalized the data we fed it into this vector store got the similarity searches done and started processing the model contracts out of it and this itself this particular framework itself was built on a Python streamllet and directly integrated with the data bricks apps so that's precisely what my my colleague was talking about some of the things which was spoken about today in the keynote session was the data bricks apps And uh this is something which we have been using and it's started to give me a feel that we are just seeing what we have been doing so far. Um before going is um Mark has a has a quick view on it. Sebastian this slide is for you. So did I get it right? Um we basically take our project code repository and all the context we have there all the comments everything. We take the data bricks metadata like convoluted lineages, schemas, everything we feed into the LLM and out we get a DVT model contract YAML file. Is that it? Absolutely. That's precisely the approach what we wanted to use and um just to take you uh through the journey of it. You can see a sevenstep approach but I'll just like split it into uh a two very broad classification. What is the offline mode? Uh where we normalize the data like I mentioned we don't want to feed the data as it is uh we need we need a proper chunking before we embed the data and that's the offline mode where we uh prepare the data in the offline. So what what we need we need the models we need the all the the best practices guides and we need the complete uh metadata information to build the model contracts. Of course in the generation phase we had to give the proper semantics to it and that's something which automatically happens and in the second phase uh where we talk about uh the the retrieval uh information uh we used a more of a rack approach um once the offline information were completely uh prepared using the rack retrieval uh the every column which is built for this model goes through a process of uh completely navigates through uh the similarity searches on the vector store and uh it for example if you have a table with 10 columns or something like that um when we decide if this particular column needs to be very validated for not or not uh there are several factors which needs to think about I mean how is it used across all the models is it really worth having the notal verification done and what is the impact business impact for it if we don't do this and this is something which we have completely vectorized and this vectorization process automatically finds the similarities gets to uh the core where it generates uh generates the model. So uh at the developer every developer he creates the data transformation model and every time uh he generates and he directly chooses the model and everything happens under the hood creates the model contract for him and this is the approach uh we used and uh we started to use this as a AI agent. Yeah. So basically this this this doesn't have to happen manually. This this is something which I was talking about every time we invest building it for tons of models it's going to take a lot of time. So an agent would automatically do it. No developer has to do create it and it happens in the under the hood. So um this is uh more or less uh how how the entire uh things works. We really wanted to show some kind of live demo but unfortunately I mean we are kind of worried about technical glitches. So we bought some screenshots which we can probably quickly share with you guys. Um so like I said this is streaml application which we built. Uh and this streamllet application um in in in in the background directly connects with the vector store. It does uh all the processing and also in the retrieval phase connects to one of the LLMs. And I think what I also would like to mention is for embedding itself there are several uh LLM models which which is in the mosaic you can make it serve and you can start using it and uh when we look at the application uh this is this is how it more or less looks. Um on the on the left hand side you can see uh the the DBD models itself. Um there are two two parts to it. You can see one is the SQL part and the one is just clearly the name of the columns uh and the description nothing more. It doesn't tell if these columns needs to be verified or validated for any kind of not null constraint referential con constraints. If if yes, how does it happen and so on so forth. It doesn't do anything like that. This is what happens in the background. So when we actually generate it, this is what you see. It automatically fetches which particular column needs to be validated for what. For example, if it's a primary key, we need to validate for uh some kind of uniqueness. Uh referential integrity has to automatically look for uh where exactly the child tables or the child and child and so on so forth. Uh just just all of you know in data bricks data tables uh the constraints are not enforced. So basically we we really cannot uh see that enforcing and the the only way where we can actually validate is using such tests and um enforcing is uh is one part another one is testing is another part. So if you really want enforce we need this bundle contracts and this automatically generated for us and and streaml application was working on completely and uh just just that uh you guys know uh the the uh embedding itself looks like this. Um the the great part about mosaici framework is uh it's not that we just have to use only the embeddings which means only the vectors. You can also have keywords in combination with vector. So this makes it even more sophisticated to actually build an uh rag application which combines both the context. Uh I guess they're they're making it even more better uh in the next days. But uh to to the extent of what I've used this combination really empowers that the the entire journey of it and we can uh get the the entire rag application uh really really running faster and uh the the traversing of uh the similarity searches happens uh pretty pretty rapid and uh as you can see it's pretty straightforward to to create a vector vector search index. Uh all that you need to do is use the UI go create a vector search index and get this up and running. You can either get it uh running in a trigger mode or continuous. uh for our purpose trigger was more than sufficient but there are some use cases where we wanted to have some real time um searches vector searches uh that that could be helpful just that it's pretty compute intensive uh to use continuous mode and that's something which probably comes out of my experience uh and this needs to be sized uh very well this happens in a serverless mode so basically we need to have a proper observability to see what cost otherwise you know there huge amount of data uh we probably would be running out of uh running out of billing so but Satish would that not mean like we have additional efforts like maintaining the app hosting it operations we need to have different people building the apps absolutely as you can if I go back um and it's still local host is it yeah it's local host yeah so it's a streamlit application and which means uh you need an app developer And normally if you if you're having a data platform it's intensively data engineers on board or data scientists on board we really need to need to needed to build something uh an additional team to take care of it which is an additional effort and all those things and that is precisely what what Mark is talking about is an additional investment yeah and that's when we started to look at uh the data bricks apps so this something which was also shown in the keynote today. Uh probably a lot of people have seen what we did is the streamlit application what was built completely directly deployed on the database apps. So um the code is completely checked into a repository move the code and started to deploy uh these streamllet applications uh which was working. So I mean you can also use a flask or any any other Python application. It works without any issues and it started to uh build we can deploy it or can schedule the deployment and so on and so forth. And what happened now we can see in the address bar it's not local anymore it directly creates the URL for you with uh with the proper workspace address or something like that HTTPS and so on so forth. So we really didn't have to worry about it and that's that made our life a lot seamless I would say. uh we didn't have to worry about hosting the application uh creating any kind of uh any kind of uh new team to manage the application and that's that's the point what what brought us and Mark what was what was your thought about it when we implemented it? Yeah. I mean like uh super cool and in the end it was blazing fast, right? I mean just a few few days it was if I'm not mistaken. Absolutely. Yeah. So it's like really taken away the operations responsibility in a sense. So the think about any kind of data retrieval as well, right? So often times you have to talk to people that are like developing applications. At least in our company that's away and like this can really help you like being more independent and actually fixing those potential gaps that you see in your data asset field and not having to think about hosting it, maintaining it, all these operations related stuff as we mentioned before. It's just taken away from you and that's like really really cool. And the time to market and time to release it's a lot shorter than before. And what did you think about going going ahead? I mean when we can implement a model contract with an agentic AI approach we could as well do data testing also as an agent and think about creating an agent which is already in process where we thinking about uh of course using langraph using a graph rack or something like that to uh extract the information more clearly more precisely uh also help if there is a violation if there is any kind of uh if there is any validation issues uh to support the developers do their validation much easier and so on and so forth and that's that's where we wanted to go ahead and before that I we would really like your thoughts and inputs and how you wanted to please please also uh feel free to post your comments in in one of the one of the mode and we would really like to know if that's something which you have been doing yeah thanks that being said thanks Thanks for hanging with us. I know it's tight, but there are three main takeaways for you today. First one is supply chain is a complex field. There's lots of irritations. You need to have your data right in order to get them resolved. Common sense. Secondly, data bricks is like a great asset for us in order to make sure that we get our data right and that we can have scale workloads that we can like do lots of different things that we couldn't do before. Maybe you also say, "Yeah, okay. I know I'm at the data and I summit. I know this already. But last but not least, maybe you remember when I was asking you to close your eyes and you thought maybe that's what's wrong with with this guy. He likes chips and he's asking me to think about coffee. Close your eyes again. Do me that favor. Now again, think about the coffee. Maybe when you leave the room may or maybe in a few days or maybe if you leave the room, you will have forgotten about the talk. But I want you to remember one thing which is if in a few two three weeks from now you're sitting at the coffee table, you have your favorite coffee and you scroll through your news app and you read something about newly imposed tariffs or any other kind of thing that is happening. Think about this talk and think about that you should make sure and you should leverage DBT and if not DBT at least mosa to easily launch agents and apps, deploy them and make your life easier. not just doing those super fancy use cases but maybe also thinking about how you can leverage them in order to increase your data reliability and robustness. Thank you. This is the end of the call or of the talk. Um would be really cool if you could complete the survey. We are always learning from feedback. And that being said, if you have any questions, uh feel free to step up to the microphones or catch up with us if you're more comfortable um or if you have some deeper discussions. We are really eager to hear your opinions, your use cases and would be happy to exchange. Thank you. Thank you.

Original Description

This session will showcase Bosch’s journey in consolidating supply chain information using the Databricks platform. It will dive into how Databricks not only acts as the central data lakehouse but also integrates seamlessly with transformative components such as dbt and Large Language Models (LLMs). The talk will highlight best practices, architectural considerations, and the value of an interoperable platform in driving actionable insights and operational excellence across complex supply chain processes. Key Topics and Sections Introduction & Business Context Brief Overview of Bosch’s Supply Chain Challenges and the Need for a Consolidated Data Platform. Strategic Importance of Data-Driven Decision-Making in a Global Supply Chain Environment. Databricks as the Core Data Platform Integrating dbt for Transformation Leveraging LLM Models for Enhanced Insights Talk By: Marc-Alexander Frey, Project Lead Logistics Innovations, Robert Bosch GmbH ; Satish Karunakaran, Development Lead , Robert Bosch GmbH Here's more to explore: Databricks named a leader in the 2024 Gartner® Magic Quadrant™for Cloud DBMS: https://www.databricks.com/resources/analyst-paper/databricks-named-leader-by-gartner An open, unified approach to your data, BI and AI workloads: https://www.databricks.com/product/databricks-sql See all the product announcements from Data + AI Summit: https://www.databricks.com/events/dataaisummit-2025-announcements Connect with us: Website: https://databricks.com Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc Facebook: https://www.facebook.com/databricksinc
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Databricks · Databricks · 0 of 60

← Previous Next →
1 Building AI Agent Systems with Databricks
Building AI Agent Systems with Databricks
Databricks
2 Databricks Workflows
Databricks Workflows
Databricks
3 Automate Unity Catalog Upgrade with UCX Part 1: Overview
Automate Unity Catalog Upgrade with UCX Part 1: Overview
Databricks
4 Automate Unity Catalog Upgrade with UCX Part 2: Installation
Automate Unity Catalog Upgrade with UCX Part 2: Installation
Databricks
5 Automate Unity Catalog Upgrade with UCX Part 3 - Assessment
Automate Unity Catalog Upgrade with UCX Part 3 - Assessment
Databricks
6 Automate Unity Catalog Upgrade with UCX  Part 4 - Group Migration
Automate Unity Catalog Upgrade with UCX Part 4 - Group Migration
Databricks
7 Table Migration and Catalog Design with UCX | Part 5
Table Migration and Catalog Design with UCX | Part 5
Databricks
8 Setting Up Azure Access for UCX Table Migration | Part 6
Setting Up Azure Access for UCX Table Migration | Part 6
Databricks
9 UCX Table Migration: Creating Catalogs and Schemas | Part 7
UCX Table Migration: Creating Catalogs and Schemas | Part 7
Databricks
10 Automate Unity Catalog Upgrade with UCX  Part 8: Code Migration
Automate Unity Catalog Upgrade with UCX Part 8: Code Migration
Databricks
11 Streaming to Kafka Just Got Easier with DLT Pipelines
Streaming to Kafka Just Got Easier with DLT Pipelines
Databricks
12 Data Engineering From Data to Dashboards with DABs: Crunching the Cookies Dataset
Data Engineering From Data to Dashboards with DABs: Crunching the Cookies Dataset
Databricks
13 Epsilon helps businesses connect with their consumers using Databricks Data Intelligence Platform
Epsilon helps businesses connect with their consumers using Databricks Data Intelligence Platform
Databricks
14 Unilever transforms operations with GenAI using the Databricks Data Intelligence Platform
Unilever transforms operations with GenAI using the Databricks Data Intelligence Platform
Databricks
15 ActionIQ enables businesses to unlock customer data with the Databricks Data Intelligence Platform
ActionIQ enables businesses to unlock customer data with the Databricks Data Intelligence Platform
Databricks
16 Mixed Attention & LLM Context | Data Brew | Episode 35
Mixed Attention & LLM Context | Data Brew | Episode 35
Databricks
17 Inside Databricks SQL: Engineering innovation with Hans
Inside Databricks SQL: Engineering innovation with Hans
Databricks
18 Inside Databricks: Engineering innovation with Michael Armbrust
Inside Databricks: Engineering innovation with Michael Armbrust
Databricks
19 The Money Team at Databricks: driving revenue and customer growth
The Money Team at Databricks: driving revenue and customer growth
Databricks
20 Unity Catalog unveiled: engineering data governance at scale
Unity Catalog unveiled: engineering data governance at scale
Databricks
21 Create a view in Databricks and share it with Power BI using Delta Sharing
Create a view in Databricks and share it with Power BI using Delta Sharing
Databricks
22 NDUS leverages Databricks Data Intelligence Platform to revolutionize higher education management
NDUS leverages Databricks Data Intelligence Platform to revolutionize higher education management
Databricks
23 Démo Databricks de AI/BI
Démo Databricks de AI/BI
Databricks
24 EMEA Data + AI World Tour 2024
EMEA Data + AI World Tour 2024
Databricks
25 GenAI: The Shift to Data Intelligence - Customer Panel on Industry Use Cases
GenAI: The Shift to Data Intelligence - Customer Panel on Industry Use Cases
Databricks
26 GenAI: The Shift to Data Intelligence - Ft. Ash Jhaveri, VP of Reality Labs Partnerships at Meta
GenAI: The Shift to Data Intelligence - Ft. Ash Jhaveri, VP of Reality Labs Partnerships at Meta
Databricks
27 Virtue Foundation leverages the Databricks Data Intelligence Platform to advance global health
Virtue Foundation leverages the Databricks Data Intelligence Platform to advance global health
Databricks
28 Announcing Synthetic Data Generation in Mosaic AI Agent Evaluation
Announcing Synthetic Data Generation in Mosaic AI Agent Evaluation
Databricks
29 AI/BI Dashboards Embedding - A tutorial
AI/BI Dashboards Embedding - A tutorial
Databricks
30 Bayer transforms global data management with the Databricks Data Intelligence Platform
Bayer transforms global data management with the Databricks Data Intelligence Platform
Databricks
31 Databricks at AWS re:Invent 2024
Databricks at AWS re:Invent 2024
Databricks
32 Hive Metastore and AWS Glue Federation in Unity Catalog
Hive Metastore and AWS Glue Federation in Unity Catalog
Databricks
33 Data + AI World Tour Paris 2024
Data + AI World Tour Paris 2024
Databricks
34 Retail reimagined: Currys data-first strategy to driving growth and improving operations
Retail reimagined: Currys data-first strategy to driving growth and improving operations
Databricks
35 Mixture of Memory Experts (MoME) | Data Brew | Episode 36
Mixture of Memory Experts (MoME) | Data Brew | Episode 36
Databricks
36 Verana Health Data Curation and Innovation with Databricks and AWS
Verana Health Data Curation and Innovation with Databricks and AWS
Databricks
37 Securing SaaS Applications: Obsidian Security on Their Journey with Databricks and AWS
Securing SaaS Applications: Obsidian Security on Their Journey with Databricks and AWS
Databricks
38 Twilio Eng VP on Data Intelligence & AI at AWS re:Invent 2024
Twilio Eng VP on Data Intelligence & AI at AWS re:Invent 2024
Databricks
39 Chegg Eng SVP on Data-Driven Approach to Student Success with Databricks and AWS
Chegg Eng SVP on Data-Driven Approach to Student Success with Databricks and AWS
Databricks
40 Ibotta Personalized Rewards Innovation with Databricks and AWS
Ibotta Personalized Rewards Innovation with Databricks and AWS
Databricks
41 Simplify AI governance with #databricks AI Gateway
Simplify AI governance with #databricks AI Gateway
Databricks
42 Databricks SQL and Power BI Integration
Databricks SQL and Power BI Integration
Databricks
43 Databricks Serverless SQL Warehouses
Databricks Serverless SQL Warehouses
Databricks
44 7 West powers audience growth with the Databricks Data Intelligence Platform
7 West powers audience growth with the Databricks Data Intelligence Platform
Databricks
45 Secret to Production AI: Tools & Infrastructure | Data Brew | Episode 37
Secret to Production AI: Tools & Infrastructure | Data Brew | Episode 37
Databricks
46 Skyflow CEO on Data Privacy with Databricks at AWS re:Invent
Skyflow CEO on Data Privacy with Databricks at AWS re:Invent
Databricks
47 Databricks Clean Rooms Product Demo
Databricks Clean Rooms Product Demo
Databricks
48 Dun & Bradstreet Enrichment & Monitoring, powered by Delta Sharing & Databricks Marketplace
Dun & Bradstreet Enrichment & Monitoring, powered by Delta Sharing & Databricks Marketplace
Databricks
49 Unpacking Libraries in Databricks
Unpacking Libraries in Databricks
Databricks
50 Providence uses an AI agent system from Databricks to help doctors improve their communication
Providence uses an AI agent system from Databricks to help doctors improve their communication
Databricks
51 How State Street Uses AI to Transform Millions of Trades Daily
How State Street Uses AI to Transform Millions of Trades Daily
Databricks
52 Vevo Therapeutics CEO on Curing Disease with Data at AWS re:Invent
Vevo Therapeutics CEO on Curing Disease with Data at AWS re:Invent
Databricks
53 Over Architected with Nick & Holly: Databricks updates for Feb 2025
Over Architected with Nick & Holly: Databricks updates for Feb 2025
Databricks
54 The Power of Synthetic Data | Data Brew | Episode 38
The Power of Synthetic Data | Data Brew | Episode 38
Databricks
55 Use Databricks Lakehouse Federation to break down data silos
Use Databricks Lakehouse Federation to break down data silos
Databricks
56 AI's rugby score: National Rugby League rallies fans with analytics and unified data
AI's rugby score: National Rugby League rallies fans with analytics and unified data
Databricks
57 Open Variant Data Type in Delta Lake and Apache Spark
Open Variant Data Type in Delta Lake and Apache Spark
Databricks
58 How would you sort Ætheldred in the alphabet using Databricks?
How would you sort Ætheldred in the alphabet using Databricks?
Databricks
59 A guide on how to operationalize the Databricks AI Security Framework (DASF)
A guide on how to operationalize the Databricks AI Security Framework (DASF)
Databricks
60 Future-Proof Your Asset Performance Management with Generative AI - Field Assistant Live Demo
Future-Proof Your Asset Performance Management with Generative AI - Field Assistant Live Demo
Databricks

This video showcases Bosch's journey in consolidating supply chain information using Databricks platform, highlighting best practices and architectural considerations for an interoperable platform. Viewers will learn how to integrate dbt and LLMs for enhanced insights and operational excellence. The session emphasizes the strategic importance of data-driven decision-making in global supply chains.

Key Takeaways
  1. Identify supply chain challenges and the need for a consolidated data platform
  2. Design an interoperable data platform using Databricks
  3. Integrate dbt for data transformation and LLMs for enhanced insights
  4. Implement machine learning models for operational excellence
  5. Analyze data for actionable insights and operational excellence
💡 An interoperable data platform is crucial for driving actionable insights and operational excellence across complex supply chain processes

Related Reads

📰
I Built My Second ETL Pipeline. This Time, I Started Thinking Like a Data Engineer
Learn how to build a production-ready ETL pipeline with Python, Docker, PostgreSQL, and Kestra by thinking like a data engineer
Towards Data Science
📰
JuiceFS Sync for PB-Scale Data Transfers: Resumable Sync, Encryption, and Bandwidth Control
Learn how to efficiently transfer large volumes of data using JuiceFS Sync, which offers resumable sync, encryption, and bandwidth control, ideal for PB-scale data transfers.
Dev.to AI
📰
How Airflow is using AI to make data engineering more resilient, not more complex
Airflow uses AI to make data engineering more resilient by detecting data drift, resuming failed pipelines, and fixing issues automatically, reducing complexity and improving reliability.
Medium · AI
📰
What Can We Do When Memory Becomes the New Bottleneck in Data Engineering?
Learn how to overcome memory bottlenecks in data engineering using Pandas chunking, Dask, and Polars, and why it matters for processing large datasets
Towards Data Science
Up next
A Moment Frozen in Time | Arnav Iyengar | TEDxJenks Youth
TEDx Talks
Watch →