How to Migrate from Teradata to Databricks SQL

Databricks · Beginner ·🔄 Data Engineering ·1y ago

Key Takeaways

Migrate from Teradata to Databricks SQL using Lakebridge, Unity Catalog, and Databricks SQL for a cloud-born platform with faster innovation and feature releases. The migration involves breaking down the environment into objects, defining lineage, and processes for conversion, and using tools like Lake Bridge Analyzer and Legbridge converter for automation.

Full Transcript

Hi everyone, thank you for being here. Uh this session is uh how to do successful migrations off of terod data. And uh before we get started, I want to see a show of hands. How many of you guys are still using terod data and why? I'm just kidding. Yeah, that's uh it was a great technology. I was I'm a longtime employee of Terod Data previous same with my partner here but we'll get into it. Uh a little bit of uh statements you probably have seen uh some of the uh learning if you wanted to probably have seen it in other sessions as well to uh kind of get you familiarized with data bricks and uh everything else. Uh I have been uh around the data warehousing for over 25 years. Uh been a terod data employee uh prior to data bricks uh for seven years and a mostly a data warehouse consultant prior to that. Uh and uh been around the data warehouse world with different technologies and different industries. uh but definitely extremely uh happy to be part of data bricks and I believe this is the revolution and this is what we're going to be seeing this platform for a long time with the levels of innovation and I want to my partner to introduce himself yeah so I'm Fabian I've been datab bricks for a bit above one year before that I've I've been at Vertica as a CSM for two years and before that I was at Terodata for five years as a professional services consultant. I did most of my consultancy into utilities and yeah before that I have a long career I won't I won't make you sleep with it. So on the agenda side uh we want to really uh strategize and plan the migration choosing the right approach right architecture and the execution plan is very important as part of the process and there are ways different ways which lake bridge is one of them that's going to get introduced in tomorrow's keynote is to create a good solid inventory of all the objects that you want to move you know whether it's your views SQL store procedures and The migration methodology is what we're going to get into uh in a little bit later and having a good data validation plan as part of very important as part of the process. The challenges that you probably have, you know, encountered that keeps you still within terod data boundaries is how hard it is to migrate off of terod data because of all the objects that are living there, especially store procedures and the data type conversions whether you're using time series whether you're using temporal and the code conversion overall is going to be a very big challenge which at the end of this presentation I'd like to give a demo of how we think using LLMs could really help out the process. We do have a product uh that is being introduced as it's called Lakebridge that does a code conversion but uh we also want to showcase how you could use the data bricks platform to do the migration successfully and then the benefits it should be obvious by now is that definitely you know we're a platform born in the cloud we're not just a VM from an on-remise system that's been you know running in the cloud environment and definitely taking advantage of all the cloudy features that you know uh we have along with soup to nuts basically all everything that you need for your data and AI and analytics needs and then of course you know I've been blown away I was part of terod data for many years every feature that was about to come out has so much problems and it took so long and it was mostly paperware before it actually made it out Every week within data bricks I'm seeing a new feature and it's uh it's a testament to the way how fast to innovation the path is and uh how actually product is connected to the customers when it comes down to it. So as far as strategizing uh the plan and uh you know we all know it's it's hard you know having legacy data warehouses on premise systems and uh examples of what you see here is that uh you know everything is built around the black box because it makes sense to have a proprietary storage. it makes sense for them to deliver good performance uh not in an open-source fashion and uh and that's that was one of you know the major success at the time that on premise systems and data warehouses and traditional data warehouses were in place they delivered the really really good performance out of what they were uh and definitely we want to get into the benefits of moving out of uh uh terod data into data bricks Absolutely. So the first one is the TCO reduction. Don't let get lured into one into this one. This is something we've seen from tera customers that have moved to datab bricks. They told us okay oh you are cheaper. But this is not the value of the migration. You don't migrate just to reduce your TCO. If you do that it will be a waste of the datab bricks platform because we can do much more. I think you've you've all seen the keynote this morning. Everything we are building into data bricks to do everything that data data connection machine learning geni uh apps uh leg base it's it's just stacking up stacking up stacking up. So if you say okay I want to move out from Terza to Darrix just to to have a TC reduction because the comput is cheaper honestly you're missing the point but it happens to be there so it's good thing but it's not the main driver the enhanced data science analytics the faster time to insight because now you can just chat with your data it's it's amazing the scalability and performance terata has a great uh warehouse house resources manager TSM it's great but in the cloud we have elasticity we can spoon compute on the need it's just much more efficient because you can have the best workload system you want when you are 100% CPU it will uh it will lag everywhere and you will have you will get into prediction issues here we just spend on your compute because the the architecture the leouse with the separation of the storage of the compute allows us to do and the the good governance with Unity catalog you know it's uh we are just adding new features and everything we're adding it go through unity cataloges so you have a new features leg base okay go run the permissions through unity catalog you want to do an app permissions through catalog you want to do data sharing you want to do model machine learning and and and this is actually very huge and Yeah, the left part the the one with the with the black box. So this this is terata again great solution. I love my time terata. I'm sure Maran did as well. But it's only tackle a part of the problem. And with the data expanding days and after days we need more solution. We need more products. We need more tools and we need to get that federated with a proper governance. And I won't go through all the all of this because you you know it you've seen it in the in the keynotes or in the slides. But just to I want you to make I I want to make sure you understand moving of terata to datab bricks is not just moving the data warehouse. You are embracing the future. You are embracing all new features all the connections everything we announced this morning. Some of them I was surprised as well. So it's amazing. It's a it's really a new a new platform. So basically the migration plan we do see it as three different patterns of migration. Whether you would want to do a lift and shift where you're kind of want to derisk the outcome by moving everything you have into a data bricks environment. That's a possibility. uh definitely you're just transferring your technical debt from one platform to another not the best approach but sometimes because of the budgetary reasons and the time to value you would be forced to do it this way whereas your ETL logic remains the same your architecture remains the same your data model remains the same and you're just moving from one platform to another definitely a possibility we've done it in the past with different customers that choose to take this path and kind of modernize uh you know one business process at a time in a later kind of a phase uh when they would get to it. But typically once you migrate it be really hard to go back and reassess and re rearchitect. Modernizing definitely is what we recommend. uh you know now that you have this platform available to you with all of its capabilities why not use it for its features really rich features that it has and I'll get into some of that on the demo as well and this is where you would update your data model you would what we call a medallion architecture similar to somewhat what terod data offers from a staging tier from a core tier from a semantic site uh but changing rearchitecting the data model where it makes sense for data bricks environment. And this is where you really want to take advantage of the ETL features that we have as opposed to probably using another tool or system where you could but we do have all of those orchestration and capabilities in house. And a hybrid migration is where you lift and shift your current environment as new use cases come to to life. you would want to modernize those use cases based on the best practices that we have. So the main thing that we see as uh the first step is having you know identifying your business use cases probably within your terod data environment you do have a customer success manager assigned. I used to be one of them at Teroda at some point before becoming a solutions architect. uh is that uh there is a consumption analytics that that typically is run for you to define how you consume your resources, business processes, business departments using terod data and which parts of terod data that it's it'd be really good to use that to define your business use cases and pertinent uh tables, views, objects and everything to to kind of set up this uh you know inventory of or plan of attack for for your execution plan. And then of course the ETL logic whether it's a you're using a partner to run ETL or you're using store procedures and and using an ETL tool as a wrapper which is a very common pattern. It definitely is something to look into and and separate out because those are going to be the most difficult uh to attack. And then the database artifacts whether it's store procedures udfs or external procedures and uh materialized views uh equivalent let's say join indices and everything else that you have uh to organize those in a fashion that you could treat it as an object and we do have a tool that that could do that. Lakebridge is the tool to to do the this and then uh of course defining your business processes with all of these pertinent objects that you want to convert. It's a it's a this is a good reference with all of the DBC tables and the PDCR tables that you want to collect your inventory based off of. uh so you know whether it's your columns you know your your you know the monitoring log monitoring users and database and systems and you know what are the you know uh tables that you want to look at but basically as you will hear in the keynote as well the tool that we have uh which is called the lake bridge analyzer it looks at all of these objects and it really breaks down your terate environment into objects that you have levels of difficulty lines of code. So it's a good you know view of the landscape and uh from there you could define your lineage and processes that that are related to each other that you want to convert. So from this perspective is defining a phase decommissioning approach that uh you do have your business use cases and as time goes on you're converting them and you're decommissioning uh your specific business use case but you're enhancing newer use cases based on the the hybrid approach that we were talking about and uh introduce them as the newer architecture, modernize architecture and uh uh make it a repeatable pl process uh from the get- go. Yeah. So when you think about doing a migration of technology, you are thinking okay so in tera I'm doing the staging then I'm doing the data warehousing then I'm doing the data mats. So I will do the same. So this is what uh we call ingestion or ETL first and basically you will need to move your sources your loads and your jobs by one by one. Uh why I said this one is great from like engineer point of view but this is not my recommendation. My recommendation because what you want to achieve is to migrate everything is to work with a serving layer first. And why is that? uh serving serving layers is like the best data you have. You have cleansed it, you have prepared it, you have pre aggregated it, you are serving it into PowerBI or Tableau or click or any tool you like and moving that to to datab bricks first um you're not like in a tunel in a tunnel of migration that would last for two years. you say okay I want to migrate to to datab bricks just move the serving layer and within within weeks you can start you can onboard your uh your your customers you can on board your colleagues and uh with the serving layers having the best data you can start also doing genai you can start doing machine learning you can start to use the other features that are new with datab bricks coming from terod data so the challenge here is to make the proper reconciliation So you work with the data m first and after behind the scene you will build the wool pipeline one one step at a time to be able to to have everything on datab bricks do your final testing and then you can decommission the terod data so this is I I find this to be a good reference uh that we've shown customers that uh in order to derisk the migration it would always be uh better to move your goal tier uh your semantic tier into data bricks environment so you don't interrupt the processing you actually uh make sure you're all all of your SLAs's are addressed uh from a presentation tier but every box represents unlocking the features that we have upon migration to to to the you know as time goes on to the different you know data products s from bronze to silver, from source systems to bronze and everything you will hear in the keynote with the lakeflow connector that we have CDC capabilities lakeflow editor by just typing you could just build out your ETL processing you will hear that tomorrow uh so all of those capabilities uh as you go down the list of the phased approach of you know kind of a safe way of uh you know going down the migration path it becomes really helpful So I've listed out some of the migration challenges that I've seen on my uh you know I've on my path with different customers definitely uh you know that the idea of uh using lineage and using terod data logs to your benefit of how tables and objects are being used. Group them together. get rid of things you don't need based off of the metadata table, last use table, column and and time stamp and things like that to be able to uh really not you know kind of bring about or or migrate things that are no use and uh data validation is very very important to automate that. Again, Lake Bridge as a tool offers that as well. uh you know ter T terod data semantic views terod data loves the idea of using views because of the really high cost of storage um and it works really well by the way uh with multi-join tables it works really well with the storage mechanism that they have but uh recommendation is that most probably using materialized views on our end which I'll show you a demo would be a better approach here uh storage is extremely cheap and the method of updating uh materialized views is all implicit and you don't have to you know write code for it or uh put any scheduling in place. It's all managed within the DT pipelines that we have. uh the most important part of it is how to migrate the data uh using the TPT access module for S3 if it's an on-prem system uh to move it into the the you know the storage environment the uh blob storage or S3 environment or GCS uh or if it's a Vantage that's already in the cloud Vantage cloud link or Vantage enterprise take advantage of the NOS connector uh this way you could just uh move things into the cloud environment that uh we could just pick it up from there and and it becomes a very seamless process. Yeah. So some challenge when you think migrating from Terza to data bricks you say okay I have store procedures it's a challenge but we have one thing which is super great I don't know if you've seen all the announcement we have been uh giving uh giving SQL scripting away and we're working on pro store procedures and we are using the SQL PSM standard which happens to be the same than the standard used by terod data uh it's I think it's the only only procedure standards that exist now and we are leveraging it. So this is an example. So this is not a tool. This is a this is me. I I've typed this code on the left. This is a terod data code. So it's basically it's a small cursor just to to divide and conquer and I I had to build a string to do to do um aggregation aggregation and to make sure I had the good partitioning kicking in and the good uh the good primary indexes in terata. I had to build I had to build a string and execute it. Doing the same in data you can keep exactly the same logic. You see the code is very look alike. The execute immediate would work also in databicks but you don't need it uh because the way we are doing the data layout in databicks with liquid restoring is very efficient and we we don't need like to pass the heart strings uh to have the same query performances. And by the way, this is just to keep the structure of this store procedure. Uh in datab bricks, you will do a join here. You won't do a loop. You will do a join for sure. But I wanted to show if you have some teratas store procedure, you can move it almost as is into datab bricks from a language point of view. So tomorrow there will be a lot of announcement about leg. We are allowed to talk about it but we are still uh letting you some some things to figure it out. But basically any any code you have in ter data would it be TPT views B tech store procedure anything you have you can pass it to to the legbridge converter and it will uh it will create for you the workflows the SQL scripting the store procedure or the pi spark you will be able to navigate this is configurable to to your output expectations but this is a great tool And you you will have most of the announcements tomorrow. Uh this is just some snippets from uh so this is from Legbridge. So this is just some uh some code from terod data. I won't bother you too much with it. This is on the left terod data code. On the right a notebook uh done by leg brbridge converter. I'm just showing you but I don't run run through it. And now the demo. Mean built an awesome demo and I'm promising you you will enjoy it. Sounds good. So in order to do this demo um I established a clearcape analytic environment which is very pretty quick to come up. Uh that's just an instance of uh terod data that I'm connecting to. And uh what I want to showcase here is that using data bricks notebook and I'll make this publicly available uh for you guys to also play around with. The most recommended way is that use our unity catalog to make a connection as a foreign catalog to terod data environment. That feature is extremely rich. I mean by just one simple code and connecting to that environment you'd be able to create a foreign catalog and see all the objects that are within the uh uh within your uh data terod data environment. An example of that would be uh this in this case that I've created my terod data catalog from that clear escape environment. I would be able to see all of the DBC tables, see all of the everything that is part of any database that Terod data has that I have access to. So the credentials of course is important but the fact that you could use this to your advantage when you do a migration for data validation running queries against terror data running it in data bricks see the difference and making sure that uh you are doing the correct uh validation. So with this I'm not going to run this I I could I just might as well run it. it basically looking at the DBC info table going across running pushing down the SQL to ter data environment and getting information back of course because of our connection and the instance that we have you know it's an it becomes a spark job basically and now it's getting the version of the terod data that I'm using and uh everything that I need so from this I could easily select from all the tables that there is on a specific database as you could within in a terod data environment. I just loaded a u demo database that uh terod data has that demo database you select from it. There's three tables that comes out of it. I would be able to uh run queries as such everything being pushed down to terod data and coming back. So uh one thing that I want to note here is that okay now I'm running this query here for example it's going across but it it's hitting a syntax that uh I took a terod data SQL statement it's hitting a syntax that we don't have we don't have the top statement the good thing is with the assistant it really would help you out so you could either go to the assistant and would say fix it for me or you could just diagnose error it basically says well that's a ter SQL syntax top doesn't exist replace it with limit and by replacing it now you have the correct syntax so this is a really powerful feature to go back and forth to do uh conversions uh on the fly and I'll get into more detail of how you could do this in a whole batch format uh you could create views against terod data objects within data bricks without having to bring the data over. So in this case I just went ahead and said okay what why don't I create a view performance is not the major thing here just getting data out uh so creating a view made sense for example in this case selecting from the view in this case I said well why don't I create an object here that is within data bricks and join it to terod data uh and create a view out of that in this case I create a dimension table within uh data bricks and then We'll go to the SQL editor side of it and we want to create a materialized view. Basically, uh I take that dimension table, join it to a terod data catalog table and uh I basically create an object that joins the two on the fly and uh materialize the information. The refresh cycle and the timing of it is all under our control. You could have it to be every hour, every 5 seconds or depending on your set of data available, it could do a refresh for you. So that's another way of doing it. Basically, you don't want to go through the headache of building ETL pipelines and it makes sense from a performance standpoint to bring the materialize everything within data bricks. You could do it through this. You could always refresh it. And then the interesting part of it is that uh you will have um an automatic lineage or information. So right there on the top AI is making a suggestion as to the metadata that it wants to uh assume. You could change that, edit that. It's it makes it pretty uh interesting. You see the view definition, you see the sample data, you see the permissions, you see uh the policies, you see the history of usage, you see the lineage, which is very interesting to me is that automatically within a lineage graph, you would be able to see a terod data object come into the mix along with a data bricks object building out that materialized view that we have a particular schedule for. And um the good part of it is that uh you also uh have ability to see the insights meaning how how much this table is being used joined all the information that typically you would have to kind of search the log files. You would have all of that information and the quality data quality or data profiling step here is very important. Let's say for whatever reason you are not getting what you want to get from both systems. Why not set up a threshold here? And typically uh with interata environment you would build out a full audit table that runs in a uh uh you know execute immediate like uh you know dynamic SQL that runs and it tells you every step. This is a pre-built type of a methodology to add all of your data profiling or data quality rules and then upon exceeding a threshold you would get notifications and things like that. So it could get really involved with uh the boundaries and the governance you set up on the data. So going back to the notebook uh I uh basically wanted to kind of introduce this whole idea of genai uh you know when it comes down to uh the migration uh I've created a much much more uh complex version of this uh that I could share you know with you guys for a customer that was they have uh half a pabyte of uh storage there the storage over 10,000 different objects uh and then they wanted a you know faster way of converting and validating the results and whatnot. So as a just to showcase the capability um I created four you know relatively complex or you know medium complex city level in that uh clear escape environments and now uh this is a definition of the few views that I created and the tables that I've already had in that database from a terod data standpoint so this is just purely getting it from terod data for a faster type of a process, I uh create and replace a table out of it. So you get all of that uh definitions and DDLs. And this is the meat of the activity. basically uh using our AI query function which I really really recommend everyone to use is that you could select your own model of choice. Uh and you could say in this case my recommendation is to use sonnet claude is the best from from a coding standpoint and conversion standpoint specifically 4.0 0 this is a 3.7 that cloud is our partner. So you're going to automatically have that endpoint within your environment. So all you have to say is cloud 4.0 and then be specific about the prompts that you want to do the conversions on. In this case I uh basically ask the LLM itself what would be the best prompt you would ask you to to do a conversion for t migration. This is what it ca it came up with is something that's the list was pretty long but a sample I put it here. First of all it's organized. Uh second you could uh my opinion you could have a very collaborative way of uh having a document attached with all the prompts that do the conversion and have different groups add to it as they encounter problems. And then coming down with the uh the rest of the output, you could be very specific about the data type that it generates, the maximum number of tokens if you want to put a limit on it. And the temperature is basically how deterministic this output is. The higher you get to one, it's very creative. If you put it at zero, it becomes more deterministic. So it comes back with the same result. And it basically all it's doing is that in a very fast batch format fashion is taking every row for that DDLs that you saw and converting it to a datab bricks object and basically uh data bricks SQL in this case. So this is the output basically. Now you could run this in data bricks. It works and it's converting. Uh with other notebooks that I have I put a validation on the top of it. Basically go through every row assume assign a explain plan to it to make sure that it runs and it's not missing anything. And then make this process repeatable. as you find out that LLM is for example missing a data type, you add that specific data type as a conversion. Or let's say you have a full list of tables that you want to convert, you add that as a document that this thing will use and it does a very very accurate job. So basically in summary, we we showcased how to create a foreign catalog with all the objects. uh we selected the objects within terod data. We did a mixture of terod data along with data bricks objects in a view materialized that information into a materialized view so it could be faster accessed with different schedules and then showcasing LLMs how it could be uh of help to uh to do this data conversion and once you play around with the AI query you'll see how easy it is to use and and how simple it is to uh achieve what you're you're trying to do and again On my other notebooks, I've tackled really complex store procedures and uh streamlining the output of the conversion into multiple SQL statements that could be called from a workflow. So you don't have to translate a store procedure that has a monolithic type of a logic into another store procedure which we do offer. But my recommendation is taking advantage of conversion of a store procedure to workflows which has got lots of features and and uh more modularized and you know more definitely uh maintainable when it comes down to it and uh really encourage you guys to take a look at this as an option. So any questions or anything? Uh I will add just one thing. Sure. The world demo you see is done is done in SQL. It was create views, create metal views, running AI query statements all in SQL. So coming from Terata, you're probably more confident with SQL than bypark. It's fine. You can do everything in SQL in data bricks red showcase. Yeah. Yeah. It's a it's definitely a uh it you know from another customer perspective we had they're they're seeing 10x benefits of migration when it comes down to it with the less amount of time that they have to spend on the SQL conversion and manual type of inspection of uh you know and when it comes down to the validation of it. Yeah.

Original Description

Storage and processing costs of your legacy Teradata data warehouses impact your ability to deliver. Migrating your legacy Teradata data warehouse to the Databricks Data Intelligence Platform can accelerate your data modernization journey. In this session, learn the top strategies for completing this data migration. We will cover data type conversion, basic to complex code conversions, validation and reconciliation best practices. How to use Databricks natively hosted LLMs to assist with migration activities. See before-and-after architectures of customers who have migrated, and learn about the benefits they realized. Talk By: Fabien Contaminard, Sr. Specialist Solutions Architect DWH, Databricks ; Mehran Golestaneh, SSA, Databricks Here's more to explore: The best data warehouse is a lakehouse: https://www.databricks.com/product/databricks-sql How to migrate legacy data warehouses: https://www.databricks.com/resources/ebook/migrate-your-legacy-data-warehouse-databricks See all the product announcements from Data + AI Summit: https://www.databricks.com/events/dataaisummit-2025-announcements Connect with us: Website: https://databricks.com Twitter: https://twitter.com/databricks LinkedIn: https://www.linkedin.com/company/databricks Instagram: https://www.instagram.com/databricksinc Facebook: https://www.facebook.com/databricksinc
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Databricks · Databricks · 0 of 60

← Previous Next →
1 Building AI Agent Systems with Databricks
Building AI Agent Systems with Databricks
Databricks
2 Databricks Workflows
Databricks Workflows
Databricks
3 Automate Unity Catalog Upgrade with UCX Part 1: Overview
Automate Unity Catalog Upgrade with UCX Part 1: Overview
Databricks
4 Automate Unity Catalog Upgrade with UCX Part 2: Installation
Automate Unity Catalog Upgrade with UCX Part 2: Installation
Databricks
5 Automate Unity Catalog Upgrade with UCX Part 3 - Assessment
Automate Unity Catalog Upgrade with UCX Part 3 - Assessment
Databricks
6 Automate Unity Catalog Upgrade with UCX  Part 4 - Group Migration
Automate Unity Catalog Upgrade with UCX Part 4 - Group Migration
Databricks
7 Table Migration and Catalog Design with UCX | Part 5
Table Migration and Catalog Design with UCX | Part 5
Databricks
8 Setting Up Azure Access for UCX Table Migration | Part 6
Setting Up Azure Access for UCX Table Migration | Part 6
Databricks
9 UCX Table Migration: Creating Catalogs and Schemas | Part 7
UCX Table Migration: Creating Catalogs and Schemas | Part 7
Databricks
10 Automate Unity Catalog Upgrade with UCX  Part 8: Code Migration
Automate Unity Catalog Upgrade with UCX Part 8: Code Migration
Databricks
11 Streaming to Kafka Just Got Easier with DLT Pipelines
Streaming to Kafka Just Got Easier with DLT Pipelines
Databricks
12 Data Engineering From Data to Dashboards with DABs: Crunching the Cookies Dataset
Data Engineering From Data to Dashboards with DABs: Crunching the Cookies Dataset
Databricks
13 Epsilon helps businesses connect with their consumers using Databricks Data Intelligence Platform
Epsilon helps businesses connect with their consumers using Databricks Data Intelligence Platform
Databricks
14 Unilever transforms operations with GenAI using the Databricks Data Intelligence Platform
Unilever transforms operations with GenAI using the Databricks Data Intelligence Platform
Databricks
15 ActionIQ enables businesses to unlock customer data with the Databricks Data Intelligence Platform
ActionIQ enables businesses to unlock customer data with the Databricks Data Intelligence Platform
Databricks
16 Mixed Attention & LLM Context | Data Brew | Episode 35
Mixed Attention & LLM Context | Data Brew | Episode 35
Databricks
17 Inside Databricks SQL: Engineering innovation with Hans
Inside Databricks SQL: Engineering innovation with Hans
Databricks
18 Inside Databricks: Engineering innovation with Michael Armbrust
Inside Databricks: Engineering innovation with Michael Armbrust
Databricks
19 The Money Team at Databricks: driving revenue and customer growth
The Money Team at Databricks: driving revenue and customer growth
Databricks
20 Unity Catalog unveiled: engineering data governance at scale
Unity Catalog unveiled: engineering data governance at scale
Databricks
21 Create a view in Databricks and share it with Power BI using Delta Sharing
Create a view in Databricks and share it with Power BI using Delta Sharing
Databricks
22 NDUS leverages Databricks Data Intelligence Platform to revolutionize higher education management
NDUS leverages Databricks Data Intelligence Platform to revolutionize higher education management
Databricks
23 Démo Databricks de AI/BI
Démo Databricks de AI/BI
Databricks
24 EMEA Data + AI World Tour 2024
EMEA Data + AI World Tour 2024
Databricks
25 GenAI: The Shift to Data Intelligence - Customer Panel on Industry Use Cases
GenAI: The Shift to Data Intelligence - Customer Panel on Industry Use Cases
Databricks
26 GenAI: The Shift to Data Intelligence - Ft. Ash Jhaveri, VP of Reality Labs Partnerships at Meta
GenAI: The Shift to Data Intelligence - Ft. Ash Jhaveri, VP of Reality Labs Partnerships at Meta
Databricks
27 Virtue Foundation leverages the Databricks Data Intelligence Platform to advance global health
Virtue Foundation leverages the Databricks Data Intelligence Platform to advance global health
Databricks
28 Announcing Synthetic Data Generation in Mosaic AI Agent Evaluation
Announcing Synthetic Data Generation in Mosaic AI Agent Evaluation
Databricks
29 AI/BI Dashboards Embedding - A tutorial
AI/BI Dashboards Embedding - A tutorial
Databricks
30 Bayer transforms global data management with the Databricks Data Intelligence Platform
Bayer transforms global data management with the Databricks Data Intelligence Platform
Databricks
31 Databricks at AWS re:Invent 2024
Databricks at AWS re:Invent 2024
Databricks
32 Hive Metastore and AWS Glue Federation in Unity Catalog
Hive Metastore and AWS Glue Federation in Unity Catalog
Databricks
33 Data + AI World Tour Paris 2024
Data + AI World Tour Paris 2024
Databricks
34 Retail reimagined: Currys data-first strategy to driving growth and improving operations
Retail reimagined: Currys data-first strategy to driving growth and improving operations
Databricks
35 Mixture of Memory Experts (MoME) | Data Brew | Episode 36
Mixture of Memory Experts (MoME) | Data Brew | Episode 36
Databricks
36 Verana Health Data Curation and Innovation with Databricks and AWS
Verana Health Data Curation and Innovation with Databricks and AWS
Databricks
37 Securing SaaS Applications: Obsidian Security on Their Journey with Databricks and AWS
Securing SaaS Applications: Obsidian Security on Their Journey with Databricks and AWS
Databricks
38 Twilio Eng VP on Data Intelligence & AI at AWS re:Invent 2024
Twilio Eng VP on Data Intelligence & AI at AWS re:Invent 2024
Databricks
39 Chegg Eng SVP on Data-Driven Approach to Student Success with Databricks and AWS
Chegg Eng SVP on Data-Driven Approach to Student Success with Databricks and AWS
Databricks
40 Ibotta Personalized Rewards Innovation with Databricks and AWS
Ibotta Personalized Rewards Innovation with Databricks and AWS
Databricks
41 Simplify AI governance with #databricks AI Gateway
Simplify AI governance with #databricks AI Gateway
Databricks
42 Databricks SQL and Power BI Integration
Databricks SQL and Power BI Integration
Databricks
43 Databricks Serverless SQL Warehouses
Databricks Serverless SQL Warehouses
Databricks
44 7 West powers audience growth with the Databricks Data Intelligence Platform
7 West powers audience growth with the Databricks Data Intelligence Platform
Databricks
45 Secret to Production AI: Tools & Infrastructure | Data Brew | Episode 37
Secret to Production AI: Tools & Infrastructure | Data Brew | Episode 37
Databricks
46 Skyflow CEO on Data Privacy with Databricks at AWS re:Invent
Skyflow CEO on Data Privacy with Databricks at AWS re:Invent
Databricks
47 Databricks Clean Rooms Product Demo
Databricks Clean Rooms Product Demo
Databricks
48 Dun & Bradstreet Enrichment & Monitoring, powered by Delta Sharing & Databricks Marketplace
Dun & Bradstreet Enrichment & Monitoring, powered by Delta Sharing & Databricks Marketplace
Databricks
49 Unpacking Libraries in Databricks
Unpacking Libraries in Databricks
Databricks
50 Providence uses an AI agent system from Databricks to help doctors improve their communication
Providence uses an AI agent system from Databricks to help doctors improve their communication
Databricks
51 How State Street Uses AI to Transform Millions of Trades Daily
How State Street Uses AI to Transform Millions of Trades Daily
Databricks
52 Vevo Therapeutics CEO on Curing Disease with Data at AWS re:Invent
Vevo Therapeutics CEO on Curing Disease with Data at AWS re:Invent
Databricks
53 Over Architected with Nick & Holly: Databricks updates for Feb 2025
Over Architected with Nick & Holly: Databricks updates for Feb 2025
Databricks
54 The Power of Synthetic Data | Data Brew | Episode 38
The Power of Synthetic Data | Data Brew | Episode 38
Databricks
55 Use Databricks Lakehouse Federation to break down data silos
Use Databricks Lakehouse Federation to break down data silos
Databricks
56 AI's rugby score: National Rugby League rallies fans with analytics and unified data
AI's rugby score: National Rugby League rallies fans with analytics and unified data
Databricks
57 Open Variant Data Type in Delta Lake and Apache Spark
Open Variant Data Type in Delta Lake and Apache Spark
Databricks
58 How would you sort Ætheldred in the alphabet using Databricks?
How would you sort Ætheldred in the alphabet using Databricks?
Databricks
59 A guide on how to operationalize the Databricks AI Security Framework (DASF)
A guide on how to operationalize the Databricks AI Security Framework (DASF)
Databricks
60 Future-Proof Your Asset Performance Management with Generative AI - Field Assistant Live Demo
Future-Proof Your Asset Performance Management with Generative AI - Field Assistant Live Demo
Databricks

Migrate from Teradata to Databricks SQL using Lakebridge, Unity Catalog, and Databricks SQL for a cloud-born platform with faster innovation and feature releases. The migration involves breaking down the environment into objects, defining lineage, and processes for conversion, and using tools like Lake Bridge Analyzer and Legbridge converter for automation. This migration can result in 10x benefits, including less time spent on SQL conversion and manual inspection, and automated validation of mi

Key Takeaways
  1. Define lineage and processes for conversion
  2. Move sources, loads, and jobs one by one
  3. Migrate serving layer first
  4. Build ETL processing and CDC capabilities
  5. Decommission Teradata
  6. Create a view in Databricks to access Teradata data without moving it
  7. Use a materialized view to join Teradata data with a Databricks dimension table
  8. Automatically refresh the materialized view with a schedule
  9. Use AI to suggest metadata and assume lineage
  10. Set up a threshold for data quality and profiling rules and get notifications when exceeded
💡 The migration from Teradata to Databricks SQL can result in 10x benefits, including less time spent on SQL conversion and manual inspection, and automated validation of migration.

Related Reads

📰
Windmill for Data Engineering: TypeScript/Python Scripts, Flows & Self-Hosted OSS
Learn how Windmill simplifies data engineering with TypeScript/Python scripts, flows, and self-hosted OSS, streamlining orchestrators, internal tools, and secret management
Medium · Python
📰
I Built My Second ETL Pipeline. This Time, I Started Thinking Like a Data Engineer
Learn how to build a production-ready ETL pipeline with Python, Docker, PostgreSQL, and Kestra by thinking like a data engineer
Towards Data Science
📰
JuiceFS Sync for PB-Scale Data Transfers: Resumable Sync, Encryption, and Bandwidth Control
Learn how to efficiently transfer large volumes of data using JuiceFS Sync, which offers resumable sync, encryption, and bandwidth control, ideal for PB-scale data transfers.
Dev.to AI
📰
How Airflow is using AI to make data engineering more resilient, not more complex
Airflow uses AI to make data engineering more resilient by detecting data drift, resuming failed pipelines, and fixing issues automatically, reducing complexity and improving reliability.
Medium · AI
Up next
A Moment Frozen in Time | Arnav Iyengar | TEDxJenks Youth
TEDx Talks
Watch →