Analytics Deep Dive: Data Streaming, Querying and Sharing with AWS - AWS Virtual Workshop
Key Takeaways
This video covers building robust analytics applications and architectures using AWS services, including data streaming, querying, and sharing with AWS Kinesis Data Streams, Amazon Kinesis Data Analytics, and Amazon Redshift.
Full Transcript
foreign [Music] welcome to today's session on analytics deep live data streaming querying and sharing with AWS myself Kalyan janaki I'm a senior big data analytics consultant and today I am joined by Uday Narayanan senior solution architect at AWS as part of today's agenda we will start our session by discussing about modern streaming data analytics architectures on top of AWS then Uday will talk about query Federation with Athena and redshift and other data other data sharing concepts of redshift and redshift ml features let's do a quick recap of key components that make up streaming Analytics the first component is of the source as an example it can be a click stream data coming from mobile devices or web applications as a source then you have a streaming ingestion layer so this is where you have an application tile that is constantly collecting The Source data such as click stream data and Publishing it to the same storage layer right so the stream storage can be anything like services like Amazon cases or Amazon msk and on the on the right side you can see there is a stream processing so once you have a stream data that's been collected so you can process the stream data in continuously for the sake of your real-time applications and finally the process data can be stored into your destinations like data warehouses data links for for future Analytics foreign architecture that brings all the components that was we discussed in previous slide with Amazon Kinesis data streams you can build custom applications that process and analyze streaming data for special needs within seconds the data will be available for your applications to read and process from the Stream what I would like to emphasize is the component consumption part of it on the right side as you can see we have so many number of options for consuming from Kansas data streams you have Amazon Kansas data analytics you can use spark running on top of elastic map reduce or you can use Kinesis data Firehouse to read the data from streaming data on top of your Kansas data streams and load into multiple destinations and you can build custom applications like consumer applications using Kinesis SDK and deploy them on top of Amazon ec2 instances or you can use database Lambda for serverless consumption processing moving on to next streaming service in our AWS streaming portfolio is Amazon Kansas data analytics Amazon cases data analytics is the easiest way to transform and analyze streaming data in real time with Apache link Apache Flink is an open source framework an engine for processing data streams Kansas data analytics reduces the complexity of building managing and integrating Apache Flink applications with other AWS services it integrates with several AWS services and support custom connectors recently uh Kinesis data analytics has launched KDA Studio Kinesis data analytics studio is a completely serverless and allows you to create stream processing applications very quickly with few clicks you can launch a serverless notebooks and query the data streams and get results in seconds you can write code in the notebook in form of SQL python or Scala to interact with the streaming data and get the response for your streaming queries within seconds once your core is ready for production you can convert your notebook into an application and transition with a single click to a streaming application that can process GBS of data per second without having to manage any servers the next service in the streaming portfolio that we'd like to talk is Amazon manage streaming for Apache Kafka Amazon manage streaming for Apache Kafka is a AWS streaming data service that manages Apache of car infrastructure and operations making it easy for developers and devops managers to run Apache of applications and Apache Kafka connectors on AWS without having to need to become experts in operating Apache Kafka so uh msk connect Amazon msk is a managed Kafka with a tight integration with many AWS services you can use AWS Lambda to uh to do a serverless consumer applications that can read data from msk and do a real-time processing you can use Amazon ec2 to launch your producer and consumer applications you can use Amazon Kinesis data analytics and deploy the Flink applications to perform real-time streaming analytics on top of Amazon msk we also have a managed msk connect where you can bring in your own connectors and run on top of msk connect infrastructure and interact streaming data with uh other third-party systems so with that having look at the different streaming services in the streaming portfolio of AWS let's review few common streaming architectures patterns enabled by Kinesis family and msk so going into access log streamings one of the standard streaming use case that we see many customers use access access log streaming access logs basically you can collect logs from any locations like web servers application servers and also so if you have a logs from AWS applications like cloudfront or web application file for Fiverr logs that are that I have been delivered to Amazon S3 you can get a put notification to an image on Q service and you can have a Lambda subscribe to that notification and push all these log files into Amazon cases data Firehouse Amazon Kinesis data Firehouse is a streaming service in a Kinesis family which can load the streaming data into multiple destinations it supports loading data into Amazon redshift Amazon open search Amazon S3 and also it can deliver to custom HTTP endpoints and it supports Dynamic partitioning so that you can partition data in high style for Effective data likes so in our use case the log logs logs messages or log events that are reaching to our Genesis data Firehouse can be loaded into your data Lake for archival and future analytics or if you have a custom log security events that you want to do a reporting you can load them into Amazon redshift and have an Enterprise data warehouses build the reports on top of it or if you want to perform some real-time windowed aggregations you can have Kinesis data analytics consume the data and perform Analytics and finally uh one of the biggest use case for log is a search solution so users wants to have this logs into a search service like Amazon open search so Kinesis data Firehouse can load this data into Amazon open search and you can use kibana visualization for searching and visualizing your log files another architecture is for real-time reporting so so you might have many your SAS applications or application databases and you can Implement any CDC Solutions using Amazon DMS or you can use a Kafka connect running on top of msk connect and have this data streamed to Amazon msk so the source can be any applications like micro services or your application servers directly interacting with an Amazon msk2 once you have streaming data coming into Amazon msk you can have a Lambda processor that constantly consumes modify the data and send it over to your Genesis firehouse and once you have data in kinesi's Firehouse canis's Firehouse completely manages of having these streaming data delivered to Amazon S3 data Lake once the data is in the SD data Lake you can do more ETL so if you if you have a requirement to do Transformations you can make use of AWS glue to transform the data and load it into your data warehouse running on top of redshift and have a quick site reports run on top of it once this has all of this has been deployed your the data coming from your application so internet applications can be sent over to your data warehouse in real time and your reports are running on fixa it can be refreshed in real time moving over to nextq use case is near real time search we have many customers that use the power the dynamodb because of its power and scale so as an example use case you can take an example of a real-time gaming application that's using dynamodb as a backend so you can stream the changes happening in your dynamodb to Kinesis data streams and have a Kinesis data Firehouse delivered and open such service so that you can build an end-to-end real-time such solution and the change is happening in your dynamodb table can be automatically shipped in a real time to an Open Source service and UK your users can use as a open search for searching these events with that uh we'll go ahead and take a look at the demo whatever this demo will be showing a sample system application we would be using msk cluster to stream The Click stream data and then we'll be using Kinesis data analytics application that process the data from cluster Aggregates and then sent back the aggregated data to the same msk cluster for other topics and also I mean Amazon open search for search and visualization Analytics moving over to an msk console this and uh this is an UI file Amazon msk console and let's see how you can create a cluster when you click on the cluster you have options for create and custom create let's take a look at quick create a custom create for more options then you'll have to provide the name of your cluster then you have two projecting options you have an option for serverless and provision let's take a look at the provisioned option here uh the first step would be used to select the Kafka version I'll go ahead and select the default recommended version from msk team and then I have an option to provide the type of your broker node and number of zones you want your Kafka cluster to be deployed across and uh number of Brokers per each zone for the for larger distribution and then you have an option to provide the initial storage that you want to assign for the broker nodes you have an option to customize the configuration of your broker nodes but in an example I'll be using default msk configuration moving over uh because you have provided three azs you'll have to provide networking details I'll select and provide the details of my VPC here so you'll have to provide three different subnets for Three Ages okay then you'll have to choose a security group for your msk cluster so that your clients have can access so this is where you'll have to open the security group for your clients who have access and for produce and consume the data we're going to security settings um uh AWS Amazon whiskey provides um three authentication modes uh it's IAM role-based authentication status clam using password uh that is stored in your secrets manager and also Mutual TLS authentication with the TLs certificates managed by AWS certificate manager and data is by default encrypted in transit between the broker nodes and also you can have two modes between clients and Brokers if you can have a TLS encryption enabled and also you can for your testing purposes you can have a plane test encryption enabled when it comes to monitoring there are different levels of monitoring we can have a basic monitoring enhanced metrics at local level topic level or individual partition levels then msk supports the open monitoring with parameters so that you can get this jmx exporter and have this jmx Matrix sent over to your other application monitoring softwares or third-party softwares then you can select the logs so you can have your broker logs delivered to Cloud watch Amazon S3 or Kinesis data Firehouse in this case we choose um cases um cloudwatch and then finally you'll have to go ahead and click create on the cluster but as cluster would take typically 15 minutes to create I have a cluster that's already been running named as msk cluster click stream so we'll be using this cluster we'll go ahead and see a bunch of commands on this cluster I have an ec2 instance with the Kafka client application installed on it let's go ahead and take a look at a list of all topics so I have already created a couple few topics in this Kafka so example topic is our input topic where we'll be streaming our click stream data and we'll have other other topics for Department aggregation results of Department level uh click user ID level and user session level and moving over let's go ahead and start producing the data so I have a sample Java producer application that runs and produces the click steam data to our to our msk cluster so this starts a Java application in the background and continuously uh streams sample stream data so now that we have the data data being streamed to msk Cluster let's go ahead and see how we can deploy the the currencies data analytics application so this is the interface for Kansas data analytics I have already have a Flink application that's running so if you click on the application it's currently in the running State and it has all the access to stream the data from msk and going into the configuration of application you can take a look at the properties so I have provided the details for bootstrap servers and other other topics and elasticsearch endpoint for it to write the data to so now that uh our Flink application is Con is running we can go ahead and take a look at the details of our of our data in elasticsearch kibana so I have a kibana endpoint running so let's open it so let's select the visualize section and we have already a department count that has been enabled and click on the department count as you can see we have just started the streaming event and that's where you can see the bump in the graph on the department count that has gone up all of a sudden so you can always refresh it back and to take a look at last four hours and you can see the data flowing for last four hours I've been trying this POC uh here for last couple of hours here and there and you can click on the refresh that send our today's demo on streaming analytics and I would like to hand it over to udai thank you Kalyan and hello everybody my name is uden aranan and I am a Solutions architect at AWS today I'll be covering three different topics with you uh we'll do some presentations and uh also do a demo for each of those topics uh we'll start off with uh query Federation so with query Federation uh you can facilitate on-demand access to data from multiple distributed sources in a single query so traditionally consolidating data from different disparate data sources required ETL processing to bring the data together into a shared format or to a Central Storage such as a data warehouse you would then run analytics on top of that data with query Federation you reduce or even eliminate the need to do ETL because you're acquiring the data in place Federated queries make it easier for data scientists and data analysts to analyze data using SQL which is a widely used query language and with query Federation it also becomes really simple to produce hybrid analytics and visual visualizations as well AWS supports Federated querying via two analytic services the first one is Amazon Athena for those who are not familiar Amazon Athena is a serverless interactive query service that can process unstructured semi-structured and structured data sets it uses Presto with a full standard SQL support and works with a variety of data formats like CSV Json orc parquet and abroad now how does Athena do query Federation so Athena uses uh data source connectors that run on AWS Lambda to run Federated queries our data source connector is just a piece of code that can translate between your data source and Athena Athena supports connectors to multiple AWS Services as well as on-premise data stores and it gives users the ability to run SQL queries on top of that data with Athena Federated query you can run SQL queries across data stored in relational databases non-relational databases object stores or even custom data sources and as you can see on the screen uh you know we have connectors to nosql services such as dynamodb elastic cache documentdb relational databases like RDS Aurora and redshift and we recently announced support for several new data connectors including other Cloud providers and isvs so some of the new connectors that we have are sap Hana teradata snowflake Oracle and bigquery and these connectors are developed open sourced and fully supported by AWS and there is no cost to using these connectors for you as a customer and Federated queries in Athena you know enables many different use cases uh we'll just quickly look at a couple of those use cases now so the first is you know you can combine and consolidate data from many different data sources and query it all in place and as I said earlier you can use familiar SQL to join data across multiple data sources for quick analysis and store the results in Amazon S3 for subsequent use so think of a use case where you have a like operational data in MySQL you have some other data in a dynamodb database and then you have some other data let's say in Amazon redshift and you've been asked to consolidate all this data to do some analysis uh so with Federated queries in Athena you can connect to all of these data sources write a single SQL statement to join all of this data and then the results of that can be written out to Amazon S3 which can be further used for machine learning or any kind of reporting analysis uh you can also use Query Federation to do some light ETL work where you can connect these different data sources bring them all together and write out this data into Amazon S3 as I was just talking about all right so uh let's now go into the AWS console and check out a quick demo of how this works so I'm going to go to Amazon Athena here yeah all right so uh before we start uh let's talk about what we are going to see in this demo right so here uh we're looking at uh orders data which is currently loaded in an Amazon Aurora mySQL database uh for those who are not familiar Amazon Aurora is a global scale relational database service from AWS uh it supports MySQL and uh or it comes in MySQL and postgres compatibility uh in this case we are using uh an uh Aurora mySQL database we also have a supplier table which is uh also in the same Aurora mySQL database and then we have a parts table which is in Amazon dynamodb and we have a line items table which is in Amazon which is in Amazon EMR and it's running on edgebase and then finally we have a countries data which is uh on running on elastic cache on on redis right now the business users come to us and say hey I want to know the total profit per country for each year and in our case because of the way the data is stored it isn't easy because all of these data are stored in these different data stores so this is where we can use Federated queries to you know bring all of this data together and do the analysis so the first step here is uh to create connectors that we talked about earlier right so as I mentioned earlier Athena uses data source connectors that run on AWS Lambda to run Federated queries and US data source connector is just a piece of code that can translate between your data source and Athena so uh there are two ways for you to create a data source connector and I'll show you both the ways uh the first one is to use the serverless application Repository so uh we'll create a dynamodb connector now so we'll go into the applications and then we can search for dynamodb like this and then we'll pick the one which is the AWS verified version so I'll click on that and it brings me to this Amazon or the AWS Lambda page where you know we need to provide it with some application settings so uh I've already created this uh before the webinar but I'll show you how to do it so in this case I'm just going to call it uh webinar uh we need to give it a spill bucket so as you know it's querying the data if you need to write it out like some temporary storage uh it does that to Amazon S3 so this is my bucket name that I'm going to provide it uh then I need to provide it with a catalog name and this is the name that we will be using when we run uh the queries in Athena uh rest of it can stay default and then we'll just uh give it the spill prefix again this is the prefix in Amazon S3 where it is going to write the data and then finally I'll just say I acknowledge it and we'll deploy so uh now you know what this is doing is it actually uh creating a Lambda function uh which will be used when we run Athena queries to Federate with dynamodb it calls it Lambda function and that's how the data comes to Athena uh but as I said uh I've already created a function beforehand so if I just sort uh I've already created one instead of Dynamo webinar uh I just call it Dynamo and in the interest of time I'm going to keep uh keep uh going ahead with this uh with this demo the second way to create a data a connector is directly from the Athena page or Athena console page so when you come to Athena you can go into Data sources and you can say create data source and you see all of these data sources that we just saw in the in the presentation you know you can select whichever one you want to Federate with in this case as I mentioned my data source is in MySQL so I'm going to select MySQL I'm going to click next I'll give it a name so we can name it MySQL over here and uh you provide it with connection details right so if a Lambda function hasn't been created you can just click on create Lambda function and it brings you back to the same page that we just saw with dynamodb but in this case I've already created uh the MySQL Lambda function beforehand so I'm just going to click on that and we can just click next and create data source and our data sources available or the data source connector is available right so now once the connectors are created a running queries is as simple as running any any normal Athena query that you would do so in this case uh let's just quickly run uh you know a query to connect to the mySQL database supplier table so if I just go there and I say run so I I at this time it's actually using the Lambda function and it is connecting to the mySQL database and you see the results are are here right uh we'll do one more uh let's just do uh dynamodb and I'm going to run it and in this case in dynamodb we are running a query on the parts parts table and you know your the important thing to understand is it's a nosql database but we are still running SQL queries on top of it using uh the Athena Federated query and again so this is what our data in dynamodb looks like and just to confirm that this is the same table that we have in dynamodb I'll just go into the dynamodbit dashboard explore items part and you will see that you have a part key brand comment container and if I come back here you have a part key brand comment container right so it's the same table so you're running SQL from Athena into uh into the dynamodb table now uh going back to the use case where you know our business wanted to get the total profit per country for each year uh I have this query here so as you can see if you look at this particular sub query we are uh uh querying from the dynamodb table uh or the dynamodb part table then we have the supplier table coming from MySQL we have the line item table coming from hbase which is running on Amazon EMR we have the parts table coming from dynamodb as well uh we have the orders table coming from MySQL and then we have the Nations table coming from redshift or coming from redis uh and then we have these uh like uh where Clause that's kind of joining all of this data together and then in the select statement we're just getting the nation information the year and then doing some calculation to calculate the sum uh the total amounts and then outside the sub query we are just summing the amount to get the total profit so if I just run this query uh as you can see right instead of bringing all of this data together into a Central Storage we are just able to run one query within Athena and using Federated queries we are able to access the data in those different data stores uh and we get our results back so you know we get it by country by year and the total sum of profit all right so uh that's uh the Athena Federated query demo uh we'll go back to the uh presentation for the next topic which is uh Federated queries using amazon.shift so Amazon redshift is a cloud data warehouse provided by AWS and by using Federated queries with Amazon redshift you can query and analyze data across operational databases data warehouses and data lakes and again the idea of Federated query is the same in redshift like we talked about in Athena right it's querying uh data in place from these different data stores so this uh capability enables a new data warehouse pattern which is live data query in which you can seamlessly retrieve data from postgres SQL or MySQL databases or build data into a late binding view which combines operational data local data stored on Amazon redshift and historical data in Amazon S3 now if you look at this diagram here uh I don't want to go into too much detail into how this redshift architecture works but you you have a client on the top uh and you're connecting to a redshift cluster when you're connecting to a redshift cluster you're connecting to the leader node uh leader node is the one that does all of the coordinating and then you have a compute nodes and you can have any number of compute nodes depending on the size of your cluster and compute node is where all of the work uh for you know running the query is happening the results are sent back to the leader node that with who then sends it back to the client right so now when running Federated queries uh redshift first makes a client connection to the RDS or Aurora database instance from the leader node to retrieve the table metadata from a compute node amazon.shift issues sub queries with predicate push down and retrieves the result rows or you know retrieves the data back Amazon redshift then distributes the result rows among the different compute nodes for further processing and once the processing is done the result is sent back to the client uh so with redshift query Federation uh you can integrate the data on your Amazon redshift cluster with operational data in RDS or Aurora MySQL or postgres to perform powerful Analytics so as discussed uh you know with this feature you can perform analytics on your operational data without having to move the data into your data warehouse in redshift or running into any kind of delays with complex ETL processing and ETL tasks uh you're running these queries directly on your data source and hence have access to the most up-to-date data and once you are able to access the operational data stores you can bring the data from multiple data stores and your data warehouse and your data Lake all together by joining them and building reports which can then you know uh be used for doing any kind of ml analysis or Downstream analysis uh business intelligence reports and so on and again the important Point here is that all of this is being done by querying data in place without having to move data from one place to the other with ETL all right so now uh let's go back to the console and just take a look at how redshift query Federation works right so for this particular case I have uh I'm going to uh go here and actually open the RDS console and we'll take a look at uh the data uh the database that we have created uh for this right so if I go into my databases I have this database lab uh database that I've created it's a serverless database and it is a postgres Aurora postgres right so if I just do a like go into my query editor and I quickly connect to it and I need to get my and I'll talk about what the secrets manager is doing in this in this whole thing but for now I'm just going to use this to connect and my database name is host address right so uh as of right now I have uh this uh serverless database and I have one table which is a customer table so if I just go and just do a quick query on this customer table you will see it currently has four rows in it right so this is my data that looks like now I'm going to keep the demo for the redshift Federation very simple uh all we are going to do is we're going to go to a redshift cluster and try to access this data from there so uh how do we do that so let's go into my redshift cluster here and I'm going to use this marketing cluster for this and I'm just going to reconnect and the first step is uh you create an external schema so in this case I'm going to name this external schema as postgres we need to provide information of where is this schema coming from like what kind of schema it is or what is the source for this particular schema and in this case it is a postgres database and in this case the database name is also postgres now this URI is the connection information for your uh for your database right so if I go back into my RDS console and if I go into my databases and I click on this database lab you will see that you know it has this endpoint which is database lab cluster and a whole bunch of information and then amazon.com and that is the same thing here database lab cluster whole bunch of information and then amazon.com uh you then need to provide it with an IAM role this IAM role is giving redshift permission to access uh the RDS database and then finally you have the secret Arn and this is where all of my uh username and password and all of that information is stored right so if I go into my secrets manager you will see that I have this secret uh created uh this is the secret Arn which is what we have on the redshift uh query as well and then if I look at the secrets that's there uh it'll show that you know the username is postgres the password is this value it's a postgres engine this is the host the port and the database identifier which is the database lab uh which is the exact same value as what we have here okay and again uh if you look at the secrets Arn uh this is the same value that we are going to provide over here all right so enough talk let's just run this query now okay so now the external schema is created uh and it threadshift is now able to connect to our postgres database so if I just run this query on the customer table uh you see those same exact four rows that were there in uh in uh in the art in RDS they show up here as well so now just to do another quick test I'm going to go back into my RDS console uh into my query editor and I'm going to run an insert statement just to insert a new row and it succeeded now if I come back to my redshift cluster and do a select star for that table uh the you know the fifth row also shows up so as you can see right with this you don't have to worry about doing any kind of an ETL to move the data from your relational data stores into redshift as data is inserted and updated or deleted in your operational store uh you have immediate access to that data here all right so now uh let's go back to our presentation and on to our next topic which in this case is redshift data sharing now data sharing provides live access to data so that your users are always seeing the most up-to-date and consistent information as it's updated in your data warehouse you can securely share live data with Amazon redshift clusters in the same or different AWS accounts as well as across different AWS regions but before talking about uh redshift data sharing let's look at how customers used to share data between clusters before this feature was available and the challenges they faced many customers who use redshift have multi-cluster deployments in order to share data between clusters customers have to manually unload the data and copy the data from one cluster to the other the producer cluster is where the data is getting copied from and consumer cluster is where the data gets copied to now on the producer side uh this process to manually unload copy uh manually unload and copy the data can become cumbersome and expensive and it usually becomes cumbersome when the number of clusters uh in the organization grows uh you also need to incorporate security and governance into this to make sure the consumers are only getting access to the data they need and nothing more uh there are challenges on the consumer side as well because you need to reconstruct the data that is coming from the producer cluster you need to manage the schema changes there are also challenges on the uh you know there are all the challenges with the stale data because the ETL job that is running is not running continuously right uh so whenever the ETL job runs the data is accurate until it runs the next time the data starts becoming stale uh which means you cannot really provide real-time insights to the business uh and then there's also concerns around incomplete and inconsistent views of the data and then you're creating data silos all of these results in business losing trust in your data because of inconsistent results uh this is my redshift data sharing comes in and helps solve these challenges so data sharing was launched as a simple and secure way to share live data across multiple relationship clusters so this allows customers to access live data in a transactionally consistent fashion without having to deal with ETL processing or data copies the owner of the producer cluster has the flexibility to share the data at either a database level schema level table level view or a SQL user defined function level the consumer cluster gets immediate access to the data there are no delays this means as and when there is an update in the producer cluster the consumer cluster is getting immediate access to the data and the data is transactionally consistent so uh when a query is run on the consumer cluster you will see the latest data data sharing also provides secure and govern collaboration you can share data within your organization or with your external customers or partners the producer cluster will always maintain control of the data set at any point if access is revoked for the customer for the consumer cluster that consumer immediately loses access to the data and as I said earlier with redshop data sharing you can share data between redshift clusters within the same AWS account across multiple AWS accounts and across multiple AWS regions as well now let's just quickly talk about how this works right so with redshift data sharing uh you know or Redtube data sharing is built on Amazon redshift ra3 nodes and redshift managed storage uh with these uh ra3 node types you know you can scale your storage and compute independently of each other as you can see in this picture uh we have a producer cluster on the left which is sharing some objects with the consumer cluster on the right the data is stored in the redshift managed storage layer when the consumer cluster is running uh or when the consumer cluster is running the queries on the shared object the compute from the consumer cluster is being used but the whole data access is happening through the managed storage layers so there is no impact to the producer cluster here uh the producer cluster and the consumer clusters are just regular redshift clusters so you know the consumer cluster has read-only access to the objects shared by the producer cluster but at the same time the consumer cluster can have its own private data that it can read and write to from the from across perspective the producer pays for managed storage and the consumer pays for the consumer cluster data sharing does not come with any cost associated with it now uh there are different use cases for uh data sharing uh I'll talk about one such use case uh and this is around workload isolation where you need to support diverse business critical workloads easily and cost effectively while still maintaining the SLA requirements so with data sharing you can rapidly onboard new analytics workloads now think of a scenario where a developer reaches out to a data warehouse administrator saying they need to do some analysis for the sales data in the production redshift cluster the administrator does not want to give the developer access to production so the administer can administrator can tell the developer to spin up their own redshift cluster in a Dev environment and then share the data they need with the developer cluster this way the developer gets read-only access to the data within minutes of requesting it the consumer does not need to worry about where the data is coming from and do any kind of ETL uh with this approach we are also preventing the creation of data silos the second aspect of workload isolation is uh you can size and scale individual workloads according to the performance requirements so as you can see in the picture we have an ETL cluster uh dashboard cluster and a data science cluster depending on the performance needs all of these clusters can be sized differently and can have different node types and number of nodes as well and finally you can pause and resume clusters as needed so uh let's say the ad hoc query cluster is only being used for a few hours in a day you can pause the cluster during the times it's not being used to save costs all right so now let's again go back to the console and uh check out uh the data sharing capability in practice all right so now uh here we'll be using two different clusters so I have this redshift producer cluster which is in the Oregon region and I have the same marketing cluster that we had used in the past uh which is in the Northwest Junior region so what I'm going to do is I'm going to share some data from the producer cluster and my marketing cluster is going to be the consumer of the data right so uh now in this case if I look at this redshift cluster uh in the public schema it has nine tables uh but uh from as a producer uh you know we want to share some data with the marketing team but we don't want to share all the tables with them right so what we decided was okay we can create a view uh based on the information that the marketing team needs and share The View with them right so I'll just create this uh orders marketing View and if I just do a quick query for that uh we'll uh check out what the data looks like so it has all of this information that needs to be shared with the marketing team now uh the first step to create uh like for data sharing the first step that we need to do is to create a data share so in this case we'll create a data share and call it uh the marketing share and a data share is just a unit of sharing data in Amazon deadshift now once the data share is created we add objects to the data share so in this case we're going to add the public schema this is where uh the view is created and then we'll add the orders marketing view that we just created right so these two objects have been added to our data share uh now we can just quickly take a look to see what our data share looks like right so if I run that you'll see a marketing share has been created and it is an outbound data share which means this cluster is the producer and it is being shared with some other cluster that would be the consumer of this that's what outbound means here uh we can also check what objects are being shared by this data share and we'll see the same schema and view that we just added and let's check out what consumers we have and we don't have any consumers because we have not added any consumers yet now in order to add consumers uh to this data share uh you run this command which is Grant usage so you run Grant usage on data share marketing share which is the data share we created and then we give it a namespace like which cluster are we going to share this data with now this namespace comes from the redshift console so if I go back into my redshift console this is my marketing cluster that is my consumer and then there is this cluster namespace right here right it starts with Phi nine and it ends with B1 and uh that's what you see here it starts with finite and it ends with B1 so this is telling uh redshift that okay I created this data share and I want to share this data share with the marketing cluster okay so this takes care of the producer side of things so this is all you need to do on the producer side now if we go to the consumer side uh and I go into uh this tab here uh first I will just quickly check to see what data Shares are available so in this case the marketing share is available and it is an inbound data share which means in this case the marketing cluster is the consumer uh the next step is to create a database which is a local database uh to the shared objects right so again here we are oops so in this case we're going to run this query to uh create a database called consumer marketing from the data share marketing share which was shared with us of namespace this and this is the cluster namespace for the producer cluster next we need to Grant usage on this data share to aw to the users who are using this redshift cluster and that's pretty much it now you know the data is available for us to run queries here right and again the important point to understand here is that uh this particular view was shared with this marketing cluster from the other uh cluster the other producer cluster right so the other producer cluster still manages uh the data you know they have uh the ability to revoke access uh marketing cluster is just getting access to the data because it was shared with this particular cluster okay now let's actually go and check to see what happens if the producer cluster revokes the access so I'm going to switch my tabs I'm in my Oregon region and this is my producer cluster and I'm going to say okay I don't want marketing cluster to have access to this data anymore so I'm going to revoke the access right now I'm let's switch back to our marketing cluster who is a consumer and then rerun this query that just worked like a minute back right and now I get an error saying that the cross region data share does not exist and why is that because if I come to the data shares I don't see any data share anymore because the access was just revoked if I go back to my producer cluster and I again Grant access and then I switch back to my marketing cluster that is my consumer and I will run this query I see the marketing share again I see the inbox inbound share type and when I read on this query everything should just work as it was working before all right so uh that brings me uh to the end of my data sharing uh demo I just have one last topic uh so we'll go back to the presentation and we'll talk about Amazon redshift ml so Amazon redshift ml uh makes it easy for data analysts and database developers to create train and apply machine learning models using familiar SQL commands in Amazon deadshift so with redshift ml uh you know you can take advantage of Amazon sagemaker which is a fully managed machine learning service without learning new tools or languages and uh you know just using simple SQL statements you can create and train Amazon sagemaker machine learning models and then you can use those uh models in your Amazon redshift uh data warehouse and again you can access them by running simple SQL statements uh so because that shift ml allows you to use standard SQL it is easy for you to be productive uh with New Years new use cases for your analytics data and uh you know to get started all you need to do is like run a create model command which uh when you when you run the create model command redshift is integrating with Amazon sagemaker and it's creating the models for you and once the models are created uh you can use those models in your select statements and if you already have uh like models created machine learning models that have already been created uh redshift ml also supports uh bring your own model so you can bring your own models in here as well right so just a quick overview of how this works so uh when you're creating the model or when you're training your model uh you're just running the create model command uh and you provide it with uh either a table name or a SQL statement and then you provide it with a Target uh column so in this case uh I'm trying to do a churn analysis of how many customers are leaving a certain company so the label uh column is where that data is stored so it'll either have a value of true or false so you need to provide it with a Target column and finally you give it a function name uh which in this case I call it predict underscore customer underscore churn uh when you run this command redshift is integrating with Amazon sagemaker autopilot uh autopilot is like going through different models doing some hyper parameter tuning and it comes up with the best model for this particular data now once the model is created uh you can use that model in your uh SQL statements and you do that by running uh predict statement like you you run that uh to predict the data or to do some predictions on the model you just use the function uh called predict underscore customer underscore churn which is the same function name that we've provided in the pane section and then you just provided with the fields or the columns that go as an input to that function when you run the SQL statement you get the results back right so as you can see two simple SQL statements and you have a machine learning model that you can use in your uh in your analysis uh so there are different use cases but some of the common use cases are to do churn prediction that is what we just talked about if you want to understand the engagement rates for certain campaigns that you're running fraud detection uh Revenue prediction customer behaviors of you know how customers are behaving based on certain campaigns uh all of that can be done using Amazon redshift ML and again this is just a small subset of use cases there are like very many different use cases out there for this particular feature all right so uh my last demo for today I'm just going to go back into my console here and uh just we will go and do a demo for Amazon redshift so uh in this case I have uh claims data that I have already created uh it's oh sorry claims data table which I've already created and this is what my data looks like but the important part here is uh there is a fraud column uh in in this uh in this data set right and I want to identify the fraud column for future data that is coming in so the first step in any machine learning uh use case is you uh split your data into training sets and test uh data sets so here I'm going to create like run this query to create a training data set uh you know so for any uh like the first 7 500 rows go into the training data set and the rest of the data goes into my test data set right so I just created two different tables for this now as we talked about it earlier uh I'm going to run this statement which is going to create the model uh so I'll just run it so here it is creating this particular model based on this particular SQL statement and again it is querying the training table uh my target column is fraud because I want to identify the fraud and this is my function name and then you provide it with an IAM role and SD bucket where all the model artifacts are stored right so if I just quickly run this show model command you will see that the model is currently in training this training will can go on for like uh you know an hour or couple of hours depending on your data set which is why I have already run this once before and if I do a show model for that you will see my model state for this one is ready it's the exact same model I followed the exact same steps right now once the model is created you can run uh uh inferences on top of that data right so here I'm just running like getting some information from my table and then I'm using this function insurance fraud model which was the function name that I provided when I created the model I provided with these three columns and I wanted to calculate the fraud and again I'm doing this on the test data set so the machine learning model was not run on this data set so this is brand new information for us right and it's doing this uh it hits the model it uh you know based on the model it determines okay for this particular user uh this transaction is not a fraud but for this one this transaction could potentially be a fraud right so as you saw like you know just writing a couple of simple SQL statements you're able to build a powerful machine learning models here uh in uh in Amazon redshift using the redshift ml feature all right so uh that brings me uh to the end of uh the presentation today we covered a lot of different topics uh so we have sharing some additional resources that you can use uh to you know uh check out the topics that we talked about and go in a little bit more detail so uh this page provides the features uh or additional resources for the query Federation feature for Athena and redshift and then uh we also have some additional features additional resources for the data sharing and redshift ml uh features that we just talked about as well all right uh so with that I'd like to thank you all for attending uh have a great day everybody thank you [Music]
Original Description
In this hands-on workshop, you'll learn how to build robust analytics applications and architectures using AWS services. Harness the power of federated querying into various data sources enabling you to share data consistently within organizations, build streaming ingestion and pipeline mechanisms for real time analytics and create, train, and build machine learning models right within your analytics architecture for predictive insights out of all your data.
Learning Objectives:
* Objective 1: See how to analyze all your data in place, without moving or copying the data with Amazon Redshift and Amazon Athena.
* Objective 2: Understand how the Amazon Redshift integration with AWS Data Exchange and Amazon SageMaker brings together predictive insights through SQL based machine learning models, working on shared data.
* Objective 3: Learn how to maximize the value of streaming data to unlock new insights with Amazon streaming data services.
***To learn more about the services featured in this talk, please visit: https://aws.amazon.com/big-data/datalakes-and-analytics/ Subscribe to AWS Online Tech Talks On AWS:
https://www.youtube.com/@AWSOnlineTechTalks?sub_confirmation=1
Follow Amazon Web Services:
Official Website: https://aws.amazon.com/what-is-aws
Twitch: https://twitch.tv/aws
Twitter: https://twitter.com/awsdevelopers
Facebook: https://facebook.com/amazonwebservices
Instagram: https://instagram.com/amazonwebservices
☁️ AWS Online Tech Talks cover a wide range of topics and expertise levels through technical deep dives, demos, customer examples, and live Q&A with AWS experts. Builders can choose from bite-sized 15-minute sessions, insightful fireside chats, immersive virtual workshops, interactive office hours, or watch on-demand tech talks at your own pace. Join us to fuel your learning journey with AWS.
#AWS
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
Playlist
Uploads from AWS Developers · AWS Developers · 0 of 60
← Previous
Next →
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
Using Microsoft Active Directory across On-premises and Cloud Workloads
AWS Developers
What is Cloud Computing with AWS? | Hebrew Webinar
AWS Developers
Best Practices for Getting Started with AWS | Hebrew Webinar
AWS Developers
Best Practices for Using AWS Identity and Access Management (IAM) Roles
AWS Developers
Building Scalable Web Apps | Hebrew Webinar
AWS Developers
Dev & Test on the AWS Cloud | Hebrew Webinar
AWS Developers
Storage & Backup on AWS | Hebrew webinar
AWS Developers
Disaster Recovery on AWS | Hebrew Webinar
AWS Developers
AWS Israel News | Episode 1
AWS Developers
Security Best Practices on AWS | Hebrew Webinar
AWS Developers
Ready: Introduction to AI on AWS | Hebrew Webinar
AWS Developers
Set: What is ML for developers? | Hebrew Webinar
AWS Developers
Go!: Building your own ChatBot with Amazon Lex | Hebrew Webinar
AWS Developers
And Beyond: Amazon Sagemaker | Hebrew Webinar
AWS Developers
Building API-Driven Microservices with Amazon API Gateway - AWS Online Tech Talks
AWS Developers
Understanding AWS Secrets Manager - AWS Online Tech Talks
AWS Developers
Best Practices for Building Enterprise Grade APIs with Amazon API Gateway - AWS Online Tech Talks
AWS Developers
Build, Train and Deploy Machine Learning Models on AWS with Amazon SageMaker - AWS Online Tech Talks
AWS Developers
AWS Israel News | Episode 2 | re:Invent
AWS Developers
AWS Floor28 News - January
AWS Developers
AWS Floor28 News - February - Hebrew
AWS Developers
AWS Floor28 News - March - Hebrew
AWS Developers
AWS Floor28 News - April - Hebrew
AWS Developers
AWS Floor28 News - May - Hebrew
AWS Developers
Authentication for Your Applications: Getting Started with Amazon Cognito - AWS Online Tech Talks
AWS Developers
AWS Floor28 News - June - Hebrew
AWS Developers
AWS Floor28 News - July - Hebrew
AWS Developers
Enriching your app with Image Recognition and AWS AI Services - AWS Webinar - Hebrew
AWS Developers
Personalize, Forcast, and Textract - AWS Webinar - Hebrew
AWS Developers
Managing Your ML Development Lifecycle with Amazon SageMaker - AWS Webinar - Hebrew
AWS Developers
Running your ML code in Amazon Sagemaker - AWS Webinar - Hebrew
AWS Developers
Get Started in Minutes with Amazon Connect in Your Contact Center - AWS Online Tech Talks
AWS Developers
AWS Floor28 News - August - Hebrew
AWS Developers
AWS Floor28 News - September - Hebrew
AWS Developers
Deep Dive on Amazon EventBridge - AWS Online Tech Talks
AWS Developers
Advanced Serverless Orchestration with AWS Step Functions - AWS Online Tech Talks
AWS Developers
Living on the Edge - an Introduction to Amazon CloudFront and Lambda@Edge - Hebrew Webinar
AWS Developers
AWS Floor28 News - October - Hebrew - YouTube
AWS Developers
What's New with AWS Storage - AWS Online Tech Talks
AWS Developers
How to Build a Compelling Migration Business Case Using TSO Logic - AWS Online Tech Talks
AWS Developers
Configuring and Managing Amazon S3 Replication - AWS Online Tech Talks
AWS Developers
AWS Floor28 News - November - Hebrew
AWS Developers
Using Relational Databases with AWS Lambda - Easy Connection Pooling - AWS Online Tech Talks
AWS Developers
AWS Floor28 News - December 2019 - Hebrew
AWS Developers
AWS Floor28 News - January 2020 - Hebrew
AWS Developers
Top 10 Data Migration Best Practices - AWS Online Tech Talks
AWS Developers
How to Use Azure Active Directory with AWS SSO - AWS Online Tech Talks
AWS Developers
AWS Tips & Tricks - Amazon Redshift Advisor - Hebrew
AWS Developers
AWS Tips & Tricks - Amazon Redshift Elastic Resize - Hebrew
AWS Developers
AWS Tips & Tricks - Amazon Redshift Spectrum - Hebrew
AWS Developers
AWS Tips & Tricks - Savings Plans & Cost Explorer - Hebrew
AWS Developers
AWS Tips & Tricks - Amazon Redshift Concurrency Scaling - Hebrew
AWS Developers
AWS Tips & Tricks - Training Models with Amazon SageMaker - Hebrew
AWS Developers
AWS Tips & Tricks - Auto Model Tuning with Amazon SageMaker - Hebrew
AWS Developers
AWS Tips & Tricks - Amazon Comprehend - Hebrew
AWS Developers
Understanding High Availability and Disaster Recovery Features for Amazon RDS for Oracle
AWS Developers
Amazon Forecast – Forecasting - From Months to Days (Hebrew)
AWS Developers
Visualize your data with Amazon QuickSight (Hebrew)
AWS Developers
Amazon Kendra (Hebrew)
AWS Developers
AWS Floor28 News - AI/ML Special Edition
AWS Developers
More on: Data Literacy
View skill →Related Reads
📰
📰
📰
📰
Entity Resolution: Why "Show Me Everything About This Customer" Is So Hard
Dev.to AI
Dari Membuat Program Pendeteksi Hujan Sampai Hak Cipta: Apa yang Saya Pelajari Tentang Data &…
Medium · Programming
How to Query Databricks from Salesforce Apex (Without Copying a Billion Rows)
Dev.to · Md Mohiuddin
Can a Data Science Course Really Change Your Career in 2026?-IABAC
Medium · Data Science
🎓
Tutor Explanation
DeepCamp AI