Workload-optimized and AI-powered infrastructure

Google Cloud Tech · Intermediate ·🧠 Large Language Models ·2y ago

Key Takeaways

Optimizes AI workloads on Google Cloud using workload-optimized and AI-powered infrastructure

Full Transcript

[Music] all right well hello it's really great to be here in person with all of you in Las Vegas thank you all for joining uh my name is Mark L and I lead the compute and AIML infrastructure product teams at Google cloud and while it's only been eight months since our last next it feels like our product teams have delivered at least at least two years worth of innovation since then so I'm really looking forward to sharing all of that with you today uh by the way you may be wondering what's here under the white uh curtain so that's a surprise for a little bit later uh so hang tight for that all right now at Google Cloud we always start with our customers so let's start with what we're hearing from all of you you told us that you want to accelerate your company's overall transfer information and capabilities and at the heart of this are three categories of applications first everyone sees the transformational opportunity available with AI and many of you are already integrating AI Services into your applications second you're looking to rapidly build new Cloud native applications and third traditional applications which for many of you represent the vast majority of the workloads and are the foundation upon which your business runs and as technology leaders you have to enable all of this while your core responsibilities are becoming even more complex and while it has always been responsible for reliability security performance cost and sustainability this is taking on a new and even more critical role in the area era of AI enabled services for example if you're taking an existing Foundation model and you're fine-tuning it and grounding it with your own Enterprise data how can you do this securely to protect your corporate data assets or if you're doing inferencing of a large multimodal model how can you ensure great performance and responsiveness right the difference in responsiveness of 10 seconds versus a few milliseconds is huge in terms of the customer experience or finally how do you understand the cost to serve those models in an application and then drive that down so you can scale your customer base profitably having the right infrastructure really really matters more than ever and it's pretty clear to us that the old approaches are not going to meet the new challenges we've never seen the types of Demands these AI workloads are placing on our infrastructure this chart that you see here shows the growth in the size of the number of parameters in large Lang language models they're increasing in size by 10x per year over the last 5 years 10 to the 5th that's a lot of zeros uh so it's pretty clear we can no longer just throw more general purpose Hardware at the problem because the benefits of Moors law are starting to slow down for example transistor count is still doubling every two years roughly but no longer at fixed cost or fixed power as a result we need a more intentional approach to Cloud infrastructure so let's talk about how we're innovating at Google Cloud to lead the way in addressing these new customer requirements and challenges our strategy at Google cloud is centered around the idea of delivering workload optimized infrastructure we look at the unique needs of each and every workload and then we design and optimize at a systems level across Hardware software and services to deliver against those outcomes that I talked about before for and that our customers care deeply about and within this we're also using AI to dramatically simplify infrastructure operations for our Cloud users so let's now get into how this approach and this strategy is resulting in concrete benefits for each and every type of workload starting with AI workloads but before I go there first let me share some broader context on our approach to AI at Google Cloud more and more more customers are choosing Google as their strategic AI partner because of our unrivaled expertise across all the relevant domains it starts with industry defining research in fact the original Transformers paper which was the which is the foundation for all the large language models that are taking the World by storm today came from Google research we then channel that research into Cutting Edge models such as Gemini and many more but it's not just about our models it's also about embracing the ecosystem and open Frameworks and Google is a great partner and great contributor here however the models by themselves are not enough you also need to scale them safely in production applications and we have unique experience running applications that serve literally billions of consumers around the world and to do all of this you need to build on the best infrastructure so now let's dig into our AI optimized infrastructure Google Cloud AI hypercomp is our endtoend systems architecture for AI workloads it's based on over a decade of experience designing and scaling some of the world's most advanced AI infrastructure and services it starts with performance optimized Hardware across compute storage and networking delivered as ultrascale data center infrastructure leveraging Technologies such as high density deployments liquid cool pooling Optical switching and more within this cloud tpus gpus and CPU extensions provide the widest range of accelerated AI compute options next through our open software layer we enable extensive and optimized support for all the leading ml Frameworks including Jacks and pytorch and the ability to scale AI workloads across phys physically desparate Hardware deployments with multi slice training and multi-host inference with Google kubernetes engine or gke integration for orchestration to help developers accelerate and maximize the utilization of that underlying Hardware finally we want to make it simple and easy and costeffective to consume these Services by offering highly flexible consumption models in addition to traditional options we also are introducing new models such as Dynamic workload scheduler which is designed specifically for The Unique needs of AI workloads ultimately the purpose of the AI hypercomp architecture is to improve system performance and efficiency and this is something we apply internally within Google as well overall we find that with this approach we can run more than twice the effective efficiency relative to Baseline Hardware only techniques AI hyper computer can also be flexibly deployed our open approach ensures that we are ready to engage at all layers of the stack for example uh earlier this year we announced a broad partnership with hugging face to allow developers to use Google Cloud infrastructure to further open innovation in AI this allows developers to train and serve thousands of hugging face models using vertex AI and gke and as the AI landscape evolves we'll continue to enable an ever expanding array of ecosystem Solutions bringing all the benefits of AI hypercom Compu to our partners and our customers in whatever way they want and need to leverage it based on the strength of our offerings and our strategy Forester recently recognized us as the clear leader in AI infrastructure Solutions and we scored a perfect five out of five on 17 of the 19 evaluation criteria you see here uh not too bad but my boss is asking when we're going to finish off those last two so we'll see how that goes next year now at the heart of the a hypercomp computer architecture is our industry-leading portfolio of accelerators spanning gpus and tpus covering the full spectrum of your AI workload needs and we have some really exciting news in this area so let's dive in here now let's start with an update on tpus TPU stands for tensor Processing Unit a highly specialized and optimized processor designed by Google that enables some of the fastest and most efficient machine Learning Systems in the world we've been developing and delivering tpus for over a decade now and many of alphabet's Key Products including search Youtube Gmail as well as state of-the-art Google AI uh models such as Gemini are powered by tpus serving billions of users around the world all right so now we get to see what's under the the curtain here so since we're in Vegas VOA here is the tpv 5p uh so you get to see up close in person it has uh two times the flops of the previous generation and three times the high memory bandwidth you can see the the TPU chips there uh on the board so you know these uh these chips and these uh systems are uh pretty impressive but we're get really interesting is how we deploy these at massive scale in our data centers but rather than just talking about that I'd like to actually show you um so recently some of our product managers shot a video inside a Google data center housing these systems let's watch it together now and we'll see if we can hear them over all the noisy data center uh fans uh going so with that let's go to the video to show you how the AI hyper computer comes to life in a data center let's think about how we set up a training toob for machine learning now this is where it all begins with the performance optimized infrastructure that is at the very core of AI hyper computer a scheduler or an orchestrator kicks off a training job that creates a compute cluster for AI workloads these instances use accelerators like Nvidia gpus or Google Cloud tpus this is a Compu engine rack filled with A3 VMS powered by nvidia's h100 tensor core gpus our Custom Design 200 gbit per second intelligence processing units or ipus create a separate data channel for gpus to communicate with each other and bypass that TPU host a TPU is a tensor Processing Unit each chip has tensor cords with four matrix multiplication units or mxus to do the actual arithmetic ppus are custom built inhouse they're specifically designed for artificial intelligence it's all in there on a single machine ppus can have one four or eight chips and multihost workloads can use hundreds or even thousands of these chips together as a pod or a pod SCE multi-slice training jobs can even use tens of thousands of ships this is all amazing everything in here is at a scale that it's just hard to wrap my brain around and you know when we move to OCS for networking we created the opportunity to reduce power consumption by up to 40% this high performance energy efficient power dense infrastructure offers a path to more sustainable it for he us I think it's time we need to get back to Las Vegas we hope that you all enjoyed this look into what makes AI hypercomp Compu a reality back to you mark all right all right so uh so thank you uh Chelsea and Hera truly amazing stuff uh so today I'm really pleased to share that TPU v5p our most powerful and scalable TPU to date is now generally available along with TPU v5e our most efficient TPU to date which is also GA as you heard from Chelsea and Hera TPU v5p can scale to tens of thousands of chips connected over a lightning fast OCS network with an optimized software training stack resulting ultimately in the ability to train large language models nearly three times faster than the prior generation this means customers like Salesforce light tricks essential Ai and Google's own Deep Mind research can train massive models in just weeks instead of months iterating faster and realizing better time to value now to enable choice and optionality for our customers we're also delivering AI optimized infrastructure based on Nvidia gpus we're pleased to announce that next month customers will be able to access our latest Nvidia h100 based instance A3 Mega compared to the prior generation A3 Mega will feature twice the GPU to GPU networking bandwidth to accelerate training and inference workloads for some of the largest models in the world in in addition we're bringing the next generation of nvidia's newest Blackwell gpus including b200 and Grace Blackwell 200 With Envy link 72 to Google Cloud starting early next year deployed on Google's fourth generation of advanced Warehouse scale liquid cooling the Blackwell generation gpus will also feature the next generation of Nvidia networking Google Cloud instance based on blackw gpus will of course uh be delivered as high performance reliable elastic cloud services uh with the full capabilities of our AI hypercom computer architecture to power the next generation of AI training and serving now shifting to open software open source software is a key Catalyst for much of the Innovation happening today and Google cloud is a top contributor and innovator here last year we open sourced Max te text to provide performant and customizable llm training on jacks on tpus now we're expanding it to support Nvidia gpus as well delivering industry-leading llm training performance with Jacks across both tpus and gpus next we're introducing Jetstream this is an open-sourced throughput and memory optimized llm inference engine for xla devices starting with tpus that offers up to three times higher performance per dollar on Gemma 7B and other open models by using Advanced Techniques like continuous batching quantization and efficient attention algorithms we also continue to advance pytorch xlaa with new techniques delivering up to three times higher training performance for image diffusion models all right now on to storage so storage and data integration is fundamental to every stage of the AI life cycle as you see here gpus and tpus are incredibly valuable and hungry resources and require high performance storage to keep them wellfed and highly utilized and this data needs to be highly resilient and secure so we're significantly expanding our storage capabilities and portfolio to meet a broad range of AI use cases first we've optimal optimized parallel store which is an ultr low latency parallel file system used for training models we've improved the caching of training data from cloud storage to parallel store delivering up to 3.9 times faster training and fine-tuning times saving our customers both time and money next file store is also well suited for models that require low latency but with smaller file sizes because file store is an NFS solution all the data is immediately available to all of the GPU and TPU instances in a cluster enabling thousands of instances access to the same data and delivering up to 56% faster training times last year we also introduced cloud storage fuse this enables you to mount a cloud storage Google Cloud Storage bucket as if it were a file system cloud storage fuse combines the high availability and low cost of GCS with a simple standard file system interface and now we've enhanced even further with cashing as well keeping the data closer to the accelerators and improving training Times by up to 2.3 relative to Common Alternatives finally I'm excited to announce a new block storage offering tailor made for AI called hyperdisk ml suitable for inference and training across hundreds or even thousands of instances hyperdisk ml can serve up to 1.2 terabytes per second of aggregate throughput by automatically replicating and caching the data across the storage servers hyperdisk ml delivers up to 12 times faster model load times for inferencing than common Alternatives so huge capabilities across the storage portfolio now networking also has a fundamental role to play in Ena in large scale Ai and ml training and inference first to enable large scale training we're enabling pedabytes per second of bandwidth between data center zones second for inference we're enabling load balancing algorithms which are aware of the TPU and GPU utilization and can automatically distribute traffic optimally across that set of instances to reduce cost next with our compute clusters within our compute clusters our titanium-based Network offloads enable high bandwidth low latency data transfers across compute nodes in a cluster and third we're enabling large data pipes with cross Cloud interconnect and dedicated interconnect so you can access and integrate data from any Cloud across on-prem or other public clouds to Google Cloud for training or fine-tuning for more detail on our cross Cloud networking and Google distributed Cloud uh which brings Google cloud services to Edge locations my colleague sain Gupta had a great Spotlight session earlier today and if you weren't able to join that live I'd highly recommend that you watch it when you get a chance on demand all right last as I mentioned earlier uh we're also enabling new consumption models specifically for AI workloads for use cases like small training jobs fine-tuning experimentation and others you want want to be able to specify the amount of capacity and the type of capacity you need for how long you need it and then be guaranteed it's going to be there when you need it and that your job will run to completion and of course you only want to pay for it when you're actually using it makes sense so to support these types of use cases we recently introduced Dynamic workload scheduler or DWS DWS supports two modes of operations first with calendar mode you can specify the amount type and location of resources you need and the specific calendar days that you need them and we will Reserve those resources for you second with flex start mode you also specify what you need and we'll reserve the resources across our Fleet and start the corresponding job as soon as possible customers like two Sigma are using DWS as you can see from the quote here very effectively to maximize their cost savings while still still accelerating their AI workloads all right so we've covered the AI hypercomp computer and the Investments we're making here but what really matters is how this comes to life for our customers that's why I'm super excited to invite Salesforce to join us on stage here now to share how they have used these capabilities within their organization so please join me in welcoming shush and sho from Salesforce hi great to see you thank you hi great to see you thanks for joining all right fantastic so uh welcome and thanks you both uh for joining so uh shush shisha maybe we'll start with you um so could you please start by sharing a little bit more about what you both enable for Salesforce thanks Mark it's great to be here with you and share our experience with all of you we enable and provision infrastructure trusted compute for AI research teams to support Cutting Edge research development and deployment we actively engage and collaborate with our internal and external stakeholder Community such as gcp leveraging their expertise to propel us to new levels of scalability efficiency and Innovation fantastic so thanks thanks for sharing that um now maybe over to you so I know um you know maybe take us back a few years in time when you were starting before you had Google cloud and then how you started to take advantage of Google cloud and our partnership to enable new capabilities uh in your environment sure Mark we started with Google cloud in 2018 and at that time the a research team was relatively small and predominantly relied on Standalone Computing stances for training their research models and for stories they used volumes that were not scalable and had lower throughput capabilities and as a team and business expanded challenges arose in managing Standalone instances for individual resource scientists we knew this is not where our time was best spent and we decided to move to gcp and then the team turned to Google kubernetes engine as one of their first services on Google cloud with gke it allowed us for dynamic provisioning of instances autoscaling and seamless integration with Cloud monitoring and research models are trained and hosted on gke where each researcher can schedule their workload on multiple flavors of compute engine instances and data is stored in Google file store in instances where each part gets mounted with either a shade or individual file store and all these resources are provisioned in a VPC Network which has which have multiple subnets with large CER ranges to accommodate enough compute resources all right that sounds really great and it's just wonderful to hear about your journey uh with Google Cloud uh and the uh the evolution of your infrastructure um so now maybe can you share a little bit more about the software and the Frameworks that you enable for your researchers and your developers and the work you've done to help them be more productive absolutely so the researchers have much better platform to work with now and their flow goes like this first the ml scientist submits their yaml configuration to gke with desired Docker image and resource configuration such as CPU GPU and memory then GK Provisions the compute nodes for the configuration in the yl and pulls the container image from container registry onto the node and creates a pod and once the desired workloads get completed kubernetes engine autoscaler Dre Provisions the compute engine nodes and on the software Frameworks we use both pyo and Jacks which are both popular machine learning libraries with specific strengths and we used Jacks mainly on tpus and there was no need for customizations we were able to use it right out of the box on Google cloud and for gpus we use py torch yeah that that's great to hear you know that that choice and flexibility is is really important for a lot of our customers um that's why we're making the Investments we talked about before around Jacks and both pytorch so it's really great to hear how you're using that within within your environment so you know since that initial deployment maybe you can share a little bit more about how this environment has evolved over the last few years sure starting in 2019 we we began utilizing Google CL platforms tpus to scale our models to billions of parameters in size since then we have used TPU V2 V3 and V4 and as our data volumes continue to grow we also saw an increase in throughput by utilizing file store High scale instances right that that that's great so thank you so it's it's great to see how you you're using you know virtually every aspect of the AI hypercomp architecture um from gke policy based to G uh gpus and tpus to Jackson pytorch um really fantastic to see how you've enabled that um you know in your environment and um you know shushma maybe back to you uh before we close here um so looking into the future a little bit can you share more about what's next uh for your environment uh on Google Cloud Sure Mark we're thrilled to make such rapid progress and foundational research through AI hypercomp computer uh we are using the latest generation TPU we 5 in fact just last week uh we got the results from our early testing we're seeing um the training time cut down almost in half for our foundational models and we are still improving we recently started using Google Dynamics workload schedul to make it easy to access Google gpus we're really excited to continue our research using conversational AI natural language processing multim morality and time series intelligence to solve problems for our Enterprise customers running on Google has enabled us to accelerate product Innovation thank you for the partnership well thank you both as well for the great partnership uh thank you for your investment in Google cloud and thank you most for taking the time to be with us today and share your journey with us it's been really great and we look forward to an even deeper partnership going forward thank you so much thank [Applause] you all right so AI is clearly top of mind for all of us right right now but we also know that for many of you the vast majority of your workloads are the ones that the ones that use to keep your business running are Cloud first applications and Enterprise applications so now we're going to shift gears to look at what's next for those workloads from Dev test to Mission critical applications and everything in between we're delivering workload optimized systems across compute storage and network for each and every workload and we're really excited to announce the most significant expansion of products and capabilities across this portfolio as you see here we'll go now through each of these starting with titanium unlike competing approaches that only use off hods offloads and accelerators on the host titanium uniquely combines smart accelerators on host with a tier of offhost scaleout servers this this approach allows us to scale performance Beyond what's possible on a single host and it enables our Market leading price performance for products like hyperdisk today we're pleased to introduce three new capabilities enabled by titanium first the ability to perform non-disruptive or hitless maintenance complimenting our live migration capabilities second security capabilities such as protection against Data Theft due to exfiltration and preventing access to user data and user encryption keys by operators third it enables new high- performance bare metal compute instances which I'll describe in more detail later now onto our new compute instances we're excited to introduce C4 and N4 the Next Generation general purpose VMS powered by fifth generation Intel Zeon processors these are the first such VMS from any leading Cloud hyperscaler C4 provides a dedicated slice of hardware for each VM offers larger shapes and is engineered to provide consistently high performance and a highly controlled maintenance experience as a result C4 delivers 19% Better Price performance compared to other leading Cloud providers and also has up to 80% better CPU responsiveness for real-time workloads C4 is ideal for your most Mission critical performance and latency sensitive workloads on the other hand en4 perfectly balances cost and performance by leveraging Google's Innovative Dynamic Resource Management capability which enables us to efficiently utilize those underlying compute infrastructure resources and pass those efficiency gains on to our customers in the form of lower prices as a result N4 delivers lower Computing costs overall while still offering 18% Better Price performance compared to our prior generation N2 instances by leveraging C4 and N4 in combination across all your general purpose workloads you can achieve the optimal level of performance for every workload while also reducing your overall TCO across the fleet customers like paloalto networks are enthusiastically adopting C4 and N4 to make the most of these benefits now while general purpose compute instances are versatile and cover a broad range of workloads some applications have even more specialized requirements applications such as scaleout databases logging and other storage dense applications require large capacities of local storage with very high iops to meet these needs we're excited to announce that our new Z3 storage optimized instances are now GA they support up to 36 terabytes of high performance local SSD with 6 million read and 6 million right iops customers like Aeros Spike are using Z3 to offer their services with both very high performance and amazing cost Effectiveness for storage heavy workloads now moving along some of our customers and our isv partners need tightly controlled process execution with direct access to the CPU or some cases want to use their own hypervisor this fundamentally requires bare metal hardware but delivered as a cloud service to meet this need we're really pleased to announce C3 metal instances with bare metal instances you get the entire resources of the compute server directly available for your workloads making the server inherently single tenant for Optimal Performance Partners like nanic red hat repet and many others are excited to leverage C3 metal instances to enable their services to their customers next let's talk about uh possibly one of the most Mission critical workloads in the world sa Hana to support the highest end inmemory databases such as sap Han deployments we're now announcing our new X4 memory optimized Cloud native bare metal instances with configurations up to 32 terabytes of memory X4 is our largest memory optimized instance and with an industry-leading three and a half n single instance SLA X4 instances are ready for the most Mission critical sap deployments now sap and Google Cloud are Partners but sap is also a customer of Google Cloud infrastructure with their rise offering the sap ECS team operates sap Hana environments for many customers on Google cloud and testing by this team has been very very positive as they prepare to use this X4 system for their sap rise offering for our mutual customers now moving on Google cloud has a long history of Designing custom processors for specific needs in fact since 2015 we've released five generations of tpus two generations of video coding units and three generations of tensor processors for Pixel phones now we're bringing that same level of engineering prowess and Innovation to general purpose data center Computing for many customers this is where some of the most difficult and critical decisions are being being made you need to balance competing priorities maximizing performance while enabling interoperability while also lowering costs and meeting sustainability and efficiency goals that's why today we're really excited to announce Google Axion processors our first Google custom-designed arm-based CPUs and uh actually someone let me borrow one from the keynote this morning so I also have one to show here live in person pretty cool um so we'll put that away for later um we're going to be building on Axion as the heart of upcoming general purpose compute instances and as you heard earlier today Axion Google Axion is a breakthrough in general purpose Computing Axion instances deliver up to 30% better performance than the fastest general purpose arm-based instances available from other leading clouds axon instances also offer up to 50% better performance than comparable current generation x86 based instances and are 60% more energy efficient than comparable x86 based instances Axion will be fully supported by Google services such as gke data proc Cloud SQL and others and will also power many of our internal Google consumer services because Axion is based on the latest industry standard arm compute designs core compute designs you can quickly deploy arm compatible applications with support from the large and ever growing arm isv ecosystem and be confident that will it will work out of the box on Google Axion and many of our customers are excited to take advantage of these new capabilities as you see here now continuing with the range of compute options Google Cloud VMware engine or GCV for short is a very popular choice for customers who want to rapidly move their VMware workloads to Google Cloud hundreds of Enterprises running on Google Cloud VMware engine have achieved better elasticity better uptime and enjoy the benefits of cloud native networking and proximity to all the Google cloud services recently broadcom and Google announced that Google Cloud will be will be supporting the new VMware cloud found Foundation byol licensed portability option this is huge customers who have the new vmor Cloud Foundation subscription on Prem can now save over 30% by avoiding the need to buy new licenses you can just bring your existing VMware licenses and only pay Google for the underly GCV Hardware instance in addition you'll continue to have the option to purchase a licens included GCV service from Google we're excited to be the only Cloud to offer these savings and this flexibility to our customers so many great compute options but we're not quite done yet let's also talk about storage again but this time in the context of general purpose workloads starting with block storage last year we shared how titanium offloads enable hyperdisk to deliver better storage performance than any other leading Cloud today I'm pleased to announce General availability of hyperdisk advanced storage pools now historically block storage capacity was stored in silos in volumes for each individual VM where it was often significantly underutilized with our new storage capacity pools customers can purchase a single shared pool of capacity across workloads by using a single storage pool combined with thin provisioning of the storage to the VM and with dup and compression within that common shared storage pool we can achieve dramatically higher storage utilization and over 50% cost savings for customers in typical scenarios Google Cloud again is the only leading Cloud that offers such storage capacity pools now let's talk about file storage last year we introduced Google Cloud netup volumes this this is a great option for customers that have standardized on Neta on Prem and want to bring those workloads and that data to Google Cloud today we're announcing three new capabilities here first policy based file tiering that enables you to tier data to colder and less expensive storage based on file access times while still providing great performance when it matters ultimately delivering up to 60% cost savings second you can now create massive petabyte volumes and we're tripling the performance the throughput performance to 12 gigabytes per second and third we're introducing a new flex service level for Less performance sensitive applications enabling flexible storage volumes from 1 Gigabyte to 100 terabytes with up to one gigabyte per second throughput all right so as I said at the beginning we have a lot of new products and a lot of new capabilities to share and we've covered a lot of ground here today now I suspect some of you in the audience may be thinking hey this is great but how do I know what to use across all of this and where do I even start over the last year we've heard from many of you that you spend too much time learning and operating clouds so that brings me to our newest offering Gemini Cloud assist this is our new generative AI assistant for Google Cloud operators we're thrilled to bring AI assistance to help you design optimize and operate Google Cloud let's welcome Jeff Welsh to the stage to give us a peek at it thanks welcome thanks for coming you well hello everybody thanks for coming today and spending this time with us I'm very excited to talk about Gemini Cloud assist today because to me it feels like we're at the precipice of maybe a New Era of computing an era where we don't even yet know what the possibilities AE and await for us and can change our our Computing capabilities so in order to help communicate this I'd like to do some demonstrations and so these are going to be live Demos in the interest of time and so you don't see my typos I'm going to copy and paste some commands over but what I want to do is demonstrate how we can create some resources maybe some Simple Resources more complex resources and then also highlight how we can operate and optimize your environment and hopefully after this you'll understand how we can really help you use and learn about Google Cloud both for the new and power users Al like so let's go ahead and jump right into that and let's say I want to create a a virtual machine to run a web server a simple virtual machine instance I might start with the prompts such as let's see how do I create an inexpensive VM for running a web server now when I've submitted this prompt to Gemini it's trying to understand what it is I intended to do now I've intentionally left this prompt a little bit vague so it is going to actually ask me for clarification so it has identified of our virtual machine families there are two that it would recommend for running an inexpensive uh web server the E2 and wow that just announced and for instance which of course I'm going to select because I think titanium is bringing some great benefits to to my use case so now that I've provided N4 as my follow-up Gemini has the data that it needs in order to um fulfill my requests and help me understand how to create that instance and so what Gemini has done is created a g-cloud command that I could copy and execute in any terminal or use our native Cloud shell integration to execute directly in my Pantheon instance and so by pasting this command line in I can edit this modify it or just go ahead and hit go and have that instance be created so while it's great that I can create an individual instance it's rare that an application really runs on a single instance generally you're going to run on an application across multiple instances and so I might want to create a load balanced and autoscaling infrastructure to support this application I could do that with gcloud but I'm also a proponent of infrastructure as code and a fan of terraform and so I'm going to ask Gemini for help to create a a load balancer that will Auto scale across VMS using a managed instance group and have some firewall rules installed so this process is uh requires creating multiple resources through g-cloud this is about an 11 command example but with terraform I'm able to create this environment with a simple piece of code that I can copy out now this has created an instance template that defines the configuration of virtual machines when they're configured it's created an instance group that allows me to manage these at scale and all of the associated load balancer and networking artifacts required such as external IP addresses Etc so now I can execute this this terraform code in my development Pipeline and and have that environment created so it's great that and I can create resources but we can also help you manage your environments let me just refresh my page here and go to my instance ah sorry hello full screen's getting me here instant scripts is what I'm looking for so let's say I've deployed an environment an application running in that load Bal for example with autoscaling I have a an environment here where maybe I'm trying to prepare for an upcoming event where I expect that I'm going to have an increase in traffic maybe my application is going viral or my um I'm releasing a new product such as the N4 and I'm expecting an increase in traffic so first I can look at my instance group and I can validate my configuration that I'm requesting my instances to Auto scale at 60% CPU utilization I might want to tune that and but before I do that I really want to understand what is the actual performance of my instances Gemini can help there as well and more than just create resources it can interact with my environment and in real time time query my deployments understand what it is the performance actually is so in this example again there's two managed instance groups and it has asked me to specify which one I would like and I'm going to choose Wiki frontend which is the one we just highlighted now in the back end Gemini is actually quering my Telemetry data and my apis to understand and display for me the average CPU utilization of the instances in that instance group over in this case the last two weeks and so from this I can see my average CPU utilization is say around 30% so I have some Headroom here I might want to optimize this deployment and say allow my instances to support up to 70% of traffic and then also maintain my cost by restricting the total VMS to 12 I can ask um Gemini for support in this regard as well to say help me set the target CPU utilization at 70% and the max V at 12 so again Gemini now with this very specific command can tell me exactly how to perform that now it has created a g-cloud instance or a command to do so I can copy this to the command shell and execute this just as I could create an instance now in this example Gemini is executing against my running environment it's not telling me how to create anything new it's actually modifying my environment you can see the command succeeded and if I refresh the page here and I scroll over you'll be able to to see that the target utilization is now at 70% so Mark with four commands we've been able to create an instance create a more complex compound uh resource such as a Autos scaling load balancer and investigate and tune our environment wow that that's uh that's fantastic awesome Jeff thank you thank you so um you know you showed us the entire life cycle with Gemini Cloud assist must feel pretty good to be done with the demo now oh it's a load off my shoulders that's sure yeah so uh should we head down to the blackjack tables after this yeah let's do it oh wait actually I forgot you're not quite done optimizing yet from what I saw in the demo there yeah you're right um but you know the good thing is today we announced Google Cloud assist gem Cloud assist for mobile as well so let's wrap the Sim goit what a surprise all right that's awesome thank you all thanks Jeff that was great all right so uh so thank you Jeff um and thank you all for taking time to learn more about our workload optimized infrastructure portfolio collectively as an industry we're at a very very exciting time I hope you found this session useful for you and for your business and thank you all and I hope you enjoy the rest of next thank [Music] you

Original Description

Join this session to hear how customers are building and running AI workloads at scale, while also optimizing their enterprise and cloud native applications with infrastructure(across compute, networking, and storage), purpose built for their workload. Learn how to increase productivity and efficiency with innovation in every layer of our supercomputing architecture, AI Hypercomputer, including TPUs and GPUs. Plus, see first hand how AI-powered assistance with Gemini for Google Cloud is making infrastructure operations easier than ever. Expect to leave with workload optimized infrastructure best practices for your organization. Speakers: Mark Lohmeyer, Srinath Reddy Meadusani, Sushma Gundlapally, Jeff Welsch Watch more: All sessions from Google Cloud Next → https://goo.gle/next24 #GoogleCloudNext SPTL205 Event: Google Cloud Next 2024
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Google Cloud Tech · Google Cloud Tech · 0 of 60

← Previous Next →
1 I’m going for it #GoogleCloudCertified
I’m going for it #GoogleCloudCertified
Google Cloud Tech
2 I had to get #GoogleCloudCertified
I had to get #GoogleCloudCertified
Google Cloud Tech
3 Be better overall at what you do #GoogleCloudCertified
Be better overall at what you do #GoogleCloudCertified
Google Cloud Tech
4 Cloud Monitoring on our radar #Analysis #Uptime
Cloud Monitoring on our radar #Analysis #Uptime
Google Cloud Tech
5 Introduction to Generative AI Studio
Introduction to Generative AI Studio
Google Cloud Tech
6 How to use Github Actions with Google's Workload Identity Federation
How to use Github Actions with Google's Workload Identity Federation
Google Cloud Tech
7 Introduction to Responsible AI
Introduction to Responsible AI
Google Cloud Tech
8 Networking updates and CDMC-certified architecture
Networking updates and CDMC-certified architecture
Google Cloud Tech
9 Create and use a Cloud Storage bucket
Create and use a Cloud Storage bucket
Google Cloud Tech
10 How to digitize text from documents
How to digitize text from documents
Google Cloud Tech
11 Faster analytical queries with AlloyDB
Faster analytical queries with AlloyDB
Google Cloud Tech
12 Next ‘23 sessions and FaaS Wave
Next ‘23 sessions and FaaS Wave
Google Cloud Tech
13 Introduction to Assured Open Source Software
Introduction to Assured Open Source Software
Google Cloud Tech
14 BigQuery Cost Optimization: Storage
BigQuery Cost Optimization: Storage
Google Cloud Tech
15 BigQuery Cost Optimization: Compute
BigQuery Cost Optimization: Compute
Google Cloud Tech
16 BigQuery Cost Optimization: Select Queries
BigQuery Cost Optimization: Select Queries
Google Cloud Tech
17 Remote Field Equipment Management with Manufacturing Data Engine
Remote Field Equipment Management with Manufacturing Data Engine
Google Cloud Tech
18 Supercharging your applications with Cloud SQL Enterprise Plus
Supercharging your applications with Cloud SQL Enterprise Plus
Google Cloud Tech
19 Vector Support on our radar #GenAI
Vector Support on our radar #GenAI
Google Cloud Tech
20 Architecting a blockchain startup with Google Cloud
Architecting a blockchain startup with Google Cloud
Google Cloud Tech
21 Kubernetes and multitasking updates!
Kubernetes and multitasking updates!
Google Cloud Tech
22 GKE: Using Kubernetes Events
GKE: Using Kubernetes Events
Google Cloud Tech
23 How to configure firewall rules for Cloud Composer
How to configure firewall rules for Cloud Composer
Google Cloud Tech
24 Vertex AI Embeddings API + Matching Engine: Grounding LLMs made easy
Vertex AI Embeddings API + Matching Engine: Grounding LLMs made easy
Google Cloud Tech
25 Geospatial analytics on our radar #EarthEngine #BigQuery
Geospatial analytics on our radar #EarthEngine #BigQuery
Google Cloud Tech
26 Ensuring requests are set in Kubernetes
Ensuring requests are set in Kubernetes
Google Cloud Tech
27 Cloud Next 2023, Google research program, and more!
Cloud Next 2023, Google research program, and more!
Google Cloud Tech
28 How to migrate projects between organizations with Resource Manager
How to migrate projects between organizations with Resource Manager
Google Cloud Tech
29 How to run #MySQL in Google Cloud
How to run #MySQL in Google Cloud
Google Cloud Tech
30 #GenerativeAI for enterprises and #Next2023
#GenerativeAI for enterprises and #Next2023
Google Cloud Tech
31 How Google Photos scales to store 4 trillion photos and videos
How Google Photos scales to store 4 trillion photos and videos
Google Cloud Tech
32 Google Cross-Cloud Interconnect (Demo 2)
Google Cross-Cloud Interconnect (Demo 2)
Google Cloud Tech
33 GKE Cost Optimization Golden Signals: Introduction
GKE Cost Optimization Golden Signals: Introduction
Google Cloud Tech
34 GKE Cost Optimization Golden Signals: Workload Rightsizing
GKE Cost Optimization Golden Signals: Workload Rightsizing
Google Cloud Tech
35 GKE Load Balancing: Overview
GKE Load Balancing: Overview
Google Cloud Tech
36 GKE Load Balancing: Best Practices
GKE Load Balancing: Best Practices
Google Cloud Tech
37 Disaster Recovery in GKE
Disaster Recovery in GKE
Google Cloud Tech
38 How to configure IP masquerade agent in GKE Standard clusters
How to configure IP masquerade agent in GKE Standard clusters
Google Cloud Tech
39 Enable and use GKE Control plane logs
Enable and use GKE Control plane logs
Google Cloud Tech
40 Compliance in Australia with Assured Workloads
Compliance in Australia with Assured Workloads
Google Cloud Tech
41 Creating budgets and budget alerts in Google Cloud #FinOps
Creating budgets and budget alerts in Google Cloud #FinOps
Google Cloud Tech
42 Cloud SQL Enterprise Plus on our radar #mySQL
Cloud SQL Enterprise Plus on our radar #mySQL
Google Cloud Tech
43 What's Next for Google Cloud?
What's Next for Google Cloud?
Google Cloud Tech
44 How Loveholidays scaled with Contact Center AI
How Loveholidays scaled with Contact Center AI
Google Cloud Tech
45 What is fleet team management in GKE?
What is fleet team management in GKE?
Google Cloud Tech
46 Troubleshoot VPC Network Peering
Troubleshoot VPC Network Peering
Google Cloud Tech
47 Introduction to DocAI and Contact Center AI
Introduction to DocAI and Contact Center AI
Google Cloud Tech
48 Cloud Run Direct VPC egress explained
Cloud Run Direct VPC egress explained
Google Cloud Tech
49 Database deployment options in GKE
Database deployment options in GKE
Google Cloud Tech
50 Analyze cloud billing data with #BigQuery
Analyze cloud billing data with #BigQuery
Google Cloud Tech
51 Tips to becoming a world-class Prompt Engineer
Tips to becoming a world-class Prompt Engineer
Google Cloud Tech
52 Serverless is simple. Do I need CI/CD?
Serverless is simple. Do I need CI/CD?
Google Cloud Tech
53 Accelerating model deployment with MLOps
Accelerating model deployment with MLOps
Google Cloud Tech
54 How Hawaii's Department of Human Services scaled with CCAI
How Hawaii's Department of Human Services scaled with CCAI
Google Cloud Tech
55 Pricing API on our #Radar
Pricing API on our #Radar
Google Cloud Tech
56 How Recommendations AI for Media can boost customer retention
How Recommendations AI for Media can boost customer retention
Google Cloud Tech
57 Troubleshooting: Node Not Ready Status
Troubleshooting: Node Not Ready Status
Google Cloud Tech
58 One weekend until Cloud Next 2023!
One weekend until Cloud Next 2023!
Google Cloud Tech
59 #GoogleCloudNext starts tomorrow!
#GoogleCloudNext starts tomorrow!
Google Cloud Tech
60 #GoogleCloudNext will be demand!
#GoogleCloudNext will be demand!
Google Cloud Tech

Related Reads

📰
We let Qwen rewrite our scoring algorithm — but only through a clinical-style gate
Improve a scoring algorithm using Qwen3.7-Max through a clinical-style gate with human oversight
Dev.to AI
📰
I Built a 100% Offline AI Research Assistant for Reading Research Papers
Learn how to build a 100% offline AI research assistant for reading research papers using Python and RAG, and discover the benefits of keeping your research data local.
Dev.to AI
📰
Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics
Learn to build production-grade LLM evaluation pipelines to automate testing and catch hallucinations before deployment
Dev.to AI
📰
Why LLMs prioritize high-signal analytical networks and how to secure citations in an AI-driven…
Learn how LLMs prioritize high-signal analytical networks and secure citations in AI-driven research
Medium · AI
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →