Essential Scale-Out Computing by James Cuff

CS50 · Advanced ·📰 AI News & Updates ·11y ago

Key Takeaways

Essential Scale-Out Computing by James Cuff, covering advanced computing concepts and infrastructure, including processors, servers, storage systems, and high-speed networks, with a focus on scale-out computing and its applications in various fields.

Full Transcript

[Music] hi good afternoon everyone um my name is James cuff I'm the assistant Dean for research Computing here at Harvard University and uh today I'm going to talk to you about why scale out Computing is essential so I guess first up who is this guy why am I here why am I talking to you I have a background in uh scientific Computing and uh research Computing stretching back to uh the United Kingdom um the welcome trust serer Institute for uh the human genome and then more recently in the United States uh working at the broad and other uh esteemed uh places of learning such as such as Harvard um I guess what that really means is that I'm a recovering molecular biophysicist um so what right have I got to tell you about uh scaleout computing there's a however 18 years or so I've just seen the most dramatic increases in scale complexity and overall uh efficiency of computing systems when I was doing my PhD at Oxford I was pretty excited with a 200 MHz silicon Graphics machine with 18 GB of storage and a single CPU times have changed if you fast forward now we're spinning over 60,000 CPUs here at Harvard many many other organizations are spinning many more um the important takeaway from this is that scale is now not only inevitable it's happened and it's going to continue to to happen so let's for a moment kind of rewind and talk very quickly about science and my favorite subject the scientific method if you to be a scientist you have to do a few key things if you don't do these things you cannot consider yourself a scientist and you um uh will struggle uh being able to to uh understand your area of discipline so first of all you would formulate your question you generate hypotheses but more importantly you predict your results you have a guess um as to what the results will be and then uh finally you test your hypothesize hypothesis and uh analyze your results and so this scientific method is is extremely important in Computing um Computing of the uh both prediction and being able to test your results are a key part of what we need to do in the scientific method these prediction and testings are the real two cornerstones of the scientific method and each require the most significant advances in modern computation the two pillars of science are that of theory and that of experimentation and more recently Computing is often mentioned as being the third pillar of science so if you students are watching this you have absolutely no pressure um third pillar of science no big deal Computing kind of important uh so glad that this the Computing part of uh computer science uh course 50 all right so enough in the background I want to tell you the plan of what we're going to talk about um today I'm going to go over some history um I'm going to explain why we got here I'm going to talk about some of the history of the Computing here at Harvard um some uh activities around social media uh some green things very passionate about uh all things green um uh storage uh computer storage how chaos affects scaleout systems and distributed systems in particular and then I'm going to touch on some of the scaleout hardware that's required to be able to do Computing at scale and then finally we're going to wrap up with some uh some awesome science so let's take a minute to look at our actual history computing's evolved so since the 60s all the way through to today we've seen basically a change of scope from centralized Computing to decentralized Computing to collaborative and then independent Computing and right back again and let me annotate that a little bit when we first started off with computers we had mainframes they were inordinately expensive devices um everything had to be shared uh the Computing was complex you can see it filled rooms and there were operators and tapes and all sorts of of worry clicky spinny uh devices around the 70s early ' 80s you started to see an um impact of the the deck uh uh VX machines so you're starting to see Computing start to appear back in Laboratories and become closer to you um the uh rise of the personal computer uh certainly in the uh' 80s early early uh uh part of the the the decade really changed Computing um and there was a clue in the title because it was called the personal computer uh which meant belong to you um so uh as the evolution of computing uh continued people realized that their personal computer wasn't really big enough to be able to do uh anything of any Merit or significant Merit in science and so folks started to develop Network device drivers to be able to connect PCS together to be able to build clusters and so this begat uh the era of the baywolf cluster um lenux exploded as a response to proprietary um operating system both costs and complexity and then here we are today where yet again we are faced with uh rooms full of computer equipment uh and the ability to swipe one's credit card and get access to these uh Computing facilities remotely and so you can then see in terms of uh history um impacting How We Do Computing today it's definitely evolved from machine rooms full of computers through some personal Computing all the way right back again to machine rooms full of computers so um this is my first uh my first cluster um so 200000 we built a computer system uh in Europe to uh effectively annotate uh the human genome um there's a lot of Technology listed on the uh the right hand side there that unfortunately is no longer with us it's passed off to the great uh technology in the sky um the uh the machine itself is probably equivalent of a few decent laptops today and that just kind of shows you however we did carefully annotate the human genome and uh and both protected it uh with this particular paper in in nature from um uh the the concerns of of the data being public or private um so this is awesome right so we've got a human genome we've done Computing I'm feeling very pleased with myself I rolled up to Harvard in 2006 feeling a lot less pleased with myself um this is what inherited this is a departmental um mail and file server you can see here there's a little bit of tape that's used to hold the system together this is a license and print server I'm pretty sure there may be passwords on some of these Post-it notes um not awesome pretty far from awesome and so I realized this little uh chart that I showed you at the beginning from sharing to ownership uh back to sharing that we needed to change the game and so we changed the game by providing incentives and so human beings as this little Wikipedia article says here are purposeful creatures and the study of incentive structures is Central to the study of economic activity so we started to incentivize our faculty and our researchers and so we incentivized them with a really big computer system so in 2008 we built a 496 processor machine um 10 racks couple hundred kilowatts of power what I think is interesting is it doesn't matter where you are in the in the cycle this is this this amount of power in compute um the power is the constant it was 200 Kow when we were Building Systems in Europe it's 200 kilowatts um uh in 2008 and that seems to be the quantor of small uh University based uh Computing systems so Howard today fast forward I'm no longer Sad Panda quite a happy panda we have uh 60 odd, thousand load balanced uh CPUs and they're climbing dramatically we have 15 pedabytes of storage also climbing again this 20 kilowatt increment we seem to be adding that every six or so months lots and lots of virtual machines um and more importantly about 1.8 megawatts of research Computing equipment and I'm going to come back to this later on as to why I now no longer necessarily count how much CPU we have but how big is the electricity bill 20 other so uh dedicated research Computing staff and more importantly we're starting to grow our gpgpus um I was staggered uh how um how much of this uh is is being added on a day-to-day basis all right so history lesson over right so how do we get there from here let's look at some uh modern scaleout compute examples um I'm a little bit obsessed with the size and scale of social media um there are a number of extremely successful um large scale uh Computing organizations now on the planet providing support and services to to to to us all so that's the disclaimer and I want to start with the number of ounces in an Instagram it's not actually a leader to a joke it's um uh it's not even that funny actually come to think of it but anyway we're going to look at anounces in Instagram and we're going to start with my bee and a flower I was at Sherburn Village and I took a little picture of a bee sitting on a hour and then I started to think about what does this actually mean and I I I took this picture off my phone and counted how many bytes are in it and it's about 256 kilobytes which when I started would basically fill a 5 and a/4 inch floppy and started to think well that's cool and I started to look and do some research on the network and I found out that Instagram has 200 million ma was actually that sure what a ma was and a ma down here is a monthly active user so 200 million miles pretty cool 20 billion photographs it's quite a lot of photographs 60 million new photographs each and every day coming out at about 002 gig per photo that's about 5 pedabytes of disc just right there and that's really not the central part of what we're going to talk about that is small potatoes or what we used to say in England tiny so let's look at the real elephant in the room unique faces again let's measure in this new Quant called a Mau uh Facebook itself has 1.3 billion ma WhatsApp which I hadn't even heard of until recently it's a some sort of messaging service is 500 million miles um Instagram which we just talked about 200 million miles and messenger which is another messaging service uh is also 200 million miles so so To Tot all that up you know it's about 2.2 billion total users clearly there's some overlap um but that's equivalent to a third of the planet and they send something on the region of 12 billion messages a day and again there's only 7 billion people on the planet not everyone has a smartphone so this is this is insane numbers and I'm going to argue that it's not even about the storage or the compute and to quote the song it's All About That graph here's our lovely Megan Trainer down here singing about all the bass not she also has quite a bit of bass herself 27 well 280 million people have seen this young lady singing her song so my argument is that it's all about the graph so we took some open source software and and started to look at a graph and this is LinkedIn so this is um Facebook for old people um and so this is my LinkedIn Gra I have 1 12200 or so nodes soall friends and uh here's me at the top and here's all of the inter connections now think back to the Instagram Story Each one of these is not just a photo it has a whole plethora of connections between this particular individual and many others this is either this Central piece is either a bug in the graph drawing algorithm or this may be David ma not sure yet so anyway so you can redraw the graphs in all sorts of ways h geffy github.io is uh where you can pull that software from it's really cool for being able to organize uh communities you can see here this is Harvard and various other places that I've worked because this is my um uh work-related data so just think about the complexity of the graph and all of the data that you pull along with so meanwhile back at friendface right we looked at the Instagram data that was of the order of five pedabytes no big deal still quite a lot of data but no big deal in the greater schem of things found this article on the uh the old internet uh scaling the Facebook data warehouse to 300 pedabytes that's a whole different game changer now when you're starting to think of data and the graph and what you bring along with and there Hive data is growing of the order of 600 terabytes a day now you know well then right I mean 600 terabytes a day 300 babyes they're also now starting to get very concerned about how to keep this stuff and to make sure that this data stays around and uh this gentleman here uh Jay Perth is looking at how to store an exabyte of data just for those of you who watching along at home um an exobyte 10 to the 18 it's got its own Wikipedia page it's that big of a number um that is the size and scale of what we're looking at to be able to store data and these guys aren't marking around they're storing that amount of data so one of the clues that they're looking at here is uh data centers for So-Cal Cold Storage which brings me to being green and here is kmit he and I agree it's extremely difficult to be green um but we give it our best try kmit can't help it he has to be green all the time can't take his greenness off at all so green Concepts few kind of Core Concepts of um of greenness when it relates to computing the one that is the most important is the longevity of the project product if your product has a short lifetime you cannot by definition be green the energy taken to manufacture a dish drive a motherboard a computer system a tablet whatever it may be the longevity of your systems are are a key part of how green you can be the important part as all of you are building software or algorithms algorithms Posh word for software right so your algorithm design is absolutely critical in terms of how you are going to be able to um make quick and accurate computations uh to use the least amount of energy possible and I'll get to this in a little bit data center design you've seen that we already have thousands upon thousands of machines sitting quietly in small D corners of the world uh Computing um resource allocation how to get to the compute to the storage to the through the network operating systems are a key part of this and a lot of virtualization to be able to pack more and more compute into a smaller space and I'll give you a small example from uh research Computing we needed more ping more power and more pipe we needed more bigger better faster computers I needed to use less juice and we couldn't work out how how to do this I don't know if the hashtag go west has probably been used by by the Kardashians but anyway go west so um and we did we picked up our operation and we moved it out to uh Western Massachusetts in a a small miltown called Holo uh it's just north of chicke and Springfield um we um we did this we did this for a couple of reasons um the the main one was that we had a very very large Dam and this very large dam is able to put out 30 plus Mega wats of energy and it was underutilized at the time more importantly we also had a very complicated Network that was already in place if you look at where the network goes in the United States it follows all the train tracks uh this particular piece of network was owned by our colleagues and friends at Massachusetts Institute of Technology and it was basically built all the way out to Route 90 so we had a large river tick uh route 90 tick we had a short path of 100 miles and a long path of about 1,000 miles we did have to do a very large Network splash as you can see here to uh basically put a loop in to be able to connect to Holo but we had all of the requisite infrastructure ping power pipe life was good and again Big Dam so we built basically the Massachusetts green high performance Computing Center this was a labor of love through five universities MIT Harvard UMass North Eastern and Buu 5 megawatt day one connected load we did all sorts of cleverness with airide economizers to keep things green um and we built out 640 odd racks um uh for dedicated for uh research Computing it was um an old uh Brownfield site so we had uh some Reclamation and some tidy up and some cleanup of the site and then we started to build uh the facility and boom lovely facility with um uh the ability to run sandbox Computing to have conferences and seminars and also a massive uh data center floor here is uh my good self I'm obviously wearing the same jacket I maybe only have one jacket but there me and John goodu he's the executive director of the uh of the center standing in the in the machine room floors which as you can see is pretty dramatic and it goes back a long a long long way I often play games driving from um Boston out to um to Holio pretending that I'm a tcpip packet and uh I wonder I do worry about my latency uh driving around in my my my my my car all right so so that's the that's that's the green piece so let's just take a minute and think about Stacks so we're trying very carefully to build data centers efficiently Computing efficiently make good selections for the Computing equipment um and uh and deliver more importantly our application be it um a messaging service or um a scientific application so here are the stacks right so physical layer all the way up through application hoping that this is going to be a good part of your course the OSI um uh seven layer model is basically you will live eat and breathe this uh throughout throughout your Computing uh careers um this whole concept of physical infrastructure wires cables data centers links and this is just describing the network and up here is well obviously this is an old slide because this should say HTTP because nobody cares about simple male transport protocols anymore it's all happening in the in the HTTP space so that's one level of Stack here's another set of stacks where you have a server host a hypervisor a guest binary and library and then your application or in this case a device driver a Linux kernel native C Java virtual machine Java API then Java applications and so on and so forth this is a description of a virtual machine holy Stacks Batman think about this in terms of how much compute you need to get from what's Happening Here all the way up to the top of this stack to then be able to do your actual delivery of the application and if you kind of rewind and start to think about what it takes to provide a floating Point operation your floating Point operation is a sum of the sockets the number of cores in the socket a clock which is how fast can um uh the the uh uh the clock turn over 4 GHz 2 GHz and then the number of operations you can do in a given Hertz so most microprocessors today do between four and six flops per clock cycle and so a single Core 2 and 1/2 gig clock has a theoretical performance of about a mega flop all right give or take but as with everything we have choices so an Intel to Core 2 N halum sandybridge Haswell AMD take your choices Intel atom all of these processor architectures all have a slightly different way of being able to add two numbers together which is uh basically their purpose in life um must be tough there's millions of them sitting in data centers now though so but anyway so um flops for what this is the big thing so if I want to get more of this uh to get through this stack faster I've got to work on how many floating Point operations a second I can do and then given them what unfortunately folks have thought about this so there's a a large contest every year to see who can build the fastest computer that can diagonalize a matrix it's called the top 500 uh they pick the top and the best 500 computers on the planet that could diagonalize matrixes and you get some amazing results um a lot of those machines are between 10 and 20 megaw they can diagonalize major C's inordinately quickly um they don't necessarily diagonalize them as efficiently per watt so uh there was this big push to look at what a green 500 list would look like and here is the list from June there should be a new one very shortly and it calls out uh I'll take the top of this uh this particular list as two specific machines um um one from the Tokyo Institute of Technology and one from Cambridge University in the United Kingdom and these have pretty staggering megaflops per watt ratios this one's 4389 and the next one down is 3,631 and I'll explain the difference between these to um in the in the next slide but these are these are moderately sized test clusters these aren't um these are just um 34 Kow or 52 kows there are some larger ones here this this particular one at the Swiss National superc Computing Center the take-home message for this is that we're trying to find computers that can operate efficiently and so let's look at this this this this top one um cutely called the KFC and um uh a little bit of advertising here this particular food company has nothing to do with this at all it's the fact that this particular system is uh soaked in a very uh clever um oil based compound and so they got the chicken fryer um uh monik um when they first started to build these types of systems but basically what they've taken here is a number of blades put them in this um sophisticated um mineral oil um and uh then worked out how to get all the networking in and out of it then not only that I've put it outside so that it can exploit outside air cooling which pretty impressive so you have to do all of this Shenanigans to be able to get this amount of compute delivered for small wattage and you can see this this is the shape of of where things are are heading the challenge is that Reg regular air cooling um is the economy of scale and is driving uh a lot of the uh the the development of of of both regular Computing and high performance Computing so this is pretty disruptive I think this is fascinating um it's a bit messy when you trying to swap the dish drives but it's a really cool idea so uh not only that uh there's a whole bunch of work being built around what we're calling the open compute project and so um uh more about that a little bit later but the industry is starting to realize that the flops per watt is becoming important and you as as as folks uh uh here um as you design your algorithms and you design your code you should be aware that your code can have a knock on effect when Mark was sitting here in his dorm room you know writing you know Facebook 1.0 I'm pretty sure you know he had a view that it was going to be huge but how huge it would be on the environment is is a big dealio and so all of you all could come up with uh algorithms that could be the next challenging thing for folks like me trying to run systems so let's just think about real world power limits this paper by landow this is not a new thing 1961 this was published in the IBM Journal this is a the canonical irreversible and heat generation in the Computing process and so he argued that machines um inevitably perform logistic functions that don't have single valued inverse so so the the whole part of this is that back in this 60s folks knew that this was going to be a problem and so the L limit said 25° C sort of canonical room temperature um the limit represents 0.1 electron volts but theoretically this is the theory computer memory operating at this limit could be changed at one billion bits a second don't know about you but not come across many one billion bits a second data rate exchanges and the argument there was that only 2.8 trillions of a watt of power uh ought to ever be expended all right real world example this is my electric bill I'm 65% of that lovely data center I showed you in this particular time this is back in June um uh last year uh I've taken an older version so that uh we we can sort of anonymize it a little bit I was spending $45,000 a month for uh for energy there so the reason being there is that we have um uh over 50,000 processes in that in in in that in that room so could you imagine your own residential electricity bill uh being that high but it was for 199 million W hours over a month so the question I POS it is can you imagine Mr Zuckerberg's electric bill um mine is pretty big and I struggle um and I'm not alone in this there's a lot of people uh with big data centers and so I guess full disclosure my Facebook friends are a little bit odd so I my Facebook friend it's the primeville data center which is one of uh Facebook's largest uh newest uh lowest energy Data Center and and they they post things that are fun to me things like power utilization Effectiveness as in how you how effective is the data center versus how much energy you're putting into it how much water are they using what's the humidity and temperature and they have these lovely lovely plots I think this is an awesome Facebook page but I guess I'm a little bit weird anyway so one more power thing um research Computing that I do is significantly different to what uh Facebook and Yahoo and Google and other on demand fully um uh always available uh services so I have the advantage that when is saw New England and is New England helps set the energy rates for um uh for for the region um and it says it's extending a request to Consumers for voluntary to voluntary conserve High uh energy because of the high heat humidity and this was back in um on the 18th of July and so I happly Tweet back hey is saw New England green Harvard we're doing our part over here in research Computing and this is because we're doing science and as much as people say science never sleeps science can wait so we are able to qus our systems take advantage of great rates on our energy um uh Bill uh and and help uh the entire New England region by shedding um uh many uh megawatts of uh of load so that's a that's a unique thing that differs about scientific Computing data data B data centers uh and uh those that are um uh in full production 247 so let's just take another uh gear here so I want to discuss chaos a little bit and I want to put it in the uh the oses of of storage so for those that kind of were struggling getting their head around what pedabytes of storage look like this is this is an example and this is the sort of stuff I deal with all all the time each one of these little fellas is a four terabyte hard drive so you can kind of count them up we're getting now between uh one to one and a half pedabytes in a standard industry rack and we have rooms and rooms as you saw in that earlier picture with John and I full of these uh these these ranks of equipment so it's becoming very very easy to build massive storage arrays it's mostly easy inside of Unix to kind of count up how things are going so this is uh counting how many Mount points have I got there so that's 423 intercept points and then if I run some sketchy orc I can add up that in this particular system there was 7.3 pedabytes of available storage so that's a lot of stuff and storage is really hard and yet for some reason um this is an industry Trend whenever I uh talk to our researchers and our faculty and say hey um I can run storage for you uh unfortunately I have to recover the cost of the storage um I get I get this business and people reference new egg or they reference Staples or how much they can buy a single terabyte dish drive for so this you'll note here that there's a there's a clue there's one dish drive here and if we go back I have many not only do I have many I have sophisticated interconnects to be able to stitch these things together um so the risk associated with these large storage arrays is not insignificant in fact we took to the internet and we wrote a little a little story about a well meaning uh mild manner director of research Computing happen to have a strange English accent uh trying to explain to a researcher what uh the ners score backup folder actually meant uh it's quite a long little story um good four minutes of um of Discovery and and note I have an awful lot less bass than the lady that sings about all the bass we we're quite a few counts lower but anyway this is an important thing to think about in terms of what could go wrong so if I get a dish drive and I throw it in the Unix machine and I start writing things to it there's a magnet there's a drive head there's a the ostensibly a one or a zero being written down onto that device Motors spinny twirly things always break think about the things that break it's always spinny twirly things uh printers dish drives Motor Vehicles Etc anything that moves is likely to break so you need motors you need drive firway you need sasada controllers wires firmware on the sasada controllers lowlevel blocks pick your um uh storage controller file system code whichever one it may be um how you Stitch things together um your virtual memory manager Pages dram Fetch and stores then you're getting up the stack which is kind of down the list on this one algorithms users and if you multiply this up I don't know how many there's a lot of places where stuff can go sideways right I mean that's an example of bad math but it's kind of fun to think of how many ways things could go wrong just for a dish drive and we we're we're already at 300 pedabytes right so imagine the number of dish drives you need in 300 pedabytes uh that can go wrong not only that so that's storage and that alludes to um the person I'd like to see enter stage left which is the chaos monkey right so at a certain point um it gets even bigger than just the dis Drive problem and so uh These Fine ladies and gentlemen that run a uh streaming video service realize that their computers were also huge and also very complicated and also providing service to an awful lot of uh a lot of people they've got 37 million members and this slide's maybe a year or so old um thousands of devices you know billions of hours of of video they log billions of events a day and you can see you know most people watch the Telly later on in the evening and and it far outweighs everything and so they wanted to be able to make sure that this service was up and reliable and working for them and so they came up with this thing called chaos monkey it's a piece of software which when you think about talking about the title of this whole um uh presentation scale Out means you should test this stuff it's not good just having a million machines so the nice thing about this is chaos monkey is a service which identifies groups of systems and randomly terminates one of the systems in a group aome so I don't know how about you but if I've ever built a system that relies on other systems talking to each other and you take one of them out the likelihood of the entire thing working diminishes rapidly and so uh this piece of software runs around Netflix's infrastructure luckily it says it runs only in business hours with the intent that Engineers will be alert and able to respond so yeah these are the types of things that we're now having to do to perturb our Computing environments uh to introduce chaos and to introduce complexity so who in their right mind would willingly choose to work with a chaos monkey hang on he seems to be pointing at me well I guess I should um cute um but the problem is is you don't get the choice the chaos monkey as you can see chooses you and this is the problem with Computing at scale is that you can't you can't avoid this it's it's it's an inevitability of of complexity and of scale um and of of of of of our Evolution uh in some ways of of computing uh expertise and remember this is one thing to remember chaos monkeys love snowflakes love snowflakes a snowflake we've explained the chaos monkey but a snowflake is a server that is unique and special and delicate and and individual and will never be reproduced um we often find snow slake servers in our environment and we always try and melt snowflake servers but if you find a server in your environment that is critical to the longevity of your organization and it melts you can't put it back together again so chaos monkey's job was to go and terminate in es if if if the chaos monkey melts the snowflake you're over you're done all right I want to talk about some Hardware uh that we're seeing in terms of sort of scaleout activities too and some unique things that are in and around the science activity we are now starting to see remember this unit of issue this rack so this is a rack of gpgpus so general purpose Graphics processing units um we have these located in our data center 100 or so Miles Away um this particular rack um is about 96 teraflops of single Precision math able to deliver out the back of it um and we have order um uh 130 odd uh cards in an instance um uh uh that uh that that we are uh multiple racks of this instance so this is interesting in the sense that um the general purpose Graphics processes are able to do mathematics incredibly quickly for very low amounts of energy so large uptick in the big uh scientific Computing areas looking at uh Graphics processing units in a big way so I ran uh some M Collective uh through our uh puppet infrastructure yesterday um very excited about this we we're short of just short of a pedop flop of single Precision just be be clear here this little multiplier is 3.95 um double Precision math would be about 1.2 but my Twitter feed looked way better if I said we had almost a pedop flop of um single prision gpgpus um but it's it's it's getting there it's getting to be very very impressive and why are we doing this because quantum chemistry among other things but we're starting to design some new photov voltaics and so um Alan asuz who's a professor in chemistry my partner in crime for the last few years we've been pushing the envelope on Computing uh and the gpgpu is ideal uh technology to be able to do an awful lot of complicated math uh very very quickly so with scale comes new challenges right so huge scale you have to be careful how you wire this stuff and we have certain levels of obsessive compulsive disorder these pictures probably drive a lot of people nuts and what and cabinets that aren't wired particularly well uh drive on network and facilities engineers nuts plus there's also air flow issues you have to contain so these are things that I would never have thought of with with scale comes more complexity this is a new type of file system it's awesome it's a petabyte it can store 1.1 billion files it can read and write at 13 uh gigabytes and 20 Gigabytes a second gigabytes a second um so it can unload you know terabytes in no time at all um and it's highly available and it's got amazing lookup rates you know 220,000 lookups a second and then many different people building these kind of systems and you can see it here um graphically uh this is one of our file systems that's under load quite happily reading at just short of 22 uh gigabytes a second so that's cool so complexity so with complexity and scale comes more complexity right this is one of our many many Network diagrams um where you have many different chassis all supporting up into a main core switch connected to storage connecting to uh low latency interconnects and then all of this side of the house is just all of the management that you need to be able to um address these systems from a remote location so scale has a lot of complexity with it all right change gear again let's go back and have a little spot of science so remember research Computing I'm this little shim little pink shim between the faculty and all of their algorithms and all of the cool science and all of this power and Cooling and data center floor and networking and big computers and service desks and health desks and so on and so forth and so we're just this little shim uh between that what we've started to see um is that the the world's been able to build these large uh data centers and be able to build these large computers we've gotten pretty good at it what we're not very good at is this little shim between um the research and uh the bare metal and the technology and it's hard um and so we've been able to uh hire folks that live in this world and more recently we uh spoke to the National Science Foundation and said you know this scaleout stuff is great but we we can't get our scientists onto these big complicated machines and so there have been a number of different programs where we really um were mostly concerned about trying to see if we could transform the campus infrastructure there are a lot of programs around um National centers and so ourselves our friends at Clemson um University of Wisconsin Madison Southern California Utah and Hawaii kind of got together to look at this problem and this little graph here is the long tale of science right so this is it doesn't matter what's on this axis this axis is actually number of jobs going through the cluster so there's 350,000 over you know whatever time period these are our usual suspects along the bottom here in fact there's Alan aspik who we just talking about tons and tons of compute really affect knows what he's doing here's another lab that I'll talk about in in a moment John kovak lab they've got it they're good they're happy they're Computing great Sciences getting done and then as you kind of come down here there are other groups that aren't running many jobs and and and why is that is it because the computing's too hard is it because they don't know how to we don't know right because we don't we've never gotone looked and so that's what this project is all about is locally within each of these regions to look to um Avenues where we can engage with the faculty and researchers actually in the bottom end of the tail and understand what they're doing so that's something that we're actually passionate about and that's something that science won't continue to move forward until we solve some of these um these edge cases other bits of science that's going on everyone's seen the large hron collider awesome right this all this stuff all you know ran out at um at Holio we built the uh the very first science that happened in Holo was the collaboration between ourselves and Boston University so it's really really cool this is a fun piece of science for scale um this is a digital access to a sky Century at Harvard basically is a plate archive if you go down Oxford um uh Garden Street sorry um you will find one of the observatory buildings is basically full of about half a million plates and these are pictures of the sky at night um over a hundred years and so there there's a whole rig set up here to digitize those plates take pictures of them register them put them on a computer and that's a paby and a half just right there one one small project these are other projects this this pan Stars project is doing a um full wide panoramic survey looking for uh near Earth asteroids and uh transient celestial events as a as a molecular biophysicist I love the word transient Celestial event I'm not quite sure what it is but anyway we're looking for them and uh we're generating the terabytes a night out out of those those those telescopes and that's not really a bandwidth problem that's like a FedEx problem so you put the put the storage on the on the van and you send it wherever it is um bicep is really interesting so background Imaging of cosmic extra Galactic polarization when I first started working at Harvard seven or so eight years ago um I remember working on this uh uh this this project and never really kind of it didn't really sink home as to why polarized light from the cosmic m microwave background would be important until this happened and this was John kovak who I talked to before you know using Millions upon millions of CPU hours um uh in in in our facility and others to basically stare into the Insight of the universe's first moments after the big bang and trying to understand Einstein's general theory of relativity it's mindblowing that our computers are helping us unravel and stare into the very Origins of like why we're here so you talk about scale this is this is some serious scale the other thing of scale is that that particular project hit these guys and this is the response curve for bicep rc. fast this was our little server and you can see here life was good until about here which was when the announcement came out and you have got literally seconds to respond to the scaling event which corresponds to this little dot here um which ended up shifting four or so terabytes of data through the web server that day um pretty hairy um and so these are the types of things that can happen to you in your infrastructure if you do not design uh for scale um we uh had a bit of a scramble that day to be able to span out enough web servers to keep the uh the site up and running and and we were successful this is a little email that's kind of cute this is um uh ma to Mark vburger and and L hunis who's a a faculty member here at at Harvard more about Mark later but I think this this one sort of sums up kind of where the Computing is uh is in in in research Computing so hey team since last Tuesday you guys racked up over 20 8% of the new cluster which combined is over 78 years of CPU in just three days and I said it's still only just Friday morning so this is pretty awesome happy Friday then I give them the the the the data point and so that was uh that was kind of kind of interesting so remember about about Mark He he'll come back into the picture in a little bit so scaleout Computing is everywhere we're even helping folks look at how the NBA uh functions and where people are throwing balls from I don't really understand this game too too well but um seeming it's a big deal as hoops and balls and and money and so um our database we um we we built a little you know 500 way parallel processor cluster a couple of ter of ram to be able to build this um uh for um for Kirk and uh and his team and and they're doing uh Computing at a whole L other way um this is a project we're involved with that's absolutely fascinating around neuroplasticity connectomics and genomic imprinting and three very heavy hiding areas of uh of of research um that uh that we we we fight with on a on a day-to-day basis and the idea that um our our our our brains are under plastic stress when we are um when we are young and much of our adult behavior is sculpted U by experience in infancy and so this is this is a big dealio and so this is work that's funded by the National Institutes of mental health um and we are trying to basically through a lot of large data and Big Data analysis kind of peer into um our our our human brain through a variety of different uh techniques so I want to just stop and and and kind of just pause for a a little moment the the challenge with remote data centers is it's far away it can't possibly work I need my data close by I need to do my research um in my lab and so I I I kind of took an example of um uh a functional magnetic resonance imaging uh data set from a our data center in in western Mass and connected it to my desktop in in Cambridge and I'll I'll play this little little video hopefully it kind of work okay so this is me kind of going through and checking my gpus are working and I'm checking that vncs up and this is a clever VNC this is a VNC with 3D pieces and so as you can see shortly this is me spinning this brain around and I'm trying to kind of get it oriented and then I can move through many different slices of uh of of MRI data um and this the only thing that's different about this is it's coming over the wire from Western Mass to my desktop and it's rendering faster than my desktop um because I don't have uh a $4,000 graphics card in my desktop um which we did have out think West M of course I'm trying to be clever I'm running GLX gears in the background whilst doing all this to make sure that I can stress the graphics card and that it all kind of works and and all the rest of it but the important thing is is this is 100 miles away and you can see from this that there's no obvious latency and you know things are uh are holding together uh fairly well um and so that in of itself is an example and some insight into how Computing and scaleout Computing is going to happen we're all working on thinner and thinner devices our uh our our use of tablets is increasing so therefore my carbon footprint is basically moving from what used to to do that would have been a huge machine under my uh under my desk to what is now a facility could be anywhere um could be anywhere at all and yet is still able to bring back um high performance uh Graphics to uh to my desktop so so um getting near the end remember Mark well smart smart lad is Mark he he decided that he was going to build a realistic virtual universe that's quite a a project when you think you've got to pitch this I'm going to use a computer and I'm going to model the 12 um million years after the big bang to represent a day and then I'm going to do 13.8 billion years of cosmic Evolution all right um this actually needed a computer that was bigger than our computer and it spilled over onto the national resources to our friends down in Texas and uh to the National facilities this was a lot of compute um but we did a lot of the simulation locally to make sure that the uh the software worked and the systems worked and it's days like this when you realize that you're supporting science at this level of scale that people can now say things like I'm going to model a universe and this is his first model and this is this team's first model um there are many other folks that are going to come behind Mark who are going to want to model with higher resolution with more specificity with more accuracy and so in in the last couple of minutes I just want to show you this uh this video of of Mark and laz's that uh to me again as a life scientist is is is is kind of kind of cute so um so this at the bottom here to orient you this is telling you the time since the Big Bang so we're at about7 billion years and this is showing the current update so you're seeing at the moment uh dark matter and the evolution of um of the fine structure and early structures um in our known universe and the point with this is that this is all done inside the computer you know this is a set of parameters and a set of physics and a set of mathematics and a set of models that are carefully selected and then um uh carefully connected to each other to be able to model the interactions and you can see some starts of some gous explosions here and gas temperature is changing and you can start to see the structure of the the visible Universe uh change um and the important part with this is each little tiny tiny tiny dot is a piece of physics and it has a set of mathematics around it informing its friend and its neighbor so from a scaling perspective these computers have to all work in concert and talk to each other efficiently so they can't be too chatty they have to store their results and they have to um uh continue to inform all of their friends and you can see now this model is getting more and more complicated there's more and more stuff going on there's more and more material flying around and this is you know what the early Cosmos would have looked like it was a pretty hairy place there explosions all over the place powerful um collisions um and you know formation of of heavy metals and elements and and these these big clouds smashing into each other with extreme force so now we're you know 9.6 billion years from uh from this initial exposure you're starting to see things that kind of calm down a little bit just a little bit because the energy is now starting to relax and so the mathematical models are have got that in place and you're starting to see coalesence of different elements and starting to see this thing kind of come together and slowly cool and it's starting to look a little bit more like the night sky a little bit and it's Qing now 13.2 billion years and we're kind of done and then what they did was that they took this model and then looked at uh the visible universe and uh and basically then we able to take that and overlay it with you know what you can see and the Fidelity is is is staggering as to how accurate the computer models are of course the astrophysicists and the research groups need even better Fidelity and even uh higher resolution but if you think about what I've been talking to you uh today through this little Voyage Through um both storage and structures of networking and and stacks um you the the the important thing is you know is scaleout Computing essential that was my original hypothesis back to our scientific method um I hope that I at the early part of this I would predict that I'd be able to explain uh to you about scaleout Computing um and we kind of tested some of those uh hypotheses as we went through this uh this conversation and I'm just going to say scaleout Computing is essential oh yes very much yes and so when you're thinking about your codes when you're doing your cs50 final projects when you're thinking about uh your legacy to humanity and the resources that we need to be able to run these computer systems think very carefully about the flops per watt and think about the chaos monkey think about your snowflakes don't do one-offs reuse libraries build reusable codes all of the things that the tutors have been uh teaching you in this uh in this in this class these are fundamental uh aspects they're not just lip service these are are are real things and if any of you want to follow up with me I I am obsessive with the with the Twitter thing I've got to somehow give that up uh but a lot of the background information is uh on our research computer website at rc. fast. harvard.edu I try and keep a Blog up to date with modern uh uh uh Technologies and um uh how we do distributed computing and so forth and then um our uh uh staff are always available through uh aib bot.org and aibot is our little helper um he uh often has little contests on his website too where you can try and spot him around campus he's the friendly little face of um of research Computing and I'll kind of um uh wrap up there and uh thank you all for your time and um hope you'll remember that um that scaleout Computing is it's a real thing um and there are a lot of people who've got a lot of prior art will be able to help you and uh all of the best of luck with your uh future endeavors in making sure that our Computing both scales uh is High performant and um helps Humanity more than anything else so thank you for your time

Original Description

Each day you interact with thousands upon thousands of processors, servers, storage systems and high-speed networks. You don't see them, and you don't physically touch them, but they are there, making everything happen behind the scenes. Everything is powered by advanced computing, from your morning news, movie and video streams, phone conversations, currency, financial markets, pharmaceuticals, navigation, traffic, weather, email and of course all of our social media updates. Each of us consumes vast amounts of data and computation on a daily basis. We also continue to push the boundaries of our science and discovery. Using ever more complex computer models to peer into the darkness of space or to understanding the genetic basis as to why were are human. All of this needs computing for it to work correctly, and it also needs advanced infrastructure and distributed computing architectures to work quickly. James Cuff is the Assistant Dean for Research Computing here at Harvard. His group runs more than sixty thousand high performance computing processors and more than fourteen petabytes of storage for science. On a global scale, this system is tiny. However, he will show you real world examples of the advances in computation science, physical infrastructure and distributed computing systems we are using each day, whether you are a particle physicist trying to reverse engineer the very fabric of the universe – or maybe you are just updating your selfie... So what will you learn from this seminar? You are all designing software for your final project. Facebook for example, was originally designed as a small single server PHP application. In order to make it scale to today’s hundreds of thousands of servers and billions of users took years. James will explain how both datacenter and systems architectures that now surpass electrical power usage of 10-20 megawatts – (enough to power more than 20,000 houses, nearly half of the City of Cambridge) enable today’s applications
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from CS50 · CS50 · 27 of 60

1 Hello, World: Hadi Partovi
Hello, World: Hadi Partovi
CS50
2 Content Distribution and Archival in a Digital Age
Content Distribution and Archival in a Digital Age
CS50
3 CS50 2014 - Week 1
CS50 2014 - Week 1
CS50
4 CS50 2014 - Week 3
CS50 2014 - Week 3
CS50
5 CS50 2014 - Week 0, continued
CS50 2014 - Week 0, continued
CS50
6 CS50 2014 - Week 4
CS50 2014 - Week 4
CS50
7 Week 3, continued
Week 3, continued
CS50
8 Quiz 0 Review
Quiz 0 Review
CS50
9 CS50 2014 - Week 3, continued
CS50 2014 - Week 3, continued
CS50
10 CS50 2014 - Week 7
CS50 2014 - Week 7
CS50
11 CS50 2014 - Week 7, continued
CS50 2014 - Week 7, continued
CS50
12 Breaking Through The (Google) Glass Ceiling by Christopher Bartholomew
Breaking Through The (Google) Glass Ceiling by Christopher Bartholomew
CS50
13 Introduction to Amazon Web Services by Leo Zhadanovsky
Introduction to Amazon Web Services by Leo Zhadanovsky
CS50
14 CS50 2014 - Week 9
CS50 2014 - Week 9
CS50
15 How to Build Innovative Technologies by Abby Fichtner
How to Build Innovative Technologies by Abby Fichtner
CS50
16 Light Your World (with Hue Bulbs) by Dan Bradley
Light Your World (with Hue Bulbs) by Dan Bradley
CS50
17 Building Dynamic Web Apps with Laravel by Eric Ouyang
Building Dynamic Web Apps with Laravel by Eric Ouyang
CS50
18 CS50 2014 - CS50 Lecture by Steve Ballmer
CS50 2014 - CS50 Lecture by Steve Ballmer
CS50
19 CS50 2014 - Week 10
CS50 2014 - Week 10
CS50
20 This is CS50 with Steve Ballmer?
This is CS50 with Steve Ballmer?
CS50
21 Meteor: a better way to build apps by Roger Zurawicki
Meteor: a better way to build apps by Roger Zurawicki
CS50
22 Data Analysis in R by Dustin Tran
Data Analysis in R by Dustin Tran
CS50
23 Data Visualization and D3 by David Chouinard
Data Visualization and D3 by David Chouinard
CS50
24 CS50 2014 - Week 6
CS50 2014 - Week 6
CS50
25 Build Tomorrow's Library by Jeffrey Licht
Build Tomorrow's Library by Jeffrey Licht
CS50
26 CS50 2014 - Week 9, continued
CS50 2014 - Week 9, continued
CS50
Essential Scale-Out Computing by James Cuff
Essential Scale-Out Computing by James Cuff
CS50
28 iOS App Development with Swift by Dan Armendariz
iOS App Development with Swift by Dan Armendariz
CS50
29 Sam Clark Leads Yale Students on Tour to CS50 at Harvard
Sam Clark Leads Yale Students on Tour to CS50 at Harvard
CS50
30 3D Modeling and Manufacture by Ansel Duff
3D Modeling and Manufacture by Ansel Duff
CS50
31 CS50 2014 - Week 5, continued
CS50 2014 - Week 5, continued
CS50
32 hello, world
hello, world
CS50
33 CS50 2014 - Deep Thoughts - Hash Table
CS50 2014 - Deep Thoughts - Hash Table
CS50
34 CS50 2014 - Deep Thoughts - Binary Tree
CS50 2014 - Deep Thoughts - Binary Tree
CS50
35 CS50 2014 - Deep Thoughts - Scratch
CS50 2014 - Deep Thoughts - Scratch
CS50
36 CS50 2014 - Deep Thoughts - MySQL
CS50 2014 - Deep Thoughts - MySQL
CS50
37 LaunchCode Visits CS50
LaunchCode Visits CS50
CS50
38 CS50 Live, Episode 100
CS50 Live, Episode 100
CS50
39 CS50 Field Trip to Google
CS50 Field Trip to Google
CS50
40 This is CS50 AP
This is CS50 AP
CS50
41 Week 4: Monday - CS50 2011 - Harvard University
Week 4: Monday - CS50 2011 - Harvard University
CS50
42 Week 2: Wednesday - CS50 2011 - Harvard University
Week 2: Wednesday - CS50 2011 - Harvard University
CS50
43 Week 1: Wednesday - CS50 2011 - Harvard University
Week 1: Wednesday - CS50 2011 - Harvard University
CS50
44 Week 11: Monday - CS50 2011 - Harvard University
Week 11: Monday - CS50 2011 - Harvard University
CS50
45 Week 3: Wednesday - CS50 2011 - Harvard University
Week 3: Wednesday - CS50 2011 - Harvard University
CS50
46 Week 12: Monday - CS50 2011 - Harvard University
Week 12: Monday - CS50 2011 - Harvard University
CS50
47 Week 1: Friday - CS50 2011 - Harvard University
Week 1: Friday - CS50 2011 - Harvard University
CS50
48 Week 3: Monday - CS50 2011 - Harvard University
Week 3: Monday - CS50 2011 - Harvard University
CS50
49 Week 10: Wednesday - CS50 2011 - Harvard University
Week 10: Wednesday - CS50 2011 - Harvard University
CS50
50 Week 2: Monday - CS50 2011 - Harvard University
Week 2: Monday - CS50 2011 - Harvard University
CS50
51 Week 9: Monday - CS50 2011 - Harvard University
Week 9: Monday - CS50 2011 - Harvard University
CS50
52 Week 7: Monday - CS50 2011 - Harvard University
Week 7: Monday - CS50 2011 - Harvard University
CS50
53 Week 5: Monday - CS50 2011 - Harvard University
Week 5: Monday - CS50 2011 - Harvard University
CS50
54 Week 5: Wednesday - CS50 2011 - Harvard University
Week 5: Wednesday - CS50 2011 - Harvard University
CS50
55 Week 7: Wednesday - CS50 2011 - Harvard University
Week 7: Wednesday - CS50 2011 - Harvard University
CS50
56 Week 8: Monday - CS50 2011 - Harvard University
Week 8: Monday - CS50 2011 - Harvard University
CS50
57 Week 9: Wednesday - CS50 2011 - Harvard University
Week 9: Wednesday - CS50 2011 - Harvard University
CS50
58 Week 8: Wednesday - CS50 2011 - Harvard University
Week 8: Wednesday - CS50 2011 - Harvard University
CS50
59 Week 10: Monday - CS50 2011 - Harvard University
Week 10: Monday - CS50 2011 - Harvard University
CS50
60 Week 2: Wednesday - CS50 2010 - Harvard University
Week 2: Wednesday - CS50 2010 - Harvard University
CS50

This video covers essential concepts in scale-out computing, including the basics of distributed systems, parallel computing, and cloud computing, with a focus on designing and implementing scalable solutions for various applications, including artificial intelligence.

Key Takeaways
  1. Understand the basics of scale-out computing
  2. Design distributed systems for scalability
  3. Implement parallel computing solutions
  4. Optimize system performance for big data and AI applications
  5. Deploy and manage cloud-based infrastructure
💡 Scale-out computing is essential for supporting the growing demands of artificial intelligence, big data, and other applications that require high-performance computing and large-scale data processing.

Related Reads

Up next
MARKET MUST REOPEN: NYSE trader RECALLS Wall Street's Message After 9/11
Fox Business Clips
Watch →