This is why Salesforce services went down

Hussein Nasser · Intermediate ·🔧 Backend Engineering ·5y ago

Key Takeaways

Salesforce services experienced a 5-hour outage due to a DNS update issue, which was resolved through the efforts of the Salesforce team using a microservices architecture with load balancing and health checks, and tools like Envoy proxy.

Full Transcript

on may 11 2021 salesforce services went down for five hours and uh the cause of this outage was as always dns we've seen this many many times how dns essentially affects your entire infrastructure and can you down but the question you might have is who's saying if venus goes down and you bring it back up does it really take five hours to for the services to pick app back up again that's the question i try to answer in this video so in this video i'm gonna go through the executive summary that uh the salesforce team have provided how did what exactly services went down and the cause obviously the dance and then i'm gonna give my opinion about why does it take so long when dns goes down how about jump into it all right guys how about we read the executive summary of why and what happened with salesforce back on may 11 2021 and uh then we're gonna discuss exactly how the dns essentially affect your services right all right on may 11 2021 at approximately nine o'clock the evening universal coordinated time utc the salesforce technology team became aware of service disruption across salesforce production instances the disruption impacted the ability for users to log into their salesforce environment within the core salesforce services marketing cloud commerce cloud government cloud experience cloud ruku pardot i don't know what part it is uh velocity in addition the status to salesforce.com trust site was also unavailable we've seen this many many times guys every time there's an outage the means to inform the customers about an outage is out is in outage too so customers have no idea that what really happened they just seen an outage and you can't really as as a company you can even inform them because the status website relies on a piece of infrastructure that goes down with the outage we've seen this many many times microsoft google amazon now salesforce every time so that's right now i think people realize it and companies realize that they need to separate the status infrastructure from the actual core infrastructure this is very very important right because you need to let your users know i mean the easiest ways to go off let's try to everybody i don't know how that feels really if i'm microsoft and i tell people about an outage in twitter i don't know about that it's like it's like everybody goes to twitter when things goes uh the outage and that that might be okay but i don't know how i feel about that let me know what you think guys trust site was also unavailable and customers were unable to log support cases so even they want to log support cases they couldn't some customers may have also experienced issues with multi-factor authentication during the incident as as we became aware of the disruptions and investigation began across the sales force technology organizers and resolve the disruption the technology team confirmed the impact of the production host which was impairing diagnostics and analytics as he was high confidence that the issue was triggered by an internal emergency domain name system change so the salesforce team made a change to the dns right that's where we know that triggered out but how let's read through the technical details a little bit i just need to know what exactly that and then we can talk about it we can we can kind of puzzle things together technical details the salesforce change the salesforce changed to the dns impacted multiple data centers affecting the sales for service for customers the status.for salesforce.com and salesforce authenticator i think this is the app that does the multi-factor authentication over the following hours the team worked to restore the service blah blah blah we're going through that all this jazz took took five hours to do exactly i'm interested in more a little bit when we do an update to a dns service like you fix the bug or you improve introduce a feature in the dns right what really happened let's read more root cause analysis the preliminary post-incident analysis by the source sales force technology team determined that a dns update triggered dns services to stop and not restart immediately there you go because when you do an update like i mean update i'm going to restart the service immediately but it it stopped but did not restart immediately that means it was only out for few a moment i don't know how long they didn't mention but that moment caused all this cascading effect so let's speculate like we always do in this channel and let's go back and because i think we're ready enough so let's go back all right guys so obviously kudos we start always with coolest all the engineers that work in salesforce and work very hard to bring the service back up and running i always say that because i don't want to be in their shoes because i was once upon a time a few few years back and that's not fun right getting cold midnight things go wrong you stay over hours just to fix the outage not fun so kudos good job guys now let's analyze why dns going out for a few like that's what we that's what we read we read it's like it's only a few moments right it didn't understand immediately that's what i can deduce from that statement right so we made an update to the service obviously to update an update you stop the service to take effect right unless you're using elixir or erlang so you can do hot swapping in the code that's what whatsapp does right they can change process while it's running that's just fantastic man right but uh the rest of us have to stop the process and restart it or do a rolling restart to pick up the change that's what we do all the time right but when we restart the service and it didn't restart immediately what will happen i mean yeah dns goes down for a few moments like what happens well in a micro services architecture or you don't have to call it micro services architect just a service oriented architecture where services are relying on each other to talk and you have a load balancing in place that tag these services with the health checks right if you try if a service tries to communicate to another service and it couldn't in this particular case because of a dns not available you queried the dns is not available and it says as a result the service will think that hey the other server is actually down so i'm going to try again and again obviously i can't find an ip address because maybe there's a service discovery that is not telling me what this ip address is and then you try again you try again and at the end of the day the service will say that hey by the way that the server is just down and the load balancers as a result will have the same effect if you make a request to a load balancer and load balancers tries to communicate to the back end service it will try multiple times if it's down if it couldn't communicate no matter what what error it gets which we talked about in another video we load balancers and proxies are nefarious for throwing this generic error that's called bad gateway i talked about it right here if you want to learn more about it but bad gateway means anything can go wrong from dns from tls handshake couldn't be established from a bad certificate on the back end of from a service taking too long to respond from a service actually giving you an invalid response like you're connecting connecting to the back end but giving you a malform response all of this is bad gateway and as a result bad gateway errors result into the back end being marked as unhealthy so future requests will never go to that so you will try another service that load balancer will right another seven and obviously it's going to be down because the dns is down well the load balancer is not smart enough i guess in this particular case to know that oh it's just the dns it's not really the back end you cannot possibly figure that out okay and and and the proxy proxies and the services are not smart enough to to look behind it obviously right because if you if you can't look up dns all bets are off right so for that period of time let's say one minute took one minute or two minutes all of these services are starting to just go down i don't by any when i mean go down they are marked as down marked as bad so services won't get any requests and as a result authentication will fail everything will fail and even if your dns is back up and running hmm it won't help you because guess what the load manager receiver request yeah it will it will start connecting to one service right again or because all of them are bad so maybe the auto scaling will kick in and say we'll spin up a new services and those right will pick up the new dns update and you all of a sudden there are hungry requests flood of waiting cued up requests just hitting that poor new service and obviously it's gonna go down and and we've seen this many many times when an outage happens right once a service goes up and that immediately goes down because the service cannot handle that potential law that's why what i like about envoy in particular as a proxy is there there's a feature called i talked about envoy guys check it out here invoice they thought about this they said okay we need to introduce panic mode panic mode means if all my back-ends are suspiciously unhealthy really i really don't care i'm just gonna forward the request anyway if those load balancers continue to request and forward the request to the back end right despite this dns that at least try then it would have maybe succeeded it will not have been removed from the backend fleet as a healthy as an unhealthy service right that's one of the advantages of one feature that can be easily added to all products and i believe all proxies should have this feature at least disabled by default give me an option to enable it panic mode is really really powerful again guys all of this here is a speculation we don't know the actual root cause analysis they are still working on it against kudos for the sales force team but that's why it takes time when dns goes down microsoft we've seen it like a few few weeks back i think microsoft went down because of dns once the dns has gone tough luck unless you have great resilience software it's gonna really be hard to bring it back up all right guys that's it for me today i'm gonna see you in the next one if you like this content subscribe i talk about back and stuff that's what i do and i'll see you in the next one guys stay awesome goodbye

Original Description

Salesforce services went down as a result of a DNS update, let us discuss how can tiny DNS unavailability cause a severe outage of 5 hours. From salesforce "On May 11, 2021, at approximately 21:08 Universal Coordinated Time (UTC), the Salesforce Technology team became aware of a service disruption across Salesforce production instances. The disruption impacted the ability for users to log into their Salesforce environments within the core Salesforce services, Marketing Cloud, Commerce Cloud, Government Cloud, Experience Cloud, Heroku, Pardot, and Vlocity. In addition, the status.salesforce.com Trust site was also unavailable, and customers were unable to log support cases. Some customers may have also experienced issues with Multi-Factor Authentication (MFA) during the incident. " Resources https://help.salesforce.com/articleView?id=000358392&type=1&mode=1 Support my work on PayPal https://bit.ly/33ENps4 Become a Member on YouTube https://www.youtube.com/channel/UC_ML5xP23TOWKUcc-oAE_Eg/join 🧑‍🏫 Courses I Teach https://husseinnasser.com/courses 🏭 Backend Engineering Videos in Order https://backend.husseinnasser.com 💾 Database Engineering Videos https://www.youtube.com/playlist?list=PLQnljOFTspQXjD0HOzN7P2tgzu7scWpl2 🎙️Listen to the Backend Engineering Podcast https://husseinnasser.com/podcast Gears and tools used on the Channel (affiliates) 🖼️ Slides and Thumbnail Design Canva https://partner.canva.com/c/2766475/647168/10068 🎙️ Mic Gear Shure SM7B Cardioid Dynamic Microphone https://amzn.to/3o1NiBi Cloudlifter https://amzn.to/2RAeyLo XLR cables https://amzn.to/3tvMJRu Focusrite Audio Interface https://amzn.to/3f2vjGY 📷 Camera Gear Canon M50 Mark II https://amzn.to/3o2ed0c Micro HDMI to HDMI  https://amzn.to/3uwCxK3 Video capture card https://amzn.to/3f34pyD AC Wall for constant power https://amzn.to/3eueoxP Stay Awesome, Hussein
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from Hussein Nasser · Hussein Nasser · 0 of 60

← Previous Next →
1 Extending ArcObjects (IGeometry) - 01 - Getting Started
Extending ArcObjects (IGeometry) - 01 - Getting Started
Hussein Nasser
2 Extending ArcObjects  (IGeometry) - 02 - The Document, The Map and The Layers
Extending ArcObjects (IGeometry) - 02 - The Document, The Map and The Layers
Hussein Nasser
3 Channel Update - New Book, New Job, New Videos
Channel Update - New Book, New Job, New Videos
Hussein Nasser
4 Learn Programming with VB.NET - 01 - Getting Started
Learn Programming with VB.NET - 01 - Getting Started
Hussein Nasser
5 Learn Programming with VB.NET - 02 - Classes and Objects (Part 1)
Learn Programming with VB.NET - 02 - Classes and Objects (Part 1)
Hussein Nasser
6 Learn Programming with VB.NET - 03 - Classes and Objects (Part 2)
Learn Programming with VB.NET - 03 - Classes and Objects (Part 2)
Hussein Nasser
7 Learn Programming with VB.NET - 04 - User Interface
Learn Programming with VB.NET - 04 - User Interface
Hussein Nasser
8 Learn Programming with VB.NET - 05 - By Value v. By Reference
Learn Programming with VB.NET - 05 - By Value v. By Reference
Hussein Nasser
9 Learn Programming with VB.NET - 06 - Variable size, 32 bit vs 64 bit
Learn Programming with VB.NET - 06 - Variable size, 32 bit vs 64 bit
Hussein Nasser
10 Learn Programming with VB.NET - 07 - Conditional Statements
Learn Programming with VB.NET - 07 - Conditional Statements
Hussein Nasser
11 Learn Programming with VB.NET - 08 - Inheritance
Learn Programming with VB.NET - 08 - Inheritance
Hussein Nasser
12 Learn Programming with VB.NET - 09 - Strategy Design Pattern
Learn Programming with VB.NET - 09 - Strategy Design Pattern
Hussein Nasser
13 Learn Programming with VB.NET - 10 -  How did I learn programming
Learn Programming with VB.NET - 10 - How did I learn programming
Hussein Nasser
14 IGeometry 2016 Retrospective - Channel Update
IGeometry 2016 Retrospective - Channel Update
Hussein Nasser
15 Javascript by Example - The Vook
Javascript by Example - The Vook
Hussein Nasser
16 Vlog - Keep your servers close and your database closer
Vlog - Keep your servers close and your database closer
Hussein Nasser
17 Vlog - Client/Server Programming Languages
Vlog - Client/Server Programming Languages
Hussein Nasser
18 Javascript By Example L1E01 - Getting Started
Javascript By Example L1E01 - Getting Started
Hussein Nasser
19 Persistent Connections (Pros and Cons)
Persistent Connections (Pros and Cons)
Hussein Nasser
20 Javascript By Example L1E02 - Building the Calculator Interface
Javascript By Example L1E02 - Building the Calculator Interface
Hussein Nasser
21 Happy new Year from IGeometry!
Happy new Year from IGeometry!
Hussein Nasser
22 Synchronous v. Asynchronous
Synchronous v. Asynchronous
Hussein Nasser
23 Javascript By Example L1E03 - Displaying the Digits on Calculator Screen
Javascript By Example L1E03 - Displaying the Digits on Calculator Screen
Hussein Nasser
24 Show Your Work. Blog, Vlog, Write, Create and Develop!
Show Your Work. Blog, Vlog, Write, Create and Develop!
Hussein Nasser
25 Relational Database Atomicity Explained By Example
Relational Database Atomicity Explained By Example
Hussein Nasser
26 Javascript By Example L1E04 - Operators, All Clear with Arrow Functions
Javascript By Example L1E04 - Operators, All Clear with Arrow Functions
Hussein Nasser
27 What Comes First, User Experience or Software Architecture?
What Comes First, User Experience or Software Architecture?
Hussein Nasser
28 Javascript By Example L1E05 -  Evaluate the Calculator Expressions with eval
Javascript By Example L1E05 - Evaluate the Calculator Expressions with eval
Hussein Nasser
29 Fastest Way to Learn Programming Language or Technology
Fastest Way to Learn Programming Language or Technology
Hussein Nasser
30 Javascript By Example L1E06 -  Fix Leading Zero Bug with Conditions
Javascript By Example L1E06 - Fix Leading Zero Bug with Conditions
Hussein Nasser
31 Stateful vs Stateless Applications (Explained by Example)
Stateful vs Stateless Applications (Explained by Example)
Hussein Nasser
32 Javascript By Example L1E07 - Running our Calculator on the Mobile Phone
Javascript By Example L1E07 - Running our Calculator on the Mobile Phone
Hussein Nasser
33 Advice for New Software Engineers and Developers
Advice for New Software Engineers and Developers
Hussein Nasser
34 Why JSON is so Popular?
Why JSON is so Popular?
Hussein Nasser
35 Building Scalable Software - SLA, HS, VS
Building Scalable Software - SLA, HS, VS
Hussein Nasser
36 Vlog (Istanbul) - Datacenter Proximity
Vlog (Istanbul) - Datacenter Proximity
Hussein Nasser
37 Should Software Engineers Learn Bleeding-Edge Technologies?
Should Software Engineers Learn Bleeding-Edge Technologies?
Hussein Nasser
38 Do Developers Build Bad User Interfaces/Experience?
Do Developers Build Bad User Interfaces/Experience?
Hussein Nasser
39 Learn By Doing.
Learn By Doing.
Hussein Nasser
40 I Wrote Bad Front-End Code That Broke Chrome
I Wrote Bad Front-End Code That Broke Chrome
Hussein Nasser
41 My Story
My Story
Hussein Nasser
42 Vlog - Horizontal vs Vertical Scaling
Vlog - Horizontal vs Vertical Scaling
Hussein Nasser
43 Can User Experience Help Build Better Rest API?
Can User Experience Help Build Better Rest API?
Hussein Nasser
44 Reverse engineering Instagram in flight mode
Reverse engineering Instagram in flight mode
Hussein Nasser
45 The Benefits of the 3-Tier Architecture (e.g. REST API)
The Benefits of the 3-Tier Architecture (e.g. REST API)
Hussein Nasser
46 Stateless v. Stateful Architecture (Podcast)
Stateless v. Stateful Architecture (Podcast)
Hussein Nasser
47 The evolution from virtual machines to containers
The evolution from virtual machines to containers
Hussein Nasser
48 Proxy vs. Reverse Proxy (Explained by Example)
Proxy vs. Reverse Proxy (Explained by Example)
Hussein Nasser
49 Canary Deployment (Explained by Example)
Canary Deployment (Explained by Example)
Hussein Nasser
50 No Excuses
No Excuses
Hussein Nasser
51 Synchronous vs Asynchronous Applications (Explained by Example)
Synchronous vs Asynchronous Applications (Explained by Example)
Hussein Nasser
52 What is an Asynchronous service?
What is an Asynchronous service?
Hussein Nasser
53 Difference between Client Polling vs Server Push in Notifications
Difference between Client Polling vs Server Push in Notifications
Hussein Nasser
54 Software vs. Hardware AdBlockers (Explained by Example)
Software vs. Hardware AdBlockers (Explained by Example)
Hussein Nasser
55 HTTP Caching with E-Tags -  (Explained by Example)
HTTP Caching with E-Tags - (Explained by Example)
Hussein Nasser
56 Simple Object Access Protocol Pros and Cons (Explained by Example)
Simple Object Access Protocol Pros and Cons (Explained by Example)
Hussein Nasser
57 Nodejs Express "Hello, World"
Nodejs Express "Hello, World"
Hussein Nasser
58 Reverse Engineering Instagram feed
Reverse Engineering Instagram feed
Hussein Nasser
59 Popup Modal Dialog with Javascript and HTML
Popup Modal Dialog with Javascript and HTML
Hussein Nasser
60 MIME and Media Type sniffing explained and the type of attacks it leads to
MIME and Media Type sniffing explained and the type of attacks it leads to
Hussein Nasser

The video discusses the 5-hour outage of Salesforce services due to a DNS update issue and how the team resolved it using a microservices architecture and tools like Envoy proxy. It highlights the importance of service reliability and robust system design. By watching this video, viewers can learn how to design and implement resilient systems and troubleshoot DNS issues.

Key Takeaways
  1. Identify potential single points of failure in system infrastructure
  2. Implement load balancing and health checks to mitigate issues
  3. Use tools like Envoy proxy to forward requests despite unhealthy backends
  4. Design systems with service reliability in mind
  5. Troubleshoot DNS issues to prevent widespread outages
💡 A small DNS update issue can cause a severe outage if not properly mitigated, highlighting the importance of robust system design and service reliability measures.

Related Reads

Up next
/dev/push: An Open Vercel Alternative to Ship Your Apps Quickly
Ian Wootten
Watch →