YouTube transcripts

8,000 failures overnight. Postmortem.: video thumbnail

8,000 failures overnight. Postmortem. transcript

Arjay McCandless · @arjay_the_dev

Published August 14, 20267:3019.7K views

Watch this video on YouTube

Transcript analysisComputed from the caption text

Words

1,789

Runtime

7:30

Speaking pace

239wpm

Reading time

7min

239 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.

Opening (first 30 seconds)

So, last night I'm sleeping and I get an email around 2:00 a.m. that there was 8,000 failures on one of my webhooks and obviously this is not what you want to happen. Um but it means that my monitoring and alerting worked and my recovery systems actually were able to automatically recover this stuff. So, I want to walk you guys through a little bit of a postmortem for how you can design your systems to automatically recover from failures like this and walk you through the exact failure method in this case so that hopefully you and I both can learn from this. And if you don't know me, my name is RJ. I'm a former Amazon software

120 words, the words spoken in the first 30 seconds at 239 words per minute.

Sentence shape

MeasureThis transcript
Sentences108
Average words per sentence16.6
Longest sentence50 words
Questions asked5
Sentences containing a number17

Most used terms

  • event11
  • events10
  • google10
  • system10
  • retry9
  • actually8
  • server8
  • superbase8
  • try8
  • email7
  • google pub7
  • pub7

Filler phrases

35 in total: like 11 · actually 8 · um 5 · basically 4 · kind of 4 · literally 1 · right? 1 · you know 1.

A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.

What this transcript is

Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.

Transcript

So, last night I'm sleeping and I get an email around 2:00 a.m. that there was 8,000 failures on one of my webhooks and obviously this is not what you want to happen. Um but it means that my monitoring and alerting worked and my recovery systems actually were able to automatically recover this stuff. So, I want to walk you guys through a little bit of a postmortem for how you can design your systems to automatically recover from failures like this and walk you through the exact failure method in this case so that hopefully you and I both can learn from this.

And if you don't know me, my name is RJ. I'm a former Amazon software engineer. Then I went into the startup space for a little bit and now I run my own company where I have a couple SaaS products that serve tens of thousands of users. Um and yeah, I'm kind of just sharing my lessons from building and maintaining those projects from a big tech software engineer's perspective. So, if you're interested in learning about system design, scaling products, stuff like that, you're in the right place.

So, let's get right into it. So, let me start by walking you through how my project works when everything is going well and there's no failures. So, one of the things that my product does is it allows me to integrate directly with your Gmail. This is a restricted Google scope. I went through a cybersecurity audit and got it approved. So, I have access to a user's mailbox if they grant me permission. So, what happens is when a user gets an email in a thread that we're tracking, that publishes an event to Google Pub/Sub.

And Google Pub/Sub will try to push an event to my server to handle that webhook. My server will then try to correlate that event with a mailbox of somebody who has an account created with me. Right now, I'm using Superbase as my backend which stores users' mailboxes and information like that. So, in the working case we'll be able to correlate their mailbox ID with their email and their account, correctly surface that data, and store it for the user, return successfully back to Google Pub/Sub with a 200 saying, "Hey, we've successfully processed this event." And this is a very standard event-driven architecture for anything like this.

So, what happened in this case is that Superbase was actually having an outage of some kind and they haven't released a full postmortem because this was literally last night that I'm making this video. So, um I don't know exactly what happened, but basically I know that my connection to Superbase was being refused. And this was going on for about an hour. And during this time, when Google Pub/Sub sent an event to my server, my server would respond with a 503 error code.

And what happens here is untracked events started piling up within Google Pub/Sub. Now, Google Pub/Sub has automatic retries in place. This is one of the couple things I have in place to help mitigate this kind of failure mode. Um but unfortunately, what I didn't have configured was a cap on my retries. So, what happened is basically the same events got stuck replaying again and again and again. So, what started with 10 failures escalated into 100 and then a thousand.

This is what's called a retry storm. Now, normally this mechanism actually does protect me because let's just say there is one event that gets rejected from Superbase or my server happens to fail in some case, Google Pub/Sub will retry on those events and that will be fine. Those events will succeed on the second, third, fourth try. Network errors are guaranteed in any distributed system, so you should always have retries of some kind up to, let's just say, five times.

And this is one of two resilient mechanisms I have. So, this one actually hurt me here. Um and let me explain the fix. So, when you have a retry mechanism, it's very important that you have something called exponential backoff. And sounds complicated, it's extremely simple. Basically, every single time you retry, you should set a timer and wait before you retry again to give the server time to recover from whatever intermittent failure it might be facing.

This is a really simple example of how that might work. You try something, it fails, wait 1 second, try it again. Fails again, wait 2 seconds, try it again. Fails again, wait 4 seconds, try it again. Fails again, stop. Because if we keep retrying here, that creates that retry storm where we're just constantly retrying or the period gets so long that it's not going to retry again for 5 years. And who cares about that?

And one other pro tip is instead of having this just be 2 seconds, this should be like 2 seconds plus or minus some random number. And you should do this in all cases because otherwise your retries can get synced exactly. And then you have a bunch of failures coming at the exact same second and then they retry again exactly 2 seconds later and then they retry again exactly 4 seconds later and they stack on top of each other.

This is a problem. So, you want to add something called jitter or a little bit of randomness in between exactly how long the retries take. So, in this case I did have exponential backoff but I didn't have a cap. So, it ended up generating a ton of requests. But, here's what actually ended up saving me. This is my second layer of resilience that I have built into my system. So, we missed all those events from Google. The re- tries were infinitely happening.

Things were bad. But, eventually Superbase recovers. And I actually have a cron job that runs every hour to try to sync up any events I might have missed from Gmail with my database. And this is very, very standard in any event-driven system. You know that you're going to miss events. So, you want to have the ability to replay those events or somehow sync between the original source of truth and your system that doesn't just rely on these events that can fail.

So, for me this is a simple job where my server goes to Gmail and says, "Hey, were there any new emails since our last sweep?" If it says yes, then we get those emails and we pass them to Superbase. Of course, this entire operation is item potent. So, if we already found the email, it doesn't matter if I send another event to the server. If you're running any kind of event-driven system, I highly recommend having some kind of reconciliation job like this.

It saved my life on many occasions. And quick plug, if you're interested in system design, I highly recommend my app DevMax. I basically give you new system design questions every single day plus access to my Discord community where you can learn directly from me and other people who are interviewing, building projects, and mastering system design. I also have a weekly league which I recently introduced where people are getting super competitive and trying to climb to the top.

So, if that sounds interesting, you can check it out on the App Store or Google Play. I'll link it in the description down below. Back to the video. So, now let's walk through the final event timeline with the reconciliation job. So, Superbase downtime begins at let's just say 2:07 and between 2:07 and 3:05 I have 8,000 errors. At this point I got an email somewhere in here but it didn't wake me up because I was sleeping and this product currently only has around 1,000 users, most of which are US-based.

So, this wasn't a huge deal for my customers. In fact, they didn't actually experience any downtime in the product itself. And around 3:05, Superbase recovers and at 4:00 my hourly cron job runs, which restored everything that might have been missed inside Gmail. So, when I woke up in the morning, everything was actually in a good state, which is amazing. This is like the best case for a failure in a distributed system.

Now, one thing I always like to do after events like this, and this was standard to Amazon, is a list of action items. So, what am I personally going to do to make sure that this doesn't happen again in the future or how can I improve what happened here? And to me, the number one lesson is I need to introduce exponential back off plus a cap. I'm going to set mine at three requests because I know the reconciliation job can easily fix stuff and if my customers are waiting for an hour instead of 20 or 30 seconds to get their data, it's not that big of a deal.

People are slow to check their emails. And of course, this is a product level decision, right? I know that email checking is a slow behavior. So, it doesn't really matter to me that we're able to reconcile it within 5, 10, 15 minutes. I'm fine with it being an hour because his email is slow. If this was something more important like a text message or an alert for an emergency weather service, then I would go about approaching this differently.

But, because it's not, I'm fine with this taking longer. So, think about your product when you make decisions like this. How long should it take for you to recover and what is the impact if you don't? And the general lesson from this is your dependencies are going to fail. You should design your systems that it recovers without you. This is obvious, but so many people don't do this and just depend on manual intervention to get their systems back up.

If I didn't have this in place, I would have been manually running scripts at 6:00 a.m. as soon as I woke up, which is a miserable experience and I've done it many times. And I don't want to do it now. Anyway, so I wanted to keep this quick. Let me know if you guys found this useful or you want to see me break down more postmortems for my own systems in the coming months here. My apps are scaling pretty quickly, so I assume there's going to be a lot of failures.

Subscribe so you don't miss the next episode and we'll see you guys next time.

The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.

Use this transcript

Three free tools that work on the material around a video like this one. No signup, no login.