Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

JavaScript Mastery · @javascriptmastery
Words
5,876
Runtime
38:00
Speaking pace
155wpm
Reading time
24min
155 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
You can build a full application now. Front end, backend, database, deployed, and all online. And with AI, you can do it faster than you ever could. So here's the uncomfortable question. If anyone can build the app now, what's left that actually makes you a good engineer? Because building the thing was never the hard part. The hard part is what happens when the app you built gets popular. When one server isn't enough, when database can't
78 words, the words spoken in the first 30 seconds at 155 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 442 |
| Average words per sentence | 13.3 |
| Longest sentence | 37 words |
| Questions asked | 16 |
| Sentences containing a number | 11 |
Most used terms
Filler phrases
30 in total: like 17 · actually 8 · kind of 2 · I mean 1 · right? 1 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
You can build a full application now. Front end, backend, database, deployed, and all online. And with AI, you can do it faster than you ever could. So here's the uncomfortable question. If anyone can build the app now, what's left that actually makes you a good engineer? Because building the thing was never the hard part. The hard part is what happens when the app you built gets popular. When one server isn't enough, when database can't keep up, and when it's slow and you don't know why.
That's system design. And most tutorials teach it as a pile of scary words. Load balancers, replicas, sharding, caching thrown at you all at once on day one. all for an app that has 50 users. We're going to do the opposite. Hi there, I'm Adrian and here's the plan. We start with one server and one database, which is already enough for way more users than you think. Then we break it on purpose over and over. And every time it breaks, we fix exactly one thing.
So that by the end, you'll have built the whole system and you'll know why every piece is there and what it costs to add it. So let's break some things. This is your app. one server running your code and one database holding your data. If you've built anything and put it online, this is what you have. And I want to say something about this diagram before we do anything about it. This is enough for thousands of users, not dozens, thousands.
Most applications that exist right now look exactly like this and will never need to look like anything else. That is not a failure and it's not something you have to grow out of. So, let's break it. Let's say you build a web app for buying concert tickets, but for real, just like how the big players like Ticket Master, Subhub, or Seedge Geek do it. You deploy it, you test it, and everything works perfectly. And then a big show goes on sale.
It's 10:00 in the morning if you're lucky and about 5:00 p.m. on a Friday if you're not. And thousands of people, if not tens or hundreds of thousands of people open up your site at the exact same moment to grab a ticket, and your app falls over. It's slow to crawl, requests hanging, and some never finish at all. And nothing is wrong with your code. There's no bad queries and no bugs. The problem is simpler than that.
Your app is running on a single server and a single server can only do so many things at once. So when thousands of requests arrive together, they pile up in a queue waiting their turn. And that queue gets so long that requests just time out before they are ever reached. So the question is, how do you handle thousands of people at once instead of having a single server that's drowning under them? Well, you have two options and they're called vertical scaling and horizontal scaling.
And you'll pick between these two for real one day. So, let's properly understand both. Start with vertical as it's the obvious one. You get a bigger server, more processing power, more memory, same server box, just a stronger and larger one. And this is the right first move almost every time. You change nothing about your code. You pay for a bigger machine and you're done for the afternoon. If a bigger server fixes your problem, buy the bigger server and get back to work.
And to put a number on it, a $20 a month server handles hundreds of requests a second for a typical app. That is millions of requests a day on one cheap box. But it runs out. There is a biggest machine you can rent and it gets very expensive long before you reach that. So a bigger box buys you time, but it doesn't buy you forever. Which brings us to the second option. Instead of making one box stronger, you stop making one box do all the work.
So you run several ordinary servers and split the people between them. This is scaling horizontally. more machines sharing the load. And I want to be very clear about what these three servers are because this trips people up. They are not three different parts of your app. This is not one server for payments, one for messaging, and one for something else. These are three identical copies of your app running at the same time. the same code running three times and each copy can handle any request on its own.
And to be clear, that doesn't mean that you deploy your app three separate times. You ship it once. You just tell your platform to run three instances of it. And if you're on something like Versel, it is quietly already doing this for you. And this approach is the one that actually keeps going. Too much traffic again next year? Well, just add more copies as there is no ceiling in the same way because you're not depending on one machine being big enough.
But the moment you have more than one server, you have a question that you didn't have before. A request comes in from a user. Which server does it go to? Something has to stand in front of your servers and hand each incoming request to one of them. spread the people out evenly so no single server gets overwhelmed while the others sit idle. And that thing is called a load balancer. A request comes in, the load balancer picks a server and the server does the work.
If one server is busy, send the next request to a different one. And if you add a fourth server next month, the load balancer starts using it. The users never know how many servers there are or which one they got. Engine X is a classic one and if you deploy on something like Versel or Railway, you already have a load balancer. You just never had to set it up or think about it. But we also have to talk about the cost because everything we add from here has one.
You just added a new box that everything flows through which means that it is now a single point of failure. If the load balancer goes down, it doesn't matter that you have three healthy servers behind it, nobody can reach them. So in real systems, the load balancer itself also has a backup which is more stuff to run. And there is also a second cost which is quieter and it breaks something that you can't even see. For the load balancer to send any request to any server, all your servers have to be interchangeable.
Any of them has to be able to handle any request and right now they are not. So the next section is where that bites. So now we are running multiple instances of our app and that creates a problem that you didn't have with just one. A user logs in on server A. They browse. They pick a ticket and they hit buy. That request goes to server B and server B immediately tells them you're not logged in. Even though they just logged in a second ago, nothing is broken.
But that happens the moment you run more than one server and do not handle sessions properly. So, let me show you why. First, what does being logged in even mean? When you log in, the server checks your password and then it needs to remember that it is you on every request after that. So you're not retyping your password on every single click. So it makes a note. This person is logged in as Bob. It hands you a session ID, which is a random string that points at that note.
Your browser sends that session ID with every request. The server looks at its node and it knows it's you. The question that matters is where does the server keep that note? And by default, the easy answer is that the note lives in the server's own memory. So now if you try to put that together with three servers, you log in, the load balancer happens to send that request to server one. So server one makes the note Bob is logged in in one server's memory.
But a moment later, you click on something that's a new request. The load balancer sends this one to server two and server two has never heard of you because the note is on server one. So server two looks in its own memory, finds nothing and does the only thing it can. It asks you to log in again. That is the random log out. You're not being kicked out, but you're rather being sent out to a different server that doesn't have your note.
And with three servers, roughly two out of every three clicks will land on a server that doesn't know you. So this is what I meant at the end of the last chapter when I said that servers were not interchangeable. This is the thing that was not interchangeable. Your login was stuck on a specific server. So the fix is a rule. The servers are not allowed to remember anything about you between requests. No server keeps a note in its own memory.
Instead, the note goes somewhere all of them can reach, a separate place off to the side that every server can read from and write to. Now, it doesn't matter which server you land on. You log in, the note gets written to that shared place, and then on the next click, whatever server the load balancer picks, that server reads the same shared place, finds your note, and knows it's you. A server that keeps nothing about you in its own memory is called stateless.
Every server is now interchangeable, exactly like the load balancer needed. Any server can handle any request because none of them are keeping secrets the others cannot see. And that shared place is very often something called Reddus, an in-memory storage that is extremely fast to read from, which matters because you're hitting it on every request. And hold on to that name because it's going to come back later for a completely different job.
And let's not forget to talk about the cost because every request now does one extra thing. It reaches out to that shared store to find your note before it can do anything else. That is a little bit of time added to every single request and one more piece of infrastructure that has to be up because if that shared store goes down, nobody can be logged in at all. So at this point you made all your servers interchangeable which you needed to do and you paid for it with one more lookup everywhere and one more thing that has to stay alive.
So now the front of our system is solid load balancer several stateless servers and everyone gets served no matter how many people show up and all of them are talking to one database. So you can probably guess where this is going next. Now, what if the app is slow again? But this time, when you look, every server is healthy. None of them is at capacity. There are no bad queries and nothing seems to be broken, which is confusing until you notice the thing that all servers have in common.
They are talking to the same one database. We scale the servers, but we never scale the thing behind them. So there are two separate problems hiding here. So let me walk you through them one at a time. The first problem is that the database will only talk to so many people at once. And that's a fact about databases that many people don't know about. A database will only hold so many open connections at the same time.
Each connection, each open line from a server to the database costs it memory and attention. So, there's a hard limit on how many it'll accept at once. And that limit is lower than you would guess. Not millions, often just a few hundred on smaller hosted plans, sometimes even a few dozen. Now, count how many things are trying to connect. Three servers and each server opens several connections so it can handle several requests at once.
That adds up fast. And if you're on serverless, which a lot of you are on something like Versell, it gets worse in a way that surprises people. When a request comes in and there is no free instance ready for it, the platform spins up a fresh copy of your function. And each fresh copy opens its own new connection. So a big spike means hundreds of copies appearing at once, each grabbing connections. And they fill every slot in seconds.
Oh, and here's one tell that'll make this recognizable very quickly. The symptom isn't slowness. It is errors. Your ticket app was completely fine, but then a popular show went on sale, the traffic spiked, and suddenly you're getting database connection errors right at the moment things were going well. That's what this is. So, the fix for this problem sits between your servers and your database. And it is almost boring just how simple the idea is.
Instead of every server opening and closing its own connections whenever it likes, you keep a small fixed set of connections open at all time and everything shares them. A request needs a database, so it borrows one from the set, uses it, and hands it straight back for the next request to use. This is called a connection pool. The database only ever sees that small fixed number of connections no matter how many servers or requests you have.
You stopped overwhelming it with the sheer number of open lines. PG bouncer is a common one for Postgress and most managed databases now ship a pooler that you can just switch on. That one fixes the errors, but it doesn't fix the other problem, which is the one people usually mean when they say that their databases cannot keep up. And that's the second problem. The database is genuinely out of capacity. Say you have the connections under control and the database is still maxed out.
It is genuinely doing more work than one machine can do. Before you do anything clever, just do one quick check. Are your common queries actually using indexes? If your app looks users up by email on every login and there is no index on the email column, the database is scanning every row to find one person every single time. That one missing index looks exactly like a capacity problem which it is not. So add the index and the load can just disappear.
So check that one first as it is free. But say you have checked it and the indexes are there. The database is simply out of room. Now look at what it is spending all of its efforts on. Think about how you actually use an app like this. Any social app you open it and scroll. You read post after post after post. And once in a while maybe you write something, a post, a like, a comment. So for every one thing you write, you read hundreds of things.
And that is true of almost any app out there. I mean, just imagine Tik Tok. You're reading videos, right? You're scrolling through different content that is served to you. So reads massively outnumber the writes, which means the database is spending nearly all of its efforts answering the reads. So if you could somehow take the reading load off of it, you would fix nearly all of the problem. So the fix is make copies of the database.
One database stays the original. The only one you're allowed to write to and everything you add or change goes there. Let's call it our primary database. And then you make copies of it. The copies are read only. You cannot write to them, but you can read from them all day. And whenever something changes on the primary, it sends that change out to the copies so they stay up to date. These copies are called read replicas.
Now the work is spread out. Writes go to the primary reads. The huge majority of your traffic get shared across the replicas. Each machine is doing a fraction of what one machine was drowning under. And again, we got to talk about a cost because we keep adding new rectangles to it. Uh, but yeah, this one has a cost that is different from the others. Because it doesn't just cost you money or a moving part, but it quietly makes your app a little wrong.
Here's why. When you write something to the primary, it takes a moment, usually tiny, but not zero, for that change to reach copies. So, picture this. You post something, that write goes to the primary. Immediately your app loads your feed and that read goes to a replica, but the replica has not yet received your new post. So you look at your own feed right after posting and your post isn't there. You refresh and then it's there.
Nothing was broken. The copy was just a half second behind. That gap has a name, replication lag, and it is the price of read replicas. Your data is now slightly out of date on copies for a moment sometimes, which is usually completely fine. Somebody seeing a post half a second late doesn't matter, but sometimes it does matter. You seeing your own post missing feels broken. So for those specific cases, apps deliberately read from the primary instead of a replica.
And that is the real lesson here, the one worth taking over everything else in this video. You just made your app faster and slightly wrong on purpose. And you chose which parts are allowed to be slightly wrong and which parts are not. So with that, the database is handled. writes to the primary reads spread across the replicas connections pulled. But some of these reads are still expensive. There is one particular kind of question that we ask the database over and over thousands of times and get almost the same answer every time.
And computing the same answer thousands of times is just wasteful. So next we stop doing that. Take the follower count on a profile. To get it, the database has to count every follower record for that person. And if someone has 2 million followers, that is the database counting through 2 million rows to produce one number. Now that number is on their profile, which means that every single person who opens that profile makes the database count all 2 million. again thousands of times an hour the same expensive count for a number that barely moves.
That is just wasteful. We're making the database do hard work over and over again to get the answer we already had a second ago. So instead of computing it every time, we compute it once and we keep the answer somewhere fast. So the next person who asks gets the saved answer in the database is never touched. That saved answer kept somewhere fast and close is a cache. The flow is simple. A request comes in asking for the follower count.
First we check the cache. Is the answer already sitting there? If yes, we hand it straight back and we never go near the database. If no, then we ask the database. We get the number and this is the important part. We put it in the cache on the way back. So the next person gets it for free. And this is fast in a way database cannot be because a cache keeps his data in memory rather than reading from disk. And that is why the answer comes back in a fraction of the time.
And this is very often reddus. The same Reddus I mentioned back when we need a shared place for login sessions. Same piece of tech doing a second job which is worth noticing. You didn't add a whole new system to your stack. The thing that you already ran for sessions can hold your cash too. So what's the cost? Well, the caching has the sharpest cost in this whole video. The moment you save a copy of an answer, that copy can go out of date.
Someone gets a new follower. So the real count in the database is now one higher, but the cache is still holding the old number. So everyone opening that profile sees the old count until the cache is updated or thrown away. So your follower count is now sometimes a little wrong and deciding when to throw away that saved answer. so it doesn't stay wrong for too long is genuinely one of the harder problems in the whole field.
There's an old joke among engineers that the two hardest things in computing are naming things and knowing when to clear your cache. So here's the real skill, which isn't really technical. Caching trades correctness for speed on purpose. And you have to decide which pieces of your app are allowed to be a little wrong for a little while and which are not. A follower account that is 30 seconds out of date, nobody will ever notice it and nobody's harmed.
Cash it hard. But the balance in somebody's bank account, well, that can never be even slightly out of date. You do not cash that ever. And I want to stop on that decision for a second because it is the whole point of a video like this. Deciding that a follower account can be 30 seconds stale, but an account balance cannot be stale even for a moment. That is not a coding decision. An AI will happily write that caching code for you and it'll write it well.
But what it will not do is make the best call because that call is not about code. It is about your product and your users and what they will and will not accept. It needs someone who understands the whole system and the people using it. And that is the part of the job that isn't going anywhere because you can outsource the typing but not the decision-m. And with that in mind, the expensive reads are now cached. The database is doing far less work, which means that our app is quick front to back.
But there is one more kind of slow and it is a strange one because it has nothing to do with reading the data at all. Think about a typical signup flow. A user fills out the form, clicks sign up, and expects the account to be created right away. But creating the account might not be the only thing that your server has to do. You might also need to send a verification email, track the sign up in your analytics, or set up some initial data for that new account.
And if you do all of that before sending the response, the user is just sitting there staring at a loading spinner waiting for everything to finish. And that's more than just slow because sending that verification email is not something your server does itself. It has the email to an outside service and waits for that service to say sent. So your signup is now only as fast as that email service. If it is having a slow day, your signup is slow today.
And if it is down, your signup fails, too. The account couldn't be created because the verification email couldn't be sent. you have tied whether someone can even join your app to whether some separate email company happens to be up right now. And that is a bad thing to tie together. And you've already seen this. You just didn't know it. When you sign up for something and the verification email takes a few seconds to land or you upload a video and it says processing and lets you carry on, that is exactly this.
The delay is not the app being slow, but it is the app being built correctly. So the idea is this. Creating the account is the only part the user actually needs to happen before you answer them. The verification email doesn't have to happen in that same moment. It just has to happen soon. So you split the work. The instant the account is created, you answer the user, done. You're signed up. And the email, you don't send it right then.
You write down a note that says, "Send a verification email to this person." And you drop that onto a list. That list is called a que. It is just a line of jobs waiting to be done in order. And then something separate, not the server handling the user's request, but a different process whose only job is to work through that list, picks up that note and sends the email. a moment later while the user is already happily looking at their loggedin homepage.
And that separate process is called a worker. The server's job is to answer the user fast and drop the slow work onto the queue. And the worker's job is to quietly chew through that work in the background. And it isn't just verification emails, password resets, notifications, receipts, any slow job you do not want a user waiting on. All of it goes on the same queue and the worker handles it in the background. For a lot of apps, this is a tool called Bool MQ and it runs on Reddus, the same Reddus you already have for sessions and your cache. three jobs now in a single piece of technology.
So, if you'd like me to do a video on Reddus and everything it offers to help you with your system design, let me know in the comments down below. But yeah, that is the whole win. The signup becomes instant and it no longer cares whether the email service is even up. If it is down, the note just waits in the queue until it comes back up. The account was already created either way. And let's not forget to talk about the cost first.
The work now happens later, which is exactly why that verification email takes a few seconds to show up. For a verification email, nobody cares. You just need to know that done now means we have promised to do this and not this is finished. Second, you have two more things to run and you have to keep both the queue and the worker alive. And third, jobs can fail. the worker tries to send the email and the email service is down.
So you need to plan for that try again in a minute and if it keeps failing put it somewhere so a human can look at it later. A job you dropped on a queue and never checked on is a job that you're quietly not doing. So what we've done up to this point is real serious architecture. Most apps never need more than what we have on the screen right now. But there is one last thing that can break and it's not like the others.
Everything we fixed today was about handling more people. This last one isn't about people. It's about how much data you have. Remember those read replicas? Everyone is a full copy of your database, which quietly assumes that your whole database fits on one machine in the first place. But what if it does not? What if the data just keeps growing until it's too big for any single machine to hold? A copy cannot save you now because there's no machine big enough to copy it onto.
So you do one thing we have avoided all video. You stop keeping all your data in one place and you split it across several databases and each one holds only a part of it. But that immediately raises a problem. If your data is spread across, say, three databases, how does your app know which one to put a new user in? And later, how does it know which one to find them in again? You cannot just drop them anywhere. If you save a user randomly, you would have to search all three databases every time to find them.
And that is slower than the one database that you started with. So obviously you need a rule. You pick one column, something that every record has. For users, the obvious choice is the user ID and you use that ID to decide which database they go on. The simplest possible rule is to take the user ID and divide it by the number of databases you have. So if you have three databases, divide by three and the remainder tells you where they go.
If the remainder is zero, database 1, remainder one, database 2, and remainder two, database 3. Now your users are spread evenly across all three. Each database holds roughly a third of them. So no single machine has to hold everyone anymore. One quick thing, if your ids have letters in them, like a UYU ID, you first run them through a hash function, which turns any value into a number, the same number every time. And then the exact same divide by remainder rule applies.
And each of these databases is called a shard. And this whole approach, splitting your database across shards by a rule, is called sharding. And here's the part to actually make it click. And the part that while I was doing the research, I saw most other teachers skip. Watch what happens when a request comes in for a user number five. The app runs the same exact rule. 5 / 3, remainder 2, which is database 3. It goes straight there.
It doesn't search. Doesn't check the other two. It just knows. And that is the whole trick because the rule is consistent. The same user always lands on the same shard. So you always know exactly where to look. Writing data and finding it again both use the same rule. One shard every time. So you have essentially split your lookup time and your reading time and your writing time by three times. But now with this cool thing you just implemented, there is still something that you gave up and it is a real cost.
As long as a question is about one user, it is fine. One shard, you get it fast. But some questions are not about one user. Say you want to count how many users you have in total. There is no single shard that knows the answer because no shard has all the users. So you have to ask all three, get their separate counts, and then add them up yourself. A question that used to be one simple query is now three queries and some additional assembly.
And it gets worse the more shards you have. 10 shards means asking 10 databases and combining 10 answers. Anything that has to look across all your data, counting, sorting everything by date, searching everyone now has to touch every shard. And these are called crossshard queries. and avoiding them is the whole art of sharding. It is why choosing that rule which column you split on is such a big decision. Pick well and almost every query stays on one shard.
But if you pick badly, all your common queries will have to hit every machine and you have built something slower than what you initially started with. Oh, and there's one more thing. Once your data is split like this and all your code is written around it, going back is its own painful migration. So this is the most powerful tool in the whole video and still the one you reach for last by a long way. Teams put it off for years and they are right to you do it only when the data simply will not fit any other way.
So that is sharding as an idea. Doing it for real on a live app with millions of people already using it is a different level of hard. The migration alone, moving all that data onto shards without taking the app down is a serious story. Notion did this. They went from one database to sharded at over 200 billion rows while the whole product stayed alive. And if you want to see how that actually played out in the real world, tell me in the comments because that one deserves its own video.
So, let's take a look at what we built. We started with one server and one database. More people showed up than what one server could handle. So, we added servers and a load balancer in front of them. The servers could not remember who was logged in. So we moved sessions to a shared store. They were all hitting one database. So we pulled the connections and split reads across the replicas. The same expensive answers were being computed over and over again.
So we cached them and slow work was making users wait. So we moved it to a queue for a worker to handle in the background. That is a real system and it is genuinely how a huge number of apps are built and AI can build every one of these boxes for you. But knowing which ones you actually need and what they cost is the part that it cannot that is the job now. And that job working with AI on real systems making these calls running the whole workflow from scoping to architecture to review.
That is exactly what our new agentic engineering course teaches. Not just typing out prompts, but the actual judgment. We even built a skill called architect that walks you through exactly these decisions on your own project. I'll leave the link below in case you want to check it out. And if you're here for practicing for job interviews, anyone can name these things, but the one who gets hired says why they would pick one over the other and what it costs.
That is system design. Not the boxes, but the order and the trades we make. This was a new type of video that I try to do on the channel. So, if you'd like to see more videos on system design, let me know in the comments down below. That's it for this one, and I'll see you in the next one. Have a wonderful day.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.