Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

theSeniorDev · @therealseniordev
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
34:142.4x the video's typical replay level
you still have is that this page is not interactive yet because you haven't went through building the virtual DOM and attaching, for example, in the case of React, your virtual DOM to the actual DOM pre-built. And that's why you need to hydrate. And so, basically, you render that HTML page and then you have
Said at 34:08
Most replayed moment #2
32:432.4x the video's typical replay level
takes a lot of time. So, if you want to have a very performant website or a website that's ready to be crawled by a search engine, then client-side rendering is not the best choice for you. An alternative to this, if you have a static website where there's not a lot of interactivity, is to pre-render it on
Said at 32:36
Most replayed moment #3
35:222.1x the video's typical replay level
about data fetching and I know we are front-end engineers, but it's very important that you also can work at a data layer. As I said before, we are moving towards front-end engineers being more full stack. And that is server-sent events. And all this has to do with real-time communication. When it comes to real-time
Said at 35:15
The graph counts replays. It does not show where viewers stopped watching.
Words
7,923
Runtime
38:02
Speaking pace
208wpm
Reading time
33min
208 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
As AI is getting better at writing front-end code, your only choice to stay relevant and competitive as a front-end engineer is to level up your system design and architecture skills. The problem is that most system design content is way too focused on the back-end while completely ignoring the front-end. So today I'm going to break down every front-end system design concept you need to know from micro-front-ends architectures to system design patterns like back-end for front-end and then go into mono-repos and rendering strategies like server-side rendering. I'll also show you what is the best way to use AI and agentic coding
104 words, the words spoken in the first 30 seconds at 208 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 378 |
| Average words per sentence | 21.0 |
| Longest sentence | 111 words |
| Questions asked | 5 |
| Sentences containing a number | 6 |
Most used terms
Filler phrases
85 in total: basically 37 · like 21 · actually 10 · you know 9 · kind of 7 · uh 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
As AI is getting better at writing front-end code, your only choice to stay relevant and competitive as a front-end engineer is to level up your system design and architecture skills. The problem is that most system design content is way too focused on the back-end while completely ignoring the front-end. So today I'm going to break down every front-end system design concept you need to know from micro-front-ends architectures to system design patterns like back-end for front-end and then go into mono-repos and rendering strategies like server-side rendering.
I'll also show you what is the best way to use AI and agentic coding to deploy these systems to production. Now let's start with our first front-end system design concept, micro-front-ends. Now the traditional front-end application will start as a front-end monolith. That means pretty much everything is in the same place. And when it comes to micro-front-ends, it basically means that we will split this application into independent front-end applications and then put them all together.
But before we even talk about micro-front-ends, we need to understand microservices and that is a back-end architecture style. And so looking at our client-server model, both the server and the client will be monoliths. That is single applications. And at that point we keep adding features to it because you're in the initial phase of the project. But as you scale, more and more people will have to work on the same code base and your development team becomes huge.
And I've worked with developers that told me their development team was 40 people and their daily stand-up took 1 hour and a half just for everybody to speak for 1 minute and a half. And that's just not sustainable. So that's where microservices and micro-front-ends come into place because this kind of architecture doesn't scale from a development point of view. Now you've probably heard about the term two pizzas team where Jeff Bezos said that team shouldn't be bigger than five to nine people, which is as many as you can feed with one pizza.
Now the two pizzas team becomes a one pizza team with AI. Because nowadays you can do the same work with a lot less people using agentic coding. And so, the ideal size of the team becomes smaller and smaller. Now, do remember that even if coding agents like Cloud Code drive implementation time to zero, they do increase the verification time. So, being an AI savvy front-end engineers it means you design system that decrease the verification time and minimize architectural drift while they maximize velocity.
And so, it's not anymore about implementing fast, but rather building a system that is extremely easy to verify. And we'll see how that changes architecture later on. But, keep this principle in mind. And so, in the ideal case we are able to split our monolithic team into feature teams that own a specific feature. And those teams are full stack and they work with coding agents, but they can deliver independently and that's how we can keep our product moving forward.
And so, we'll have these different feature teams, but it's also shared infrastructure and that's owned by the platform team. You might have front-end tooling or pure infrastructures like the servers. So, you might have a team for example that is building a design system that the feature teams are using. And all those teams are smaller and they all work with coding agents. That is the new norm in development teams. Now, according to Conway's law, the software architecture of a system will look like the organization of the people.
So, if what we want is smaller independent teams, then we need to somehow break our system into small independent vertical slices. And this truly starts in the back-end where we take our monolith and we split out the independent modules to build microservices. Those are independent small back-ends, micro back-ends that expose their own API and can be deployed independently. The code base is independent and you have independent teams that don't need to talk to each other extending those and they only communicate between each other through APIs.
And so, a typical microservice blueprint will look like something like this. We have an API, there's a business logic layer, then there's a persistence layer and usually a database. In this case I added a Postgres DB, but it could be a no sequel DB. And so, this is the building unit of your application, and big applications will have thousands of those in productions. A couple of years ago, I used to work for this financial company, and we had around 1,000 microservices in production, and you had different teams only different microservices.
My team owned 13 of them. And by the way, if you want to see where do you stand across the full stack, there's a free assessment that you can take in the link below, and you basically understand how much of this full stack knowledge you actually have, where is the gap, and we added a new AI section because a lot of companies are starting to ask AI questions in front-end engineering interviews, also. So, take the assessment and you can see more or less where do you stand in the market and what gaps you need to close.
Link it's in the comments. Now, moving on, we split our back end into microservices, but our front end is still a monolith. So, even if our back end teams now could ship independently and be much smaller and go much faster, the front end is pretty much the same. And the problem with the front end is that you have so many people pushing code to the same client that it's so easy for someone to make a mistake or bring the whole system down or change a global CSS rule that then affects everybody.
So, as of now, our front end team is still very, very big and totally not sustainable. It gets to a point where we just cannot scale because there's just too much communication overhead between all those people needing to coordinate as they work on the same code base. And that's why we go back to micro front ends, where we have a front end monolith and we split it in different independent front end applications. Now, how do we put all those together?
Well, basically, we'll have a shell. And the micro front end shell is the one that takes care of the global functionality, things like auth, routing, language, global state, you know, are the users logged in or not, it's all handled by that. And all the others model applications, they live inside this one. They're basically loaded by this micro front end shell. And this micro front end shell, it's also on passing this global state inside those micro front ends.
But remember, those applications are now deployable independently. That means you might have a domain that it's header.theseniordev.com, where, for example, we would host only our header application, and you could load that independently. And then you have the product page and the cart page, but then you have the shell that puts them all together. So, if you look at a big software company like amazon.com, you could basically split the header, and then maybe take out the product page and the payment into two different micro front-ends.
But, to be honest, I feel like they have even more granularity there, and they probably split their front-ends into even smaller and more focused micro front-ends. So, once you apply micro front-ends and microservices together, you end up with a micro front-end that will speak to one or several microservices in the same domain. And this is what we call a vertical slice. And this system can now be extended by an independent team.
So, we can start releasing changes independently. Our releases are not a bottleneck for another team, and we can work against the APIs of other teams. This is how pretty much every big software company operates. And this is what in software architecture they call vertical slicing, where your features are vertically owned by one team. And I'm going to go even farther and tell you that if you are an engineer, and you're doing only front-end or back-end, you have to move as fast as you can towards becoming a vertically integrated engineer, basically a team of one that can work with a coding agent across the full stack.
And we'll see in a second why that happens. But, that's the transition you're going to need to go through to survive in today's market. So, when it comes to traditional front-end engineering and front-end development, we see that moving towards two extremes, where you're either more of a full-stack person that works in a feature team, so they work across these vertical slices, and that's why you need to know a bit more about full-stack.
By the way, we do have full-stack videos on this channel. If you want to know, as a front-end engineer, transition slowly towards a full-stack. Or, you're going to be really good at the front-end, and then you're going to the infra team, and you start building it at the shell, or you build a design system, and that's where you do spend a lot of time, you know, doing traditional front-end work with just CSS, and building the building blocks that the feature teams will implement.
And pretty much both of those people will work with coding agents. Now, a quick tip about AI and AI coding, micro front-ends and microservices decrease the blast radius of a production bug, and they could use the cognitive load of code changes because you have smaller and localized PRs. You also get less context, which means your AI coding tool would be much more efficient because you work on smaller pieces and the risk is less.
So, it's a really good strategy when you have a lot of people pushing a lot of code very frequently as we do with AI to move towards micro front-ends and microservices because the risk is lower and it decrease what we call the verification time. It's just easier to go on the quality when the surface area of the changes you are making is smaller. So, micro front-ends and microservices are usually the way to go if you're building fast with AI.
Now, let's move to our next concept, which is only API gateway. Now, whenever we have a micro front-end and a back-end service, we need to end up making a lot of requests. And those requests have what we call edge functions. They need caching, HTTPS, authentication, content negotiation, rate limiting for security, and so all this functionality at the API level it's pretty repetitive. So, having to implement this both on the front-end side and on the back-end side every time we plug in a new microservice to the micro front-end ends up adding a lot of overhead.
And it becomes a lot harder to do this if you have a lot of back-end services. So, a typical solution is to add what we call an API gateway, and that would be your gate into the back-end systems. And the advantage of an API gateway is that the client would only implement the edge functions once. So, they basically only implement the HTTPS handshake once, or caching, or rate limiting, and then once the request goes to the API gateway, it gets forwarded to the back-end services, which don't have to care about all this functionality.
They can just care about their own logic. And all this usually lives in a virtual private cloud, which means it's totally secure because the only way to go into it is through the API gateway. The other cool feature when you have an API gateway is that you can make the communication between the client and API gateway HTTPS, but the communication between the microservices and the API gateway HTTP because you don't need that security anymore cuz you're in a closed environment.
And the advantage here is performance because HTTPS it's likely less performant than the HTTP because you need more round trips to do the HTTPS handshake. So, in general, once you have a couple of microservices in production, it's a good idea to add an API gateway both for security and performance and also less complexity in the front end. Now, let's move on to our next pattern, very similar to the API gateway, we have the backend for frontend.
Now, let's remember our setup from before where we have our client and our API gateway and all our microservices. The problem here is that the client has to make a call to all these different microservices that might have a different API and there's so many fetch calls. And if you're in the frontend team, whenever you need a new feature that needs even the slightest backend change, you need to go and talk to that backend team and figure out if that will be a priority for them, add it to their backlog, and maybe something gets done.
So, it ends up being very, very slow. And one of the biggest challenge that some of the engineers we work with at the senior dev that work in bigger companies is that it's so hard for them to get things done because they have to talk to key different backend teams that have their own backlog and their own priority. So, they just don't want to implement that small API field that they need. So, better solution is to somehow own your backend as a frontend team, as a client team.
There's also a lot of complexity when you have to implement all these different APIs because they all look different, you need a different client and SDK for all these different APIs. And so, that's more frontend complexity. So, the solution for that is to add what we call a backend for frontend. This is a backend that will take the frontend request and then just forward it to the microservices. But the advantage here is that whenever you have to implement a feature, you can act as a full-stack developer when you're in the front-end team.
So, basically, you have all those microservices, you integrate, like to change the way you equals them from the back-end for front-end, and then you implement your feature, and you kind of own end-to-end the client-side feature. And the back-end teams, they can just work in isolation on the different microservices. So, the front-end team ends up owning both the client and the back-end for front-end, and back-end teams they develop pure back-end services.
And this is very important for you as a front-end engineer because it means that you do need to know, at least at the high level, how to extend a microservice, how to build a back-end for front-end. You need to know API design, and you need to know a little bit about GraphQL, for example, which is a great technology to build back-end for front-ends. This is a requirement for all the front-end engineers that you see working at bigger companies, which are usually the ones that also have the best conditions and the most exciting work.
So, even if you're in the front-end, make sure that you can also get things done in the back-end. Now, the advantage with BFFs is that you can adapt them to a specific client. So, if in the future we end up building a mobile app that has totally different requirements to our desktop app, we don't need to duplicate API endpoints. So, with back-end for front-ends, every client will have its own dedicated back-end, which means every client team can work independently and is completely decoupled from the back-end services that can just focus on their own service.
And so, basically, whenever the desktop client, for example, goes to gateway, they get redirected to the desktop back-end for front-end. And for mobile, we have a completely different API, a completely different mobile back-end for front-end. It uses the same back-end services, but it might expose a total different API to consume them, just because the data needs in mobile are usually different from desktop. Where in desktop, you want to fetch a lot of information at once, but in mobiles, we have smaller screens, so you need different records.
And if you try to put all those things together in a single API, it will become bloated, or it will end up not satisfying one of the requests. For example, you'll have a great API for desktop, but when it comes to mobile, you'll fetch too much data, a lot of data you don't really need because the interface is a bit different. Or if you make the smaller endpoints, when it comes to desktop, you'll have to make many requests to get the same data.
So, you'll have over-fetching or under-fetching, and it's a lot better if you split those things. On to the next concept, which is load balancing. Now, going back to our client-server model, we had a server and we end up having clients. But, the problem is we might get a lot of users. And so, you have all those clients making requests to a single server. Now, a typical Node.js server is pretty powerful. It can usually satisfy up to 2,000 to 10,000 concurrent requests with a well-optimized Node.js server.
But, when you go beyond that, you might need different ways to scale it because a single server has its own limits. And the easiest way to scale it is by load balancing. And that means you basically will create identical instances of your servers and then have this component that's an application load balancer splitting traffic between them. Now, load balancing is a topic by itself. There's different ways to split traffic and different criterias.
And as a front-end engineer, you don't need to go so deep into it. But, do make sure that you know about it. Ideally, you're even capable of setting up a small load balancer using traditional web servers like Nginx and Docker Compose, you can have this set up on your local machine. Anyhow, the important thing is that you can reason your way through it. Most cloud providers like AWS or Google Cloud, they allow you to provision a load balancer in seconds.
So, don't worry about it. It's very uncommon that you need to manually set up one, but it's important that you know about it. Now, congrats on making it this far. Make sure you subscribe so you don't lose any updates in the future, and let's move on to our next concept, which is container systems. So, we basically distributed our architecture into micro front-ends and different microservices, but the problem is that deploying all this would be a headache.
Now we need to build up and provision infrastructure and pipelines, and they all might use different technologies. You might have a Python microservice and then a Node.js one. You might have a Vue.js application or a Next.js application. And this is all very complicated to get to production. So, in order to standardize the deployment, we can actually use Docker. And Docker is technology that allows you to package your application to Docker image.
And the way you do that is that you take your code, and then you have a Dockerfile where you kind of declare your recipe of how we should package that code. And basically, based on that, you'll create a Docker image. The Docker image will contain all the application code and then the runtime. Let's imagine you're using Next.js, then it will contain Next.js, will have the runtime, which is Node, and then the operating system, which is usually Linux.
And that is a full Docker image, and the advantage there is that whenever you find a host that runs Docker, you can run that image. You don't need to worry about the Node.js version or if they need to install PHP and all the dependencies that your application has. It really comes packaged all together. This is Docker image. You run it, open this port, you have a front end running. You don't need to know about what's inside.
And this is wonderful for DevOps teams because all of the sudden they can take all those images and push them into a container orchestration system. A container orchestration system usually has a container deployment pipeline, which will run these containers. And so, running a container is not as easy as it sounds. You might need load balancing. Uh you might want to run several instances in parallel and be able to, you know, if a container fails, spin up another one really fast.
So, all this kind of hard work it's built into systems like Kubernetes. You probably saw it in job office. Now, do you need to know Kubernetes as a front end engineer? No. But you do need to explain, you do need to know about it, and you will see a lot of front end positions that mention either Kubernetes or container systems or a ECS, which is the AWS alternative to Kubernetes, in the job description. Don't be scared.
You don't need to become a DevOps engineer by tomorrow, but you should be able to, at a high level, understand where it fits in your architecture. Finally, a more front-end concept, a CDN, a content delivery network. So, what is a CDN? So, going back to a client-server model, imagine you are a client, you want to go on a website, you usually go to the server and get some static files. Static JavaScript and CSS, and you download that and you run that on your web browser.
Let's imagine in a more hypothetical case that you are a user from the US and you want to visit the application that is based in Europe. For you to get the JavaScript and CSS and all the HTML, you have to go all the way to the Atlantic Ocean and come back. And that young trip adds latency, and there's no physical way you can work around it. No matter how performant your application is, there is the limit of the speed of light, because data only travels as fast as the speed of light, and over long distances, light is very fast, but it still will add around, let's say, 200 ms to 250 ms on every request of latency.
So, a solution to that is to your client decrease the distance between you and where those files are hosted. And the easiest way is to use a content delivery network. And so, basically, a CDN will be a network of these edge locations that are placed all over the world, and what happens is that the server will push the static assets there, and whenever you make a request, you'll be redirected to the edge location that is closest to you.
Now, this mechanism on how exactly are you getting redirected to that edge location that is closest to you, it's very interesting, and I do want to make a video about it, but I'm not sure if it's something you're interested in. If you want a video about it, let me know in the comments. It's a bit more technical, and it's not something you'll get in interviews. But, to be honest, it's really interesting. So, if you want me to make a video about it, just give me an excuse by letting me know in the comments, and I'll go ahead and do that.
Now, a CDN is a distributed cache. So, the technical term for whenever you get the asset from the CDN, it's a cache hit. And whenever the server pushes a new version, that's called cache invalidation. And this concept is very closely related to what we call cache busting, which is a mechanism that module bundles use to invalidate your assets. So, we make sure that when you deploy a new version, users really get the latest one, not a previous version that is probably still sitting somewhere in the CDN.
I do have other videos in the channel talking about it, so I'm not going to go deep into it. But make sure you're able to relate those things across the stack. Now, keep in mind that a CDN is the fastest and most cost-effective to increase web performance by serving optimized assets with the right cache policy out of the box. Because nowadays, CDNs do a lot more than just placing the asset close to the client. They also compress it, and they also take care of the caching policy.
So, it's a really cheap way for you to fix most performance problems. Now, let's talk about design systems. And going back to our micro front-end discussions, we have a feature team, that's a product feature team, and then we might have the payments feature team, and they work separately and release independently. Those are completely separate product teams. And the problem there is that you might get code repetition or what we call visual divergence.
You can have this silos mentality. And so, basically, a button in the product page will look differently than a button on the payment page. And so, you will lose what we call visual coherence, and people will start to notice that those are actually separate front-ends because they look differently. And that's not good from a product perspective, but it's also not good from a technical perspective because you have too much repeated code.
And so, the solution is to have a design system. And in a design system, the first thing you do is to define your design tokens. That's basically your team information. What's your primary color, what is your border definition, your fonts families, and so on and so forth. And the modern way to do that is you add them as CSS custom properties, basically variables, at the global level in your CSS, probably in the shell micro front end, and then it gets fed into all the other micro front ends.
But, you can go a bit farther and build reusable components that then the feature teams can use to assemble their features. So, you could build inputs and buttons, and basically they would consume that UI package and use it in all these different micro front ends. So, basically what you achieve is that the UI looks consistent, but also you don't repeat your code. The cool thing about the design system is you can take care of accessibility, you can have all those components unit tested, and basically you are applying at an architectural level the do not repeat yourself dry principle.
And going back to AI coding, a solid design system really makes the difference between generating some component slob and having inconsistent styles and bugs that need a lot of rework, and really having a consistent, reliable coding agent output. Trust me, I've been building a lot of AI in the last couple of months, and the first thing I do when I make a new project is to tell to extract from whatever the design is the design system, because then you feed that into your coding agent through different sessions, and you still get consistent output.
If you don't do that, the agent will make things up, and your UI would just look different, and it'll be obvious that it was live coded. Now, our next concept is a design to code MCP, that's a model context protocol server. And so, basically, you remember we have our design system, and normally you would import that to start building with it. But, nowadays, let's imagine that your designers built your design system initially in Figma, and then you implement it in your library.
You can use the Figma MCP server with a coding agent to very quickly assemble features into the feature teams, into the vertical slices of your product. And this is the kind of workflow that most companies are moving towards. So, if you are a front end engineer, you got to make sure that you know how to use an MCP server, you know what an MCP server is, and ideally, whatever design tool your team is using, you can plug an MCP into that or you might have to even build that connection yourself.
And then plug that into your AI coding agent like Cloud Code, which is the most used in enterprise or Codex. Now, a quick note on CSS architecture and design tokens. What you can do to take things even further is to apply the Atomic CSS methodology and with your design tokens, create atomic classes that you can use in your code base. So, basically, you define what the border would be, but then you create a dot border class that has that property.
And then the only thing that other developers have to do is to use that class. And this is how Tailwind CSS works. So, keep this in mind. Tailwind CSS is an implementation of the Atomic CSS architecture style for CSS. And there's three more architectural styles, which I won't get into right now, but who knows, maybe I'll make a video about it later on. Next front end system design concept, the monorepo. So, basically, our applications right now are sitting in all these different repos where you have a GitHub repo for the shell, one for the payment micro front end, inventory micro front end, and it's all spread out into thousands of repos.
And the problem there is that you end up with different code styles, different dependencies, and different quality standards. Some people might be using TypeScript, they might be using different linter configurations, and it's so easy to have what we call architectural drift or code style drift where two projects diverge too much. What's the problem with that? If you're a developer and you change teams, you have to relearn everything.
So, you don't really leverage standardization. And the way to is to put everything into a single repository. So, basically, you have a big repository that contains all the other small applications, and you have tooling that works with both of them. So, for example, when you run NPM run build in a monorepo, all your applications will build individually. Now, when it comes to AI coding, a monorepo gives coding agents the context they need to make changes across service boundaries.
So, for example, if you're building a micro front and you realize that you run into a reusable use case. You can easily extract that and propose it as a component into your design system. That might be a different repo. And you can do that with the coding agent in a single session if you have everything in a mono repo. If not, you have to somehow feed those two different repositories to your coding agent and everything becomes harder.
So, combining micro frontends and microservices with a mono repo, it's usually the most efficient way to work with AI. Now, one thing I'm really excited about is also MC PUI. So, MC PUI is basically combining the traditional website with an LLM-powered app, and that's basically the glue in between them, the duct tape, let's say that. And so, basically, in a chat application, you usually send a chat query, and then that goes to the LLM, and the LLM will answer back, but in this case, it can actually answer with UI components.
It can actually render products, for example, in the answer, not only text. But to do that, it has to somehow talk to the backend, and then we need to render some components. And that is what MC PUI solves. And this is an example I found in a recent website where I was looking for venues for events because we are organizing our annual in-person meetups at the Senior Dev, so we're going to meet all the engineers we work with around Europe or in the US, and I was looking for different cool venues where we could actually meet up.
And so, I was talking to this chat UI, and all of a sudden, after a couple of questions, it started rendering those venues, and it's asking me for input, and then it would go into a deep search and give me even more venues. And as you see there, it can even render a map. And all this happens in a chat application. And I think this is the direction front end will go with LLMs, where we will integrate LLMs in web applications would be.
A lot of people were saying that there's no more need for UI now that we have a chat. I disagree. I think the UI is a very useful way to communicate things, and I think you need front end developers, but we'll be able to combine the LLM approach with the traditional web approach. So, this component was rendered because the front end parsed it from the answers of the LLM. At a high level, the way this works is that the model harness will provide in the context the tool registries and all the MCP servers and then as a user prompt, and the LLM will send an answer back to the tool UI with the text answer, but also with an instruction to render a certain div, and then the tool, the web application has to follow that and render it.
And just to really bring this to the code, the way we declare a resource or an MCP UI is we basically tell the LLM, "Hey, there's this resource, and you can use it in this case scenario, and this is the answer you can give." And then our front end has to take that and parse that and then render it. If you want me to make an in-depth video on how exactly MCP UI works, let me know, but this is one of those patterns that if you know really well, you'll really stand out because I believe it goes way beyond the current AI hype and it's actually something very useful for specific edge cases, and you'll see a lot of applications actually implementing this hybrid solutions.
Now, to wrap up this talk a bit about performance, and the most important pattern that you'll ever see out mentioned in job descriptions for front end engineers, it's the Core Web Vitals. And so basically the Core Web Vitals are three metrics that quantify the three dimensions in which we measure the performance of a website, and those are the loading speed, the interactivity speed, and the visual stability. And the three Core Web Vitals that measure this are the Largest Contentful Paint, the Interaction to Next Paint, and the Cumulative Layout Shift.
Basically, those are measured by Google, and they tell us what a good number would be. So, when you look at the Largest Contentful Paint, it would be the time it takes from when you hit enter to when the largest element in the web page is rendered. The Interaction to Next Paint, it's slightly different, and it has to do to when you interact. So, you do something, and then when do we repaint the UI? And the CLS is basically how much the UI changes when it loads.
And so, the LCP and CLS have to do with the initial render, and the INP has to do with the re-renderings, which is very important when you talk about component frameworks. Now, to understand those, you need to understand the critical rendering path, which is all the steps you go from downloading some HTML to actually showing something on the screen. And that involves building the DOM, and then building the CSSOM, then building a render tree, then computing the layout tree, which is basically a tree of where all your nodes are and the positions and the width, and then transforming that into what we call paint operations that goes through to the GPU, then going to the composite phase, which has its own complexity, and I'm not going to talk about it, but just to summarize this, all those steps will happen, and then you have some re-renders because we use component frameworks, you get data, you start re-rendering, you finally finish, and then that's when you paint the LCP.
So, all that time is measured in the LCP, and the bottom line here is, if you're shipping a lot of JavaScript, if you're shipping a lot of CSS, if you have to fetch a lot of data, and your server is slow, your application will be slow, and you'll get a very poor score in the LCP. The other thing that will happen if your application is poorly optimized is that you'll have layout switching all the time as you start loading things because CSS comes too late, and then fonts come in, and then some data comes in, and the browser will take screenshots of that and try to figure out if you are moving things too much.
That would be the cumulative layout shift. And finally, the interaction to next paint is basically whenever you have a user event, and you have to go to the reconciliation and re-render after state update in your framework, and then you start again, you modify the DOM, and that triggers a re-paint. So, all that time it's quantified as the interaction to next paint. So, basically, if you have very slow re-renders or you're re-rendering too many components when users do something, then you have a very slow IMP.
You can measure those metrics like house, and they go a lot deeper into this video on this channel about it, so make sure you check out that one. Our next concept also has to do with performance, and it's called splitting. Now, traditionally, a module bundler will take all our JavaScript and put it together in a single big file. But loading that single file will totally mess our core web vitals because we load too much JavaScript.
So code splitting allows us to split our JavaScript to where it's being needed so we can really ship only the JavaScript that's needed to a specific page. And the easiest way to code split is by out. So basically, you would only ship to the slash login page the components that are specifically needed for the login. And if you have a dashboard that's very heavy with a lot of graphs, for example, you don't ship all that.
So you selectively ship your JavaScript to where it's being needed instead of putting it together all in a single file. And all this is achieved with a module bundler like Webpack or Vit, which understands your bundle, splits it, and then dynamically loads it based on the path you're on, working together with your application router. And the higher overarching mental model here is lazy loading, which is the opposite of eager loading.
Eager loading means you are really loading everything when a user lands on the page. And lazy loading means you load things as they needed. So for example, you might load certain things on scroll, or you might load certain things when they visit a certain page, or you might load certain things when they click on something. So you kind of wait for that user interaction, and when they land on the page, you really only ship the things they need.
And all this is done in order to make this core web vitals better and working across the critical learning path. And finally, let's talk about rendering strategies. And nowadays, we usually work with modern component frameworks like React, Vue, and Angular. And the problem with those is that when you land on the page, you see a white screen. And the reason for that is because, you know, you load that empty HTML, and until you don't really run whatever render function they have, you don't really see much on the screen.
That's what they call client-side rendering. So in an SPA architecture, in a single-page application with client-side rendering, you'd go, get your static file, and then you need to go get some dynamic data, and then finally render. And all that takes a lot of time. So, if you want to have a very performant website or a website that's ready to be crawled by a search engine, then client-side rendering is not the best choice for you.
An alternative to this, if you have a static website where there's not a lot of interactivity, is to pre-render it on the server and ship it already rendered. So, when your client goes to get the static files, they already get HTML CSS and they don't need to run so much JavaScript. Now, again, this only works with static sites. If your static site will change often, let's say you have a blog and you want to publish new articles, then what you can do is incremental static generation.
That means you only regenerate the pages that have changed. So, basically, your CMS will trigger a rebuild when you add a new blog post and that will get the dynamic data and go to a build pipeline and regenerate the only portion of the static files that change. So, it's a bit of a partial rebuilding of the website. The advantage here is not only that makes the build faster, but if I'm a client and I already downloaded part of your CSS and JavaScript, but I don't need to download re-download only the parts that change.
Now, in most cases, this is built into a framework like Next.js, so you never have to worry about this yourself. And finally, we have server-side rendering. And in server-side rendering, the client will make a request to the front-end server, but then the front-end server will request the back-end server and get some data and then render the application on the server and then send it back pre-rendered. So, the client receives a full HTML page.
So, you don't have this problem of the white screen. The problem you still have is that this page is not interactive yet because you haven't went through building the virtual DOM and attaching, for example, in the case of React, your virtual DOM to the actual DOM pre-built. And that's why you need to hydrate. And so, basically, you render that HTML page and then you have to execute your JavaScript, create internally the virtual DOM and then that virtual DOM is attached to the existing HTML markdown.
That's what we call hydration. And finally, when you're hydrated, you might have to do some extra data fetching, so you might still need to go to the back. And again, this is one of the most complex approaches, so be careful with it. It's only useful whenever you need very, very fast performance or you need SEO. And a lot of people and a lot of companies jumped into this and they're doing server-side rendering, but that's like building an F1 car to go to the groceries.
It's over-engineering and it creates a lot of problems and then everything gets slower. There's so many issues that will appear when you have this kind of setup. Technology to make it work is very complex. Bugs are harder to solve, so you want to stay away from it. And I really prefer simple solutions unless your use case really needs the highest performance. Oh, finally, let's talk about data fetching and I know we are front-end engineers, but it's very important that you also can work at a data layer.
As I said before, we are moving towards front-end engineers being more full stack. And that is server-sent events. And all this has to do with real-time communication. When it comes to real-time communication that is not following the request-response cycle that I showed until now, you basically have to give it to it. You can do polling, you can do web sockets or server-sent events. And so polling would mean that you keep calling a specific endpoint until something happens.
So, let's say we have a transaction that's processing, I can keep calling the status endpoint until it becomes completed. That's very easy to do with plain JavaScript and set timeout. There's nothing complex about it. The problem is you're calling your server way too much. And so it doesn't scale really well. You can have race conditions. So, it's a very simple solution, but not the most performant one. The next alternative would be web sockets, where you open a channel between the server and the browser and you can send updates and they can send you updates back.
The problem here is that again, it's a lot of overhead and it's extremely good whenever you have bi-directional communication. Like you're working for a chat application, for example, where both the client and the server would keep sending chunks of messages. Now, for most use cases, this is not really needed and again, it's complex and it's very intense on the server. It needs a lot of resources. Now, with AI, we do need real-time communication, but it's only one direction.
Because usually when you send a query to a chat application, you then just wait and they start sending you tokens back. So, it's the server sending a lot of messages, but you usually only send one. So, there's a lot of asymmetry between the client and the server. And the way to make this happen is with server-sent events, where you send a text message and that will create a conversation and on that endpoint, you're able to receive updates from the server.
So, you receive all these tokens, but you don't send so many. And this approach is the one used by most LLM applications. If you've ever used the OpenAI NPM package in a React application to build a chat app with an LLM, that's exactly what they use under the hood. And you can actually figure this out if you go to your network when you use ChatGPT or Claude and find that conversation request and you'll see that the answer to that it's all these event streams.
So, you're getting chunks of the answer. And that's how you build these cool UIs for LLMs. It's not WebSockets and it's not pulling, it's the server-sent event API. Make sure you look it up because it is something that will help you build product software products, software applications with AI embedded. Thanks so much, folks. If you want to go even deeper, make sure you check out these two videos on our channel. One of them is about all the front-end architectures that you need to know as a front-end engineer and the other one is about in-depth concepts about web performance, CSS, component frameworks that you really need and that show up in interviews at the senior level.
And I'll see you in the next one.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.