Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

NVIDIA Game Developer · @NVIDIAGameDeveloper
Where viewers went back to watch this video again, from YouTube's public Most replayed graph, lined up with what was said at that moment.
Most replayed moment #1
45:166.5x the video's typical replay level
from before. It has one bounce of GI and then falls back to sampling the environment map. What we would really want to do here is to have multiple levels of multiple bounces of global illumination, but that would multiply the cost, and that would be prohibited for most games. Instead, with SHARK, we will dispatch a
Said at 45:09
Most replayed moment #2
35:446.4x the video's typical replay level
screen space ambient occlusion, looks like in our sample project. Uh it is a straight out of the Jimenez paper, and it is very similar to what many games already have. And this is the same scene with RTAO on it. It is exactly the same denoiser. I just
Said at 35:37
Most replayed moment #3
15:296.4x the video's typical replay level
Okay, now let me show you our most recent work based on mega geometry which is a level of detail system for foliage. Uh let's start by seeing that in action. So what you're seeing here is a scene that consists about consists of about 60
Said at 15:22
The graph counts replays. It does not show where viewers stopped watching.
Words
10,797
Runtime
59:34
Speaking pace
181wpm
Reading time
45min
181 words per minute, the same as the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Uh welcome. My name is Martin. Uh I lead the uh ray tracing team at Nvidia. This is my colleague Carmelo uh from the DevTech team. Uh and we'll talk about path tracing. So, path tracing has become uh the state of the art for graphics in games. Since we introduced RTX in 2018, uh over 200 games have shipped with ray tracing and about 15 or so with path tracing now. Uh and we've got more big titles coming this year. Uh the example screenshots I have here are from
91 words, the words spoken in the first 30 seconds at 181 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 590 |
| Average words per sentence | 18.3 |
| Longest sentence | 86 words |
| Questions asked | 40 |
| Sentences containing a number | 35 |
Most used terms
Filler phrases
636 in total: uh 379 · like 65 · um 51 · you know 47 · actually 39 · right? 23 · sort of 12 · basically 10 · kind of 9 · I mean 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Uh welcome. My name is Martin. Uh I lead the uh ray tracing team at Nvidia. This is my colleague Carmelo uh from the DevTech team. Uh and we'll talk about path tracing. So, path tracing has become uh the state of the art for graphics in games. Since we introduced RTX in 2018, uh over 200 games have shipped with ray tracing and about 15 or so with path tracing now. Uh and we've got more big titles coming this year. Uh the example screenshots I have here are from Black Myth: Wukong uh and Resident Evil that just came out with really beautiful lighting.
Uh the purpose of this talk will be to focus on some of the key ingredients that we believe uh are what make a great state-of-the-art path and in the foreseeable future. Uh I'll cover some of the core technologies uh that go into this uh and then Carmelo will cover uh an overview of the light transport algorithms that are involved. Uh we'll start with a quick mention of DXR 1.2 which Microsoft just released 2 weeks ago or so.
Uh this has been in preview for about a year since last GDC. Uh and it standardizes two very important features for path tracing performance uh that we introduced with Ada. Um the first is opacity micro maps or OMMs which are basically hardware accelerated uh is basically hardware accelerated alpha testing for ray tracing uh to reduce the overhead that you get with any hit shader execution. Uh there's a lot of material on this topic out there.
We have blog posts and talks and all sorts of things. So, we won't spend too much time uh on it today. Uh the second feature is shader execution reordering or SER uh which is really the tool to tackle the divergence that is inherent in pretty much any path tracing workload. Uh every or nearly every game every path traced game today uh uses SER through the NVAPI uh and we see speed ups in the range from 20-30% sort of on the lower end all the way up to about 3x uh on the path tracing regime.
So, it's uh really uh you know can be a very high impact feature uh and in some cases these speed ups are achieved with really trivial minimal code changes. Uh so, I'd like to thank Microsoft for the collaboration on getting these features standardized. Uh and we'll look at SER in a bit more detail. Um so, we'll do a super quick recap of what the feature does and what it looks like for those who haven't looked at this at all.
Uh and the API is actually really simple. Uh first, it lets you split your uh traditional trace ray call uh into this uh what I have in pseudo code here, these these two calls where you use a different call to uh trace the actual ray, uh store the result in a hit object, and then uh call a new function that we call invoke uh to actually trigger the shading uh the the closest hit or miss shader. Uh and this sequence you see here behaves exactly like the traditional uh good old trace ray, right?
So, what is the point of that? Well, the point is that we can now insert in the middle the real magic which is uh reorder thread. Um and what that does is actually pretty remarkable uh because it shuffles around uh the physical threads on the machine on the GPU uh with respect to the hit object that we gave it. Um and so, after the function returns, uh it uh you know your warp uh that comes back is more likely to have neighbors uh neighboring threads that have the same intentions as uh the current thread that you're that we're looking at.
Uh and so, once we uh invoke the hit or or execute the closest hit shader after that, uh we're likely to do that much more efficiently because we've now in other words extracted coherence and reduced divergence. This is the basic idea. It's the basic pattern, right? It's a very powerful concept uh that's expressed in a very simple API. Um now looking at this sequence, you might wonder why the original trace ray didn't just do this under the hood.
Uh and that would have been possible or it is possible. The the current APIs uh actually allow for that. Uh we designed it this way. But, uh it turns out that by splitting up these calls, splitting the trace ray function into three different calls, uh this is a is a better sweet spot for API abstraction level, right? So, uh it gives uh gives us the right mix of application flexibility. It lets the application uh inject app site knowledge into the reordering decisions and into the scheduling decisions that the driver or the implementation uh can possibly have.
Uh and so, we have this this better balance of application side flexibility uh combined with uh the the implementation or the hardware vendor actually being optimized uh being able to optimize uh the reordering under the hood. And we do that constantly, right? So, we uh as an example in Blackwell, uh we were able to increase uh the granularity uh or the the precision of the reorder um to give a uh uh you know better coherence extraction and therefore lead to more performance.
So again, this is the basic pattern uh that you can get started with an SER, but let's look at some more interesting examples uh and also at the same time address some of the common misconceptions uh that we see uh with SER. Uh and we'll start with uh first of all, you don't have to choose between SER and ray queries. That's a common one that we see. Uh in fact, they go very well together, right? As you see here, you can actually take a ray query.
So, imagine there's a ray query uh before this code and turn that into a hit object and then use that hit object to drive the reordering. Uh we actually recommend doing that in in in many cases. Uh ray queries are great for things like shadow rays or simple visibility rays uh and that can be a a very useful uh sort of pattern. Uh the only requirement here is that this happens in a ray gen shader uh not a compute shader because compute shaders don't support the reordering.
Uh the second important concept is reorder hints. Uh this is the main mechanism by which the application uh communicates its intention uh to the reordering implementation. So, uh you know a hint is essentially a couple of bits that are encoded in uh a uint as you see here uh that is then passed in addition to the hit object to the reorder call. Uh and the information stored in the hints and the hit object get sort of mixed together.
You can imagine it like a sort key that gets combined from the hints and the the hit object into the sort key that actually drives the reordering. Uh one of the most effective uh ways to use these hints is basically this example here. If you imagine this code is in a in your main path tracing loop, right? Where we bounce the rays around, um it's very useful works in almost every path tracer to encode the loop exit condition uh into a hint to make sure that the the warps that uh exit the loop uh agree on the all the threads in the warp agree on exiting the loop together and all the warps all the threads in the warps that stay in the loop uh are coherent as well.
Uh that can give some very nice speed up and it's a very uh simple uh simple optimization. So, if there's one SER optimization takeaway from this talk, this should be the one you you remember. Um reorder hints can be useful for other things too like anything that you can predict after a reorder any code flow. Uh for example, we have we often see closest hit shaders making you know having branches based on whether the material is emissive or has a clear coat or something like that.
Uh if there is a possibility to to encode that into a hint before the shader is executed, that can be very useful as well. Uh and then uh third, we have the uh the power of having these uh three calls separate. Uh the trace reorder invoke separately is that you can arrange them any way you want uh and any any way you see fit. Uh you're not required to call all three of them. For example, uh it it can be useful to to leave out some or or rearrange them anyway.
So, in this example, we're not actually calling reorder as we're not actually calling invoke, but we're inlining uh the shading code into the ray gen program. And that can be you know can can increase coherence further, can reduce instruction cache pressure, uh and depending on the material system or shading system that your engine has, uh that can be the right thing to do. We can also vary uh vary this or or mix and match things.
Like in this example, we've uh you know pulled some of the common code out of a closest hit shader into like a common light evaluation. So, if you have a you know closest hit shader that at the end of each closest hit shader, you sort of execute the same light lighting code, but the material uh above that is different, uh it can make can make sense to pull that out into the ray gen shader again with the same goals as before.
And this exact pattern actually we've seen in in you know real titles shipping uh giving giving nice speed ups. Uh there's interesting things you can do with the hit object itself. Uh you can point it at any uh shader table entry with custom logic other than the you know unlike the the fixed function logic that's currently in the APIs. And then call invoke once or maybe even twice. There's you know anything goes. Uh whatever is is useful.
Uh what's not shown here is you can also call reorder thread without uh tracing any rays at all. Uh there's a variant of reorder thread that takes just a hint. Uh takes up to bits on Blackwell and no hit object. Uh and so, that let lets you reorder threads based on you know whatever you happen to have whether it's a ray hit or something else in a single line of code. Uh and unfortunately, we don't have time to go into much more detail about SER.
I just wanted to get these you know couple of ideas out. There's much more to say about all this uh especially optimization techniques and so forth. Uh I'll point to the talk that Louis Balval gave yesterday uh in case you missed it, it's going to be on on YouTube soon. He had a great case study with more optimization details and ideas and real world uh uh case studies in Doom the Dark Ages. Uh overall, the point I was trying to make here is you know, be creative in exploring SER instead of just using this initial pattern we saw.
Experiment with it. Every application is a little bit different. But the API does make it really easy to try out different strategies and see what works. Okay, let's switch gears to a different technology that is a bit more recent but we think no less groundbreaking and that's mega geometry. Uh we introduced this about a year ago uh and it completely really changes how you can think about the limitations of geometry and BVHs in ray tracing.
Uh mega geometry consists of two orthogonal features. One is clusters where the input into a BVH build is no longer a list of individual triangles but is now a list of clusters where each cluster is you know, a local group of about 100 to 200 or so triangles uh which immediately reduces the the load on the BVH build by the order of 100x which makes makes it obviously a lot faster, right? Uh there's a concept inside of the clusters feature that that we call templates uh that is uh given given a preselected topology you can create and instantiate clusters just by inputting vertices.
It's a little bit like a BVH refit. Uh and we can do this at extremely high speeds. So this runs at nearly memory bandwidth performance. I told on a 5090 we get well over a terabyte a second throughput by this which translates to about 30 to 40 billion triangles a second that we can build these clusters from. Uh so this is amazing for all sorts of tessellation use cases in animation. Uh the second feature is the partitioned TLAS or partitioned top level acceleration structure.
This addresses a common bottleneck where your top level acceleration structure once you feed it with too many objects and too many instances the build of that TLAS can quickly become the bottleneck. And that's because it can't exploit the fact that most of the scene is actually static most of the time right it always builds the entire structure from scratch and partitioned TLAS addresses that. Uh what it is is an is an persistent data structure that can be incrementally updated.
Uh and the way it works is that the application assigns partition IDs to all the instances in the scene ideally in a way that has little spatial overlap to keep trace performance high. Um To a good you know, good best practice is to aim for something like a couple hundred to a thousand or so instances per partition. Uh and then as the PTLAS gets built really only partitions that had any changes you know, objects removed or or or added or or moved uh only those have to be built and everything else is just not touched.
And this saves a ton of time making the build much much faster if not everything in your scene moves and this way we can support millions of objects in a scene. Um what both of these APIs have in common is that they're both indirect and batched. So batch meaning they just operate on many things at once with a single call and indirect meaning they get their arguments from GPU memory so all of this can be driven by the GPU with extremely little CPU overhead.
You can basically manage the BVHs for your entire scene with a handful of actual calls CPU side into the driver. Uh these APIs are available today through NVAPI and Vulcan extensions but we're actively working with Microsoft on standardizing this technology in in DX. This is making very good progress. The the work in progress spec is actually public since yesterday uh and the preview is is aimed for late summer or so this year and actually retail release end of the year.
So if you're interested interested in this I recommend you check out the the work in progress spec. Here's an example of mega geometry in action. This is a demo that you might have seen last year. It [clears throat] shows how how mega geometry can be used in Unreal Engine to implement fully ray traced Nanite geometry uh and integrated with Nanite's streaming system in the engine. Uh Nanite requires thousands many thousands of BVH builds per frame at you know, high high polygon counts and this would have been completely infeasible with with the previous APIs.
Here's another example. This is something that we really look forward to to seeing in games and that is clusters used for dynamic tessellation of objects. Um Here in this example we path trace real time displaced subdivision surfaces where every patch that you see on this dragon is dynamically tessellated and displaced. Every single frame the the object here is or the scene is tens of millions of triangles and we rebuild the BVHs for the entire thing from scratch every frame.
Uh this is using the cluster templates that I mentioned earlier to make these builds super fast and you can actually find this sample on GitHub not with this not with these exact assets cuz they're they're huge but uh it's it's the exact same code. Okay, now let me show you our most recent work based on mega geometry which is a level of detail system for foliage. Uh let's start by seeing that in action. So what you're seeing here is a scene that consists about consists of about 60 million plants with about 1 million trees and over 200 different species of plants.
Uh the terrain is about 5 by 5 km and there's no streaming going on. All of this is in memory. Uh you don't really see any popping or the other you know, typical LOD issues uh and everything in the scene is modeled as geometry down to individual pine needles. There's no alpha maps or cards or anything like that. Um the larger trees here like this one have up to or or over 10 million polygons in some cases. Uh my favorite statistic about it is if you flatten the entire scene into a triangle list at full LOD you would get over 5 trillion triangles out of it.
Um we have fully dynamic lighting of course this is path traced uh pixel perfect shadows. You saw everything can be uniquely animated. You saw the the spaceship earlier sort of casting a or blowing wind over over trees. Uh and what we're targeting here is a level of performance and also memory requirements but also authoring workflow you know, choices that is really usable in real world games not just as a tech demo.
Of course here we cranked up all the settings and everything or all the you know, environment but uh the basic system is really meant for real world game engines and that is important because this tech is coming to The Witcher 4. Uh so I want to give a big shout out and thank you to CD Projekt Red who also did provide the tree assets that you're seeing for us to work with. Um all of this is built on the existing mega geometry APIs that I mentioned.
There's no new APIs there's no new hardware required and we will open source this later this year and publish you know, publish much more detail about it in the coming months. Uh we will take a quick peek under the hood just to get a basic understanding of how things work. So in order to get these large amounts of geometric detail into the trees I mentioned 10 million polygons or so per tree uh they're modeled by using smaller objects like these twigs and instancing them many times across the tree.
So each of tree will have like a couple hundred or few thousand instances. Uh and this makes it very memory efficient. It's actually the same approach that Epic uses for their Nanite foliage system as well. Uh and we have you know, typically about a dozen or so twigs individual pieces per larger tree. Uh for the animation we have a skeleton. Each twig or mesh is attached to the bones of that skeleton in a rigid way. It could conceivably be vertex skinned as well for the very near close ups but we haven't actually found that necessary at this at these small levels of you know, small objects that we have with these twigs.
Um now if we look at you know, think through this a bit we have I said about a million trees right each tree has maybe a thousand instances. Uh if we just place these in the scene very quickly we have you know, a billion instances a million trees trees times a thousand instances. This will super quickly blow out our our memory budget even just on the on the instance data. Um it would also be very inefficient to trace this because if you have a tree that has a thousand instances at the horizon covering only 10 pixels that's a lot of work to traverse for the ray tracing unit and it's completely unnecessary if your tree covers only a few pixels.
So we'll need to to add some level detail to this. We need to reduce the instance count so we have any chance at all to fit in memory. Uh so what we do is we merge these individual meshes into fewer and fewer meshes with the far further LODs until we just have one instance per tree. Uh and at one instance per tree we can place a couple million of those without any problems thanks to the partitioned TLAS. Uh we simplify the animation skeleton accordingly along with this in a way that you know, makes the animation look consistent across different LOD levels.
Uh but now we introduce the second problem because now we merged all these meshes that were nicely instanced before into unique combinations. So now we've you know, new geometry uh that is, you know, replicated many times essentially, um and that that that blew out the memory budget again, right? So, we we can't we can't just leave it like that. We have to come up with something else. So, what we need is a uh a representation of geometry that is really lightweight in memory, uh simplifies well visually, uh meaning we can't just do the typical mesh triangle simplification, that doesn't work so well with foliage, um and it has to be efficient to ray trace, of course.
Uh and our solution to that is the opacity micro maps we touched on earlier, uh where uh we generate the the LODs offline and rasterize essentially the uh geometry of each mesh into a set of automatically positioned OMM triangles. Um and remember, OMMs are basically alpha maps, right? So, I mentioned alpha there's no alpha maps earlier, that's for the authoring process and LOD zero. The later LODs essentially have these OMM uh alpha maps, if you want.
Um so, this reduces the triangle count from an individual twig for from a few thousand or so to a few dozen, plus uh plus the OMM bitmaps. Um [snorts] so, these OMMs have multiple benefits, right? They're fast, they're hardware accelerated, they're small, uh but they're also reusable, so you can uh you can imagine that we can take a triangle with an OMM bitmap in it and sort of replicate that over a tree multiple times and position the same OMM multiple times, uh without actually having to duplicate the uh the OMM data itself.
So, in a way, it's like you can imagine uh it's a bit like a separate or or an additional level of transform that you get out of the BVHs. Um So, this reuse is what saves us a lot of memory. Uh we then take this idea uh and reuse not only the bitmaps across or the OMM bitmaps across the tree multiple times, but also we reuse it across multiple LOD levels. Uh so, we take the same bitmaps and at later LOD levels, uh just do the same thing again.
And then we take that idea further and put the actual triangles uh into clusters and reuse the clusters over multiple LOD levels as well. So, there's a lot of reuse in this whole system, uh which really brings down memory consumption. Um now, you might wonder because we're reusing all this stuff that represents the actual geometry, does it even How does that simplify things? Uh and the answer is it kind of doesn't, right?
The geometric uh the geometry itself that you're seeing starting at LOD one uh is essentially the same for all the for all the LODs, but the main point was to reduce instance count, not geometric complexity. Uh and so, all the way down to the last LOD, we actually uh do the same thing. We reuse that same uh geometry with one exception. Um Now, we still need to manage uh millions of instances, uh each frame we have to select and swap LODs, uh manage, you know, the scene traversal, uh manage different instance budgets, do memory memory management and all these things, uh animation.
And as you can probably guess, this is where the partition TLAS comes in. Uh we use a quad tree that's sort of visualized here to guide the partition selection uh of the partition TLAS. Um and again, we target about 1,000 or so instances per partition. We manage all this logic on the GPU. Uh we use, you know, partition TLAS to update only the relevant portions of the scene, so we have heuristics that decide whether something needs or doesn't need to be animated based on view frustum and and what's what's important to the scene.
Uh overall in this demo, uh the PTLAS has about 100,000 partitions and there are 3 million active instances, uh and we update about 80,000 or so on a typical shot uh of these instances per frame. The terrain is pretty interesting, too. Uh we we tessellate everything dynamically uh and displace it every frame, just like you saw in the Dragon demo. Uh we we use uh templates cluster templates for that. Um And then, uh you might have noticed I I mentioned 60 million plants in the beginning, but only 3 million instances in the PTLAS.
The rest of it, we you use this to our advantage and just sort of stamp uh very far away, very small plants into the height map uh of the terrain. Uh that gives us nice silhouettes uh and very, you know, plausible visuals for super far away stuff, and it costs zero instance count. Um so, on the right side, you see a visualization of that, where the the red parts are bright parts of the height map are actual instances and the rest is just sort of stamped in the height map.
Okay, let's take a look a quick look at performance. Uh And remember, these are work-in-progress numbers, uh and it reflects where we currently are at with this specific demo in this specific environment, right? It doesn't represent uh actual performance of any actual game. Um on the left, we have a breakdown of the most important regimes in the frame. As you can see, most of the performance or most of the work goes into the path tracing itself.
We do a pretty simple two-bounce path tracer here. The partition TLAS update is around a millisecond or so. And the uh and DLSS, of course, we use DLSS RR for denoising and upscaling, uh you know, is is another chunk here. And then everything else in the blue section is uh you know, everything other things that happen, like this includes the terrain tessellation, BLAS builds, uh animation, all the GPU management of instances, all these other things.
Um with DLSS RR in quality mode, which is upscaling from 1440p to 4K, we hit about 80 to 90 frames a second on a 5090, uh and about 60 frames, almost 60 frames uh in 1440p on a 4070. Um and so, we can fit this on a 12 GB card. Uh the total memory we see in the application is about 5 GB for the all the track buffers and data structures. Uh you know, Windows will tell you it's 9 GB for other driver uh related uh buffers and then DLSS.
Uh so, so this is roughly, you know, where we are at today for a landscape with 60 million plants. Um The partition TLAS takes about 900 MB and and a little bit of scratch memory. Uh so, I think this, you know, just to give you an an idea of what this what this currently looks like. Another interesting data point here is uh if we had a G-buffer that we filled just with primary rays, this would take about 1 to 1.5 milliseconds on a 5090, which I think is kind of a remarkable data point, uh and actually gets me to my last topic, uh which is uh a question that comes up relatively regularly now, uh and it is you know, is rasterizing your G-buffer still the best choice?
Uh because clearly the performance uh or the the the data point I had on the last slide, about a millisecond or so, uh for primary rays, indicates that it's at least not completely infeasible to think about ray tracing of primaries. And thanks to mega geometry, we can build all the BVHs at full detail, uh right? There's no uh it's no longer a blocker to have, uh you know, instance count or or LOD uh be downscaled for uh what what's in the BVH.
Uh so, if we did this, what would be the actual benefits? Um and the details of this are of course engine dependent and and title dependent, uh but it the idea opens up a couple of interesting uh paths of exploration, I would say. Uh so, for example, uh it might be possible to avoid some data duplication between the raster path and the ray tracing path, uh especially for static objects, you might be able to throw away your input meshes entirely and only keep the BVHs.
Um In in other cases, you might have different data layouts between raster and ray tracing and maybe get rid of one of those. Uh scene traversal is also often done twice, once for ray tracing, once for once for raster, LOD decisions, uh frustum culling, these sort of things uh could be reduced to uh a single run. Um there might even be a way to do this incrementally, um because in ray tracing, right, you have your BVH, the BVH you had last frame is still there in the next frame, so you might be able to take advantage of that.
Um CPU overhead we already touched on uh with the mega geo APIs, uh if your scene is managed uh entirely on the GPU, that can reduce uh CPU overhead a lot. Uh and then some camera and lens effects just get potentially really either cheap or easy to implement, you know, things like sniper scopes is like a common one, right? And then maybe crazy display setups, like triple screens or curved screens that you could adjust the uh the projection for pretty easily.
Um Geometric mismatches between primary and and secondary rays uh is another kind of common one uh that's a little bit painful, right? Often, we have lower uh LODs in your BVHs and for for your secondary rays, you need to use aggressive ray offsetting, you miss small shadow details. That kind of thing goes away completely. Uh your primary sampling becomes more flexible. Uh you can choose to do adaptive sampling or implement foveated rendering ideas, uh define your own anti-aliasing patterns, also such things such of things become possible.
Uh non-rendering rays uh can be interesting, right? Uh audio, physics, collision detection, that kind of thing. If you already have absolutely everything in your BVHs, then this becomes almost free cuz normally, these types of things don't need that many rays. Um and finally, uh tracing might actually just be faster than rasterizing. We've seen uh we we start to see cases where this is the case, uh even if you ignore the other benefits that are potentially there as well.
We'll look at one example really quickly. Uh this is a test application that we have, where we took the uh this Nanite demo that I showed you earlier or the content of that demo, brought it into a different application and just compared the raster and ray tracing performance. Uh these are full frame timings for just a G-buffer laydown, but includes all of the scene traversal BVH builds, everything to render basically a G-buffer.
Uh and you can see that it's kind of close like on the left side. So, green is timings for for ray tracing and blue is is rasterization. Uh you can see on Turing it's not so close, right? It's about 2x difference where where raster is faster. But on the other architectures it's getting pretty close and the difference is very often in sub-millisecond areas. So, you can kind of start thinking about, you know, is it worth to spend, I don't know, half a millisecond or so to just get those the the the simpler ray tracing only path going and and get some of those benefits.
And remember this includes the BVH builds. So, once you have that, your BVHs are done and are essentially free. So, something like, you know, shadows for any type of secondary ray, um for example, a shadow in the far distance on a tiny light where you would have never bothered to render a shadow map, you can just trace a couple of rays and it doesn't basically cost anything. So, this isn't to say people should start deleting their raster paths, uh but I I think it is worth to start playing around with these ideas if you're in the position to implement a renderer or work on an engine or title, and and see if this works for you, right?
Is there benefits that you can that you can extract from that? All right. So, in summary, we've kind of raced through a whole bunch of topics. DXR 1.2, SCR, and mega geometry are all technologies that we think every path tracer going forward that's built today should make use of. Uh we looked at a preview of our open-source foliage system that will come later this year and that is coming to The Witcher 4 as well. Uh and we touched on this slightly more exploratory idea of what if you could just trace your primary rays.
All right. Thank you for your time. I will now hand it off to Carmelo. >> [applause] >> Thank you, Martin. Um in this part of the talk, we're going to take a high-level, high-speed look at how to build a modern path tracer from scratch. We will uh talk about how to actually go from having almost no ray tracing to full real-time path tracing, and we're going to break the problem down into smaller chunks that you can tackle one by one.
There are a few things that I would like you to take away from this part. First, path tracing is not magic. It's built from components that you probably already have in your engine and know how to use. Second, progress is incremental. You don't need to go from rasterization to path tracing overnight. Third, path tracing can can scale across a wide range of hardware. I have run these examples that you will see all the way from an RTX 3060 to a 5090, and it can probably run on lower spec.
And in general, I want you to get a good idea of where the complexity really is and what the trade-offs are. Every game's renderer is going to be different and it's going to have different trade-offs, but understanding where the complexity really is and where to spend your effort will lead you to a more successful implementation. Now, I considered giving this talk as a roadmap to path tracing, but since this is GDC, I think we're going to do a skill tree.
Every skill in this tree represents represents a new challenge, a new technique or concept that you will need to master in order to implement modern path tracing in your game. And in this presentation, I will walk you through each of them and why they're important to path tracing and their impact on image quality. To follow along in this adventure, I'm going to implement every one of these skills in a sample project.
This sample loads the Bistro scene, which is a common scene you see in graphics research, but we've added a ton of things on top of it. Lights, skin meshes, and even tiny spaceships to make it more representative of a real game content. This is still a really simple scene for illustration purposes, but the learnings here actually translate to more complex scenes and have indeed been shipped in multiple times already.
Since we have to start somewhere, I'm going to assume that you already have some things in your engine. Almost every game these days has all of these components, so this should be a really reasonable starting point. You will need some form of a G-buffer. You can ray trace it like Martin mentioned, or you can rasterize it. Doesn't matter. You should also have some form of screen space ambient occlusion and maybe reflections, typically a combination of screen space with some blurring and maybe cube maps as fallback.
Um and you will need to have basic ray tracing support. You know how to build a BVH, how to trace a ray, how to bind shaders to a shader table. Doesn't need to be fancy, just working. I'm not going to go into the details of these because they have been covered extensively in the past, but I will have reference slides at the end. Finally, it would be nice if you have some form of temporal anti-aliasing or upscaling. Ideally, you have DLSS integrated or even ray reconstruction.
Now, our path tracing adventure starts with ray traced ambient occlusion. Seriously. The simplest way to dip your toes into path tracing is to add ray traced ambient occlusion to your game. RTAO is great because you just need BVHs and ray queries. You don't even need to set up your shader tables yet. And it's conceptually simple, but it has all the blocks of modern path tracing. Uh you can learn about Well, actually, there's something more.
If you have any form of screen space AO already, you're honestly like 80% there. Most of those effects already have some form of spatial temporal accumulation, usually to amortize the cost over many frames, and you can just reuse that. You don't even need to write that denoiser from scratch. Take out the screen space part and put a real ray in there. This is what GTAO, a common form of screen space ambient occlusion, looks like in our sample project.
Uh it is a straight out of the Jimenez paper, and it is very similar to what many games already have. And this is the same scene with RTAO on it. It is exactly the same denoiser. I just swap the screen space ray for a real ray. Let me show the difference again. SSAO and RTAO. This denoiser is from 2016 and it was probably designed to run on a PS4 or something like that. This is super cheap. You can learn about uniform rays versus cosine distributions with it, which is one of the most common forms of importance sampling, something that you will see a lot of in path tracing.
And you can learn about inline ray queries. Knowing how to use this will be critical for optimizing your path tracer, especially if you use them inside ray generation shaders and combine them with SCR. Think you mastered AO? Move on to GI. This is where you start exposing your materials and shaders in a shader table. Other than that, this is pretty much the same as AO. It just has more color channels. This is what a single bounce of GI looks like in our scene.
Some performance tips here. Keep your ray payload small. Can you get away with just normals and base color in your payload? Do that. Can you simplify your materials and lighting? Avoid using normal maps, maybe sample lower mips. All of that will help with performance. And absolutely use SCR. Uh GI rays are highly incoherent, and SCR will give you your coherence back and will fill your works. I have seen speedups of up to 40% in ray tracing cost just by enabling SCR alone on GI.
Not to mention, moving to ray traced GI can save you a lot of memory compared to baked alternatives, and RAM isn't getting any cheaper. Also, a cool trick here. Uh you will notice throughout this talk I don't actually mention transparent surfaces like glass. The reason is that I just put everything in a single G-buffer, transparent and opaque. The thing is glass surfaces don't really have a diffuse component to them, so tracing GI rays would be a waste.
Instead of that, for glass, I just trace a refraction ray inside the GI pass and get transparency almost for free. This particular trick may or may not work for your game, but path tracing is full of opportunities to cut costs like this, so be on the lookout for them. And at this point, you can probably spend a lot of time developing a denoiser, but I would advise against that. The denoiser that you write for each individual component will not be the same one that you would write for a full path tracer.
So, I would recommend that you speed run the main quest and get to path tracing first. Skip the denoising and just use ray reconstruction for now. That's actually how I took most of these screenshots. Um getting to path tracing early will give your production team something to work with, and it will give your artists a tool that they can use to develop the game look without having to rely on obsolete workflows or baking stuff.
The way that way the rest of the team can work on the game with realistic expectations about image quality and cost while you work on scalability and denoising. However, if you do decide to go on a side quest and start learning about denoisers, check out Nvidia's real-time denoiser library, or NRD. NRD has been fine-tuned over many years, and it offers multiple state-of-the-art denoisers that you can drop into your game.
Now, back to the main quest. Our next challenge is ray traced reflections. There are already a lot of resources out there about ray traced reflections, so I will not go deep into this one, but I will have some references at the end again. Just like with AO, your engine may already have some form of screen reflections, so you can reuse all that code and just remove the screen space part and drop in a real ray. This is what basic SSR looks like in our scene.
Notice a lot of the black you see is just SSR failing to find a result on screen. And this is ray traced reflections. You immediately see how much of your actual scene is captured by the rays. However, the real reason why implementing RTAO reflections is important is because it allows you to practice BRDF sampling. Ideally, you would implement some form of Eric Heitz's visible normal distribution paper. It is an absolutely beautiful paper that teaches you how to efficiently choose the ray directions for ray trace reflections.
And it even comes with code, so you can copy paste it and get it in your engine in like 10 minutes. BRDF sampling is one of those pillars of path tracing that you'll come back to again and again. So, take your time to play with it and understand it. You will use it a lot in path tracing. Now, armed with our knowledge of GI and reflections, we are ready for a real challenge. Stochastic direct lighting. This is where you encounter some of the most intimidating terms like next event estimation, ray guiding, uh ray guiding, and all the math and statistics.
The simplest way you can do stochastic lighting is to take a list of your existing lights, choose one randomly, and trace a shadow ray to it just like in AO. If the ray doesn't hit anything, shade that light, multiply by the total number of lights in the scene, and you're done. That takes us from the top image where we shade every single light for every single pixel and didn't get any shadows to the bottom image where you can't really see anything.
Um this scene has about 100 analytical lights, uh but only a few of them are actually visible from this point of view. So, choosing one at random is not going to work very well. We'll have to do better than this. Like technically, if you wanted to, you could say now that you have a path tracer. A single bounce next event estimation real-time path tracer to be exact. I just don't think you're going to ship this in video games.
However, this introduces one very important concept. That multiplication by the total number of lights, that's your unbiased contribution weight. Uh all it is doing is dividing by the probability of choosing this one light. See, if we have shaded all end lights in the scene, well, we would have end times more radiance. So, if we only pick one of them, then we need to multiply by that total number of lights, so we keep the same average radiance.
It's basically what an unbiased means. As we do more advanced sampling, computing this this weight will get more complicated, but the basic concept is still the same. This is not complicated maths, it just has complicated names. And with that, we are ready to face our first mini boss, the reference path tracer. This is the basic path tracing loop. Looks something like that. And at every bounce, you evaluate lighting from one random light.
This is called next event estimation, and it is exactly what we just did for direct lighting. Then you sample the BRDF to choose a new ray for the next iteration. And from that new direction, only a fraction of incoming light will actually make it into your path. The rest of it will be either absorbed by the surface or scattered some other direction. Those losses are modeled by the throughput. The more bounces you do, the lower your throughput will be.
Eventually, either one of your rays won't hit anything and you will sample the environment map or you will run out of bounces and just return the radiance that you have. The trickiest part of that loop is choosing your next ray direction. If you only use a cosine distribution like in AO, your specular reflections will be missing because you will never choose the right specular direction by chance. But if you only sample the BMDF like we did in reflections, then that will introduce unnecessary noise to the diffuse component.
The solution here is to do multiple important sampling, which is to say you choose whether to follow the specular path or the diffuse path randomly, and then use the inverse PDF to compensate for the fact that we can only take one of those two paths. Here I use the probability of 50/50 basically, so I just multiply by two at the end, but you should try other other options as well. You can try sampling proportional to the your specular color and your diffuse color, or even proportional to some Fresnel term.
And that's it. Accumulate the result over many frames, and you have your reference. Stop right there. The more features you add, the less reliable your implementation will be. And you're not going to ship this anyways, so take the win and rejoice the fact that you have now written a full path tracer which can produce really pretty pictures. Now, the noise hurts our resampling here. You want to be able to trust this. If possible, compare it with simple scenes against an an external tool to make sure that you got your implementation right.
So, now you have a path tracer. You're almost done, right? Just one more little thing to do. Easy. It used to be that every graphics tutorial out there would be something like, create a window, draw a triangle, and then do everything else. And I feel like with path tracing, we do something similar. We tell you to build your BVHs, render your reference path tracer, and then just make it faster and ship it. Obviously, that's where all the work is really at.
But if you made it this far, you really have all the skills you need to beat this game. So, let's go on. Now, let's leave our reference path tracer aside. We don't want to touch it anymore. And let's go back to our real-time render. The first level two skill is radiance caching. In this case, I'll be using SHARK. But you can try other things like neural radiance caching. SHARK, or spatial hash radiance cache, is a cheap way to reuse information between different paths.
The more you can reuse information between different rays, the better your image quality and your performance will be. More reuse will give you less noise, which in turn lets you spend less time and and resources denoising your image. How does SHARK help us reuse information? Well, in this case, we're going to use it to add multiple bounces of GI to our scene. This is again our RTGI implementation from before. It has one bounce of GI and then falls back to sampling the environment map.
What we would really want to do here is to have multiple levels of multiple bounces of global illumination, but that would multiply the cost, and that would be prohibited for most games. Instead, with SHARK, we will dispatch a separate path before GI and trace multiple bounces of light, but do so only for a fraction of pixels, maybe one out of every 16 or one out of 25. SHARK will record each of those bounces into its own cache, and then when we get to GI, instead of doing full lighting, we can just sample from the cache.
And that looks something like this. Notice how much more lighting we're able to recover with the extra bounces that SHARK gives us. So, back and forth again. Single bounce, multi bounce. Also notice that SHARK is not just a performance optimization for multi bounce. SHARK can average information from different bounces of different rays into a single cache entry, so it's also reducing your variance in the image. It is reducing the noise that you get.
It is giving you more bounces, and it makes your path tracing faster. So, it is really an amazing tool for every path tracer. And while we're on the topic of really powerful tools, let's bring back something that we did in our reference path tracer, multiple important sampling. This is again the same exact code uh that we had for MIS in our reference path tracer. We are reusing pieces that we already practiced. But instead of having a separate dispatch for GI and reflections, we'll use a single ray and shader.
Each pixel will then use MIS to choose whether the ray direction will follow a specular or diffuse path. Only the ray direction is actually in the A files code. The actual ray tracing code and lighting will be shared later, and that will help us with the variance, especially in lower-end diffuse where SCR might not be available. At the end of your shader, write your radiance to the corresponding buffer and zero to the other one.
But remember to compensate for the probability of choosing each path. Again, in this case, I did 50/50, so I just multiply by two. Putting all that together, our multi bounce indirect lighting goes from being just diffuse lighting to a full indirect path tracer at half the cost. Unlike SHARK, MIS will increase your noise. This is a common theme in path tracer uh in path tracers. You can usually make things faster if you're willing to deal with more noise, but you can reuse that noise if you find more ways to reuse information.
So, ultimately, more reuse makes your game faster. And on that note, let's talk about the elephant in the room, reservoir sampling. If you can only spend your time learning one concept about path tracing, this is the one. In modern path tracing, reservoirs are one of the most powerful ways to reuse information. And the simplest way you can do reservoir sampling is doing it inside your pixel. This is it. That's reservoir sampling.
The idea is actually pretty straightforward. Instead of a single light, we're going to pick many, and for each of them, we're going to compute how important we think it is for our current pixel. Then we will choose one of those many lights based on that importance and trace a ray towards it. Just like we did in when we first implemented direct lighting. While the code in this slide is technically the simplest possible reservoir sampling you can write, it is also useless.
If you assign the same weight to all lights, you're no better than uniform sampling. To really get the benefits of reservoir sampling, you need a target function or a target PDF like this. This is what defines your implementation of resampling. Your target function can be as simple as a dot product or as complicated as your full BRDF with light attenuation profiles and everything. Target PDFs are really a balance between performance and accuracy.
The faster your target function, the cheaper it is to evaluate many candidates, but the more accurate it is, the less noise it will introduce in your image. You You should spend time curating a good target function and fine-tuning this balance. Let's see an example. This is the sample scene where we're basically giving the same weight to every light. This is equivalent to not having a target function or to doing uniform sampling.
This is it with simple dot product. Uh even that will get us a lot more of the image. You can start to see something, but there's still a lot of noise. And this function includes uh square distance attenuation, spotlight attenuation profile, and also the approximate light intensity. Uh Uh the difference is actually significant, right? You can start making out most of the scene, and if you did give this to a denoiser, you can probably get a good result.
And in fact, if I give this to ray reconstruction, you get this. And remember, this is still a single ray per pixel. And we get pixel-perfect shadows for every single light without ever rendering a shadow map. In fact, this is probably cheaper than what your shadow maps are in many GPUs today. Not to mention that CPU complexity that you would save. Now would be a good time to introduce our second side quest, emissive triangle lights.
As we just saw, you don't need to make every emissive into a triangle light in order to do path tracing, but there is another trade-off here. If you treat all your emissives as third-class citizens and make one emissive area light per triangle, that will cause that will add cost and complexity to your sampling, which means that now you can afford a smaller number of initial sampling initial candidates uh in your local sampling.
It also means your total number of lights will probably grow into the thousands, significantly lowering the chance that you will find the right light for your pixel with just a few samples. That again will increase noise. On the other hand, making your emissives into lights means that you will at least sample them explicitly instead of relying on your BRDF sampling to hit them by chance. Explicit sampling can reduce your noise significantly.
That's right. Making emissives into triangle lights can both increase or reduce the amount of noise in your image. And the particulars will depend on your game and scene. So, you will have to find the right balance by playing with the different options. Try different things. Maybe not all emissives need to be triangle lights. Maybe only the bigger ones or the ones that move with your vehicle or character. And with that, we made it.
You really have all the pieces now to put together a state-of-the-art real-time path tracer for your game. You beat the final boss, completed the game, and finished your run-time path tracer. Congratulations. Of course, after you beat the game, there's always the end game. And you may have noticed that there are a couple side quests that I haven't mentioned yet. These are ReSTIR DI and ReSTIR PT, which are two variations of the full ReSTIR algorithm.
ReSTIR does reservoir sampling, but not just within your pixel. It does it across neighbor pixels and across multiple frames. This double reuse can actually result in exponential reductions in the noise of level um for your image, and significantly reduce the amount of work needed to denoise it. Particularly when you have really high numbers of lights, like in the hundreds or thousands. If you want to learn more about ReSTIR DI and ReSTIR PT, take a look at the latest release of our RTX DI SDK, which contains implementations of both, as well as a sample project.
And if you want to learn about even more advanced path tracing techniques, check out RTX PT SDK, which is also available today. With that said, we have completed our skill tree, and it's time for questions. Thank you. >> [applause] >> So, [clears throat] you mentioned uh you got briefly through the BVH section, uh especially on getting your render your path tracer a thousand times faster. Um but I've seen a lot of papers and a lot of uh progress towards like new systems of having curating, updating, and creating BVHs and ray traversal.
Uh how important do you think that is on that whole schema of getting path tracers to run in real time? Um I do think it's a key uh component, like I mentioned in the mega geometry section. Uh I think this, you know, what what we see with shipping titles today is usually the sort of downscaled version of your scene that's in the BVH, and this is what we're trying to get away from and sort of move forward onto all the full detailed geometry in the BVHs.
Uh we do think that the new APIs and and the mega geometry uh tech is what enables that, uh but we do see it as a key uh step forward, yes, absolutely. Thank you. Sure. Thank you for your presentation. Next talk. Uh I have worked on offline path tracers or real-time ray tracing techniques uh like by a real-time path tracer or ReSTIR DI, uh but uh the gap between David two and the final boss in your skill tree uh feels very big.
Uh today's real-time ray tracing techniques uh relies heavily on screen space resampling, uh that is ReSTIR-related technique and denoising. This means uh invalid samples from screen changes uh such as disocclusion areas, uh light flickering, or high-speed moving objects uh introduce a lot of noise due to lack of samples. So, it might be no problem on Amazon Lumberyard scene or RTX 3090, uh but when running uh 30 FPS uh on weaker CPU, uh that becomes yes problem.
Uh what do you think? Uh Sorry, it was a little bit hard to understand your question. I think you were asking about denoising. Uh Is that Uh uh uh lack of samples uh due to uh disocclusion areas uh because uh ReSTIR DI uh resampling uh in is invalid in disocclusion area. Well, that is actually one of the places where SHaRC, for example, or neural radiance caches can help you because you seed them with all the bounces of your of your path.
Some of them will actually be hitting surfaces that were disocclusions, right? So, your cache will be warmed up with samples for those even before you see the surface in the main view. Uh okay, thank you. So, uh it means so high-speed moving object is also uh also solved by these techniques. Yes, it helps with that, too. Yes. Uh okay, thank you. Hello. Really nice talk. I love the gamification of of path tracing as fun.
Uh one maybe side quest kind of on top of the final game. Uh I was wondering if uh you had any thoughts on where like neural um neural neural network um techniques would fit into this entire image. I know for instance, there's like NRC is a thing that's been going on, like what else is happening? Yeah. Uh there's actually a bunch of places, right? We have ray reconstruction, this completely a neural-based approach. We have Well, we had DLSS in general forever.
Uh we do have neural radiance caches. There are now neural shading as well that you can perform uh that actually gives you not just the shading of the final surface, but also gives you probability distributions that you can use to sample your next ray direction. There are uh complete neural materials. There is uh neural texture compression. All of those little pieces can actually fit inside the path tracing pipeline.
Got you. Next. Hey, thanks for the talk. Um I had a question about the primary visibility ray test. Uh was transparency included in that implementation at all? Um No, not in this particular test, no. In that case, do you at least have any thoughts on handling that, or do we can we keep shoving it to the side? >> Yeah, we we see implementations do this already, actually, uh for their primary transparent surfaces. Uh there there's cases where primary rays are already traced, uh and then use that, you know, as as a follow-on to to actually simplify the way that uh transparency is handled, which is transparency always painful uh in any approach, right?
Whether it's raster or even in ray tracing, it's a little bit more uh uh more work. Uh but it it helps a little bit to uh to unify uh the approach, you know, you don't have to necessarily separate out transparency the same way that you do with raster. Do you think it'll impact the performance gap between raster and a primary visibility ray? Uh on transparen- transparency uh specifically, I don't have any data on it, uh and like with everything else, it's going to be extremely scene-dependent and and, you know, shading-dependent and everything, so.
Okay. Um first of all, thank you so much for the talk. My question is more about ReSTIR PT. So, I'm trying to write an implementation of it in Falcor, and then looking at the original implementation from the paper, the edges are visibly noisy, and also like if you switch the camera really quickly, then the entire image becomes very noisy until the like spatial reuse can catch up, but the edges are like remain noisy because you can't really do effective spatial reuse, so I was wondering like how much or like you how can you deal with the noise on the edges?
Like does the RTX SDK like deal with that in any way? Uh sure, yeah. Uh the way we usually do that is I mean, the best option is always ray reconstruction. It really is the one that gives you the best result. If you don't have support for ray reconstruction, the recommendation is to use Reflex, which is one of the denoisers that we ship in NRD. Uh that has adaptive variance-based uh filtering. So, in the disocclusion areas, like the edges of the screen, because you will have more temporal variance cuz they were just visible, then you expand the kernel for spatial filtering.
Okay. Thank you. Right. I think we're out of time. Thank you very much. Thank you.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.