Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Roblox · @Roblox
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in Roblox's most watched videos.
Most replayed moment at 18:33
3.1x that video's typical replay level
frame rate optimization part of what we helped grow a garden with. So, who's that Pokémon? Has anyone worked out what this thing is yet? It's a coconut. So, for reference, a Roblox sphere is supposed to have around 432 triangles.
Said at 18:26
The graph counts replays. It does not show where viewers stopped watching.
Words
7,275
Runtime
37:02
Speaking pace
196wpm
Reading time
30min
196 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Please welcome Peter Yang, Andrew Grim, Ryan Barbachia from Brook Haven, Brad Bower from AdoptMe, and Tyler Gennas from Paradox Games. [Music] All right. Hey everyone. Uh, my name is Peter. I'm really excited to talk to you today about how our top studios use analytics and experimentation to grow their games. I'm really lucky here to be joined by my colleague Andrew, uh, Ryan from Brook Haven, Brad from Adomi, and of course Tyler from Tower Defense Simulator. Uh, let me advance the slides. Okay, so here's our agenda today. So, first Andrew and I are going
98 words, the words spoken in the first 30 seconds at 196 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 466 |
| Average words per sentence | 15.6 |
| Longest sentence | 140 words |
| Questions asked | 35 |
| Sentences containing a number | 50 |
Most used terms
Filler phrases
267 in total: uh 127 · um 49 · like 32 · you know 31 · right? 10 · actually 6 · kind of 6 · sort of 4 · basically 1 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Please welcome Peter Yang, Andrew Grim, Ryan Barbachia from Brook Haven, Brad Bower from AdoptMe, and Tyler Gennas from Paradox Games. [Music] All right. Hey everyone. Uh, my name is Peter. I'm really excited to talk to you today about how our top studios use analytics and experimentation to grow their games. I'm really lucky here to be joined by my colleague Andrew, uh, Ryan from Brook Haven, Brad from Adomi, and of course Tyler from Tower Defense Simulator.
Uh, let me advance the slides. Okay, so here's our agenda today. So, first Andrew and I are going to walk through our new config and experimentation products. And then we're going to spend most of our time with our top creators sharing real examples and tactics of experiments that they ran to grow engagement, monetization, and acquisition for their games. So, our vision for analytics is to hopefully guide all of you to take the most important actions to grow your games faster.
And we'll be shipping a lot this year, including these new insight reports that try to summarize and aggregate all the metrics and player feedback that you have. But today, we really want to focus on configs and experimentation. Um, and I want to invite my colleague Andrew here to talk about why we think these products are a gamecher for all of you. All right, thanks Peter. So, I'm Andrew Grim, an engineer here at Roblox, and I think every single one of us in this room is obsessed with the same question.
How do I make my experience better for players? So last year at RDC we launched funnel analysis. It was all about understanding where in our games players dropped off. Uh but today I want to talk about the next step. What can we do about that? So let's say we want to build a better fatoule. We see people dropping off and we get it ready. We're ready to ship. We're going to hit publish. Brace for impact. What can go wrong?
Let's check our error stats. Oh my god, there's so many errors. Okay, what do we do? We jump into studio. I can try to yank out the whole feature. I can try to find the bug and fix it. Maybe I shut the whole game down if it's bad enough. So, was there a better way that we could have done this and prevented the panic? Instead of a painful hot fix, we could have used a runtime config and specifically a feature flag. Think of it as a remotec controlled if statement around our new tutorial.
Gives us an on-off switch. When there's those errors appeared, we could have flipped that switch and instantly disabled the feature for everyone. No editing in studio, no server restarts, the problem would have been solved in seconds. Roblox internally uses uh remote configs extens ext extensively. We make hundreds of changes every day across all our surfaces. So now that we've got a safer way to release our feature, how will we know if it actually improves our metrics?
We could wait a week and look at before and after comparisons of our analytics to see how it affected things. But what about all the external factors that also affect our games? There's weekends, holidays, what if a popular streamer played our game? These are all called confounding factors and can easily mislead you if you're just comparing time periods. What we could use here is experimentation and specifically AB testing.
Can I get a show of hands for who's familiar with AB tests in general? Okay, cool. Like 2/3. That's awesome. And how many people have actually ran an AB test inside a Roblox game? Okay, also a good amount. That's great. Um, so a true experiment neutralizes confounding factors by showing more than one version of something to random players at the same time. You can isolate the impact of your change from all the outside noise.
So the goal of an experiment is to be confident that if we ran it again, we'd get the same results. So I'm going to go over the basics of an AB test. We've got something we want to test. In this case, this button that says click me. And we're not getting enough clicks on it. We want to find a better version. So we're going to come up with a hypothesis. We want to increase the urgency. And in this simple example, we're just going to change the button to red because that we think that'll drive some urgency.
So, we've got our new version. We're going to call these the control and the variant. Now, once we run the test, we're going to get some results. Um, in this example, we're going to say we had a,000 users in each group. The first one had 30 clicks and the second group had 35. So, while we did see an improvement, it's pretty small, especially for a relatively small sample size. So, the question is, if we ran this test again, do we think we'd see the same results or would they flip-flop?
We don't really know, right? So, I'm going to go through a couple more examples. Let's say I've got two just basic coins and I want to flip them and see how many times they each come up heads. So, in the first example, I'm going to do that test. First one comes up five, second one comes up heads seven times. If I ran that test again, those results are just totally random, right? Everybody knows that a regular coin is going to come up heads about half the time.
So, another experiment. My friend Peter here has a coin that I'm going to say he got through dubious origins. I think it's weighted. He says it's just lucky. Um, so let let's run the test and see how it goes. We got the exact same results. And we know from the first test that this is pretty much just random with a small sample size. So, what can we do about that? Let's flip them a thousand times each. Okay, we're getting somewhere.
We still got pretty much the same results as far as percentages, but experiment 3 is obviously much more convincing. Intuitively, we can see that when you run it this many more times, it starts to move from lucky streaks into measurable patterns. But we're still just using intuition. Can we go a step further? We can. And the solution here is math. Now, don't worry. I'm not going to show any formulas. We're going to try to do all the math for you.
But there's the idea of statistical significance. And I'm gonna say stats sig for short because it's a mouthful. Um using outcomes to measure if the results were random chance or not. You can pick a confidence level and then calculate a confidence interval. This is the range where we think the true value of that test would fall. As you increase sample size, you decrease and shrink the confidence intervals. Um so in our example here, if the range spans the 0%, it's not stat sig.
If it's fully above like the platinum bundle, it's about 4% better than our control. it's stat and if it's fully below the gold bundle in this example is about 2% worse. So let's go back to the earlier tests and remember with only 10 flips we've got this huge confidence interval which is pretty much what we'd expect. But with a thousand flips it shrinks becomes precise and we can show that it's about 40% better than a regular coin.
So I can finally prove that Peter's lucky coin was in fact loaded. I knew it. Okay. So, I've shown you why we need a way to safely deploy code and more importantly a scientifically valid way to measure our impact and learn what actually works. To make confident decisions, we need to run experiments. And to show you some tools that can make all of this possible, I'm going to hand it back to Peter. All right. Thanks, Andrew.
Okay, so first we're working on configs to let you roll out features and update in-game values in real time without restarting your servers, kicking out any players. and you'll be able to create and edit configs in both studio and creator hub. Right? And then we're busy working on experimentation to let you measure causal impact from both your in-game changes and also your matchmaking configuration changes. And um yeah, so let let me give you guys a quick tour of the product, right?
So um you know suppose we're working on a dungeon RPG and uh our demon retention is below benchmark and upon closer examination we look at the funnel and we see that like you know 33% of players are even make it into our first dungeon right so our experiment hypothesis is that we can grow D1 retention by you know launching a by testing a new onboarding dungeon that's guided and way more fun. So let's create our experiment.
So, let's create a 50-50 experiment. 50% of users see the new onboarding dungeon and 50% see the current one. And we'll set D1 retention as our goal metric. Uh, and we'll run the experiment. Uh, we set a duration for about 14 days or about 2 weeks. Now, a quick tip here is that, you know, a larger roll out and a longer duration will let you detect smaller static lifts than uh, you know, if if you didn't set set it properly.
So for example, you can you can detect a 1% lift in de retention instead of only a 5% or above lift. So let's run our experiment and while it's running you can monitor it progress in the experiment list and when the plan duration is reached when the 14 days is reached we will show a decision needed status on the experiment and you can click on that to view the results. Now um very important tip here and I make this mistake all the time is like you know the day after you launch your experiment you're like you know let me just look at the results you know I won't do anything I look at the results and then you look at the results and it's like oh oh crap it's like you know 30% lift or 30% dropped in retention and you're like okay I'm going to ramp this now and like try to avoid doing that right because that's usually like novelty effect or something else is going on you definitely want to wait until your numbers are static or you know ideally at least wait two weeks before you start making choices about what you want to Um so here are the results for this example experiment.
Um it's a little bit hard to see but you can see all the metrics here and you see here that uh D1 retention is up you know 37% and session time is down 2% and and both are static like if if there's like shading in the background that means it's static. Um and you know often you have to you know kind of make trade-offs here right where like one thing is up static and another thing is down static. And in this case, you know, it's pretty obvious because retention is up so much that it's pretty obvious that we just ramp the experiment to all our users so they can all see the new onboarding flow and we can realize our lift.
Um, so that's just a quick example, but um I want to hand it over to Ryan now to talk about a real example for Brook Haven. >> Yeah, >> thank you. Hi everyone. Uh, my name is Ryan Bbachia and I am the product lead on Brook Haven. Uh, for those of you who aren't familiar, Brook Haven is a top role playinging experience where users can own houses, drive cars, customize their avatar, um, roleplay any scenario, and be whoever they want to be in the small town of Brook Haven.
Uh, today I'm excited to be sharing some of the experiments we've run, including one of our most important features, the avatar editor. Uh, as you might expect, the avatar editor allows players to customize their avatar. Uh, players can equip any item from the UGC marketplace for free within the Brook Haven experience. Character customization is core to any role playing experience. In fact, it's so important in Brook Haven, it's used by over 70% of our users every day.
So, what's the problem? Well, the avatar editor was due for some improvements. Players have been asking for new features such as 3D clothing. Uh, we were running on an old back end that needed to be updated. and we wanted to add some UI and UX polish. So, we decided to set up an AB test. Uh the reason we decided on an AB test for this feature is that we could understand the impact the test would have on our metrics before we rolled it out to all of our users.
Uh and as with any experiment, we had a hypothesis. Uh our hypothesis for this was that by launching this new avatar editor, we'd see a lift in both playtime and early player retention. Um for the experiment setup, we had two groups. uh our control, which would be the avatar editor with no changes made to it, and the new editor, which would have all the features I mentioned before, the 3D clothing, the new backend, and the UI and UX polish.
Before I get into the results, there's a couple of pointers I want to share with you all when you start getting started with experimentation in your own games that I think can be really helpful. The first is going to be informing your players. Um, we've learned from our experience that many players on Roblox aren't familiar with experimentation. In fact, many of them find it surprising that themselves or one of their friends is having a different experience from one another.
Sending a simple message in your Discord communities or in your update log can go a long way in building the transparency and trust that players want, especially as you continue to experiment more in the future. The second point is going to be around play testing. We all play test every day. Anytime we package up a build, we're testing it to make sure that it's working bugree in the way we intend. Sometimes when features are really important or they're different, we want to make sure we test with our players first.
We give them a version of the build. We get feedback from them. That's great. These types of tests are awesome. And um they're even more valuable in the context of play test in the context of experimentation. Let me give you a quick example. Let's say I took the avatar editor and I didn't play test it with my players and I just rolled it out as an experiment. Uh, I took the wise words of Peter and I didn't look at the results for about a week or so.
Uh, and finally I take a peek and I see that my results are actually negative and I'm quite surprised by it. Um, and I go and talk to my communities and they say, "Well, you know, the new categories you chose for 3D clothing didn't really make a lot of sense. I was actually kind of confused." Well, if I'm hearing in my communities, I'm probably experiencing it with my other players. This is something that could have easily been caught in a play test.
And so I highly recommend triing this when you go get started with experimentation. So now you've informed your players, maybe you've run a play test with them, um, and you're getting ready to roll it out. Two more things to consider. The first is going to be your group sizes. There's lots of ways to think about group size, the number of variants you have and other things and other factors. But one thing on our team we like to think about is risk.
Not every feature is made the same. Something like the avatar editor I might consider to be relatively high risk. 70% of our users engage with it every day. It's a feature that players really love. I want to make sure that when I go to roll this out, I can minimize any negatives that might come from it. So, starting with a smaller group size might be beneficial because I can always ramp it later once I'm confident with the results.
The second point is going to be time. And Peter touched on this, Andrew touched on this, we're all going to touch on this, but letting your test sit long enough is critical. Getting stat sig is the entire reason why we test. It's literally confidence in the numbers that we have. Uh, and to give you an idea of how important it really is, um, I have an example here from Brook Haven of what it took for us to test and build, uh, the avatar editor.
So, by late April, we kicked off development and by miday, we had our first public play test, right? The way we did this was that we sent a link to our users with access to a place. They were able to play it with the D with the new editor being the test variant. Um, and so we also launched a uh Google form that players could provide to us so we had structured feedback that we could uh collate and address. We ran a second public play test to give players additional time to give more feedback and get the fine tuning in.
By early June, we were ready to release and we went live at 10%. So 90% of our users were still in the control group and 10% had the new editor. About a week later, after going looking at some of the numbers, we felt confident and ramped it to 50%. At this point, we're just looking for stats sig. We wanted statig in our in some of our variables like D7 retention, which takes a long time to get. And so, we waited two more weeks and by early July, we were ready to launch and we rolled out 100% to all of our players.
In total, this is 28 testing days, including two public play tests. It might seem like a long time, but the weight's totally worth it. We saw a 14% increase in uh new user play time. We also saw an 8 1/2% increase across day 0 through day 7 retention. And surprisingly, we also saw a drop in ARPD. Um this is something as Peter mentioned, you always have to make trade-offs, but given the lifts that we saw here, we were totally okay with it.
It was a big win for our team, and our players are really happy with the results as well. With that, I'm going to hand it off to Brad. Awesome. Thanks, Ryan. Hi, everyone. I am Brad Bower. I am the game director for Adopt Me, which is a pet care and collection game uh on Roblox for anyone who hasn't had a chance to try it out yet. So, you know, in addition to a lot of the uh major features that you can put out there, AB tests are also incredibly useful for those smaller quality of life uh updates that you'd be looking to do.
Uh for today, I wanted to cover one of these that we did in the past year uh in Adopt Me. And so, uh to give some context and background in here, Adopt Me was having an issue where bad actors were uh impersonating other players. They would come in, copy their username uh of of someone, they would get their friend to show up uh and then try to trade them uh do a trust trade with them to hand over a pet. Therefore, the player will lose their pet uh to the bad actor.
And so for us, we said we've got to we got to do something about this. We can't keep letting this happen. Um and so we came up with a number of solutions. Uh and the reason we decided to go with an AB test on this is that trading in AdoptMe normally directly reflects into engagement for us. When we see engagement or when we see trading go up, engagement usually follows. And so we wanted to be incredibly careful when we were touching this metric where we know we're directly going to bring trading down.
Um, and so the uh our hypothesis and the the solution that we went with was that we wanted to try to intro in introduce a a delay when someone was new to the game. You wouldn't be allowed to trade for the few first couple of hours. So I'll show what those kind of looked like and what the players were seeing here. Um, so as you can see on the left, you know, our control group was everything was normal. You'd be able to jump into the game, trade as usual.
Uh but in our test group, the player would instead be presented with a message that said you can't trade for the next however long was remaining. Uh and so this would directly interfere with anyone trying to pull off this sort of a a scam within the community. And now our hope with this is that we would see scam reports going down and at the same time we were really hoping that we would at least have minimal negative impacts to the rest of our stats.
And so let's uh before we get into the actual um uh results of this, I want to take a moment just to look at some of the data that we were getting back and to really encourage everyone to take a a moment to dive deep into the data as you get it. So what you can see up here is there's obviously a a major difference between two of these groups at the very start and then as it goes on over time you're starting to see them kind of come back together. uh what this graph is showing on the bottom on the x- axis there is you're seeing the se first length uh first session length time uh and then on the y-axis is the total play time.
So what you're seeing with the control group is basically within the first few seconds of the game uh players are suddenly for some reason in the control group sticking with that at a much higher rate and then over time coming back to normal. So something isn't feeling right about that. And we were trying to figure out, you know, what's going on here. Uh and so we dug in deeper to try to see, you know, what what's happening.
The answer ultimately the people that we were targeting were the uh were exactly what that group is inside of the control. Those are the bots and the alt accounts. They had a way to be able to see if they were on the correct account uh or within a test group or not uh right away. And so we did that without that investigation. we probably would have looked at this and said we're having some really negative impacts but we were able to figure out no there's something more happening here.
So as you are looking at your data consider segmenting it you know uh that could take the form of new players versus veteran players or maybe looking at your data based on different player types. For Adop Me that would take the form of maybe traders or the collectors or builders in the game and seeing how does the data move based on those different player types. Um, and one other piece is to just always consider if something doesn't make sense.
You do have the option to jump back in and, you know, run another AB test, create another hypothesis. It may make sense in order to use that time to figure out exactly what's happening and will ultimately provide you with that extra insight regardless that you'll be able to use going forward. So, going back to the results that we saw, uh, this was a huge success for us. Uh scam reports went down across the game by 5 and a.5%.
Uh you know, we were just hoping that we weren't going to go negative on our other stats, but we ended up seeing playtime go up a percent. Uh and revenue even went up about 2%. So small AB tests, even like this one with just a control and a and a variant on a small portion of the game, can really have massive impacts for engagement for your audience. And with that, I will pass it over to Tyler. >> Thank you. So, I'm Tyler, the studio director at Paradox Games.
Um, we're always optimizing tower defense simulator to keep players engaged. And one area we wanted to focus on was matchmaking. So, the problem we had was matchmaking times were taking too long and that caused some frustrating players and leading to drop offs. So, the tickets had too many parameters and creating some niche cases when players were combined with game mode, difficulty, and party size. So our hypothesis was to reduce these parameters such as simplifying the difficulty options and shortening uh pairing times and increasing average play time.
To test this, we ran an experiment to compare the standard multi- difficulty setup against a streamlined version. In our experiment setup, we split our users in a 60% control and a 40% variant. In our control, we had players choose the difficulty up front in the matchmaking menu and they have tickets covering the game mode, the difficulty, and the party size. And there was no in-game voting. In the variant, we streamlined it by removing the difficulty from the menu, and the tickets only cover the game mode and the party size, which difficulty was later handled through in-game voting.
Turning to the results from our variant implementation, before the update, we're averaging about 4,000 matches disconnects per hour out of roughly 34,000 matches played. And right after the update went live, we saw a noticeable spike in disconnects largely tied to migration challenges. But overall, the average disconnects rates increased by 40% underscoring early friction in the new streamline matchmaking flow. And looking at this player behavior, the 40% jump in disconnects was a clear red flag.
It signal player frustration. But on the upside, we did see a 14% increase in average play time and a 4% lift in our day one retention. But a major drawback stood out. Players often quit immediately if the voted difficulty didn't match their preference, which led to creating a poor experience. Because of that, we chose to stick with the control version, prioritizing fear of disconnects and stronger community vibes over modest metric gains.
And I'll be handing it over to Ryan. Thank you. Okay, I'm back. Uh we are going to switch gears a little bit and talk about monetization. Uh some context here. So, historically in Brook Haven, uh when a user would click on something that was a pagated piece of content, something behind a game pass, uh nothing would happen. There was no feedback. Uh let me give you an example. Player clicks on the hoverboard, which is behind the premium game pass.
There was nothing. No popup, no message saying you can't equip this. Uh nothing directing them to the store. I think we can all see the problem very quickly here, right? Um this is something you would expect. you'd expect to see some sort of an upsell. Um, and so we saw this as a big opportunity, right? We called it contextual upsells. When a user is trying to use something in a particular context, we upsell them with the opportunity to buy the product.
A lot of you probably already do this in your games today. We didn't have it. Um, great. So, we're we have a very simple hypothesis, right? By adding this in, we should see a lift in both ARPD and conversion. Um, so I think it really almost begs the question. Um, this is blatantly obvious. Um, it seems like a really easy win. Uh, we should definitely roll it out, but like maybe for every day we don't have this. We might be losing money.
Why are we even bother testing this? Um, it should just be out in the wild. We should just we know it's going to work. Uh, there's two problems with that. The first is going to be magnitude. Um, I couldn't tell you how much this is going to lift by. It could be 5%, it could be 50%. but you really don't know. And so it's important to test against a control so you can make sure that there isn't some sort of external force that's driving up the revenue or whatever the metric is that you might be looking at.
And the second point, which is arguably the most important, is I could say with pretty high confidence that this is probably going to move ARPD down conversion. I couldn't tell you though that it wasn't going to move a different player metric. Let me paint you a quick picture. Take a user that's brand new to Brook Haven. They've never played before. They have no context about it. They go in and they click the hoverboard and this popup comes up and they say, "I don't have Robux in my account.
Cancel." Um, they click on something else and another popup and they cancel and by the second one they go, "You know what? I'm done with this game. I'm out." It's not hard to believe that there could be a world that players would be negatively impacted by showing them pop-ups even when it's something that's pagated. Um, and so it's important to remember to test even when things seem really obvious. Uh in this example, fortunately, looking at the results real quick, it was a great win.
Uh we saw a 10% lift in ARPD. We saw an 18% lift in payer conversion rate. And very fortunately, we also saw no no change in playtime and retention. And that was the most important thing for our team was that we could introduce this without having to make a large trade-off. Um so with that, I'm going to hand it back over to Brad. All right. So, on the monetization side, uh I really wanted to spend some time to just talk about uh running AB tests with starter packs.
I think starter packs are something that many of us use inside of our own experiences and wanted to talk about a few of the things that that we've done inside of here. So, before I get into the example, I really wanted to start by talking about some of the variables that we like to test. Um, so one of the things would be product types inside of AdoptMe. That would be the form of of pets versus maybe a house or offering bucks as currency to players.
Uh, we like to also look at things like cost. Do you start with a really low cost? Where's maybe you go higher? Where's the right uh spot for from a spend uh standpoint? Uh, presentation is another huge one. What is the player going to see? What's the art that that pops up on them that they're going to be looking at to make their decision of if they want to purchase or not? When do you prompt them? Are you going to uh throw it at them the moment that they join into the game?
Maybe it's right after they finish the tutorial, or maybe it's after a really positive experience in the game that you say, "Now is the right time uh to to prompt them with this." And then finally, we also like to look at offer duration. Is this something that you're going to say, "Hey, let's go with the FOMO route. uh you've got about 10 minutes to decide if you're going to uh to make this purchase or is it something that maybe you give them a week to decide so that they get a chance to become familiar with your game and to really understand why this is such a great offer to them.
Uh so let's go ahead and look at one of the examples we did. I think Peter might have spoiled a little bit of this one yesterday for us, but uh so Adopt Me uh very recently this year, we had updated the presentation portion of our our starter packs. And with that update, uh, it became very obvious that maybe one of our products in here, uh, wasn't as, uh, cracked up as it should be. As Peter said yesterday, I really liked your joke.
I thought that worked well. So, uh, thought I'd steal that. Um, so yeah, we had the cracked egg, as you can see up here, and it really just I don't know anyone that would look at that and be like, I really want to buy a cracked egg. That's the thing that I need in my uh my inventory. So, we said, let's go ahead and try out something new. So, we brought in instead what is our royal egg in the game. It has a nice crown on it and it also comes with another advantage which is that we were able to tell people that it was guaranteed uncommon or better which is important when we're talking new users who again are not exactly familiar with the game.
Seeing something like uncommon or better pet is uh is terminology that at least is more known and kind of gives everyone a starting point uh when they're trying to judge the value on it. So let's go ahead and take the results and and take a look at the results on this. Uh so our new user Arpoo jumped up 6% from uh from this change. Our new user conversion went up 6% as well and we had a nice surprise uh D7 retention uh jumping up 2% uh on top of all of that.
So overall big success for us. We were very happy with where this ended up. One tip that I do want to give here as we were all talking about stat sig uh is how important that is with this sort of a test. So, you know, the number of spenders in your game is already a very small number of your player base. New users is a small portion of your player base as well. Put those together, the new user spender is a very very small segment of your users.
Uh so you have to give this time to bake when you do this. Uh for example, this example for us for adopt me took us just over a month to get to statistical significance. Uh so you need to give it that time. Don't kneejerk react. uh you will see particularly the first day, the second day, it may jump up 40%, might drop 60% the next day. It takes time for that to level out. Give it that time before you make your decision on that.
So, uh with that, I'd like to hand it over to Tyler to jump us into the next section. >> Thank you. >> Okay, last but not least, let's talk about experiments to prove you new user acquisition. So thumbnails are the front door to your game. So let me walk you through an experiment that we ran with our PVP update back in July. So the problem we wanted to set to solve was maximizing our QPTR to help drive user acquisition.
So our hypothesis was simple. If we follow recent trends like handdrawn styles, we'll see a lift in QPTR and new users. To test this, we ran an experiment. So, our control, we had a 3D render of the new game mode packed with action and competitiveness. And our variant had a hand-drawn style focusing on fewer characters and a cleaner style. So, the results were clear, our variant delivered a two times higher QPTR than our control.
And we as we continue to refine our thumbnail strategies through iterative testing which by experimenting with art styles elements like vibrant scenes, character focus and theatic updates we have consistently lifted QPTR. For example, one of our variants had a 1.2x increase. Another was a 1.5 and our top performer delivered a two times boost over our baselines. So the key takeaway here is iterative thumbnail testing works.
It strengthens homepage recommendations and drives traffic to your game. And now I'll hand it off to Brad. Thanks, Tyler. So, I think I'm going to really just reiterate a lot of what Tyler just nailed on this is uh with the example that we recently used. This was back in the spring. We were running one of our big spring festivals. It was a cherry blossom festival that we put in. And this is a perfect uh chance for us to do some thumbnail testing as well.
Uh for some context, I think Adopt Me, we always have somewhere between three to five thumbnails running at any given time. Just really trying to gather as much data and information as we can. Roblox's entire uh thumbnail testing system works so seamlessly uh that it really is just a great way to be gathering information. So in the spring we ran this uh you can see we said, you know, we're going to really show off our new miniame, our new pet.
And so you can see that in the variant on the right. Uh we were going for a real action shot and thought that's exciting and it's going to pull players in. Uh so we used our Kai Jr. over there knocking over the blocks which was part of the miniame and we put that up against one of our uh historically high performers which is the dog walking which really fits well with AdoptMe. I think it when you think about a pet care game taking a dog for a walk is about as uh quintessential as it gets.
Um, but our hope was, you know, this is really going to pull people in and get that attention and and deliver on that. So, how did we do? We lost on this one. Uh, so we ended up finding that our control group uh was performed 5% better than the the thumbnail. So, I think there were a lot of key learnings that came out of this that um that everyone can take and apply to their own uh games. So, up here you'll see a bunch of different thumbnails that we've tested out.
I think you can tell we really like the dog or I guess the the player base really likes the dog uh because it is all over. So the first thing and I think this is one of the big lessons that we learned out of this is that you have to remember your audience. The thumbnail tests are typically being shown on home recommendations and what types of players is that actually showing up for. It's typically to new players, people who have not come into your game before.
So those players, they don't really care about a new miniame or even that we're running a spring festival necessarily. What they are trying to judge at that moment is is this a game I want to play in general. Um the other big thing, make it relevant. You know, I think it's really easy to make a clickbait thumbnail. Uh I think we've all been there where we've been baited into a clickbait thumbnail and uh find out that the game inside of there is not what the cover was uh showing off.
This is really detrimental to your game. you may get a higher QPTR, but you're probably not going to be able to succeed on all of the other stats that uh that the algorithm and that the players ultimately are looking for uh inside of there. And the last piece is to really stand out. Roblox has so many amazing experiences and so many thumbnails are trying to grab a player as they are on that homepage. So, what are you going to do to be able to stand out and make sure that someone sees yours and wants to click through it?
Um, so now you've all had a chance to kind of look at this. I'm sure your question is, you know, which one of these is is Adopt Me's highest performer. So, uh, we'll go ahead and reveal it. We'll see if your guess was right when you looked at it. So, our winner, this is the the pet collection thumbnail is the one that, uh, typically performs about 1.7 times better than our average thumbnail. And I think it's clear why when you look at it.
It is hitting exactly what the game is about, about pets, about collecting. It's vibrant and colorful uh and really just hits exactly what someone would be uh looking for inside of the game. So with that, Peter, I hand it back over to you. >> All right. All right. Thanks, everyone. So, we just covered some real experiments that our top creators here run to grow engagement, monetization, and acquisition for their games.
And experiments help you test changes safely, make datadriven decisions, and ultimately help you grow your game faster. Right now, experiments are super powerful, right? But remember that your creativity should still be in the driver's seat. Um, you know, experiments will give you signal, but you should feel empowered to decide what you want to build next for your game, right? There's other signals, too, like player feedback and other signals you should consider as as well.
So with that, uh, we have a QR code here and we're, you know, everyone's hard at work on this feature and, um, hopefully we'll launch it later this year. And if you want to get early access or help us test this feature, help us test the exper feature, uh, please take a picture of the QR code and, uh, fill out a quick application form and we'll get back to you as soon as possible. And, um, Andrew has also printed these stickers, confidence boost, uh, that, you know, we can hang out in the back there.
I don't think they're officially sanctioned by Roblox, but Andrew is very passionate. So, so, uh, come get some stickers after this. Yeah. Thank you. [Music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.