Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

The Andrew Faris Podcast · @andrewfarispodcast
Words
10,477
Runtime
52:28
Speaking pace
200wpm
Reading time
44min
200 words per minute, between the 181 median and the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Drew Maronei is the founder and CEO of Intelligjams. One of my very favorite tools, software tools in the e-commerce space. Has been pretty much since the first time I heard Drew on Andrew Udarian's podcast years ago talking about his background and why he was founding Intelligence and what Intelligence could do. Intelligence is one of those tools that uh is [music] has been used by a lot of my clients for a very long time and uh therefore has been a sponsor on this show for a very long time. This is not a sponsored episode. I actually
100 words, the words spoken in the first 30 seconds at 200 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 582 |
| Average words per sentence | 18.0 |
| Longest sentence | 264 words |
| Questions asked | 72 |
| Sentences containing a number | 55 |
Most used terms
Filler phrases
562 in total: like 275 · um 82 · uh 57 · actually 53 · you know 36 · I mean 16 · kind of 15 · right? 13 · sort of 12 · basically 3.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Drew Maronei is the founder and CEO of Intelligjams. One of my very favorite tools, software tools in the e-commerce space. Has been pretty much since the first time I heard Drew on Andrew Udarian's podcast years ago talking about his background and why he was founding Intelligence and what Intelligence could do. Intelligence is one of those tools that uh is [music] has been used by a lot of my clients for a very long time and uh therefore has been a sponsor on this show for a very long time.
This is not a sponsored episode. I actually wanted to get Drew on to talk about some elements, not so much of the tool itself of Intellams, but about how people badly run tests, all the mistakes people are making when they're running tests. I see some of them because most of my clients are running Intelligjs and so I see a lot of tests come through. Now, most of my clients are doing great tests, of course, but I still see and hear chatter about all kinds of results that are done poorly.
So, if you have any kind of testing cadence in your business where you are running split tests of any kind, whether they're incrementality tests or split tests on your website or whatever, this episode is for you. I want to think a little bit methodologically about how you get great results from your split test. And I thought, who better to do that with than the guy who sees probably more split tests than anybody around?
That is Drew Maronei. So, let's get into it with Drew right now from Intel Gems and talk about how to run better tests that give you real results [music] and uh make the make the most out of the tools that you're buying. Hello, Drew. >> What's up, Andrew? Thanks for having me. >> Good to see you. Yeah, good to see you. Thanks for doing this. I I really thought there was nobody better to talk to about this than you because I once once the idea for the episode came to my mind, I thought like I just wonder I've told you for a while that you guys need to build an AI therapist into into Intellig >> to to solve probably the biggest of all testing problems.
Uh which is the desire. It should it should it it's an AI therapist that should should force you like it should what it should really do it's not therapist it should it should lock your screen the moment you screenshot a test >> after the first like six conversions have come in and it's like five of them are on cell A or five of them are on the variant and one of them's on the control and you're just like I mean I don't even >> 500% gain >> or or five 400% loss you know. >> Yeah.
Yeah. Yeah. I I don't even I don't even see it. I there's no way there's no way the control side can overcome this massive deficit and you're just you're so tempted when that happens to to set to to just >> early screenshot boom send it to Twitter. >> Yeah. And so I you need a tool that that tells people like hey you can't screen it should be unscreenshottable. You shouldn't be allowed to do it. So you can't put it on your internal Slack.
You can't whatever. And the thing is for me this actually is a nice lead in I think to like what the thing is. I I have a clear rooting interest in every test I've ever run. Um, I am I am so deeply biased and it's because I, you know, I launch a test or whatever and uh and I'm I'm like, "Oh, this is gonna work. I'm gonna talk to the customer this way." >> That's right. Yeah. So, um, anyway, uh, yeah. So, solve that. Get get rid of the get make it so that people can't screenshot their tiny test. >> Therapist is for when your your group is losing early on [laughter] and it's like, "Hey, let's manage this anxiety a bit. like give it a couple days.
Uh >> yeah. Yeah. Yeah. >> Yeah. >> Okay. Can I tell here's the actual thing that triggered this episode for me. So I want to start here. I I'm I'm actually curious broadly what what bad testing methods you guys see and and hear from people as you talk to customers. But um but the the thing that I actually noticed first that really made me think about this was I was thinking about the concept of replicability. What I was actually thinking about was the replication crisis in academic literature which is a real thing that people may or may not know about is the idea that >> sort of people create these uh you know various tests, >> scientific looking >> right >> conclusions, >> right?
But if you know anything about that literature, what happens is that is that many many tests never actually get replicated or or cannot be replicated. And therefore this result that looks really strong in some paper in some journal somewhere that got picked up by the New York Times to tell you why the thing you ate is definitely going to kill you or why the thing you ate is definitely going to add 20 years to your life or whatever. uh or why you know what makes you a good or bad parent based on these things that you did you know or whatever right >> or red wine and chocolate is good for you and bad for you or good for you then bad for you >> yeah coffee yeah right all these things um so um so so but the problem is that in many of those tests they they either cannot replicate the result or not and I I think the same thing is almost certainly happening in people's uh people's website split tests which is that they run a test they see a result they maybe see a big result and they never think to replicate the test.
Um, so say anything you want about that. Tell me that I'm dumb and that people shouldn't replicate the test because they're marketers, not scientists publishing in academic journals. Or tell me that that's the reason why people make a lot of big mistakes. Or tell me anything else you want. What do you think about that? What do you think about my thesis there? >> Uh, okay. Okay. So if I repeat back the thesis, >> yeah, >> it's that there is likely also a replicability problem if we went and tried to do that among split tests. >> Yes.
That and and the problem specifically is that people think they have gotten a result that has been definitive in their business and then they make all kinds of conclusions. Oh, it's because the customer wants this and not that and I just I identified that variable which actually that that inference has another number of other problems too. But uh because of that uh and and but actually the result just was probably biased towards stat sig or something like that and for whatever reason what the especially if it's a really important thing in their business like you know again you guys can uh let people test price or something like that they they make the change and they never replicate it and therefore the result is actually not is actually not real.
Uh, and and if they were to try to replicate it, they wouldn't be able to. That's >> Yep. >> That's the That's the idea to me. >> Great. Um, I think there's a lot there and I I I think there's a lot of cases in which you're probably right. I also don't know if that's necessarily a bad thing, but like it kind of depends on >> the type of test. So, >> um, I don't I just wrote down a bunch of things so we can because I could go on for this for like two hours and want to say >> vaguely organized.
Um, number one, I think we should talk about like different test types that probably have different bars of rep replicability you should have. Um, two, I want to talk about statistical significance and a trade-off versus speed. So, there is a replicability versus speed consideration which we can talk about like overall testing programs. >> Yeah. Um you mentioned an interesting thing there which is like the stories around test results that we tell ourselves which I think is a really powerful concept uh that can be used wisely or correly um because data alone doesn't tell you a story um and uh yeah like how testing works in organizations.
So >> great let's talk about replicability right you may go into intelligence and uh you see this group has a 90% probability to be best that's how we reflect it like a it's a basian analysis that says hey based off prior and what we've seen so far what's the what's the confidence level that this group is better it's actually a different concept than what we all learned in high school which was the 95% confidence interval um like oh there's 95% chance we're in here.
This is hey we're just like weighing the probability that this group is to the right or above group B on conversion or on revenue producer on cost producer whatever >> we're gonna we're going to run into my own limitations here but let me try let me try to clarify that point because I think this is important. Are you saying that intelligence is not set up to give you essentially an RCT style like >> it's not a p value >> not a p value reading right what it actually is is a basian probability that your test is right >> we are looking at like sorry it's probably not a great word format but I'll verbalize as I go so let's say that you know we're looking at what is the true conversion rate values and we're we're often looking at profit per visitor which gets more Okay, look like we've observed a bunch of clients for for group A which are binary conversions and a bunch of clients for group B and both of those create some curve some cloud of possibilities that the where the true conversion rate should be right so we've only been running this for a couple days we've seen these like or or AOV you can imagine would be a distribution like this and so what we're looking at as you get more data the cloud of possibilities constricts, right?
It's like I have a 100,000 data points. I'm pretty sure what the average is going to if we're running off five like okay, yeah, we we know that we're on on Mars and not Earth, but like could be anywhere on this planet. Yeah. >> And with a basian approach, what you're basically doing is you're taking these different examples, control, group A, group B, and you're you're mapping those clouds against one another, and you're saying, what's the percent overlap?
So, like, you know, in this case, we actually have two predefined cloud boards. Like, you know, these clouds are pretty far apart. There's some overlap, but like that overlap only represents 5% of the time. So in that case, we are like 95% confident that this group on the right is better. Now that doesn't tell you how much better, right? That's a separate qu I feel like I'm I'm going deep in the we here already. >> That's great. >> Um we're saying is 90 cassette likely that new group one is to the right is a higher number than control. that does not necessarily tell you is it 1% better, is it 5% better?
And that's where you then have an interval of like possible outcomes. >> So when you run a test for longer, you're gaining two things. You're gaining more confidence in what is likely to be best. And that is like you get there faster, right? Because we're just >> just A or B. So So forget how much better or not it is. You just have a a more you're more likely to have a winner. and a loser. Exactly. And now this is and I think I'm anticipating what you're about to say, but I this is another thing I see people do all the time, which is >> sell B the variant beat the control and it's and it had a 20% lift in profit per visitor and and so they interpret that to mean I am now going to get 20% more profit for every visitor.
But I believe what you are about to say is actually no. The only confidence we have so far is that at that 95% level or whatever that the variant wins, but you don't actually know by how much. >> Yeah, cuz I mean there's a the control group the true value could be in its own range. >> Yeah. >> The winning group it could be in its own range. So like we may be you could be 90% confidence better but actually have like we don't know whether it's a.1% improvement or a 10% improvement.
And when you run it longer, again, we're constricting these clouds of possibilities. They're getting tighter, which gives us a better read on what to actually expect after the test one. And so, to repeat the point you made, there's replicability of will would this win again? Easier to get to the replicability of like it says 20% better. Is that what I should believe? Well, it's not going to be replicable if you didn't wait to tighten that window because that that window could be 1 to 20.
So, um I'll pause there, but like that is kind of probability to be best is truly just like what is the odds that it is the best of the options. And you can also do pair-wise. So, if you're testing three groups, it's like, well, I know that B is 90% more likely to be better than control. I know C is 90% more likely to be controlled, but it's actually kind of a toss up between B and C. >> So, okay. So, let's take this is great.
Let's take this is exactly what I want to do. And people some people are going to turn this off right now, and I know they're going to because they're going to go, "This is too nerdy." But this is the whole point. >> The whole point 2x 2x >> exact No, exactly. It's because people just want an answer. But what I think is part of what motivated me to think about this is like if I want actual long-term benefit for my business, I have to think a little more carefully.
And I think there is there is in the emphasis on speed. There is sometimes an unwillingness to go like I don't know almost like why do I need to go that fast? Why can't I wait for two more weeks and get a better result in this test and learn a little more deeply and critically have more confidence in whatever I learned? Wouldn't that actually be better for my business? Now, I again I don't want to put words in your mouth, but what you believe here?
Because sort of when I'm talking theoretic like this, of course, that could be true. Like you can sort of say like, "Yeah, of course I can." But you also may have plenty of examples to where you say, "Actually, >> screw it." Like take the example and pile up more tests, you know? Um but but I I just think that like there's sort of this danger here which is that so many tests are done so poorly that actually people are not really learning and and the biggest danger is uh what I would call negative transfer which is where you you like believe a test and you're actually wrong.
It's actually telling you the it's actually telling you information that or you misinterpret it and it's actually you're you're you you actually uh push the wrong cell forward and now you actually hurt your business by running the test. So um so anyway the um the let's just just to wrap up the example you gave earlier um the solution it sounds like to get more confidence in the actual percentage difference is more time and more results typically.
So not only do I know a winner or a loser but also like how much better or worse it is. Do you think the solution has more time or not? >> It's always going to be more time. We don't recommend I mean like this is where it gets tough as a software provider and not a service provider like we can't tell I can't force behavior through the app. >> If we were doing this ourselves test would run two weeks always even if you have enough data in the first two weeks we would hold because there is value in getting a representation like early data is extremely noisy.
We have some analysis on this and like it it takes time to dial in and like we could say, "Oh wow, it's five to one. He's actually quite confident." Like the other thing is early in tests there's actually a sampling bias. The people who are more likely to see your test first are people who are on your site more frequently. That means probably people who are in the middle of the buying journey, people who are returning customers versus okay, there's a steady drip of cold people coming in each day.
So like there is fundamentally the sampling when you look at a test after 3 days it is a different representation of the group than after you look at two weeks and like you may be starting a test on a Wednesday and Wednesday Thursday Friday behavior is going to be very different than Saturday Sunday. So get an even mix like two weeks ensures you have a solid sample and compresses some of this amplitude. So run it for longer.
That said this is always a confusing topic for people. it may be the test just doesn't make a difference. So, if you're still seeing it at like 50%, 51, 45, and we give you kind of like we call it the ESPN chart, like the probability as it changes throughout the game, throughout the test, if that's bouncing around the middle, you're not going to find it's just like, okay, this doesn't really make a difference. Like, conversion rate went up, but AOV went down.
It kind of nets out. So, um, there is like a little bit you can ask our our agent in the app whether you should run it for longer, but like oftentimes it's just, hey, this isn't going to create a change. Shut it down. Just like pick which one is better for your business and that I mean we could talk about that, but there's a whole art to like if these two things were tied, what would you pick? There's a reason basically all of my clients use IntellJams for their split testing.
That's because it is the split testing tool of record in the e-commerce space. You hear from Drew in this episode. Really smart guy thinking really well about this kind of stuff. And he designed the tool to reflect that. Is testing way beyond the basics. You are running every test through the lens of not just am I driving more conversion rate on my AB test, but am I actually driving more profit? [music] And even if you're a supplement brand or something like that, you can even check this at the level of subscription.
So am I getting more subscribers from cell B versus cell A of the test, whatever it is. You also of course with intelligence can test all the things that really move the needle in your business. The things like a price of your product, first product, first uh new customer order, [music] discount, free shipping threshold, shipping cost that you're charging customers, all of those things which really have a big impact on customer behavior.
I've got a brand that I'm starting with that's actually launching from $0 and is using us to help them launch their brand. And we're telling them from day one be running intelligence tests with landers, etc. to try all of those different things that good brands try and to test them from the beginning because we want to know how are these things going to impact customers behaviors because I believe they will impact them a lot and of [music] course we want to know how they will work at the level of profit.
So uh so intelligence just is one of those tools that if you're growing e-commerce brand and you're serious about performance marketing you should have it. You should install it. It's very fast and easy to install. And with my code, Ferris 20, you get 20% off your first three months. So, you can test it and try it out for yourself. Ferris 20. F A R I S20. Uh, for 20% off your first 3 months. Go to intelliggeems.io or follow the link in the show notes.
Fantastic tool built by people thinking really, really well about this who are really helpful in the process. [music] Go check it out right now. There's so much there. The the first thing I want to ask about is the timing. So I think it's another thing people don't understand uh which is which is that there is that having a timed element to the test people I think what I what I hear normally people say is like they they say all I cared about is getting a stat sig result once I got statistical significance and I hate even saying those word it's impossible to say those two words next to each other without mling >> statist sig I'm calling it stat sig um people will say I've got stat sig So, okay, then that means I can trust the result.
Um, you actually didn't just advocate that. You advocated instead something which was at least run it for two weeks. Um, can you talk about the the difference between those two approaches and if maybe maybe it is okay to just go satsig like what I mean what do you think? Yeah, I mean it is important to have a defined period of time up front that you were going to try to commit to, right? Because on let's say this idea of the ESPN chart.
So, you know, we're getting samples. >> The ESPN chart, just to be clear, is your guys's readout on which test is winning or losing like in a like if you're watching a a game, a sports. So, so let's say that like we have control versus new. This is the start of your test and we are looking at like throughout the test what is more likely to win because we could take this measurement at any time. And so early on you see big swings.
Oh, someone scores a touchdown. We got five five orders on the first group in the first day and like that chart is going to be all over the map like very someone scores a touchdown and there's a pick six. all of a sudden we went from like 80% likely to win to 30% likely to lose. Um, and what you tend to see is that it either evens out over time like dino noises they this is a pretty even match or it solidifies in favor of one of the groups like we get more data the noise compresses we get a better signal.
Um, again it's probably a good one to to watch and not listen to if you have the option. Um, >> yeah, >> but that that steadies out. If you are just looking for, let's say that you said, I need 85% confidence. That's my bar. What you are doing is just saying you're you're drawing a line here and you're saying, "All right, as soon as anything crosses that 85% threshold, I'm just ending the test." That is that is the equivalent of I'm just looking for 85%.
Yes, >> but like that can be really early on when we are still in a very high noise, high amplitude environment. It's the same as P hacking, which is the problem with a bunch of these SC like academic papers, which is like >> I'm just going to like cut and cut the data until I get like a p value that that fits with what I want. >> Yes. >> So, you know, it is important to decide upfront what level of confidence am I comfortable with because that is a real consideration.
If you want to be 99% confident in everything, God bless you. But you're going to run the slowest damn testing program I've ever seen. If you want to be 70% confident and solve for speed, also totally legitimate. Like you'll you'll be able to get a lot more reps in and and potentially that's very valuable for your business. And this is where like types of test matters. We talk strategy. >> Yeah. >> So so you should have a perspective on how confident do I want to be and we can talk about that. >> Yeah.
You also need to have a perspective on how long do I need to run this test to get a minimum detectable effect. Like there's people who've written about this much better than me and you should look up like minimum detectable attack effect and length of test to find great content on that. Um but like you can kind of say, hey, if I get 10,000 visitors in this, will I have enough to make a good decision that I like at 85% confidence?
We're going to have that in the app shortly. We like don't have that exactly, which is why we give two weeks as the rule of thumb, but you really should go in with both of those perspectives in mind. >> Yeah. I uh and the time element, just to come back to that, uh so it does a couple things. One of them is it is it sort of resists the P hacking problem, but also uh and and that is a really big problem. We kind of went over that fast, but for me, like I joked about it in the beginning, but I have so much bias going into >> test the first hour five verse one.
Wow, we're going to 5x the business. >> Yeah. Right. Right. And I have so much bias going in that like I am so tempted like when I see it move that way. And then you actually add in things like like my bias towards my boss or client or whatever it is. And it's like, well, if I can go put in front of them, you know, like this outcome. But the the knowledge is the prize at the end of the day here. That's that's the actual treasure at the end of the rainbow is like is like you >> you know, you >> truthfully answering the question is like what you're trying to do. >> So valuable.
So, this is part of why I'm advocate, this is why I wanted to have you on is to advocate for like the idea of get more robust reliable results because maybe there's actually a go slow to go fast thing here where it's like where again I'm not saying 99% confidence. I I agree with you. That's probably over the top and it's going to be really slow. But uh but there's some level here at which I would say if you can pile up a bunch of things that you really have great knowledge of and also that's in an environment where most likely actually probably lots and lots if not most tests and I'm curious what you'd say about this will actually have no result like a lot of you know a lot of tests will probably will show no difference um or just very minimum difference then then that's actually the giant win and if you can pile up I mean I would even think like if you could get six great results in a year six results where you were like really confident and they were well constructed tests and they really moved the needle on important things in your business.
That for me would be like a pretty big win. That would be like man you could you could really learn. >> Okay. But but here's here's the problem. >> Yeah. >> With that approach and why I think it needs to be a little more nuanced. 30% of tests win. >> So, >> right, >> you run six tests in a year and four of them now. >> I'm not saying I'm not saying six tests. I'm saying wins. I I'm saying six really uh really bankable results.
They're not even necessarily wins. They might be losses, but if you came up with because I'm just assuming a lot of tests you're going to get some sort of middle tier kind of number, right? You're you're going to get a relatively small impact. So if you could get six and that would be like a low-end maybe maybe you'd want something like one a month or something where you felt really great about the result. But I'm just assuming take two week timeline you run that means two tests a month.
Let's just call it like that, right? If that's a basic thing, two tests a month. If you came at the end of the year with six of those 24 tests you ran giving you really really really clear results that felt really good and really definitive in certain direction I think that would be pretty good and that's part and I think for some people they would say that's too slow whatever definitely not just six tests you should be running more tests than that but like but you know something like that I think those if you're testing good needlemoving things not just little stuff but good needle moving things that strikes me as really good and that's kind of what that's that's my thought but I don't know maybe am I still being too conservative >> I think that would be a great outcome I think if you found six really solid needle moving results, like that's that's pretty great.
If each of those are getting even just a percent or two, like that that compounds. Um, >> yeah. >> And it's like extra profit that you can take back into scaling your acquisition or, you know, just taking off the bottom line. Um, I think it's worth calling out that different tests require different levels of precision and science, right? So, like what I was playing around with here, again, sorry, not a visual medium, but like if you imagine a 2 by two >> Oh, it's on YouTube.
True. >> Big swings versus more incremental things. A big swing. I'm going to totally >> Yes. >> redo my my landing page versus >> I'm going to um change the customer testimonials on the landing page on the side >> and then >> kind of like high risk or like more one-way door like a little bit harder to reverse. So, I'm going to change my prices. Okay. like that is going to flow down to your ERP and probably your marketing and like it's not a it's one value in Shopify, but there's other considerations before you make that call. >> Absolutely. >> Versus low risk like whatever this doesn't work and just change it back next day.
It's like like one line somewhere. Um if it's incremental and low risk, just do it probably. Um like you know roll it out see what happens or um like we see there on lowrisk incremental and big swings like an approach called a canary test. So it's more like hey I want to I want to change my um my nav. Uh, I think we're not set up for like great product discovery. We're going to have more product lines coming out this year.
Like we kind of need to change the nav to set ourselves up better. Um, what you're testing there is more just that the new nav structure doesn't completely tank. Um, or an example I worked with a brand recently on like they were actually just shifting the PDP to a better liquid file that gave them a lot more customizability down the road and they wanted to test it because like hey is this performing? Are there any bugs that we're missing? >> But it was basically just like let's make sure this doesn't drop our conversion off the cliff. >> And in those cases you're actually not you're not seeking a big win. >> No.
Yeah. >> You don't need 95% precision. you're just trying to be like, can I check the box and make sure that like the canary in the coal mining coal mine's not about to collapse um and I can roll with it versus as you get to like big swing more one-way door. So, we're going to raise our prices 30%. We are going to um roll out a completely new theme with everything brand new, which you should probably never do. >> Yeah. >> Okay. that we're going to have a higher threshold of confidence going into that test.
So I may prefer 95% for some tests and be fine with 80% on others and have a different approach. So like I think it's good to have testing as part of your workflow, as part of your development workflow. Like hey, we're rolling new stuff out. We're rolling out new landing pages. We're rebuilding the theme. We're doing this. Like let's just have testing be super easy, consistent. It helps us make decisions. It helps us make sure we're not like train's not going off the tracks and then we're going to have these more structured bigger bets that we will have a higher degree of confidence on.
Um that is like I think a great way to approach it. Um, and then if it's in super incremental stuff, if you have enough traffic, your traffic is the budget for your testing program. >> Is it the traffic >> every visitor? >> Um, conversions ultimately. >> Yeah. >> Um, so like your conversions are your budget for your tests. Like you really >> you can mix tests a little bit, but kind of like each conversion should be a main signal for one test.
And so spend that budget wisely. Like you know, you're thinking of your risk adjusted returns on your testing program. Like is it better to run a test that has a 20% probability of having 30% uplift or a bunch of tests that have like 40% probability that have like of having 2% uplift? >> Um yeah, >> I don't know if I if I went too deep into extraction then there, but >> I am continuing to hire in the Philippines with my friends at more staffing.
You should be too because more staffing gives you an opportunity to connect to the best Filipino talent available for e-commerce people with deep e-commerce resumes, deep resumes across digital marketing and uh supply chain and all kinds of areas that are relevant to e-commerce businesses. And more staffing connects you to those people and understands who to find, how to get them, how to set them up for success in your business because they were born out of US-based e-commerce businesses that had hired a bunch of Filipino talent.
And so they understand the dynamics, the cultural dynamics, the workflow dynamics, all the things required to set both parties up for success. [music] They know how to do it and they've done it for my business. They will keep doing it for my business. They should be doing it for yours. The value proposition here is actually really, really simple. Your dollars just go a lot farther to getting top of the market talent in the Philippines, a country I should say, that has 80 million people, pretty much all of whom grow up speaking English.
And your dollars go a lot farther hiring talent there than they do in the US. Uh and and if you are trying to run a business with a lean opex and yet maintain great talent, it's just an incredible way to work. I [music] have been so grateful for my experiences talent in my business and beyond. I expect you to be as well. I just love it. [music] Uh huge, huge fans of the talent that I've been able to work with. Go check out more staffing right now.
Keep your opex lean in your business. Even as you add incredible talent, go open up a search. See what candidates come back for a position that you're hiring for right now. Before you go hire that talent locally, think about whether or not there's a possibility of doing that in the Philippines or or maybe even thinking a little bit proactively about what positions there are there, what positions in your business you can fill.
More staffing.co is the place to do it. More staffing.co. Get started today. I think the point that you're pushing there of of generally speaking trying to conduct I mean it sounds like the best the best test you can conduct is big swing low risk. That's the ideal test. Uh because because you could potentially make a really big impact without it being a one-way door. Um but yeah, I like >> shipping thresholds like yes, do that all day.
So easy to change >> can make a massive difference for your AOV >> and therefore your profitability and like I'm confused why more people don't. >> I agree for uh and by the way, here's a little trick on that for the people. This was not original to me. is told told to me and Taylor separately a long time ago in Taylor Holiday. Um from from some smart marketer somewhere, I don't remember who. Uh uh when you set your free shipping threshold, don't check your average order value.
Check your modal order value. Uh because almost nobody's buying at your average order value. Uh because what it represents is a wide range of orders, right? If you got a $50 product, but some people buy two, um then maybe your average order value turns out to be 80 bucks, but nobody is spending $80. So, don't set your free shipping threshold at 80. Set it at 55 or 60 or or whatever and then and see if you can get people to go over that line.
Um, so anyway, so just when you're thinking about your shipping threshold, yeah, you should think that way. Um, >> secondly, >> um, secondly, >> and people can do that. We just added if you go to sitewide analytics and order distribution, if you use Intelligjams, you can now find the mode and just like launch test from there because I think that point >> is so easily lost. >> Yeah. Um the uh another one that I would think would fit that actually would be your sort of new customer first purchase offer.
Um it's there's a little more oneway door there because a lot of times your ads reference your offer or something like that. But um but like if you're offering 10% off for every new customer, what happens if you go 20? What happens if you go 10% cash back? You know, what happens if you go stack discount? Those kinds of things. Again, at some point you may have to change some ads around, >> but that can be another big needle mover that can be a real profit suck.
I always I always call out Drew Sinaki here. Just just piling up the Drews and Andrews here. Uh Drew Sinaki. Uh I've heard him talk about how he used to whenever he used to buy e-commerce businesses, the first thing he would do is go just kill the 10% off for new customers offers because he just found it was almost entirely a drain on margin. He said it just it didn't actually push anybody to buy. It just it just hurt marginalized. >> Yeah.
So anyway, that'd be another one that I think would be a good a good test. Um, okay. Having said all that, maybe this is a good chance to jump into the conversation of are there any other sort of just like high value tests that you think brands ought to run? And if so, give me the framework for how they ought to run them like the two week versus just get a minimum versus whatever. Anything else that you see is just like why doesn't every You just mentioned the free shipping threshold.
What other ones should should people do just like that? >> Yeah, I I mean there's pricing I think is still underutilized. like it's what we started at as a company and uh >> there's been a lot more people running pricing tests in the last year thanks to all tariff madness which we're about to hit like the one-y year anniversary of and we're still in it but I just think that should be a pretty like always on practice for folks.
There are challenges to to changing price but you can build a process around it. And I find that like if you have someone who owns it, it is just so core to how your customers perceive you, what they buy, why they buy, how your P&L works. Um, so that is a a big one. Um, the second thing is frankly the welcome offer test. I we've seen it actually work well and there's a trade-off of you know you you do gather emails. How valuable are those?
But um we can test that welcome offer and often you can find ways to whether through the popup but more often through like the landing page that you're driving people to. Is it a free gift? Is it a you know a buy more save more? Are we changing the subscription discount? I mean there's so much depth that you can get to there and changes immediately help you scale. >> Yeah. um offers more generally. I mean we call it offer testing but there's so many flavors of it of I think people under like this more relevant in apparel but like clearance discounting like you should go get data on that can be a very powerful way to go figure out elasticity and like >> convert to cash faster.
Just standard bundles and buy more save like forget the new customer offer. Like often if you get people to put more stuff in the box, the shipping cost is going to be the same. >> Yep. >> That helps your total margin on that order. So like let's boost the AOV. What can we do on >> like we just launched like upsells on the checkout page and it's such a no-brainer. Um >> everybody underrates how much of a margin suck last mile shipping is.
Everybody under it's just it's just a huge part. It's often costs more than the product for some brands. Not maybe not often, but for some brands costs more than the product. It's it's uh it's a huge deal. >> So yeah, I mean those are those are some of my favorites. Um >> what other mistakes are people making with their testing? What are like the common things that you see? >> So we talked about don't run it for long enough.
Um I think we've talked about misinterpreting results or like having the wrong uh yeah ingoing expectations depending on the type of test. I'm going to list two things. We may add to this, but um number one is not testing frequently enough, like allowing for dead time, and number two is not having a strategy um that you were trying to fulfill with your testing program. >> Let's do those in reverse order because it makes sense to me that you would want to do the timed the timing of tests on the back end of how you set up your strategy.
So, let's talk about your strategy first. How do you think about that? Look, some stuff it's like, yeah, if I can like improve conversion on my my main new customer landing page, like go for it. That's just and that's good all day. Figure out ways to do it. But like people will ask me what's the right price? And I'm talking about price here, but it's it's relevant to all testing programs. I'm like, well, what what is important to the business, right?
Like I can I can tell you what is margin maximizing today, but are you trying to grow? So, are you trying to get as many people in as possible up front? Are you willing to take a loss or is first order profitability or break even? Do you have a subscription program in which case we're actually down to take a loss up front? Are you trying to exit in a year or two? In which case, we're just juicing the bottom line as much as possible.
Are we trying to um introduce new product lines later this year so we are getting new personas of a particular type that we think we can boost LTP on later? From a technology perspective, are we trying to make the site faster? Are we trying to make other products more discoverable? Like some things are are always good, but I always encourage founders. I'm like, well, what is the biggest rate limiter in your business?
Is it finding new customers, finding them efficiently? Is it introducing them to new products to booleell TV? And you can think of rate limiter like technologically if you think of your site as a product. But like that should be step zero. What are the unlocks for the business? What are my goals for the year strategically? Okay, how do we go test different unlocks for that? You know, like the test I mentioned earlier of this new PDP structure was a great very useful test that if you looked at it from the outside be like I don't get it at all.
But it helped them validate their entire technical direction for the ne rest of the year which let them introduce new products faster which let them be more relevant to more personas and cross-ell later in the life cycle. And so like you have to testing should be used to answer the most important questions of your business. Like one of one of my investors talks about it like the domino theory like okay great if if we want intelligence to be a public the public company.
That's a huge domino. We can get leverage by knocking over smaller small dominoes at the beginning that knock over bigger ones. And so you work backwards. You're like what needs to be true along that way? All we need to find this many customers. we uh find them all over. Okay, how many can we find on Shopify? How many can we find here? And you get down to these like testable hypotheses that must be true to hit your end ambitions and that is how you should think about a testing program.
Like yeah, >> these are the important questions that help me >> unlock the future I want to create. Um, and it sounds pedantic, but like people who come in with a really specific perspective on the things they need to solve, where they have bottlenecks in the business are more successful because if that first test doesn't work, we can try more and we have a really clear thing on the problem to solve. >> Um, so when I talk about strategy, it's like make sure we're answering questions that are important.
I had a coaching call earlier today with a with a brand that we were talking about a bunch of different tactical stuff and eventually the conversation shifted towards like well wait what are we actually trying to do like as a brand like where do we see >> where do we see the brand's future actually trying to get to and who is the customer we're trying to reach and how will we know if we've reached them like because you could put up the revenue number because that's the easy measurement right which is like in some way it's it's like okay that this is the revenue number we have to But there's this other question which is sort of like okay well who what what customers are going to buy the products that make it so that we get to that revenue number and are we reaching them now are we not is it like what about product line A versus product line B are they doing the things that we want them to do and there's some things where you're like you know you you go with your gut as a marketer and say I think if we talk this way and maybe you you know you plug your product into claude and say help me come up with some messaging and then you do all that stuff but but theoretically there needs to this alignment between an actual thing you're trying to accomplish and then the tactics which I would say intelligence and testing like this is a tactic.
It's a part of the way to do that in service to the larger strategy. And I think a lot of brands just actually aren't very clear about that at all. And it and it ends up meaning that you sort of have this wheel spinning around a bunch of a million little small tactics, you know, just like so you do a test over here because somebody told you you should do price testing and you do an ad over here because somebody told you you should do a post-it note ad and like whatever, you know, and I I mean, listen, I love tactics.
I think they all matter and I I I uh I don't ever want to sound like an anti-tactics guy because I mean I produce tons of content about tactics. >> That's right. That's how you how you achieve the goals, but the tactics should be >> in service of a voice. >> Yes. Yeah. So, I I like that point a lot. >> Kim from my team calls it like let's let's make sure it's like progress and not just motion. Like we can like be very busy, but are we making progress towards a goal? >> I really like that.
I like that way of saying that a lot. Yeah. I think I think that's a uh an underrated thing with you guys because I think that people don't think about Intelliggeems as uh something that fits into like split testing as something that fits into a true strategy. I think they just think of it as like a a le >> some people see it as a strategy of its own. It's like okay but yeah like >> right yeah I don't think so. Yeah >> but why? >> Yeah.
Okay. Then you mentioned the other one which is which was uh lining up test. You kind of started to talk about it but I don't want you to say anything else about sort of the timing of tests and and not having I I think just like if I think of if I were to turn how much money will I make from my testing program into a function, the inputs are how many tests could I run? What is the expected outcome if a test wins and what's the percent of tests that win?
Right? It's like this number of tests I have a 20% chance of them winning. Uh and they have an average 10% uplift. If you break those down, the number of tests, well, how many visitors do we have, right? Which is exogenous, like or how many orders, um, that is going to impact how many tests we can fit into the calendar. >> But then there's also like how much dead time do I leave between tests and how many tests can I run in parallel?
And so our best customers certainly never have dead time between tests. It's like I always have something running >> because they have a road map presumably. >> Exactly. That is testing different things towards a strategy because that dead time is just wasted budget. Like we have wasted orders that we could have been learning from with dead time. So what is the next thing? I have it queued up. Boom. Even if we need to analyze this one, I'll end it and I've got the next one going.
So I'm making the most of that budget. As you get to scale, you can even use mutually exclusive tests. So like >> totally >> hey you know I have enough traffic to get there in five in seven days. Drew told me we need 14 days. So okay let's split the traffic in half. >> Yeah. >> Run two different tests in that period completely isolated so we get clean results. >> And um you know now we've run in parallel and every time even if those tests don't win like you know in a good program 30 to 40% like are winners.
We're learning from the ones that didn't win. We're both like not making a bad decision, but that will fuel our hypothesis of what's to test next. It loops back to the road map. So, we start eating up that test win percentage. So, um yeah, I I like that. >> Maybe maybe another way of maybe another way of framing that is that in some ways the easiest variable to manipulate is number of tests. Like it's actually it's sort of easier to to manipulate that variable than it is the outcome of a given test.
So just like >> you'll get a better intuition for what's potentially impactful and how likely are things to win. The the more you get to practice. >> Organizational focus is so hard. It's just occurring to me as you're talking like the thing that would stop somebody from doing this is that somebody has to sit down. They have to have the strategic direction. They have to have the plan. They have to then plan out like what the tests are.
They have to have somebody whose job it is to execute it. They have to have a reminder. The test ends on this day. Go set the next one. you know, there's a bunch of things you have to do to do that well, and it's so much easier sounding than it actually is. Um, yeah. >> Well, because what I'm what I'm advocating for typically involves >> different people at the or you know, like in a great >> program, we have >> the acquisition marketers, the retention marketers, the head of ecom or head of products, potentially merchandising if that's a role, like all using intelligence towards different purposes and like it's hard to coordinate. >> Yeah.
Do you want to did I not ask you anything that you that I should have asked you as we wrap up here? >> Okay, I'll hit on I'll hit on a couple. I I I didn't cover a couple of the things I wanted to rapid fire. >> Number one, the stories around why your tests work and didn't work are important. If you see a result that's counterintuitive and you cannot put a story on it, like I would be skeptical of the data. Um or hey, maybe this is 95% confident does mean like 5% of the times it's it's not correct. >> Um and and I think we >> we don't do a good job of of handic like handicapping that of like one out of 20 is not not that infrequent.
Um it's like you know you go to >> So stories are important. They also help us digest the data. But what I see people do is they get like so deep in telling themselves a story about the test results and they start like looking at new verse returning and by segment and by pay channel. And now all of a sudden we're looking at sample sizes that are like tiny. So use the stories but like just check yourself as you start segmenting data that you actually have enough volume for that to be meaningful.
Um, two, I like embrace the fact testing is a great organizational tool, right? Like maybe we're not surprised by the results, but it helps us get to a decision. >> And um, >> yeah, >> like organizational focus is hard. I I mean the way we initially started testing was that my last company, >> the CEO was super skeptical just from an intuitive basis and he had great intuition for the business >> of a bunch of changes and it was like we had to >> test things to to get decisions through and that created very good discipline on the team.
But it was also very helpful for me as a like decision making and and political not in a bad way tool to >> drive pace of decision-m so we're not just stuck in like >> abstract debates forever. >> Um >> that that's actually part of what I meant earlier by saying like six really reliable tests in a year is like it's possible that those tests are actually conf are losers, you know, but they confirm really clearly the answer for exactly the reason. thing we've been talking about for two years.
We can just like put to bet >> which is also why I'm like only semi joke when I say we should put a betting system into intelligent and bet your co-workers. >> Um >> that's I don't that's a that's a very good idea. >> Yeah. Yeah. Just sell some scores. Um the last thing I'll say is like all of this is potentially upended by agentic optimization. And so if brands want to talk about that I would love to. But like there's both the aging can be embedded with the expertise around the statistical knowledge and like we've already started on the path you can find it in the app >> but um it can also undertake a lot of the effort that it takes to run tests but from hypothesis building to codewriting to the analysis and like I don't know I I I it's going to have very interesting implications. one one example actually I think a lot of the stuff in the incremental box incremental lowrisk like hey we're going to personalize every landing page depending on what segio segment someone is in >> it's not worth doing if someone has to go sit there and rip through 50 slash like someone just doesn't want to do that >> a machine will happily make those personalizations and okay maybe >> a 2% gain is not worth doing >> for a person but if a machine can do that 20s times all of a sudden now we're at, you know, like potentially a 4% gain.
We lose a little bit and like stacking up some of these incremental wins may make more sense in a world where effort is cheaper. So, we're thinking about it a lot, have a lot of like product. I mean, we're just trying to jam with folks who are doing cool things. Um, and uh, yeah, it's going to all this is going to be quite different in a few years. Intelligence.io IO is the place to go get Intelligjam yourself if you want to run some of these tests and because you guys are a sponsor of this Ferris 20 for 20% off your first three months.
Thanks Drew. Appreciate this a lot. Thanks Andrew. [music] Big thanks to Drew for doing that episode. Uh as I mentioned, go to intelligjs.io to go follow up with them and thanks to them for sponsoring this show for so long. Ferris 20 gets 20% off your first three months. Don't forget to also follow up with my friends at more staffing morestaffing.co. / aaf huge fan of more staffing and uh and yeah, grateful to both of them for sponsoring the show.
Uh for everything I'm doing, you can go to afgrowth.com and if you are interested in working with me and my team, you should fill out the intake form there. Tell me a little bit about your business so I can understand better what you're [music] looking for, whether or not we are the right fit. I might have a recommendation for you even if we're not. So, uh so [music] it's worth just a couple seconds to tell me about your business so I can learn about it and uh and we can see what the possibilities are.
Of course, you should you should subscribe wherever you're watching or listening. You should comment with any questions, any thoughts you have. I read and interact with all of those. So, go do that right now. Um, huge thanks to you for listening to watching. I'll see you next time.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.