Back to Blog
Research

We Bet Our Own Hook Study Would Replicate on Small Channels. It Did Not.

July 24, 2026•12 min read•By Prepublish Team

Two weeks ago we published a study of 349 YouTube hooks comparing the openings of each channel's biggest hits against its biggest flops. Three patterns survived that study's own attempts to kill them. Then we bet those patterns would replicate on small channels, wrote the predictions down in a public registry before collecting a single hook, and ran the test on channels between 1,000 and 20,000 subscribers: 96 analyzable channels, 941 videos.

Here is the verdict table, in the exact format we committed to before the data existed:

Pattern from the July studyVerdict on small channels
Openings that engage the title's promise are more common among hitsCAN'T TELL
Openings that never give a reason to keep watching are more common among flopsCAN'T TELL
Claim-or-result-first openings are more common among hitsCAN'T TELL

"Held up" means the same within-channel association appeared in this sample, not that changing an opening caused more views. REVERSED means the pattern clearly flipped direction after correction; CAN'T TELL includes results compatible with no association and effects too small for this study to resolve.

None of the three earned a YES. None reversed either. The study did not recover effects as strong and consistent as July's. Under the preregistered assumptions built from July's per-channel win and tie rates, each test had at least 94 percent power at 96 channels, so this non-replication argues against patterns of the original strength existing on small channels. It does not prove the true effects are zero; the limitations section says exactly what this design can and cannot rule out.

We are publishing this because we promised to before we knew the answer. That promise is the point of this article as much as the numbers are.

Why we bet against ourselves in public

The July study had one weakness its readers found immediately: the smallest channel in it had about 21,000 subscribers. Most creators asking whether the findings apply to them are under that line. So the obvious follow-up was to re-run it on genuinely small channels.

The less obvious part is what we did before running it. Every "we analyzed 10,000 videos" article you have read shares a dirty secret: the analysis came first and the claims came after. Given a dataset and freedom to explore, you can always find a slice that flatters a conclusion, publish the slice, and quietly drop the twenty slices that showed nothing.

To make that impossible for ourselves, we froze the entire study before collecting any outcome data: the three predictions, the 98 enrolled channels with their exact hit and flop videos, the classifier prompt, the statistical tests, and the rules for every verdict. The frozen registration is public and timestamped at osf.io/2zq9p. It cannot be edited. Anything this article claims can be checked against what we said when we did not know the answer.

Because we asked three main questions, a result earns YES only if it remains convincing after accounting for all three chances to get lucky. Nothing here got the chance to be rounded up.

What we actually found

The July study's six measurable gaps between hit and flop openings looked like this, and here is what happened to each on small channels (percent of openings showing each signal, hits vs flops):

Signal in the first 45 secondsJuly hitsJuly flopsSmall-channel hitsSmall-channel flops
Engages the title's promise87.0%77.9%79.2%80.4%
Contains a specific number63.3%52.3%47.2%47.4%
Never gives a reason to stay10.2%16.3%21.0%19.8%
Opens with a concrete claim33.3%25.0%22.2%19.8%
Opens with a context dump18.6%26.2%30.6%29.3%
Opens with a greeting7.3%11.6%16.6%23.3%

Read the small-channel columns row by row. Five of the six gaps collapsed to roughly nothing. The 9-point title gap is gone. The 11-point specific-number gap is down to two tenths of a point. The confirmatory tests, run per channel the way the registration requires, agree: between 39 and 44 of the 96 analyzable channels are exact ties on each of the three confirmatory patterns, meaning those channels' hits and flops showed that feature at the same rate. The remaining channels split close to evenly.

That is the finding. On the features this study measured, the openings of small-channel hits and flops looked similar: none of the three confirmatory opening features consistently separated hits from flops in this sample.

One row did not collapse, and it is the last one.

The greeting signal, an exploratory note

Greeting openings ("hey guys, welcome back to the channel") were the only signal pointing the same way as July. On small channels, 23.3% of flop openings start with a greeting against 16.6% of hits, and at the per-channel level the pattern leans the same way in 33 channels against 11. This was a preregistered exploratory outcome, which means the registration forbids us from giving it a verdict or a headline, and we honor that here: it gets no verdict. It is the one line in this study consistent with July, it is the strongest per-channel split in the data, and small channels greet far more often than big ones to begin with.

Association, not causation, and exploratory besides. Treat it as the one thread worth watching in a future study, not as a proven rule.

Was the effect hiding in one niche?

A fair criticism of the July study was that it pooled niches, and a pooled result can hide a strong pattern in one niche canceled by its opposite in another. So here is the same per-channel scoring split by niche, using the frozen classifier's genre assignments. One honesty rule first: this breakdown was not preregistered, so it gets no statistical tests and settles nothing on its own. It is here to show what the pooled numbers are made of.

Each cell reads channels supporting the July direction / against it / tied:

NicheChannelsTitle promiseNever gives reasonClaim first
Gaming commentary3210/14/89/14/910/7/15
Video essay307/9/147/7/169/7/14
Finance225/4/137/6/98/4/10
Tech70/4/31/2/41/3/3
Education43/1/02/1/10/2/2
Explainer10/0/10/0/11/0/0

The table did not reveal one large niche obviously driving the pooled result: finance leans mildly toward the July patterns, gaming commentary leans mildly against, and the three smallest slices are too thin to lean anywhere. Across the three larger niches, tie rates run from 25 to 59 percent of channels. A post-hoc split this size cannot rule out niche-specific effects, especially in the smaller slices; what it can say is that no big slice looks like it is quietly carrying a July-sized pattern.

The composition itself is worth knowing: two thirds of the enrolled channels are gaming commentary and video essays. That reflects which small scripted channels YouTube search actually surfaces, and it is listed in the limitations for exactly that reason.

What a small creator should take from this

The honest reading, which is also the useful one:

First, the hook-anatomy advice you are following was measured on channels far bigger than yours, and this study could not find its patterns at small-channel size despite being built to find them. If you have been agonizing over opening formulas at 3,000 subscribers, this study is permission to relax: in this sample, none of the three confirmatory opening features consistently separated the videos that took off from the ones that did not.

Second, that makes those opening features a weak place to look when you ask why a video flopped. This study did not measure what does decide it, but topic choice and packaging (the title and thumbnail that got the click) are the obvious suspects, and both operate before a single second of your video plays.

Third, the boring caveat that stops you from over-rotating: this study measured which videos get views. It says nothing about whether viewers who clicked stay to the end. Retention is a different measurement on different data, and the July finding that concrete openings hold clicked viewers is untouched by this result. Do not read "openings do not decide small-channel hits" as "openings do not matter for anything."

Method, in full

The registration at osf.io/2zq9p contains every rule; this is the summary. A fixed grid of YouTube searches across the same six scripted niches as July surfaced 1,579 channels. Four mechanical gates reduced them to 98: subscribers between 1,000 and 20,000; at least 15 long-form uploads aged between 60 days and 24 months with median views of at least 300; scripted content confirmed by a frozen classifier that never saw view counts; and public captions. Every exclusion is logged with its reason. Within each channel the top and bottom 5 eligible videos by views formed the hits and flops, locked at enrollment before any transcript was collected. Two channels fell below 3 usable videos per group after caption drops and were excluded by the preregistered attrition rule, leaving 96. Collection produced 948 classified videos from the 98 enrolled channels, 31.9 million combined views; the analysis and every pooled percentage in this article use the 96 analyzable channels, 941 videos. By classifier genre: 32 gaming commentary, 30 video essay, 22 finance, 7 tech, 4 education, 1 explainer.

Classification was blinded: the model saw only the title and the first 45 seconds of transcript, never the views, the channel, or the group. Per channel, each pattern's rate among hits was compared with its rate among flops; channels counting as wins, losses, or ties fed exact two-sided sign tests with Holm's correction across the three confirmatory hypotheses. The per-channel counts (wins/losses/ties): title promise 25/32/39, never-gives-reason 26/30/40, claim-first 29/23/44. Corrected p-values were 1.0 for all three (uncorrected 0.43, 0.69, 0.49). A preregistered robustness check re-ranking every video against its 10 nearest-in-time uploads produced the same non-result, as did splitting channels by median views (300 to 999 vs 1,000+). The greeting result: 33 channels flop-leaning vs 11 hit-leaning, uncorrected p = 0.001, which survives correction across the five exploratory outcomes but receives no verdict by rule. Speaking pace and words-to-first-payoff were as identical between groups as they were in July (medians: 168 vs 170 wpm, word 15 in both groups).

Limitations, honestly: the sample is "small scripted channels surfaced by fixed YouTube searches," not all small channels; YouTube removed chronological search in 2026, so a neutral frame is not available to anyone. The search was widened twice during enrollment, both times before any outcome data existed, and both widenings are documented in the registration. The niche mix skews heavily toward gaming commentary and video essays, so niches like education are represented by a handful of channels and their rows in the niche table mean little. Views are a noisy outcome at this scale; a channel whose median video gets 400 views has hits and flops separated by numbers small enough that luck plays a large role. And this was not an equivalence test: CAN'T TELL does not prove zero. The design was strong for July-sized patterns and much less informative about smaller effects, and with only three to five usable videos per side and binary labels, exact ties are mechanically common rather than independent proof of sameness.

The data

Every row is public: download the CSV with all 948 collected videos (the analysis used the 941 from analyzable channels), their URLs, views, channel medians, and every classification, so any row can be checked against the actual video. The frozen registration with the full protocol, the locked channel list, and the analysis code is at osf.io/2zq9p.

We said whatever comes out, ships. This is what came out.

Frequently asked questions

What is a preregistered study?

A study whose hypotheses, sample, and full analysis plan are written down and frozen in a public timestamped registry before any outcome data is collected. It prevents the most common failure of data articles: exploring a dataset until something publishable appears, then presenting that slice as if it had been the question all along. This study's registration was frozen at osf.io/2zq9p on July 22, 2026, before any hooks were collected.

Does this contradict the original 349-hook study?

No, and that is the interesting part. Both results can be true at once: the patterns were found above roughly 20,000 subscribers, and a study powered to detect effects that size could not find them below it. The two studies used the same classifier, the same metrics, and the same per-channel design on different populations. What this result does undercut is the assumption that hook advice measured on large channels automatically transfers down to small ones.

Does this mean hooks do not matter for small channels?

It means the measurable anatomy of an opening is very unlikely to decide which of a small channel's videos gets views. It does not say openings are irrelevant to keeping the viewers who clicked, which is a retention question this study did not measure. The one opening signal that repeated across both studies: greeting openings were more common among flops, in 33 channels against 11 here.

Why publish a study where nothing was confirmed?

Because the registration committed to publishing whatever came out, and because an informative null is useful: under the preregistered assumptions the study had at least 94 percent power to detect effects the size of the originals, so failing to find them argues against patterns that strong on small channels, even though it cannot prove the effects are zero. Selective publication of confirmations is how the advice ecosystem got into its current state.

Want to analyze your own scripts?

Start Script Analysis