Back to Guides
Scripting11 minUpdated Sep 19, 2026

How to Teach an AI Agent Your YouTube Channel Voice

Adjectives do not transfer to a writing assistant, and quoted examples do. The rules worth handing over are already in your transcripts: how an opening starts, how long the sentences run, which words the channel avoids, and when the point arrives.

TL;DR

Adjective briefs fail because they constrain nothing an assistant can check. An evidence brief works because every rule carries a verbatim line from a published transcript, so a draft can be compared against the source. The seven extractions below come from one real transcript; the generator does the same work in roughly six minutes.

Try it on your own script

Paste your draft below. You get your hook, structure, and pacing scores, a script-level attention-risk map, and the single biggest issue quoted from your own lines. Free, no login.

Free · No login · See a sample audit first if you prefer.

Key Takeaways

  • Adjectives do not transfer to a writing assistant
  • What transfers is already in the transcripts: the shape of the opening, the length of sentences, the words the channel never uses, and when a point arrives
  • Extract those, attach a verbatim line to each rule, and the agent gets constraints it can check
  • Every quote field longer than 8 words in a generated channel skill is checked against the supplied transcripts, and a quote that cannot be found is removed rather than shipped
  • Sentence-like units in the measured transcript average 13.2 words, with 48 of 327 running four words or fewer and 10 running 35 words or longer
  • The Channel Skill Generator returned @maxinomics ready with four files in 6 minutes 30 seconds, reading 12 videos and using 9 of their transcripts

Key Statistics

  • The Channel Skill Generator returned @maxinomics ready with four files in 6 minutes 30 seconds, reading 12 videos and using 9 of their transcripts, measured on 18 and 19 September 2026.
  • Every quote field longer than 8 words in a generated channel skill is checked against the supplied transcripts, and a quote that cannot be found is removed rather than shipped, per prepublish-be/internal/infrastructure/ai/gemini_channel_skill.go.
  • The auto-generated caption track for Maxinomics' AI Was Supposed To Take Your Job. Why Hasn't It? holds 4,330 spoken words across 1,222 seconds, about 213 words per minute, measured with yt-dlp 2026.08.19 on 19 September 2026.
  • Sentence-like units in that transcript average 13.2 words, with 48 of 327 running four words or fewer and 10 running 35 words or longer.
  • Prepublish's 349-video hook study found that hits engage the title's promise 87% of the time against 78% for flops, while both reach the point at about word 14.

How to Teach an AI Agent Your YouTube Channel Voice

Adjectives do not transfer to a writing assistant. What transfers is already in the transcripts: the shape of the opening, the length of sentences, the words the channel never uses, and when a point arrives. Extract those, attach a verbatim line to each rule, and the agent gets constraints it can check.

Why telling an AI to write in your voice does not work

"Conversational but authoritative" describes a mood, and a mood fits almost any sentence. An assistant given that instruction, plus its own default register, has nothing to satisfy and nothing to fail.

A quoted line gives it a comparison to make before it writes the second sentence, and a banned string gives it something to search for. An assistant reads a skill file as text and cannot hear your delivery, so emphasis and pace never arrive. Diction and structure do.

The adjective brief and the evidence brief, side by side

Both briefs were written by hand for Maxinomics. The first is the shape most creators send.

Write in the Maxinomics voice. Conversational but authoritative. Smart and accessible.
Make it engaging from the start. Use clear, simple language. Build to a strong conclusion.
OPENINGS. Start inside a scene with a named place and an unnamed person. Do not name the
video's subject in the first sentence.
  Example, 0:00, verbatim: "On a sweltering June day in front of a velvet curtain in
  Barclays Bank in North London, the man took a piece of paper out of his pocket and
  pressed it into the beckoning slot on the base box."
  First reference to the present day, 0:40: "exactly like the conversation we're having
  today." Subject, 1:00, inside a question.

THESIS. Do not answer the title early. The answer arrives at 16:38 of 20:22 and carries its
own uncertainty marker.
  Example, 16:25: "And that statement I lean is critical. You should check my work. You
  should assume I'm biased."

BANNED. "in this section", "as we will see", "before we begin". Search the draft for each.

The difference is whether a failure is detectable.

Line in the briefWhat it rules outDetectable failure
Conversational but authoritativenothing specificnone
Do not name the subject first, example attachedopenings that start with the subjectcompare the first sentence to the quote
The answer arrives at 16:38 of 20:22drafts that state the thesis in minute onefind where the claim lands
Banned: "in this section"one exact stringsearch the draft

The first line cannot produce a failure, which is why drafts written from the adjective brief come back interchangeable.

The seven things to extract from your own transcripts

Everything below uses one real transcript: Maxinomics, AI Was Supposed To Take Your Job. Why Hasn't It?, published 1 July 2026, 20:22 long, 1,002,212 views when read on 19 September 2026. Captions were pulled with yt-dlp 2026.08.19 as the auto-generated English track, 4,330 words at about 213 words per minute. Quotes appear as printed, errors included.

1. Where the opening starts

By hand: copy the first 200 words of ten transcripts into one document, one row per video, and mark the sentence where the video first names its own subject. Keep only rules that hold in seven of the ten rows.

The Maxinomics transcript opens inside a scene with a place name and no mention of AI:

On a sweltering June day in front of a velvet curtain in Barclays Bank in North London, the man took a piece of paper out of his pocket and pressed it into the beckoning slot on the base box.

The present day arrives at 0:40 ("exactly like the conversation we're having today") and the subject at 1:00 inside a question. Narrative structures covers scene-first openings, and the transcript cannot show whether those 40 seconds are narration over footage or a person on camera.

2. Sentence rhythm

By hand: strip the bracketed markers, count the words in every sentence-like unit, and record the mean, how many run four words or fewer, and how many run 35 or longer.

The transcript holds 327 units with a mean of 13.2 words, 48 at four words or fewer and 10 at 35 or longer. At 0:29 the channel runs 50 words without stopping:

The murmurs turned into conversations, which turned into rumors, which turned into the nearly universal belief that the people whose job it was to hand cash to everyone through a window were definitely, without a question, about to be out of a job, exactly like the conversation we're having today.

At 0:42, ten words: "This was the birth of the ATM, and it was obvious." The recogniser places the punctuation, not the writer, so treat 327 as an approximation.

3. Vocabulary and banned words

By hand: search the transcripts for terms you believe you repeat and write the counts down, then search a list of script clichés and put whatever returns zero on the ban list. Keep it literal and short, because its only job is to be searchable in a finished draft.

Across this transcript: "you" 80 times, "I" 62, "we", "our" and "us" 35, "question" 7, and "so" opening a sentence 8 times.

Zero hits does not prove the channel never uses a phrase, only that the set you pulled does not contain it.

4. When the point arrives

By hand: find the sentence where the video answers its own title, note the timestamp as a fraction of the runtime, and repeat across ten videos. Write the rule as a timing and a hedge.

The title asks at 1:00. The answer arrives at 16:38 of 20:22:

So if I were to conclude, if I were to bet, put money on what will happen, the safest bet to make, the most likely answer to is AI going to take my job is no. But someone using AI will take your job if you do not use AI.

Thirteen seconds earlier the channel marks its own uncertainty: "And that statement I lean is critical. You should check my work. You should assume I'm biased." How to write a YouTube script covers where a thesis can sit. Placement does not separate hits from flops in Prepublish's 349-video hook study: 87% of hits engage the title's promise against 78% of flops, and both reach the point at about word 14.

5. Recurring segments

By hand: read the first three and last three minutes of ten transcripts, write down anything that appears by name, then search the set for that name.

Maxinomics closes with a named segment, introduced at 17:00 as "So here are the footnotes, the part that everybody keeps asking for." The sponsor read is a second segment, opening at 1:02 with "This video is brought to you by Communitier Coffee. More on them later." and closing at 12:01 with "Thank you to Cometeer for the coffee and for sponsoring this video."

Both spellings sit in one caption track. Copy a quote as printed, because a cleaned quote stops being evidence you can search for later.

6. Transitions

By hand: extract every sentence that ends in a question mark and every sentence that opens with "So", "Now" or "But", then read them in order. On the page they look like filler. In the video they are the seams.

At 3:51: "we must ask the question, what exactly is a job?" At 16:11: "Augment or replace? That is the question every single one of us is wondering."

The rule: change subject by asking the viewer a question, never by announcing the next section. The caption track cannot show the pause or the cut that marks the seam to a listener.

7. The call to action

By hand: search for "subscribe", "comment", "link", "below", and the name of anything you sell, and record every timestamp.

"subscribe" appears once, at 19:55 of a 20:22 video, inside the sign-off: "Thank you for watching. Like and subscribe, and I will see you when I get back from vacation. See you." There is no mid-video ask. A brief that says "include a call to action" invites an assistant to place one at the end of the first act; one that says a single ask, inside the sign-off, with the example attached, does not.

How to write the file the agent reads

The file is an Agent Skill: a directory with a SKILL.md and optional reference files. The specification requires a name of lowercase letters, numbers and hyphens up to 64 characters matching the parent folder name, and a description of up to 1,024 characters. Write the description in the third person, because it is the only part injected into the assistant's context before the assistant decides whether to open the file (Agent Skills specification, Anthropic's Agent Skills overview, read 19 September 2026).

The body carries one section per extraction above. A rule without a verbatim line does not go in: without a line, the rule is an impression, and an impression is what the assistant already has. Keep the body under about 5,000 tokens and push evidence into references, which cost nothing until they are read (Anthropic's authoring guide, read 19 September 2026).

Prepublish's generator caps it tighter, holding SKILL.md to 480 lines against the vendor recommendation of 500 and the name to 48 characters against 64 (prepublish-be/internal/domain/entity/channel_skill_render.go).

The automated path

The Channel Skill Generator takes a channel URL or a handle and returns the bundle. Three runs were measured on 18 and 19 September 2026. @maxinomics came back ready with four files in 6 minutes 30 seconds, reading 12 videos and using 9 of their transcripts, for a 9,623-byte SKILL.md. @veritasium took 5 minutes 35 seconds with 12 videos and 11 transcripts for 9,395 bytes, and a rerun of @maxinomics took 6 minutes 05 seconds. The pipeline reads the 12 most recent long-form uploads, caps each transcript at 8,000 characters and the set at 60,000, requires 4 usable transcripts, and allows one run per IP per day against a seven-day per-channel cache. Quote grounding runs in code: ai.ValidateChannelSkillQuotes checks every quote field longer than 8 words against the supplied transcripts, and ai.StripChannelSkillUnverifiedQuotes removes one that cannot be found rather than shipping it, recording the removal in the provenance block. The bundle is a SKILL.md beside references/hooks.md, references/titles.md and references/transcript-excerpts.md, confirmed with unzip -l, and the tool is at the Channel Skill Generator.

How to test whether the file worked

A test that means anything holds videos back. Pull ten transcripts, extract the rules from eight, and keep the two most recent uploads out of the set, because a rule taken from a video cannot judge a draft about that video.

  1. Load the assistant with the skill file and nothing else, and ask for the first 150 words of a script for the title of one held-out video.
  2. Put that opening and the video's real opening in one document with the labels removed. Write the order down before you read either one.
  3. Give the pairs to three people who watch the channel and ask which one the channel published, without saying that one was machine-written. Repeat for the second held-out video and two earlier ones, for four pairs.

Failure looks like this: readers name the machine-written opening in four pairs out of four. Two out of four is a coin toss. Record how often they pick the real opening at all, because a pair where the published one is unidentifiable carries no information. A second signal takes a minute: search the draft for every string on your ban list. One hit means the file did not carry the rule.

The whole procedure fits in ten minutes. The hook analyzer scores an opening against the same first-party dataset the study above came from, and it measures recognisability rather than quality.

What a transcript cannot capture

The transcript has words and not delivery. It records that a line was said, not how loudly, where the emphasis sat, or how long the pause lasted before the verdict. The channel's own footnotes make the same point: "Seth Laupus and Shinpei Ashen handmade all the incredible graphics you saw."

The timestamps are the timing of the words, not of the edit. A sentence can be delivered across three shots, and the caption track cannot show which words sat over which picture. Auto-generated captions also contain errors: this one spells the sponsor two ways and mangles the channel's own handle inside a URL. Extraction survives that, because shape and repetition hold across a misspelling.

The largest limit sits above all of this. A voice file constrains diction and structure and does not supply an idea. A draft that hits the mean unit length and opens inside a scene will still read as the channel saying nothing if the idea underneath it was not worth the 20 minutes.

Frequently asked questions

How do I teach an AI to write in my voice for YouTube?

Extract rules from your own transcripts instead of describing your voice. Pull the last ten transcripts, then record four things: where the opening starts, how long the sentences run, which words you repeat or avoid, and how late the video answers its own title. Write each rule with a verbatim line from the transcript underneath it, and put the file where the assistant loads it as a skill. A rule with a quote can be checked; an adjective cannot.

Why does telling an AI to write in my voice not work?

Because a description of a voice gives the assistant nothing to measure. Conversational but authoritative fits thousands of sentence orders, so the draft satisfies the instruction and still sounds like every other draft. A quoted line gives the assistant a comparison it can make before it writes the next sentence, and a banned term gives it a search it can run on its own output. The instruction is not too vague, it is unverifiable.

How do I extract a writing style from transcripts?

Work in columns instead of reading videos one at a time. Copy the first 200 words of ten transcripts into one document, one row per video, and mark where the subject first appears in each. Then count words per sentence across the same set, search for the terms you believe you repeat, and note the timestamp where each video answers its own title. Write a rule only for what holds in seven of the ten rows.

How many transcripts do I need to extract a channel voice?

Ten gives you enough rows to see which habits repeat and which are one-off. One transcript is enough to find an opening pattern in that video and not enough to know it is a pattern. The automated route takes the same idea further: the Channel Skill Generator reads the 12 most recent long-form uploads, caps each transcript at 8,000 characters, and requires at least 4 usable transcripts before it runs.

What should a channel voice file contain?

A name and a description that state what the skill does and when to load it, then one section per extracted habit. Every rule carries a verbatim line from a transcript and the timestamp it came from. Keep the body under about 5,000 tokens, and move bulk evidence such as thirty openings into a reference file that the assistant reads only when a task needs it.

How do I test whether the voice file worked?

Hold back the two most recent videos and leave them out of the extraction. Ask the assistant for the first 150 words of a script for a held-out title, then put that opening and the real opening of the same video side by side with the labels removed. Ask three people who watch the channel which one the channel published. If they identify the machine-written one every time, the file is not carrying the voice.

What can a channel voice file not do?

It cannot carry delivery. A transcript records words, not emphasis or pace or the pause before a punch line, and it never shows the edit. It also cannot supply an idea. A draft can match the sentence length and land its verdicts in the right places while saying nothing worth 20 minutes, and no brief fixes that. The file removes a category of wrong drafts, not the work of having something to say.

Related Guides

Free tools to put this into practice

Want to see how this reads on real channels? Browse the channel breakdowns. Each one compares script patterns across a channel's own higher-viewed and lower-viewed uploads, quoted from the transcripts.

See where your next script leaks viewers

Paste your script, get your scores and the biggest leak for free. No login.