What Is Inside a Generated Agent Skill File: A Real One
A generated agent skill file is a folder holding SKILL.md and a references directory, and the run torn down here turned 12 public uploads and 11 transcripts into four markdown files, with SKILL.md at 8,827 bytes. The brief is worth reading. Two of its sections describe things the writer never received.
What is inside a generated agent skill file
The Agent Skills specification requires one SKILL.md with YAML frontmatter and permits any other files beside it, recommending that the main file stay under 500 lines with the detail moved into reference files (agentskills.io specification, read 19 September 2026).
| File | Bytes in this run | What it holds |
|---|---|---|
SKILL.md | 8,827 | Frontmatter, twelve sections, provenance |
references/hooks.md | 566 | Opening lines with real view and duration numbers |
references/titles.md | 1,677 | Every title analysed, with date, length and views |
references/transcript-excerpts.md | 13,747 | The passages the quotes came from |
Measured on 19 September 2026 by downloading the bundle from the local stack: the four files as one markdown document come to 24,968 bytes with four file markers, and the zip is 11,926 bytes holding exactly four entries under maxinomics-voice/. The same folder passes the reference validator published with the specification, which prints Valid skill: maxinomics-voice.
The run this teardown is based on
The bundle is the newest ready row in the local run database, served from the seven-day cache when I requested it on 19 September 2026. The row was inserted at 23:41:30 UTC on 18 September 2026 and reached ready at 23:45:20 UTC, so the pipeline ran for 3 minutes 50 seconds. It read 12 videos and used 11 transcripts, none of them manual and all of them automatic captions, wrote with gemini-3.5-flash at prompt version 1, and recorded its own frontmatter timestamp as 2026-09-18T23:45:19Z.
| Channel | Result | Wall clock | Videos read | Transcripts used | SKILL.md |
|---|---|---|---|---|---|
| @maxinomics | ready, 4 files | 6 min 30 s | 12 | 9 | 9,623 bytes |
| @veritasium | ready, 4 files | 5 min 35 s | 12 | 11 | 9,395 bytes |
| @maxinomics (rerun) | ready, 4 files | 6 min 05 s | 11 | 11 | 9.4 KB |
| the run above | ready, 4 files | 3 min 50 s | 12 | 11 | 8,827 bytes |
The first three rows are the measured runs in the evidence pack for this batch, on 18 and 19 September 2026. The fourth is this one. Run time moves with how many caption pulls succeed and how long the model call takes, which is why the spread between them is nearly three minutes.
Section by section, and how far to trust each one
Every line below is real output, quoted from the bundle. Trust verdicts are per section, because the sections are not built the same way.
The frontmatter is the part a machine writes rather than the model:
---
name: "maxinomics-voice"
description: "Load this skill when writing scripts for Maxinomics to capture its analytical, history-driven, and economics-focused storytelling style."
metadata:
channel_url: "https://www.youtube.com/channel/UCUEmaeh13ai1ivT5wzz5Lhg"
videos_analyzed: 12
generated_at: "2026-09-18T23:45:19Z"
generator: prepublish
---
The name is derived by code from the handle through SkillFolderName and validSkillName, which reject anything outside ^[a-z0-9-]{1,48}$ rather than repairing it, so a name that would break the folder layout never reaches the file. videos_analyzed and generated_at are stamped from the row, and the prompt states that the caller stamps them, so the model cannot write a different number there. The description is the sentence the model wrote, limited by the prompt to 30 words naming the channel and when to load the skill. The Agent Skills specification treats that description as the text an assistant matches a request against, which makes it the most consequential model sentence in the file and the one worth editing if it reads wrong.
The twelve body sections arrive in a fixed order, and the evidence behind each one is not equally strong:
| Section | What produced it | How far to trust it |
|---|---|---|
| Identity and niche | Model reading of titles, descriptions and transcripts | A hypothesis. Useful as a first draft of your own positioning |
| Audience | Same, with no audience data in the input | A plausible guess. The model has never seen a viewer |
| Voice and diction rules | Model rule, with a transcript quote attached | The quote is machine-checked. The rule and its explanation are not |
| Hook patterns | Checked quote plus a model analysis sentence | Quote real, analysis soft, which the file admits in two places |
| Structure and pacing template | Model mapping of chapters and durations onto beats | Directional. The beats do not cover the videos they came from |
| Title conventions | Model reading of the title list | Conventions checkable against the same file. The examples are code-checked |
| Thumbnail conventions | Model inference with no image data in the input | The weakest section. Nothing in the payload shows a thumbnail |
| Recurring segments and catchphrases | Checked quote for a segment, unchecked counting for a phrase | Segment evidence solid, frequency notes unreliable |
| Topics covered and gaps | Model reading of the transcripts | Solid on coverage, opinion on the gaps |
| CTA patterns | Model wording with an unchecked frequency note | Wording can be real, the count can be attached to the wrong line |
| Banned moves | Absence claims over 11 transcripts | Patterns, not rules. Phrasing needs a human |
| Draft checklist | Model checklist over the whole brief | The most usable part of the file |
| Provenance | Numbers stamped from the row and the snapshot | Highest trust. The model prose here is its own list of gaps |
Here is the voice section, which is the one place where a quote reaches the file only after a machine confirms it exists in the transcripts:
**Rules**
- Use physical analogies to explain abstract economic or technical concepts. (It grounds complex ideas in everyday objects that the viewer can easily visualize.)
> If the Ghawar oil field is like the sun, a well in Texas is like a bottle rocket.
The hooks section shows both halves of the trust problem. In SKILL.md the promise is given a time, and in references/hooks.md the same opening carries real view and duration numbers:
> A simple coil of copper wire, one battery, one magnet. This is the world's simplest motor.
China Found Something Better Than Oil · 27:26 · 4.6M views · https://www.youtube.com/watch?v=BXLGV0Sj0n8
It starts with a simple, tangible physical reality before scaling up to global geopolitical consequences.
Promise paid off around 1:10.
Those four numbers are looked up from the snapshot stored with the run by videoByTitle, which matches the title the model wrote against the videos that were actually analysed. A title that does not match exactly produces a hook with no citation at all. The analysis sentence and the 1:10 are model output. The provenance block says what the 1:10 is worth:
The exact seconds to payoff for several hooks could not be precisely determined from the chapter listings alone, so a representative estimate was used based on the transition to the core topic in the transcript.
Our own study of 349 hooks found that hits and flops reached their first payoff at a median of word 14 and word 15 (the 349-hook study), so the missing measurement is not a serious loss. The risk is the reverse: a number that looks measured when the file itself calls it representative.
The beat table is the section most likely to be read as measurement, and it does not add up:
| Beat | Typical start | Typical length | Purpose |
| --- | --- | --- | --- |
| The Physical or Historical Hook | 0:00 | 1:10 | Introduce a simple physical object, historical event, or data point that serves as the anchor for the episode. |
| The Core Paradox or Question | 1:10 | 2:00 | State the central economic mystery or counterintuitive trend that the video will investigate. |
| The Historical Backstory | 3:10 | 6:40 | Go back in time to trace how the current system, technology, or geopolitical situation was built step-by-step. |
Three beats cover 0:00 to 9:50. Nothing in that table says what fills the rest, and the title table inside the same bundle lists durations from 13:25 to 27:26 while a pacing note states that episodes run between 13 and 27 minutes. The prompt told the writer to put its best reading of transcript order into the beats and to record the limitation in evidence_notes when the timings do not show a beat. The evidence notes do not mention it, so the gap is left for a human to notice.
Two sections are written from data the writer never received. The thumbnail section reads:
- Features clean, high-contrast graphics often depicting maps, industrial machinery, or key political figures.
- Uses minimal text, focusing instead on clear visual metaphors representing the economic conflict.
Those lines describe images. Nothing the writer receives is an image. The per-video payload carries a video id, title, description, upload date, duration, view count, tags, chapters, a caption source, a transcript opening and a transcript body. The prompt names what the model may use: titles, descriptions, tags, chapter lists, view counts, durations and transcript excerpts. The thumbnail conventions are inferred from titles and topics. They are the least verifiable lines in the bundle and the ones worth checking against the channel yourself.
The catchphrase and call-to-action sections carry counting claims with no machine check behind them. This run lists a catchphrase as This video is sponsored by and counts it in 11 of 11 videos, and its call-to-action row pairs the wording This video is sponsored by Zapier. More on them later. with the same count. The excerpts in the same bundle name Zapier once, Tastytrade twice and DraftKings once. The template is a sponsor line that recurs, so a count of 11 is believable for the template; attaching it to the Zapier sentence is not, and nothing in the pipeline checks it. An earlier bundle for the same channel, generated at 2026-09-18T17:07:05Z according to its own frontmatter, listed a different catchphrase altogether, It's not the going, it's the coming back., counted in 1 of the 11 videos. Two runs of one channel hours apart disagreed about what the creator repeats.
The rest of the brief is steadier than it looks. Title formula examples are printed only when titleInSnapshot finds the example among the analysed videos, so the formula list shows real titles or no example at all. The draft checklist is the part a drafter can act on without checking anything:
- [ ] Open the script with a concrete physical object, simple setup, or historical scene.
- [ ] State the central economic paradox or counterintuitive trend within the first two minutes.
- [ ] Include a detailed historical backstory that traces the origins of the modern industry or conflict.
The provenance block is the section to read first, because its numbers come from the run rather than from the model:
- Videos analysed: 12 (2025-12-23 to 2026-09-11)
- Transcripts used: 11 (0 manual, 11 auto-generated)
- Public channel data read on 2026-09-18 with yt-dlp.
**What the data could not show**
- Audience retention, click-through rates, and private analytics are not visible in the supplied material.
- The exact seconds to payoff for several hooks could not be precisely determined from the chapter listings alone, so a representative estimate was used based on the transition to the core topic in the transcript.
- Some call-to-action examples in the transcripts are brief sponsor transitions rather than direct viewer actions like subscribing or liking, which is reflected in the CTAs section.
One subtlety in the first bullet: 12 videos were analysed and 11 produced a usable transcript, and the titles table in references/titles.md has 11 rows because it is rendered from the snapshot of transcript-bearing videos. The file never says that, so a reader who counts rows and compares them with the frontmatter sees a missing title.
How a quote that cannot be found gets removed
The check runs in code, on the JSON the model returned, before any file is rendered. scanChannelSkillQuotes walks four fields, voice.rules[].example_quote, hooks[].opening_quote, structure.template[].example_quote and recurring_segments[].example_quote, and skips any value of 8 words or fewer because a short span is generic enough that a coincidental match proves nothing (channelSkillQuoteVerifiedAboveWords). channelSkillQuoteAppears normalizes case, punctuation and whitespace, and tests each transcript on its own so a quote cannot match by straddling two videos.
ValidateChannelSkillQuotes returns one message per failure, naming the field, the word count and the rejected text. The writer may make three attempts (channelSkillQuoteAttempts). From the second attempt the prompt carries a correction that names each failure and repeats the contract. The span must be copied character for character, contractions stay as written, no punctuation is added, and where no span supports the point the field is left empty and the gap is recorded in evidence_notes. If the last attempt still fails, StripChannelSkillUnverifiedQuotes empties exactly those quote fields, the caller appends a line recording how many were cleared, and the run still returns. The rule, the beat or the hook that the quote supported stays in the file. A brief whose quotes are empty is honest; a brief with a fabricated quote attributed to a real creator is not.
The boundaries matter as much as the mechanism. Short quotes are unchecked. Catchphrase phrases, call-to-action wording and the description and analysis sentences the model wrote are never checked at all. The title formula examples are checked by a different rule, against the snapshot rather than the transcripts. Anything in the file outside those four fields and the title formulas carries the same trust as any other sentence a language model wrote.
What a channel skill file cannot contain
The exclusions are deliberate. The prompt tells the writer that it cannot see audience retention, click-through rate, watch time, revenue or private analytics, and that nothing in the material explains why a video performed the way it did. The rendered file carries no performance field, no retention estimate, no view prediction and no statement about platform treatment.
Delivery is absent for the same structural reason: captions are text. Pauses, emphasis, tone, music, on-screen text and B-roll are not in the payload, so a brief cannot tell you how a script should sound, only what it says. Editing decisions are equally out of reach, because chapters and durations are the only trace of the edit the run collects.
Thumbnails are the third exclusion. A section exists, and it is written from titles and topics. No image, thumbnail URL or alt text reaches the writer, so the section cannot report what is in a thumbnail, only what the titles suggest a thumbnail might be.
Finally, the file cannot speak about the audience. It has no demographic, no comment data, no subscriber history and no per-video outcome. Every claim it makes about viewers is an inference from how the scripts are written.
What to add before you use the file
- •Read the provenance block first and decide whether the date range and the manual-to-automatic caption split describe the channel you think you are modelling.
- •Replace or delete the thumbnail section, since it was inferred without images.
- •Re-word the banned moves as descriptions of what the channel does not do in the videos supplied, because the file states absences as rules.
- •Check every counting claim, starting with catchphrase and call-to-action frequency notes, against the videos yourself.
- •Treat hook timings as estimates unless the file shows a chapter time, and ignore them if the provenance block calls them representative.
- •Read the draft checklist, which needs no verification, and keep it. That list is the part of the bundle with the highest ratio of value to risk.
Generate a bundle for your own channel at /tools/youtube-channel-skill-generator. It is free, needs no signup, is limited to one run per connection per day, and returns the stored bundle for seven days instead of spending a run.
Limits of this teardown
One channel and one run, at prompt version 1 and one model. Two bundles for the same channel generated hours apart disagreed about the catchphrase and did not agree on file size, which is the model varying, not the renderer: the renderer is pure and takes no wall clock, so the same stored result always produces the same bytes.
I could not re-run the quote check on this bundle. The run keeps the transcript openings in its input snapshot, 11,999 characters for the 11 videos, and not the bodies the model was shown, so verifying the check from the stored row alone is not possible. The evidence for the check is the code path and the fact that the file renders at all.
The trust verdicts here are for this bundle and this section list. A different prompt version can change what each section is built from, and the counts that are unchecked today are the ones to watch if that changes.
Try it on your own script
Paste your draft below. You get your hook, structure, and pacing scores, a script-level attention-risk map, and the single biggest issue quoted from your own lines. Free, no login.
Free · No login · See a sample audit first if you prefer.
Frequently asked questions
What is inside a generated agent skill file?
A folder named after the skill, holding SKILL.md and a references directory. SKILL.md opens with YAML frontmatter and then carries twelve sections: identity, audience, voice rules, hook patterns, structure and pacing, title conventions, thumbnail conventions, recurring segments and catchphrases, topics covered and gaps, call-to-action patterns, banned moves and a draft checklist. It closes with a provenance block. The reference files hold the evidence, and each one is written only when it has something to carry.
How big is a generated channel skill file?
The bundle torn down here had SKILL.md at 8,827 bytes, references/hooks.md at 566 bytes, references/titles.md at 1,677 bytes and references/transcript-excerpts.md at 13,747 bytes, as listed by the zip it ships in. Three runs measured for this batch produced SKILL.md files between 9,395 and 9,623 bytes. The size tracks how much evidence the model kept, not the length of the channel, and the body is capped at 480 lines with the excess pushed into the reference files.
Are the quotes in a generated skill file real?
They are checked in code, with two limits. Every quote field longer than 8 words in voice rules, hooks, structure beats and recurring segments is matched against the transcripts the run pulled, and a quote that cannot be found is emptied and the removal is recorded in the provenance block. Quotes of 8 words or fewer are not checked at all, and the catchphrase, call-to-action and title formula fields sit outside the check. Treat short quotes and those three fields as model text.
Can a generated skill file describe thumbnails?
Not from the images. The writer receives titles, descriptions, tags, chapter lists, view counts, durations and transcript excerpts, and no thumbnail file, URL or alt text is in that payload. The thumbnail section is therefore inferred from titles and topics, which is why it reads plausibly and why it is the least verifiable part of the brief. Check it against the channel yourself before you act on it.
Do I need to edit a generated channel skill before using it?
Yes, and the file tells you where. The provenance block lists what the data could not show, which in the run torn down here included estimated hook timings and call-to-action examples that are sponsor transitions rather than viewer asks. Fix the thumbnail section, re-word the banned moves as descriptions of the channel rather than instructions to a drafter, and verify any counting claim such as how often a phrase recurs, because those counts are model arithmetic and are not machine-checked.
Does a channel skill predict how a video will perform?
No, and the writer is instructed to stay away from performance. The prompt tells the model it cannot see audience retention, click-through rate, watch time, revenue or private analytics, and the file carries no performance field. The brief describes how a channel writes across its recent uploads. What that implies for a new video is your reading, not the file's, and no section of the bundle makes a claim about views, growth or platform treatment.
How current is a generated skill file?
As current as the twelve uploads it read. The provenance block names the upload date range and the day the public data was read, which in the run torn down here was 2025-12-23 to 2026-09-11 with the data read on 18 September 2026. The frontmatter carries a generated_at timestamp, and the service caches a finished bundle for seven days per channel. A channel that changes format makes the file stale, so read the date before trusting the conventions.
Related Articles
We Analyzed 349 YouTube Hooks. Most Hook Advice Did Not Survive.
The first 45 seconds of 349 videos from 36 channels, 254 million combined views, overperformers against flops. Speaking pace and time to payoff showed no difference. Concreteness did.
YouTube MCP Servers Compared: 20 Public Servers, Verified
Twenty public MCP servers connect an AI client to YouTube, and eighteen of them fetch transcripts of videos that already exist. Only one accepts a script that has not been recorded.
YouTube Transcript Extensions Compared: 15 Listings, Verified
The YouTube transcript extension category splits in two. Summarisers send the caption text to a server and return a digest, which needs an account and a provider that holds your text. Extractors read the caption track YouTube already served to your player and hand you the text.
Free tools to put this into practice
Want to analyze your own scripts?
Start Script Analysis