Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Simon Høiberg · @SimonHoiberg
Words
3,480
Runtime
20:08
Speaking pace
173wpm
Reading time
15min
173 words per minute, between the 160 25th percentile and the 181 median of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Most AI agents work great in a demo and then completely fall apart when [music] you give them a real job inside your business. The research comes back without sources. The code is super sloppy and the content sounds so generic that you end up rewriting the whole thing yourself. All those amazing use cases you saw on Twitter and YouTube just doesn't work that way for you. So, you reach for a better, more expensive AI model and almost kill yourself squeezing every [music] output token
87 words, the words spoken in the first 30 seconds at 173 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 256 |
| Average words per sentence | 13.6 |
| Longest sentence | 37 words |
| Questions asked | 3 |
| Sentences containing a number | 2 |
Most used terms
Filler phrases
31 in total: like 15 · actually 12 · kind of 2 · basically 1 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Free, no account. See where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most. Or run it on the words above first.
Free · No login · See a sample audit first if you prefer.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Most AI agents work great in a demo and then completely fall apart when [music] you give them a real job inside your business. The research comes back without sources. The code is super sloppy and the content sounds so generic that you end up rewriting the whole thing yourself. All those amazing use cases you saw on Twitter and YouTube just doesn't work that way for you. So, you reach for a better, more expensive AI model and almost kill yourself squeezing every [music] output token you can.
And it's so unnecessary. Listen, I run multiple SaaS products and a YouTube channel with Tiny Tea. AI agents help us with support, product development, and marketing, including creating these [music] videos. I have moved those agents between GPT, Kimmy, GLM, Fable, and even a few much less capable models. They all produce different work, and none of them gets to decide what our company accepts, because it's not really about the model.
It's about the harness you create around them. We have five layers of custom harness we built around our AI agent team. And in this video, I'll show you those five layers and how to apply them. So, changing the model does not also change your company standard. Okay, let's start with the most annoying problem I know everyone has. You can give an agent a task and a set of rules, and the agent can repeat those rules back to you, and then still break a bunch of them while doing the work.
You tell it that it clearly broke the rules you gave it, and it'll casually say, "You're absolutely right. That was on me." Infuriating. And here's the weird part. Even if you show it the finished work afterwards and ask it to look for mistakes, it can often spot the exact things it just did wrong without you even telling it. The model understands the standard. It just has a hard time doing the job, following every instructions, and judging the result at the same time.
So, we split that into separate jobs. The first layer is an eval classifier. If you've written software before, think of it like a unit test for AI work. A unit test checks that one function gives you the result you expect. An eval checks one piece of work against one standard and returns a simple verdict. Some evals are completely deterministic. They are normal code checking something objective, like whether a required field is missing or a test failed.
Other evals need judgment. This is when we use another AI model as the classifier, and setting this up is much simpler than it sounds. We call a model with a system prompt that gives it one job, inspect this input against these rules and classify it. The producing model sends in the work. The classifier gets that work, the rule it has to check, and a few examples that should pass or fail. Then it reduces the result to one class, usually approved or rejected, and gives us the reason and the evidence behind that decision.
You can use a score instead, too. Either way, the output stays narrow and predictable. Let me show you how we use this for these YouTube videos. Agents help us write the scripts, and every script is split into chapters. Before a chapter moves forward, we run three independent checks. One looks for generic AI writing. We call this the slop gate. Stuff like it's not X, it's Y, the part everyone's missing, using cliché metaphors, cringe AI buzzwords, and so on.
The second one checks whether the chapter sounds like something I would actually say. It compares the writing with a large collection of my previous scripts and outputs whether the chapter feels familiar in terms of style of writing and tone of voice. And the third checks whether it does its job in the video. An intro has to create tension and make the promise clear. A middle chapter has to move the argument forward and make you want the next part.
It uses previous scripts, but also more official sources on how to write well for YouTube. If one check rejects the chapter, we get that exact reason back. The first check is a good example of a deterministic eval. It uses a blacklist for phrases like this is the part where and simple regular expressions to catch writing patterns like it's not X, it's Y. No AI is needed for that check. Normal code can find it. The second and third checks need much more taste and judgment.
So, these are using AI. We give a model the chapter, its job in the video, and examples of chapters that worked or failed. Then it judges whether the chapter actually sounds like me and does its job. These evals often send the writing agent back to redo its own work more than a hundred times on average before a script ever reaches me. Imagine having to get involved every time. I would have to sit there, point out some stupid little writing problem, wait for the revision, find the next one, and keep doing that more than a hundred times before I could give it meaningful feedback.
Now the entire loop happens automatically between the writing agent and the evals. The script is still rarely perfect when it finally reaches me, but it is infinitely closer. I can spend my time on the argument, the taste, [music] and whether the video is actually good instead of nitpicking the same small mistakes again and again. We use the same principle with all work. We have evals everywhere for YouTube scripts, code, [music] support replies, research, literally everything.
Without a proper setup like this, even GPT-5.6 and Fable will produce sloppy work and drag you into a tedious, hair-pulling review session. Whereas, with a proper setup like this, you can have agents run on the older Opus models or even smaller models like Quinn or DeepSeek and still produce very good work. Now, the second problem is more sneaky. An agent can finish a task without returning the information the next step actually needs.
And this one is nasty because the result often looks complete. The writing is clean, the answer sounds confident, nothing is obviously broken until another agent or one of your employees tries to continue the job. Then the whole handoff falls apart. A critical piece is missing. The evidence was never saved. The agent gives you a conclusion but doesn't show enough of the work for anyone to continue. Now, someone has to repeat the task just to reconstruct what done was supposed to include.
Completely useless. The agent didn't necessarily ignore the task, but you gave it a vague definition of finished work so it filled in the blanks for you. Think about it. Would you ever hire a contractor with an agreement that says, "Renovate the office and make it look kind of nice."? Of course not. You agree on what gets delivered, which evidence you need, the format, and what has to be finished before the job is complete.
Agents need the same thing. So, layer two is a schema contract. A schema contract sounds very technical. It really isn't. It tells the agent exactly what to return in a format our software can check. OpenAI has an implementation called structured outputs. You give the model a JSON schema, which is basically a list of fields, what belongs in each one, and which fields cannot be skipped. Let me give you an example. A customer reports a bug.
One support agent investigates the issue and then hands it to a technical agent or human support staff. That handoff cannot just say, "The customer cannot log in. Please investigate." What is the next agent or employee supposed to do with that? We need the customer ID, we need the account ID, we need the exact steps that reproduce the problem, we need what should happen, what actually happened, and the relevant logs or traces from that failed request.
All of this should be defined in the schema contract. Now, marking a field as required only means the field has to exist. So, we add tighter rules, too. The IDs have to match the right format, the reproduction steps cannot be empty, and the handoff has to include at least one log or trace. Normal software checks all of that. If the account ID is missing, the reproduction steps are empty, or the agent forgot the traces, the handoff is immediately rejected.
That boring little check saves the technical agent from opening the escalation and realizing it has to start the entire investigation again. Now, it gets the context and evidence in predictable places. It can actually work on the bug instead of interrogating the previous agent. It also puts the claimed cost and the logs in predictable places. So, a separately designed e-val can judge whether those logs are actually relevant and support the claimed cost without digging through a long conversation and guessing what the support agent meant.
We use schema contracts for research, code delivery, sponsor qualification, and production handoffs, too. The fields change with the job, obviously, but the principle stays the same. The agent does not get to invent its own definition of done. Of course, the schema can still be filled with garbage. The logs might come from the wrong request, or an agent might interpret them incorrectly. That's why we still run the e-vals separately after validating that the delivered work is contractually valid.
So, once an agent returns complete work with evidence, people usually make a much more dangerous mistake. They give it access to everything. They connect the general database tool, hand over broad API credentials, and then write, "Only use it when necessary" in the prompt, and hope the model behaves. That's insane. Normally, you wouldn't even give human employees those kind of privileges, and the agent now has to guess which operation to use, and one bad guess can reach far beyond the job you gave it.
Imagine hiring someone to collect a box from one storage room and giving them the master key to the entire office. Nobody does that. Their key card opens the room they need. Some doors still require approval. Security people call this the principle of least privilege. NIST's definition is simple. You only get the access you need for the job. So, layer three is purpose-built tools. A purpose-built tool exposes one narrow company action.
Normal software handles the permissions, validation, filtering, and formatting around it. To the model, the tool is just a named action with a description and a few allowed inputs. The model can request a tool call. That's just a structured request to run the action. Before anything executes, our software checks the inputs, the agent's permissions, and whether a human has to approve it. We use this with partnership inquiries for this YouTube channel.
Companies contact us about sponsoring a YouTube video, and an agent helps qualify the conversation and prepare the reply. That agent has a small set of tools. It can list our active sponsor conversations, it can fetch one specific conversation when it needs the context, it has access to our content schedule and allowed price ranges. It can prepare a reply for us to approve. It can't browse every unrelated record in the company, pull a random customer conversation, or get impatient and send the sponsor reply behind our backs.
No tool exists for any of that. Same with product development. No AI agent can accidentally push broad code changes to production or write directly to the production database. Mistakes are simply not possible. The system prompt isn't our security boundary. The capability simply isn't available. We built narrow tools like this for support, publishing, development, analytics, and finance. Each agent gets the tiniest set of actions its job actually requires.
This is particularly important with frontier models like GPT-5.6 and Fable cuz they are incredibly smart and very good at finding their way. And they are often very, very eager to go out of their way to handle the task you gave them. And that's exactly where these hard blocking boundaries are crucially needed. Now, a quick side note because there's another piece people forget. None of these layers helps if the software your agents need has no proper way for them to use it.
The agent [music] either can't do the work or you give it some horrible workaround that ruins these boundaries and defeats the whole purpose. This is [music] why every SaaS product in my portfolio is being built for humans and agents with [music] proper APIs, MCPs, and connectors for OpenClaw and CloudCode. If you want your AI agent team helping with content, FeedHive lets agents and humans create, review, and publish together.
AidBase does the same for customer support. It has a public API, customizable ticket forms, and inboxes. So, agents can join the support flow without [music] forcing your customers or your team into some weird homemade setup. And here's the best part. You don't have to pay a monthly subscription for this. With Founder's Deck, you can get lifetime access to FeedHive and AidBase plus two other tools for a single one-time [music] purchase.
You pay once and you keep access to all four tools forever. No monthly subscriptions for every product. Check out the description. I'll leave a link to Founder's Deck here. Now, safe tools still don't guarantee a safe process. Models are probabilistic. Give one the same job twice and you can get two different decisions. That's useful when the agent needs to interpret research or rewrite a weak paragraph. It's absolutely horrible when the agent decides that an approval feels optional that day.
A convenient model might skip a fail check, publish early, or move forward because the draft looks good enough. And once the action happens, an apology from the agent is worth nothing. In fact, hearing my agent say, "You're absolutely right. That's on me." honestly just pisses me off even more. Think about trained pilots. Pilots use checklists even after thousands of hours in the cockpit. They still make judgment calls, obviously.
But takeoff doesn't depend on whether the pilot happens to remember every critical step that day. We do the same thing around our agents. So, layer four is deterministic workflows. The AI makes judgments inside the job. Normal software controls the sequence of the jobs. Take product development. Say an agent is fixing a bug in the FeedHive post composer. The most dangerous result is a completely reasonable-looking patch that changes something near the bug without ever proving what caused it.
So, the first workflow state is reproduction. A state is just where the job is right now. Before the task can move forward, the agent has to reproduce the bug and return the exact steps plus evidence of the failure. Then it has to write a regression test and show that the test fails. An eval or a human reviewer checks whether that test actually captures the original bug. Only after that approval can the workflow move on into implementation and let the agent change the code inside its isolated workspace.
When the agent says the fix is done, the same regression test has to pass. The broader test suite has to pass, too. And the agent has to run the original reproduction steps again and attach the result. If any of those checks fail or the evidence is missing, the workflow sends the task back. The agent can decide how to fix the code. It cannot decide that reproducing the bug was unnecessary, conveniently swapping an easier test, or declare victory because the patch looks sensible.
Once all of that passes, the change moves into review. We can inspect the actual diff, the failing test before the fix, and the passing evidence afterwards. Only our approval lets the workflow merge the code. Those allowed moves are called transition rules. They define which state can follow another one. That means software enforces the checklist instead of leaving it as another instruction the model can interpret or ignore.
We use deterministic workflows for support, escalation, sponsor replies, social publishing, and these YouTube videos, too. The model gets freedom where judgment is useful, but the company keeps control of what is allowed to happen next. I don't care how intelligent or confident the model sounds. It never gets to improvise our company process. The last layer came from watching the same stupid AI mistakes return after we had already fixed them.
An agent would fail. We would correct the output, add another instruction to the prompt, and move on feeling like the problem was solved. Then we change the model or one of the tools a few weeks later, and the exact same issue came back. Now you get to pay for the same lesson twice. Someone has to find the failure again, explain it again, and patch the system again. And if nobody catches it this time, the bad work moves forward as if you never fixed anything.
We had this with research sources. An agent returned a claim with all the fields from the schema contract. It looked complete, but the source did not actually support the claim. Fixing that one result is easy. You replace the source, tighten the prompt, and continue with the work. Three weeks later, we change the model, the new model makes the same mistake in a completely different document, and nobody remembers the exact case that exposed it the last time.
Aviation doesn't rely on every future crew remembering one old incident. The lesson becomes part of the checklist, the training, or the aircraft design. We wanted the same thing for our agents. So, layer five is failure replay. Developers call this regression testing. You keep an old bug as a permanent test so a future change can't bring it back. So, when that research source failed, we saved the original research task.
We saved the bad claim, the source that failed to support it, and the exact reason our eval should have rejected it. The first replay is super simple. We feed that exact bad claim and source back into the current eval. It still has to reject the future failure we already know about, so whenever we update that eval, we run the old cases again before the change goes live. Then, we test the producing agent. Imagine we want to move research to a smaller query model.
Before that model touches a new video, we give it the original research task again. It has to produce fresh research with the required claim, URL, and supporting excerpt. Then, the eval checks whether the excerpt actually proves that claim. If the model fabricates the source, quotes the wrong passage, or gives us another confident claim with no real evidence behind it, the switch fails right there. We adjust the prompt or the harness, run the same painful case again, and keep doing that until the old failure stays dead.
We do this for all the failures that have actually hurt our work. Missing delivery evidence, the wrong tool being selected, agents trying actions outside their permissions, and sloppy writing patterns, too. Public benchmarks can tell me whether a model looks impressive on somebody else's test, but before I change a model, I care more about our own failure library. Can it survive the mistakes that have already cost us time inside my company?
This is the test I care about before I let any new model anywhere near my business. I know, if you're new to this, it probably sounds like a lot. But, I can tell you for sure, whenever you see companies doing really well with AI, it's almost never because of the model and always because of the quality of the harness. This is what makes the entire difference between fighting AI and using AI to produce high-quality outputs on par [music] with what a human employee could do.
Better model can improve the output, but it still doesn't own the company standard. Your company does. >> [music] >> If you already have an agentic system like Open Claw or Hermes set up, I suggest using this video as a reference. Pull out the transcript, give it to your agent, and have them help you set this up. >> [music] >> Ask it where these layers would benefit your company operations. Then, have them set it up and tailor it to your use.
And if you want access to agent-friendly tools for social media automation, customer support, and much more, check out Founder's Stack. The link is in the description below.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.