Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 11:58
4.6x that video's typical replay level
issues. Uh I also invented OS certification. I just close the tracker whenever I want, so I have my life back. So, does this work? Yes, sort of. >> [laughter] >> Which leads me to act three, slow the down. Everything's broken.
Said at 11:52
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
The graph counts replays. It does not show where viewers stopped watching.
Words
2,978
Runtime
18:59
Speaking pace
157wpm
Reading time
12min
157 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Well, thank you for joining the final day of the conference here. I know it's been a long week. But I'm excited to talk about a topic here that is very near and dear to me. Software factories. Um You this is probably not a new concept to any of you. I'm guessing if you're all here, you're very familiar with software factories. You know, Ramp blogged about their inspects several months ago. I think maybe December last year.
79 words, the words spoken in the first 30 seconds at 157 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 174 |
| Average words per sentence | 17.1 |
| Longest sentence | 70 words |
| Questions asked | 14 |
| Sentences containing a number | 0 |
Most used terms
Filler phrases
131 in total: uh 33 · um 33 · like 23 · kind of 17 · actually 15 · you know 8 · I mean 1 · right? 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Well, thank you for joining the final day of the conference here. I know it's been a long week. But I'm excited to talk about a topic here that is very near and dear to me. Software factories. Um You this is probably not a new concept to any of you. I'm guessing if you're all here, you're very familiar with software factories. You know, Ramp blogged about their inspects several months ago. I think maybe December last year.
I forget exactly when. It's been a little while. And that really kind of like excited the entire industry of thinking like, oh, we can actually build software very differently now that AI is extremely capable of writing code. Um a lot of other companies have followed suit. Um and there's kind of like a standard set up here. Uh you have a sandbox that you put your code into. Um you put a an AI agent in there. Maybe a Claude code section or an open code.
Um and with a prompt, it generates a PR and you merge it. Um You might be building this already [clears throat] yourself. We have taken a slightly different approach at WorkOS. Um and the motivation for what we are building is that we see a lot of the kind of success metrics um really focusing on output, right? Like you'll see posts on X and blogs talking about the percentage of PRs or the number of pull requests or how much code uh is being generated by AI moving into production.
Um and this is just really output metrics that um you know, this can actually disguise how well these systems are working. Percentage of PRs or number of PRs may be lying in overall increase in PRs. It's very hard to distinguish whether the output is actually driving outcomes. And so that's kind of where we focus our efforts when we're thinking about a software factory. It's like if we're going to go put engineering time into building this level of automations and there's a bit of work that needs to go into this.
Full-time teams dedicated to this. How do we want to think about the success of this and whether this is actually driving value for our organization, our engineering organization? And so we measure instead of just code output, we're looking at outcome metrics. And that is really about like are we accelerating our ability to deliver features? We think that the dream of the software factory is that it gives each of our engineers a small engineering team for themselves.
And so the consequence of that really needs to be that we're building more and we're shipping more and that as we embark on complex features, those can get built a lot more a lot more quickly. So what does this actually look like? Well, we kind of started with the same thing you've seen in other other factories. We start with a sandbox we built on top of Cloudflare. We put an open code model router in there. We're feeding it prompts.
Um, and we kind of very quickly kind of ran into this situation where this wasn't driving more outcomes. This was not an incremental increase or an exponential increase over engineers just driving cloud code on their laptops. It was actually pretty indistinguishable for us. And so we really kind of started to think about, well, how can we embed our engineering processes into the factory itself? It's not sufficient that the factory is producing code.
We actually want it to take a lot of the other work that our engineers do to produce products and automate that as well. So, we kind of split this into two parts. One is we have the system we call TARS. This is a way for users to interact with the coding agent. And it is embedded into the tools that we use. So, not just Slack, but also Linear and GitHub. We subscribe to webhooks through TARS so that TARS can actually track the progress of projects in addition to generating the code outputs.
And then we built a separate system called Horizon. And this is the infrastructure orchestration layer. This is a lot more similar to Inspects and Minions and and some of the other systems that other companies have built. But, it sits in front of an MCP gateway. And I'm going to talk a little bit more about why that MCP gateway has been really transformational for us. So, by feeding our webhooks, activities that are happening in other uh source code systems and project tracking systems, we're starting to get to that level of autonomy where our factory is performing product engineering work, not just code work.
So, what's an example of this? In Linear, we can define tickets that have dependencies. So, one ticket blocks another ticket. Pretty common for, you know, breaking down uh large units of work into smaller ones. Because uh TARS is getting webhooks on ticket completion, it can automatically pick up the next ticket in a cycle. And so, we can do through our planning process construct like a map of how this plan we think this plan is going to get executed and TARS can start to execute on that autonomously.
The other thing that we have is in between those steps when a ticket is completed, we can ask TARS "Can you reevaluate the linear project and let me know if it's missing now tickets?" Because as you do work, you're learning about where the gaps in your plan. And we want to continuously keep our our plan fresh by using the agent itself. Because the agent is determining that there may be missing pieces to what we initially planned for the project and the goal that we have in mind.
So, we run this product engineering culture at WorkOS. This is where engineers are responsible for a lot of the product functions. We don't have product managers on teams today. Um and one of the rituals as part of this product engineering process is we create a hilltop document. It's a PRD and it is intended to both define what is the purpose of this project, but also incorporate um what are customers talking about?
Like where are we seeing the need for this unit of work? Uh it looks at competitive analysis. Like are there similar products out in the market today that we can draw inspiration from? We start to bring design early design screens into this. Um and we outline like what are the major milestones. And so, by kind of consolidating and canonicalizing this information into a document, this is something our product engineers have been doing uh throughout the entire uh company.
Um now we can ask give this resource to an agent and an agent can break this into units of work. Uh and so, this has been a really uh powerful way for us to take existing processes and encode them into the factory themselves. Like I mentioned that it's listening for web hooks. I'll show an example of what this looks like, but it it can see that the hilltop document through linear tickets has been reviewed and approved and by virtue of it being approved can start work on this project automatically.
So we don't need a human to be shepherding this agent through every single step of the life cycle. So we created an agent specifically for this part of the process. We call it the PM. And it is doing the first draft of the hilltop based on the brief specification that we give it. It's adding context to that in addition to reading the human reviews and then it picks up the implementation and breaking that into tickets.
And there's an opportunity for humans to stay in the loop for each part of this process. I mean it's very common for a human to intervene or comment on a ticket that's generated by AI, give it further guidance or refinements. Um but it's really about that like cold start problem where we can kind of get over the hump of creating all these different resources. Okay, so what does this actually look like? So here's a project I kicked off last month two months ago.
Um we're adding a new API to products I lead called Vaults. And with a command here in um in Slack I can kick off this project with a simple description, usually a few sentences of what I'm trying to accomplish. Tarsin goes in and creates all these resources for me. I don't need to go and create the project in linear. I don't need to create a draft of the notion document. Um all of the decision logs and the open questions.
Um Tars is creating a first pass at that. And so this then gives me an easy framework for me to step into as the lead product engineer and start giving more specification, rounding out areas that aren't well defined with knowledge that I had know about what we're trying to accomplish with the project, and then hand it back to Tars uh for execution. So, here's an example of uh the project. Um and it's like I said, it sets up these first milestones around our product engineering process.
So, we haven't done the hilltop review yet here. When that's done, we'll uh we'll we'll mark this ticket as complete and Tars will pick that up and move on to the next stage of project implementation. So, what we found is like a lot of teams are now doing more project brief and hilltop work because our agent can seed all that information. It's that blank page problem with writing. If you can come into a document that already has information, and you know, these agents aren't perfect.
There are times where it grossly overestimates what we're trying to accomplish with this project, and we have to cut out a lot of the scope that it comes up with. But, that's fine. That's a lot simpler for an engineer to add input into into rather than them spending time to set up all these primitives themselves. And then we can continue to drive this work, you know, through Slack, through Linear, um and ultimately uh results in PRs in GitHub.
The other thing that this lets us do is we don't necessarily always need to use our coding agent to implement uh pieces of work. Uh Devin is very popular. We have a lot of folks that really like using Devin. They can point Devin to these tickets and the same documentation to give Devin the context it needs to go and implement different parts of the work. Uh same with the local cloud code if you're just using Opus in a local harness.
Um it can grab through MCP all of these uh all of these pieces of documentation and use that as context for its work. And then, of course, we have people that collaborate in these project channels as well. You'll notice our security team will come in and look take a look at new projects, weigh in on the security implications, and so forth. So, I alluded to we we kind of early on in the the development of our factory is created our own MCP gateway.
Um, we call it our context engine. Um, this connects into all of our internal systems, but it also builds system prompts and context around what how to navigate these tools and when to use these tools. So, it is connected into Snowflake, which is our main data lake. We have semantic tables that we've built in Snowflake that describe product utilization or customer conversations. And then, we can provide an agent through our MCP server in the tool description list of here are the tables, this is the content they contain.
If you're trying to answer questions about this type of content, write a query for these tables. And so, it gives us a little bit of orientation to an agent to understand how to use how we use our tools, how we've organized linear in Snowflake, and give some guidance on where to find information. What has been really surprising about this is we actually now use this MCP server across a lot of other different internal tools.
We open it up for people to query directly from Slack to do data or customer analysis. And so, while we initially kind of envisioned that this would just be a way to connect our agent or coding agent to all of our systems. Uh this has been an incredible piece of leverage for a lot of our internal teams, and we're building a lot of other tools on top of this MCP gateway. So, if you're just getting started with thinking about a software factory, it's well worth your time to invest in an internal MCP gateway server that both connects all your tools, has descriptions that can tell the agent how to use those tools and how you specifically organize the information in them, and I think you'll find that that ends up being useful in a lot of other use cases.
Like I mentioned, we try to automate other parts of our software development. Um we have like bug bug requests that come in through Slack. Uh so, TARS through webhooks can listen to those and take a first pass at triaging and implementing opening PRs to fix bugs. Uh we have a lot of our customers uh in shared channels in Slack, as well. Uh we found TARS to be pretty useful to uh triage support requests that we're getting through them, because it can look at the code, it can often understand like where the customer might be running into problems with our product uh from the code perspective.
Um I showed you an example of how we kind of organize major features around this. Uh and then the last place I think is really exciting, we're all trying to get to the self-improving software, or the self-driving software. So, we're using uh TARS to actually build out our own uh sandbox infrastructure. So, we want to move off of some of the sandboxes as a service and own that infrastructure layer ourselves. Um Some of the reasons we have for that is we want really deep control over the session information, and uh be able to move workloads around different parts of our infrastructure.
Uh so what we're building out next is our memory layer and we want this to be an evergreen context uh about what every person in the company is working on, what team do they sit in, what products are they responsible for, and then at an organization level, what is the semantics about Work OS and how we do our work. And this we want to be able to plug into both our software factory, but lift that context out of our factory and into our other AI tools as well.
So, I think this is kind of our our major takeaway of where we've seen a lot of folks invest effort just around outputs and where we think we can actually get that human exponential value out of our factory. Can we be actually delivering values to our customers? We think a lot about shipping and customer impact and so we want to encode that into their our software factory itself. Um certainly worried about, you know, introducing instability into our products.
So, all right, can we measure uh defect rate and other um you know, time to recovery uh metrics? Um and then a lot of anecdotal, like we want to see our engineers using TARS as a sign that this is giving them value and accelerating their work. We want to see that they're moving off electing to move out of their local harness and into a cloud sandbox. And then the real goal is we can use this information because we own the infrastructure, we can see what's happening in the infrastructure.
We use this information to self-improve the factory. We want it to learn, we want it to get better about how it writes code. We want to see where um our engineers can improve their skills. There's so much moving quickly in AI. It's kind of like a constant race to keep up with the latest um you know, tips and techniques. We can actually point agents at sessions and see where are the gaps and how people are using our factory.
Where is there a skill that we should be building because the agent made a mistake. What skills are no longer relevant? Maybe the code has changed to a degree where a skill that we wrote six months ago is now obsolete. So, that continues verification. Semi-online is where we kind of see a big advantage to actually owning all the infrastructure ourselves. So, say it for the third time, it's worth repeating. Sandboxes are great.
It's cool to run an agent and have it open a PR. We really think about our software engineering practices and processes and we want to encode those in automation. I'm Ryan. I'm one of the engineers here at WorkOS. We have a booth right opposite the speaking area. If you are building software factories, I would love to talk to you. I would really love to hear how you're handling authorization. This is something I don't talk about because we haven't figured it out yet for ourselves, but maybe you all you all have some insights.
So, please come by the booth, talk to me, would love to chat. Thank you. >> [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.