YouTube transcripts

How to Build a Software Factory for AI Coding Agents: video thumbnail

How to Build a Software Factory for AI Coding Agents transcript

Boundary · @boundaryml

Published August 28, 20261:10:4560.2K views

Watch this video on YouTube

Transcript analysisComputed from the caption text

Words

15,403

Runtime

1:10:45

Speaking pace

218wpm

Reading time

64min

218 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.

Opening (first 30 seconds)

You start with the normal software factory where a human builds the thing and then you swap it out with agent builds the thing. There's a couple different options, right? You could like build everything inhouse and then there's like the other side of the spectrum which is like buy everything. >> In terms of buy versus build, where do you land? >> The real answer is like you should have as an engineer choice and the ability to like work in open systems and plug these things together. So this was a really fun episode where we talked about software factory design patterns. We've talked a lot about

109 words, the words spoken in the first 30 seconds at 218 words per minute.

Sentence shape

MeasureThis transcript
Sentences920
Average words per sentence16.7
Longest sentence165 words
Questions asked68
Sentences containing a number23

Most used terms

  • uh129
  • yeah128
  • um85
  • basically70
  • actually62
  • code57
  • stuff47
  • build44
  • run43
  • okay41
  • agent39
  • lot39

Filler phrases

1,271 in total: like 788 · uh 129 · um 85 · basically 70 · actually 62 · kind of 58 · right? 21 · I mean 19 · you know 19 · literally 12 · sort of 8.

A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.

What this transcript is

Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.

Transcript

You start with the normal software factory where a human builds the thing and then you swap it out with agent builds the thing. There's a couple different options, right? You could like build everything inhouse and then there's like the other side of the spectrum which is like buy everything. >> In terms of buy versus build, where do you land? >> The real answer is like you should have as an engineer choice and the ability to like work in open systems and plug these things together.

So this was a really fun episode where we talked about software factory design patterns. We've talked a lot about the broad software factory. And today we zoomed in on this part where it's like agents building the thing and testing the thing and getting feedback. We talked a lot about how the boundary in-house uh software factory works and the architecture of it and contrast it to like sort of full stack cloud agents where you just buy everything and everything is managed for you.

Um, so the architecture here was really interesting and then we kind of dug into like the four layers of the enterprise software factory, what the stacks are and some actual like implementation examples of like how you can mix and match parts of this stack and how different like open- source and closed source products like let you bring your own compute, your own dev environment, build or buy the harness and build or buy the uh control plane and orchestration.

So a lot more detail there. We talked a lot in detail about the boundary software factory. Super fun. And then uh as a little bonus at the start you can see our take on the Financial Times article of uh hey is OpenAI taking market share from Anthropic and why and why are we also codeex build. So super fun episode uh excited to get into it. Uh let's go. Hey folks this is AI that works. We talk about AI that actually works uh beyond the demo for real production use cases enterprise whatever it is you want something that actually works well and is reliable.

This is uh all we do. We do a lot of systems design. We do a lot of context engineering. Um, the very first uh conversations that led to the term context engineering existing happened on this show. Uh, I'm going to stop yapping. I'm Dex. I'm the CEO and co-founder of Human Layer. We build a multiplayer coding agent workspace. And Vibb here is >> I'm Vibb. I'm one of the co-founders at Boundary and we build a programming language designed for agentic coding. >> Amazing.

Check it out. boundaryml.com. Is that right? Did you guys not get baml.com? Uh sadly, Bank of America has a slightly bigger war chest than us. >> We'll get it. That's your That's your goal, right? It's like by the end of the uh >> We are slightly bigger than them on SEO for Google, though. So, we do beat them on in almost every region. Yeah. So, if you Google Bam, we're now rank one or rank two and rank three and they are on like page two, I think. >> Um >> amazing.

We have go Google BAML uh and help their SEO even more. Click the Boundary ML website. Go check it out. Programming language for agents. I've been building a bunch of internal tools in it. It's a lot of fun. Uh you can try human layer at humanllayer.com. Um someone asked me today about this article. Did you read the uh Financial Times article about how uh Fable and Opus are actually like losing market share and Codex is taking over. >> So [clears throat] like if you think about it, this is kind of interesting.

I mean, I'll tell you honestly, like my mom knows nothing about Claude or Anthropic as a company, she only knows Chad GBT. >> Yep. >> If she ever wants to go ahead and build an agentic stuff, she would only ever use Shad GBT cuz it's the closest she can get. So, I suspect it's just a relationship of building a consumer brand versus an enterprise brand. like no matter what you do at some point if more people get exposed to your topic the consumer brand will grow bigger >> and that's I suspect that might be it >> I completely disagree I think >> really >> I think I saw the change in like December and January of this year where all of the best agent decoders the ones that are like >> 6 to 12 months ahead of the like SF meta they all switch to codeex So, okay, that's a totally separate thing.

Uh, like Codeex is definitely I mean I switched to Codeex, too. >> Like, and I think >> No, yeah, because it's better. >> Well, I was actually hesitant for a long time cuz I could only get Claude to work, >> but for some reason I I kept on seeing I was like I saw Theo switch to Codeex, Dylan switch to Codeex. A lot of these people on Twitter that I was watching only use codec switch. Peter, well, >> that's a Peter was early.

I had a conversation with Peter Steinberger in October of last year. >> Peter have totally different incentives. >> This was way before openclaw or anything. He was like when he was I mean he was working on like 40 side 50 side projects. This is before openclaw launch but I was like yeah we built this whole research plan implement system for claude and like it makes Claude so much better and he was basically his take was just like this was codeex five maybe uh maybe >> I don't know if codex is better than a fable back then or back then.

No, it was Codex back then is not better than Faber now, of course. But he's basically like all that research upfront [ __ ] that you have to tell Claude to do, you're just compensating for the fact that you're using a worse model and if you like you ask Codeex to do something, it will just read stuff for longer, which is exactly what you want. >> Yeah. Um, so have you guys switched fully to Codex? >> We I so I I will use Opus models for like generic knowledge work type stuff like running the company and like like managing pipeline and things like that.

Um, it's just like a little bit better at like tool use and things like that. Like barely noticeable. It's just like what I'm used to. Um, but 90 or 95% of the features I write are written with uh Terra or Soul. >> You know what's so funny? >> Yeah. >> Um, you know what I want to see being done is you know how people are like talking about startups like there's all these people on Twitter's like uh vibes slopping about how much revenue they make and people are like just post the revenue or it doesn't count.

What I kind of want to see with people talking about AI models is I kind of want to see people posting their token usage or it doesn't count. It's like you can't say you're a Codeex man unless you like actually like for example you you can't say you're a CEX man until you post your token usage and prove that you're a CEX man cuz so many people just follow the hype without doing it. So I can never tell like for you I know you actually do it like in your case, right?

That's Oh yeah, show us. >> I'll show you the token usage. There you go. >> Let's see. >> This is cool. >> Um, let's find Yeah. cost cost by model. I mean, so this is by spend. >> Um, but yeah, we're using a [ __ ] ton of codecs. >> There you go. Um, yeah. It's also interesting. It's not like you have stopped using the Opus models. It sounds like you use it more targeted though. >> It's a more like things that are UI focused will still use Fable or Opus.

Like, okay, >> if I have Fable credits, I will let I will let Soul drive and then if I find a shitty UI, I will I will I will kind of like, okay, throw this out and then we'll we'll have it either invoke a Fable sub agent or I'll just start a new session with Fable. >> Yeah. I mean, what I find is actually like so I've been doing a lot of UI work myself as well, and for that work, I also end up using like the Opus models because like I feel like for whatever reason, Codex sounds a lot more like AI than Claw does when it's writing content >> when I'm writing things with it.

And like or at least maybe I just know how to tweak codeex better to sound more like me or claude to sound more like me than I know how to tweak codeex better. Codex is kind of like but for engineering task it's not even a question like codeex is execution codeex is better. >> Yeah. Whether it's codeex or claude like I don't use I don't let the models write much. I basically like here's a transcript of a thing. I'll feed it like a transcript of an episode of this show and I will um give it a um I'll be like write the headers and write a little snippet like a a markdown comment about what goes in this section and the reference to where in the transcript it's from.

And I'll have it write one little literally like two or three sentences at a time. And I'll polish them. And in one session I will have it write the thing. I will change it. I will be like, "Read what I changed it to. Do the next section. Don't [ __ ] it." And then like after four or five iterations, you start it starts in one session, you can start to give it a little bit of a sense of a tone in the patterns to follow. >> And then sometimes I'll flush that out to a doc that is like here's all the anti-atterns that you do by default.

But it's like my writing process is like I cannot stand like three pages of Claude or or codec slop. Like I just I won't read it. I will. So, like I have to do it really small increments at a time, otherwise it's overwhelming. >> Because otherwise, you just don't read. >> Otherwise, I'm just like, "Oh, I [ __ ] hate this. I might as well write this from scratch." And then I'm like, "Well, I don't have time to do that." Like, how do I get how do I get leverage from AI?

How do I get it to like kind of like unblock my writer's block without me actually having to like without actually like having to fix every single issue that it that it introduces. >> I found that I can't really have my whole team switch. I still don't have a good way to convert any humans to like force them to go do the >> like to either model like to codeex or to claude. And the reason is like >> I think a large part of how these AI models work today and these AI workloads work is like they're kind of personal to every human and like for example like obviously like we land on two different spectrums of how we like to engineer.

I I like to sit and think really hard about the problem for a while and like try and like remove as many roadblocks up front and you do a lot more iterative approach. Neither approach is necessarily bad or good, but like clearly like if I and like it's also not like both of us live in one world at all times. We live on the spectrum. We just have a tendency on one end or the other. But like if we for if you were forced to always plan upfront and do every single amount of planning up front and I was forced to always be iterative, we would kind of not enjoy the work as much.

And I think that's the problem with cloud and codeex is you as an engineer got used to a workflow. It's like [ __ ] now I have to change again. [ __ ] I have to change again. I have to like get the tweaks and the nuance of the model right every single time. >> Yeah, there's intuition that you have to build up and you need it's almost like a separate mode of being like you're like, "Okay, I'm going to be less productive and the the the benefit is not what I actually produce.

It is going to be the the intuition I build." >> Have you ever switched keyboards? Like every time you switch keyboards, like you're just like slightly less productive for a while. Like when I first used the Kinesis, I was like >> I literally just lost three months of my life. >> Like because I could not type like I used to type. I just feel so dumb for that three months until I finally got used to it. And I think that's kind of what switching models and providers are like. >> And yet I can't convince you to try Vim. >> Uh Vim is so hard.

Vim is I cannot for the life of me understand HJKL. Like what the like this was designed for an era with no mice. >> Can I show you? Can I show you something? >> Okay, while you do that, let me read some of the questions that are popping up. Um, Matas was asking, "Do you guys use Soul with sub agents or just soul alone?" I use Fable for code writing, research, and web search sub agents and it beats Soul by a mile, but within the same harness um uh without using uh sub agents like I don't know how you make anything happen.

So, here is my most recent workflow for big stuff is like >> basically like I do some research and create a product outline back and forth and build a canvas that looks uh something like let me see if I can pull something up. Uh oh, I updated um let me find a good PRD. Uh show me a workflow. >> V3 task launcher. No, not this one. Actually, this probably has it in there. The Yeah, the PRD. So, I will work back and forth and build like a bunch of like interaction mocks with the model.

Like, okay, here's how it works. Here's the different parts of this. Here's the visual stuff that we need to add. >> Not fully in like again like these are HTML mockups. They don't have to be faithful. It's just like get the direction >> and then I will have Soul and Fable implemented overnight. I'll literally queue up a fable session and I'll be like implement and then I'll cue another prompt that is like okay review it.

Okay, another prompt and it's like use the codeex CLI to have C codeex review it and then compact and then do it again and then do it again and basically like queuing up like 15 prompts to let them just argue it out overnight >> and then I look at what they built in the morning and I look at at a human in the loop like polish okay make this look like this this thing is broken and then once that's ready I'll go make the PR and if it's 20,000 lines of code I will actually work back and forth with the models are really good at this like how do I break up this work into separate PRs where do the migrations go and then we merge like a one to 3,000 line PR at a time. >> Okay.

But I've Okay, so I've got a question for you. When you do this workload, um >> what's your wall time of landing it? >> Like all the work? >> It's It's hard to say because I am I will have like bursts of like two days of a lot of coding and then I'll have like four days of like doing like CEO [ __ ] >> Um >> that's fair. >> The feature that this is about I think This was finished. This part took probably like a day or two on and off with doing other stuff.

And then this took like a day. And then I ship the first five PRs, maybe like one a day in addition to doing other things. And then Kyle is picking up the last three PRs and like polishing it and getting it ready to like turn on for customers. >> Oh, you're literally the manifestation of a CEO that says, "I vibe coded something. Can my engineer go implement it?" That's what you're doing. >> I implemented the first five pull requests.

I laid the I laid all the GR and then we got to PR number five and he's like I don't like how this is architected. I'm like great. You can fix it. >> But it was digestible, right? That PR where we had to like >> I'm just No, this is the thing though is like prototyping is great and like it's a great way to figure out what you want. It's a great way to see what's possible. But like I'm not asking anyone to review a 20k line PR because it's it's like I'm not I'm not gonna like have the same bar of quality.

The person who's reviewing it is not going to have the same bar of quality and they're going to they're going to resent it. So you got to make these like models are really good at making these things digestible. We don't do this for everything. I'd say 90% of what we ship is small enough that you could just do it in one pass. But if it's huge and it's prototype and it's experimental and it's just like I need to create the vision for what this is beyond just those HTML mockups, then it's super valuable. >> Yeah, I actually agree.

I find it's very like kind of what I'm hearing you say and I'm going to repeat it back to you and see if it this is what this is what I heard >> which is you basically go through you get an idea you spend a lot of time on like model auto looping to produce a pretty good spec. Then you implement the whole spec end to end and it produces basically a slot PR. >> But the slot PR allows you to decide whether or not the features are exactly what you want them to be and you keep iterating on slot PR. >> Then you basically >> this is pretty cheap. >> This this this part is very cheap in terms of human time and we have subs and it's like we use those tokens anyways. >> Um but the engineering happens here.

[clears throat] >> Yeah. Yeah, but once you have the slot PR, then you say, "Okay, this is a slot PR. Use this as the spec and say that you were going to implement this from scratch. How would you redo it?" >> And like what are the faces down into it pieces? >> Yep, that's right. It's like how well it's like how do you break this down into digestible chunks and like what are things that we can incrementally ship to users that maybe add incremental value or uh or at least like easy to digest and review. >> I want to show you something really cool actually.

Uh this is kind of something uh I know I know we were talking about a bunch of stuff >> but I actually um so we've talked about this a little bit >> which is on the idea of like um okay I'm going to screen share. We've talked about like uh our one of the things that we're running into now uh for context is like we're kind of running into this issue where we get a bunch of like tiny PRs and like the amount of tiny work is like overwhelmingly large.

Uh and each of those tiny work items will improve the system. It just takes a while to go do all of them. So like we've been uh we've been trying to make our software factor a little bit better. And I wonder if you've seen this kind of thing before and if you guys have considered this where basically like users report feedback. >> Yep. And every single feedback that comes in, we basically like the idea is you can't trust any feedback ever because all feedback is basically untrustworthy because users are reporting stuff who knows how.

So we like immediately check if the feedback is already fixed on the latest version. >> And if it is, we can just notify that user really quickly. If it's not, we basically try and search if it's like an existing repro or new one. >> And these are like basically like we put as like percentages of like how correct is every single subsystem meant to be here. like this this will likely be wrong at some cadence because like you know like you might like maybe this feedback report in like you think it's fixed but it's a flaky issue >> dduping is going to be buggy like I've never seen a good dduping agent for this and the key ones are bad agents are >> bad if it ddupes half of the things then you've saved yourself a lot of time >> exactly so then you do this then you create like a really really good repro and I think we can spend a lot of engineering to make this repro really really good uh just by like producing like n others then you actually create an issue So the idea is like an issue isn't created from feedback like issues are created post feedback and then you just constantly check basically on every single PR every single one is run and we just check if every issue is fixed always. >> It's very cheap.

It's like CPU time uh to measure it. We don't have like web browsers or anything to go check. Uh for some stuff we do and that isn't >> and you just like keep basically like prioritizing the work. You assign a human to it. You basically communicate with the system and you either like put a plan together or you like just fix it and then you just like do the whole merging system. But like we tried to I'm have you I mean you guys have built tried building something like this before, right? >> Where was the bottom? >> This is I mean like >> you have at the very top an important piece of text which is for issues with clear code repros. >> Uh where do I where do I have that?

Yeah. Yeah. I mean >> at the very top above feedback reported like if it's easy to reproduce then this whole thing works. If it's hard to reproduce, then uh then you need something else >> or you need you need to send an agent to see if it can create one. >> Yeah. Yeah. And that's kind of what the whole like creator repro is cuz like I don't trust any feedback files by users. So we kind of just need to spend a burn our own tokens to try and produce a clear repro and you can't produce for clear repro then you assign a human to it.

So like this is the thing we've been working on internally and like it's like the next version of agent tries baml but for like human file feedback and I was like >> I was going to say this this feels shaped very similar to the infrastructure you wrote for agent tries BAML and now it's just like cool let's add another like input stream which is human tries BAML >> or humans agent tries BAML. >> Yeah, it's that one more so than [laughter] anything else.

But it's like it's fascinating because like this PR review part is actually the thing that we're really trying to bottleneck and like I actually learned this concept from you when we were talking about this in the team. What we were talking about a lot is like how do we push the human time as low as possible. So you'll notice like a large part of this is like assign a human to this like a shepherd is like down here and like really this is assignment but then we actually ping the shepherd when this happened in Slack.

We're like hey here's a bug fix. Here's the repro. Here's the bug fix. go read the PR or this is really hard. Here's a plan. Can you work can you either like download the plan into your local agent or like tell me this plan is good and I'll go implement it. >> What's the um what's the uh accuracy rate of your classify issue difficulty? Cuz that seems like an is that a human task? >> No, we're going to do an agent task and I think we'll actually what we're going to do with this is actually a twofold idea and like we're still playing around with this.

So, we're building like evals for this actually internally. >> Yeah. >> Um, and now we have enough issues as a database that we actually like have eval which is kind of nice. Like our first PR just went out and I think where they go like what's cool about this approach that we're doing. >> By the way, I only see your Excal tab. I don't know if you're trying to show us something else. >> Uh, let me pull this over here.

What's cool about this PR for example that we're making now is all these PRs. This is the one rule that we have which is like you have to push repros onto here. >> Yep. >> So you actually have to give us metrics for every single version of this pipeline. So then we can decide whether or not it's good or not. And like this is the dduping PR for example. Like obviously the dduping PR we know is low. It's lower than we expected.

So like we want to do it even better. >> Yep. >> Right. But like >> Yeah. So we don't actually know what the XT rate here is yet. So, we're still going to be playing with this um to figure out what the accuracy rate is, but I'm hoping we can get to at least like 60% correct and that is enough. And we can measure it always by lines of code. If the fix is more than 60, if the fix ends up being mclassified as a small issue and ends up being 500 lines of code, we just we can just force classify it as the other one.

So, it's not that hard. >> Yeah, that makes sense. >> Um >> I'm going to show you one. None of this was what today's ep Oh, one more. Let's go. >> So, uh, if you come, >> we got a little distracted as as you guys could see >> from the chat. If you type Vimtutor, >> it is a interactive guide that teaches you Vim. And if you do this for 10 minutes a day, uh, you're welcome. >> No, no, no. Amigos, we're we're we're in the voice era.

You don't need Vim. You just need voice. >> Voice tutor. Um, >> voice t. That's what [clears throat] I need. I need a voice that tells me how I should go talk to talk to other agents. >> Make you more clear. Uh yeah, me and Kyle were joking about like giving the agents a/ bro skill that they can give back to humans of like that was super unclear. I need you to like uh speak more like articulate yourself more clearly. >> Honestly, that would probably fix context problems more than anything else would.

It's probably a good idea. >> Yep. Cool. Okay. I want to talk about today our topic which is software factory design patterns. So we've talked a lot about the software factory in various different versions and like cool you start with the normal software factory where a human builds the thing and then you swap it out with agent builds the thing and you have all these pieces that like make up your your your software factory and then you make the agent test the thing and then like this part is now fast but the review part is still slow and like we've talked about this a bunch of times.

Um, I want to come and talk actually a thing that we haven't really done is like zoom in on this part of the factory. Um, which is like agent builds the thing. And the way that I think about this is basically like there's a couple different layers. We've talked to a lot of different teams and we'll talk about kind of like there's there's a couple different options, right? You could like build everything inhouse, uh, which I think is probably what your team has done.

Um and then there's like the other side of the spectrum which is like buy everything. So build everything in house might make like what do you guys use for your like um sandbox execution like where where do your cloud sessions run >> for like agent tries? >> Uh we we buy MacBooks and Mac minis. >> Okay. So your compute is like agent sessions run on MacBooks >> and like basically when an agent session runs on demand like it I imagine it sets up some kind of developer environment. >> Uh yeah and stuff like this. >> Yeah.

But it's all kind of already pre-installed on the MacBook. So like part of this is like speed is why we have this >> cuz like this core stuff is preset up and there's no boot up time in that case. >> Yep. >> Right. Um, and that's kind of nice. And then the other reason that we use MacBooks is because like we just had some extra MacBooks lying around. >> Yep. >> And then for the harness, you use cloud code primarily. >> Cloud code uh cloud code and codeex. >> I don't know.

I don't know what the latest one is, but it's agnost. It's like we it's just like CLI you swap out. >> Yeah. So you use the CLI with some sort of some sort of headless mode that like streams out JSON or something, right? >> Yeah. Exactly. Uh we use headless mode that stream.json via the file. It dumps all the chat logs like the chat transcript, you know, like the JSONL file. So we just read that >> and then Yeah. So then you just tail tail the JSONL file or you like you scoop it up once the session is done. >> Yeah, exactly. >> Um cool.

And then like do you have customizations that you add on there? I call it like the outer harness, right? So you have like skills you're injecting. Do you inject any MCPS or anything like that? >> We actually have like actual while loops that we have that we've built around this. So like for example to drive things to completion. >> Um >> Okay, cool. >> So like like for like we have a pull request bot like code rabbit that we use. >> Uh we want to merge after code rabbit.

That's a while loop. That's like a hard-coded Y loop that actually says this. And like if the Y loop hits like max iterations like three, >> then what we do is like but then we boot it out and we ask a human to take a look. Cool. >> That makes sense. Um >> uh >> we don't merge automatically. Uh we do have a human hit merge. >> Yeah. Yeah. Okay. But it's like fix code rabbit until mergeable. >> Yeah. Like basically do not notify a human until code rabbit is happy. >> Uh and then we have >> or until three iterations get hit. >> Yeah.

And then uh we have another thing that we're actually about to add now in this new loop that I built which is like our CI/CD takes a long time. So we want a human to hit that they're happy and after a human hits that they're happy this agent is supposed to just constantly merge against our main branch and keep it up to date so it can just mer drive it to merge. >> I see. So you're merging from main into the other thing and like basically looping this part. >> Yeah.

So after the human is happy it's basically like I I want this to merge. just babysit this thing until it merges in cuz we have a merge queue and like sometimes conflicts happen etc. >> So this is the last part that we're adding now and we think this is just going to help. It's basically just a better a it's an agentic merge queue. That's how you can think about it. This >> Yep. Okay. And where does this run? >> Um and actually all this is uh going to be it's the same machine.

It's the same thing, right? >> Um so does this run like so this is this is separate from from agent tries BAML, right? This is separate from the factory stuff. This is like running on someone's workstation. >> No, this actually we don't want it to run on your workstation. We want it to run on a non on a on a work on a like a sandbox type machine. >> Okay. So, do you run it in GitHub actions or does it run on the same MacBook where the agent >> runs?

Everything is running on the same MacBook. >> Yeah. >> Okay. So, you just have a pool of MacBooks that run any generic agent task. >> Yeah. the agent loop for example. I bet that one could run as like a pure web layer because like GitHub is automerge. So we could just do like automerge and then go do it and then but right now like that's kind of how we're thinking about it. >> Okay. Um and then like my last question is like okay so you have this MacBook and then you have you know your you know Rust tool chain etc set up. >> Yeah.

Uh, and then you have your CC/codeex like your like inner harness. You have your like >> outer harness. And then you basically have like I guess like how does how does this system know that there's a new thing to work on? Is it pulling some queue or like >> there's some there's some way to turn this it runs a web server on it locally >> and then we use some like herder or something. So, we're running like a web server. >> Okay.

So, there's a there's a web server that you can on each MacBook. >> Yeah. Yeah. And then there's like another way that we I forgot the app we use. We use some app that basically makes that app that makes that thing like >> uh accessible >> like a tail scale or enrock or something. >> It's not either of those, but it's something else. I don't remember it off the top of my head. I would um but we use something that makes it basically like behave like a remote server that's like deployed somewhere.

You have some sort of like privately accessible like thing you can open in a browser. >> Yeah. Well, it's it's a REST API. So like there's no we don't build a UI for it. I mean we could but like we could sort of react off of it but like yeah it's a rest API. >> Okay. So what sends requests to this? >> Uh now we have yeah well now we have another web server that we actually host that's a pure rest endpoint. So now there's a web app that we build.

Um yeah, and this web app is the thing that like listens to linear, listens to Slack, listens to kind of everything and that communicates with these. >> Okay, cool. So this is actually like on on the host. >> Yeah. Yeah, exactly. >> And then you have basically a cloud web server that receives like web hooks from Slack and Linear. >> Yeah. And GitHub and all the other stuff. >> Yep. Okay. >> Yeah. And it also has its like own endpoint that you can communicate with if you want. >> Yeah.

So you have this thing as kind of like what I would call your like dispatcher. >> Yeah. Yeah. It's also the thing that has a database and everything on there. >> Yeah. >> Um so like >> so yeah this is the this is basically what we call like the orchestration layer. Um, and basically the the the kind of thing I was like want to get across is like you can buy all of this, right? If you use like cursor cloud agents or cognition or whatever, you get you get all of that, right? >> Mhm. >> Uh, >> yeah.

Every single one of these is viable. 100% viable. >> And so like basically what I what I get into is like what is the what I'm trying to like elaborate here is like what is the um what is the what are the building blocks? What are like the core interfaces inside the software factory so that if you decide there's some part of it that's hard and you want to build certain things and buy other things like what does that all look like?

And I think I've sent you this picture before but I'm going to I'm going to draw it out here for the for the fans. >> Yeah. Well, people >> here I'll let you draw it out and then I'll just chat. >> Yeah. People are asking them questions. I'll just answer them along the way. What steps you use to run evals? Honestly, the eval system is pretty simple and you're seeing that we're making more rigorous in the PR that I showed earlier where now we actually have like benchmarks that we're trying to set of like we download past issues.

We see if they reproduce and what the rate is. So like now we actually just build tests around them and we have like benchmarks for like what percentage of tests pass and we have like a pass rate and stuff that we're building out. Um all the code is actually open source. You'll be able to go see it if you want. Uh do you use containers with G Visor? Uh we do not use containers. I mean these are all trusted systems. So the reason for example that you saw earlier in my in my Excala draw that I was showing is like raw feedback coming from a user is completely untrustworthy.

We run our own prompts to create repros and those repros are things that we trust. So we never execute raw code off the prompt up there. We don't want to. So this is why it's like slightly more trustable. So we can just run it directly on the machine. Uh, I guess someone could try and Bitcoin mine it, but I think Claude Code and stuff are like smart enough to like not try and do that and we don't really accept remote code. >> Uh, and then do we have a sandbox or Dro?

No, we don't use any sandboxes or anything else. We just run it directly on the MacBook because again, like I said, it's trustworthy. So, we don't have any issues on this. >> Yep. And so, like again, you can buy or build that. You could throw it in Daytona. You could throw it in your Kubernetes cluster. You could throw it ETP, Freestyle, whatever it is. Uh, but you need a place for the aven isolates, right? Like some people have done this thing with jetpash where like the agent runs it doesn't even have a real Linux file system.

It's literally just tools that look like a file system. >> Yeah. And then Simon asked the interesting question like how do you treat issues that are dependent on external sources, data providers and such. The thing is like we don't um lucky for us we don't have that much of a problem because we're building a programming language. So a lot of stuff is like either self-reroducable or not and it's very clear when it's not and if it's not we'll just bump it up to a human. >> Like external data source will just go to a human and a human can make some educated decision on that kind of problem. >> Yep. >> I think I think the folly that a lot of people make is they try and make a system that is fully automatic.

That is much harder than getting a system that's 95% automatic. So like that's how we're we're aiming for that latter bit. We're not trying to be fully automatic. We sadly, not sadly, I guess luckily will actually keep employing humans for a while because I like people and I like making friends >> if for no other reason. >> Um, yeah. So, anyways, yeah, the the the comput stuff is basically like you can build it, you can manage your own compute, whether it's EC2 or your own like Kubernetes cluster or literally you have a stack of MacBooks in the office like you can buy it and run it in your infrastructure.

So I know like people like Daytona will like basically give you like kind of a black box and you get an API and just be like hey I need a sandbox and they'll just spit it up for you but it runs in your AWS cloud or in your data center or on your like rack and office >> or you can buy where it's just like I don't want to run any infrastructure. I have like here's my stuff and here's like the vendor stuff and I just send a request and all the execution happens like in their thing. >> Yeah.

Like basically it's like you can do as much BOC as you want. >> Yep. >> Uh and this is like a very well-established domain. And >> I'm gonna ask you a question. >> How do you decide how much BYOC you want? >> So this is this is kind of what I want to get at is like if you understand these like four layers or five layers of the stack, you can make the decisions based on your needs. Um, and I think the most interesting part of this whole thing that I think is the most controversial is the idea of the dev environment.

Um, because the dev environment has many layers, right? It's like, okay, like what like language run times do I need to actually be able to uh like compile and test my software? There's like how do I like test it from the outside? So for BAML is very verifiable, but if I need to run a web app, I need a way not just to like send work to this workstation, but I need to be able to go like look at what the agent is building and see a preview of it, right? >> You have like a lot of a lot of larger teams have like internal services where like you might run like four layers of your stack when you're developing, but actually the app has like 50, you know, 50 plus services that it needs to run and these are all shared. >> Yep.

Uh, you know what's >> like shared dev? We call it like a dev cloud or something like this. >> You know what's um Okay, I'll I'll let you keep going. Actually, I want to keep I want to see where this goes. >> Yeah. Yeah. Yeah. Yeah. Uh just in the interest of time. Um, basically my take is like once you're going into the vendor's cloud, if if it's running here and you need to test it, like telling telling Daytona or whatever, hey, I need to install like I need to use this base Docker image cuz it has Rust pre-installed at a specific version that we like is easy, but when you get to the point where like I need to like poke holes in my in my uh in my like cloud to allow this running thing in someone else's compute cloud to access uh the shared things that my stack pack needs cuz we have a 100 repos and I don't want to run all of them in the sandbox.

Uh this becomes like quite a bit of friction and the the need to basically like debug oh the rust thing is not working on this thing because they're using like a fake Linux kernel that doesn't support this sys call or something. Uh basically like uh it it can cause a lot of friction and my one of my thesis is like you are going to want to own this uh unless you're building like tiny little toy nex.js JS apps and you're like, you know, in the like majority like in distribution set of stuff, you're going to start to hit friction unless you can own this in a really clean way.

Does that make sense? >> I agree. >> Cool. Um, the next layer is the harness. So, this is like basically and I I think we probably split this into like in yours we kind of split it into like inner hardness versus outer harness. you have like the core cloud code or codeex or um this could be like or you can like or amp or devon or factory. You can like buy this from a vendor who builds a harness or you could like build your own on top of like open code PI custom things.

And basically like you can you can you can kind of say like okay cool there's the inner harness is really thin and we build a bunch of like outer harness stuff or you can basically buy an inner harness that comes with a ton of stuff like a browser and testing and stuff. Uh oh my god I can't type. And then you can have your outer outer harness be like very thin where it's like just your skills and stuff. Does that make sense? >> Yep. >> Yeah.

You have options here how you do [clears throat] compaction, how you do testing. Yeah. >> Uh what's fascinating here is like I don't know if anyone in the audience has ever worked at a like Google or Facebook, but if you have then you might see a lot of this stack is literally something that you have probably worked on before. And Google, for example, they buy you a really nice MacBook, but you never code on your MacBook.

You always get what's called a cloud top. And every single time you get a new environment, you actually set up a whole new cloudtop and you just set that up repeatedly. And Facebook's the same way. And like that's just how your default coding. And like same thing here, like when >> no one ever has no one ever has code checked out on their workstation because like they basically don't want you to take the code outside the building.

Well, that's one part of it, but also it's just like too big. Like you can't have that. And then the other thing that's really fascinating to me is like >> whenever you make like web services and you go do them, you get default URLs that are attached to you. So like my uh uh my I don't I don't even remember what my LDAP was. I'm pretty sure it's like what my user ID was. Like I got a subdomain that's like myeldap.c.google or some internal Google domain.com.

I could send that around to people. So if I ever build like an internal UI, by default it's sharable in my whole company if I want it to be. And you get >> you can just share that URL and a PM could go look at the web app or whatever it was. >> Exact. And what that means is like any service I want to point to, I can have my own version of the service or I can have a hosted version of the service that I point to by default and it's always kind of just works >> um in that way.

And what's funny to me is it feels like we're all trying to reinvent that kind of golden stack where like you go in, you press a button, you get a cloud top issue to you, you can start coding on it, it does the work, you can access it from anywhere. You can see all the code on it from the same editor that you're the same editor can just switch to like different machines really really fast. >> And then um on top of that you get all these like nice little web systems where like it just points the right database, right URLs and like sharing and everything is provisioned.

You can't share URLs publicly unless you explicitly escalate them to a public URL that's sharable. So, it's internet only. >> Anyway, I just find it so interesting that like we're doing all this stuff and like there are a couple companies that have done this like perfectly already by necessity of their work. But it's also it's like there's something to be said for like you might not be Google and you might be able to get away with basically like a full stack cloud agent that does all of this kind of like for you because you don't need that flexibility. >> Yeah.

Yeah. Yeah. Exactly. >> Um cool. The last piece here is I think the most interesting and underserved piece. I'll take the pitch part of this out but is basically the control plan, the orchestration layer. Uh, and so like basically the idea here is like you want to be able to dispatch new work. You want to be able to see what's you want to be able to look at the session traces. You want to be able to look at the plans and architecture docs that are being made.

You want to be able to like schedule things to run every night or on a cron or in response to web hooks. You want to be able to like review and iterate on the code itself. uh sort of like a PR shaped thing but like doesn't necessarily need to happen in GitHub. There's like permissions and audit and who can talk to what and who can see which services and like managing spend and budgeting. And then there's like this idea of compounding engineering.

I we'll start high level. We'll dig into like whichever one of these you think is the most interesting. Compounding engineering is we talked about this on a an episode like a month ago of like how do we how do we build a memory system where it's like hey if all your engineers across your whole team are yelling at Claude all day and they're all saying the same thing or codecs or whatever it is how do you incorporate that into your uh your outer harness basically. >> Yeah.

Um this is I I I mean if I were to go and write the layers of stack I could not come up with a better representation of the stack. Um >> there is there is one argument to say that like the dev environment is actually part of the outer harness and it doesn't it doesn't sit underneath the harness. It's actually part of the harness. But uh >> I disagree. >> We'll we'll we'll avoid that that semantic right now. >> I I I disagree.

I think it's I think it's better to have the separate layer because it's like when I get a company laptop for example, there's a bunch of crap I have to install. I just want that installed. >> Yeah. There. And like to me like what else goes in the dev environment is like identity provisioning like that all goes into the dev environment not the harness and that's why I think um >> you mean like authenticating to internal services and integrations and stuff. >> Yeah.

Just like identity attached to that environment. There's some sort and that implicitly gives you internal services. It gives you access to certain API keys, gives you access to certain scopes like all the stuff you might expect. >> Um but that belongs in the dev environment layer. It's not in the infrastructure layer. it's not in the harness layer. Um because like I want my harness to not have to think about certain provisioning. >> Yep. >> Um and >> one other thing I think is interesting on the dev environment side, then I want to flip it to the audience and hear like which parts of this stack are the most interesting to drill into.

Um are you familiar with the idea of like pets versus cattle? >> Okay. While while you describe a new uh concept to me, um I will read the comments. Describe the concept to me. >> So yeah. So >> I have no idea what this is. basically this idea of like you kind of like have some number of server you like prod web01 you have prod web 02 you have like a set of servers that have like roles in your infrastructure >> okay >> uh and like if one of them goes down it's like cool I got to go rebuild prod web 03 cuz the disc filled up or whatever it is >> okay what's cattle >> and cattle is basically this idea of like you don't care if it goes away you've designed your >> oh yeah it's like web previews for example web previews on PRS or cattle. >> Exactly.

So it's basically just like web and then you have some like you know >> random UID or something. Yeah. Yeah. >> Yeah. And then you just kind of have infinite of these and these are usually cattle is more like okay we have a script that can provision these. No one has to go SSH into the server and set it up. >> Uh and like basically just on demand. So like I would say your MacBook setup today is actually more of a pet setup.

And I actually think this is a great place to start. I think you can get very far with like basically like if if you bought a new MacBook and you wanted to add it to the stack, you would go someone would go like pull it up and manually set it up. You don't have a button you can push that makes it ready to receive like workloads, right? >> Well, there's a repo that you download, you start running uh some start some script and it just provisions it automatically. >> Yeah.

Um >> but you do have to download the repo and go do that. We made it. It's not like someone has to go basically chef or >> you need the hardware. >> You need the hardware to go run it. Uh which means and we own the hardware. >> Yeah. So when I worked it, we basically like we built a system to turn everybody's workstation in like undergrad in college. We we worked it for the library and we basically built a system to turn all of the workstation from pets into cattle >> cattle.

And we had this like literal station in the basement where like whenever someone's machine got corrupted or got a virus or whatever it was, instead of going in and trying to fix what was going on, we would literally pop two drives in the in the bay, clone from one drive to the other, bring them a new drive, put it in their machine, and be like, "Cool, go download your [ __ ] from the from the file share." >> Yeah. Yeah.

No, I I I agree that dev environment should be more cattle than pro uh than bats. 100%. I actually like when we build our dev environments and like we have a similar stack to you. I run some stuff in EC2. I run some stuff in XC.dev. >> We actually treat it like pets for now because we don't have [ __ ] thousands of work workflows starting up and down. I just leave one server running all day >> and when I need to run stuff on it, I run stuff on it. >> When I need to authenticate to GitHub again, I just log in and reauth to GitHub.

It's >> another great analogy for everyone else is like I would say like AWS Lambda functions or cattle. like an EC2 container you own is more pet than cattle. >> It's it's kind of in the middle. Uh but it's it's kind of in that same kind of workflow. >> Um we've got a couple questions uh in here that I want to go over. >> Uh it's like um it's a weird network cluster. My factory ended up as a tail scale sprawl. Tunnels to extend.

Uh tunnels are Yeah, I mean yeah, we don't use tail scale for that reason. Like it's kind of annoying. Um Sedarth has some things. It's like uh I sort of understand how software factories can behave in a single repro. How do you deal with multiple repros when you have like multiple things to be completely honest sadly my recommendation is uh you're going to be sad because you're not using a monor repo. So the best thing you can do and I know Dexter has helped teams that have tons of like micro repos teams with 80 repos and or 100 repos or whatever it is.

So like my recommendation would be like simulate a monor repo as much as you can and you can do that in many ways. You can get like a checkout of you can go to a folder where you sim link all the monor repos in there. You sim link it and treat it like a monor repo. You can have an actual monor repo that actually like just contains a list of sim links to every single repo that you have and it just like >> with sim links >> like cloud can traverse your workspace pretty easily.

So if I come here and look at the code here we literally this is our template for this. We don't use this anymore cuz this is built into human layer now. But uh yeah, we basically say like if you're using especially if you're doing research and you want to go find the stuff, it's like cool, you're in a coordination repo for multiple repos. They're all like one level up basically. And so the way the way that you lay things out on disk is basically like source repo one, repo 2 coordination.

They don't need to be in the same git tree. You don't need subm modules. You don't need sub trees. You don't need sim links. As long as the model knows all of the things that might be in scope and you give it a like part of your workflow that is like go do research and understand what matters to this uh this is like it's really not like a product thing. It's literally just like a cool just tell your model where the other [ __ ] is and like >> you still uh while you say that actually you still don't need one like control file that tells you the organization structure and the layout system that you're going to have. >> Yeah.

Right. So this is like but like it doesn't have to be sim linked into every repo. You don't need every repo to have like a doubly linked like two-way reference to all the other ones. We just run all of our session you just run all of your sessions from one coordinate and this can be a separate repo with just this cloud MD in it >> or it can be or it can be one you know the repo that everyone uses the most and then sometimes you touch other repos well you put all the config there and even if you're not touching that main repo you just you have one canonical place where you start agent work from and then you give it the ability to find those other repos.

The sandbox thing gets a little trickier if you're doing on demand sandboxes because like do you really want to clone 200 repos onto the box every single time? Well, that's going to make your thing slower and you need more disk and things like this. The alternative is like make the human pick which repos are relevant, which works to a point, but as soon as I want to like take advantage of the fact that agents can help me be productive in code bases that I'm not an expert in.

Now we're back to like, oh, if I missed a thing that was important, then I'm back to like, oh, I need to go talk to the person or know, I need to have that intuition about the codebase to know that. And like we really want to >> enable you to leverage agents to do this. So, this is another time where it's like >> you could do cattle, but then you're bulk cloning all the repos all the time and have to keep them all up to date.

Or you could do pets and just have them all checked out. >> Yeah. Um, what I would the other thing I would recommend is please don't use subm modules. >> Uh, when you use subm modules, subm modules are horrible because they pin the specific commits in every repro. You're going to be a sad sad cookie. Unless you're >> sub trees are good though. >> I know a lot of people who are who are skilled with this stuff that do a lot of work with get sub trees and I've actually played with them and like I think I think it's an option that makes >> I haven't actually used sub trees.

So, I'd be I I'll have to go check it out and see how they're different. How are they different? Do you know off the top of your head? >> I I did the tutorial like a month or two ago. I I I can't I I can't do it justice today, but it's it's somewhere it's a it's kind of like a sub a work tree and a subm module had a baby. >> I see. Okay. Yeah. Cuz the biggest problem I have with subm modules, it's just like you you have to like commit in two places every time you update the every time you update the module, the subm module and their actual repo.

And it's just like not fun. Yeah, sub trees have a little bit of that. But I think I think also I've heard from people who do a lot of this and like who especially I think Mos Mike Hostetler was on an episode like six months ago where he was showing his like Ralph loop with git sub trees set up to help people basically do agent work across all their reps. >> You want to pull up your thing really fast? I have a couple other interesting you scaladraw that you had.

I had a couple of things I was going to do. >> People have got some other interesting things of like uh where do you see the software factory design pattern evolving in the near in the near term. Honestly, I think the biggest thing that's going to happen is more people like actually what's funny is I should actually just paste our our workflow into here because it's basically the same. It's like a more micro version of a part of this stack. >> Yeah. >> So, I'll put that in here really fast.

But like when I think about this, when I think about the evolution, I'm curious about your thoughts on this evolution. >> Yep. >> Uh let me Sorry, I'm coming to the Excala draw and I'll post it down here. There you go. It's at the bottom part. >> Nice. Um when I when I go think about this um uh when I the way I see this evolving is one it's going to become a lot more frequent. Like every company is likely going to build some sort of internal agent that does something like this.

And like I personally don't think the whole thing is viable. Uh, and the reason I don't think the whole stack is viable is because every company has some weird quirk about their um internal stuff. That is how their setup works and everything. So, it's like you just can't know everything. Like for us, like someone asked earlier, how do you deal with issues with external resources? Well, like >> we're lucky we don't have that problem or it's so rare for us.

We don't have to design around that system. On the other hand, >> this is >> this is why I'm so obsessed with like defining these building blocks. Yeah. And like defining really good interfaces between them. So like the interface here is is an interesting one. Like ACP tries to do this uh but ACP is kind of >> on hit X on Riverside again please. >> I can't see it full screen and I want to see Dexter's diagramming and max potential. >> Yeah.

So like between the control plane and the harness you have something called ACP as an option. It's quite narrow. Um I think we need like a thicker version like a wider version of this interface >> for a like coding harness to talk to a UI or an editor or a web app. Basically it's like a protocol for things to talk along. You also have things like AGUI which again is great for broadcasting the UI and the events but actually like the >> the challenge I have is neither of these support hooks which means like something that is like job it is to life cycle a harness and like its events and do special things in response to like I don't know in human layer we have a bunch of hooks of like >> yeah go ahead >> do you know why though I think it's because every single harness is literally basically its own bespoke UI. >> Yep. and like converting every like claude code and codeex do not have the same hooks and pi has totally different hooks. >> Open code has a plug-in system that's like the hooks are like a completely separate concept. >> Yeah.

And I think this is kind of where it like all these things die which is that it's not a >> there's no standard for this there and and rightfully so like every harness has different trade-offs. Like clawed code is much much more like you bring it and it's good like it's guaranteed to be good. >> Yep. >> Whereas like uh pi is more like you have to configure every little thing and that allows you to have >> like it comes with control but you have to you have to build more stuff or but also you have to build more stuff. >> Yeah.

Exactly. Right. And I think that's why like you can't really do this cuz the abstraction layer in which you represent everything there is just not that good because it and and it by design can't be that good at least in my eyes uh not right now but also like you know I used to think about web and then I finally learned about react I was like oh [ __ ] they actually did solve this hard problem by inventing this concept of state turns out state is all you need and everything else kind of just devolves from that in a really nice way and like whether you like react or angular doesn't really matter but like web web systems generally are state driven engines and that was a novel concept. >> So like you just got to think about what's the boundary here to think about and I don't know what it is and I think that's why it feels like you guys I know do a lot of work of wrapping every single one of these harnesses and making it good to plug into. >> Yep.

Yes sir. So um I think that's most of what I had. I mean are there more are there more comments in the chat? I mean like yeah the the interface between the harness and the dev environment is something like the Linux kernel and like I don't know the dev environment the infrastructure >> infrastructure is mostly like how do you go set it up it's like what APIs do you have what control do you have how do you go build this >> there is something called compute SDK that tries to be this which is basically like one library for oh it's become a they made decided to be a benchmark instead of a instead of an SDK but it used to be >> like it doesn't really matter for that layer.

Like it's the same reason that you don't like yes it's nice to be able to abstract over like GCP and AWS but thing is once you're a big enough company you're basically either a GCP shop or AWS up no one really needs >> to swap stuff over that often and when you do use >> the AWS interface is already pretty good. >> Yeah, that's kind of what I mean. Like if you're using Daytona then Daytona interface is good. If you're using like GCP, the GCP SDK is good and the GCP CLIs. >> So like all these things ship CLI.

Go ahead. >> We used to have this rule of like at Sprout, we said don't wrap interfaces with other interfaces or don't wrap APIs with other APIs. It was like okay cool. like we have the official Twitter SDK and then I'm going to make our version of the Twitter client which just wraps some of the methods that we need and all the data is just passed through with a little bit of like transformation and it's way better to just like >> just call the official SDK at every call site.

You don't need to like put an abstraction in front of it if it all it does is like actually create >> like okay now I have to like do extra work every time I want to add a new call to a new endpoint on Twitter. That's kind anyway that was my spiel about where I think the interface stuff was going. In terms of buy versus build, where do you land? Obviously, I know you're biased, but that's okay. >> I mean, I think the idea is like I think when you use a f fully verticalized cloud agent, some of the challenges you will hit is there's friction in defining your dev environment.

There's friction in having limited like choices on compute. I think what Devon and Cognition are doing is really interesting in terms of like letting you run what they call like outposts or workers where it's like okay you can run the compute and the outpost will bring the harness and the tooling into your level of compute and you and the harness can kind of co-own the dev environment. Uh, and so like things like that I think we're going to see more and more of and like the the conversations we're having with a lot of people and like trying to figure out is like what are the right interfaces and I think what a lot of people are starting to realize one more time or I can sure >> most people are either like buy the whole thing or build the whole thing and I think the answer is like you should be able to instead of having like it's like it's like composition over inheritance right it's like rather than having to be like okay whatever layer I buy at I have to buy everything below it.

And so it's like, oh, if I want orchestration and I want to be able to ping agents from Slack, I either have to build the whole freaking thing or I have to buy the whole freaking thing. And like the the real answer is like you should have as an engineer choice and the ability to like work in open systems and plug these things together. >> Uh so like what's interesting for us is like uh our a you're Devon is kind of at these three layers, right? >> Yeah. >> And they do the infrastructure too.

No, they do the cloud sandboxes. Yeah, you can move that. >> Yes, but I think they kind of own they don't want to own all of it. It's like optional for them. >> Uh I I think in most cases they they want to own it. Uh just because it's a better experience. Like it's easier if I can just sign up and like send a Slack message and see work start getting done versus like the price. >> Huh. >> For enterprise customers, I think they just run the whole thing on prem.

Um but yeah, enterprise customers. Yeah. >> So enterprise wants to bring the compute and probably bring parts of the dev environment. Yep. >> Yeah. Like for us, what's interesting is like when we build our agent drives BAML stack, what we're really building is this this kind of we're building like up here and we're kind of like we don't we explicitly do not want to own this layer. We don't want to own >> because you want to use the harness that everyone else is using >> and like really well I want to make it swappable cuz yeah cuz the harness is going to change like all the time.

And really I don't really want to own the compute sandbox. I just do this because of like convenience. Like whether we run on a box or not, I don't really care. >> But these are the parts that I I find valuable. It sounds like >> the layer that you're trying to help is like, hey, you for human layer >> uh human layer, >> you just wanted to buy the orchestration, >> but you wanted to bring your own harness. >> Yeah. >> This is kind of what you're trying to help plug into. >> Yeah. and Devon is there and like cloud cloud automations like cloud background stuff is like cool you can do like cloud code as a building block >> or you can bring the entire cloud code st Oh oh that's not good or you can like >> I got it >> basically I got I got stop I don't want that one you can have like basically like clawed uh what do they call it like automations or I think they call them routines or something but it's like basically like where they'll bring the compute environment and stuff too. >> Yeah.

And I'm going to do one more thing really fast. Um, put this up here. And like I would say like there Claude has cron jobs and stuff that you can do and they're trying to like plug into this part over here, right? >> Yeah, exactly. >> Craw where they're trying to be like >> so you can if you want to be the full Cloud stack and like uh I know Pearson Marks does this the the Jelly Pod guy. He was demoing this at the last thing. like he uses the full cloud stack and like he knows it's separable but he likes cloud code as a harness.

He thinks the dev environments there are good enough for him and the compute on demand is great and he can schedule automations on it and so he his you know composed stack happens to be all in the cloud ecosystem. >> Exactly. [clears throat] Um I'm going to go read some comments because we got a bunch more really fast. >> Some people are building both. They're like fully multiplayer. It sounds like there's a product there that's kind of cool. if you can drop in the chat if it's interesting.

Uh given that we have a choice that it stack would a lampstack equivalent emerge I think so I actually think we will get an equivalent to a lamp stack here over time. Um >> you mean nextJS next.js superbase and uh work. >> No no no no I actually I don't know about that because that's not really what this is about. This is like the harness the dev environment the infrastructure deployments like dev environment is much more closer to docker for me >> like the infrastructure >> application stack but the yeah >> the yeah like for this like >> software factory stack >> software factory stack yeah exactly uh ATB is like a thing that we call like agent tries BAML which is like our own software factory that we built to automatically improve the BAML programming language and then yes these videos are recorded if you want to check them out they're on the YouTube channel I'll drop them off in a bit >> we post them on Twitter as well. >> Yeah.

And then is there an open-source project for control plane orchestration? I think the reason that there isn't an open source project for control plane orchestration right now um is because like one like no one you don't want to use someone else's control plane because it's so easy to write code now unless you actually get all the integrations built with it. And right now everyone is everyone that's doing that to some degree wants to make a product and they have some value to serve.

So like they want to sell it to you and that's not an unreasonable request. There's a lot of good design and intuition that goes into it and like >> there is um we met the open inspect guy last last week but he's that's open source but it's another full stack. His his is like okay cool we bring I think it's like basically like it does all of this and you bring the compute to run it on basically. >> Yeah. Exactly. And like I that's what I mean like there is stuff but it's so easy to build now that I would just build like for us for example like we just decided the two layers of the stack we care about and we just plug and play and like you ask me like why would like I actually have thought like should we use human layer stuff for part of our control plane and like the only reason we don't is because like >> they don't have an API surface that I can't like plug into my slack into our GitHub issues into our BAML CLI.

So now we have to build this whole layer around this and I think the input control plane Oh yeah, but the point is like we still have to build an input control plane to define what we consider an issue. Like maybe we could use human layer for like the good once an issue is good, but for that untrustworthy feedback that we get from a another user's agent. >> Where do we run that? Like that's a and this is the same reason that we don't just use linear pure linear for this anymore because like we have two different versions of issues and when you ask like is there an open source project the problem is every company's working style has different definitions.

Yeah, this is David Cross's thing of like software should be extensible now that agents can write code. Like why are you going to make me plumb some like wissy wig like integration manager of like cool when a linear web hook comes in transform it in this way and then do a workflow. It's like just let me write the code for that. >> Yeah, exactly. And Matas is asking like do I see an open source I do personally see an open source solution happening and like I suspect companies like you will probably eventually support something like this because like it's in everyone's best interest to go do that and if you have a good prod stack you would build a good open lure solution but I think it's really hard to upfront build an open source solution like part of what I really admire about the work that Dra has done is that like they're working with real companies and real companies have real products and instead of optimizing for one company what they can do is they can take the learnings across like 10, 20, 30, 40, 50, 100, a thousand companies and build the best interface that really serves across all of these.

And then you can imagine like an open source thing being much more u um able to being built like you can't build React the first time the web was invented. You [laughter] really have to watch a lot of stuff happen right and wrong before you see that happening. >> Well, and I think step one is the interface layer. like MCP did this really interesting thing where it created this really explosion of ecosystems that let both like agent client and agent harness builders and integration builders like mix and match with each other really easily.

And I think basically like there's a lot of things that we're working on is like less like what is the what does the product look like at any one of these layers and more like what are the open interfaces that enable people to have choice and be able to buy versus build any part of this stack for the parts they want control over without having to uh >> without having to uh to to go all in on like let's say the cloud stack or the the codec stack or the open inspect or build the whole thing themselves.

And it still is to show like how nice do Claude and Codex want to play with each other. >> Yeah. >> Like if they don't want to play nice with each other, they will keep breaking APIs and open. >> This is why Claude doesn't support agents MD cuz they're like we want Claude MD to be the Cloud thing and everyone else can use the other file, but we're special. >> Yeah. And like they've just decided that's the case and it's annoying, >> but like we'll just have to kind of see.

And yeah, Eric has this great thing. like the browser wars because the thing is like in the past the companies believe owning the browser was worth it and it turns out like it was right. If you own the browser you can make a lot of money >> by doing that. Um and owning mobile was worth a lot of money. I think a lot of people are betting that owning the harness is going to be worth a lot of money. You want to see something one last thing that's kind of interesting before uh we close it out for the day. >> You talked about a question from Eric real quick and then yes which is is human layer composable decomposible?

So, so human layer owns the uh orchestration layer and it can also we have a harness uh but it's swappable. You can use cloud code, you can use our harness, you can use codecs. Um we kind of have noped out of the like we'll eventually have managed cloud sandboxes and stuff but our kind of take is like if you're serious and like the customers that we work with all of them would prefer a good interface on top of compute and dev environment that they own basically. >> Yeah.

Exactly. Yeah. It's just way easier. like I don't want to deal with setting it up on someone else's infrastructure. >> You got one more thing. >> I I'll show you extendable software. I was going to make this a whole podcast. So if it's interesting, let us know and we can make this a whole thing. Um I've been playing around with this concept a lot. >> Um so when I think about extendable software, I think the best way to extend software is let agents write code. >> It's called code mode.

You've probably seen it. So in this case, here's like a string that an agent wrote. It does some code. It does some stuff. And I'm basically just going to start mcking with this string in really interesting ways. So I'm going to pass in some code that I'm additional code that I'm going to dynamically generate. So like if I do like name two uppercase just run this really fast and I'll run uh dexter. So when I run this I have compilation errors.

U let me just do one grap instead. H we changed our API to be only reflect. So I had to replace all these when I run this code. Oh, never mind. I got nothing good to show. I apologize. >> Uh this is broken for whatever reason. Uh but anyway, we'll do an episode on this >> about how to run like dynamic code generation, how to actually have your app be extendable from a really extend from a really interesting perspective.

We I just upgraded BAML version so I can't show this yet I guess. uh I should stop running local dev is what I've really learned. But um I think extendable code is definitely going to be one of the most interesting features out uh out ever because like like for example you in human layer how great would it be if people could define their own columns and their own like you could be like hey here's a small little code workflow that I want to run >> for doing this and like they can just go say like oh here's an agent loop that I want to add on because that's kind of what cloud automations is effectively but just done kind of shittily like a cron job a chron Crown job is basically like code that says while true at this time run this piece of code.

That's like the ultimate crown job. >> Yeah. In a scalable, you know, fault tolerant way where like if the server goes down at that time, it still runs a minute later on another server. >> Yeah. Yeah. So that's the hard part about the execution. But the extendable part is what I really want because if I can only write the workflows that you let me write, >> I'm kind of sad. But if I can write the workflows I can write on your stack and I can buy the reliability from you, then buying the reliability is totally worth paying for, >> right? >> I mean, I think it comes into like this idea.

I mean, in in our own code bases, I know many people have talked about they have like viable zones and they have no vibe zones, right? They have the the functional core and then they have an interface and things like pragmatic code can sit on top of it. And there's no reason not to basically create space for users to deploy their cut like instead of user files a bug and if it's in the no slop zone our factory builds it chips a feature goes in the next release.

It could literally be like in real time it's just like cool these parts of this have you know sandbox execution adapters where your agent writes code and that modifies your app in real time. But getting the interfaces right and figuring out okay what is the core versus what is the code and what are the interfaces that are available in that code is is uh it's an interesting space that yeah I think I think we've both been exploring a little bit uh and excited to I don't know I I don't have any like cool like demos of that yet but I think Ree from Exeutor actually built something like this where like the way you customize and configure exeutor is just by dropping TypeScript files in your directory.

Yeah. I I think there's like no other way around this is like long term what you want to do is you just want to be like here's my code and every piece of software effectively turns into like a harness. >> Yeah. >> And open code like the whole point of those things is like hey you want to customize the behavior. Cool. Have your have your agent write code and it live reloads into the agent and now you have new extensions. >> Yeah.

Exactly. Because like nothing else I think is like that's like the most interesting stuff in my in my eyes about like how people build apps. It's like the like React built like the uh interactive web and I think we're about to have like just like extendable software as like a common place like VS Code amazing because you can add extensions. Tails Tailwind comes out you can have it work in VS Code without VS Code updating >> and that's like magic. >> Um cool we're way over time. >> Oh [ __ ] we are.

It was a really fun conversation today. I love today's conversation. It's been so long since we've had one of these. Yeah, had a good time.

The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.

Use this transcript

Three free tools that work on the material around a video like this one. No signup, no login.