Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

AI Engineer · @aiDotEngineer
This video has no Most replayed graph yet: YouTube shows one only once a video has enough views. These are the moments viewers replayed most in AI Engineer's most watched videos.
Most replayed moment at 18:45
4.8x that video's typical replay level
that they're they're changing they're changing things in the database not yet. You want to run them through the ontology first and make sure that works. Okay. I only got an I've got I've got another I've just a short time. I'm going to try to show you some of the things that um that you can
Said at 18:37
Most replayed moment at 16:13
3.6x that video's typical replay level
method signatures, the program layout and the call stacks. So here's some examples. I don't think you'll be able to read this one, but this is like the level of abstraction we're at. It's how we're actually going to lay this stuff out and how these systems are going to interact. Dylan Mulroy from Cloudflare talks a
Said at 16:06
Most replayed moment at 13:49
2.3x that video's typical replay level
Fable uh and it runs into an unknown, ask it to log it, right? So that um you uh you can see where the deviations happened and then you can sort of figure out why as well, you know? It will usually give you some context about what happened.
Said at 13:43
The graph counts replays. It does not show where viewers stopped watching.
Words
10,422
Runtime
51:30
Speaking pace
202wpm
Reading time
43min
202 words per minute, above the 201 75th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Welcome to this uh fireside chat. Um we're going to I have with me uh Theik Shihipa and Cat Woo from Anthropic. We are going to be diving deep into claude code and we'll probably talk a little about this Fable thing that's been going on that that's that's been out there in the news. Actually, on the subject of Fable, literally a minute and a half ago, Fable came back. Like Fable is now available to me. So, if you all want to run out of the room and start using up your Fable credits, I I wouldn't hold that
101 words, the words spoken in the first 30 seconds at 202 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 549 |
| Average words per sentence | 19.0 |
| Longest sentence | 166 words |
| Questions asked | 77 |
| Sentences containing a number | 24 |
Most used terms
Filler phrases
730 in total: like 355 · um 110 · uh 91 · you know 75 · actually 31 · kind of 20 · right? 19 · sort of 13 · I mean 12 · basically 3 · literally 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
Welcome to this uh fireside chat. Um we're going to I have with me uh Theik Shihipa and Cat Woo from Anthropic. We are going to be diving deep into claude code and we'll probably talk a little about this Fable thing that's been going on that that's that's been out there in the news. Actually, on the subject of Fable, literally a minute and a half ago, Fable came back. Like Fable is now available to me. So, if you all want to run out of the room and start using up your Fable credits, I I wouldn't hold that against you.
But we're going to have a great conversation. So, please please stick around. Um but yeah, so um please welcome uh The Cat for me. Thanks for having us. Yeah, we timed it for the the chat for sure. Yeah. >> Yep. Yep. This is this is why it's all happening. Um this year has been somewhat absurd. Uh I it's amazing. Claude Code came out in February of last year. It's under a year and a half old and it was a bullet point on the Claude Sonnet 3.7 launch.
I'd love to hear from you. How has your how has what you do on a day-to-day basis changed in the past year now that we have these coding agents that actually work for us? >> I remember when we first came out with Cloud Code and Sonnet 37, you would give it this task and you would have to closely monitor every single little thing that it tried to do. Uh I remember I would read every permission prompt extremely carefully.
I would frequently say no. I would always say no no no like did you check this file? Did you check that file? And now it's been incredible with every model generation. I feel like we've all gotten a chance to just take a step back, delegate a lot more of the like menial implementation to claude. And it just freed up a lot of our time to think about more creative work like what is the right experience that we should be providing to our users now that we know cloud code can implement a lot of it.
And now with Fable, it's just a totally different step change improvement. um we see for a lot of our use cases that you can actually oneshot a ton of features with Fable now. So it it's been amazing to see the transition and to go through this with all of you in the community. >> Yeah, I mean I I think I remember the first text I got about cloud code. One of my best friends was like, "Oh, you need to go try Claude Code." And it was about just when Opus 4 came out and I I tried it and I was like, "Oh, shit." Like I I need to work at Anthropic now, you know?
And that was Opusport which I mean great model but like yeah you were permission prompts and uh yeah I think it's kind of crazy how much amnesia we have I think where I'm like oh like auto mode has always been here right like I I don't even remember pressing like yes and allow. Um and yeah I think for me the big thing that I'm trying to push myself is like oh we have to do like high higher quality work than we've ever done before.
You know like the the outputs are like like incredibly high quality. like I've been using it to edit videos a bunch and I'm like okay it has to meet the very exacting demands of our brand team and in a couple hours or we just can't do it you know and so um yeah I think that's how I'm sort of like trying to shift with Fable where it's like okay the best work we've ever done like faster than we've ever done it before >> I've certainly been finding that myself like software engineering is getting harder because the the level of ambition of the th stuff we can take on has gone up like I I have such higher expectations of myself now that I I have these tools to back me up which is fun but it's al it's a lot of work. >> Yeah.
It's all the thinking is you know it's it's tiring. >> And so what's a piece of conventional software engineering that was true a year ago that you don't think holds anymore in this new world? I think one of the biggest shifts that we're seeing in the ENG skill set is I think you know two years ago it was pretty typical for a product manager to go talk to a bunch of customers and over the course of six months align with like cross functional teams on some PRD and then write this like thorough spec um and uh dock on how exactly we'll implement this before the first line of code gets written and now things are like completely turned the opposite way.
I I think for a lot of engineers, the the push I would give to a lot of folks in the room is to develop more of your business sense and product sense on what is it that we should build because now that this the timeline between having this idea and building it is so much shorter. It's down from six to 12 months to maybe even a week. That means all of us need to have better taste on what is it that is worth building. what is it that will actually inflect the businesses that we're working on.
So I think it's like an increase in value on product taste and business sense and a bit lower on execution in in most in most product domains. Of course for infra uh there there's still a very heavy emphasis on making sure all the details are right. >> Yeah. I I think for me uh it's like rewrites are now good. You know what I Like I I think that like >> the worst thing you could do is now actually fine. >> Yeah. Exactly.
Exactly. Like all the like especially Yeah. mythical man stuff like never rewrite like I I'm a pro rewriting now you know like if you have a good test suite I think actually the rewrite forces you to like make sure you have a good test suite but >> um I think that like what people I think underount is like a codebase is a spec and maybe it's the only copy of the spec that you have right because like no one knows every branching part of the codebase and and yeah you can take this as like an artifact and like distill it or create other versions of it obviously like yeah we rewrote bun in rust And uh you know it works great like you know it's it's live for me right now.
Um >> but you're not shipping clawed code on bun in rust yet right or is that >> internally we have. >> Wow. Oh that's exciting. Yeah. >> Yeah. Yeah. >> I think what you're saying there about um yeah the uh I'm sorry I just lost my train of thought. But yeah, the um the rewrites thing has been really interesting for me as well because you can almost come up with a good test suite and then spin up three implementations and pick which of those implementations was the most accurate.
I'm doing a lot more prototyping now. I've always been a prototyper and now I pro I prototype things on my phone during the conference just so I've got something that I can pick up later on and that's working now which is kind of extraordinary. So the other big launch recently was claude tag which that's what a week old now I think at least for the rest of us. Claude tag I understand that's being used in anthropic by non by non-engineers a great deal.
What kind of things are non-engineers doing with claude tag? >> So claw tag is a claw that lives in your team's collaboration tools. We launched it last week within Slack. The thing that's different about claw tag is it's multiplayer by default. And so once you add cloud tag into a Slack channel, you can chime in, your teammates can chime in, and you can collaborate together on the PR. The other big difference is that it's proactive instead of reactive.
So you can tell claw tag, hey, monitor every bug report in this channel, put up a PR to fix it, and tag the engineer who last touched this part of the codebase, and it'll do it for the lifetime of the channel without you having to manually tag it in. And then the third big shift that we've seen is um we've added team memory into this. So if you tell quad tag your preferences in the channel, it'll remember this for every future post.
So if you wanted to um always debug like if you always wanted to debug outages, but you don't want it to debug warnings, um just tell it that in natural language in the channel and it'll remember it for you and everyone else on your team. Internally we see quad tag as the evolution of quad code. So we see this as a large shift in how we work internally. Um claw tag currently lands 65% of our product PRs >> for for all of anthropic and for cloud code or just for cloud code. >> Uh this is just for our product engineering team.
So our internal version of claw tag lands 65% of our product PRs right now. And this is a huge shift. This is like this is more than 50% [laughter] of our PRs. Um, and the way that we actually see people split work between cloud code and cloud tag is cloud code is still the best place for your most complex tasks when you're interactively iterating with the agent. But claw tag is great for having it work proactively on your behalf so that you no longer need to um manually uh kick off quad codes for for all of the bug reports that might come up for features that you're working on. >> Yeah.
And for like non-coding cases like I I I think we've seen people use claw tag just uh like for example before this before this talk we asked claw tag like hey when is Fable releasing? We wanted to make sure that like you know we'd line it up with the announcement. Um, and so Cloud Tag would search our Slack and and look at, you know, who who's been saying what. Uh, so as a search engine for your company is really valuable.
Uh, it has all the context for your product. So you can ask it like metrics related questions. And oftent times when you're making decisions, you want it to be informed by like, you know, what do the metrics say? And then so like you hook it up to your event store. Um, I've seen like our marketing team do things like, oh, like, hey, tell me about this feature. And, you know, they're not programmers, but Claude is a programmer.
I can clone the code base and be like, "Oh yeah, this is like, you know, the feature. This is what it looks like. This is a recording of me using the feature, you know. So like um yeah, it just enables a whole wide variety of things." And I think we're still early on in figuring that out. Yeah. >> Well, I feel like this is one of the fascinating things about the Claude Code story is you use Claude Code to build Claude Code and you've been doing this since presumably before the public launch of launch of Claude Code a year and a half ago.
And yeah, one of the problems I've had with coding agents is I get how to use them as an individual, but I'm not really clear on how I use that in a team in a team environment. It sounds like Claude Tag is is your current answer to that sort of team collaborative layer for this stuff. >> Exactly. And a large percentage of our sessions are actually multiplayer right now. So that means maybe I say, "Hey, I think we should like um implement this new feature in co-work." And I'll tag in claw tag to do a first pass at it.
And then I'll tell cloud tag, hey, just like share a recording of of your final implementation. And then I'll tag in design to take a look and they'll nudge it and then they'll pass it on to edge to like take it to the finish line and get it out to prod. And so it's been this like very fluid experience. We're still trying to iron out what the social dynamics are for steering the same session, but we found that people just like observe how others use it and then follow those social norms and it's actually been pretty easy for pretty intuitive for us to integrate claw tag into our teams.
Yeah, I I think it's great for like Yeah. teaching people and also like kind of kind of reducing slo because because you know like if you see someone just be like hey at Claude fix this or something you're like uh you know like I I think there's like some societal or not societal like just like uh like the fact that everyone is seeing you use cloud together sort of levels up how you use claude as well. So >> right you you want to do work that you're proud to do in public that where the quality doesn't doesn't fall off the cliff.
So since we're talking about using claude to build CL well how the process for building claude how do you deal with the hardest problem in all of engineering it's prioritization right how do you decide which features are worth building and shipping when building a feature is so much more inexpensive now this is the hard thing [laughter] so there's a few ways we approach it one is we dog food our products every single day whenever there's something that we want to be able to do in our products that we're not able to instead of finding a different solution we fix our product uh so that it can support this case.
Uh we have a very heavy dog fooding culture internally. So before we are able to share our products with everyone in the world, we share it with everyone within anthropic and we share it with some early customers who give us very honest feedback about it. The more brutal the better and we iterate until people love it. So we we have a internal bar for the number of active users and the the amount of retention a feature has to have before we share it with the world.
And because this bar is very clear, every engineer knows what they're trying to hit. And I think this also levels up our polish because if the feature isn't polished, people will churn and then we shouldn't ship that feature. >> Do you have an example of a feature which surprised you? you you rolled it out and the engagement was off the charts and it it became was something that was unlikely to be shipped that actually turned into a a real product thing. >> I do have one.
So a lot of folks on our team love remote control. So remote control lets you connect use your mobile device or quad in the web browser to connect a local quad code session running in your CLI. I never have this need because I just kick off the task directly on mobile and it runs in a cloud session and doesn't use my local environment. I think this is because I'm doing very easy coding tasks. But this is something where I didn't totally understand it.
I was like, hey, people should just set up their remote dev environments. But in practice, once we rolled out remote control, everyone like so many people who I talk to are like, "Okay, now what I do every night is I um a lot of people tell me that they just plug their lo their laptop into a power charger, close the screen or like open a bunch of remote control sessions, lock the screen, and then use their mobile phone from their couch to to control quad code." And so this has been this flow that we're now leaning into that I didn't originally get, but now I do.
I do exactly that. I get so much so much work on my laptop done from more comfortable environments because yeah, I can remote control it now. That's really fun. Um, how does code review work? Are you pres are you reviewing does a human being review every line of production code that makes it into clawed code? And if not, what are you doing? How do you keep the quality up? >> Sure. Yeah. Um, it it varies on the task a lot.
So, so we um for important areas we have code owners, right? And so, uh the system prompt is kind of an example where we have a code owner uh you really need to like uh submit your uh you know you need to get their approval. Um and then uh >> so I guess the code anyway is directly responsible for the quality of that area of the code. >> That's right. Yeah. Yeah. >> And they need to approve the PR that touches it. >> Right. >> That's right. uh we have code review our like uh our uh code review git b you know review everything and and so um that like goes on every PR and often times like that's doing the bulk of the review.
Um I think that like uh we I something I've seen on the team is like for more complex PRs you might make like an artifact to explain the PR so that other people can then review. Um and uh yeah we just invest a lot into verification CI/CD things like that to make sure that like you know uh anytime anything fails like we have a test we have like uh a really robust environment cloud can control cloud code and test it you know what I mean so um yeah there is a uh yeah there's just like a multi-pronged approach to like code review I think you have anything to add >> in general we are trying to move to a world where humans don't need to be in the loop And so for the most critical core changes to the core of quad code and other the cores of other products, there is always a code owner and they do manually review all the changes.
But increasingly for the changes that are at the outer layers, we actually have quad code review fully review those. That sounds pretty scary but there we've had this six plus monthl long process to get here and I think there are like baby steps that you take to build up trust with code review. So in the beginning we would have human review for everything and then increasingly we would say okay for code changes that touch these files code review is catching a 100% of the issues there.
So we actually don't need a human to be manually reviewing those. Um and then also when we have incident review, we look at the PRs that cause the incident and we say okay how do we update code review to catch that and then we also take those PRs and add it to an eval set to make sure that our future changes to code review never regress that metric. So it it is a big like removing humans from the code review loop is a big step forward.
I think it can sound scary and it's not something that you can do overnight, but it is something that you can do through like many months of investment in the infrastructure to give you the confidence that code review is catching everything that you care about. >> So, it's interesting you mentioned building trust in the models because that's something I found is like I know that Opus 4.8 if I ask it to build me a JSON endpoint that runs a SQL query and outputs JSON, it's just going to get it right.
Like that that that's not something I have to review closely. But then a new model comes along and I still don't know how do I build trust in fable quickly that it's not going to mess things up that Opus didn't. Is that something that you you you have to think about much the those new model like how does the new model affect your intuition for for what it can do and what it can't do? So the main reason that we're building up this eval base over time is so that new models can be a drop-in replacement because what we do when we have a new model is we run the whole eval set and we make sure that for example Fable is strictly better than Opus 48 and that gives us the confidence to drop it in. >> And are those model evals for Anthropic as a whole or are these claud code team specific evals that you're using? >> We have both.
So we we have evals on our team and we run code review across every repo within Anthropic and so we have evals for that and for things like auto mode we not only have evals across every user within Enthropic but we've also commissioned multiple external testers to red team this to create environments with prompt injections and malicious inputs and make sure that automode doesn't let any of those pass. So for claude code itself and this is a challenge I've had with stuff I'm building.
I want to know if the system prompt improvement I made actually improved the product. Right? That's that's the sort of most basic form of product specific eval. And I still don't have a great feel for how to do that. Is something is that something that you're doing such that you have complete confidence that this tweak that you've made to the system prompt does result in in better in better output? We don't have complete confidence but we do a lot to make sure that we don't regress performance.
So the starting point that we have is we have a suite of external evals that we trust and we complement that with an even larger suite of internal evals that we trust. Uh to start we mainly optimize for capability. So given given a complete definition of a task and the full codebase does claude make the right decisions and fully fix the the bugs and pass all the tests. So that's the starting point and that's the thing that we optimize for because it is like most directly what users want.
But there's a lot of like behaviors that impact how users feel when they work with quad code. For example, people really don't like it when cloud code says it's like time to go to sleep. [laughter] or people really don't like it when it says like, "Hey, I finished two out of five parts. Like, do you want me to continue?" Like, "Yes, please continue." Um, and so we're building up a set of behavioral emails to catch these.
And as we get user feedback, please be loud with us about your user feedback. As we get user feedback, we just rank, okay, these are the priority issues. And we go down one by one and build evals for each of them. So, it's not 100% coverage, but we try to it it is a priority for us to increase the coverage. And how much overlap is there between the how much interaction is there between the Claude code team and the teams at Anthropic who are training the models in the first place?
Is is that quite a close collaboration? Now >> across Anthropic, we all work quite closely together. So um we we meet often to talk about like what do we expect the next generation of models to be able to do. I think our research team has also been amazing about just like showing this publicly. So we often talk in our blog posts about how we're targeting everinccreasing longer horizon work, how we train quad itself to be honest, harmless, and helpful.
Um, we also put a lot of effort into making sure that it's aligned with your intent, even if your intent is expressed in a fuzzy way. Of course, try your best to be specific about what you want. So Quad um has all the context, but even when you're you're not specific, uh we we teach Claude to make good assumptions and yeah, I I think it's been a it's been a productive partnership. >> And so um Derek, you this morning you mentioned that the system prompt for Claude Code has been reduced by 80%.
Because of Claude Fable, can you go into a little bit more detail about what that looks like? What kind of things have you been able to drop? Yeah. So it wasn't just Fable, it was uh Opus 4.8 as well. And um yeah, going forward the future models, but we do sort of um we have different system prompts for different models now. Um I think that like uh some of the patterns we saw is that um we were over constraining Claude, right?
So I think the initial like maybe Opus 4ish kind of models wanted a lot of examples and uh removing examples was extremely helpful because it was just more creative than like uh you know the examples we gave it. >> That's really because one of the top prompting tips I give people is give it examples like examples are the easiest ways. If that's no longer true that kind of breaks my prompting model a little bit. >> Ex yeah same same here.
I think I was surprised to hear that. I I think that now it's more about like sort of the shape of what you give the tools to Claude and like yeah your system prompt and things like that. Um the other thing we did is we um we we try and give it more context and fewer like do not do this you know because like I I think that um it's just a very strong impulse to Claude and especially if that uh conflicts with user instructions later on that can be like extremely confusing to Claude, right? because you're like, "Oh, like I've got this skill that says this and the system prompt says this." Um, and so we try and like uh have fewer hard constraints and more just like sort of context and just like fewer instructions overall.
Um, yeah, I think it's definitely a science. It took a bunch of like evals to build. I'm not sure if you had anything else on uh the lean system. >> I think in general when you're prompting these models, you should always think about like are there edge cases to the instruction that I'm giving it? And when we went back and we reviewed all the instructions in the cloud code system prompt, we found a few cases where yes, this statement is like 90% true, but there's like a real 10% of cases where this is not true.
And we didn't want to constrain the model or like confuse it into thinking, hey, it should always do this. Like one good example is verification. Everyone here wants claw to verify its work. Um, and we had some instructions in the prompt that just said um, if you make a front-end change, always verify. But you know, there there is a limit to it. Like for example, if you if it's changing copy from one string to another string and the user tells says it says like just make a quick fix and update the test, maybe you don't want to verify.
And so we we've also adjusted our wording from saying always verify verify verify verify to like hey most of the time when you're doing front-end work you can't always understand the full experience by hitting the backend endpoints. So like when when you make like more changes to the user experience uh please run the app locally and actually in fact that instruction probably isn't even good because what is a large change like maybe it wants to change it test it for small changes too.
Um, in general, whenever you give a prompt to the model, you should always think about the ways in which it could be misinterpreted by like a well-intentioned uh other user or human in order to better understand how the model might interpret it and in order to make sure that you can soften the prompt such that it is actually 100% accurate because you are giving this prompt to the model 100% of the time. But what's fascinating about that is you're relying on the model's judgment.
And that's got to be a opus fable level thing. Like models a year ago did not have the levels of judgment necessary to decide if they were going to like test a change or not. That's absolutely fascinating. But that does break down if you're building for a wide range of models and trying to run the cheaper models for cheaper tasks and stuff. >> We actually have a different system per model now because of this very reason.
So, it's only our uh most frontier models that have this 80% token decrease and the older models actually still have the full system prompt. >> Do you think Fable and Opus are smart enough to be able to prompt Haiku with more details because they understand that Haiku has less judgment, has less taste? >> We haven't been able to evalu, but we don't have any hard data to show it. I think there's a tough thing with um smaller models sometimes because like uh you know we saw this with um just like sometimes the the larger models can be more token efficient on a hard problem than the smaller models.
And so uh you know there's like a little bit of that uh intuition to build about like you know sometimes you really just want frontier intelligence almost all the time you know um but it's uh yeah like the paralle curve shifts you know and so it's hard to find. Yeah. >> I mean that's something I found fascinating. I feel like a year ago I did not trust a model to write a prompt. Like today the good models are very good at prompting.
Like a lot of my prompts are written by models which feels absurd but it actually works really well. And something that helped me come to terms with that was thinking about sub agents which is entirely about a clawed model setting up a prompt for another claw model so that it knows what to go and do. Yeah, I think workflows are actually a really good example of this because it's like cloud not just prompting a single sub agent, but it's like pro prompting like the orchestration of many sub aents and each one of them gets like, you know, a very detailed prompt.
So, it's like almost like a level above like, you know, just spawning a sub agent. Um, so yeah, it's quite good at that. Yeah, I've also been using on my personal machine like giving it the Gemini API and being like, oh, like here, generate images. And it's so good at it's way less lazy than I am at prompting an image model, you know. So, um yeah, it's just Claude prompting Claude all the way down. Yeah, >> I think Claude also wrote the prompt for the workflow tool. >> For the workflow tool, I've read that prompt.
It's a good prompt. I mean, that's actually um a frustration I have with Anthropic generally is you publish the prompts for Claude. There's there's a web page with them on, but you don't include the tool prompts and the Claude code prompts. I still have to run a proxy to intercept them. I would love it if the Claude code prompts were deliberately published because they're the documentation. They're how you know what the tool can do and how it works. >> I'll write down that feature request. >> Please do. >> I'll have Claude tag do it. >> And also the diffs like every now and then I'll do I'll diff the older and the newer prompt and that's how I learn the capabilities of the new model.
I'm really looking forward to seeing what this 80% reduction actually looks like. >> Yeah, this is on me. I have to make a post about this in detail. Yeah. >> So, what's your bar? Let's talk about tools a little bit. Claude code is basically a just bag of to a big bag of tools. What's your bar for introducing a new tool? How do you decide when it's worth doing that additional engineering at that level? >> Do you want to take it?
Because you introduced one of the best tools we have. >> Yeah. I mean, yeah, it's like my career peaked when I introduced the ask user question tool. I I think so. Um, it it's really hard. I I I think is the the especially for some tools like ask user question is claude's tool to ask you and so it's hard to eval and and sometimes it's more of a user preference thing. So like um especially back then we had fewer evals.
It was very ant fooding based um or sorry dog fooding is yeah ant fooding is you know our ant version of that. Um but yeah I mean I think overall we've been trying to trend towards fewer tools. I think the last set of tools we introduced were like the task tool I think um and try and give cloud more general versions to do this. Um >> right one of the most interesting tools is the file editing tool which but you can have file editing as a tool or you can tell teach it tell it to use said and GP and and do things that way.
What's the latest evolution of your file editing tool? >> Uh I think we still have one but like for example we removed our GP and uh other search tools. glob tools for just like native uh like bash and so uh yeah we still have one I I think this is kind of like I said in my uh talk earlier that the models are kind of like more of a biology than a physics you know and so like you know it's hard to like especially tool design I think is quite hard and I'm not sure if actually if cat disagrees and is like oh like we should no there's like a science to the ebal of it but I'm sort of like yeah tool design is more of an art maybe or like a biology Yeah, >> I think I largely agree, but that I think in general as we introduce more tools, it like in general we try to keep the cardality pretty low and make sure that every tool we add has a distinct function from every other tool so that Claude can very easily distinguish when to call each. for file edit.
Actually, the reason that we have file edit is because we can render it because back back in the day we used to show or I guess we still do. Um, so we show people when quad makes a file change and there's this nice like dedicated UI that just says uh do you approve this edit to this file? And the reason that we had a dedicated file edit tool was so that we could deterministically know that quad was making a file so we could show people this nice UI.
And for a lot of the new users who are onboarding, I think they still really like this experience. So, we've kept it around. But for a lot of us who are on auto mode right now, or hopefully you're not on YOLO mode. But anyway, for a lot of us right now, I don't think actually it matters and we probably could just remove file edit and we'll be totally fine. >> So, let's talk about auto mode or let's talk about safety and security in general.
Like I am deeply aware of the risks of prompt injection and there there are so much bad things can happen if somebody else tells my clawed code what to do. I still mostly run clawed code in yellow mode and feel incredibly guilty about it. What's the advice within anthropic for safely running clawed code? Like what do you tell people to do? >> Why not auto mode? >> I am starting to use auto mode and I don't understand it enough to get how safe it is.
But yeah, that's as of maybe three weeks ago, I'm defaulting to auto mode. >> Okay, so broadly within Enthropic, almost every single person uses auto mode. It is the best way to do longunning work in quad code while being safe. We've done extensive bashing. We have thousands of eval. We've commissioned many red team teamers to create adversarial environments in order to trick cloud code into doing bad actions. and we've mitigated every single issue that they found.
And so we're going to publish some evals in the in the coming weeks, but we we've pretty much mitigated every attack. >> That is a big claim. That's very exciting. If that holds up, >> we we will we'll share the evals for it so uh f folks can assess, but we've been extremely diligent about identifying all the ways in which quad might mess up and then updating auto mode to counter it. It doesn't catch 100% of things. Um I don't think any Yeah, that would be way too strong of a claim.
But for the main categories of risks that we're concerned about like prompt injection, data exfiltration, um the risks are far lower than the average human reviewer. >> So, and oh yeah, a little bit on how auto mode works. I I think it's useful to build this mental model. So, whenever Claude is doing a turn, uh there's a or a bash call, uh there's a sonet classifier that is judging the tool and also the context of the conversation, your instruction, right?
And so there are some things around like uh permissions which are dependent on your request right so you don't want to give git push like permissions all the time but if you say hey push this to github you want it to do it right and so auto mode will or if you say don't push you want it to deny it right and so auto mode will do that particular thing happens to me a lot where it's like oh like auto mode like I like claw tried to do this because it's very helpful and proactive and auto mode saw like oh like you know uh you know don't do this and it like surfaced it.
So it's good at like the dynamic permissions that you yourself give inside of the prompt which I think is really important. Um it also works well with our sandboxing infrastructure because like sandboxing is one of those things where there are so many different edge cases. Uh and it's hard for like us to deterministically follow them. But if you we have a sandbox and something needs to escape the sandbox like a you know a network request auto mode can then look at that request and be like oh hey does this like you know does this make sense right and just allow that in. >> So I hadn't realized auto mode is interacting with the networking sandbox as well. >> Yeah exactly.
So it's also part of sandbox. Um yeah >> it it interacts with any permission prompt that the user would otherwise see. >> And how old is auto mode? like I feel like as a feature that I had access to, it's it's only a couple of months old, right? >> We've been using it within Anthropic since January. >> Okay. >> So, we've been hardening it for for quite a while. And it's obviously Anthropic is extremely focused on safety and security.
And so, we've been working broadly across our um alignment and safeguards teams in order to enable the rollout internally, build out these eval even more robust before sharing it out with the world. I think my only my main problem with automotive is I don't understand it deeply enough. Like for anything that's looking after my security, I want to know as much as I can about how it works and what it protects me against and what it doesn't so I can decide like how much I can trust it. >> I I think yeah Dell was working on a post about this.
So um a little bit just a little bit more about auto mode. This is also the reason cloud tag is so so good, right? because cloud tag uses auto mode and like you can imagine that like one of like I've heard a lot of like build versus buy questions on on a Slackbot. I'm like please you probably shouldn't build your own AI Slackbot, you know, like there's so many attack vectors, you know what I mean? And like um like you have a a feedback channel that like you know users can post feedback into and now your bot is reading it, right?
And and so I think that like this the work we've put in with auto mode and you know we have a general Swiss cheese defense of like security, right? We also like yeah you know like RL against this stuff and things like that. Um I think this is really what makes cloud tag work. It like just works seamlessly with your permissions and yeah like you know we you don't want to be prompt injected in your Slack. Yeah. Yeah. >> Do you have any are there any more security things in the pipeline beyond that go beyond auto mode?
Um I I I think I mean I I think we're very secure like like so we with claude tag you can uh peri can provision your own like sort of credentials for claude so it doesn't need to act on your uh on your behalf. You can have like claude as an identity and and that also makes it easier to audit and inspect what claude is doing. >> Well I guess cuz claude tag is influenced by anyone who can talk to it. So it's it's got a much wider pool of people who are telling it what to do. >> That's right.
Yeah. Yeah. And of course we have probes as well like with like mythos and uh sorry with fable um and that's also like a downstream effect of our safety and research work. And I think this is the moment where you sort of see AI like Anthropic being an AI safety company really paying off when you like you know we really want Claude to be able to run in an aligned way over long periods of time and like yeah automode has to be basically flawless for this to work right and it's sort of like all downstream of our like you know our being an AI safety company.
We also launch trusted devices for the remote control users out here who want to be safer. Um, and for all of our remote environments, we support uh credential injection. So if you want like quad code to be able to access data dog but you don't want quad code itself to hold the data dog credential you can set up um our identity uh credential management system so that the data dog credentials are only usable by the agent but not accessible by the agent.
So we insert it on the fly when the agent tries to make a data dog request. This is that that token the proxying trick, right? Where the proxy knows anytime somebody calls this an API. Dog.com address with a token, replace the token with the real thing. I love that pattern. I'm seeing that in a whole bunch of places. It feels so obviously right to me. Let's talk a little bit about the human element. And you touched on this in the keynote this morning, but um a lot of people are feeling a sense of loss now that so much of what they considered to be their role in in building software is is is being subsumed by the models.
Um how do you think about that? Like um how firstly how has the past year and a half thought changed the way you think about your own craft and the value that you add? >> Yeah, I I think for me and I think Kat is always such a good reminder. Ken and Boris are such good reminders of like you have to be more ambitious. They're always like, you know, you have to like like, you know, we're growing so fast, we have to be on the edge, we have to do like the best work we can.
Um, I think that that's kind of like a constant reminder for me where I'm like anytime I'm like kind of like slow on something. I'm like, okay, can I do it faster? Can I be more ambitious here? Um, I think the like point on and like oftentimes the answer is claude because claude is getting better as you go. So, I'm like, "Oh, the last time I tried this, I it was with the previous model or something." Um, I think with your point on loss, I think this is real.
And I I I do feel that like if you're only trying to do the same work you were doing before LLMs and now it's like a prompt, it it is like I think kind of a sad feeling. And I think the way you offset that is by being more ambitious, right? And you're like, you know, like I love like I think Jared is such a good example where he's like hand wrote all of the Zig code in his Oakland apartment in like a year barely left his house, you know, and then uh and he had so much fun doing that and now I see him like rewrite all of Bon into Rust and he's having so much fun doing that, right? and it's like so much more ambitious and that's how how like he sort of like offsets that and I think just generally being like okay like how do I do the bigger thing and and and do more and I think success is fun you know and that's how I kind of like >> the it's changing your ambition it's changing what you do because what you did before is a lot easier in quotes but now we can take on these bigger challenges >> yeah I think there's just like on average everyone has things they wish they did you know what I mean and that they were better at and I think now it's Let's let's do it, you know, like let's Yeah. >> And Cap, what does that look like from a sort of product management perspective? >> I feel like the product role just changes every single month and it's very much just identifying, okay, what are um like all the PMs on our team are like this like mix of engineer, designer, PM.
Um most of the engineers on our team actually used to be uh full-time engineers in the past. And so for us it really means like plugging in whenever there's any kind of gap. So if it's like okay we have this idea and we didn't inspire any engineer to go build it then like we should just build it and then put it into a notebook and inspire people to take this to production. Or if the designs look a little off, well let's let's take a page that's similar and do a first pass design and uh tag in tag in someone who's very detail oriented to like fill in the gaps. or if we notice that like hey now now our team uh and our product adoption is a bit bigger within the company and more people need to know what's coming down the pipe for quad code quad tag and co-work uh what we do then is like okay let let us like automate figuring out our whole launch calendar let's automate getting those status updates asynchronously so we're not bugging people and then let's figure out okay these are our three internal announce channels and make sure that our updates there are fully detailed and to the point.
And so for us, it's very much just understanding what is the gap right now between a great idea and getting something to our customers and then how do we automate it as much as possible. >> It sounds to me like with product management, there's always more to do, right? It's I feel like one of the things that makes me feel good is I've never worked at a company that didn't have a backlog of a thousand things they wanted to do and didn't have the resources to take on.
Um, so what's a moment when Claude has surprised you? Like when when when the model has done something that genuinely surprised you that you you didn't think it would be able to do that. >> Yeah, I mean I I've posted a lot about uh cloud video editing, but like like most recently I I gave a talk at the ACM Agentic conference and uh I was like, "Hey guys, do you have the edited video? I'd love to post it and share with my coms team." And they're like, "Oh, it's taking so long." And I'm like, "Okay, could you send me the raw files?" So they send me the video of me talking on stage, the video of the deck and the audio file, and they're like, "Good luck." And and so I like give this to Claude.
I give it my HTML deck as well, and I'm like, "Hey, can you just like edit this together, you know?" And um what it does is like honestly incredible. Like I I'm ready to ship it. Like so it transcribes the entire video. Um it notices that sometimes the video of my deck is a little bit weird. there's like a popup of like an auto update in the middle and it's like, "Oh, I probably shouldn't use the video of your deck. Actually, what I'm going to do is I'm going to slice up and figure out which um slide you're on and instead use your h the HTML source, right?" And so, it's displaying the HTML source.
Then, it's got a video of me, but you know, like I'm only taking up a small part of the stage. And so, it's cropping dynamically where I am on the stage. Um, and like I'm I'm pacing. So, it's like tracking me as I'm pacing. And I've got like this crop of me, the the deck, and then like it's transcribing what I'm saying. >> This was Fable, right? >> This Fable. Yeah. Yeah. Absolutely. Yeah. >> Um and it's just like like, you know, it was a good prompt, but it was a oneshot prompt.
And and then I asked it to like add some like interesting animations and graphics, and it I was just like blown away kind of, you know, and it just like does all this stuff. It does ffmpeg, it does remotion, it does. >> So I have to ask the follow-up. What can't it do? What are the things where you're still disappointed? You're waiting for Claude Fable 6 to to figure out for you. >> I wanted to have better design and UX taste. >> Uhhuh.
[laughter] >> Like I feel like it's now at the point where if I give it if I write out a prompt with a detailed spec of how I want a feature to behave, it will usually behave that way. But, you know, the the paddings might be off or like the the interface is it's just not delightful yet. Um, I think it kind of leans on like existing best practices for apps for how apps are designed, but I feel like for Frontier AI products, there's so many new interaction experiences that we still have to have yet to design. >> There's an Opus aesthetic.
You can look at something and go, "Yeah, that was designed by Opus." It' be good if we could move beyond that. >> Yeah. Yeah. It like I I'm very excited for future models to hopefully be like interaction design thought partners. Hm. Uh what can't it do? Um I I think I would love to see it, you know, interact more with the real world like okay like can it do this like you know can it solve science right? Like can it like orchestrate you know the experiments and there's some amount of coding that goes into that but there's also this like other taste of uh you know like the broader world that it needs.
So >> Claude Science is a new product that just came out a few days ago right? >> Yeah but I have no context on it. >> I was going to ask is that part of Claude code or is that a separate separate sex section? >> It's our partner team. >> Gotcha. Try it out, though. >> Um, so we've got I've got a couple of closing questions. Um, which parts of Anthropic's company culture do you think uniquely help Anthropic be productive with these tools that other companies should steal?
What are the cultural hacks that people should be should be adopting from you? >> I'll share one and then you go. Um, I'll share one for quag. So claw tag works best when you have it in a public channel and when most of your channels are public. Claw tag is able to search across all public channels to get as much context as possible to give you the highest accuracy answer. And it's only able to do this if it has access to everything. >> Yeah, I mentioned this in my keynote, but I I think I it's so important to me.
I want to reemphasize like I I think the co-founders say like we don't negotiate against ourselves, you know, and I think this is really important where you're like you can imagine trade-offs in your head and talk yourself out of doing something ambitious, you know? Um or you can just try and do the ambitious thing. And I think that like we're just so often being like, okay, what if we just did it? Like what if like, you know, like is this a real trade-off or not, right?
Or like and if so, like why? like where's the proof that it's a real trade-off and not just like it sounds reasonable, right? And so I think just yeah like you know make the trade-offs show themselves to you. Be as ambitious as you can. >> That's so because that goes against I've got 25 years of software experience that says the default answer should be no. Everything is a trade-off. Everything has a cost. And now we're having to reimagine all of those intuitions.
It's it's kind of fascinating. Okay. And so final question for both of you. What is something what what's one of your favorite absurd things that you've built with Claude just because you could build it? >> I can go. Well, you think um I'm working on a 2D Street Fighter fighting game uh with me as a character and like my friends as well. Um and it uses cloud code to prompt you know Gemini and honestly the sea dance model is pretty good like uh to make like video animations.
Um, and uh, it works great like like it's so good at prompting. It's like, you know, it can verify like the frames to check if this was a good animation. >> Is this Street Fighter 2 level 2D sprites that you're generating? >> Yeah, exactly. Yeah. Yeah. Yeah. Like like 2D sprites. The animation looks amazing. And it can also figure out hitboxes. It can be like, oh, you know, your fist is like here, I'll draw the JSON hitbox.
Yeah. Yeah. It's like incredible. Yeah. Yeah. So, um, I don't know if I'll put this out, but it's uh >> I feel like we need a screenshot at least. This sounds I can make a screenshot happen. Yeah. Yeah. >> Mine is much more simple. Um I'm a big rock climber and a lot of my friends climb and so we have this little app that we built with quad code that where we just log all the projects that we're working on and we also uh go outdoors together a lot.
So we have uh quad do all this research with workflows. Workflows is amazing. Like we brand it as a coding tool but it's amazing for doing deep research for travel. Um, I also plan our team off sites and it's good at finding venues that can fit all of us. Um, it yeah there it has a lot of uh side benefits. But anyway, um I also use workflows to just research all the climbing destinations that we might want to go to, what has direct flights from where all of us are located.
It goes to mountain project and finds all the climbs that are in our grade level. It finds the Airbnb and it actually maps out like I I don't like hiking and so I care a lot about it having a very short approach. So very short walking distance from where the car parks to where the rock actually is. And so it it filters for this and so I like with existing apps I have to like manually click through mountain project but with this I just put in all of our preferences and it's just a custom app for us. >> So you're you're basically vibe coding Jira for mountain climbing. >> Exactly.
That's that's pretty fantastic. Um, you know what? We have time for a couple of audience questions. Um, if you want to sh if you want to come forward and say them to me and I will repeat them so everyone can hear them. But yeah, um, tell you what, anyone who gets here first gets to ask a question. Sorry for people at the back. Hey um actually >> yeah uh my question is that um do you have any near plan to build more uh eval uh tools for us to build eval data set or anything like that and um more observability tools to monitor the performance of agents and uh workflows.
We've considered building eval tools, but I think the limiting factor actually tends to be that it takes a long time for customers to build really high quality evals. And so I think the tooling is less of the constraint and more of the skill set of how do you build a great eval? And that's an area where we're excited to both invest internally and also hopefully we can share some of the best practices externally. >> Hey Katar, uh my name is Sai.
Uh so my question was because I'm more interested in the memory and the multiplayer. How does how is memory being designed? So two questions right uh so how is memory being designed today I I assume it's around files. So and second part of it is have you thought about thinking in an orthogonal direction where you would actually need a data store to store these memory instead of files to scale it better. So I think that's my question. >> Yeah right now for cloud tag the memory is channel specific.
So every claude in that channel has a shared memory and then you know the instances have a session um but like the session can contribute back to main memory. We do a lot of memory research and it's you know can be kind of unintuitive like what what is the right way to do memory. Uh but yeah we're always working on this. So yeah >> uh yeah I mean we're always running me memory experiments. I don't have you know anything to yeah like how it works right now in cloud tag is a markdown file per channel.
Yeah. >> Thank you. >> So I'm afraid we're out of time. Please join me in thanking Cat and Tariq and we will be around for more questions um in the hallway. >> Thanks guys. >> Thank you. [music]
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script. No signup, no login.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.