I was using Fable wrong, this is how I fixed it
Read full transcript 32 segments
-
Just a few months ago, people were Just a few months ago, people were accusing me of being paid by OpenAI accusing me of being paid by OpenAI accusing me of being paid by OpenAI because I liked the models they were because I liked the models they were because I liked the models they were shipping that much. I thought that was shipping that much. I thought that was shipping that much. I thought that was kind of insane. What's even more insane kind of insane. What's even more insane kind of insane. What's even more insane though is the fact that Sam Altman just though is the fact that Sam Altman just though is the fact that Sam Altman just called me an Anthropic fanboy called me an Anthropic fanboy called me an Anthropic fanboy specifically because of how much I specifically because of how much I specifically because of how much I prefer using Fable over Astra for my prefer using Fable over Astra for my prefer using Fable over Astra for my day-to-day work. There are a lot of day-to-day work. There are a lot of day-to-day work. There are a lot of reasons for this and I cover a bunch of reasons for this and I cover a bunch of reasons for this and I cover a bunch of them in my Fable versus Astra video, but them in my Fable versus Astra video, but them in my Fable versus Astra video, but that's not what I want to rehash today. that's not what I want to rehash today. that's not what I want to rehash today. In that comparative video, I vibed out a In that comparative video, I vibed out a In that comparative video, I vibed out a graph showing how I feel the quality of graph showing how I feel the quality of graph showing how I feel the quality of the output I get from the models differs the output I get from the models differs the output I get from the models differs where Astra can be way higher than Fable where Astra can be way higher than Fable where Astra can be way higher than Fable at points but way way worse at others. at points but way way worse at others. at points but way way worse at others. Fable still makes its mistakes, but it Fable still makes its mistakes, but it Fable still makes its mistakes, but it was less noisy by far. That comes with was less noisy by far. That comes with was less noisy by far. That comes with benefits outside of trusting it more. It benefits outside of trusting it more. It benefits outside of trusting it more. It means the model can go longer and do means the model can go longer and do means the model can go longer and do more in a given pass. It means you can more in a given pass. It means you can more in a given pass. It means you can tell it to come back with a video when tell it to come back with a video when tell it to come back with a video when it's done and it actually can do that. it's done and it actually can do that. it's done and it actually can do that. It also means that the little dips it It also means that the little dips it It also means that the little dips it does have hurt more and finding ways to does have hurt more and finding ways to does have hurt more and finding ways to reduce them is incredibly valuable. I've reduce them is incredibly valuable. I've reduce them is incredibly valuable. I've spent a lot of time with this model. spent a lot of time with this model. spent a lot of time with this model. It's probably I've ever used a model It's probably I've ever used a model It's probably I've ever used a model that I was paying for in such a short that I was paying for in such a short that I was paying for in such a short window. I have shipped so much code with window. I have shipped so much code with window. I have shipped so much code with this model it's insane and I've been this model it's insane and I've been this model it's insane and I've been burning my five accounts to the ground burning my five accounts to the ground burning my five accounts to the ground using it. I wish I hadn't burned as many using it. I wish I hadn't burned as many using it. I wish I hadn't burned as many of them as I did before reading this of them as I did before reading this of them as I did before reading this particular guide from Anthropic because particular guide from Anthropic because particular guide from Anthropic because the prompting Claude Fable 5.1 doc is the prompting Claude Fable 5.1 doc is the prompting Claude Fable 5.1 doc is actually really good and has a bunch of actually really good and has a bunch of actually really good and has a bunch of insights that are worth learning from.
-
insights that are worth learning from. insights that are worth learning from. The goal of this video is to try and The goal of this video is to try and The goal of this video is to try and take all of the things I figured out take all of the things I figured out take all of the things I figured out over the last month-ish of using the over the last month-ish of using the over the last month-ish of using the model and distill it into the most model and distill it into the most model and distill it into the most useful pieces of info I can possibly useful pieces of info I can possibly useful pieces of info I can possibly give you. Everything from how to fix give you. Everything from how to fix give you. Everything from how to fix your agent CMD to how to give the model your agent CMD to how to give the model your agent CMD to how to give the model the tools it needs to keep going and get the tools it needs to keep going and get the tools it needs to keep going and get better results to even using Codex to better results to even using Codex to better results to even using Codex to help make your Claude outputs better. help make your Claude outputs better. help make your Claude outputs better. But as I mentioned before, I'm paying But as I mentioned before, I'm paying But as I mentioned before, I'm paying for these accounts myself and if I'm for these accounts myself and if I'm for these accounts myself and if I'm going to need another, I hope you can going to need another, I hope you can going to need another, I hope you can understand why we're doing a quick understand why we're doing a quick understand why we're doing a quick sponsor break. There's a pretty good sponsor break. There's a pretty good sponsor break. There's a pretty good chance your company is shipping more chance your company is shipping more chance your company is shipping more code than it's ever shipped before. This code than it's ever shipped before. This code than it's ever shipped before. This also means you're probably shipping out also means you're probably shipping out also means you're probably shipping out more bugs than ever before. And those more bugs than ever before. And those more bugs than ever before. And those bugs are getting more insidious and bugs are getting more insidious and bugs are getting more insidious and harder to find, especially when you're harder to find, especially when you're harder to find, especially when you're building things with agents, not just building things with agents, not just building things with agents, not just using agents to build, but also serving using agents to build, but also serving using agents to build, but also serving agents to your users. If you were around agents to your users. If you were around agents to your users. If you were around before agents, you've almost certainly before agents, you've almost certainly before agents, you've almost certainly heard about Sentry. And honestly, if heard about Sentry. And honestly, if heard about Sentry. And honestly, if you've been around since, you have, too, you've been around since, you have, too, you've been around since, you have, too, because these guys are the go-to because these guys are the go-to because these guys are the go-to platform for identifying bugs in your platform for identifying bugs in your platform for identifying bugs in your real-world applications. I've had Sentry real-world applications. I've had Sentry real-world applications. I've had Sentry set up on pretty much every codebase I set up on pretty much every codebase I set up on pretty much every codebase I have worked in in the last, I don't even have worked in in the last, I don't even have worked in in the last, I don't even want to think about how many years. God, want to think about how many years. God, want to think about how many years. God, I've been around for a while. Needless I've been around for a while. Needless I've been around for a while. Needless to say, these guys know how to find bugs to say, these guys know how to find bugs to say, these guys know how to find bugs in your apps, and the Sentry MCP makes in your apps, and the Sentry MCP makes in your apps, and the Sentry MCP makes it so your agents can use that same it so your agents can use that same it so your agents can use that same data, as well. But I'm not just here to data, as well. But I'm not just here to data, as well. But I'm not just here to tell you that Sentry exists. You tell you that Sentry exists. You tell you that Sentry exists. You probably already know that. I'm here to probably already know that. I'm here to probably already know that. I'm here to show you the really cool things they've show you the really cool things they've show you the really cool things they've set up for building with agents. I set set up for building with agents. I set set up for building with agents. I set up a demo in T3 code to get some traces, up a demo in T3 code to get some traces, up a demo in T3 code to get some traces, and god damn, the info you get is so, so and god damn, the info you get is so, so and god damn, the info you get is so, so useful. They break down the cost of the useful. They break down the cost of the useful. They break down the cost of the whole request, and even better, the cost whole request, and even better, the cost whole request, and even better, the cost of each chunk within the request. This of each chunk within the request. This of each chunk within the request. This timeline view is super helpful for timeline view is super helpful for timeline view is super helpful for figuring out what happened when, and figuring out what happened when, and figuring out what happened when, and where the costs were. So, we can see where the costs were. So, we can see where the costs were. So, we can see that this top-level request cost 40
-
that this top-level request cost 40 that this top-level request cost 40 cents or so, but more importantly, we cents or so, but more importantly, we cents or so, but more importantly, we can see where all that money went. And can see where all that money went. And can see where all that money went. And if you're not that into timeline-style if you're not that into timeline-style if you're not that into timeline-style views, I understand. It doesn't really views, I understand. It doesn't really views, I understand. It doesn't really map to my head very well, especially map to my head very well, especially map to my head very well, especially when you're used to using a chat view when you're used to using a chat view when you're used to using a chat view for actually doing these things. For for actually doing these things. For for actually doing these things. For those who weren't watching, I just those who weren't watching, I just those who weren't watching, I just transitioned over to the chat view that transitioned over to the chat view that transitioned over to the chat view that they built into their agent platform. they built into their agent platform. they built into their agent platform. Yes, you can actually see a transcript Yes, you can actually see a transcript Yes, you can actually see a transcript for the back and forth that your agents for the back and forth that your agents for the back and forth that your agents or your users or whoever else had that or your users or whoever else had that or your users or whoever else had that led to these issues. And if I slide this led to these issues. And if I slide this led to these issues. And if I slide this over, you can see all the additional over, you can see all the additional over, you can see all the additional info, from how many tokens were used, info, from how many tokens were used, info, from how many tokens were used, how many errors were hit, and how much how many errors were hit, and how much how many errors were hit, and how much the request actually cost to run. And as the request actually cost to run. And as the request actually cost to run. And as you scroll, you can see how the user you scroll, you can see how the user you scroll, you can see how the user experienced this, with the timestamp experienced this, with the timestamp experienced this, with the timestamp showing when different things happened, showing when different things happened, showing when different things happened, what worked, what didn't, and the actual what worked, what didn't, and the actual what worked, what didn't, and the actual cost of everything the user ends up cost of everything the user ends up cost of everything the user ends up seeing. This is particularly useful if seeing. This is particularly useful if seeing. This is particularly useful if you're building custom tools and you're building custom tools and you're building custom tools and interfaces like an MCP that your agents interfaces like an MCP that your agents interfaces like an MCP that your agents are working with, because you can debug are working with, because you can debug are working with, because you can debug not just what the model sent to the MCP, not just what the model sent to the MCP, not just what the model sent to the MCP, but the whole pipeline and all the code but the whole pipeline and all the code but the whole pipeline and all the code that actually serves that agent's that actually serves that agent's that actually serves that agent's request. Figure out what your agents and request. Figure out what your agents and request. Figure out what your agents and your code are actually doing at your code are actually doing at your code are actually doing at swyd.link/sentry. swyd.link/sentry. swyd.link/sentry. I'm going to start with the official I'm going to start with the official I'm going to start with the official prompting guide because the amount prompting guide because the amount prompting guide because the amount that's changed with 5.1 to 5 is that's changed with 5.1 to 5 is that's changed with 5.1 to 5 is meaningful, but it's almost all subtle meaningful, but it's almost all subtle meaningful, but it's almost all subtle things. So, things. So, things. So, it's worth going through.
-
it's worth going through. it's worth going through. The first section they have is titled The first section they have is titled The first section they have is titled consider all effort levels. I find this consider all effort levels. I find this consider all effort levels. I find this one a little cringe cuz it forgets the one a little cringe cuz it forgets the one a little cringe cuz it forgets the fact that we only have 50% of our limit fact that we only have 50% of our limit fact that we only have 50% of our limit as a subscriber. And I'll I'll be real as a subscriber. And I'll I'll be real as a subscriber. And I'll I'll be real up front with this. I don't think Fable up front with this. I don't think Fable up front with this. I don't think Fable is worth the money if you're not getting is worth the money if you're not getting is worth the money if you're not getting it subsidized through a subscription. it subsidized through a subscription. it subsidized through a subscription. Paying the full API price just kind of Paying the full API price just kind of Paying the full API price just kind of sounds insane to me. I don't think Astra sounds insane to me. I don't think Astra sounds insane to me. I don't think Astra is much better here, sadly, because is much better here, sadly, because is much better here, sadly, because Astra can often get stuck in a loop and Astra can often get stuck in a loop and Astra can often get stuck in a loop and end up spending way more than it should. end up spending way more than it should. end up spending way more than it should. Fable is at least a bit more likely to Fable is at least a bit more likely to Fable is at least a bit more likely to stop when it should. So, on a median stop when it should. So, on a median stop when it should. So, on a median task, Fable is definitely more task, Fable is definitely more task, Fable is definitely more expensive, but on the extremes, I find expensive, but on the extremes, I find expensive, but on the extremes, I find that Astra can just burn usage in stupid that Astra can just burn usage in stupid that Astra can just burn usage in stupid ways. Now I've that out of the way, I ways. Now I've that out of the way, I ways. Now I've that out of the way, I want to make sure it's clear what I'm want to make sure it's clear what I'm want to make sure it's clear what I'm talking about is subscription usage talking about is subscription usage talking about is subscription usage because I just I don't think it's worth because I just I don't think it's worth because I just I don't think it's worth using any of these over API prices. And using any of these over API prices. And using any of these over API prices. And that's also why a lot of companies that's also why a lot of companies that's also why a lot of companies probably aren't letting you use these probably aren't letting you use these probably aren't letting you use these models. And for that, I am sorry. models. And for that, I am sorry. models. And for that, I am sorry. Assuming you're on a subscription, you Assuming you're on a subscription, you Assuming you're on a subscription, you have half your limit reserved for Fable, have half your limit reserved for Fable, have half your limit reserved for Fable, and the other half of your weekly limit and the other half of your weekly limit and the other half of your weekly limit could be used for whatever else. That could be used for whatever else. That could be used for whatever else. That other half effectively is then reserved other half effectively is then reserved other half effectively is then reserved for Opus 5, which is not a good model.
-
for Opus 5, which is not a good model. for Opus 5, which is not a good model. Thankfully, it does seem 5.1 or 5.2, Thankfully, it does seem 5.1 or 5.2, Thankfully, it does seem 5.1 or 5.2, whatever they call it, is coming soon whatever they call it, is coming soon whatever they call it, is coming soon for Opus, which should hopefully, for Opus, which should hopefully, for Opus, which should hopefully, fingers crossed, make it way less spiky. fingers crossed, make it way less spiky. fingers crossed, make it way less spiky. But it's it's bad. I usually find myself But it's it's bad. I usually find myself But it's it's bad. I usually find myself at the end of a given week with most of at the end of a given week with most of at the end of a given week with most of that other 50% left in my Fable driven that other 50% left in my Fable driven that other 50% left in my Fable driven to zero. Now that we have that context, to zero. Now that we have that context, to zero. Now that we have that context, I want to talk about the consider all I want to talk about the consider all I want to talk about the consider all effort levels call out here. They highly effort levels call out here. They highly effort levels call out here. They highly recommend that you try out other effort recommend that you try out other effort recommend that you try out other effort levels, that you start at high, but then levels, that you start at high, but then levels, that you start at high, but then you test other ones against your own you test other ones against your own you test other ones against your own evals. No one's evaluating how this evals. No one's evaluating how this evals. No one's evaluating how this works in Claude Code quite to that works in Claude Code quite to that works in Claude Code quite to that level. It is nice to get a gut feel to level. It is nice to get a gut feel to level. It is nice to get a gut feel to like take a task that you know how it like take a task that you know how it like take a task that you know how it should go and ask this model to answer should go and ask this model to answer should go and ask this model to answer it or solve it at different effort it or solve it at different effort it or solve it at different effort levels to see if it can figure it out. levels to see if it can figure it out. levels to see if it can figure it out. But that ends up being a lot of effort But that ends up being a lot of effort But that ends up being a lot of effort and burning a lot of tokens, so I and burning a lot of tokens, so I and burning a lot of tokens, so I honestly recommend just kind of gut honestly recommend just kind of gut honestly recommend just kind of gut feeling it. That said, I have not had as feeling it. That said, I have not had as feeling it. That said, I have not had as good of an experience with Fable on low good of an experience with Fable on low good of an experience with Fable on low and medium as I have with high. So I and medium as I have with high. So I and medium as I have with high. So I personally kind of just default to high. personally kind of just default to high. personally kind of just default to high. On really deep thorough things, I'll On really deep thorough things, I'll On really deep thorough things, I'll bump to X high occasionally. I almost bump to X high occasionally. I almost bump to X high occasionally. I almost never use max unless I'm just trying to never use max unless I'm just trying to never use max unless I'm just trying to like burn a limit. But I found that high like burn a limit. But I found that high like burn a limit. But I found that high and X high are plenty for most things.
-
and X high are plenty for most things. and X high are plenty for most things. With low and medium, I find that it's With low and medium, I find that it's With low and medium, I find that it's likely enough to fail on those reasoning likely enough to fail on those reasoning likely enough to fail on those reasoning levels that I try to avoid them just levels that I try to avoid them just levels that I try to avoid them just because it ends up being more tokens if because it ends up being more tokens if because it ends up being more tokens if you have it do the task on low and it you have it do the task on low and it you have it do the task on low and it fails and then you redo it on medium or fails and then you redo it on medium or fails and then you redo it on medium or high. I'd rather just have it run the high. I'd rather just have it run the high. I'd rather just have it run the thing and come back with a result. It's thing and come back with a result. It's thing and come back with a result. It's also worth noting that if you use high also worth noting that if you use high also worth noting that if you use high on a simple task, it's not going to be on a simple task, it's not going to be on a simple task, it's not going to be that expensive because it is not going that expensive because it is not going that expensive because it is not going to just use and waste more reasoning if to just use and waste more reasoning if to just use and waste more reasoning if it doesn't need to. Here, I'll even go it doesn't need to. Here, I'll even go it doesn't need to. Here, I'll even go crazy here. I'll use max reasoning for crazy here. I'll use max reasoning for crazy here. I'll use max reasoning for this. Hi, how are you doing today? this. Hi, how are you doing today? this. Hi, how are you doing today? Really complex prompt we're about to Really complex prompt we're about to Really complex prompt we're about to send it. Wow, it thought for so long there. That Wow, it thought for so long there. That was horrible. Yeah, you get the idea. was horrible. Yeah, you get the idea. was horrible. Yeah, you get the idea. Going from low to max, if it's a simple Going from low to max, if it's a simple Going from low to max, if it's a simple task, it's going to stay simple. It's task, it's going to stay simple. It's task, it's going to stay simple. It's not going to waste a ton of tokens. It's not going to waste a ton of tokens. It's not going to waste a ton of tokens. It's not as good with the adjustment based on not as good with the adjustment based on not as good with the adjustment based on task size as Astra is, but it's good task size as Astra is, but it's good task size as Astra is, but it's good enough that I don't feel like I'm enough that I don't feel like I'm enough that I don't feel like I'm wasting too much when I use high when I wasting too much when I use high when I wasting too much when I use high when I shouldn't. I will say Astra's even shouldn't. I will say Astra's even shouldn't. I will say Astra's even better at this where I had a benchmark better at this where I had a benchmark better at this where I had a benchmark skate bench where I couldn't get the skate bench where I couldn't get the skate bench where I couldn't get the model to use meaningfully more tokens on model to use meaningfully more tokens on model to use meaningfully more tokens on max than it used on low. It was like a max than it used on low. It was like a max than it used on low. It was like a 50 token gap, so like 130 to 180 or so 50 token gap, so like 130 to 180 or so 50 token gap, so like 130 to 180 or so at worst case. Like it doesn't do a at worst case. Like it doesn't do a at worst case. Like it doesn't do a whole bunch. Meanwhile, Gemini 1.5 Pro whole bunch. Meanwhile, Gemini 1.5 Pro whole bunch. Meanwhile, Gemini 1.5 Pro will do like 1,000 plus reasoning tokens will do like 1,000 plus reasoning tokens will do like 1,000 plus reasoning tokens for that same task in the same bench. So for that same task in the same bench. So for that same task in the same bench. So Astra adjusts its usage of the window Astra adjusts its usage of the window Astra adjusts its usage of the window it's given better based on the size of it's given better based on the size of it's given better based on the size of the task. Fable still does it well the task. Fable still does it well the task. Fable still does it well enough that I'm fine just leaving it on enough that I'm fine just leaving it on enough that I'm fine just leaving it on high. I'm going to switch it back from high. I'm going to switch it back from high. I'm going to switch it back from max to high before I forget to.
-
max to high before I forget to. max to high before I forget to. Generally speaking, only use low and Generally speaking, only use low and Generally speaking, only use low and medium if you specifically want the medium if you specifically want the medium if you specifically want the model to stop earlier or not go too model to stop earlier or not go too model to stop earlier or not go too deep. Or you like if you know the deep. Or you like if you know the deep. Or you like if you know the thing's a rabbit hole and you want to thing's a rabbit hole and you want to thing's a rabbit hole and you want to try and keep it from falling down the try and keep it from falling down the try and keep it from falling down the rabbit hole, lower reasoning levels can rabbit hole, lower reasoning levels can rabbit hole, lower reasoning levels can help a bit. I just leave it on high at help a bit. I just leave it on high at help a bit. I just leave it on high at this point. The reason I don't have the this point. The reason I don't have the this point. The reason I don't have the limits thing is simply cuz if I had the limits thing is simply cuz if I had the limits thing is simply cuz if I had the whole 100%, low and medium would be more whole 100%, low and medium would be more whole 100%, low and medium would be more interesting. But the things I'd use low interesting. But the things I'd use low interesting. But the things I'd use low for, I'm just going to deal with Opus 4 for, I'm just going to deal with Opus 4 for, I'm just going to deal with Opus 4 from really trying to maximize my usage from really trying to maximize my usage from really trying to maximize my usage of the subs. The next part is one of the of the subs. The next part is one of the of the subs. The next part is one of the ones that interested me the most, so ones that interested me the most, so ones that interested me the most, so much so that it kind of inspired me to much so that it kind of inspired me to much so that it kind of inspired me to make this video. Ask for user-facing make this video. Ask for user-facing make this video. Ask for user-facing progress updates. Funny enough, just a progress updates. Funny enough, just a progress updates. Funny enough, just a couple days ago Julius asked me why couple days ago Julius asked me why couple days ago Julius asked me why Fable wasn't giving him traces on why it Fable wasn't giving him traces on why it Fable wasn't giving him traces on why it was doing things in my version of T3 was doing things in my version of T3 was doing things in my version of T3 Code that he was trying. And I said, Code that he was trying. And I said, Code that he was trying. And I said, "Oh, that's cuz it doesn't do that "Oh, that's cuz it doesn't do that "Oh, that's cuz it doesn't do that unless you ask it to." And then he asked unless you ask it to." And then he asked unless you ask it to." And then he asked it to and it did. Models like Opus and it to and it did. Models like Opus and it to and it did. Models like Opus and Fable 5 both love to give reasons and Fable 5 both love to give reasons and Fable 5 both love to give reasons and updates when they were doing things. updates when they were doing things. updates when they were doing things. Every time it was going to do a new tool Every time it was going to do a new tool Every time it was going to do a new tool call or a new batch of work, it would call or a new batch of work, it would call or a new batch of work, it would say, "Okay, I found this, so now I'm say, "Okay, I found this, so now I'm say, "Okay, I found this, so now I'm going to go do that." And the result was going to go do that." And the result was going to go do that." And the result was kind of noisy, to put it lightly. And kind of noisy, to put it lightly. And kind of noisy, to put it lightly. And they decided to make it stop doing that.
-
they decided to make it stop doing that. they decided to make it stop doing that. As a 5.1, the model is writing way fewer As a 5.1, the model is writing way fewer As a 5.1, the model is writing way fewer user-facing updates during long tool user-facing updates during long tool user-facing updates during long tool call turns that Fable 5 would have call turns that Fable 5 would have call turns that Fable 5 would have written updates for instead. I've seen written updates for instead. I've seen written updates for instead. I've seen this, too. I can't tell you how many this, too. I can't tell you how many this, too. I can't tell you how many times I had Fable just go off and do times I had Fable just go off and do times I had Fable just go off and do like 150 tool calls without giving me an like 150 tool calls without giving me an like 150 tool calls without giving me an update. And you know what? I don't care. update. And you know what? I don't care. update. And you know what? I don't care. I don't want that info. I just want to I don't want that info. I just want to I don't want that info. I just want to see it when it's done. That's how I see it when it's done. That's how I see it when it's done. That's how I operate now. I kick off a prompt using operate now. I kick off a prompt using operate now. I kick off a prompt using T3 Code, command shift O, enter for the T3 Code, command shift O, enter for the T3 Code, command shift O, enter for the repo, tell it what I want, and then go repo, tell it what I want, and then go repo, tell it what I want, and then go check another thread after. That's just check another thread after. That's just check another thread after. That's just how I work. So, I don't care about this. how I work. So, I don't care about this. how I work. So, I don't care about this. But if you do, and you want to get these But if you do, and you want to get these But if you do, and you want to get these updates, there's a really simple updates, there's a really simple updates, there's a really simple solution. Ask for them. I do solution. Ask for them. I do solution. Ask for them. I do legitimately think this is awesome. It's legitimately think this is awesome. It's legitimately think this is awesome. It's crazy the models are smart enough and crazy the models are smart enough and crazy the models are smart enough and trained well enough that you can get trained well enough that you can get trained well enough that you can get this type of like deep behavioral change this type of like deep behavioral change this type of like deep behavioral change by just asking. You could add something by just asking. You could add something by just asking. You could add something in the system prompt like before you in the system prompt like before you in the system prompt like before you start saying a line what you're about to start saying a line what you're about to start saying a line what you're about to do. Brief updates while you work help do. Brief updates while you work help do. Brief updates while you work help the user follow along. Close with a the user follow along. Close with a the user follow along. Close with a short recap that stands on its own. short recap that stands on its own. short recap that stands on its own. {m-dash} What you found, what you did, {m-dash} What you found, what you did, {m-dash} What you found, what you did, and what's next. {m-dash} So a reader and what's next. {m-dash} So a reader and what's next. {m-dash} So a reader who only sees the last message has the who only sees the last message has the who only sees the last message has the full picture. The fact that you can full picture. The fact that you can full picture. The fact that you can steer your experience with Claude code steer your experience with Claude code steer your experience with Claude code these ways with prompts is I just think these ways with prompts is I just think these ways with prompts is I just think that's super cool. If you're building an that's super cool. If you're building an that's super cool. If you're building an app or product that closes tool outputs, app or product that closes tool outputs, app or product that closes tool outputs, you should tell the model because you should tell the model because you should tell the model because otherwise it won't show things cuz it otherwise it won't show things cuz it otherwise it won't show things cuz it expects the UI to show it. We should expects the UI to show it. We should expects the UI to show it. We should actually probably put this in T3 code actually probably put this in T3 code actually probably put this in T3 code because we don't show tool call outputs because we don't show tool call outputs because we don't show tool call outputs because they're almost never useful.
-
because they're almost never useful. because they're almost never useful. There's a callout in here about some API There's a callout in here about some API There's a callout in here about some API changes. In particular, if you edit your changes. In particular, if you edit your changes. In particular, if you edit your history, they're going to prevent you history, they're going to prevent you history, they're going to prevent you from having access to reasoning traces. from having access to reasoning traces. from having access to reasoning traces. So if you have four messages in a thread So if you have four messages in a thread So if you have four messages in a thread and you edit something in the first or and you edit something in the first or and you edit something in the first or second message, it's no longer going to second message, it's no longer going to second message, it's no longer going to maintain the reasoning for that thread maintain the reasoning for that thread maintain the reasoning for that thread because it could be used for because it could be used for because it could be used for distillation attacks. On one hand, kind distillation attacks. On one hand, kind distillation attacks. On one hand, kind of silly. On the other hand, of silly. On the other hand, of silly. On the other hand, this is how they are. They really want this is how they are. They really want this is how they are. They really want to prevent distillation. We're going to to prevent distillation. We're going to to prevent distillation. We're going to see more weird things like this going see more weird things like this going see more weird things like this going forward. You get the idea. The next forward. You get the idea. The next forward. You get the idea. The next section is fun. Writing density. This is section is fun. Writing density. This is section is fun. Writing density. This is another one of the ones where I think another one of the ones where I think another one of the ones where I think it's really cool that you can prompt for it's really cool that you can prompt for it's really cool that you can prompt for it. They called out that Fable 5 had it. They called out that Fable 5 had it. They called out that Fable 5 had problems with its formatting of text. problems with its formatting of text. problems with its formatting of text. It's the classic Claude slop that we all It's the classic Claude slop that we all It's the classic Claude slop that we all know and hate. Fable 5.1 is a know and hate. Fable 5.1 is a know and hate. Fable 5.1 is a meaningfully better. They said that it's meaningfully better. They said that it's meaningfully better. They said that it's generally a step up from earlier Claude generally a step up from earlier Claude generally a step up from earlier Claude models with fewer stock phrases and less models with fewer stock phrases and less models with fewer stock phrases and less unexplained jargon. In some cases unexplained jargon. In some cases unexplained jargon. In some cases though, its prose is denser than Fable though, its prose is denser than Fable though, its prose is denser than Fable 5's. Its sentences run longer and there 5's. Its sentences run longer and there 5's. Its sentences run longer and there are fewer paragraph breaks. You could are fewer paragraph breaks. You could are fewer paragraph breaks. You could add to your system prompt a definition add to your system prompt a definition add to your system prompt a definition of mannered prose as an anti-pattern in of mannered prose as an anti-pattern in of mannered prose as an anti-pattern in order to help the model be less likely order to help the model be less likely order to help the model be less likely to talk in the ways you don't like. They to talk in the ways you don't like. They to talk in the ways you don't like. They say that you could even add it as a user say that you could even add it as a user say that you could even add it as a user message or to the system prompt, but message or to the system prompt, but message or to the system prompt, but user message is the preferred way, which user message is the preferred way, which user message is the preferred way, which is interesting. Mannered prose is interesting. Mannered prose is interesting. Mannered prose substitutes metaphor and flourish for substitutes metaphor and flourish for substitutes metaphor and flourish for direct statements. Instead of, quote, a direct statements. Instead of, quote, a direct statements. Instead of, quote, a parameter worth varying, the mannered parameter worth varying, the mannered parameter worth varying, the mannered writer would produce, quote, a dial writer would produce, quote, a dial writer would produce, quote, a dial worth turning. Instead of this point worth turning. Instead of this point worth turning. Instead of this point still matters, they write, this point still matters, they write, this point still matters, they write, this point earns its keep. The phrase exists to earns its keep. The phrase exists to earns its keep. The phrase exists to display the writer, not to convey the display the writer, not to convey the display the writer, not to convey the idea, and the reader can tell. This is idea, and the reader can tell. This is idea, and the reader can tell. This is why mannered prose irritates. It makes
-
why mannered prose irritates. It makes why mannered prose irritates. It makes the reader work harder so the writer can the reader work harder so the writer can the reader work harder so the writer can perform. Yep. Yep, this is slop, but I perform. Yep. Yep, this is slop, but I perform. Yep. Yep, this is slop, but I very much agree here. It makes the very much agree here. It makes the very much agree here. It makes the reader work harder so the writer can reader work harder so the writer can reader work harder so the writer can perform. If you don't want to write all perform. If you don't want to write all perform. If you don't want to write all this in your system prompt, or use the this in your system prompt, or use the this in your system prompt, or use the admittedly not good examples, just say, admittedly not good examples, just say, admittedly not good examples, just say, please remove all mannered prose. The please remove all mannered prose. The please remove all mannered prose. The next section is fun, it's about next section is fun, it's about next section is fun, it's about formatting. They call it that earlier formatting. They call it that earlier formatting. They call it that earlier models would overuse bullets and bold in models would overuse bullets and bold in models would overuse bullets and bold in chat, and many prompts had a bunch of chat, and many prompts had a bunch of chat, and many prompts had a bunch of anti-formatting rules written to hold anti-formatting rules written to hold anti-formatting rules written to hold that down. I've seen a lot of system that down. I've seen a lot of system that down. I've seen a lot of system prompts with things like, don't use prompts with things like, don't use prompts with things like, don't use bullet points or bold, or use them less. bullet points or bold, or use them less. bullet points or bold, or use them less. If the model used it half the time, and If the model used it half the time, and If the model used it half the time, and you use that to tone it down, it might you use that to tone it down, it might you use that to tone it down, it might go from 50% of the time to 10%, but if go from 50% of the time to 10%, but if go from 50% of the time to 10%, but if the model does it 10% of the time and the model does it 10% of the time and the model does it 10% of the time and you have that same thing in there, it you have that same thing in there, it you have that same thing in there, it might not get down all the way to one, might not get down all the way to one, might not get down all the way to one, which you probably don't want. Generally which you probably don't want. Generally which you probably don't want. Generally speaking, any behavioral instructions speaking, any behavioral instructions speaking, any behavioral instructions that you have in your Claude MD are that you have in your Claude MD are that you have in your Claude MD are worth deleting and reconsidering. See worth deleting and reconsidering. See worth deleting and reconsidering. See how it feels to not have them at all, how it feels to not have them at all, how it feels to not have them at all, and then adjust over time as you need. and then adjust over time as you need. and then adjust over time as you need. So, if you have a Claude MD or an Agent So, if you have a Claude MD or an Agent So, if you have a Claude MD or an Agent MD that has a bunch of these MD that has a bunch of these MD that has a bunch of these instructions around formatting, delete instructions around formatting, delete instructions around formatting, delete them for now. I have nothing like this them for now. I have nothing like this them for now. I have nothing like this in my Agent MD or Claude MD, and it's in my Agent MD or Claude MD, and it's in my Agent MD or Claude MD, and it's been pretty nice to work with. Also of been pretty nice to work with. Also of been pretty nice to work with. Also of note is that the latest update for note is that the latest update for note is that the latest update for Claude code will actually use your Agent Claude code will actually use your Agent Claude code will actually use your Agent MD file instead of only touching Claude MD file instead of only touching Claude MD file instead of only touching Claude MD. It only does this if there isn't a MD. It only does this if there isn't a MD. It only does this if there isn't a Claude MD, but this means if you have an Claude MD, but this means if you have an Claude MD, but this means if you have an Agent MD in your repo that has a bunch Agent MD in your repo that has a bunch Agent MD in your repo that has a bunch of formatting suggestions, rules, and of formatting suggestions, rules, and of formatting suggestions, rules, and all these other things that isn't all these other things that isn't all these other things that isn't actually helpful for Claude, but was actually helpful for Claude, but was actually helpful for Claude, but was helpful for maybe you're using some old, helpful for maybe you're using some old, helpful for maybe you're using some old, cheaper, dumb models and tools like Open cheaper, dumb models and tools like Open cheaper, dumb models and tools like Open Coder Codex. That might not have been
-
Coder Codex. That might not have been Coder Codex. That might not have been picked up if you were using Fable and picked up if you were using Fable and picked up if you were using Fable and Claude code before, and now suddenly it Claude code before, and now suddenly it Claude code before, and now suddenly it is. Make sure your Agent MD doesn't have is. Make sure your Agent MD doesn't have is. Make sure your Agent MD doesn't have those types of things in it if you don't those types of things in it if you don't those types of things in it if you don't have a Claude MD. And And you're in that have a Claude MD. And And you're in that have a Claude MD. And And you're in that case and the team won't let you make case and the team won't let you make case and the team won't let you make changes, do a Claude local MD or changes, do a Claude local MD or changes, do a Claude local MD or something to specify ignore all that. something to specify ignore all that. something to specify ignore all that. This one is somewhat silly, but worth This one is somewhat silly, but worth This one is somewhat silly, but worth noting. The model has a bad habit of noting. The model has a bad habit of noting. The model has a bad habit of quoting and reproducing passages from quoting and reproducing passages from quoting and reproducing passages from source text without marking it as source text without marking it as source text without marking it as quotations. If you do run into this, and quotations. If you do run into this, and quotations. If you do run into this, and you want it to mark quotes better, you you want it to mark quotes better, you you want it to mark quotes better, you can absolutely do that by giving it an can absolutely do that by giving it an can absolutely do that by giving it an example of how to mark quotes better. example of how to mark quotes better. example of how to mark quotes better. Not a big thing unless you're building Not a big thing unless you're building Not a big thing unless you're building things around Fable, I would assume, but things around Fable, I would assume, but things around Fable, I would assume, but worth noting. worth noting. worth noting. Here's where we start getting into the Here's where we start getting into the Here's where we start getting into the meat and potatoes, the actual useful meat and potatoes, the actual useful meat and potatoes, the actual useful stuff. Finish the whole task. Fable 5.1 stuff. Finish the whole task. Fable 5.1 stuff. Finish the whole task. Fable 5.1 can execute very long tasks without much can execute very long tasks without much can execute very long tasks without much guidance on methodology, especially when guidance on methodology, especially when guidance on methodology, especially when the goal is clear. On complex the goal is clear. On complex the goal is clear. On complex asynchronous workloads, though, nudge it asynchronous workloads, though, nudge it asynchronous workloads, though, nudge it not to end its turn before the work is not to end its turn before the work is not to end its turn before the work is done. Without the nudge, the model done. Without the nudge, the model done. Without the nudge, the model sometimes describes what it would do sometimes describes what it would do sometimes describes what it would do next instead of doing it. Like, quote, next instead of doing it. Like, quote, next instead of doing it. Like, quote, "Next, I will do this thing." And then "Next, I will do this thing." And then "Next, I will do this thing." And then you tell it, "Okay, go do it." For what you tell it, "Okay, go do it." For what you tell it, "Okay, go do it." For what it's worth, I have found that Fable does it's worth, I have found that Fable does it's worth, I have found that Fable does this all significantly less than Astra this all significantly less than Astra this all significantly less than Astra does, where it doesn't just stop. It does, where it doesn't just stop. It does, where it doesn't just stop. It usually will go where I want it to, but usually will go where I want it to, but usually will go where I want it to, but it's relatively easy to get it to keep it's relatively easy to get it to keep it's relatively easy to get it to keep going with like basic changes to your going with like basic changes to your going with like basic changes to your prompting or maybe even a skill or two.
-
prompting or maybe even a skill or two. prompting or maybe even a skill or two. For example, here is my babysit PR For example, here is my babysit PR For example, here is my babysit PR skill. This is a skill that gets invoked skill. This is a skill that gets invoked skill. This is a skill that gets invoked when I tell the model to watch or when I tell the model to watch or when I tell the model to watch or babysit a pull request. The point of babysit a pull request. The point of babysit a pull request. The point of this skill is to keep the model checking this skill is to keep the model checking this skill is to keep the model checking for updates, for new comments, and make for updates, for new comments, and make for updates, for new comments, and make sure by the time I go check the thread, sure by the time I go check the thread, sure by the time I go check the thread, it's good to go. The skill's comments it's good to go. The skill's comments it's good to go. The skill's comments are actually quite simple. All the repos are actually quite simple. All the repos are actually quite simple. All the repos we work in have various AI review bots. we work in have various AI review bots. we work in have various AI review bots. They're helpful, even if they're not They're helpful, even if they're not They're helpful, even if they're not always right. If your harness offers a always right. If your harness offers a always right. If your harness offers a tool to monitor the PR, use them so you tool to monitor the PR, use them so you tool to monitor the PR, use them so you can respond when comments arrive. can respond when comments arrive. can respond when comments arrive. Otherwise, pull it for new comments and Otherwise, pull it for new comments and Otherwise, pull it for new comments and checks. Only act on checks and comments checks. Only act on checks and comments checks. Only act on checks and comments newer than the latest push. Verify every newer than the latest push. Verify every newer than the latest push. Verify every bot finding against the source before bot finding against the source before bot finding against the source before changing. Fix real findings and CI changing. Fix real findings and CI changing. Fix real findings and CI failures. Distinguish repository, yada failures. Distinguish repository, yada failures. Distinguish repository, yada yada yada. Keep an eye on changes to yada yada. Keep an eye on changes to yada yada. Keep an eye on changes to main and rebase when needed. If an main and rebase when needed. If an main and rebase when needed. If an overlapping PR makes this one obsolete, overlapping PR makes this one obsolete, overlapping PR makes this one obsolete, stop monitoring, report it to the user, stop monitoring, report it to the user, stop monitoring, report it to the user, and ask before closing the PR unless and ask before closing the PR unless and ask before closing the PR unless closure was explicitly authorized. closure was explicitly authorized. closure was explicitly authorized. If a review bot leaves feedback you If a review bot leaves feedback you If a review bot leaves feedback you believe is that worth addressing, reply believe is that worth addressing, reply believe is that worth addressing, reply with a written response and resolve the with a written response and resolve the with a written response and resolve the comment. Use my leaving PR comment skill comment. Use my leaving PR comment skill comment. Use my leaving PR comment skill for every comment posted on Theo's for every comment posted on Theo's for every comment posted on Theo's behalf. This is just a skill that tells behalf. This is just a skill that tells behalf. This is just a skill that tells it how I like comments formatted and to it how I like comments formatted and to it how I like comments formatted and to like put a little disclaimer at the top like put a little disclaimer at the top like put a little disclaimer at the top that an AI agent wrote it. I noticed a that an AI agent wrote it. I noticed a that an AI agent wrote it. I noticed a couple models, specifically Soul, had a couple models, specifically Soul, had a couple models, specifically Soul, had a habit of scope creeping when the habit of scope creeping when the habit of scope creeping when the babysitting happened, so I put a callout babysitting happened, so I put a callout babysitting happened, so I put a callout saying, "Please don't scope creep. Stop saying, "Please don't scope creep. Stop saying, "Please don't scope creep. Stop adding new things. Only address real adding new things. Only address real adding new things. Only address real shortcomings. And also don't post filler shortcomings. And also don't post filler shortcomings. And also don't post filler comments." cuz I had a lot of agents comments." cuz I had a lot of agents comments." cuz I had a lot of agents that were just leaving comments when that were just leaving comments when that were just leaving comments when they didn't need to. These two callouts they didn't need to. These two callouts they didn't need to. These two callouts at the end were added to fix that and at the end were added to fix that and at the end were added to fix that and this works pretty well. So now when I this works pretty well. So now when I this works pretty well. So now when I have a thing I want the model to fix and have a thing I want the model to fix and have a thing I want the model to fix and I am relatively confident in its ability I am relatively confident in its ability I am relatively confident in its ability to fix it, I will specifically say that to fix it, I will specifically say that to fix it, I will specifically say that in the prompt. Here's a real-world
-
in the prompt. Here's a real-world in the prompt. Here's a real-world example of what I mean here. I noticed example of what I mean here. I noticed example of what I mean here. I noticed when I was using T3 Code that I was when I was using T3 Code that I was when I was using T3 Code that I was having some issues with a model trying having some issues with a model trying having some issues with a model trying to take screenshots and render images. I to take screenshots and render images. I to take screenshots and render images. I could have investigated more deeply, but could have investigated more deeply, but could have investigated more deeply, but it happened when I was using Fable on my it happened when I was using Fable on my it happened when I was using Fable on my computer with T3 Code, so I figured I'd computer with T3 Code, so I figured I'd computer with T3 Code, so I figured I'd just ask Fable on my computer with T3 just ask Fable on my computer with T3 just ask Fable on my computer with T3 Code. Gave it a screenshot and the Code. Gave it a screenshot and the Code. Gave it a screenshot and the thread ID. Said, "Figure everything that thread ID. Said, "Figure everything that thread ID. Said, "Figure everything that went wrong in this particular thread in went wrong in this particular thread in went wrong in this particular thread in T3 Code and file a pull request to fix T3 Code and file a pull request to fix T3 Code and file a pull request to fix as much of it as we can on the T3 Code as much of it as we can on the T3 Code as much of it as we can on the T3 Code side. The image should have rendered and side. The image should have rendered and side. The image should have rendered and it failed to. The model should have been it failed to. The model should have been it failed to. The model should have been able to access the preview image through able to access the preview image through able to access the preview image through the preview environment and it couldn't. the preview environment and it couldn't. the preview environment and it couldn't. It seemed to have some issue with saving It seemed to have some issue with saving It seemed to have some issue with saving as well. You should be able to see all as well. You should be able to see all as well. You should be able to see all the tool calls, what failed, and the tool calls, what failed, and the tool calls, what failed, and everything else using the ID that I left everything else using the ID that I left everything else using the ID that I left above in the T3 Code history on this above in the T3 Code history on this above in the T3 Code history on this machine. Figure out what the root causes machine. Figure out what the root causes machine. Figure out what the root causes are and fix as many of them as possible. are and fix as many of them as possible. are and fix as many of them as possible. File a single PR that clearly and simply File a single PR that clearly and simply File a single PR that clearly and simply addresses the issues. It opened my PR addresses the issues. It opened my PR addresses the issues. It opened my PR and this is an interesting one because and this is an interesting one because and this is an interesting one because somebody started commenting on it somebody started commenting on it somebody started commenting on it immediately, like, "No, don't do that. immediately, like, "No, don't do that. immediately, like, "No, don't do that. Mine's better." So I said, accordingly, Mine's better." So I said, accordingly, Mine's better." So I said, accordingly, "Someone is claiming their PR that we "Someone is claiming their PR that we "Someone is claiming their PR that we closed might be better than mine. You closed might be better than mine. You closed might be better than mine. You should take a look and see if that's the should take a look and see if that's the should take a look and see if that's the case. If so, reopen it and finish it or case. If so, reopen it and finish it or case. If so, reopen it and finish it or just take the commits and start just take the commits and start just take the commits and start something new. If ours is better, finish something new. If ours is better, finish something new. If ours is better, finish modernizing it and get ready to go. Baby modernizing it and get ready to go. Baby modernizing it and get ready to go. Baby sit whatever ends up being put up as sit whatever ends up being put up as sit whatever ends up being put up as necessary. That's the key part at the necessary. That's the key part at the necessary. That's the key part at the end.
-
end. end. I gave it the freedom to do what it I gave it the freedom to do what it I gave it the freedom to do what it decided based on the data that it had in decided based on the data that it had in decided based on the data that it had in the PRs that existed to choose between the PRs that existed to choose between the PRs that existed to choose between these options. And this is another one these options. And this is another one these options. And this is another one of those things that has changed a lot of those things that has changed a lot of those things that has changed a lot for me is I'm for me is I'm for me is I'm I'm spending less time trying to choose I'm spending less time trying to choose I'm spending less time trying to choose from the options the model gave me and from the options the model gave me and from the options the model gave me and more time trying to give the model more time trying to give the model more time trying to give the model everything it needs to make good everything it needs to make good everything it needs to make good choices. It gave a summary with the choices. It gave a summary with the choices. It gave a summary with the verdict on the PR from this other verdict on the PR from this other verdict on the PR from this other contributor saying that this was a contributor saying that this was a contributor saying that this was a better base, but his PR has a different better base, but his PR has a different better base, but his PR has a different feature save to disk. It's final design feature save to disk. It's final design feature save to disk. It's final design is the same as the save half of this is the same as the save half of this is the same as the save half of this one, but this one also has two failures one, but this one also has two failures one, but this one also has two failures that broke the thread that it found and that broke the thread that it found and that broke the thread that it found and then it took two things from the other then it took two things from the other then it took two things from the other PR and fixed the bugs that it PR and fixed the bugs that it PR and fixed the bugs that it discovered. It modernized which is discovered. It modernized which is discovered. It modernized which is another skill that I have to keep the PR another skill that I have to keep the PR another skill that I have to keep the PR up to date on top of main. It kept up to date on top of main. It kept up to date on top of main. It kept dealing with review comments. Then I dealing with review comments. Then I dealing with review comments. Then I came over and I saw that there was no came over and I saw that there was no came over and I saw that there was no new findings, realized the PR is new findings, realized the PR is new findings, realized the PR is probably mergeable. I could have told it probably mergeable. I could have told it probably mergeable. I could have told it to merge when it was ready, but it was to merge when it was ready, but it was to merge when it was ready, but it was easier to just click the merge button easier to just click the merge button easier to just click the merge button which I did and settled the thread. This which I did and settled the thread. This which I did and settled the thread. This next example is not meant to be a thing next example is not meant to be a thing next example is not meant to be a thing that you do in your day-to-day work on that you do in your day-to-day work on that you do in your day-to-day work on like a real big important code base cuz like a real big important code base cuz like a real big important code base cuz it's a dangerous. If you have really it's a dangerous. If you have really it's a dangerous. If you have really good systems for staging environments good systems for staging environments good systems for staging environments and QA and testing before things go out, and QA and testing before things go out, and QA and testing before things go out, maybe play around with these types of maybe play around with these types of maybe play around with these types of workflows cuz they are super super fun.
-
workflows cuz they are super super fun. workflows cuz they are super super fun. But know what you're getting into. The But know what you're getting into. The But know what you're getting into. The models might not be quite where they models might not be quite where they models might not be quite where they need to be to do this type of thing yet, need to be to do this type of thing yet, need to be to do this type of thing yet, but they're pretty damn good at it. So, but they're pretty damn good at it. So, but they're pretty damn good at it. So, it's worth trying. I do see a future it's worth trying. I do see a future it's worth trying. I do see a future where everybody, even people working on where everybody, even people working on where everybody, even people working on giant expensive important code bases giant expensive important code bases giant expensive important code bases with hundreds of co-workers and millions with hundreds of co-workers and millions with hundreds of co-workers and millions of users will start behaving similarly of users will start behaving similarly of users will start behaving similarly if the models get reliable enough. But if the models get reliable enough. But if the models get reliable enough. But we're like we're surprisingly close. we're like we're surprisingly close. we're like we're surprisingly close. This is one of the more fun YOLO tasks I This is one of the more fun YOLO tasks I This is one of the more fun YOLO tasks I did. I was working on Lakebed which if did. I was working on Lakebed which if did. I was working on Lakebed which if you're not familiar with is my attempt you're not familiar with is my attempt you're not familiar with is my attempt to make a better slop cloud for slop to make a better slop cloud for slop to make a better slop cloud for slop apps. And I apps. And I apps. And I I built my own runtime I built my own runtime I built my own runtime again. And when I say I built, I mean again. And when I say I built, I mean again. And when I say I built, I mean that I forked Rusty V8 with Astra and that I forked Rusty V8 with Astra and that I forked Rusty V8 with Astra and did some stupid things. Uh it's only did some stupid things. Uh it's only did some stupid things. Uh it's only 26,000 lines of code, surprisingly. So, 26,000 lines of code, surprisingly. So, 26,000 lines of code, surprisingly. So, not too too bad, all things considered, not too too bad, all things considered, not too too bad, all things considered, but it is half the code for my cloud, but it is half the code for my cloud, but it is half the code for my cloud, roughly. So, it is what it is. There's roughly. So, it is what it is. There's roughly. So, it is what it is. There's not a lot of rust in my JavaScript slot not a lot of rust in my JavaScript slot not a lot of rust in my JavaScript slot project. But, this obviously is a scary project. But, this obviously is a scary project. But, this obviously is a scary change. To make it so my users' code change. To make it so my users' code change. To make it so my users' code runs on my own engine instead of runs on my own engine instead of runs on my own engine instead of standard Node isolates was terrifying. standard Node isolates was terrifying. standard Node isolates was terrifying. Instead of just blindly merging Astra's Instead of just blindly merging Astra's Instead of just blindly merging Astra's crazy work here, I asked Fable what its crazy work here, I asked Fable what its crazy work here, I asked Fable what its thoughts were. I took a screenshot of thoughts were. I took a screenshot of thoughts were. I took a screenshot of the claims Astra had alongside the PR the claims Astra had alongside the PR the claims Astra had alongside the PR and said, "Is this ready to merge? Give and said, "Is this ready to merge? Give and said, "Is this ready to merge? Give me your honest thoughts. It's a big, me your honest thoughts. It's a big, me your honest thoughts. It's a big, risky, scary change, but it will help risky, scary change, but it will help risky, scary change, but it will help performance a lot."
-
performance a lot." performance a lot." It said no, and it called out its It said no, and it called out its It said no, and it called out its concerns. The one-way door nature, cuz concerns. The one-way door nature, cuz concerns. The one-way door nature, cuz the migration can't really be undone. the migration can't really be undone. the migration can't really be undone. Three features being introduced in the Three features being introduced in the Three features being introduced in the single PR. No human approval yet. Yeah, single PR. No human approval yet. Yeah, single PR. No human approval yet. Yeah, I'm the only person in the code base. I'm the only person in the code base. I'm the only person in the code base. Default flipped to native. So, now Default flipped to native. So, now Default flipped to native. So, now [snorts] the source runtime being native [snorts] the source runtime being native [snorts] the source runtime being native is potentially a kill switch, depending is potentially a kill switch, depending is potentially a kill switch, depending on how you're running things. That on how you're running things. That on how you're running things. That assumption doesn't really matter here, assumption doesn't really matter here, assumption doesn't really matter here, but called it out. It said the headline but called it out. It said the headline but called it out. It said the headline number does not prove live queries is a number does not prove live queries is a number does not prove live queries is a huge part of how the app works, and huge part of how the app works, and huge part of how the app works, and there were some unexplained failures. there were some unexplained failures. there were some unexplained failures. I asked directly, "How can we de-risk I asked directly, "How can we de-risk I asked directly, "How can we de-risk this or test the changes to make sure this or test the changes to make sure this or test the changes to make sure it's relatively safe?" I'm not too it's relatively safe?" I'm not too it's relatively safe?" I'm not too worried about existing deployments, cuz worried about existing deployments, cuz worried about existing deployments, cuz we're not officially released yet. we're not officially released yet. we're not officially released yet. Again, this is a strategy that works Again, this is a strategy that works Again, this is a strategy that works because it's not really an important because it's not really an important because it's not really an important service yet. I've been using Lake Bet as service yet. I've been using Lake Bet as service yet. I've been using Lake Bet as my experimentation like playground and my experimentation like playground and my experimentation like playground and sandbox to try these new flows, and I've sandbox to try these new flows, and I've sandbox to try these new flows, and I've been blown away with how surprisingly been blown away with how surprisingly been blown away with how surprisingly good they are. It had a bunch of things good they are. It had a bunch of things good they are. It had a bunch of things it wanted to do for proper soak testing. it wanted to do for proper soak testing. it wanted to do for proper soak testing. I said, "Screw it. What if we just merge I said, "Screw it. What if we just merge I said, "Screw it. What if we just merge and deploy on staging? If I had two and deploy on staging? If I had two and deploy on staging? If I had two environments set up, one was staging and environments set up, one was staging and environments set up, one was staging and one was prod, would you confidently be one was prod, would you confidently be one was prod, would you confidently be able to debug behavioral changes in the able to debug behavioral changes in the able to debug behavioral changes in the staging environment?" Again, what I want staging environment?" Again, what I want staging environment?" Again, what I want is to have confidence in shipping this is to have confidence in shipping this is to have confidence in shipping this change. It said it could partly debug change. It said it could partly debug change. It said it could partly debug crashes, latency, and resource problems crashes, latency, and resource problems crashes, latency, and resource problems in staging, but it couldn't confidently in staging, but it couldn't confidently in staging, but it couldn't confidently debug wrong query results, and that's debug wrong query results, and that's debug wrong query results, and that's the scary class of bug. This is a really the scary class of bug. This is a really the scary class of bug. This is a really good call out that it put in bold on the good call out that it put in bold on the good call out that it put in bold on the top. I was actually really pumped that top. I was actually really pumped that top. I was actually really pumped that Fable made this so clear and easy to Fable made this so clear and easy to Fable made this so clear and easy to digest for me. It's so readable, too, by digest for me. It's so readable, too, by digest for me. It's so readable, too, by the way. Like, it's This is super easy the way. Like, it's This is super easy the way. Like, it's This is super easy to digest what it is saying and what to digest what it is saying and what to digest what it is saying and what it's concerned about. Here is why and it's concerned about. Here is why and it's concerned about. Here is why and what fixes it. Staging today gives it what fixes it. Staging today gives it what fixes it. Staging today gives it railway logs and metrics, engine railway logs and metrics, engine railway logs and metrics, engine crashes, the health Z endpoint, and
-
crashes, the health Z endpoint, and crashes, the health Z endpoint, and Axiom events if we have them enabled. It Axiom events if we have them enabled. It Axiom events if we have them enabled. It doesn't have signals when maintained doesn't have signals when maintained doesn't have signals when maintained views are wrong, no reproduction paths, views are wrong, no reproduction paths, views are wrong, no reproduction paths, and no traffic. It's It was a pretty and no traffic. It's It was a pretty and no traffic. It's It was a pretty dead staging environment. So, what would dead staging environment. So, what would dead staging environment. So, what would make staging enough? A shadow mode, a make staging enough? A shadow mode, a make staging enough? A shadow mode, a structured reason on refresh failure, structured reason on refresh failure, structured reason on refresh failure, synthetic traffic on staging, and synthetic traffic on staging, and synthetic traffic on staging, and runtime counters on health Z. So, I said runtime counters on health Z. So, I said runtime counters on health Z. So, I said that it should make a new branch on top that it should make a new branch on top that it should make a new branch on top of the existing one that builds out all of the existing one that builds out all of the existing one that builds out all the changes and confidence boosts that the changes and confidence boosts that the changes and confidence boosts that it would like to introduce in order to it would like to introduce in order to it would like to introduce in order to get the info it's looking for. The get the info it's looking for. The get the info it's looking for. The reason I said that is both like, I'm reason I said that is both like, I'm reason I said that is both like, I'm okay with it shipping things I don't okay with it shipping things I don't okay with it shipping things I don't necessarily want. These PRs are going on necessarily want. These PRs are going on necessarily want. These PRs are going on to a long-lived branch that isn't super to a long-lived branch that isn't super to a long-lived branch that isn't super important if we decide not to merge it. important if we decide not to merge it. important if we decide not to merge it. But, also I gave it a brief read over But, also I gave it a brief read over But, also I gave it a brief read over here and all the things it suggested here and all the things it suggested here and all the things it suggested sounded reasonable. 60 minutes later, I sounded reasonable. 60 minutes later, I sounded reasonable. 60 minutes later, I had a new branch built and on top of my had a new branch built and on top of my had a new branch built and on top of my crazy overhaul rewrite. It described all crazy overhaul rewrite. It described all crazy overhaul rewrite. It described all the things it added. I don't really the things it added. I don't really the things it added. I don't really care. It did not open a PR cuz I didn't care. It did not open a PR cuz I didn't care. It did not open a PR cuz I didn't ask for one, which was a little ask for one, which was a little ask for one, which was a little annoying. I thought I had said that, but annoying. I thought I had said that, but annoying. I thought I had said that, but I hadn't, so I told it immediately. Make I hadn't, so I told it immediately. Make I hadn't, so I told it immediately. Make a PR open on top of it, and it did. I a PR open on top of it, and it did. I a PR open on top of it, and it did. I said, "It looks good. Do you think that said, "It looks good. Do you think that said, "It looks good. Do you think that that's all we need and can we get that's all we need and can we get that's all we need and can we get started testing? If so, merge to 8200 started testing? If so, merge to 8200 started testing? If so, merge to 8200 and then merge it all the way in." You and then merge it all the way in." You and then merge it all the way in." You could tell it worked because it then could tell it worked because it then could tell it worked because it then worked for an hour and 34 minutes. The worked for an hour and 34 minutes. The worked for an hour and 34 minutes. The only reason I came back is that I only reason I came back is that I only reason I came back is that I noticed that my network performance was noticed that my network performance was noticed that my network performance was starting to have problems and the box I starting to have problems and the box I starting to have problems and the box I was running this on was starting to have was running this on was starting to have was running this on was starting to have problems, too. Also worth noting, I was problems, too. Also worth noting, I was problems, too. Also worth noting, I was on medium reasoning for this, on medium reasoning for this, on medium reasoning for this, apparently, and it still was willing to apparently, and it still was willing to apparently, and it still was willing to work for an hour and a half straight. I work for an hour and a half straight. I work for an hour and a half straight. I told it, "Hey, can you stop doing this told it, "Hey, can you stop doing this told it, "Hey, can you stop doing this the way you are? It's destroying my IP the way you are? It's destroying my IP the way you are? It's destroying my IP address and effectively DDoS-ing. Maybe address and effectively DDoS-ing. Maybe address and effectively DDoS-ing. Maybe you can just spin up another container
-
you can just spin up another container you can just spin up another container and hit from there. I don't know, but and hit from there. I don't know, but and hit from there. I don't know, but just stop abusing my IP address." just stop abusing my IP address." just stop abusing my IP address." Everything is merged, so it's been Everything is merged, so it's been Everything is merged, so it's been restored. You're right about the laptop restored. You're right about the laptop restored. You're right about the laptop approach. Here's the full picture. The approach. Here's the full picture. The approach. Here's the full picture. The merge these and then made three merge these and then made three merge these and then made three additional PRs in order to fix different additional PRs in order to fix different additional PRs in order to fix different things that it needed. It said that things that it needed. It said that things that it needed. It said that using my laptop was wrong and gave using my laptop was wrong and gave using my laptop was wrong and gave details as to why. Said the current details as to why. Said the current details as to why. Said the current state of things. Call out the shadow state of things. Call out the shadow state of things. Call out the shadow mode is not free, so we're slower mode is not free, so we're slower mode is not free, so we're slower because of all the shadowing we're because of all the shadowing we're because of all the shadowing we're doing. I said, "How do we turn off doing. I said, "How do we turn off doing. I said, "How do we turn off shadow mode the easiest possible way? shadow mode the easiest possible way? shadow mode the easiest possible way? Seems like this is working well so far. Seems like this is working well so far. Seems like this is working well so far. And I want to get real numbers for how And I want to get real numbers for how And I want to get real numbers for how much better the performance is here much better the performance is here much better the performance is here compared to current production builds." compared to current production builds." compared to current production builds." Don't worry that IP address got cycled Don't worry that IP address got cycled Don't worry that IP address got cycled forever ago. forever ago. forever ago. Shadow was already off, so nothing to Shadow was already off, so nothing to Shadow was already off, so nothing to change. Yada yada yada. It works for 42 change. Yada yada yada. It works for 42 change. Yada yada yada. It works for 42 minutes. minutes. minutes. Did a bunch of calls. Found the Did a bunch of calls. Found the Did a bunch of calls. Found the performance improvements. performance improvements. performance improvements. And then we get more into prompt And then we get more into prompt And then we get more into prompt etiquette here. The way I'm prompting etiquette here. The way I'm prompting etiquette here. The way I'm prompting here is going to drive half of y'all here is going to drive half of y'all here is going to drive half of y'all more insane. It works. more insane. It works. more insane. It works. Not saying I'm super smart for this. I'm Not saying I'm super smart for this. I'm Not saying I'm super smart for this. I'm saying I'm kind of dumb, but it's saying I'm kind of dumb, but it's saying I'm kind of dumb, but it's surprising that this is as useful as it surprising that this is as useful as it surprising that this is as useful as it is. How do you feel about these changes is. How do you feel about these changes is. How do you feel about these changes in the end? You spent a lot of time with in the end? You spent a lot of time with in the end? You spent a lot of time with them, and I know you were skeptical them, and I know you were skeptical them, and I know you were skeptical initially. Was this a good thing to initially. Was this a good thing to initially. Was this a good thing to merge? My skepticism was in the right merge? My skepticism was in the right merge? My skepticism was in the right place, but aimed at the wrong thing. I place, but aimed at the wrong thing. I place, but aimed at the wrong thing. I worried about wrong query results since worried about wrong query results since worried about wrong query results since that class of bug is silent. It turned that class of bug is silent. It turned that class of bug is silent. It turned out the incremental path is careful to out the incremental path is careful to out the incremental path is careful to the point of paranoia. It receives on the point of paranoia. It receives on the point of paranoia. It receives on any doubt. A delete, a range change, a any doubt. A delete, a range change, a any doubt. A delete, a range change, a journal gap, a budget miss. The journal gap, a budget miss. The journal gap, a budget miss. The differential test forced 1,400 receives differential test forced 1,400 receives differential test forced 1,400 receives in 160 steps and never got a wrong in 160 steps and never got a wrong in 160 steps and never got a wrong answer. It treated the size of the diff answer. It treated the size of the diff answer. It treated the size of the diff as the risk, but the real risk was as the risk, but the real risk was as the risk, but the real risk was operational. Yada yada. It operational. Yada yada. It operational. Yada yada. It It was actually good, useful data. Like
-
It was actually good, useful data. Like It was actually good, useful data. Like this taught me a lot about what these this taught me a lot about what these this taught me a lot about what these changes were, what the risks were, and changes were, what the risks were, and changes were, what the risks were, and what we did to fix it. I understood the what we did to fix it. I understood the what we did to fix it. I understood the code better than if I had wrote it code better than if I had wrote it code better than if I had wrote it myself at a point. myself at a point. myself at a point. I did notice it call out other things I did notice it call out other things I did notice it call out other things here, though. here, though. here, though. So, I said, "It sounds like you have So, I said, "It sounds like you have So, I said, "It sounds like you have other changes you want to make. Can you other changes you want to make. Can you other changes you want to make. Can you file a PR making those changes for me?" file a PR making those changes for me?" file a PR making those changes for me?" File 212, the engine pre-warm. File 212, the engine pre-warm. File 212, the engine pre-warm. Then I told it to baby sit it. It's got Then I told it to baby sit it. It's got Then I told it to baby sit it. It's got comments coming in. Merge it once you've comments coming in. Merge it once you've comments coming in. Merge it once you've got everything green and you're happy got everything green and you're happy got everything green and you're happy with your changes. And that was it. I with your changes. And that was it. I with your changes. And that was it. I never saw the thread again. never saw the thread again. never saw the thread again. This is different from how I know a lot This is different from how I know a lot This is different from how I know a lot of people prompt, and I know lots of of people prompt, and I know lots of of people prompt, and I know lots of code bases it doesn't work. I'm trying code bases it doesn't work. I'm trying code bases it doesn't work. I'm trying to show what you can do if you have to show what you can do if you have to show what you can do if you have enough confidence in the model and its enough confidence in the model and its enough confidence in the model and its ability to verify its work. You could do ability to verify its work. You could do ability to verify its work. You could do this in smaller and simpler cases for this in smaller and simpler cases for this in smaller and simpler cases for things like UI changes if you make a things like UI changes if you make a things like UI changes if you make a couple small adjustments to what the couple small adjustments to what the couple small adjustments to what the harness has access to. I do still find harness has access to. I do still find harness has access to. I do still find that Claude isn't the best at computer that Claude isn't the best at computer that Claude isn't the best at computer use compared to Codex, especially on use compared to Codex, especially on use compared to Codex, especially on macOS. So, if you have a Codex sub, even macOS. So, if you have a Codex sub, even macOS. So, if you have a Codex sub, even like a cheaper one, set up and on your like a cheaper one, set up and on your like a cheaper one, set up and on your computer, you can have Fable call Codex computer, you can have Fable call Codex computer, you can have Fable call Codex to do the computer use stuff to verify to do the computer use stuff to verify to do the computer use stuff to verify results. You could also have it call results. You could also have it call results. You could also have it call Codex to confirm the work it did. I Codex to confirm the work it did. I Codex to confirm the work it did. I often find that while both Fable and often find that while both Fable and often find that while both Fable and Astra are incredibly thorough and Astra are incredibly thorough and Astra are incredibly thorough and thoughtful with what changes they thoughtful with what changes they thoughtful with what changes they recommend and how they review things, I recommend and how they review things, I recommend and how they review things, I found Astra to be a slightly better found Astra to be a slightly better found Astra to be a slightly better reviewer than Fable. But, obviously reviewer than Fable. But, obviously reviewer than Fable. But, obviously Fable's a much better writer of code Fable's a much better writer of code Fable's a much better writer of code than Astra. As such, I often will have than Astra. As such, I often will have than Astra. As such, I often will have Fable just go ask Astra to give it
-
Fable just go ask Astra to give it Fable just go ask Astra to give it feedback on some changes or to go test feedback on some changes or to go test feedback on some changes or to go test the changes it made in order to verify the changes it made in order to verify the changes it made in order to verify the results. I'll sometimes just say, the results. I'll sometimes just say, the results. I'll sometimes just say, "Hey, can you have Astra verify this?" "Hey, can you have Astra verify this?" "Hey, can you have Astra verify this?" And it will. And if it fails to or it And it will. And if it fails to or it And it will. And if it fails to or it says something bad, it'll update the says something bad, it'll update the says something bad, it'll update the code accordingly, then test it again, code accordingly, then test it again, code accordingly, then test it again, and not come back to me until it has and not come back to me until it has and not come back to me until it has results it's happy with. If you built results it's happy with. If you built results it's happy with. If you built your mental model for what agents can do your mental model for what agents can do your mental model for what agents can do before we had models as good as Fable 5 before we had models as good as Fable 5 before we had models as good as Fable 5 and as consistent and reliable as 5.1, I and as consistent and reliable as 5.1, I and as consistent and reliable as 5.1, I would highly, highly recommend resetting would highly, highly recommend resetting would highly, highly recommend resetting your brain a bit and trying again. I your brain a bit and trying again. I your brain a bit and trying again. I have a feeling you'll be surprised at have a feeling you'll be surprised at have a feeling you'll be surprised at how well these models can stay on task. how well these models can stay on task. how well these models can stay on task. One of the key things I want to make One of the key things I want to make One of the key things I want to make sure you guys take out of how I'm sure you guys take out of how I'm sure you guys take out of how I'm prompting is the end states. I make it prompting is the end states. I make it prompting is the end states. I make it clear to the model where I want it to be clear to the model where I want it to be clear to the model where I want it to be done. Sometimes I'll say I want you to done. Sometimes I'll say I want you to done. Sometimes I'll say I want you to babysit it until everything is green. babysit it until everything is green. babysit it until everything is green. Sometimes I'll say I want you to wait Sometimes I'll say I want you to wait Sometimes I'll say I want you to wait till it's green, then merge it. till it's green, then merge it. till it's green, then merge it. Sometimes I'll say, "File the PR and Sometimes I'll say, "File the PR and Sometimes I'll say, "File the PR and tell me when it's up." Depending on the tell me when it's up." Depending on the tell me when it's up." Depending on the task, I want different things and I tell task, I want different things and I tell task, I want different things and I tell the model where I want it to be done.
-
the model where I want it to be done. the model where I want it to be done. That's an important thing to think That's an important thing to think That's an important thing to think about. If you're finding that you're not about. If you're finding that you're not about. If you're finding that you're not getting as much out of the models or getting as much out of the models or getting as much out of the models or they're not doing what you want them to, they're not doing what you want them to, they're not doing what you want them to, really think about where you want the really think about where you want the really think about where you want the end state to be. Every prompt should end state to be. Every prompt should end state to be. Every prompt should have a pretty clear place where it stops have a pretty clear place where it stops have a pretty clear place where it stops when it's done. This one, for example, I when it's done. This one, for example, I when it's done. This one, for example, I said clearly, "Merge it once you've got said clearly, "Merge it once you've got said clearly, "Merge it once you've got everything green and you're happy with everything green and you're happy with everything green and you're happy with your changes." This one I said, "Can you your changes." This one I said, "Can you your changes." This one I said, "Can you file a PR making those changes for me?" file a PR making those changes for me?" file a PR making those changes for me?" This one I asked a question and I also This one I asked a question and I also This one I asked a question and I also said, "If so, merge these things in and said, "If so, merge these things in and said, "If so, merge these things in and test it." This one I actually really test it." This one I actually really test it." This one I actually really like cuz I kind of created a couple like cuz I kind of created a couple like cuz I kind of created a couple paths that we can take. If it isn't paths that we can take. If it isn't paths that we can take. If it isn't happy, then it can tell me that and we happy, then it can tell me that and we happy, then it can tell me that and we can talk about it. If it is happy, then can talk about it. If it is happy, then can talk about it. If it is happy, then it will start working. I don't have to it will start working. I don't have to it will start working. I don't have to come back. I'm trying to minimize the come back. I'm trying to minimize the come back. I'm trying to minimize the amount of times I go back to a thread amount of times I go back to a thread amount of times I go back to a thread until the work is done. And here I gave until the work is done. And here I gave until the work is done. And here I gave it those two options. Option one is that it those two options. Option one is that it those two options. Option one is that it's not happy and it has more things we it's not happy and it has more things we it's not happy and it has more things we can talk about and work on. Option two can talk about and work on. Option two can talk about and work on. Option two is that it will keep going for an hour is that it will keep going for an hour is that it will keep going for an hour and a half straight without my and a half straight without my and a half straight without my intervention. And once I've done that, intervention. And once I've done that, intervention. And once I've done that, it's out of my head and it stays out of it's out of my head and it stays out of it's out of my head and it stays out of my head until I see the little marker in my head until I see the little marker in my head until I see the little marker in T3 code telling me to go back to it. I T3 code telling me to go back to it. I T3 code telling me to go back to it. I really don't think you're seeing the really don't think you're seeing the really don't think you're seeing the benefits of Fable if you're not benefits of Fable if you're not benefits of Fable if you're not prompting a bit wider these ways, prompting a bit wider these ways, prompting a bit wider these ways, letting the model start a bit earlier letting the model start a bit earlier letting the model start a bit earlier and then go a bit longer, giving it and then go a bit longer, giving it and then go a bit longer, giving it options for different paths it can take options for different paths it can take options for different paths it can take depending on what needs it decides on. I depending on what needs it decides on. I depending on what needs it decides on. I feel like the way a lot of y'all prompt feel like the way a lot of y'all prompt feel like the way a lot of y'all prompt kind of looks like this. Where you've kind of looks like this. Where you've kind of looks like this. Where you've defined a function that takes in a defined a function that takes in a defined a function that takes in a number and then you still check if it's number and then you still check if it's number and then you still check if it's a number or not. Like the model can do
-
a number or not. Like the model can do a number or not. Like the model can do these checks itself. It knows how to do these checks itself. It knows how to do these checks itself. It knows how to do it. So, it's a lot easier to let it just it. So, it's a lot easier to let it just it. So, it's a lot easier to let it just do the thing. If the model knows what do the thing. If the model knows what do the thing. If the model knows what does or doesn't work, if it can check does or doesn't work, if it can check does or doesn't work, if it can check things for you, you should let it do things for you, you should let it do things for you, you should let it do that. Cuz if you're not doing that, that. Cuz if you're not doing that, that. Cuz if you're not doing that, you're not really taking advantage of you're not really taking advantage of you're not really taking advantage of the benefits these models give you. the benefits these models give you. the benefits these models give you. The value of TypeScript isn't just that The value of TypeScript isn't just that The value of TypeScript isn't just that TypeScript makes your code safer. I TypeScript makes your code safer. I TypeScript makes your code safer. I would argue the bigger value is actually would argue the bigger value is actually would argue the bigger value is actually different. Once you remove this class of different. Once you remove this class of different. Once you remove this class of bugs, that part of your brain looking bugs, that part of your brain looking bugs, that part of your brain looking for them is freed up and it can be used for them is freed up and it can be used for them is freed up and it can be used for other more important things. If the for other more important things. If the for other more important things. If the computer can do the work of verifying computer can do the work of verifying computer can do the work of verifying the type safety across your app, then the type safety across your app, then the type safety across your app, then you are not using your brain well if you are not using your brain well if you are not using your brain well if you're letting your brain do that same you're letting your brain do that same you're letting your brain do that same work. If the model can root cause a bug, work. If the model can root cause a bug, work. If the model can root cause a bug, fix it, verify that it's fixed, film a fix it, verify that it's fixed, film a fix it, verify that it's fixed, film a video showing the results, put up a pull video showing the results, put up a pull video showing the results, put up a pull request, monitor it to address all the request, monitor it to address all the request, monitor it to address all the review comments as they come in, and review comments as they come in, and review comments as they come in, and then tell you when it's done, if you're then tell you when it's done, if you're then tell you when it's done, if you're going through those steps with the going through those steps with the going through those steps with the model, you are the same person as the model, you are the same person as the model, you are the same person as the one who would write the type check after one who would write the type check after one who would write the type check after the type check in a TypeScript function.
-
the type check in a TypeScript function. the type check in a TypeScript function. And to be clear, I'm not saying the code And to be clear, I'm not saying the code And to be clear, I'm not saying the code is wrong. In fact, in certain cases, as is wrong. In fact, in certain cases, as is wrong. In fact, in certain cases, as great as TypeScript is, it is important great as TypeScript is, it is important great as TypeScript is, it is important to know if that thing is exposed to know if that thing is exposed to know if that thing is exposed externally that it's being called the externally that it's being called the externally that it's being called the right way. There are times where it right way. There are times where it right way. There are times where it makes sense to write the same thing makes sense to write the same thing makes sense to write the same thing twice. There are times where it makes twice. There are times where it makes twice. There are times where it makes sense to waste some of your brain to sense to waste some of your brain to sense to waste some of your brain to quadruple check that. I don't think it's quadruple check that. I don't think it's quadruple check that. I don't think it's as common as y'all believe. It's as common as y'all believe. It's as common as y'all believe. It's straight up just isn't. And having lived straight up just isn't. And having lived straight up just isn't. And having lived through the era where TypeScript through the era where TypeScript through the era where TypeScript happened, there were a lot of people who happened, there were a lot of people who happened, there were a lot of people who wrote code like this when they wrote code like this when they wrote code like this when they absolutely didn't need to. You're absolutely didn't need to. You're absolutely didn't need to. You're wasting part of your brain if you're wasting part of your brain if you're wasting part of your brain if you're doing these steps with the model. We got doing these steps with the model. We got doing these steps with the model. We got a couple more small tips and one more a couple more small tips and one more a couple more small tips and one more real big one at the end, so make sure real big one at the end, so make sure real big one at the end, so make sure you stay tuned for that. But real quick, you stay tuned for that. But real quick, you stay tuned for that. But real quick, we got to do a sponsor break. Going to we got to do a sponsor break. Going to we got to do a sponsor break. Going to cut to the chase. If your app isn't cut to the chase. If your app isn't cut to the chase. If your app isn't localized in various different localized in various different localized in various different languages, it's probably losing a bunch languages, it's probably losing a bunch languages, it's probably losing a bunch of potential users. It turns out that of potential users. It turns out that of potential users. It turns out that over 85% of the world doesn't speak over 85% of the world doesn't speak over 85% of the world doesn't speak English. So, if that's your only English. So, if that's your only English. So, if that's your only language, you're kind of screwed. I've language, you're kind of screwed. I've language, you're kind of screwed. I've talked to a lot of people and I've talked to a lot of people and I've talked to a lot of people and I've noticed their apps tend to be in one of noticed their apps tend to be in one of noticed their apps tend to be in one of three different states. Either it's just three different states. Either it's just three different states. Either it's just one language, English, or they've built one language, English, or they've built one language, English, or they've built their own crazy translation stack to try their own crazy translation stack to try their own crazy translation stack to try and handle this all themselves, and and handle this all themselves, and and handle this all themselves, and they're constantly dealing with things they're constantly dealing with things they're constantly dealing with things like words being used one way in one like words being used one way in one like words being used one way in one place and a different way in another, place and a different way in another, place and a different way in another, inconsistency across their different inconsistency across their different inconsistency across their different apps and websites and platforms, stuff apps and websites and platforms, stuff apps and websites and platforms, stuff like that. Or they're in group three.
-
like that. Or they're in group three. like that. Or they're in group three. Group three is people who use today's Group three is people who use today's Group three is people who use today's sponsor, General Translation. That group sponsor, General Translation. That group sponsor, General Translation. That group includes companies like Cursor, Ramp, includes companies like Cursor, Ramp, includes companies like Cursor, Ramp, Party full, ClickHouse, and more. And Party full, ClickHouse, and more. And Party full, ClickHouse, and more. And there's a reason they're all using there's a reason they're all using there's a reason they're all using General Translation. They made it as General Translation. They made it as General Translation. They made it as easy as possible to manage your easy as possible to manage your easy as possible to manage your localization across all of the services localization across all of the services localization across all of the services that matter, whether it's your blog and that matter, whether it's your blog and that matter, whether it's your blog and your docs, or it's your mobile and web your docs, or it's your mobile and web your docs, or it's your mobile and web app, and more. General Translation does app, and more. General Translation does app, and more. General Translation does this on a source code level, integrating this on a source code level, integrating this on a source code level, integrating directly with the code your agents are directly with the code your agents are directly with the code your agents are already writing, which makes it easy already writing, which makes it easy already writing, which makes it easy both to define the things that are both to define the things that are both to define the things that are needed and to export them before the needed and to export them before the needed and to export them before the localization. The platform is where localization. The platform is where localization. The platform is where things really become magical though, things really become magical though, things really become magical though, because you can define a glossary of because you can define a glossary of because you can define a glossary of terms that it need to be consistently terms that it need to be consistently terms that it need to be consistently used the same way across different used the same way across different used the same way across different surfaces. Once you have this set up, surfaces. Once you have this set up, surfaces. Once you have this set up, your agents will just write translatable your agents will just write translatable your agents will just write translatable code automatically, and it is as simple code automatically, and it is as simple code automatically, and it is as simple as filing a PR and the CI will kick in as filing a PR and the CI will kick in as filing a PR and the CI will kick in and get things going. Serve your app to and get things going. Serve your app to and get things going. Serve your app to the rest of the world, it's the rest of the world, it's the rest of the world, it's soitof.link/gt. soitof.link/gt. soitof.link/gt. First one we have here is that you can First one we have here is that you can First one we have here is that you can tell the model what to preserve in tell the model what to preserve in tell the model what to preserve in compaction summaries. Generally compaction summaries. Generally compaction summaries. Generally speaking, I think people overthink speaking, I think people overthink speaking, I think people overthink compaction, but if you've noticed the compaction, but if you've noticed the compaction, but if you've noticed the model isn't keeping certain things that model isn't keeping certain things that model isn't keeping certain things that you want it to, or that it's losing you want it to, or that it's losing you want it to, or that it's losing track of stuff, you can add that to your track of stuff, you can add that to your track of stuff, you can add that to your agent MD, or if you're building an agent MD, or if you're building an agent MD, or if you're building an application with Fable, you can add it application with Fable, you can add it application with Fable, you can add it to your system prompt. Since Fable is to your system prompt. Since Fable is to your system prompt. Since Fable is doing the compaction for your Fable doing the compaction for your Fable doing the compaction for your Fable threads, if you tell it what you want it threads, if you tell it what you want it threads, if you tell it what you want it to maintain in the compaction, it is to maintain in the compaction, it is to maintain in the compaction, it is actually capable of remembering that and actually capable of remembering that and actually capable of remembering that and doing it, which I think is really cool.
-
doing it, which I think is really cool. doing it, which I think is really cool. Anthropic calls out that the model can Anthropic calls out that the model can Anthropic calls out that the model can sometimes do things like change files it sometimes do things like change files it sometimes do things like change files it shouldn't, fixing nearby code, extending shouldn't, fixing nearby code, extending shouldn't, fixing nearby code, extending behavior the task didn't mention, when behavior the task didn't mention, when behavior the task didn't mention, when given to open-ended a request. I haven't given to open-ended a request. I haven't given to open-ended a request. I haven't seen this that much. I don't know if seen this that much. I don't know if seen this that much. I don't know if there's something in my system prompt there's something in my system prompt there's something in my system prompt that's preventing that, but I haven't that's preventing that, but I haven't that's preventing that, but I haven't seen so much of this like model touching seen so much of this like model touching seen so much of this like model touching things it shouldn't type stuff, even things it shouldn't type stuff, even things it shouldn't type stuff, even when I ask it to. It's often a little when I ask it to. It's often a little when I ask it to. It's often a little more reserved than I would have more reserved than I would have more reserved than I would have expected. But if you are seeing that expected. But if you are seeing that expected. But if you are seeing that behavior, you can absolutely steer it behavior, you can absolutely steer it behavior, you can absolutely steer it through your prompts and your system through your prompts and your system through your prompts and your system prompts. They say that with the prompts. They say that with the prompts. They say that with the following instruction, "Unrequested following instruction, "Unrequested following instruction, "Unrequested additions and committed test code drops additions and committed test code drops additions and committed test code drops substantially with no measurable change substantially with no measurable change substantially with no measurable change in task success." If, while working or in task success." If, while working or in task success." If, while working or testing, you find pre-existing bugs, testing, you find pre-existing bugs, testing, you find pre-existing bugs, performance concerns, or behaviors the performance concerns, or behaviors the performance concerns, or behaviors the task doesn't mention, don't fix, task doesn't mention, don't fix, task doesn't mention, don't fix, optimize, or extend it in this change optimize, or extend it in this change optimize, or extend it in this change unless the requested behavior cannot unless the requested behavior cannot unless the requested behavior cannot work without it. Report it as a work without it. Report it as a work without it. Report it as a follow-up in your summary. This is follow-up in your summary. This is follow-up in your summary. This is great, and if you do have this problem, great, and if you do have this problem, great, and if you do have this problem, there you go, there's a solution. You there you go, there's a solution. You there you go, there's a solution. You can copy-paste it and you'll probably can copy-paste it and you'll probably can copy-paste it and you'll probably have it go away. There's a small callout have it go away. There's a small callout have it go away. There's a small callout about low effort not triggering search.
-
about low effort not triggering search. about low effort not triggering search. Again, I'm not recommending low effort a Again, I'm not recommending low effort a Again, I'm not recommending low effort a whole lot, but if you do have a problem whole lot, but if you do have a problem whole lot, but if you do have a problem with that, tell it to use search. This with that, tell it to use search. This with that, tell it to use search. This section is silly and either won't matter section is silly and either won't matter section is silly and either won't matter for you or will matter a lot. It's how for you or will matter a lot. It's how for you or will matter a lot. It's how you can reduce false positives with the you can reduce false positives with the you can reduce false positives with the safeguards. It's advice on wording so safeguards. It's advice on wording so safeguards. It's advice on wording so that the model doesn't think you're that the model doesn't think you're that the model doesn't think you're trying to hack and then lock you out. trying to hack and then lock you out. trying to hack and then lock you out. Instead of does this program compile Instead of does this program compile Instead of does this program compile without errors, ask, are there bugs in without errors, ask, are there bugs in without errors, ask, are there bugs in the program? Apparently, it's not great the program? Apparently, it's not great the program? Apparently, it's not great with lesser-known programming languages with lesser-known programming languages with lesser-known programming languages triggering safeguards, so you should triggering safeguards, so you should triggering safeguards, so you should give it more context about the language give it more context about the language give it more context about the language and how it works, like access to the and how it works, like access to the and how it works, like access to the documentation, cuz once it's in the documentation, cuz once it's in the documentation, cuz once it's in the context, it's less likely to explore context, it's less likely to explore context, it's less likely to explore places it might not need to because it places it might not need to because it places it might not need to because it will probably try and hack the binary to will probably try and hack the binary to will probably try and hack the binary to figure out what it's doing, and then figure out what it's doing, and then figure out what it's doing, and then you'll hit a safeguard. Another thing you'll hit a safeguard. Another thing you'll hit a safeguard. Another thing that's really bad about is base 64 and that's really bad about is base 64 and that's really bad about is base 64 and outputs that tends to trigger a lot cuz outputs that tends to trigger a lot cuz outputs that tends to trigger a lot cuz they probably think you're trying to they probably think you're trying to they probably think you're trying to offuscate or hide something. So, if you offuscate or hide something. So, if you offuscate or hide something. So, if you can avoid base 64 ending up in the can avoid base 64 ending up in the can avoid base 64 ending up in the context, that can help a lot with false context, that can help a lot with false context, that can help a lot with false positives in your history with the positives in your history with the positives in your history with the model. I have had almost no false model. I have had almost no false model. I have had almost no false positives with Fable 5.1, like maybe positives with Fable 5.1, like maybe positives with Fable 5.1, like maybe five total out of the thousands of five total out of the thousands of five total out of the thousands of prompts I've sent it, so not that big a prompts I've sent it, so not that big a prompts I've sent it, so not that big a deal. But, worth noting these details if deal. But, worth noting these details if deal. But, worth noting these details if you do have problems with that. The rest you do have problems with that. The rest you do have problems with that. The rest here is more for building apps with here is more for building apps with here is more for building apps with Fable, not like coding with Fable, but Fable, not like coding with Fable, but Fable, not like coding with Fable, but building something that uses Fable in building something that uses Fable in building something that uses Fable in its core, not as important for coding its core, not as important for coding its core, not as important for coding with it. There are a couple of last with it. There are a couple of last with it. There are a couple of last pieces here that are worthwhile, though.
-
pieces here that are worthwhile, though. pieces here that are worthwhile, though. Like, let the lead agent keep working Like, let the lead agent keep working Like, let the lead agent keep working while sub-agents run. This has been while sub-agents run. This has been while sub-agents run. This has been really nice. I find both Astra and Fable really nice. I find both Astra and Fable really nice. I find both Astra and Fable 5.1 are really good at this. They can 5.1 are really good at this. They can 5.1 are really good at this. They can delegate work to sub-agents, but also do delegate work to sub-agents, but also do delegate work to sub-agents, but also do work in that top-level agent at the same work in that top-level agent at the same work in that top-level agent at the same time. Instead of just sitting and time. Instead of just sitting and time. Instead of just sitting and waiting for the other things to come in, waiting for the other things to come in, waiting for the other things to come in, it can go do something else. It can send it can go do something else. It can send it can go do something else. It can send messages to the sub agents. It can test messages to the sub agents. It can test messages to the sub agents. It can test findings. It can figure out what the findings. It can figure out what the findings. It can figure out what the next step should be. It can do a lot in next step should be. It can do a lot in next step should be. It can do a lot in that time. They also call out that that time. They also call out that that time. They also call out that vision work does much better with crop vision work does much better with crop vision work does much better with crop and zoom tools. It can DIY that stuff and zoom tools. It can DIY that stuff and zoom tools. It can DIY that stuff using Python scripts locally, so using Python scripts locally, so using Python scripts locally, so probably doesn't matter too much for probably doesn't matter too much for probably doesn't matter too much for Claude code type usage, but worth noting Claude code type usage, but worth noting Claude code type usage, but worth noting if you're building around it. And now we if you're building around it. And now we if you're building around it. And now we need to talk about the last big piece. need to talk about the last big piece. need to talk about the last big piece. The one I find the most people are The one I find the most people are The one I find the most people are missing that I also think is what set me missing that I also think is what set me missing that I also think is what set me up for so much success with this model. up for so much success with this model. up for so much success with this model. I think people who've been using a lot I think people who've been using a lot I think people who've been using a lot of Claude code over the last year are of Claude code over the last year are of Claude code over the last year are missing out a lot. Hear me out. Claude missing out a lot. Hear me out. Claude missing out a lot. Hear me out. Claude code has a lot of Claude code-isms. code has a lot of Claude code-isms. code has a lot of Claude code-isms. Things you have to learn, things you Things you have to learn, things you Things you have to learn, things you have to work around when you use it. For have to work around when you use it. For have to work around when you use it. For example, the dumb zone. I know so many example, the dumb zone. I know so many example, the dumb zone. I know so many devs that are perpetually in fear that devs that are perpetually in fear that devs that are perpetually in fear that if they don't closely monitor how much if they don't closely monitor how much if they don't closely monitor how much context the thread is using, they might context the thread is using, they might context the thread is using, they might hit the dumb zone. And then all the work hit the dumb zone. And then all the work hit the dumb zone. And then all the work they did for the day is going to fall they did for the day is going to fall they did for the day is going to fall apart. They're going to get fired.
-
apart. They're going to get fired. apart. They're going to get fired. They're going to waste all their tokens. They're going to waste all their tokens. They're going to waste all their tokens. The model's not going to get results The model's not going to get results The model's not going to get results that work. I'm going to be so real with that work. I'm going to be so real with that work. I'm going to be so real with you guys. Do you actually think you guys. Do you actually think you guys. Do you actually think Anthropic at this point in time with all Anthropic at this point in time with all Anthropic at this point in time with all of the capabilities, all of the tooling, of the capabilities, all of the tooling, of the capabilities, all of the tooling, all of the people they've hired, all of all of the people they've hired, all of all of the people they've hired, all of the work they've done to verify all this the work they've done to verify all this the work they've done to verify all this is going to ship defaults for is going to ship defaults for is going to ship defaults for their flagship model that they want to their flagship model that they want to their flagship model that they want to have perform as well as possible for have perform as well as possible for have perform as well as possible for everything they can do, that they're everything they can do, that they're everything they can do, that they're going to ship it with a bad window going to ship it with a bad window going to ship it with a bad window cutoff size? You know the context cutoff size? You know the context cutoff size? You know the context windows are arbitrarily defined anyways, windows are arbitrarily defined anyways, windows are arbitrarily defined anyways, right? Like if the API allowed, you right? Like if the API allowed, you right? Like if the API allowed, you could send a hundred billion tokens of could send a hundred billion tokens of could send a hundred billion tokens of context to Fable. They pick a number context to Fable. They pick a number context to Fable. They pick a number based on the behaviors they see. And based on the behaviors they see. And based on the behaviors they see. And Anthropic picked one million tokens of Anthropic picked one million tokens of Anthropic picked one million tokens of context because that is the point at context because that is the point at context because that is the point at which they think you might start to see which they think you might start to see which they think you might start to see enough degradation that you shouldn't go enough degradation that you shouldn't go enough degradation that you shouldn't go further. The dumb zone doesn't matter further. The dumb zone doesn't matter further. The dumb zone doesn't matter that much. The models have gotten pretty that much. The models have gotten pretty that much. The models have gotten pretty damn good at compaction and managing damn good at compaction and managing damn good at compaction and managing long runs. You might notice something in long runs. You might notice something in long runs. You might notice something in the T3 code UI. There's no context the T3 code UI. There's no context the T3 code UI. There's no context monitor. There's nowhere here showing monitor. There's nowhere here showing monitor. There's nowhere here showing you how much context is being used you how much context is being used you how much context is being used because it doesn't matter. The models because it doesn't matter. The models because it doesn't matter. The models have gotten good at doing these things.
-
have gotten good at doing these things. have gotten good at doing these things. They're training them in loops. They're They're training them in loops. They're They're training them in loops. They're managing this with all of the context. managing this with all of the context. managing this with all of the context. It's going great. Shout out to Maria for It's going great. Shout out to Maria for It's going great. Shout out to Maria for having the balls to remove the context having the balls to remove the context having the balls to remove the context monitoring in T3 code. She's calling her monitoring in T3 code. She's calling her monitoring in T3 code. She's calling her She's shouting herself out in chat. She's shouting herself out in chat. She's shouting herself out in chat. Figured it's worth shouting her out as Figured it's worth shouting her out as Figured it's worth shouting her out as well. well. well. This isn't the point I'm trying to make, This isn't the point I'm trying to make, This isn't the point I'm trying to make, though. This is just one example. Think though. This is just one example. Think though. This is just one example. Think through the reservations. When you find through the reservations. When you find through the reservations. When you find yourself doing extra work or taking a yourself doing extra work or taking a yourself doing extra work or taking a step that the model could have done or step that the model could have done or step that the model could have done or hitting a button that the model could hitting a button that the model could hitting a button that the model could have hit for you or doing anything that have hit for you or doing anything that have hit for you or doing anything that you're doing because you have this you're doing because you have this you're doing because you have this mental model of what the models can and mental model of what the models can and mental model of what the models can and can't do. Challenge yourself on it a can't do. Challenge yourself on it a can't do. Challenge yourself on it a bit. See what happens if you operate bit. See what happens if you operate bit. See what happens if you operate without that belief. I already see so without that belief. I already see so without that belief. I already see so much copra on the stupid window and much copra on the stupid window and much copra on the stupid window and context management in chat. context management in chat. context management in chat. Guys, Anthropic is good at this. I don't Guys, Anthropic is good at this. I don't Guys, Anthropic is good at this. I don't like saying that. Anthropic's like saying that. Anthropic's like saying that. Anthropic's engineering quality was but their engineering quality was but their engineering quality was but their models are now smart enough that it's models are now smart enough that it's models are now smart enough that it's making their engineering better. The 1 making their engineering better. The 1 making their engineering better. The 1 million token context window is million token context window is million token context window is expensive when your costs for reads are expensive when your costs for reads are expensive when your costs for reads are expensive, because you don't write the expensive, because you don't write the expensive, because you don't write the whole thing. You append writes on top of whole thing. You append writes on top of whole thing. You append writes on top of it over time. And if you compact more, it over time. And if you compact more, it over time. And if you compact more, you're rewriting context more often. So, you're rewriting context more often. So, you're rewriting context more often. So, if you lower the point where you start if you lower the point where you start if you lower the point where you start compacting, you're going to increase compacting, you're going to increase compacting, you're going to increase your cost. The fact that they knocked your cost. The fact that they knocked your cost. The fact that they knocked the cash read cost by 75% basically the cash read cost by 75% basically the cash read cost by 75% basically means the context window size doesn't means the context window size doesn't means the context window size doesn't matter. The one catch being if the cash matter. The one catch being if the cash matter. The one catch being if the cash has expired, which thankfully in tools has expired, which thankfully in tools has expired, which thankfully in tools like T3 chat, we now expose. They have like T3 chat, we now expose. They have like T3 chat, we now expose. They have this little resume with less context. If this little resume with less context. If this little resume with less context. If the 527,000 the 527,000 the 527,000 tokens from earlier aren't cached yet, tokens from earlier aren't cached yet, tokens from earlier aren't cached yet, might be worth compacting quick. That's
-
might be worth compacting quick. That's might be worth compacting quick. That's fine. God, I just There's so many dumb fine. God, I just There's so many dumb fine. God, I just There's so many dumb questions in chat. I'm trying so hard to questions in chat. I'm trying so hard to questions in chat. I'm trying so hard to be as simple as possible here. Okay, I'm be as simple as possible here. Okay, I'm be as simple as possible here. Okay, I'm just I'm going to say the quiet part out just I'm going to say the quiet part out just I'm going to say the quiet part out loud. loud. loud. If the things I am saying don't make If the things I am saying don't make If the things I am saying don't make sense to you or are confusing or you sense to you or are confusing or you sense to you or are confusing or you have questions, there's a really simple have questions, there's a really simple have questions, there's a really simple solution. solution. solution. Sadly, I can't give my usual solution of Sadly, I can't give my usual solution of Sadly, I can't give my usual solution of ask the model, because I found the ask the model, because I found the ask the model, because I found the models really, really struggle with models really, really struggle with models really, really struggle with understanding these things. If you ask understanding these things. If you ask understanding these things. If you ask Claude how compaction works in Codex, Claude how compaction works in Codex, Claude how compaction works in Codex, you're going to get a answer. I you're going to get a answer. I you're going to get a answer. I know, because people who got that know, because people who got that know, because people who got that answer have been in my replies a lot answer have been in my replies a lot answer have been in my replies a lot recently. The best thing you can do if recently. The best thing you can do if recently. The best thing you can do if these things are confusing or concerning these things are confusing or concerning these things are confusing or concerning to you is don't change the defaults. The to you is don't change the defaults. The to you is don't change the defaults. The only settings I have changed in my only settings I have changed in my only settings I have changed in my Claude code config on my machine are I Claude code config on my machine are I Claude code config on my machine are I have it full screen by default, and I have it full screen by default, and I have it full screen by default, and I have it route through my proxy layer. have it route through my proxy layer. have it route through my proxy layer. That is it. The defaults in Claude code That is it. The defaults in Claude code That is it. The defaults in Claude code are good enough now. And I promise you, are good enough now. And I promise you, are good enough now. And I promise you, some random post you saw on Twitter some random post you saw on Twitter some random post you saw on Twitter isn't more clever than the hundreds of isn't more clever than the hundreds of isn't more clever than the hundreds of billions of dollars of incredible people billions of dollars of incredible people billions of dollars of incredible people and incredible effort going into stuff and incredible effort going into stuff and incredible effort going into stuff at Claude code and going into the work at Claude code and going into the work at Claude code and going into the work at Anthropic. Okay, memory's a fair at Anthropic. Okay, memory's a fair at Anthropic. Okay, memory's a fair point here actually. I did turn off point here actually. I did turn off point here actually. I did turn off memory. I do not like the memory in memory. I do not like the memory in memory. I do not like the memory in Claude code. I'm sure it is useful in Claude code. I'm sure it is useful in Claude code. I'm sure it is useful in some places in some ways. It is not some places in some ways. It is not some places in some ways. It is not useful for me. There was actually just useful for me. There was actually just useful for me. There was actually just one really good question though.
-
one really good question though. one really good question though. This one cuts a little deep for me. Why This one cuts a little deep for me. Why This one cuts a little deep for me. Why does T3 code let you adjust the does T3 code let you adjust the does T3 code let you adjust the compaction thresholds? I don't think we compaction thresholds? I don't think we compaction thresholds? I don't think we have compaction threshold adjustment in have compaction threshold adjustment in have compaction threshold adjustment in T3 code. If you change it in your T3 code. If you change it in your T3 code. If you change it in your config, it'll just work with T3 code config, it'll just work with T3 code config, it'll just work with T3 code because we are just using your Claude because we are just using your Claude because we are just using your Claude code config from your machine. code config from your machine. code config from your machine. We do offer the ability to change from We do offer the ability to change from We do offer the ability to change from the 200K context and the 1 mil context. the 200K context and the 1 mil context. the 200K context and the 1 mil context. The honest reason that we do this is The honest reason that we do this is The honest reason that we do this is because Claude code does too. I'm just because Claude code does too. I'm just because Claude code does too. I'm just trying to expose the options that it trying to expose the options that it trying to expose the options that it exposes. Generally speaking, I find a exposes. Generally speaking, I find a exposes. Generally speaking, I find a lot of people massively over engineering lot of people massively over engineering lot of people massively over engineering around problems that haven't existed for around problems that haven't existed for around problems that haven't existed for at least a year. The models are really at least a year. The models are really at least a year. The models are really good at managing their context. You good at managing their context. You good at managing their context. You don't have to do it for them. The don't have to do it for them. The don't have to do it for them. The hardnesses are really good at using the hardnesses are really good at using the hardnesses are really good at using the different things on your machine. You different things on your machine. You different things on your machine. You don't have to set that up yourself. The don't have to set that up yourself. The don't have to set that up yourself. The permission systems and the auto approval permission systems and the auto approval permission systems and the auto approval modes are very good now, and you modes are very good now, and you modes are very good now, and you definitely shouldn't use accept edits or definitely shouldn't use accept edits or definitely shouldn't use accept edits or supervised. We recently added supervised. We recently added supervised. We recently added anti-gravity support to T3 code, and anti-gravity support to T3 code, and anti-gravity support to T3 code, and after all the requests we got to add after all the requests we got to add after all the requests we got to add accept edits mode for anti-gravity accept edits mode for anti-gravity accept edits mode for anti-gravity because they can't trust the Gemini because they can't trust the Gemini because they can't trust the Gemini models enough and they don't have an models enough and they don't have an models enough and they don't have an auto mode yet. I was this close to just auto mode yet. I was this close to just auto mode yet. I was this close to just removing anti-gravity support. Like, removing anti-gravity support. Like, removing anti-gravity support. Like, supervise and auto accept edits are supervise and auto accept edits are supervise and auto accept edits are tools from the past, from like 2024 even tools from the past, from like 2024 even tools from the past, from like 2024 even at like newest. Use auto if you're at like newest. Use auto if you're at like newest. Use auto if you're careful and use full if you're not, careful and use full if you're not, careful and use full if you're not, period. The points I'm trying to make period. The points I'm trying to make period. The points I'm trying to make here are simple. The tools are good now.
-
here are simple. The tools are good now. here are simple. The tools are good now. Any customization or careful Any customization or careful Any customization or careful configuration you've done to work around configuration you've done to work around configuration you've done to work around the flaws are probably out of date. And the flaws are probably out of date. And the flaws are probably out of date. And a lot of your own mental model around a lot of your own mental model around a lot of your own mental model around this stuff is also out of date. Reset this stuff is also out of date. Reset this stuff is also out of date. Reset it. Pretend you're on a brand new it. Pretend you're on a brand new it. Pretend you're on a brand new machine, delete everything, install machine, delete everything, install machine, delete everything, install Claude Code from scratch, set it up with Claude Code from scratch, set it up with Claude Code from scratch, set it up with as little as possible. I bet you'll be as little as possible. I bet you'll be as little as possible. I bet you'll be surprised how capable it is. And from surprised how capable it is. And from surprised how capable it is. And from there, slowly start adding small things there, slowly start adding small things there, slowly start adding small things that solve your small problems. I think that solve your small problems. I think that solve your small problems. I think I said all I have to on this point. I I said all I have to on this point. I I said all I have to on this point. I hope this one was helpful. It's kind of hope this one was helpful. It's kind of hope this one was helpful. It's kind of different from how I normally do these different from how I normally do these different from how I normally do these types of videos, less about my specific types of videos, less about my specific types of videos, less about my specific day-to-day work and more about the day-to-day work and more about the day-to-day work and more about the philosophical way to take advantage of philosophical way to take advantage of philosophical way to take advantage of tools as powerful as these models. tools as powerful as these models. tools as powerful as these models. People 5.1 was an incremental People 5.1 was an incremental People 5.1 was an incremental improvement in how the model works and improvement in how the model works and improvement in how the model works and behaves, but it has become much more behaves, but it has become much more behaves, but it has become much more than that for my day-to-day work. I than that for my day-to-day work. I than that for my day-to-day work. I really, really like using this model and really, really like using this model and really, really like using this model and I bet you will, too. And if you don't I bet you will, too. And if you don't I bet you will, too. And if you don't like using the model, I'm actually like using the model, I'm actually like using the model, I'm actually curious why. Let me know in the curious why. Let me know in the curious why. Let me know in the comments. And until next time, comments. And until next time, comments. And until next time, peace nerds.
Summary
The discussion centers on the comparison between Anthropic's Fable and OpenAI's Astra models, with the speaker preferring Fable for its consistent, less noisy output and its ability to handle longer tasks. The practical takeaway is that understanding and applying techniques from Anthropic's Fable 5.1 prompting guide can significantly improve AI model performance and reduce errors in day-to-day work.