← Back
Theo September 25, 2026 28m

Getting the most out of Opus 5.5

Read full transcript 21 segments
  1. Opus 55 has been out for a bit now, and Opus 55 has been out for a bit now, and the more I use it, the more I think this the more I use it, the more I think this the more I use it, the more I think this model might actually be really, really model might actually be really, really model might actually be really, really good. I've been very impressed with the good. I've been very impressed with the good. I've been very impressed with the code it puts out, with how nice it is to code it puts out, with how nice it is to code it puts out, with how nice it is to interact with, with how it stays on task interact with, with how it stays on task interact with, with how it stays on task for long work, as long as you don't use for long work, as long as you don't use for long work, as long as you don't use Macs. I We'll talk about that in a bit, Macs. I We'll talk about that in a bit, Macs. I We'll talk about that in a bit, don't worry. Today, I want to talk about don't worry. Today, I want to talk about don't worry. Today, I want to talk about how you can get the most out of this how you can get the most out of this how you can get the most out of this model. I've had a couple videos like model. I've had a couple videos like model. I've had a couple videos like this and they vary in performance, but I this and they vary in performance, but I this and they vary in performance, but I think it's some of the more important think it's some of the more important think it's some of the more important work we can do, especially when awesome work we can do, especially when awesome work we can do, especially when awesome people like Addiosmani, who used to be people like Addiosmani, who used to be people like Addiosmani, who used to be part of the Chrome team, recently joined part of the Chrome team, recently joined part of the Chrome team, recently joined Anthropic and is using his depth of Anthropic and is using his depth of Anthropic and is using his depth of knowledge and education capabilities to knowledge and education capabilities to knowledge and education capabilities to write something awesome like this. If write something awesome like this. If write something awesome like this. If all goes well here, not only are you all goes well here, not only are you all goes well here, not only are you going to learn about how you can use going to learn about how you can use going to learn about how you can use Opus 55 more effectively, hopefully I Opus 55 more effectively, hopefully I Opus 55 more effectively, hopefully I will as well. I trust Addie with my damn will as well. I trust Addie with my damn will as well. I trust Addie with my damn life if I'm being real. This guy is life if I'm being real. This guy is life if I'm being real. This guy is super legit and when I saw him joining super legit and when I saw him joining super legit and when I saw him joining Anthropic, I got really hyped and I am Anthropic, I got really hyped and I am Anthropic, I got really hyped and I am so excited to see what he has to teach so excited to see what he has to teach so excited to see what he has to teach us about maximizing the value we get out us about maximizing the value we get out us about maximizing the value we get out of the new Opus release. On the topic of of the new Opus release. On the topic of of the new Opus release. On the topic of maximizing value, we should take a quick maximizing value, we should take a quick maximizing value, we should take a quick break for today's sponsor. If a chef break for today's sponsor. If a chef break for today's sponsor. If a chef doesn't have access to ingredients, they doesn't have access to ingredients, they doesn't have access to ingredients, they probably can't cook a very good meal.

  2. probably can't cook a very good meal. probably can't cook a very good meal. So, why do we think our agents are going So, why do we think our agents are going So, why do we think our agents are going to make good results if they don't have to make good results if they don't have to make good results if they don't have access to 80% of the data they need that access to 80% of the data they need that access to 80% of the data they need that is currently locked away on the is currently locked away on the is currently locked away on the internet? It turns out agents need a way internet? It turns out agents need a way internet? It turns out agents need a way to get that data. They need a way to to get that data. They need a way to to get that data. They need a way to browse the web. And that's why browse the web. And that's why browse the web. And that's why Browserbase built the browser for your Browserbase built the browser for your Browserbase built the browser for your agents. Not just for going to sites and agents. Not just for going to sites and agents. Not just for going to sites and clicking through them, but for getting clicking through them, but for getting clicking through them, but for getting all the context they need from the all the context they need from the all the context they need from the internet. Whether it's through their internet. Whether it's through their internet. Whether it's through their search API to get results on whatever search API to get results on whatever search API to get results on whatever arbitrary queries they need to resolve arbitrary queries they need to resolve arbitrary queries they need to resolve or their fetch API where they need to or their fetch API where they need to or their fetch API where they need to get some data out of a page that is hard get some data out of a page that is hard get some data out of a page that is hard for them to access. How many times have for them to access. How many times have for them to access. How many times have you seen an agent write a curl request, you seen an agent write a curl request, you seen an agent write a curl request, fail, give up, and then go find fail, give up, and then go find fail, give up, and then go find something else that was incorrect? If something else that was incorrect? If something else that was incorrect? If you've ever seen Codex or Claude do you've ever seen Codex or Claude do you've ever seen Codex or Claude do this, you know how miserable it can be. this, you know how miserable it can be. this, you know how miserable it can be. You've probably already burned hundreds You've probably already burned hundreds You've probably already burned hundreds of dollars in tokens hitting these of dollars in tokens hitting these of dollars in tokens hitting these errors silently and not even noticed. errors silently and not even noticed. errors silently and not even noticed. And if your agents do need real browsing And if your agents do need real browsing And if your agents do need real browsing capabilities to do actions on your capabilities to do actions on your capabilities to do actions on your behalf, whether it's to sign into a page behalf, whether it's to sign into a page behalf, whether it's to sign into a page to book something or go find some data to book something or go find some data to book something or go find some data that's behind a payw wall or a signed-in that's behind a payw wall or a signed-in that's behind a payw wall or a signed-in state that you need, Browserbase can state that you need, Browserbase can state that you need, Browserbase can handle all of that for your agents, too. handle all of that for your agents, too. handle all of that for your agents, too. Browserbase's goal is simple. Since Browserbase's goal is simple. Since Browserbase's goal is simple. Since agents are good at calling APIs, turn agents are good at calling APIs, turn agents are good at calling APIs, turn the whole web into one. This means that the whole web into one. This means that the whole web into one. This means that they handle all the details from scaling they handle all the details from scaling they handle all the details from scaling up their servers to make sure you're up their servers to make sure you're up their servers to make sure you're always able to access what you need to always able to access what you need to always able to access what you need to catching broken flows in your apps with catching broken flows in your apps with catching broken flows in your apps with real agents clicking things and noticing real agents clicking things and noticing real agents clicking things and noticing issues to getting around captas, issues to getting around captas, issues to getting around captas, handling off, and so so much more. Ready handling off, and so so much more. Ready handling off, and so so much more. Ready for a crazy lore drop? Browserbase is so for a crazy lore drop? Browserbase is so for a crazy lore drop? Browserbase is so good that it's what Google officially good that it's what Google officially good that it's what Google officially recommends for their computer use work.

  3. recommends for their computer use work. recommends for their computer use work. This is a real Google repo demoing This is a real Google repo demoing This is a real Google repo demoing computer use, and they recommend computer use, and they recommend computer use, and they recommend browserbase. Figure out why everyone browserbase. Figure out why everyone browserbase. Figure out why everyone from Theo to Google loves them at from Theo to Google loves them at from Theo to Google loves them at soy.link/browserbase. soy.link/browserbase. soy.link/browserbase. Let's talk about getting the most out of Let's talk about getting the most out of Let's talk about getting the most out of Opus 5.5. Opus 55 works well in the way Opus 5.5. Opus 55 works well in the way Opus 5.5. Opus 55 works well in the way that you already use Claude. A few that you already use Claude. A few that you already use Claude. A few things do behave differently, though. It things do behave differently, though. It things do behave differently, though. It works for longer on its own. It tells works for longer on its own. It tells works for longer on its own. It tells you plainly what it did, and it thinks you plainly what it did, and it thinks you plainly what it did, and it thinks before every reply. This guide covers before every reply. This guide covers before every reply. This guide covers how to work with Opus 5.5 in Cloud Apps how to work with Opus 5.5 in Cloud Apps how to work with Opus 5.5 in Cloud Apps and Cloud Code, including how to prompt and Cloud Code, including how to prompt and Cloud Code, including how to prompt the model, steer long runs, and check the model, steer long runs, and check the model, steer long runs, and check your results. He starts with a really your results. He starts with a really your results. He starts with a really fun exercise. He recommends trying three fun exercise. He recommends trying three fun exercise. He recommends trying three specific things in your first sessions specific things in your first sessions specific things in your first sessions with Opus 55. The first example he has with Opus 55. The first example he has with Opus 55. The first example he has here is a thing I talk a lot about. The here is a thing I talk a lot about. The here is a thing I talk a lot about. The idea of handing over the whole task. I idea of handing over the whole task. I idea of handing over the whole task. I call this prompting wider having the call this prompting wider having the call this prompting wider having the model come in earlier and go further. model come in earlier and go further. model come in earlier and go further. And an important part of doing this is And an important part of doing this is And an important part of doing this is defining good done states. What is defining good done states. What is defining good done states. What is complete? Rather than start working on complete? Rather than start working on complete? Rather than start working on this feature, say I want you to build it this feature, say I want you to build it this feature, say I want you to build it so it has these three things. Verify it so it has these three things. Verify it so it has these three things. Verify it by showing me a screenshot and then file by showing me a screenshot and then file by showing me a screenshot and then file a PR after. That gives the model a clear a PR after. That gives the model a clear a PR after. That gives the model a clear end state. It's not done until it has end state. It's not done until it has end state. It's not done until it has made the feature. There is a screenshot made the feature. There is a screenshot made the feature. There is a screenshot of the feature and there is a PR up of the feature and there is a PR up of the feature and there is a PR up including that feature and the including that feature and the including that feature and the screenshot. Generally speaking, these screenshot. Generally speaking, these screenshot. Generally speaking, these higher tier models from anthropic and higher tier models from anthropic and higher tier models from anthropic and open AAI work really, really well if you open AAI work really, really well if you open AAI work really, really well if you tell them what done is because they'll tell them what done is because they'll tell them what done is because they'll keep going until they get there. The keep going until they get there. The keep going until they get there. The next two things are more vague, so I'm next two things are more vague, so I'm next two things are more vague, so I'm curious how he frames these as we go on.

  4. curious how he frames these as we go on. curious how he frames these as we go on. The second one was delete think The second one was delete think The second one was delete think carefully lines. You don't need to tell carefully lines. You don't need to tell carefully lines. You don't need to tell the model to think. it already is going the model to think. it already is going the model to think. it already is going to. I don't know if that really fits to. I don't know if that really fits to. I don't know if that really fits here exactly, but interesting inclusion. here exactly, but interesting inclusion. here exactly, but interesting inclusion. The last piece I don't like the wording The last piece I don't like the wording The last piece I don't like the wording for, but it is important. When long runs for, but it is important. When long runs for, but it is important. When long runs end, read what it needs from you first. end, read what it needs from you first. end, read what it needs from you first. This is an important thing, and I This is an important thing, and I This is an important thing, and I actually think the new model does this actually think the new model does this actually think the new model does this quite well, where after it does a run, quite well, where after it does a run, quite well, where after it does a run, it will tell you exactly what it needs. it will tell you exactly what it needs. it will tell you exactly what it needs. It does sometimes miss this though, so It does sometimes miss this though, so It does sometimes miss this though, so I'll often find myself prompting I'll often find myself prompting I'll often find myself prompting accordingly. Here's a real thread I accordingly. Here's a real thread I accordingly. Here's a real thread I kicked off earlier today in T3 Code that kicked off earlier today in T3 Code that kicked off earlier today in T3 Code that I think showcases what I'm talking about I think showcases what I'm talking about I think showcases what I'm talking about relatively well. This is a thread I had relatively well. This is a thread I had relatively well. This is a thread I had started on my machine testing out the started on my machine testing out the started on my machine testing out the new Opus model with Max Reasoning. Six new Opus model with Max Reasoning. Six new Opus model with Max Reasoning. Six and a half hours later, I realized I and a half hours later, I realized I and a half hours later, I realized I wasn't going anywhere. I I really don't wasn't going anywhere. I I really don't wasn't going anywhere. I I really don't think you should use Max on Opus 55. It think you should use Max on Opus 55. It think you should use Max on Opus 55. It just This thread ran for way too long. just This thread ran for way too long. just This thread ran for way too long. So, I had it on Max Reasoning, and I was So, I had it on Max Reasoning, and I was So, I had it on Max Reasoning, and I was running it on my own computer, which was running it on my own computer, which was running it on my own computer, which was even more annoying. As y'all who have even more annoying. As y'all who have even more annoying. As y'all who have been around for a bit know, I have moved been around for a bit know, I have moved been around for a bit know, I have moved pretty much all my agent work off my pretty much all my agent work off my pretty much all my agent work off my laptop and onto other machines, mostly laptop and onto other machines, mostly laptop and onto other machines, mostly my apartment, but a few remote servers my apartment, but a few remote servers my apartment, but a few remote servers as well. So, normally when I use T3 as well. So, normally when I use T3 as well. So, normally when I use T3 code, I trigger it somewhere else. But code, I trigger it somewhere else. But code, I trigger it somewhere else. But in this particular case, I only had the in this particular case, I only had the in this particular case, I only had the codebase and everything set up on this codebase and everything set up on this codebase and everything set up on this particular machine and I wanted to see particular machine and I wanted to see particular machine and I wanted to see it working too. So I ran it on here, it working too. So I ran it on here, it working too. So I ran it on here, forgot about it, went and did other forgot about it, went and did other forgot about it, went and did other things. Came back and it had been things. Came back and it had been things. Came back and it had been running for six and a half hours. I running for six and a half hours. I running for six and a half hours. I switched to high reasoning, asked what switched to high reasoning, asked what switched to high reasoning, asked what was going on. I only asked it what's the was going on. I only asked it what's the was going on. I only asked it what's the state of the work. I it didn't continue state of the work. I it didn't continue state of the work. I it didn't continue the work after there. I told it to the work after there. I told it to the work after there. I told it to continue and it did. 10 minutes later, continue and it did. 10 minutes later, continue and it did. 10 minutes later, it had finished and was in a pretty good it had finished and was in a pretty good it had finished and was in a pretty good spot. I asked how long it thinks these spot. I asked how long it thinks these spot. I asked how long it thinks these changes will take. That was for another changes will take. That was for another changes will take. That was for another demo and video. But here is the prompt I demo and video. But here is the prompt I demo and video. But here is the prompt I actually care about. Remember this

  5. actually care about. Remember this actually care about. Remember this thread was running on this laptop I'm on thread was running on this laptop I'm on thread was running on this laptop I'm on right now and I wanted it to run right now and I wanted it to run right now and I wanted it to run somewhere else. I could have put the somewhere else. I could have put the somewhere else. I could have put the manual work in here. I could have SSH to manual work in here. I could have SSH to manual work in here. I could have SSH to the other computer, cloned to the repo, the other computer, cloned to the repo, the other computer, cloned to the repo, added it to T3 code, opened it there, added it to T3 code, opened it there, added it to T3 code, opened it there, SSH in again to go dump all the SSH in again to go dump all the SSH in again to go dump all the environment variables, make sure it has environment variables, make sure it has environment variables, make sure it has browser use and everything it needs set browser use and everything it needs set browser use and everything it needs set up, and then tell it to go take this up, and then tell it to go take this up, and then tell it to go take this markdown file that I would copy paste markdown file that I would copy paste markdown file that I would copy paste over. Instead, I just told the model over. Instead, I just told the model over. Instead, I just told the model what I wanted. This is the prompt that I what I wanted. This is the prompt that I what I wanted. This is the prompt that I think does a good job of giving the think does a good job of giving the think does a good job of giving the agent an out when it needs it. I want agent an out when it needs it. I want agent an out when it needs it. I want you to kick off an agent run on my you to kick off an agent run on my you to kick off an agent run on my computer named Leftbook. You should be computer named Leftbook. You should be computer named Leftbook. You should be able to see how to connect through the able to see how to connect through the able to see how to connect through the fleet repo on this machine. This was me fleet repo on this machine. This was me fleet repo on this machine. This was me letting it know how to connect so it letting it know how to connect so it letting it know how to connect so it doesn't try to do anything sketchy to doesn't try to do anything sketchy to doesn't try to do anything sketchy to connect and also so it doesn't do connect and also so it doesn't do connect and also so it doesn't do something dumb that might cause it to something dumb that might cause it to something dumb that might cause it to hit a like security guard or flag. I hit a like security guard or flag. I hit a like security guard or flag. I want you to make sure one that the repo want you to make sure one that the repo want you to make sure one that the repo is cloned, two all the needed is cloned, two all the needed is cloned, two all the needed environment variables are there, and environment variables are there, and environment variables are there, and three that you can run opus 55 through three that you can run opus 55 through three that you can run opus 55 through cloud code and T3 code such that I can cloud code and T3 code such that I can cloud code and T3 code such that I can still see the thread how I normally do still see the thread how I normally do still see the thread how I normally do in T3 code. There's a couple pieces here in T3 code. There's a couple pieces here in T3 code. There's a couple pieces here that are important. I want you to make that are important. I want you to make that are important. I want you to make sure is me telling the model, not only sure is me telling the model, not only sure is me telling the model, not only do I want you to do these things, I want do I want you to do these things, I want do I want you to do these things, I want you to validate that you can and inform you to validate that you can and inform you to validate that you can and inform me if it fails. I also am explicitly me if it fails. I also am explicitly me if it fails. I also am explicitly giving it permission for very specific giving it permission for very specific giving it permission for very specific important things like its ability to important things like its ability to important things like its ability to copy my environment variables. If I copy my environment variables. If I copy my environment variables. If I didn't tell it it could do that, it didn't tell it it could do that, it didn't tell it it could do that, it might be concerned when it goes to do might be concerned when it goes to do might be concerned when it goes to do that and then stop and ask me. I don't that and then stop and ask me. I don't that and then stop and ask me. I don't want it to stop. I know it's going to want it to stop. I know it's going to want it to stop. I know it's going to have to do that for this work. So, I have to do that for this work. So, I have to do that for this work. So, I just told it it can so I don't have to just told it it can so I don't have to just told it it can so I don't have to give it permission. Then we get to the give it permission. Then we get to the give it permission. Then we get to the other parts I had here specifically what

  6. other parts I had here specifically what other parts I had here specifically what my goal was. My goal here is to do the my goal was. My goal here is to do the my goal was. My goal here is to do the whole implementation on that machine in whole implementation on that machine in whole implementation on that machine in one pass with however many sub aents and one pass with however many sub aents and one pass with however many sub aents and whatever else are needed with full whatever else are needed with full whatever else are needed with full computer use capability on the other computer use capability on the other computer use capability on the other machine. This is an explicit statement machine. This is an explicit statement machine. This is an explicit statement of what I want. Being explicit about of what I want. Being explicit about of what I want. Being explicit about what you want is great, but being what you want is great, but being what you want is great, but being explicit about what you don't want is explicit about what you don't want is explicit about what you don't want is even better. I don't want you to be even better. I don't want you to be even better. I don't want you to be spamming my machine with computer use spamming my machine with computer use spamming my machine with computer use requests and browser control while I'm requests and browser control while I'm requests and browser control while I'm using it. This is the most annoying using it. This is the most annoying using it. This is the most annoying thing in the world that actually made me thing in the world that actually made me thing in the world that actually made me not use browser use and computer use as not use browser use and computer use as not use browser use and computer use as much. I got so annoyed with codecs just much. I got so annoyed with codecs just much. I got so annoyed with codecs just popping windows up on my screen popping windows up on my screen popping windows up on my screen constantly. When I started using this constantly. When I started using this constantly. When I started using this remotely on other boxes, I started remotely on other boxes, I started remotely on other boxes, I started liking computer use a lot more when I liking computer use a lot more when I liking computer use a lot more when I didn't have to watch it pulling windows didn't have to watch it pulling windows didn't have to watch it pulling windows up when I'm trying to do other [ __ ] on up when I'm trying to do other [ __ ] on up when I'm trying to do other [ __ ] on my computer because I fire and forget. I my computer because I fire and forget. I my computer because I fire and forget. I throw up a thread, I let it do its throw up a thread, I let it do its throw up a thread, I let it do its thing, and then I go respond to some thing, and then I go respond to some thing, and then I go respond to some messages in Slack or I go play a game messages in Slack or I go play a game messages in Slack or I go play a game through my fork of moonlight. I do other through my fork of moonlight. I do other through my fork of moonlight. I do other things while my threads run. I don't things while my threads run. I don't things while my threads run. I don't want them interrupting the other things want them interrupting the other things want them interrupting the other things while they're running. So this specific while they're running. So this specific while they're running. So this specific request to not spam my machine makes it request to not spam my machine makes it request to not spam my machine makes it very explicitly clear that even if it very explicitly clear that even if it very explicitly clear that even if it thinks remote controlling my machine to thinks remote controlling my machine to thinks remote controlling my machine to access the other machine is a good idea access the other machine is a good idea access the other machine is a good idea that it knows not to. Then we get to that it knows not to. Then we get to that it knows not to. Then we get to what might be the most important piece.

  7. what might be the most important piece. what might be the most important piece. If anything does not work as expected If anything does not work as expected If anything does not work as expected when you move this workload over to when you move this workload over to when you move this workload over to Leftbook, don't hesitate to ask Leftbook, don't hesitate to ask Leftbook, don't hesitate to ask questions so we can get it working questions so we can get it working questions so we can get it working right. It sounds silly, but these models right. It sounds silly, but these models right. It sounds silly, but these models have been RL to [ __ ] hell and back to have been RL to [ __ ] hell and back to have been RL to [ __ ] hell and back to not give up on tasks. So if the model not give up on tasks. So if the model not give up on tasks. So if the model asks a question, the initial thing has asks a question, the initial thing has asks a question, the initial thing has been RL to do is to try and answer the been RL to do is to try and answer the been RL to do is to try and answer the question itself. So if it did have question itself. So if it did have question itself. So if it did have questions or issues with my request questions or issues with my request questions or issues with my request here, it by default might not do the here, it by default might not do the here, it by default might not do the right thing. It might ask itself, should right thing. It might ask itself, should right thing. It might ask itself, should we try another way or is this still we try another way or is this still we try another way or is this still right and then go blast through it right and then go blast through it right and then go blast through it anyways. This particular last sentence anyways. This particular last sentence anyways. This particular last sentence gives it an out. It makes the model more gives it an out. It makes the model more gives it an out. It makes the model more willing to stop and ask the question if willing to stop and ask the question if willing to stop and ask the question if it does run into a problem or it does it does run into a problem or it does it does run into a problem or it does get confused. Thankfully, it didn't run get confused. Thankfully, it didn't run get confused. Thankfully, it didn't run into any problems here. Everything went into any problems here. Everything went into any problems here. Everything went relatively well. It gave me an overview relatively well. It gave me an overview relatively well. It gave me an overview of all the things that it did in this of all the things that it did in this of all the things that it did in this process. Called out that I have some process. Called out that I have some process. Called out that I have some connection issues with Leftbook remotely connection issues with Leftbook remotely connection issues with Leftbook remotely right now, which I do actually need to right now, which I do actually need to right now, which I do actually need to figure out. I have some PRs up that figure out. I have some PRs up that figure out. I have some PRs up that should fix that. But the point here is should fix that. But the point here is should fix that. But the point here is that I gave the model exactly what I that I gave the model exactly what I that I gave the model exactly what I wanted. I upfront told it about the wanted. I upfront told it about the wanted. I upfront told it about the things that I was scared it would ask things that I was scared it would ask things that I was scared it would ask about because I didn't want it to stop about because I didn't want it to stop about because I didn't want it to stop and get permission. I just wanted to do and get permission. I just wanted to do and get permission. I just wanted to do the thing. And I also gave it an out the thing. And I also gave it an out the thing. And I also gave it an out where if it didn't have confidence in a where if it didn't have confidence in a where if it didn't have confidence in a path, it could ask me so I could hand it path, it could ask me so I could hand it path, it could ask me so I could hand it the solution. And the result was that the solution. And the result was that the solution. And the result was that after about 17 minutes, it was able to after about 17 minutes, it was able to after about 17 minutes, it was able to get all of this set up. And most get all of this set up. And most get all of this set up. And most importantly, I now had the thread from importantly, I now had the thread from importantly, I now had the thread from my other machine. Here it is. This my other machine. Here it is. This my other machine. Here it is. This thread from my other machine appeared thread from my other machine appeared thread from my other machine appeared and is behaving. And apparently it only and is behaving. And apparently it only and is behaving. And apparently it only did phase zero.

  8. did phase zero. did phase zero. Great. Let me see the prompt it ran Great. Let me see the prompt it ran Great. Let me see the prompt it ran here. Goal: Build the whole ping rewrite here. Goal: Build the whole ping rewrite here. Goal: Build the whole ping rewrite described in overhaul MD in one pass on described in overhaul MD in one pass on described in overhaul MD in one pass on this branch. Read the whole plan before this branch. Read the whole plan before this branch. Read the whole plan before you start. Follow its phase 0 to 7 in you start. Follow its phase 0 to 7 in you start. Follow its phase 0 to 7 in order. And it did phase zero and then order. And it did phase zero and then order. And it did phase zero and then stopped. I want you to hold off on all stopped. I want you to hold off on all stopped. I want you to hold off on all questions until the very end when you questions until the very end when you questions until the very end when you get through to phase 7 in completion. I get through to phase 7 in completion. I get through to phase 7 in completion. I don't want you to stop after every phase don't want you to stop after every phase don't want you to stop after every phase to ask for permission to continue. I to ask for permission to continue. I to ask for permission to continue. I want you to do the whole thing. Continue want you to do the whole thing. Continue want you to do the whole thing. Continue working on this port until the entirety working on this port until the entirety working on this port until the entirety of it is completed and don't inform me of it is completed and don't inform me of it is completed and don't inform me until you have a tail scale link I can until you have a tail scale link I can until you have a tail scale link I can click on. That is the complete rewrite click on. That is the complete rewrite click on. That is the complete rewrite so I can test it and make sure so I can test it and make sure so I can test it and make sure everything works. Don't worry about how everything works. Don't worry about how everything works. Don't worry about how it's hosted. Don't worry about how it's it's hosted. Don't worry about how it's it's hosted. Don't worry about how it's built. Just worry about completing the built. Just worry about completing the built. Just worry about completing the work as assigned. I am admittedly a work as assigned. I am admittedly a work as assigned. I am admittedly a little disappointed that that first little disappointed that that first little disappointed that that first prompt wasn't enough to get it to behave prompt wasn't enough to get it to behave prompt wasn't enough to get it to behave properly, but now it will be. Now I do properly, but now it will be. Now I do properly, but now it will be. Now I do not have to look at this thread again not have to look at this thread again not have to look at this thread again for at least an hour or two. Yay. So for at least an hour or two. Yay. So for at least an hour or two. Yay. So with that all running, let's dive back with that all running, let's dive back with that all running, let's dive back into this blog post. The first section into this blog post. The first section into this blog post. The first section is how should you ask? It starts with is how should you ask? It starts with is how should you ask? It starts with this important piece that funny enough this important piece that funny enough this important piece that funny enough opus doesn't understand. That's why it opus doesn't understand. That's why it opus doesn't understand. That's why it just had the issues there. You have to just had the issues there. You have to just had the issues there. You have to say what done looks like then let it say what done looks like then let it say what done looks like then let it run. Name the finish. Give the whole run. Name the finish. Give the whole run. Name the finish. Give the whole task in one message. Something like the task in one message. Something like the task in one message. Something like the tests pass or every endpoint is tests pass or every endpoint is tests pass or every endpoint is migrated. It needs specific migrated. It needs specific migrated. It needs specific instructions. The reason that we have to instructions. The reason that we have to instructions. The reason that we have to do this is because despite oat was 55 do this is because despite oat was 55 do this is because despite oat was 55 being much better at going on long being much better at going on long being much better at going on long multiart work it might still stop to ask

  9. multiart work it might still stop to ask multiart work it might still stop to ask for permission with a clear finish line for permission with a clear finish line for permission with a clear finish line it will know when it's done but again as it will know when it's done but again as it will know when it's done but again as you saw with that example there without you saw with that example there without you saw with that example there without that it might just randomly stop gives that it might just randomly stop gives that it might just randomly stop gives an example here of what this would look an example here of what this would look an example here of what this would look like migrate the payments endpoint from like migrate the payments endpoint from like migrate the payments endpoint from the old client to the new one done means the old client to the new one done means the old client to the new one done means every endpoint uses the new client the every endpoint uses the new client the every endpoint uses the new client the old one is deleted and the test suite old one is deleted and the test suite old one is deleted and the test suite passes stop and ask me only if a test passes stop and ask me only if a test passes stop and ask me only if a test fails for a reason that you can't fails for a reason that you can't fails for a reason that you can't explain. This is the perfect simple one explain. This is the perfect simple one explain. This is the perfect simple one two three. What you want, what means two three. What you want, what means two three. What you want, what means done, and the out if it needs an out if done, and the out if it needs an out if done, and the out if it needs an out if it does get confused or does something it does get confused or does something it does get confused or does something wrong. Next point, stop telling it to wrong. Next point, stop telling it to wrong. Next point, stop telling it to think hard. I am amazed people still do think hard. I am amazed people still do think hard. I am amazed people still do this. I've seen so many people say like, this. I've seen so many people say like, this. I've seen so many people say like, "Think deeply about this." The models "Think deeply about this." The models "Think deeply about this." The models are pretty good at knowing how much to are pretty good at knowing how much to are pretty good at knowing how much to think. In fact, a a video I kind of want think. In fact, a a video I kind of want think. In fact, a a video I kind of want to do, but I don't know how to like to do, but I don't know how to like to do, but I don't know how to like phrase the ideas and like package it phrase the ideas and like package it phrase the ideas and like package it well enough is a adittedly a crash out well enough is a adittedly a crash out well enough is a adittedly a crash out on max reasoning. Because the cool thing on max reasoning. Because the cool thing on max reasoning. Because the cool thing about these models now is that the about these models now is that the about these models now is that the reasoning levels aren't necessarily reasoning levels aren't necessarily reasoning levels aren't necessarily think this much. The reasoning levels think this much. The reasoning levels think this much. The reasoning levels are think up to this much. Low and are think up to this much. Low and are think up to this much. Low and medium are effectively saying think as medium are effectively saying think as medium are effectively saying think as much as you need until you hit this much as you need until you hit this much as you need until you hit this point and then stop. high and X high are point and then stop. high and X high are point and then stop. high and X high are much higher caps for how much thinking much higher caps for how much thinking much higher caps for how much thinking the model can do before it stops and the model can do before it stops and the model can do before it stops and gives up because the thing was too hard.

  10. gives up because the thing was too hard. gives up because the thing was too hard. Max isn't just making it so the model Max isn't just making it so the model Max isn't just making it so the model can think more. It is removing its can think more. It is removing its can think more. It is removing its ability to think less. And that's what ability to think less. And that's what ability to think less. And that's what scares me about max reasoning and why I scares me about max reasoning and why I scares me about max reasoning and why I specifically think you shouldn't use it. specifically think you shouldn't use it. specifically think you shouldn't use it. Funny enough, I was talking about this Funny enough, I was talking about this Funny enough, I was talking about this with Addie earlier and he's trying to with Addie earlier and he's trying to with Addie earlier and he's trying to figure out how they can improve Max figure out how they can improve Max figure out how they can improve Max reasoning levels in Opus 55 in Cloud reasoning levels in Opus 55 in Cloud reasoning levels in Opus 55 in Cloud Code. But the way I would recommend Code. But the way I would recommend Code. But the way I would recommend thinking about Max right now isn't it thinking about Max right now isn't it thinking about Max right now isn't it can think more. It's forcing it to not can think more. It's forcing it to not can think more. It's forcing it to not think less. You can see this very easily think less. You can see this very easily think less. You can see this very easily here with my runs with 55 using X high here with my runs with 55 using X high here with my runs with 55 using X high and max in Skatebench, which is and max in Skatebench, which is and max in Skatebench, which is admittedly kind of silly benchmark. On X admittedly kind of silly benchmark. On X admittedly kind of silly benchmark. On X high, the average tokens per response high, the average tokens per response high, the average tokens per response was 338. The average duration for a was 338. The average duration for a was 338. The average duration for a response was 6 seconds and the slowest response was 6 seconds and the slowest response was 6 seconds and the slowest was 31 seconds. On max, the tokens went was 31 seconds. On max, the tokens went was 31 seconds. On max, the tokens went from 338 average to 5,000 average, more from 338 average to 5,000 average, more from 338 average to 5,000 average, more than a 10x. Much scarier, it bumped the than a 10x. Much scarier, it bumped the than a 10x. Much scarier, it bumped the average duration to 50 seconds. So, the average duration to 50 seconds. So, the average duration to 50 seconds. So, the average max run took longer than the average max run took longer than the average max run took longer than the slowest X high run. And the slowest run slowest X high run. And the slowest run slowest X high run. And the slowest run was 20 times slower at 600 seconds. You was 20 times slower at 600 seconds. You was 20 times slower at 600 seconds. You know what the best part is? All it got know what the best part is? All it got know what the best part is? All it got out of that was one additional correct out of that was one additional correct out of that was one additional correct answer. It went from 78% to 79%. It cost answer. It went from 78% to 79%. It cost answer. It went from 78% to 79%. It cost 13 times more. It used 15 times the 13 times more. It used 15 times the 13 times more. It used 15 times the tokens. It took 20 times longer in the tokens. It took 20 times longer in the tokens. It took 20 times longer in the worst cases. And I got jack [ __ ] [ __ ] worst cases. And I got jack [ __ ] [ __ ] worst cases. And I got jack [ __ ] [ __ ] out of it. Are there things where Max out of it. Are there things where Max out of it. Are there things where Max could help? Like maybe it is falling for could help? Like maybe it is falling for could help? Like maybe it is falling for a simple trick answer, but if it has to a simple trick answer, but if it has to a simple trick answer, but if it has to reason more, it doesn't fall for the

  11. reason more, it doesn't fall for the reason more, it doesn't fall for the trap. Perhaps. I still don't think you trap. Perhaps. I still don't think you trap. Perhaps. I still don't think you should ever use max reasoning. All the should ever use max reasoning. All the should ever use max reasoning. All the other reasoning levels just change the other reasoning levels just change the other reasoning levels just change the ceiling. And X high can still be fast. ceiling. And X high can still be fast. ceiling. And X high can still be fast. It can still not do much reasoning. It's It can still not do much reasoning. It's It can still not do much reasoning. It's just a matter of how much can it do, not just a matter of how much can it do, not just a matter of how much can it do, not how much will it do. Max is a will. You how much will it do. Max is a will. You how much will it do. Max is a will. You are setting it and forcing it to think are setting it and forcing it to think are setting it and forcing it to think more. I just leave it on higher X high more. I just leave it on higher X high more. I just leave it on higher X high and don't think about it right now and don't think about it right now and don't think about it right now personally. But yeah, low kind of sucks. personally. But yeah, low kind of sucks. personally. But yeah, low kind of sucks. I'll be real there. This model does need I'll be real there. This model does need I'll be real there. This model does need to think a bit. That's why they don't to think a bit. That's why they don't to think a bit. That's why they don't have a no reasoning version. So yeah, have a no reasoning version. So yeah, have a no reasoning version. So yeah, I'm sure a lot of y'all like the the I'm sure a lot of y'all like the the I'm sure a lot of y'all like the the comment I got the most in my previous comment I got the most in my previous comment I got the most in my previous video was, "What reasoning level should video was, "What reasoning level should video was, "What reasoning level should I use? High or XI? They're both fine. I I use? High or XI? They're both fine. I I use? High or XI? They're both fine. I don't have any issue with either. I can don't have any issue with either. I can don't have any issue with either. I can barely notice the difference with barely notice the difference with barely notice the difference with either. Xi feels like it takes a little either. Xi feels like it takes a little either. Xi feels like it takes a little longer. High sometimes misses like longer. High sometimes misses like longer. High sometimes misses like smaller details that Xi doesn't. Not smaller details that Xi doesn't. Not smaller details that Xi doesn't. Not that big of a gap. I just leave it on that big of a gap. I just leave it on that big of a gap. I just leave it on higher XI. And don't tell the model to higher XI. And don't tell the model to higher XI. And don't tell the model to [ __ ] think. It knows that it should [ __ ] think. It knows that it should [ __ ] think. It knows that it should think. It is smarter than you probably think. It is smarter than you probably think. It is smarter than you probably think. Hell, it's smarter than you think. Hell, it's smarter than you think. Hell, it's smarter than you probably are. Smarter than I am. Next, probably are. Smarter than I am. Next, probably are. Smarter than I am. Next, we have add to a running task. If you we have add to a running task. If you we have add to a running task. If you remember something midrun, you can type remember something midrun, you can type remember something midrun, you can type a follow-up while it works. Why does a follow-up while it works. Why does a follow-up while it works. Why does this matter? It matters because the runs this matter? It matters because the runs this matter? It matters because the runs are longer, so going back to the start are longer, so going back to the start are longer, so going back to the start is more expensive. I have done this is more expensive. I have done this is more expensive. I have done this before. I've had times where I noticed before. I've had times where I noticed before. I've had times where I noticed the model was just going the wrong way the model was just going the wrong way the model was just going the wrong way when I was reading its outputs and just when I was reading its outputs and just when I was reading its outputs and just decided [ __ ] it, stop, new thread, copy decided [ __ ] it, stop, new thread, copy decided [ __ ] it, stop, new thread, copy paste prompt, make a few adjustments, paste prompt, make a few adjustments, paste prompt, make a few adjustments, add two more things to the bottom. Like, add two more things to the bottom. Like, add two more things to the bottom. Like, by the way, do these things as well. I by the way, do these things as well. I by the way, do these things as well. I don't do that anymore. The modern don't do that anymore. The modern don't do that anymore. The modern models, specifically Opus, Fable, Soul models, specifically Opus, Fable, Soul models, specifically Opus, Fable, Soul kind of, and Astro mostly are way better kind of, and Astro mostly are way better kind of, and Astro mostly are way better at what we call steering. That's when at what we call steering. That's when at what we call steering. That's when you hit send when it's already working

  12. you hit send when it's already working you hit send when it's already working and it steers the model in the direction and it steers the model in the direction and it steers the model in the direction of what you sent while it's running. of what you sent while it's running. of what you sent while it's running. steering used to have a pretty rough steering used to have a pretty rough steering used to have a pretty rough problem where the RL was on a per problem where the RL was on a per problem where the RL was on a per message basis. So if you said I want you message basis. So if you said I want you message basis. So if you said I want you to do tasks one, two, and four, and it to do tasks one, two, and four, and it to do tasks one, two, and four, and it started working, and you're like, wait, started working, and you're like, wait, started working, and you're like, wait, I forgot to mention three. It would then I forgot to mention three. It would then I forgot to mention three. It would then forget about one, two, and four, forget about one, two, and four, forget about one, two, and four, immediately do three, and then be like, immediately do three, and then be like, immediately do three, and then be like, okay, I finished three. And then not do okay, I finished three. And then not do okay, I finished three. And then not do the other tasks from the previous the other tasks from the previous the other tasks from the previous message. Because through its training, message. Because through its training, message. Because through its training, it effectively was taught to complete it effectively was taught to complete it effectively was taught to complete the message and ignore the history the message and ignore the history the message and ignore the history beyond how it helped the context. They beyond how it helped the context. They beyond how it helped the context. They have now been trained better about this have now been trained better about this have now been trained better about this to take these interruption messages not to take these interruption messages not to take these interruption messages not as a reason to stop previous work or as a reason to stop previous work or as a reason to stop previous work or treat it as done but as what they are treat it as done but as what they are treat it as done but as what they are steering additional things to help the steering additional things to help the steering additional things to help the model go the right way. Here's another model go the right way. Here's another model go the right way. Here's another one that's really important around one that's really important around one that's really important around design. This one actually bit me a bit design. This one actually bit me a bit design. This one actually bit me a bit when I was using the model. When I took when I was using the model. When I took when I was using the model. When I took a look at Witchai which I still think is a look at Witchai which I still think is a look at Witchai which I still think is one of like the cooler projects I can one of like the cooler projects I can one of like the cooler projects I can use to showcase model capabilities. This use to showcase model capabilities. This use to showcase model capabilities. This was built by DRA to show the front-end was built by DRA to show the front-end was built by DRA to show the front-end capabilities of different models. I capabilities of different models. I capabilities of different models. I noticed that Fable 51 had much better noticed that Fable 51 had much better noticed that Fable 51 had much better designs overall than Opus 55 did. This designs overall than Opus 55 did. This designs overall than Opus 55 did. This is like this is the Opus 55 version of is like this is the Opus 55 version of is like this is the Opus 55 version of this design and this is the Fable 51 this design and this is the Fable 51 this design and this is the Fable 51 version. I hope we can all agree that version. I hope we can all agree that version. I hope we can all agree that the Fable version is obviously better the Fable version is obviously better the Fable version is obviously better and nicer. But I saw other people saying and nicer. But I saw other people saying and nicer. But I saw other people saying Opus was a way better designer and I was Opus was a way better designer and I was Opus was a way better designer and I was confused because when I took this quick confused because when I took this quick confused because when I took this quick look here didn't seem to be the case.

  13. look here didn't seem to be the case. look here didn't seem to be the case. Turns out the thing Opus 55 is good at Turns out the thing Opus 55 is good at Turns out the thing Opus 55 is good at isn't just making a nice design when you isn't just making a nice design when you isn't just making a nice design when you say, "Hey, make it pretty." It's much say, "Hey, make it pretty." It's much say, "Hey, make it pretty." It's much better at going the right direction with better at going the right direction with better at going the right direction with design instructions. If you tell it what design instructions. If you tell it what design instructions. If you tell it what you want, and more importantly, also you want, and more importantly, also you want, and more importantly, also tell it what you don't want, it can make tell it what you don't want, it can make tell it what you don't want, it can make really good designs. I like how Addy really good designs. I like how Addy really good designs. I like how Addy framed it here. This is another example framed it here. This is another example framed it here. This is another example of like anthropic loosening their death of like anthropic loosening their death of like anthropic loosening their death grip on comms, letting the employees say grip on comms, letting the employees say grip on comms, letting the employees say things that the models are bad at. The things that the models are bad at. The things that the models are bad at. The reason this matters Opus 55 is that with reason this matters Opus 55 is that with reason this matters Opus 55 is that with no design direction, Opus 55 will fall no design direction, Opus 55 will fall no design direction, Opus 55 will fall back on a few default styles. A general back on a few default styles. A general back on a few default styles. A general instruction like avoid generic looks instruction like avoid generic looks instruction like avoid generic looks will mostly just swap one default for will mostly just swap one default for will mostly just swap one default for another. A list of specific patterns another. A list of specific patterns another. A list of specific patterns will work much better. For example, will work much better. For example, will work much better. For example, build a personal website with build a personal website with build a personal website with placeholder content. Don't use a cream placeholder content. Don't use a cream placeholder content. Don't use a cream or off-white background, italic accent or off-white background, italic accent or off-white background, italic accent words and headings, numbered one, two, words and headings, numbered one, two, words and headings, numbered one, two, three section labels, monospace labels, three section labels, monospace labels, three section labels, monospace labels, or pill-shaped buttons. I don't think or pill-shaped buttons. I don't think or pill-shaped buttons. I don't think you should have to specify that many you should have to specify that many you should have to specify that many negatives. But you get the idea. When negatives. But you get the idea. When negatives. But you get the idea. When you tell it to not do a thing, it won't you tell it to not do a thing, it won't you tell it to not do a thing, it won't do the thing. And once that's done, you do the thing. And once that's done, you do the thing. And once that's done, you can look at it and decide if you like it can look at it and decide if you like it can look at it and decide if you like it or not. If you don't like something it or not. If you don't like something it or not. If you don't like something it added, put that on the list, and then added, put that on the list, and then added, put that on the list, and then ask again. Or something I do a lot, I ask again. Or something I do a lot, I ask again. Or something I do a lot, I use a tool like Shotter where I take a use a tool like Shotter where I take a use a tool like Shotter where I take a screenshot, I draw an arrow, I point it, screenshot, I draw an arrow, I point it, screenshot, I draw an arrow, I point it, and then I just paste this into the chat and then I just paste this into the chat and then I just paste this into the chat prompt. Like, I don't know, here, paste, prompt. Like, I don't know, here, paste, prompt. Like, I don't know, here, paste, and I'll say something like, "This and I'll say something like, "This and I'll say something like, "This sucks. Make it better." That works way sucks. Make it better." That works way sucks. Make it better." That works way better than you would think. Next, we better than you would think. Next, we better than you would think. Next, we have more on steering. Steering long have more on steering. Steering long have more on steering. Steering long runs in Claude Code. The first point, runs in Claude Code. The first point, runs in Claude Code. The first point, tell it which stops you want. You can tell it which stops you want. You can tell it which stops you want. You can put a short rule in your Claude MD file put a short rule in your Claude MD file put a short rule in your Claude MD file about when to stop and ask and when it

  14. about when to stop and ask and when it about when to stop and ask and when it should keep going. The reason this is should keep going. The reason this is should keep going. The reason this is important is because Opus 55 will keep important is because Opus 55 will keep important is because Opus 55 will keep you posted as it works. On long tasks, you posted as it works. On long tasks, you posted as it works. On long tasks, it'll sometimes stop to report instead it'll sometimes stop to report instead it'll sometimes stop to report instead of going on, a summary that names the of going on, a summary that names the of going on, a summary that names the next step without taking it, an offer to next step without taking it, an offer to next step without taking it, an offer to continue, or a list of choices that continue, or a list of choices that continue, or a list of choices that don't actually block the work. It don't actually block the work. It don't actually block the work. It follows instructions that name the follows instructions that name the follows instructions that name the stops. So, name the stops that you want stops. So, name the stops that you want stops. So, name the stops that you want as well. When a step doesn't need my as well. When a step doesn't need my as well. When a step doesn't need my input, keep going. Put status notes in input, keep going. Put status notes in input, keep going. Put status notes in the same message as your next action. the same message as your next action. the same message as your next action. Stop and ask only when you can't Stop and ask only when you can't Stop and ask only when you can't continue without me or before anything continue without me or before anything continue without me or before anything destructive, deleting data, force destructive, deleting data, force destructive, deleting data, force pushing, or changing anything outside of pushing, or changing anything outside of pushing, or changing anything outside of this repo. You know what I'm going to this repo. You know what I'm going to this repo. You know what I'm going to do? I'm putting my money where my mouth do? I'm putting my money where my mouth do? I'm putting my money where my mouth is. I'm pasting this directly into my is. I'm pasting this directly into my is. I'm pasting this directly into my cloud MD. I already have a section cloud MD. I already have a section cloud MD. I already have a section around approvals here, and I'm putting around approvals here, and I'm putting around approvals here, and I'm putting at the end of this. When a step doesn't at the end of this. When a step doesn't at the end of this. When a step doesn't need my input, keep going. put status of need my input, keep going. put status of need my input, keep going. put status of same message yada yada it's the exact same message yada yada it's the exact same message yada yada it's the exact thing we just read. Now this will be in thing we just read. Now this will be in thing we just read. Now this will be in my cloudmd but I don't have a script my cloudmd but I don't have a script my cloudmd but I don't have a script that auto pushes this. I am much sillier that auto pushes this. I am much sillier that auto pushes this. I am much sillier than that nowadays. So I hop here. I than that nowadays. So I hop here. I than that nowadays. So I hop here. I switch over to opus 55 and I say I just switch over to opus 55 and I say I just switch over to opus 55 and I say I just updated the cloud MD in this repo.

  15. updated the cloud MD in this repo. updated the cloud MD in this repo. Commit it and push it out to the rest of Commit it and push it out to the rest of Commit it and push it out to the rest of the fleet. One of the rare instances of the fleet. One of the rare instances of the fleet. One of the rare instances of typeless screwing up my pros but that is typeless screwing up my pros but that is typeless screwing up my pros but that is fine. Cool. Oh, it actually does have a fine. Cool. Oh, it actually does have a fine. Cool. Oh, it actually does have a dictionary. That's nice. Hopefully this dictionary. That's nice. Hopefully this dictionary. That's nice. Hopefully this will be a meaningful improvement. I like will be a meaningful improvement. I like will be a meaningful improvement. I like that. We'll see how it helps. If a run that. We'll see how it helps. If a run that. We'll see how it helps. If a run stops with, "Want me to continue?" Reply stops with, "Want me to continue?" Reply stops with, "Want me to continue?" Reply continue. If it happens often, the rule continue. If it happens often, the rule continue. If it happens often, the rule above will help. Yep, that's why I added above will help. Yep, that's why I added above will help. Yep, that's why I added it. A rule to keep going means fewer it. A rule to keep going means fewer it. A rule to keep going means fewer stops. So, keep your own check before stops. So, keep your own check before stops. So, keep your own check before anything risky or hard to undo. The last anything risky or hard to undo. The last anything risky or hard to undo. The last line of the rule above does that. He line of the rule above does that. He line of the rule above does that. He also calls out the example of prepared also calls out the example of prepared also calls out the example of prepared programming where you might want the programming where you might want the programming where you might want the opposite. Maybe a oneline plan before it opposite. Maybe a oneline plan before it opposite. Maybe a oneline plan before it starts and a short recap at the end starts and a short recap at the end starts and a short recap at the end because you want to be more involved because you want to be more involved because you want to be more involved with it or maybe you have another person with it or maybe you have another person with it or maybe you have another person working on this with you too can be very working on this with you too can be very working on this with you too can be very useful. I don't want that. I want to useful. I don't want that. I want to useful. I don't want that. I want to kick off the thread and come back to it kick off the thread and come back to it kick off the thread and come back to it when it's done if even. So I very much when it's done if even. So I very much when it's done if even. So I very much prefer the telling it to just keep going prefer the telling it to just keep going prefer the telling it to just keep going here. The next point ask it to split big here. The next point ask it to split big here. The next point ask it to split big work across sub aents. The point here is work across sub aents. The point here is work across sub aents. The point here is that for audits, migrations, reviews that for audits, migrations, reviews that for audits, migrations, reviews across large codebases, stuff like that, across large codebases, stuff like that, across large codebases, stuff like that, splitting the work up into sub aents can splitting the work up into sub aents can splitting the work up into sub aents can make it much easier for the work to get make it much easier for the work to get make it much easier for the work to get done within the context window and for done within the context window and for done within the context window and for those different sub aents to go explore those different sub aents to go explore those different sub aents to go explore different things. The reason this different things. The reason this different things. The reason this matters is because early testers had matters is because early testers had matters is because early testers had Opus 55 coordinate parallel sub aents on Opus 55 coordinate parallel sub aents on Opus 55 coordinate parallel sub aents on long audits with little oversight and it long audits with little oversight and it long audits with little oversight and it was successful. The other thing he isn't was successful. The other thing he isn't was successful. The other thing he isn't saying here is that I've noticed Opus 55 saying here is that I've noticed Opus 55 saying here is that I've noticed Opus 55 seems less willing to spin up sub aents seems less willing to spin up sub aents seems less willing to spin up sub aents unless you tell it that it can. So I unless you tell it that it can. So I unless you tell it that it can. So I find myself often telling at the end of find myself often telling at the end of find myself often telling at the end of prompts, use sub agents and workflows prompts, use sub agents and workflows prompts, use sub agents and workflows however you choose. I trust your however you choose. I trust your however you choose. I trust your judgment there. Even then it doesn't judgment there. Even then it doesn't judgment there. Even then it doesn't always do it. So if you like know the always do it. So if you like know the always do it. So if you like know the task is better with sub aents like a

  16. task is better with sub aents like a task is better with sub aents like a giant pile of PR reviews or an audit of giant pile of PR reviews or an audit of giant pile of PR reviews or an audit of your codebase just tell it to you sub your codebase just tell it to you sub your codebase just tell it to you sub agents and it does. The next section is agents and it does. The next section is agents and it does. The next section is around checking the results. What should around checking the results. What should around checking the results. What should you do when long runs end? Look first you do when long runs end? Look first you do when long runs end? Look first for anything Claude is waiting on you for anything Claude is waiting on you for anything Claude is waiting on you for like a decision it left open or a for like a decision it left open or a for like a decision it left open or a change it wants you to approve. Then change it wants you to approve. Then change it wants you to approve. Then read the rest of the Claude summary. The read the rest of the Claude summary. The read the rest of the Claude summary. The reason this matters is because Opus 55 reason this matters is because Opus 55 reason this matters is because Opus 55 reports on its work more clearly than reports on its work more clearly than reports on its work more clearly than Opus 5 did. So it's actually worth Opus 5 did. So it's actually worth Opus 5 did. So it's actually worth reading the outputs. The updates and reading the outputs. The updates and reading the outputs. The updates and their final summaries say what they did, their final summaries say what they did, their final summaries say what they did, what they found, and what it needs from what they found, and what it needs from what they found, and what it needs from you in plain language. This is you in plain language. This is you in plain language. This is interesting. He actually suggests interesting. He actually suggests interesting. He actually suggests putting this in the cloud MD. End every putting this in the cloud MD. End every putting this in the cloud MD. End every run with three headings. Blocked on me, run with three headings. Blocked on me, run with three headings. Blocked on me, changed, and found. I'm not going to go changed, and found. I'm not going to go changed, and found. I'm not going to go that far. When I do want this, I usually that far. When I do want this, I usually that far. When I do want this, I usually just ask. I often will find myself going just ask. I often will find myself going just ask. I often will find myself going back to a really old thread and not back to a really old thread and not back to a really old thread and not remembering what it's for or what the remembering what it's for or what the remembering what it's for or what the status is. I'll just ask it. What's the status is. I'll just ask it. What's the status is. I'll just ask it. What's the status of this work? What do you need status of this work? What do you need status of this work? What do you need from me? Another one I ask a lot and I from me? Another one I ask a lot and I from me? Another one I ask a lot and I I'm amazed at how helpful this has been. I'm amazed at how helpful this has been. I'm amazed at how helpful this has been. What are the risks of merging this code What are the risks of merging this code What are the risks of merging this code today? That's probably my most common today? That's probably my most common today? That's probably my most common prompt. I'll often have a PR up from an prompt. I'll often have a PR up from an prompt. I'll often have a PR up from an agent that ran for hours, made a bunch agent that ran for hours, made a bunch agent that ran for hours, made a bunch of changes, simplified it, and it gives of changes, simplified it, and it gives of changes, simplified it, and it gives me like a 300 line of code thing that I me like a 300 line of code thing that I me like a 300 line of code thing that I don't fully understand or really care to don't fully understand or really care to don't fully understand or really care to read. And I don't know if I should read. And I don't know if I should read. And I don't know if I should bother reading it or not until I know bother reading it or not until I know bother reading it or not until I know how risky it is. So, I'll just ask, how how risky it is. So, I'll just ask, how how risky it is. So, I'll just ask, how much risk do you perceive with this much risk do you perceive with this much risk do you perceive with this change? What's the worst thing that change? What's the worst thing that change? What's the worst thing that would happen if we merge this right now?

  17. would happen if we merge this right now? would happen if we merge this right now? And I have found myself much much better And I have found myself much much better And I have found myself much much better understanding the state of my work when understanding the state of my work when understanding the state of my work when I ask questions like that. Another fun I ask questions like that. Another fun I ask questions like that. Another fun thing you can have the model do is thing you can have the model do is thing you can have the model do is review code. It can review its own code. review code. It can review its own code. review code. It can review its own code. It can review code from other agents and It can review code from other agents and It can review code from other agents and other models. I have personally found other models. I have personally found other models. I have personally found that swapping between model families and that swapping between model families and that swapping between model families and labs is actually pretty useful. I think labs is actually pretty useful. I think labs is actually pretty useful. I think that like generally speaking, I find that like generally speaking, I find that like generally speaking, I find OpenAI models to be better at review OpenAI models to be better at review OpenAI models to be better at review still. they just really dig into the still. they just really dig into the still. they just really dig into the details and won't let go of the thing details and won't let go of the thing details and won't let go of the thing until it's confident it's found until it's confident it's found until it's confident it's found everything. It'll report more stuff that everything. It'll report more stuff that everything. It'll report more stuff that doesn't matter, but it will occasionally doesn't matter, but it will occasionally doesn't matter, but it will occasionally find things that do matter that the find things that do matter that the find things that do matter that the claude models just miss entirely. So, claude models just miss entirely. So, claude models just miss entirely. So, personally, when I have Opus code being personally, when I have Opus code being personally, when I have Opus code being reviewed, I don't want just Opus reviewed, I don't want just Opus reviewed, I don't want just Opus reviewing it. I would like for Aster or reviewing it. I would like for Aster or reviewing it. I would like for Aster or even like Soul to give it a review as even like Soul to give it a review as even like Soul to give it a review as well. That said, 55 does find a lot more well. That said, 55 does find a lot more well. That said, 55 does find a lot more than previous Opus models did. I than previous Opus models did. I than previous Opus models did. I mentioned this in both my Gro and Opus mentioned this in both my Gro and Opus mentioned this in both my Gro and Opus videos. I put together a a bench in videos. I put together a a bench in videos. I put together a a bench in quotes that measures how well different quotes that measures how well different quotes that measures how well different models are able to find areas of models are able to find areas of models are able to find areas of improvement in the T3 code codebase. And improvement in the T3 code codebase. And improvement in the T3 code codebase. And I was very surprised to see that Grock I was very surprised to see that Grock I was very surprised to see that Grock 47 came in second place and Fable was 47 came in second place and Fable was 47 came in second place and Fable was quite a bit lower. The reason why is quite a bit lower. The reason why is quite a bit lower. The reason why is pretty clear when you look at the number pretty clear when you look at the number pretty clear when you look at the number of supported unresolved contradicted of supported unresolved contradicted of supported unresolved contradicted findings. Astra and Grock 47 both found findings. Astra and Grock 47 both found findings. Astra and Grock 47 both found eight things that could be improved and eight things that could be improved and eight things that could be improved and substantiated and supported its claims.

  18. substantiated and supported its claims. substantiated and supported its claims. GB6 Soul actually found nine things, but GB6 Soul actually found nine things, but GB6 Soul actually found nine things, but apparently they weren't quite as valid apparently they weren't quite as valid apparently they weren't quite as valid because the judge panel I had set up because the judge panel I had set up because the judge panel I had set up found it less quality overall. Babel found it less quality overall. Babel found it less quality overall. Babel found only five things in its two runs. found only five things in its two runs. found only five things in its two runs. That's a big difference. Babel That's a big difference. Babel That's a big difference. Babel absolutely vetted the things and they absolutely vetted the things and they absolutely vetted the things and they all matter and are worth improving, but all matter and are worth improving, but all matter and are worth improving, but it just didn't find as much as the it just didn't find as much as the it just didn't find as much as the OpenAI models and hell even Grock did. OpenAI models and hell even Grock did. OpenAI models and hell even Grock did. But also, you can see the gap with Opus But also, you can see the gap with Opus But also, you can see the gap with Opus 5 to 5.5 here. Opus 55 was almost twice 5 to 5.5 here. Opus 55 was almost twice 5 to 5.5 here. Opus 55 was almost twice as successful according to my judging as successful according to my judging as successful according to my judging system. and also didn't find any system. and also didn't find any system. and also didn't find any unresolved or contradicted things. Where unresolved or contradicted things. Where unresolved or contradicted things. Where with Opus 5, it had four supported with Opus 5, it had four supported with Opus 5, it had four supported findings and then two that weren't. So findings and then two that weren't. So findings and then two that weren't. So for every two findings that are valid, for every two findings that are valid, for every two findings that are valid, it has one that's [ __ ] The point it has one that's [ __ ] The point it has one that's [ __ ] The point I'm trying to make is hopping between I'm trying to make is hopping between I'm trying to make is hopping between model families can be useful here. And model families can be useful here. And model families can be useful here. And as much as I prefer coding with Fable as much as I prefer coding with Fable as much as I prefer coding with Fable and Opus, the review quality I get out and Opus, the review quality I get out and Opus, the review quality I get out of Astra is also really solid and maybe of Astra is also really solid and maybe of Astra is also really solid and maybe even Grock, by the way. So consider even Grock, by the way. So consider even Grock, by the way. So consider using Grock for your reviews, too. One using Grock for your reviews, too. One using Grock for your reviews, too. One more important piece with the reviews, more important piece with the reviews, more important piece with the reviews, though, and I do this a lot. You should though, and I do this a lot. You should though, and I do this a lot. You should ask the model to mark the things that it ask the model to mark the things that it ask the model to mark the things that it can't confirm so it can review and can't confirm so it can review and can't confirm so it can review and verify as much as possible. Ideally, you verify as much as possible. Ideally, you verify as much as possible. Ideally, you give it the tooling it needs to open up give it the tooling it needs to open up give it the tooling it needs to open up a browser and check your changes or run a browser and check your changes or run a browser and check your changes or run the system and use computer use to the system and use computer use to the system and use computer use to verify it or a test suite that it can verify it or a test suite that it can verify it or a test suite that it can use or build to verify the things that use or build to verify the things that use or build to verify the things that it's concerned about. So, the model it's concerned about. So, the model it's concerned about. So, the model knows the code works. But if there are knows the code works. But if there are knows the code works. But if there are things it cannot confirm for any of many things it cannot confirm for any of many things it cannot confirm for any of many reasons, whether it's like a tool that reasons, whether it's like a tool that reasons, whether it's like a tool that you need to use different hardware for you need to use different hardware for you need to use different hardware for or using an environment variable or it's or using an environment variable or it's or using an environment variable or it's a more subjective thing, if you ask the a more subjective thing, if you ask the a more subjective thing, if you ask the model to let you know what things it model to let you know what things it model to let you know what things it can't verify, it will. And this is kind

  19. can't verify, it will. And this is kind can't verify, it will. And this is kind of the point we're at now. The models of the point we're at now. The models of the point we're at now. The models know your codebase really well. They know your codebase really well. They know your codebase really well. They know how to operate really well know how to operate really well know how to operate really well autonomously. They know how to get stuff autonomously. They know how to get stuff autonomously. They know how to get stuff done. They can even verify their done. They can even verify their done. They can even verify their changes. They don't know how to work changes. They don't know how to work changes. They don't know how to work with a human yet. That's partially with a human yet. That's partially with a human yet. That's partially because the data doesn't exist. It's because the data doesn't exist. It's because the data doesn't exist. It's partially because this is still a new partially because this is still a new partially because this is still a new phenomena. But if you tell the model how phenomena. But if you tell the model how phenomena. But if you tell the model how to work with you and you tell it what to work with you and you tell it what to work with you and you tell it what you need and what you're expecting, you need and what you're expecting, you need and what you're expecting, it'll usually give it to you. Especially it'll usually give it to you. Especially it'll usually give it to you. Especially models like Opus 55, Fable 5.1, and GPT6 models like Opus 55, Fable 5.1, and GPT6 models like Opus 55, Fable 5.1, and GPT6 Astra mostly. Okay, this is a little bit Astra mostly. Okay, this is a little bit Astra mostly. Okay, this is a little bit of a silly call out. This is actually of a silly call out. This is actually of a silly call out. This is actually about how to use the cloud apps more about how to use the cloud apps more about how to use the cloud apps more effectively. And the first piece is to effectively. And the first piece is to effectively. And the first piece is to check the model picker says Opus 5.5. check the model picker says Opus 5.5. check the model picker says Opus 5.5. This is making me feel like I shouldn't This is making me feel like I shouldn't This is making me feel like I shouldn't be reading this article. be reading this article. be reading this article. So, let's skip down to what to do when So, let's skip down to what to do when So, let's skip down to what to do when messages are flagged. Obus 55 is the messages are flagged. Obus 55 is the messages are flagged. Obus 55 is the first OBUS model to launch with Fable first OBUS model to launch with Fable first OBUS model to launch with Fable level bio and cyber safeguards in cloud level bio and cyber safeguards in cloud level bio and cyber safeguards in cloud apps as well as cloud code. Most flagged apps as well as cloud code. Most flagged apps as well as cloud code. Most flagged messages move to older models and your messages move to older models and your messages move to older models and your work goes on there. Binding security work goes on there. Binding security work goes on there. Binding security vulnerabilities in source code is vulnerabilities in source code is vulnerabilities in source code is allowed in everyday health and allowed in everyday health and allowed in everyday health and educational questions should still work. educational questions should still work. educational questions should still work. These safeguards can sometimes flag These safeguards can sometimes flag These safeguards can sometimes flag legitimate work and we're tuning them to legitimate work and we're tuning them to legitimate work and we're tuning them to cut down on incorrect flags. If you're cut down on incorrect flags. If you're cut down on incorrect flags. If you're switched, here's what you'll see and switched, here's what you'll see and switched, here's what you'll see and what to do in the cloud apps. You'll get what to do in the cloud apps. You'll get what to do in the cloud apps. You'll get a notice that says you were switched to a notice that says you were switched to a notice that says you were switched to a different model. Quad answers on that a different model. Quad answers on that a different model. Quad answers on that model and the chat stays on it too. If model and the chat stays on it too. If model and the chat stays on it too. If you go back and choose 55 again in the you go back and choose 55 again in the you go back and choose 55 again in the model picker, it might work. It might model picker, it might work. It might model picker, it might work. It might flag again. Starting a new chat should flag again. Starting a new chat should flag again. Starting a new chat should avoid it. You they also have settings avoid it. You they also have settings avoid it. You they also have settings for quad where you can turn off the auto for quad where you can turn off the auto for quad where you can turn off the auto switch and instead you'll get a pause or switch and instead you'll get a pause or switch and instead you'll get a pause or an error. I'll be real, I have not hit an error. I'll be real, I have not hit an error. I'll be real, I have not hit as many flags recently. It doesn't seem as many flags recently. It doesn't seem as many flags recently. It doesn't seem like that big a deal, but to each their like that big a deal, but to each their like that big a deal, but to each their own. One thing that is called out here

  20. own. One thing that is called out here own. One thing that is called out here is that you shouldn't ask the model to is that you shouldn't ask the model to is that you shouldn't ask the model to show its reasoning. This is a little show its reasoning. This is a little show its reasoning. This is a little annoying because sometimes I want the annoying because sometimes I want the annoying because sometimes I want the model to explain why it did a thing and model to explain why it did a thing and model to explain why it did a thing and my like natural English way of doing my like natural English way of doing my like natural English way of doing that is what was your reasoning for that is what was your reasoning for that is what was your reasoning for these changes that might flag as you these changes that might flag as you these changes that might flag as you asking it to share the reasoning traces asking it to share the reasoning traces asking it to share the reasoning traces which anthropic doesn't want to do which anthropic doesn't want to do which anthropic doesn't want to do because that can be used to distill because that can be used to distill because that can be used to distill their models. They don't want that so their models. They don't want that so their models. They don't want that so they hide them. It is what it is. As they hide them. It is what it is. As they hide them. It is what it is. As such, anything that matches reasoning is such, anything that matches reasoning is such, anything that matches reasoning is risking getting a flag. I guess there's risking getting a flag. I guess there's risking getting a flag. I guess there's a cute little checklist at the end here. a cute little checklist at the end here. a cute little checklist at the end here. Make sure you tell the model what done Make sure you tell the model what done Make sure you tell the model what done looks like. Don't tell it to think hard, looks like. Don't tell it to think hard, looks like. Don't tell it to think hard, it will anyways. Design requests should it will anyways. Design requests should it will anyways. Design requests should have the styles to leave out as well as have the styles to leave out as well as have the styles to leave out as well as what you want it to look like. Charts what you want it to look like. Charts what you want it to look like. Charts and screenshots should be attached. and screenshots should be attached. and screenshots should be attached. Yeah, I I think people underrate how Yeah, I I think people underrate how Yeah, I I think people underrate how useful screenshotting is. Half the time useful screenshotting is. Half the time useful screenshotting is. Half the time I would need to give a model like I would need to give a model like I would need to give a model like context on an error. I don't copy paste context on an error. I don't copy paste context on an error. I don't copy paste the error. I just screenshot the browser the error. I just screenshot the browser the error. I just screenshot the browser and paste that. They're good at reading and paste that. They're good at reading and paste that. They're good at reading these things. And Claude Code even has these things. And Claude Code even has these things. And Claude Code even has tools where if the screenshot is too tools where if the screenshot is too tools where if the screenshot is too high res and it can't read the text, it high res and it can't read the text, it high res and it can't read the text, it will crop to the area it needs to get will crop to the area it needs to get will crop to the area it needs to get the context it needs. The other sections the context it needs. The other sections the context it needs. The other sections we've spent this whole video going we've spent this whole video going we've spent this whole video going through, but if you do want to read this through, but if you do want to read this through, but if you do want to read this and check it yourself, the link is in and check it yourself, the link is in and check it yourself, the link is in the description as always. This is a the description as always. This is a the description as always. This is a pretty fun article and I'm thankful to pretty fun article and I'm thankful to pretty fun article and I'm thankful to see once again that Anthropic and I are see once again that Anthropic and I are see once again that Anthropic and I are pretty aligned on the best way to use pretty aligned on the best way to use pretty aligned on the best way to use these models. We're now at the point these models. We're now at the point these models. We're now at the point where we need to give our agents more where we need to give our agents more where we need to give our agents more leash. They need to have the ability to leash. They need to have the ability to leash. They need to have the ability to verify their changes. They need to be verify their changes. They need to be verify their changes. They need to be told when done is done and given enough told when done is done and given enough told when done is done and given enough trust to go do the thing. I talk about a trust to go do the thing. I talk about a trust to go do the thing. I talk about a lot of these layers in a video coming lot of these layers in a video coming lot of these layers in a video coming out soon all about how Anthropic made out soon all about how Anthropic made out soon all about how Anthropic made the performance for the cloud site and the performance for the cloud site and the performance for the cloud site and desktop app three times faster because

  21. desktop app three times faster because desktop app three times faster because the systems you build to verify these the systems you build to verify these the systems you build to verify these changes are just as important as the raw changes are just as important as the raw changes are just as important as the raw source code itself. Hopefully this video source code itself. Hopefully this video source code itself. Hopefully this video will help you maximize your usage of will help you maximize your usage of will help you maximize your usage of Opus 5.5. It really is a great model and Opus 5.5. It really is a great model and Opus 5.5. It really is a great model and the more you trust it, the more you give the more you trust it, the more you give the more you trust it, the more you give it the things it needs to trust its own it the things it needs to trust its own it the things it needs to trust its own work, the further it can go and the more work, the further it can go and the more work, the further it can go and the more you can build. I hope this was helpful you can build. I hope this was helpful you can build. I hope this was helpful and until next time, peace nerds. Also, and until next time, peace nerds. Also, and until next time, peace nerds. Also, don't do not touch max mode.

No summary available yet.

View original episode ↗