← Back
Theo September 23, 2026 40m

Wow.

Read full transcript 31 segments
  1. Approximately 10 months ago, the way Approximately 10 months ago, the way that we wrote code with AI changed that we wrote code with AI changed that we wrote code with AI changed fundamentally with a single model fundamentally with a single model fundamentally with a single model release. The model, of course, was Opus release. The model, of course, was Opus release. The model, of course, was Opus 4.5, which, believe it or not, came out 4.5, which, believe it or not, came out 4.5, which, believe it or not, came out in November last year, almost a year in November last year, almost a year in November last year, almost a year ago. Now, Opus 4.5 was a really ago. Now, Opus 4.5 was a really ago. Now, Opus 4.5 was a really interesting release for a handful of interesting release for a handful of interesting release for a handful of reasons. First off, it was way more reasons. First off, it was way more reasons. First off, it was way more expensive than Sonnet, which made it not expensive than Sonnet, which made it not expensive than Sonnet, which made it not really the default choice for really the default choice for really the default choice for developers. Second off, they skipped a developers. Second off, they skipped a developers. Second off, they skipped a bunch of numbers and went straight to.5. bunch of numbers and went straight to.5. bunch of numbers and went straight to.5. But third off, despite being a But third off, despite being a But third off, despite being a revolutionary change in how we can use revolutionary change in how we can use revolutionary change in how we can use models for code, they didn't make it models for code, they didn't make it models for code, they didn't make it Opus 5, they made it 4.5. Since then, Opus 5, they made it 4.5. Since then, Opus 5, they made it 4.5. Since then, Anthropic has released some incredible Anthropic has released some incredible Anthropic has released some incredible stuff in the Fable line, but pretty much stuff in the Fable line, but pretty much stuff in the Fable line, but pretty much every other launch they've done has been every other launch they've done has been every other launch they've done has been iffy at best, especially model releases iffy at best, especially model releases iffy at best, especially model releases like Opus 5 and Sonnet 5. And I'll take like Opus 5 and Sonnet 5. And I'll take like Opus 5 and Sonnet 5. And I'll take some fault here. I fell for Opus 5 when some fault here. I fell for Opus 5 when some fault here. I fell for Opus 5 when it came out. It seemed like it was it came out. It seemed like it was it came out. It seemed like it was working really well. But when you working really well. But when you working really well. But when you combine the absolute slop it would combine the absolute slop it would combine the absolute slop it would output in text trying to describe what output in text trying to describe what output in text trying to describe what it was doing with the code that would it was doing with the code that would it was doing with the code that would often miss really important stuff, the often miss really important stuff, the often miss really important stuff, the model felt more like a liability than an model felt more like a liability than an model felt more like a liability than an actual thing you could rely on when you actual thing you could rely on when you actual thing you could rely on when you worked every day. All of this sets a worked every day. All of this sets a worked every day. All of this sets a pretty brutal stage for Anthropic who pretty brutal stage for Anthropic who pretty brutal stage for Anthropic who previously made incredible models in the previously made incredible models in the previously made incredible models in the Opus line and has since kind of let it Opus line and has since kind of let it Opus line and has since kind of let it die. All of that said, I still did hold die. All of that said, I still did hold die. All of that said, I still did hold some hope for Opus because I needed some hope for Opus because I needed some hope for Opus because I needed something to use the other half of my something to use the other half of my something to use the other half of my anthropic subs with. And I still have a anthropic subs with. And I still have a anthropic subs with. And I still have a soft spot in my heart for that moment in soft spot in my heart for that moment in soft spot in my heart for that moment in November last year where all of a sudden November last year where all of a sudden November last year where all of a sudden I was watching my computer code for me I was watching my computer code for me I was watching my computer code for me instead of just telling it what to instead of just telling it what to instead of just telling it what to change. That's why I was so shocked when change. That's why I was so shocked when change. That's why I was so shocked when I saw that this isn't Opus 5.1, it was I saw that this isn't Opus 5.1, it was I saw that this isn't Opus 5.1, it was 5.5. They set a really, really high bar 5.5. They set a really, really high bar 5.5. They set a really, really high bar for this release. And while it is

  2. for this release. And while it is for this release. And while it is admittedly far too early to know for admittedly far too early to know for admittedly far too early to know for sure, think they may have hit it. I've sure, think they may have hit it. I've sure, think they may have hit it. I've been using Opus 5.5 all day and it has been using Opus 5.5 all day and it has been using Opus 5.5 all day and it has been blowing me away. They made it been blowing me away. They made it been blowing me away. They made it cheaper. They made it less bad at cheaper. They made it less bad at cheaper. They made it less bad at writing. They made it way better at code writing. They made it way better at code writing. They made it way better at code and following instructions and just and following instructions and just and following instructions and just doing stuff every day. Is this the Fable doing stuff every day. Is this the Fable doing stuff every day. Is this the Fable Killer I've been waiting for? I can't Killer I've been waiting for? I can't Killer I've been waiting for? I can't know that just yet. But is this model know that just yet. But is this model know that just yet. But is this model blowing me away? Absolutely. I can't blowing me away? Absolutely. I can't blowing me away? Absolutely. I can't wait to show you all of the reasons why wait to show you all of the reasons why wait to show you all of the reasons why I'm so impressed after a real quick I'm so impressed after a real quick I'm so impressed after a real quick break for today's sponsor. While I was break for today's sponsor. While I was break for today's sponsor. While I was prepping to film today, I was trying to prepping to film today, I was trying to prepping to film today, I was trying to get a bunch of dashboards open. It's get a bunch of dashboards open. It's get a bunch of dashboards open. It's been a little hard for me cuz I'm kind been a little hard for me cuz I'm kind been a little hard for me cuz I'm kind of down a hand. So, I decided to let of down a hand. So, I decided to let of down a hand. So, I decided to let Codeex use computer use to go through Codeex use computer use to go through Codeex use computer use to go through all of these pages and get me signed in. all of these pages and get me signed in. all of these pages and get me signed in. And I was blown away at just how many And I was blown away at just how many And I was blown away at just how many times it couldn't do it. The biggest times it couldn't do it. The biggest times it couldn't do it. The biggest reason is that these sites don't offer reason is that these sites don't offer reason is that these sites don't offer good O solutions for agents, which makes good O solutions for agents, which makes good O solutions for agents, which makes it really hard for my agent to do it really hard for my agent to do it really hard for my agent to do anything on my behalf without me signing anything on my behalf without me signing anything on my behalf without me signing into the browser myself. If only all into the browser myself. If only all into the browser myself. If only all those companies were using today's those companies were using today's those companies were using today's sponsor, Work OS, because they built the sponsor, Work OS, because they built the sponsor, Work OS, because they built the standard for this, OMD. The goal of OMD standard for this, OMD. The goal of OMD standard for this, OMD. The goal of OMD is to give an easy way for agents to O is to give an easy way for agents to O is to give an easy way for agents to O on a user's behalf or set things up for on a user's behalf or set things up for on a user's behalf or set things up for a user ahead of time. Work OS partnered a user ahead of time. Work OS partnered a user ahead of time. Work OS partnered with Cloudflare and Firecrawl to build with Cloudflare and Firecrawl to build with Cloudflare and Firecrawl to build this standard and now a ton of other this standard and now a ton of other this standard and now a ton of other companies are adopting it and helping companies are adopting it and helping companies are adopting it and helping push it further. Companies like Neon, push it further. Companies like Neon, push it further. Companies like Neon, Recent, Parallel Monday, and more are Recent, Parallel Monday, and more are Recent, Parallel Monday, and more are already moving over. And if you're already moving over. And if you're already moving over. And if you're looking for the easiest way to adopt looking for the easiest way to adopt looking for the easiest way to adopt this new standard, work OS is probably this new standard, work OS is probably this new standard, work OS is probably your best bet. There's a reason your best bet. There's a reason your best bet. There's a reason companies like OpenAI and Anthropic companies like OpenAI and Anthropic companies like OpenAI and Anthropic trust them so deeply, it's because Work trust them so deeply, it's because Work trust them so deeply, it's because Work OS solved enterprise off, not just OS solved enterprise off, not just OS solved enterprise off, not just simple signin with Google buttons. All simple signin with Google buttons. All simple signin with Google buttons. All the pieces that real businesses need the pieces that real businesses need the pieces that real businesses need when they want to adopt your product. If when they want to adopt your product. If when they want to adopt your product. If you've never had to configure Octa, ADP, you've never had to configure Octa, ADP, you've never had to configure Octa, ADP, or Duo at your job, I envy you. It is

  3. or Duo at your job, I envy you. It is or Duo at your job, I envy you. It is hell and miserable, and I wouldn't wish hell and miserable, and I wouldn't wish hell and miserable, and I wouldn't wish it on my worst enemy. Thankfully, when it on my worst enemy. Thankfully, when it on my worst enemy. Thankfully, when you use Work OS, you get the admin you use Work OS, you get the admin you use Work OS, you get the admin portal. Just you send a single link to portal. Just you send a single link to portal. Just you send a single link to the IT team at the other company, and the IT team at the other company, and the IT team at the other company, and now they're good to go. Works is the now they're good to go. Works is the now they're good to go. Works is the only platform that is loved by only platform that is loved by only platform that is loved by developers, agents, companies, and IT developers, agents, companies, and IT developers, agents, companies, and IT teams. Figure out why atv.link/workos. teams. Figure out why atv.link/workos. teams. Figure out why atv.link/workos. Let's dive into Opus 5.5. Kind of crazy. Let's dive into Opus 5.5. Kind of crazy. Let's dive into Opus 5.5. Kind of crazy. This is one of four major model drops in This is one of four major model drops in This is one of four major model drops in the last 2 days with Gro 4.7, then Opus the last 2 days with Gro 4.7, then Opus the last 2 days with Gro 4.7, then Opus 5.5, and then GPT6 Soul and Luna, which 5.5, and then GPT6 Soul and Luna, which 5.5, and then GPT6 Soul and Luna, which of course we'll be covering in the near of course we'll be covering in the near of course we'll be covering in the near future. If you haven't hit the subscribe future. If you haven't hit the subscribe future. If you haven't hit the subscribe button, you should consider it because I button, you should consider it because I button, you should consider it because I stay up all night working on these stay up all night working on these stay up all night working on these videos nowadays cuz there's just so much videos nowadays cuz there's just so much videos nowadays cuz there's just so much to cover. I do my best to distill it. to cover. I do my best to distill it. to cover. I do my best to distill it. I'm sorry the videos are long, but hit I'm sorry the videos are long, but hit I'm sorry the videos are long, but hit the button if you don't mind. Helps us the button if you don't mind. Helps us the button if you don't mind. Helps us out a ton. Less than half of y'all are out a ton. Less than half of y'all are out a ton. Less than half of y'all are subbed and it's a great opportunity to subbed and it's a great opportunity to subbed and it's a great opportunity to stay on top of these things as they stay on top of these things as they stay on top of these things as they change constantly. So, what makes Opus change constantly. So, what makes Opus change constantly. So, what makes Opus 5.5 so cool? We'll start with what they 5.5 so cool? We'll start with what they 5.5 so cool? We'll start with what they had to say. It performs at the level of had to say. It performs at the level of had to say. It performs at the level of Fable 5.1 on most work and it costs 40% Fable 5.1 on most work and it costs 40% Fable 5.1 on most work and it costs 40% less to run than Opus 5. They open by less to run than Opus 5. They open by less to run than Opus 5. They open by saying that this is the first release saying that this is the first release saying that this is the first release since they called for pacing the since they called for pacing the since they called for pacing the frontier. I have a lot of thoughts on frontier. I have a lot of thoughts on frontier. I have a lot of thoughts on that that I'm going to reserve for a that that I'm going to reserve for a that that I'm going to reserve for a future video, so we'll ignore that for future video, so we'll ignore that for future video, so we'll ignore that for now. What I care a lot more about is how now. What I care a lot more about is how now. What I care a lot more about is how this model actually works. They this model actually works. They this model actually works. They mentioned that they had a tester that mentioned that they had a tester that mentioned that they had a tester that completed a 680,000 line of code completed a 680,000 line of code completed a 680,000 line of code migration in less than a day, which is migration in less than a day, which is migration in less than a day, which is genuinely very impressive. They also genuinely very impressive. They also genuinely very impressive. They also site its ability to fix inefficiencies site its ability to fix inefficiencies site its ability to fix inefficiencies in software. I've experienced this too, in software. I've experienced this too, in software. I've experienced this too, throwing it at some pretty brutal stuff throwing it at some pretty brutal stuff throwing it at some pretty brutal stuff to try and do performance optimizations.

  4. to try and do performance optimizations. to try and do performance optimizations. I'm impressed. They say it's really safe I'm impressed. They say it's really safe I'm impressed. They say it's really safe as well. At this point, I just trust as well. At this point, I just trust as well. At this point, I just trust Anthropic when they say this, but now we Anthropic when they say this, but now we Anthropic when they say this, but now we need to go to the cost and speed part. need to go to the cost and speed part. need to go to the cost and speed part. This is where things start to get fun. This is where things start to get fun. This is where things start to get fun. 5.5 requires less compute to serve than 5.5 requires less compute to serve than 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Opus 5, and its pricing reflects that. Opus 5, and its pricing reflects that. Our tests show that at default settings, Our tests show that at default settings, Our tests show that at default settings, it will cost 40% less than Opus 5 on it will cost 40% less than Opus 5 on it will cost 40% less than Opus 5 on typical workloads. Input and output typical workloads. Input and output typical workloads. Input and output tokens are now cheaper at $4 per mill in tokens are now cheaper at $4 per mill in tokens are now cheaper at $4 per mill in and $20 per mill out, which is 20% less and $20 per mill out, which is 20% less and $20 per mill out, which is 20% less than before. They also made a huge than before. They also made a huge than before. They also made a huge change to cash reads. Not quite as big change to cash reads. Not quite as big change to cash reads. Not quite as big as the change that they made for Fable, as the change that they made for Fable, as the change that they made for Fable, but closeish. Now it's 60% cheaper to do but closeish. Now it's 60% cheaper to do but closeish. Now it's 60% cheaper to do cash reads than it was before. Only 20 cash reads than it was before. Only 20 cash reads than it was before. Only 20 cents per mill, which is a very good cents per mill, which is a very good cents per mill, which is a very good deal. It's also about 30% faster, which deal. It's also about 30% faster, which deal. It's also about 30% faster, which I have seen from my numbers. It is much I have seen from my numbers. It is much I have seen from my numbers. It is much faster, which is nice because Opus can faster, which is nice because Opus can faster, which is nice because Opus can be a bit token hungry. This version in be a bit token hungry. This version in be a bit token hungry. This version in particular, uh, we'll talk about the particular, uh, we'll talk about the particular, uh, we'll talk about the token efficiency in a bit, don't worry. token efficiency in a bit, don't worry. token efficiency in a bit, don't worry. My favorite change by far, though, is My favorite change by far, though, is My favorite change by far, though, is communication. 5.5 communicates more communication. 5.5 communicates more communication. 5.5 communicates more naturally than prior models. Early naturally than prior models. Early naturally than prior models. Early testers found its writing to be clearer testers found its writing to be clearer testers found its writing to be clearer and easier to follow, which addresses and easier to follow, which addresses and easier to follow, which addresses some of the common feedback that they some of the common feedback that they some of the common feedback that they heard about Opus 5. It puts the most heard about Opus 5. It puts the most heard about Opus 5. It puts the most important information up front, and its important information up front, and its important information up front, and its style makes it a better work partner style makes it a better work partner style makes it a better work partner over long sessions. As one early tester over long sessions. As one early tester over long sessions. As one early tester put it, quote, "It writes the way I do."

  5. put it, quote, "It writes the way I do." put it, quote, "It writes the way I do." They also call out that being easier to They also call out that being easier to They also call out that being easier to read and follow along is actually a read and follow along is actually a read and follow along is actually a safety benefit as well as a practical safety benefit as well as a practical safety benefit as well as a practical one. They have a real fun teaser at the one. They have a real fun teaser at the one. They have a real fun teaser at the end here saying that Sonet 55 and Haiku end here saying that Sonet 55 and Haiku end here saying that Sonet 55 and Haiku 55 are going to follow in the coming 55 are going to follow in the coming 55 are going to follow in the coming weeks. That's going to be our first weeks. That's going to be our first weeks. That's going to be our first Haiku release in a year, by the way. Haiku release in a year, by the way. Haiku release in a year, by the way. Insane that it's taken them that long. Insane that it's taken them that long. Insane that it's taken them that long. They are getting absolutely destroyed by They are getting absolutely destroyed by They are getting absolutely destroyed by OpenAI on the cheaper, smaller models, OpenAI on the cheaper, smaller models, OpenAI on the cheaper, smaller models, but it does seem like they have a solid but it does seem like they have a solid but it does seem like they have a solid lead on the big ones right now. First, lead on the big ones right now. First, lead on the big ones right now. First, we have the benchmarks where literally we have the benchmarks where literally we have the benchmarks where literally all of them other than the agentic all of them other than the agentic all of them other than the agentic scientific research bench, terminal scientific research bench, terminal scientific research bench, terminal bench science, and business workflows bench science, and business workflows bench science, and business workflows with automation bench. There are a few with automation bench. There are a few with automation bench. There are a few other benches that weren't included here other benches that weren't included here other benches that weren't included here where they didn't score quite as well, where they didn't score quite as well, where they didn't score quite as well, but I'm still waiting on updated numbers but I'm still waiting on updated numbers but I'm still waiting on updated numbers from those. In particular, Deep Suite. from those. In particular, Deep Suite. from those. In particular, Deep Suite. Guys, can you update the bench? I've Guys, can you update the bench? I've Guys, can you update the bench? I've been bugging you on Slack forever. Your been bugging you on Slack forever. Your been bugging you on Slack forever. Your data is just so out of date now. H. data is just so out of date now. H. data is just so out of date now. H. Anyways, here's where we start to Anyways, here's where we start to Anyways, here's where we start to actually look at the results where we actually look at the results where we actually look at the results where we can see the cost against the score for can see the cost against the score for can see the cost against the score for various things including Terminal Bench various things including Terminal Bench various things including Terminal Bench 4 where it is slaughtering where Opus 4 where it is slaughtering where Opus 4 where it is slaughtering where Opus 5.5 Medium is scoring higher than Fable 5.5 Medium is scoring higher than Fable 5.5 Medium is scoring higher than Fable did on max and it's under half the did on max and it's under half the did on max and it's under half the price. Actually, it's a lot less than price. Actually, it's a lot less than price. Actually, it's a lot less than that. It was $20 for Fable 5.1 on Max that. It was $20 for Fable 5.1 on Max that. It was $20 for Fable 5.1 on Max and it was $2.94 for Opus 5.5 on medium.

  6. and it was $2.94 for Opus 5.5 on medium. and it was $2.94 for Opus 5.5 on medium. Yeah, it did have a slight dip on max Yeah, it did have a slight dip on max Yeah, it did have a slight dip on max and we'll be talking about that in a and we'll be talking about that in a and we'll be talking about that in a bit. I have feelings about the max bit. I have feelings about the max bit. I have feelings about the max effort level. If we look at other things effort level. If we look at other things effort level. If we look at other things like Frontier Code, you can see the the like Frontier Code, you can see the the like Frontier Code, you can see the the weirdness of this particular bench where weirdness of this particular bench where weirdness of this particular bench where things go down as often as they go up things go down as often as they go up things go down as often as they go up when reasoning levels increase, but it when reasoning levels increase, but it when reasoning levels increase, but it is a new all-time high score on medium is a new all-time high score on medium is a new all-time high score on medium and on max, which is funny. Their best and on max, which is funny. Their best and on max, which is funny. Their best score was medium. I don't know why they score was medium. I don't know why they score was medium. I don't know why they put this in here. It just weird bench put this in here. It just weird bench put this in here. It just weird bench still. But then we go to Cursor Bench, still. But then we go to Cursor Bench, still. But then we go to Cursor Bench, which is one of the most fun ones. which is one of the most fun ones. which is one of the most fun ones. Sadly, Cursor Bench no longer includes Sadly, Cursor Bench no longer includes Sadly, Cursor Bench no longer includes Astra because OpenAI has banned Cursor Astra because OpenAI has banned Cursor Astra because OpenAI has banned Cursor from using their models. But we can from using their models. But we can from using their models. But we can still see how it compares to Fable, still see how it compares to Fable, still see how it compares to Fable, which again shows Medium as scoring which again shows Medium as scoring which again shows Medium as scoring higher than Max did with 5.1. Kind of higher than Max did with 5.1. Kind of higher than Max did with 5.1. Kind of crazy that Opus 55 on Medium in most crazy that Opus 55 on Medium in most crazy that Opus 55 on Medium in most real world codebenches is showing a much real world codebenches is showing a much real world codebenches is showing a much higher score than Fable. I have a lot of higher score than Fable. I have a lot of higher score than Fable. I have a lot of feelings about how these charts are feelings about how these charts are feelings about how these charts are visualized to the point where I actually visualized to the point where I actually visualized to the point where I actually took the time to make my own alternative took the time to make my own alternative took the time to make my own alternative visualization. I'm just using the data visualization. I'm just using the data visualization. I'm just using the data from artificial analysis here, but I from artificial analysis here, but I from artificial analysis here, but I wanted to make it easier to showcase the wanted to make it easier to showcase the wanted to make it easier to showcase the actual cost gaps between these different actual cost gaps between these different actual cost gaps between these different options. Right now I have it on linear options. Right now I have it on linear options. Right now I have it on linear cost scale which really emphasizes how cost scale which really emphasizes how cost scale which really emphasizes how cheap Luna is because it's just stuffed cheap Luna is because it's just stuffed cheap Luna is because it's just stuffed in the side there when compared to all in the side there when compared to all in the side there when compared to all these other things. But I'll switch to these other things. But I'll switch to these other things. But I'll switch to log so it's a tiny bit more readable in log so it's a tiny bit more readable in log so it's a tiny bit more readable in particular the section we care about particular the section we care about particular the section we care about here. What you'll see is that this model here. What you'll see is that this model here. What you'll see is that this model is more expensive than Astra in a lot of is more expensive than Astra in a lot of is more expensive than Astra in a lot of cases but it is also proving to be much cases but it is also proving to be much cases but it is also proving to be much more effective with its results. Like more effective with its results. Like more effective with its results. Like here, Opus 5.5 on high cost more than here, Opus 5.5 on high cost more than here, Opus 5.5 on high cost more than GBD6 Astra on high, but it also scored a GBD6 Astra on high, but it also scored a GBD6 Astra on high, but it also scored a meaningfully better score getting a 53

  7. meaningfully better score getting a 53 meaningfully better score getting a 53 on artificial analysis versus the 50.9 on artificial analysis versus the 50.9 on artificial analysis versus the 50.9 that Astra got. So again, it is still that Astra got. So again, it is still that Astra got. So again, it is still not cheap, largely due to the token not cheap, largely due to the token not cheap, largely due to the token inefficiency, but god damn, these scores inefficiency, but god damn, these scores inefficiency, but god damn, these scores are insane, but also low is really bad. are insane, but also low is really bad. are insane, but also low is really bad. This model's interesting in this This model's interesting in this This model's interesting in this particular way. It feels like low and particular way. It feels like low and particular way. It feels like low and max are dangerous and generally best max are dangerous and generally best max are dangerous and generally best avoided, but medium, high, XH high, avoided, but medium, high, XH high, avoided, but medium, high, XH high, those all have seemed really good for my those all have seemed really good for my those all have seemed really good for my experience so far. They also report experience so far. They also report experience so far. They also report improvements to knowledge work, which is improvements to knowledge work, which is improvements to knowledge work, which is cool and awesome if you really like cool and awesome if you really like cool and awesome if you really like using Excel with your models. I'm sure using Excel with your models. I'm sure using Excel with your models. I'm sure this is super important for a lot of you this is super important for a lot of you this is super important for a lot of you guys. Not a thing I particularly care guys. Not a thing I particularly care guys. Not a thing I particularly care for. Yeah, good numbers. Cool. Nice to for. Yeah, good numbers. Cool. Nice to for. Yeah, good numbers. Cool. Nice to see. Here is the thing I said I was most see. Here is the thing I said I was most see. Here is the thing I said I was most excited about, though. Communication. excited about, though. Communication. excited about, though. Communication. We've made major improvements to the way We've made major improvements to the way We've made major improvements to the way opus 5.5 writes and communicates. One of opus 5.5 writes and communicates. One of opus 5.5 writes and communicates. One of the most common areas of feedback we the most common areas of feedback we the most common areas of feedback we heard about opus 5. Yeah, please explain heard about opus 5. Yeah, please explain heard about opus 5. Yeah, please explain the issue to me. This is the opus 5 the issue to me. This is the opus 5 the issue to me. This is the opus 5 example. What I found the extra drop example. What I found the extra drop example. What I found the extra drop isn't the free tier mdash. It's a isn't the free tier mdash. It's a isn't the free tier mdash. It's a regression in random commit hash. The regression in random commit hash. The regression in random commit hash. The bug bunch of code used to do a halfopen bug bunch of code used to do a halfopen bug bunch of code used to do a halfopen interval. That is absolute nonsense. And interval. That is absolute nonsense. And interval. That is absolute nonsense. And there's three m dashes in that response there's three m dashes in that response there's three m dashes in that response by the way versus opus 5.5. The extra by the way versus opus 5.5. The extra by the way versus opus 5.5. The extra drop is a bug in the billing refactor.

  8. drop is a bug in the billing refactor. drop is a bug in the billing refactor. The free tier change accounts for only The free tier change accounts for only The free tier change accounts for only $1.50 of Acme's August drop. The other $1.50 of Acme's August drop. The other $1.50 of Acme's August drop. The other $9.92 come from a bug in this particular $9.92 come from a bug in this particular $9.92 come from a bug in this particular commit. That commit was labeled quote no commit. That commit was labeled quote no commit. That commit was labeled quote no behavior change, but it stops counting behavior change, but it stops counting behavior change, but it stops counting usage from the last day of the month. usage from the last day of the month. usage from the last day of the month. And then a clear what changed the before And then a clear what changed the before And then a clear what changed the before and after. It's just so much more and after. It's just so much more and after. It's just so much more readable. They have a handful of readable. They have a handful of readable. They have a handful of examples of this. And yeah, the examples of this. And yeah, the examples of this. And yeah, the difference is pretty immediate. I've difference is pretty immediate. I've difference is pretty immediate. I've already felt this since I started using already felt this since I started using already felt this since I started using it. I actually took the time to go it. I actually took the time to go it. I actually took the time to go delete my unslop stuff from my agents in delete my unslop stuff from my agents in delete my unslop stuff from my agents in CloudMD in order to see how this model CloudMD in order to see how this model CloudMD in order to see how this model just organically talks and I don't think just organically talks and I don't think just organically talks and I don't think I'm going to bother readding it cuz it's I'm going to bother readding it cuz it's I'm going to bother readding it cuz it's doing a great job. One last thing from doing a great job. One last thing from doing a great job. One last thing from the official article before we dive into the official article before we dive into the official article before we dive into all the much more interesting stuff all the much more interesting stuff all the much more interesting stuff people are doing with it. The data people are doing with it. The data people are doing with it. The data retention policy. For those who don't retention policy. For those who don't retention policy. For those who don't know, the Fable line is unique in that know, the Fable line is unique in that know, the Fable line is unique in that it doesn't offer proper zero data it doesn't offer proper zero data it doesn't offer proper zero data retention. So companies that need to retention. So companies that need to retention. So companies that need to have vendors that don't have any of have vendors that don't have any of have vendors that don't have any of their data saved cannot use Fable. Opus their data saved cannot use Fable. Opus their data saved cannot use Fable. Opus has stayed the number one model for has stayed the number one model for has stayed the number one model for businesses and also the number one businesses and also the number one businesses and also the number one coding model in general. Both because it coding model in general. Both because it coding model in general. Both because it represents a good price to performance represents a good price to performance represents a good price to performance for a lot of businesses, but more so for a lot of businesses, but more so for a lot of businesses, but more so because the ZDR policy can be applied on because the ZDR policy can be applied on because the ZDR policy can be applied on Opus and it couldn't be applied on Fable Opus and it couldn't be applied on Fable Opus and it couldn't be applied on Fable and Mythos. This means that those and Mythos. This means that those and Mythos. This means that those companies, the ones that spend all this companies, the ones that spend all this companies, the ones that spend all this money on AI, kind of couldn't enable money on AI, kind of couldn't enable money on AI, kind of couldn't enable Fable for their employees. Opus they can Fable for their employees. Opus they can Fable for their employees. Opus they can absolutely enable. So this model is absolutely enable. So this model is absolutely enable. So this model is almost immediately going to become the almost immediately going to become the almost immediately going to become the most popular for coding with AI by far most popular for coding with AI by far most popular for coding with AI by far just from this one particular detail.

  9. just from this one particular detail. just from this one particular detail. And it seems like y'all are going to And it seems like y'all are going to And it seems like y'all are going to have a pretty good experience. Opus 5.5 have a pretty good experience. Opus 5.5 have a pretty good experience. Opus 5.5 is way, way, way better than Opus 5. is way, way, way better than Opus 5. is way, way, way better than Opus 5. Sorry about that model. Please try this Sorry about that model. Please try this Sorry about that model. Please try this one. This is a fantastic tweet for a one. This is a fantastic tweet for a one. This is a fantastic tweet for a handful of reasons. I don't want to like handful of reasons. I don't want to like handful of reasons. I don't want to like overread social media etiquette stuff, overread social media etiquette stuff, overread social media etiquette stuff, but trust me, it's worth it. This is an but trust me, it's worth it. This is an but trust me, it's worth it. This is an acknowledgement that Opus 5 missed the acknowledgement that Opus 5 missed the acknowledgement that Opus 5 missed the mark. This is like a silly mark. This is like a silly mark. This is like a silly tongue-in-cheek way to communicate with tongue-in-cheek way to communicate with tongue-in-cheek way to communicate with the audience and users that I love. It's the audience and users that I love. It's the audience and users that I love. It's very personal. like clearly didn't get very personal. like clearly didn't get very personal. like clearly didn't get through like a normal PR review. But through like a normal PR review. But through like a normal PR review. But most importantly by far, this seems like most importantly by far, this seems like most importantly by far, this seems like Anthropic is loosening their very tight Anthropic is loosening their very tight Anthropic is loosening their very tight leash on what employees are allowed to leash on what employees are allowed to leash on what employees are allowed to do on the internet. Historically, it has do on the internet. Historically, it has do on the internet. Historically, it has seemed like anthropic employees are seemed like anthropic employees are seemed like anthropic employees are scared to comment on almost anything. scared to comment on almost anything. scared to comment on almost anything. I'm not feeling that at all with this I'm not feeling that at all with this I'm not feeling that at all with this release. It seems like everyone's just release. It seems like everyone's just release. It seems like everyone's just kind of allowed to go and post now, kind of allowed to go and post now, kind of allowed to go and post now, which is a very nice cultural change at which is a very nice cultural change at which is a very nice cultural change at Anthropic, and I really hope they Anthropic, and I really hope they Anthropic, and I really hope they maintain this. I know y'all think I'm an maintain this. I know y'all think I'm an maintain this. I know y'all think I'm an Anthropic fanboy now. Well, at least Sam Anthropic fanboy now. Well, at least Sam Anthropic fanboy now. Well, at least Sam Alman does. But for the rest of y'all Alman does. But for the rest of y'all Alman does. But for the rest of y'all who aren't familiar, I have had massive who aren't familiar, I have had massive who aren't familiar, I have had massive issues with how Anthropic runs their issues with how Anthropic runs their issues with how Anthropic runs their business for a while now. So, it's very business for a while now. So, it's very business for a while now. So, it's very relieved to see these changes to how relieved to see these changes to how relieved to see these changes to how they've been operating publicly. It's a they've been operating publicly. It's a they've been operating publicly. It's a nice shift. And this whole release kind nice shift. And this whole release kind nice shift. And this whole release kind of just feels that way. Like a nice of just feels that way. Like a nice of just feels that way. Like a nice shift in the right direction. I do want shift in the right direction. I do want shift in the right direction. I do want to talk about some more benchmarks in a to talk about some more benchmarks in a to talk about some more benchmarks in a second, but I want to give you a quick second, but I want to give you a quick second, but I want to give you a quick teaser of some of the fun 3D stuff that teaser of some of the fun 3D stuff that teaser of some of the fun 3D stuff that we'll be showing in the end as well, we'll be showing in the end as well, we'll be showing in the end as well, because I know y'all love these crazy because I know y'all love these crazy because I know y'all love these crazy flashy 3D demos. This is a clone of Dark flashy 3D demos. This is a clone of Dark flashy 3D demos. This is a clone of Dark Souls that was made by Matt's editor and

  10. Souls that was made by Matt's editor and Souls that was made by Matt's editor and assistant, Alex, here. That is pretty assistant, Alex, here. That is pretty assistant, Alex, here. That is pretty nuts. Like, the fact that models can do nuts. Like, the fact that models can do nuts. Like, the fact that models can do all of this is insane. I honestly would all of this is insane. I honestly would all of this is insane. I honestly would have called this demo fake if I hadn't have called this demo fake if I hadn't have called this demo fake if I hadn't gotten similar results for my own stuff, gotten similar results for my own stuff, gotten similar results for my own stuff, which trust me, we'll cover in a bit. which trust me, we'll cover in a bit. which trust me, we'll cover in a bit. Okay, I lied. One more detail before we Okay, I lied. One more detail before we Okay, I lied. One more detail before we get into benches. Opus 55 fixed one of get into benches. Opus 55 fixed one of get into benches. Opus 55 fixed one of the most annoying things about anthropic the most annoying things about anthropic the most annoying things about anthropic models, which is that the current effort models, which is that the current effort models, which is that the current effort level was at the top of the thread as level was at the top of the thread as level was at the top of the thread as part of the system prompt. This sounds part of the system prompt. This sounds part of the system prompt. This sounds like a silly thing to care about, but it like a silly thing to care about, but it like a silly thing to care about, but it means that you would bust cash whenever means that you would bust cash whenever means that you would bust cash whenever you changed your reasoning levels, which you changed your reasoning levels, which you changed your reasoning levels, which isn't great. Now you can change isn't great. Now you can change isn't great. Now you can change reasoning levels midsession without reasoning levels midsession without reasoning levels midsession without busting cash. Very nice change. Thank busting cash. Very nice change. Thank busting cash. Very nice change. Thank you to Lydia for calling this one out. I you to Lydia for calling this one out. I you to Lydia for calling this one out. I might have missed it otherwise. Very might have missed it otherwise. Very might have missed it otherwise. Very happy to see. With all of that said, happy to see. With all of that said, happy to see. With all of that said, it's time for some benchmarks. As you it's time for some benchmarks. As you it's time for some benchmarks. As you saw earlier, Opus 5.5 now has a massive saw earlier, Opus 5.5 now has a massive saw earlier, Opus 5.5 now has a massive lead on the artificial analysis lead on the artificial analysis lead on the artificial analysis intelligence index, a 5oint jump from intelligence index, a 5oint jump from intelligence index, a 5oint jump from Fable 5.1. For reference, the gap Fable 5.1. For reference, the gap Fable 5.1. For reference, the gap between Astra and Gro 47 is not much between Astra and Gro 47 is not much between Astra and Gro 47 is not much bigger. And the gap between Astra and bigger. And the gap between Astra and bigger. And the gap between Astra and Muse Spark is roughly the same. So if Muse Spark is roughly the same. So if Muse Spark is roughly the same. So if you think Fable is way better than you think Fable is way better than you think Fable is way better than Uspark, which it is, allegedly the gap Uspark, which it is, allegedly the gap Uspark, which it is, allegedly the gap from Fable to Opus is similarly sized in from Fable to Opus is similarly sized in from Fable to Opus is similarly sized in favor of Opus, which I think more so favor of Opus, which I think more so favor of Opus, which I think more so showcases how bad these benchmarks are.

  11. showcases how bad these benchmarks are. showcases how bad these benchmarks are. But it is worth noting that Opus 5.5 is But it is worth noting that Opus 5.5 is But it is worth noting that Opus 5.5 is slaughtering basically every bench you slaughtering basically every bench you slaughtering basically every bench you throw it at. It comes out at the top in throw it at. It comes out at the top in throw it at. It comes out at the top in almost everything. In particular, in almost everything. In particular, in almost everything. In particular, in output tokens, this is one of the sadder output tokens, this is one of the sadder output tokens, this is one of the sadder things that I think is important to call things that I think is important to call things that I think is important to call out. The price changes are real. This out. The price changes are real. This out. The price changes are real. This model is cheaper, but it's not because model is cheaper, but it's not because model is cheaper, but it's not because it's more efficient. In fact, this model it's more efficient. In fact, this model it's more efficient. In fact, this model is meaningfully less efficient than Opus is meaningfully less efficient than Opus is meaningfully less efficient than Opus 5. Opus 5 on Max did 73k tokens per 5. Opus 5 on Max did 73k tokens per 5. Opus 5 on Max did 73k tokens per task. Fable did 78k tokens and Opus 55 task. Fable did 78k tokens and Opus 55 task. Fable did 78k tokens and Opus 55 max did almost 120,000 tokens per task. max did almost 120,000 tokens per task. max did almost 120,000 tokens per task. For reference, GPT6 Astra did 27k tokens For reference, GPT6 Astra did 27k tokens For reference, GPT6 Astra did 27k tokens per task. Yes, there is really a 4x gap per task. Yes, there is really a 4x gap per task. Yes, there is really a 4x gap in efficiency here. And I thought that in efficiency here. And I thought that in efficiency here. And I thought that was worth calling out because it doesn't was worth calling out because it doesn't was worth calling out because it doesn't matter if you're two times faster if you matter if you're two times faster if you matter if you're two times faster if you need four times more tokens to generate. need four times more tokens to generate. need four times more tokens to generate. And it also closes some of that cost gap And it also closes some of that cost gap And it also closes some of that cost gap which is why we see the interesting which is why we see the interesting which is why we see the interesting numbers I had here in my comparison. numbers I had here in my comparison. numbers I had here in my comparison. Opus 5.5 comes close to being the most Opus 5.5 comes close to being the most Opus 5.5 comes close to being the most expensive model that artificial analysis expensive model that artificial analysis expensive model that artificial analysis has run simply because of how many has run simply because of how many has run simply because of how many tokens it is using. Opus 5 cost $5.86.

  12. tokens it is using. Opus 5 cost $5.86. tokens it is using. Opus 5 cost $5.86. Opus 5.5 cost almost six bucks and then Opus 5.5 cost almost six bucks and then Opus 5.5 cost almost six bucks and then Fable cost $763. Fable cost $763. Fable cost $763. But this is far from the whole picture But this is far from the whole picture But this is far from the whole picture because we're only looking at the max because we're only looking at the max because we're only looking at the max reasoning level for all of these. I made reasoning level for all of these. I made reasoning level for all of these. I made a subtle change here. I added the low, a subtle change here. I added the low, a subtle change here. I added the low, medium, and high for everything. The medium, and high for everything. The medium, and high for everything. The goal here is to show you that the costs goal here is to show you that the costs goal here is to show you that the costs aren't that brutal if you're not using aren't that brutal if you're not using aren't that brutal if you're not using max, which I will about even more max, which I will about even more max, which I will about even more in just a moment. You'll see that Opus in just a moment. You'll see that Opus in just a moment. You'll see that Opus 5.5 on XH high is a lot closer to 5.5 on XH high is a lot closer to 5.5 on XH high is a lot closer to Astra's costs. But if you're willing to Astra's costs. But if you're willing to Astra's costs. But if you're willing to bump down to high, it ends up being bump down to high, it ends up being bump down to high, it ends up being about half the price of Astra on Max. about half the price of Astra on Max. about half the price of Astra on Max. And when you go down to medium, it is And when you go down to medium, it is And when you go down to medium, it is way cheaper at $134 versus the 326 for way cheaper at $134 versus the 326 for way cheaper at $134 versus the 326 for Astra or the $6 for Opus 5.5. So yeah, Astra or the $6 for Opus 5.5. So yeah, Astra or the $6 for Opus 5.5. So yeah, the cost scales much better. I didn't the cost scales much better. I didn't the cost scales much better. I didn't put low in here because I think low is put low in here because I think low is put low in here because I think low is pretty stupid. It's just not that pretty stupid. It's just not that pretty stupid. It's just not that capable considering the price. What's capable considering the price. What's capable considering the price. What's much more interesting here is the Pareto much more interesting here is the Pareto much more interesting here is the Pareto Frontier line. This is the line for what Frontier line. This is the line for what Frontier line. This is the line for what is the best price at a given level of is the best price at a given level of is the best price at a given level of performance. Opus 5.5 has swept this performance. Opus 5.5 has swept this performance. Opus 5.5 has swept this line. It is the best price for all of line. It is the best price for all of line. It is the best price for all of these scores in the 50 plus range. I these scores in the 50 plus range. I these scores in the 50 plus range. I wish that artificial analysis made these wish that artificial analysis made these wish that artificial analysis made these visualizers a bit better. I've had to, visualizers a bit better. I've had to, visualizers a bit better. I've had to, as I mentioned before, rip it and make as I mentioned before, rip it and make as I mentioned before, rip it and make my own, which, as you can guess, I made my own, which, as you can guess, I made my own, which, as you can guess, I made with Opus 5.5. It was actually very with Opus 5.5. It was actually very with Opus 5.5. It was actually very pleasant to do this type of thing with.

  13. pleasant to do this type of thing with. pleasant to do this type of thing with. It was able to rip the data from It was able to rip the data from It was able to rip the data from Artificial Analysis's homepage and make Artificial Analysis's homepage and make Artificial Analysis's homepage and make a better visualizer in not very much a better visualizer in not very much a better visualizer in not very much time at all. But there is one little time at all. But there is one little time at all. But there is one little thing I want to sneak in here as well. I thing I want to sneak in here as well. I thing I want to sneak in here as well. I have other models. I'll throw Gro 47 and have other models. I'll throw Gro 47 and have other models. I'll throw Gro 47 and 46 in here so you can see why they 46 in here so you can see why they 46 in here so you can see why they aren't really relevant anymore because aren't really relevant anymore because aren't really relevant anymore because they are worse and more expensive than they are worse and more expensive than they are worse and more expensive than any other competition in the space right any other competition in the space right any other competition in the space right now. They are scoring below at higher now. They are scoring below at higher now. They are scoring below at higher costs consistently. What I wanted to costs consistently. What I wanted to costs consistently. What I wanted to sneak in here is Mimo V2.6 because this sneak in here is Mimo V2.6 because this sneak in here is Mimo V2.6 because this model is actually performing very well model is actually performing very well model is actually performing very well for the price and is a nice end to for the price and is a nice end to for the price and is a nice end to smooth out Opus 5's line here before you smooth out Opus 5's line here before you smooth out Opus 5's line here before you hit the super cheap Luna models. I hit the super cheap Luna models. I hit the super cheap Luna models. I switch over to log. You can see even switch over to log. You can see even switch over to log. You can see even more so how unique Mimo26 is in this more so how unique Mimo26 is in this more so how unique Mimo26 is in this chart. It seems like Xiai's been cooking chart. It seems like Xiai's been cooking chart. It seems like Xiai's been cooking on their new open weight line. So I'm on their new open weight line. So I'm on their new open weight line. So I'm very excited to dive more into the Mimo very excited to dive more into the Mimo very excited to dive more into the Mimo stuff in the near future. But for now, stuff in the near future. But for now, stuff in the near future. But for now, we're talking about Opus. So yeah, just we're talking about Opus. So yeah, just we're talking about Opus. So yeah, just wanted to call that one out quick cuz I wanted to call that one out quick cuz I wanted to call that one out quick cuz I thought it was cool. Artificial Analysis thought it was cool. Artificial Analysis thought it was cool. Artificial Analysis published a little breakdown for why the published a little breakdown for why the published a little breakdown for why the model is not much more expensive despite model is not much more expensive despite model is not much more expensive despite the fact that it's doing two times the the fact that it's doing two times the the fact that it's doing two times the number of tokens. They point out that it number of tokens. They point out that it number of tokens. They point out that it would have been 80% more expensive if would have been 80% more expensive if would have been 80% more expensive if they didn't decrease the price and also they didn't decrease the price and also they didn't decrease the price and also cut the cash read costs as well. Those cut the cash read costs as well. Those cut the cash read costs as well. Those two changes were able to keep it from two changes were able to keep it from two changes were able to keep it from being a massive price hike. I like this being a massive price hike. I like this being a massive price hike. I like this call out from Edwin which is similar to call out from Edwin which is similar to call out from Edwin which is similar to what I said before about the reasoning what I said before about the reasoning what I said before about the reasoning effort levels where medium is way more effort levels where medium is way more effort levels where medium is way more efficient. He says that medium on Opus efficient. He says that medium on Opus efficient. He says that medium on Opus 55 is Fable 5.1 level intelligence. I 55 is Fable 5.1 level intelligence. I 55 is Fable 5.1 level intelligence. I think that's a bit of a reach, but it's think that's a bit of a reach, but it's think that's a bit of a reach, but it's only $1.34 per task on artificial only $1.34 per task on artificial only $1.34 per task on artificial analysis versus the 6 plus for Fable. I analysis versus the 6 plus for Fable. I analysis versus the 6 plus for Fable. I think this release is really emphasizing think this release is really emphasizing think this release is really emphasizing my disdain for those max effort levels.

  14. my disdain for those max effort levels. my disdain for those max effort levels. They result in things just running in They result in things just running in They result in things just running in loops and never finishing and burning loops and never finishing and burning loops and never finishing and burning way more tokens than are needed for the way more tokens than are needed for the way more tokens than are needed for the vast vast majority of tasks. Believe it vast vast majority of tasks. Believe it vast vast majority of tasks. Believe it or not, even I got hit with this or not, even I got hit with this or not, even I got hit with this earlier. I had a thread go for 6 hours earlier. I had a thread go for 6 hours earlier. I had a thread go for 6 hours and 30 minutes writing a markdown plan and 30 minutes writing a markdown plan and 30 minutes writing a markdown plan because I had it on max and it just ran because I had it on max and it just ran because I had it on max and it just ran in loops forever. I ended up having to in loops forever. I ended up having to in loops forever. I ended up having to stop it and then switch to XH high and stop it and then switch to XH high and stop it and then switch to XH high and ask for an update. And apparently it ask for an update. And apparently it ask for an update. And apparently it only got halfway done. So it will only got halfway done. So it will only got halfway done. So it will hopefully be able to finish in the near hopefully be able to finish in the near hopefully be able to finish in the near future. Sadly probably won't have that future. Sadly probably won't have that future. Sadly probably won't have that demo ready for this video, but I have demo ready for this video, but I have demo ready for this video, but I have plenty of others. Don't worry. Just as plenty of others. Don't worry. Just as plenty of others. Don't worry. Just as one last piece of evidence for my one last piece of evidence for my one last piece of evidence for my disdain towards the max reasoning disdain towards the max reasoning disdain towards the max reasoning effort, I ran this model through effort, I ran this model through effort, I ran this model through Skatebench alongside GPT6 Soul and Luna. Skatebench alongside GPT6 Soul and Luna. Skatebench alongside GPT6 Soul and Luna. And the thing I care about with this And the thing I care about with this And the thing I care about with this isn't how high the score was. It was isn't how high the score was. It was isn't how high the score was. It was fine. It was an improvement, but not a fine. It was an improvement, but not a fine. It was an improvement, but not a massive one. Actually, no, that's a lie. massive one. Actually, no, that's a lie. massive one. Actually, no, that's a lie. It actually scored worse than Opus 5 did It actually scored worse than Opus 5 did It actually scored worse than Opus 5 did on Skatebench. Interesting. Well, what I on Skatebench. Interesting. Well, what I on Skatebench. Interesting. Well, what I found much more intriguing with this is found much more intriguing with this is found much more intriguing with this is the gap between Xigh and Max's amount of the gap between Xigh and Max's amount of the gap between Xigh and Max's amount of tokens generated per question. On XH tokens generated per question. On XH tokens generated per question. On XH high, it only used about 330 reasoning high, it only used about 330 reasoning high, it only used about 330 reasoning tokens per question. And from X he high tokens per question. And from X he high tokens per question. And from X he high to max, which is just a one little bump to max, which is just a one little bump to max, which is just a one little bump inside of your slider, it went from inside of your slider, it went from inside of your slider, it went from 300ish tokens per task to almost 5,000 300ish tokens per task to almost 5,000 300ish tokens per task to almost 5,000 per task. I'm not exaggerating. That's per task. I'm not exaggerating. That's per task. I'm not exaggerating. That's over a 10x gap for that tiny little over a 10x gap for that tiny little over a 10x gap for that tiny little jump. And you get almost no improvement jump. And you get almost no improvement jump. And you get almost no improvement in score here at all. It's a 1% bump.

  15. in score here at all. It's a 1% bump. in score here at all. It's a 1% bump. And I still have two of these running And I still have two of these running And I still have two of these running because it just takes forever cuz it's because it just takes forever cuz it's because it just takes forever cuz it's burning so many tokens. It is a more burning so many tokens. It is a more burning so many tokens. It is a more than 10x increase in your costs for no than 10x increase in your costs for no than 10x increase in your costs for no good reason. So yeah, don't use good reason. So yeah, don't use good reason. So yeah, don't use Max on this model. It's not worth it Max on this model. It's not worth it Max on this model. It's not worth it unless you just want to watch tokens unless you just want to watch tokens unless you just want to watch tokens burn and be lit on fire. And I know rich burn and be lit on fire. And I know rich burn and be lit on fire. And I know rich coming from me, the guy who loves coming from me, the guy who loves coming from me, the guy who loves talking about their token furnace, but talking about their token furnace, but talking about their token furnace, but this model has been very pleasant in this model has been very pleasant in this model has been very pleasant in that regard. As I mentioned, I've been that regard. As I mentioned, I've been that regard. As I mentioned, I've been going ham on it all day. And remember, I going ham on it all day. And remember, I going ham on it all day. And remember, I have five Claude subs because I really have five Claude subs because I really have five Claude subs because I really like Fable 5.1. Let's refresh to see how like Fable 5.1. Let's refresh to see how like Fable 5.1. Let's refresh to see how my usage is after heavy usage all day. my usage is after heavy usage all day. my usage is after heavy usage all day. Yeah, not much. One of my accounts has Yeah, not much. One of my accounts has Yeah, not much. One of my accounts has been hit for the majority of it. This been hit for the majority of it. This been hit for the majority of it. This account got a little bit more because I account got a little bit more because I account got a little bit more because I was using it in the Cloud Code desktop was using it in the Cloud Code desktop was using it in the Cloud Code desktop app. But the vast majority has been in app. But the vast majority has been in app. But the vast majority has been in this here, which was only about 40% of this here, which was only about 40% of this here, which was only about 40% of my weekly, except for the fact that my weekly, except for the fact that my weekly, except for the fact that about half that weekly usage came from about half that weekly usage came from about half that weekly usage came from me using Fable over the last 2 days. So, me using Fable over the last 2 days. So, me using Fable over the last 2 days. So, it's actually closer to like 20ish% of it's actually closer to like 20ish% of it's actually closer to like 20ish% of my weekly, which is crazy considering my weekly, which is crazy considering my weekly, which is crazy considering how I've been going out of my way to how I've been going out of my way to how I've been going out of my way to push this model all day. I was able to push this model all day. I was able to push this model all day. I was able to get some real work done with it. I made get some real work done with it. I made get some real work done with it. I made that new visualizer that I showed that new visualizer that I showed that new visualizer that I showed earlier. I ported it multiple times to earlier. I ported it multiple times to earlier. I ported it multiple times to different places that I wanted to host different places that I wanted to host different places that I wanted to host it with additional functionality. I had it with additional functionality. I had it with additional functionality. I had it use computer use to verify it and it use computer use to verify it and it use computer use to verify it and also get the data out of the original also get the data out of the original also get the data out of the original artificial analysis site. I noticed a artificial analysis site. I noticed a artificial analysis site. I noticed a bug in Lakebed my cloud when I was bug in Lakebed my cloud when I was bug in Lakebed my cloud when I was working with this. So I told it to go working with this. So I told it to go working with this. So I told it to go through my computer, find the repo, and through my computer, find the repo, and through my computer, find the repo, and make a PR. And it was able to find it, make a PR. And it was able to find it, make a PR. And it was able to find it, realize how out of date my local clone realize how out of date my local clone realize how out of date my local clone was, update it, make a work tree, file was, update it, make a work tree, file was, update it, make a work tree, file the PR with the fixes, verify the fixes,

  16. the PR with the fixes, verify the fixes, the PR with the fixes, verify the fixes, and get it all merged for me. All of and get it all merged for me. All of and get it all merged for me. All of this combined was 1% of my weekly. Yeah, this combined was 1% of my weekly. Yeah, this combined was 1% of my weekly. Yeah, they fixed the problem. I can't believe they fixed the problem. I can't believe they fixed the problem. I can't believe I'm saying this, but if this model does I'm saying this, but if this model does I'm saying this, but if this model does turn out to be a solid daily driver, turn out to be a solid daily driver, turn out to be a solid daily driver, which it seems like it will, it is a which it seems like it will, it is a which it seems like it will, it is a massive improvement to the value you get massive improvement to the value you get massive improvement to the value you get compared to something like the codec compared to something like the codec compared to something like the codec sub. It is a little bit too easy to burn sub. It is a little bit too easy to burn sub. It is a little bit too easy to burn through your codec plans right now. And through your codec plans right now. And through your codec plans right now. And since the only options you have are since the only options you have are since the only options you have are Astra, which is capable but spiky and Astra, which is capable but spiky and Astra, which is capable but spiky and way too expensive, or Soul, which we'll way too expensive, or Soul, which we'll way too expensive, or Soul, which we'll be talking about in a new video soon, be talking about in a new video soon, be talking about in a new video soon, this is a way better value for your 200 this is a way better value for your 200 this is a way better value for your 200 bucks right now. Kind of crazy that this bucks right now. Kind of crazy that this bucks right now. Kind of crazy that this happened so quickly, but yeah, it did. happened so quickly, but yeah, it did. happened so quickly, but yeah, it did. I've only used my codec subs lately I've only used my codec subs lately I've only used my codec subs lately today to play with the new GBD6 Soul today to play with the new GBD6 Soul today to play with the new GBD6 Soul release. Not really reaching for it that release. Not really reaching for it that release. Not really reaching for it that much right now. I am just really much right now. I am just really much right now. I am just really impressed with Opus 55 for everything impressed with Opus 55 for everything impressed with Opus 55 for everything I've been doing with it today. I've been I've been doing with it today. I've been I've been doing with it today. I've been pushing this model hard to find its pushing this model hard to find its pushing this model hard to find its weaknesses. And thus far, what I found weaknesses. And thus far, what I found weaknesses. And thus far, what I found is that its weaknesses are similar to is that its weaknesses are similar to is that its weaknesses are similar to those of Fable. I mentioned yesterday in those of Fable. I mentioned yesterday in those of Fable. I mentioned yesterday in my Grock video that I made a bench that my Grock video that I made a bench that my Grock video that I made a bench that was trying to figure out how well models was trying to figure out how well models was trying to figure out how well models can dive into big code bases and make can dive into big code bases and make can dive into big code bases and make realworld suggestions for improving the realworld suggestions for improving the realworld suggestions for improving the quality of the code. And that bench quality of the code. And that bench quality of the code. And that bench actually showed Grock 47 massively actually showed Grock 47 massively actually showed Grock 47 massively outperforming relative to what it is and outperforming relative to what it is and outperforming relative to what it is and its capabilities. because I have its capabilities. because I have its capabilities. because I have actually found Grock 47 to be pretty actually found Grock 47 to be pretty actually found Grock 47 to be pretty good at being thorough with code, good at being thorough with code, good at being thorough with code, reviewing the hell out of it and diving reviewing the hell out of it and diving reviewing the hell out of it and diving deep. My suspicion for why Grock is so deep. My suspicion for why Grock is so deep. My suspicion for why Grock is so good at this is all the effort that good at this is all the effort that good at this is all the effort that cursors put into code review pipelines cursors put into code review pipelines cursors put into code review pipelines with things like Bugbot and all the data with things like Bugbot and all the data with things like Bugbot and all the data they have from that. That data is super

  17. they have from that. That data is super they have from that. That data is super useful when you're trying to fine-tune useful when you're trying to fine-tune useful when you're trying to fine-tune models to be better at digging into the models to be better at digging into the models to be better at digging into the details and reviewing code. I suspect details and reviewing code. I suspect details and reviewing code. I suspect that is why they're able to make Rock so that is why they're able to make Rock so that is why they're able to make Rock so uniquely good at this compared to other uniquely good at this compared to other uniquely good at this compared to other models. Soul also seems to be very good models. Soul also seems to be very good models. Soul also seems to be very good at this as well. Opus and Fable less. at this as well. Opus and Fable less. at this as well. Opus and Fable less. So, I will say the gap from Opus 5 to So, I will say the gap from Opus 5 to So, I will say the gap from Opus 5 to 5.5 in this bench is absolutely 5.5 in this bench is absolutely 5.5 in this bench is absolutely hilarious, nearly doubling in the hilarious, nearly doubling in the hilarious, nearly doubling in the accuracy and reliability of its accuracy and reliability of its accuracy and reliability of its findings. And Gemini 38 Flash is still findings. And Gemini 38 Flash is still findings. And Gemini 38 Flash is still meme tier as expected. So, if you're meme tier as expected. So, if you're meme tier as expected. So, if you're planning on using this model to very planning on using this model to very planning on using this model to very thoroughly review code and find every thoroughly review code and find every thoroughly review code and find every single thing wrong with it or suggest single thing wrong with it or suggest single thing wrong with it or suggest sweeping improvements to your codebase, sweeping improvements to your codebase, sweeping improvements to your codebase, might not be as strong at those things might not be as strong at those things might not be as strong at those things as OpenAI's Frontier is. But, I'm going as OpenAI's Frontier is. But, I'm going as OpenAI's Frontier is. But, I'm going to be real with you guys. This is not a to be real with you guys. This is not a to be real with you guys. This is not a task I do every day. This is a task I do task I do every day. This is a task I do task I do every day. This is a task I do once every two or so weeks, mostly as a once every two or so weeks, mostly as a once every two or so weeks, mostly as a way to test new models. It is super way to test new models. It is super way to test new models. It is super useful to have a model that can deeply useful to have a model that can deeply useful to have a model that can deeply dig into every detail in your codebase, dig into every detail in your codebase, dig into every detail in your codebase, poke every hole there is to poke, and poke every hole there is to poke, and poke every hole there is to poke, and give you real feedback on how to improve give you real feedback on how to improve give you real feedback on how to improve it, but it is not the default case and it, but it is not the default case and it, but it is not the default case and it is not the most realistic thing to it is not the most realistic thing to it is not the most realistic thing to bench. I just thought these numbers were bench. I just thought these numbers were bench. I just thought these numbers were interesting and it was one of the few interesting and it was one of the few interesting and it was one of the few things I could do and find that showed things I could do and find that showed things I could do and find that showed the weakness that I perceive with the weakness that I perceive with the weakness that I perceive with anthropic models compared to OpenAI.

  18. anthropic models compared to OpenAI. anthropic models compared to OpenAI. Still got a bunch of fun things to cover Still got a bunch of fun things to cover Still got a bunch of fun things to cover including the model's front-end including the model's front-end including the model's front-end capabilities and most importantly all capabilities and most importantly all capabilities and most importantly all the crazy demos people have been making the crazy demos people have been making the crazy demos people have been making with it. This model is insane at 3D and with it. This model is insane at 3D and with it. This model is insane at 3D and the stuff I've been seeing is the stuff I've been seeing is the stuff I've been seeing is mindblowing. Normally I would make a mindblowing. Normally I would make a mindblowing. Normally I would make a joke here about how expensive it was to joke here about how expensive it was to joke here about how expensive it was to generate all of these tests and all of generate all of these tests and all of generate all of these tests and all of these examples, but this model has not these examples, but this model has not these examples, but this model has not been very expensive for me. Regardless, been very expensive for me. Regardless, been very expensive for me. Regardless, I hope you can pardon me for a real I hope you can pardon me for a real I hope you can pardon me for a real quick sponsor break. I have two quick quick sponsor break. I have two quick quick sponsor break. I have two quick questions. First, have you ever used an questions. First, have you ever used an questions. First, have you ever used an agent without search? If you have, you agent without search? If you have, you agent without search? If you have, you know how painful and miserable it is. It know how painful and miserable it is. It know how painful and miserable it is. It basically can't do anything. My second basically can't do anything. My second basically can't do anything. My second question is the opposite. Have you ever question is the opposite. Have you ever question is the opposite. Have you ever used an agent with insanely fast and used an agent with insanely fast and used an agent with insanely fast and accurate search? I personally hadn't accurate search? I personally hadn't accurate search? I personally hadn't until I started using today's sponsor, until I started using today's sponsor, until I started using today's sponsor, Parallel, because they have the best and Parallel, because they have the best and Parallel, because they have the best and fastest search results for pretty much fastest search results for pretty much fastest search results for pretty much every single thing you can measure. every single thing you can measure. every single thing you can measure. Historically, I've really liked the Historically, I've really liked the Historically, I've really liked the search built into OpenAI, but when you search built into OpenAI, but when you search built into OpenAI, but when you compare it to Parallel, it's just night compare it to Parallel, it's just night compare it to Parallel, it's just night and day. Watch and see just how fast and day. Watch and see just how fast and day. Watch and see just how fast Parallel can get you good results. Going Parallel can get you good results. Going Parallel can get you good results. Going now, all real time, of course. It took now, all real time, of course. It took now, all real time, of course. It took just over a second for parallel. And just over a second for parallel. And just over a second for parallel. And OpenAI still going. Still going. OpenAI still going. Still going. OpenAI still going. Still going. Meanwhile, OpenAI's endpoint took over 6 Meanwhile, OpenAI's endpoint took over 6 Meanwhile, OpenAI's endpoint took over 6 seconds long. If search was all they seconds long. If search was all they seconds long. If search was all they did, they'd be one of the best options did, they'd be one of the best options did, they'd be one of the best options available. But they also have everything available. But they also have everything available. But they also have everything else you would need on the web. from a else you would need on the web. from a else you would need on the web. from a proper monitor that will send your proper monitor that will send your proper monitor that will send your agents info when things change on agents info when things change on agents info when things change on different pages to a traditional different pages to a traditional different pages to a traditional response API for when you want to response API for when you want to response API for when you want to actually get synthesized results from actually get synthesized results from actually get synthesized results from your search queries to their extract your search queries to their extract your search queries to their extract endpoints that let you send a URL and endpoints that let you send a URL and endpoints that let you send a URL and get back the data that your agents get back the data that your agents get back the data that your agents actually want and need. Even parsing actually want and need. Even parsing actually want and need. Even parsing JSheavy pages, by the way, there's even JSheavy pages, by the way, there's even JSheavy pages, by the way, there's even an MCP so you can expose Parallel to an MCP so you can expose Parallel to an MCP so you can expose Parallel to your existing agents in whatever tool your existing agents in whatever tool your existing agents in whatever tool you're using now. Every month you'll get you're using now. Every month you'll get you're using now. Every month you'll get 5,000 requests for free. And on top of 5,000 requests for free. And on top of 5,000 requests for free. And on top of that, if you sign up today, you'll get that, if you sign up today, you'll get that, if you sign up today, you'll get $80 in credit. What are you waiting for?

  19. $80 in credit. What are you waiting for? $80 in credit. What are you waiting for? Join now at soyv.link/parallel. Join now at soyv.link/parallel. Join now at soyv.link/parallel. Next, we need to talk a bit about front Next, we need to talk a bit about front Next, we need to talk a bit about front end. This one's interesting cuz the end. This one's interesting cuz the end. This one's interesting cuz the model does seem to have really good model does seem to have really good model does seem to have really good design instincts. I just don't think design instincts. I just don't think design instincts. I just don't think it's showing them very well in its it's showing them very well in its it's showing them very well in its front-end design capabilities. I was front-end design capabilities. I was front-end design capabilities. I was floored with how much better Fable 5.1 floored with how much better Fable 5.1 floored with how much better Fable 5.1 was at subtle design taste and like was at subtle design taste and like was at subtle design taste and like getting the little animations and getting the little animations and getting the little animations and details right in things that you details right in things that you details right in things that you designed with it. I have not had the designed with it. I have not had the designed with it. I have not had the same experience with Opus 5.5. I find same experience with Opus 5.5. I find same experience with Opus 5.5. I find most of these demos to be not sloppy, most of these demos to be not sloppy, most of these demos to be not sloppy, but okay. Yeah, they're a little sloppy but okay. Yeah, they're a little sloppy but okay. Yeah, they're a little sloppy if I'm being real. These are all the if I'm being real. These are all the if I'm being real. These are all the ones using the official Claude Code ones using the official Claude Code ones using the official Claude Code design skill. And if we turn off the design skill. And if we turn off the design skill. And if we turn off the skill, it's a little bit better, but not skill, it's a little bit better, but not skill, it's a little bit better, but not a lot better. Still kind of feels like a lot better. Still kind of feels like a lot better. Still kind of feels like half of them are Tailwind templates, the half of them are Tailwind templates, the half of them are Tailwind templates, the other half are this very specific like other half are this very specific like other half are this very specific like colorful colorful colorful I don't even know what the term for the I don't even know what the term for the I don't even know what the term for the style is, but like the color pastel with style is, but like the color pastel with style is, but like the color pastel with the big shadows and everything. Not my the big shadows and everything. Not my the big shadows and everything. Not my thing at all. And again to compare with thing at all. And again to compare with thing at all. And again to compare with Fable, which I think was the best design Fable, which I think was the best design Fable, which I think was the best design model by far recently, you can see even model by far recently, you can see even model by far recently, you can see even on the similar designs like this station on the similar designs like this station on the similar designs like this station one, the little animations and the taste one, the little animations and the taste one, the little animations and the taste and the page layout is meaningfully and the page layout is meaningfully and the page layout is meaningfully better. So that shows Fable still is the better. So that shows Fable still is the better. So that shows Fable still is the GOAT for a handful of things. I will GOAT for a handful of things. I will GOAT for a handful of things. I will know more as I use Opus 55 more for know more as I use Opus 55 more for know more as I use Opus 55 more for actual day-to-day work, but right now actual day-to-day work, but right now actual day-to-day work, but right now gut feel is there are definitely some gut feel is there are definitely some gut feel is there are definitely some edges where Fable is stronger. Fable is edges where Fable is stronger. Fable is edges where Fable is stronger. Fable is still slightly more thorough. Astra is still slightly more thorough. Astra is still slightly more thorough. Astra is still the most thorough, but it's also still the most thorough, but it's also still the most thorough, but it's also very spiky and weird. I think Opus is very spiky and weird. I think Opus is very spiky and weird. I think Opus is probably the right default for the vast probably the right default for the vast probably the right default for the vast majority of stuff. But you should not be

  20. majority of stuff. But you should not be majority of stuff. But you should not be as scared to reach to Fable for front as scared to reach to Fable for front as scared to reach to Fable for front end still, especially when you're doing end still, especially when you're doing end still, especially when you're doing like a marketing page or some big design like a marketing page or some big design like a marketing page or some big design overhaul. I still think Fable is a overhaul. I still think Fable is a overhaul. I still think Fable is a little bit worth it. That said, others little bit worth it. That said, others little bit worth it. That said, others are having a different experience. For are having a different experience. For are having a different experience. For example, Mia, who doesn't necessarily example, Mia, who doesn't necessarily example, Mia, who doesn't necessarily always agree with me on these types of always agree with me on these types of always agree with me on these types of things, she did a hundred HTML page things, she did a hundred HTML page things, she did a hundred HTML page generations with Opus 5.5 and said very generations with Opus 5.5 and said very generations with Opus 5.5 and said very adamantly that it is the best model that adamantly that it is the best model that adamantly that it is the best model that she's ever tested. And if I just open she's ever tested. And if I just open she's ever tested. And if I just open like the first three here, you'll see, like the first three here, you'll see, like the first three here, you'll see, yeah, this is crazy. This like paper yeah, this is crazy. This like paper yeah, this is crazy. This like paper style layout where things are very like style layout where things are very like style layout where things are very like real physical objecty. Like this is real physical objecty. Like this is real physical objecty. Like this is super skumorphic with the paper feeling super skumorphic with the paper feeling super skumorphic with the paper feeling recut. recut. recut. I was just showing you the animation I was just showing you the animation I was just showing you the animation coming in again. That's pretty cool. coming in again. That's pretty cool. coming in again. That's pretty cool. Yeah, this is stunning for a thing of Yeah, this is stunning for a thing of Yeah, this is stunning for a thing of model just one shot as part of a hundred model just one shot as part of a hundred model just one shot as part of a hundred generations or this neon rain which generations or this neon rain which generations or this neon rain which while still a little too uh while still a little too uh while still a little too uh you know that one terminal retro theme you know that one terminal retro theme you know that one terminal retro theme that the skill recommends, still feels a that the skill recommends, still feels a that the skill recommends, still feels a little bit like that. and it made a game little bit like that. and it made a game little bit like that. and it made a game of chess, which I would play, but it of chess, which I would play, but it of chess, which I would play, but it would embarrass me too much. But you get would embarrass me too much. But you get would embarrass me too much. But you get the idea. It makes decent pages. I got the idea. It makes decent pages. I got the idea. It makes decent pages. I got DM'd this fun demo from Max right before DM'd this fun demo from Max right before DM'd this fun demo from Max right before getting ready to film where he actually getting ready to film where he actually getting ready to film where he actually ported all of T3 code into Minecraft, ported all of T3 code into Minecraft, ported all of T3 code into Minecraft, which is hilarious and quite cool that which is hilarious and quite cool that which is hilarious and quite cool that he can be playing Minecraft and not have he can be playing Minecraft and not have he can be playing Minecraft and not have to leave the window to see how his to leave the window to see how his to leave the window to see how his agents are doing. This isn't a direct agents are doing. This isn't a direct agents are doing. This isn't a direct demo, but I think it showcases how demo, but I think it showcases how demo, but I think it showcases how strong the model is. Gabriel was one of strong the model is. Gabriel was one of strong the model is. Gabriel was one of the lead devs and researchers on Sora at

  21. the lead devs and researchers on Sora at the lead devs and researchers on Sora at OpenAI. He somewhat recently left to OpenAI. He somewhat recently left to OpenAI. He somewhat recently left to start his own company. He was hyping up start his own company. He was hyping up start his own company. He was hyping up all the stuff OpenAI was about to all the stuff OpenAI was about to all the stuff OpenAI was about to release and called out that Anthropic release and called out that Anthropic release and called out that Anthropic feels like you need to wait for a week feels like you need to wait for a week feels like you need to wait for a week to see if there was any catastrophic to see if there was any catastrophic to see if there was any catastrophic problem with the new models. He followed problem with the new models. He followed problem with the new models. He followed up a few hours later saying, "Never up a few hours later saying, "Never up a few hours later saying, "Never mind. Opus 55 is crazy good." Yeah. mind. Opus 55 is crazy good." Yeah. mind. Opus 55 is crazy good." Yeah. Next, we have that set of demos from Next, we have that set of demos from Next, we have that set of demos from Alex, who I really do owe the fallback. Alex, who I really do owe the fallback. Alex, who I really do owe the fallback. Sorry about that. First, he showed a Sorry about that. First, he showed a Sorry about that. First, he showed a Dark Souls demo that Opus made, which is Dark Souls demo that Opus made, which is Dark Souls demo that Opus made, which is a clone of a very popular game that a a clone of a very popular game that a a clone of a very popular game that a lot of people like to yell at. Obviously lot of people like to yell at. Obviously lot of people like to yell at. Obviously nowhere near as fine-tuned and nowhere near as fine-tuned and nowhere near as fine-tuned and welldesigned as the official games from welldesigned as the official games from welldesigned as the official games from From Software. That was a mouthful and From Software. That was a mouthful and From Software. That was a mouthful and tongue twister, but it's pretty nuts it tongue twister, but it's pretty nuts it tongue twister, but it's pretty nuts it can make something like this. I made a can make something like this. I made a can make something like this. I made a flight simulator as well for him that is flight simulator as well for him that is flight simulator as well for him that is uh a lot better than the previous flight uh a lot better than the previous flight uh a lot better than the previous flight sims I've seen people making with these sims I've seen people making with these sims I've seen people making with these models. Kind of nuts that we've made models. Kind of nuts that we've made models. Kind of nuts that we've made this much progress so fast. Like the 3D this much progress so fast. Like the 3D this much progress so fast. Like the 3D demos we saw even just a few months ago demos we saw even just a few months ago demos we saw even just a few months ago weren't even close to what people are weren't even close to what people are weren't even close to what people are making nowadays. sparks [music] sparks [music] in your eyes.

  22. in your eyes. in your eyes. >> I'm so sorry. The music's too cringe for >> I'm so sorry. The music's too cringe for >> I'm so sorry. The music's too cringe for me. I can't do the AI generated music. me. I can't do the AI generated music. me. I can't do the AI generated music. It's just not my thing. But this whole It's just not my thing. But this whole It's just not my thing. But this whole animation was designed with Claude. animation was designed with Claude. animation was designed with Claude. Apparently, it was able to make it and Apparently, it was able to make it and Apparently, it was able to make it and it actually looks very good. There's a it actually looks very good. There's a it actually looks very good. There's a lot of taste for these types of like 2D, lot of taste for these types of like 2D, lot of taste for these types of like 2D, 2.5D, and even sometimes 3D animation 2.5D, and even sometimes 3D animation 2.5D, and even sometimes 3D animation stuff, which is a a real challenge to stuff, which is a a real challenge to stuff, which is a a real challenge to get right. Bjan Bowen made some really get right. Bjan Bowen made some really get right. Bjan Bowen made some really cool demos as well. He had early access, cool demos as well. He had early access, cool demos as well. He had early access, which I didn't. Anthropic, you know how which I didn't. Anthropic, you know how which I didn't. Anthropic, you know how to get a hold of me. He was able to make to get a hold of me. He was able to make to get a hold of me. He was able to make a bunch of stuff ahead of time and some a bunch of stuff ahead of time and some a bunch of stuff ahead of time and some of these demos are super cool. He of these demos are super cool. He of these demos are super cool. He actually took the time to make a Tony actually took the time to make a Tony actually took the time to make a Tony Hawk Pro Skater style clone fully from Hawk Pro Skater style clone fully from Hawk Pro Skater style clone fully from scratch in C++. So, not even using a scratch in C++. So, not even using a scratch in C++. So, not even using a game engine apparently. And uh yeah, game engine apparently. And uh yeah, game engine apparently. And uh yeah, even if the model didn't do great on even if the model didn't do great on even if the model didn't do great on Skate Bench, seems to do pretty well on Skate Bench, seems to do pretty well on Skate Bench, seems to do pretty well on Skate Game Bench. >> And the the way the skateboard actually >> And the the way the skateboard actually goes away when you bail. Obviously, this is still far from like Obviously, this is still far from like perfect, but it's insane how much perfect, but it's insane how much perfect, but it's insane how much progress we've been making as an progress we've been making as an progress we've been making as an industry. It also seems weirdly good at industry. It also seems weirdly good at industry. It also seems weirdly good at FPS's FPS's FPS's >> is coming down the just >> is coming down the just >> is coming down the just good. Something is coming down the good. Something is coming down the good. Something is coming down the stairs.

  23. stairs. stairs. >> All of these are web demos. >> All of these are web demos. >> All of these are web demos. >> The skate game was a C++ native game, >> The skate game was a C++ native game, >> The skate game was a C++ native game, but this one's a web demo, and it seems but this one's a web demo, and it seems but this one's a web demo, and it seems much better at things like 3JS than much better at things like 3JS than much better at things like 3JS than native tech. So far, this demo is native tech. So far, this demo is native tech. So far, this demo is entirely in 3JS in the browser compared entirely in 3JS in the browser compared entirely in 3JS in the browser compared to the previous one which was a C++ to the previous one which was a C++ to the previous one which was a C++ native demo with no engine. It does seem native demo with no engine. It does seem native demo with no engine. It does seem to be much stronger when you throw it in to be much stronger when you throw it in to be much stronger when you throw it in 3JS in the browser. 3JS in the browser. 3JS in the browser. >> It's not it's like a construction light, >> It's not it's like a construction light, >> It's not it's like a construction light, >> but it does get a lot of these weird >> but it does get a lot of these weird >> but it does get a lot of these weird like flickery edges. I found it like flickery edges. I found it like flickery edges. I found it surprisingly capable of using computer surprisingly capable of using computer surprisingly capable of using computer use to identify these things and fix use to identify these things and fix use to identify these things and fix them, but not always. Shout out to Bjon them, but not always. Shout out to Bjon them, but not always. Shout out to Bjon for letting me use his examples for for letting me use his examples for for letting me use his examples for this. Really appreciate that. But now this. Really appreciate that. But now this. Really appreciate that. But now it's time for my demo. Fish slop. it's time for my demo. Fish slop. it's time for my demo. Fish slop. That doesn't look great. Oh yeah, that's That doesn't look great. Oh yeah, that's That doesn't look great. Oh yeah, that's the Gro 4.7 version. My bad. The first the Gro 4.7 version. My bad. The first the Gro 4.7 version. My bad. The first thing I noticed when I opened this thing I noticed when I opened this thing I noticed when I opened this version is how well it functions. The version is how well it functions. The version is how well it functions. The movement is great. The frame rate's movement is great. The frame rate's movement is great. The frame rate's insane. It's running at like an actually insane. It's running at like an actually insane. It's running at like an actually perfectly smooth 120 FPS on my machine, perfectly smooth 120 FPS on my machine, perfectly smooth 120 FPS on my machine, and the models look pretty solid. I was and the models look pretty solid. I was and the models look pretty solid. I was impressed with this immediately impressed with this immediately impressed with this immediately and then I realized I had made a and then I realized I had made a and then I realized I had made a mistake. I was running on medium mistake. I was running on medium mistake. I was running on medium reasoning. Oops, my bad. I was trying to reasoning. Oops, my bad. I was trying to reasoning. Oops, my bad. I was trying to get this all together quick and this is get this all together quick and this is get this all together quick and this is the default reasoning level for the the default reasoning level for the the default reasoning level for the model. So, understandable mistake, I model. So, understandable mistake, I model. So, understandable mistake, I hope. So, what I did instead after is I hope. So, what I did instead after is I hope. So, what I did instead after is I told it to refine and make it as told it to refine and make it as told it to refine and make it as visually appealing as possible all on a visually appealing as possible all on a visually appealing as possible all on a much higher reasoning level on X high.

  24. much higher reasoning level on X high. much higher reasoning level on X high. And this is what it did on top of that And this is what it did on top of that And this is what it did on top of that initial demo. Yeah, it's all over. Yeah, it's all over. I did not expect it to be even close to I did not expect it to be even close to I did not expect it to be even close to this good. I really thought OpenAI would this good. I really thought OpenAI would this good. I really thought OpenAI would maintain their lead for a while here, maintain their lead for a while here, maintain their lead for a while here, but look at the stingray. Do you know but look at the stingray. Do you know but look at the stingray. Do you know how hard it is to design something like how hard it is to design something like how hard it is to design something like that in Blender and then get the that in Blender and then get the that in Blender and then get the animation right with that many polygons animation right with that many polygons animation right with that many polygons to rig? This is not trivial. It did a to rig? This is not trivial. It did a to rig? This is not trivial. It did a great job. Obviously, there's some great job. Obviously, there's some great job. Obviously, there's some issues with like the lighting, the way issues with like the lighting, the way issues with like the lighting, the way it's like blending the light wrong when it's like blending the light wrong when it's like blending the light wrong when it moves its wings there, but like it moves its wings there, but like it moves its wings there, but like what the And the craziest thing is that like not And the craziest thing is that like not only is the movement and like the core only is the movement and like the core only is the movement and like the core mechanics like flying around in the mechanics like flying around in the mechanics like flying around in the water and everything good, they got the water and everything good, they got the water and everything good, they got the gameplay loop right as well. This is the gameplay loop right as well. This is the gameplay loop right as well. This is the first version of Fishlop 3D I've had any first version of Fishlop 3D I've had any first version of Fishlop 3D I've had any model generate where I had to like model generate where I had to like model generate where I had to like remind myself it's just a demo and I remind myself it's just a demo and I remind myself it's just a demo and I have to go film. All the others are have to go film. All the others are have to go film. All the others are like, "Oh, cool. That looks really like, "Oh, cool. That looks really like, "Oh, cool. That looks really nice." And then I close it. This one is nice." And then I close it. This one is nice." And then I close it. This one is Oh, I actually kind of see the Oh, I actually kind of see the Oh, I actually kind of see the vision for the game now.

  25. vision for the game now. vision for the game now. Yeah, it's Yeah, it's Yeah, it's it's good. it's good. it's good. The little animations when they go to The little animations when they go to The little animations when they go to eat the food, especially when they grow, eat the food, especially when they grow, eat the food, especially when they grow, which hopefully one will be ready for in which hopefully one will be ready for in which hopefully one will be ready for in a sec. Also, this like crazy cinematic a sec. Also, this like crazy cinematic a sec. Also, this like crazy cinematic cam button they have where you can just cam button they have where you can just cam button they have where you can just Oh, that's that's insane. What? Oh, that's that's insane. What? Oh, that's that's insane. What? How is it this good? How is it this good? How is it this good? Oh, first aliens incoming. Jesus. Like, what? Jesus. Like, what? I I did not think we would get this far I I did not think we would get this far I I did not think we would get this far this fast. I really didn't. All of that this fast. I really didn't. All of that this fast. I really didn't. All of that said, this model is not without its said, this model is not without its said, this model is not without its quirks. The quirks are much tamer and quirks. The quirks are much tamer and quirks. The quirks are much tamer and less worldending thus far when compared less worldending thus far when compared less worldending thus far when compared to Opus 5, but it still has them. One of to Opus 5, but it still has them. One of to Opus 5, but it still has them. One of the most frustrating ones for me is the most frustrating ones for me is the most frustrating ones for me is paranoia around the context window. Back paranoia around the context window. Back paranoia around the context window. Back in the days before compaction was good in the days before compaction was good in the days before compaction was good and reliable, models were mostly trained and reliable, models were mostly trained and reliable, models were mostly trained on data that would fit within its on data that would fit within its on data that would fit within its context window. So, if it was given a context window. So, if it was given a context window. So, if it was given a task and the task was completed before task and the task was completed before task and the task was completed before it hit the end, awesome. If it took too it hit the end, awesome. If it took too it hit the end, awesome. If it took too long and ran out of context, it would long and ran out of context, it would long and ran out of context, it would fail. The result of this was an fail. The result of this was an fail. The result of this was an accidental training the models to be accidental training the models to be accidental training the models to be scared of those limits and to try to scared of those limits and to try to scared of those limits and to try to keep things under them. OpenAI has keep things under them. OpenAI has keep things under them. OpenAI has entirely beat this out of the models.

  26. entirely beat this out of the models. entirely beat this out of the models. They just don't care anymore. They will They just don't care anymore. They will They just don't care anymore. They will do whatever they want and they will do whatever they want and they will do whatever they want and they will trust compaction to keep them on task. trust compaction to keep them on task. trust compaction to keep them on task. Anthropic less so. They relied more on Anthropic less so. They relied more on Anthropic less so. They relied more on that million token context window where that million token context window where that million token context window where OpenAI caps it to 272K. Usually, even OpenAI caps it to 272K. Usually, even OpenAI caps it to 272K. Usually, even though OpenAI models can do a million though OpenAI models can do a million though OpenAI models can do a million tokens of context, Anthropic leans on it tokens of context, Anthropic leans on it tokens of context, Anthropic leans on it more. I've had a couple instances now more. I've had a couple instances now more. I've had a couple instances now where I experienced this particular where I experienced this particular where I experienced this particular style of call out. I noticed that one of style of call out. I noticed that one of style of call out. I noticed that one of the performance improvement runs I was the performance improvement runs I was the performance improvement runs I was doing on fish slop was not making doing on fish slop was not making doing on fish slop was not making progress, at least from my perspective. progress, at least from my perspective. progress, at least from my perspective. I didn't even spin this one up, though. I didn't even spin this one up, though. I didn't even spin this one up, though. I had this cloud code on my machine SSH I had this cloud code on my machine SSH I had this cloud code on my machine SSH into another machine to trigger cloud into another machine to trigger cloud into another machine to trigger cloud code, which it was able to do. What's code, which it was able to do. What's code, which it was able to do. What's concerning here isn't even what it says concerning here isn't even what it says concerning here isn't even what it says is concerning. It called out that the is concerning. It called out that the is concerning. It called out that the Opus 55 instance that is improving the Opus 55 instance that is improving the Opus 55 instance that is improving the game is slow to commit. After almost game is slow to commit. After almost game is slow to commit. After almost three hours, all its changes are still three hours, all its changes are still three hours, all its changes are still local. All this is fine. Where things local. All this is fine. Where things local. All this is fine. Where things get sketchy is the next part of the get sketchy is the next part of the get sketchy is the next part of the sentence. Its context is 10% from sentence. Its context is 10% from sentence. Its context is 10% from autocompact. It then follows up with if autocompact. It then follows up with if autocompact. It then follows up with if it crashed that work would be lost. None it crashed that work would be lost. None it crashed that work would be lost. None of that makes sense. None of this is of that makes sense. None of this is of that makes sense. None of this is stuff the model should think about or stuff the model should think about or stuff the model should think about or care about. Its context is 10% from care about. Its context is 10% from care about. Its context is 10% from autocompaction. Who cares? The model autocompaction. Who cares? The model autocompaction. Who cares? The model compacts well. It handles compaction.

  27. compacts well. It handles compaction. compacts well. It handles compaction. Great. And it's also way faster at Great. And it's also way faster at Great. And it's also way faster at compaction, which is quite nice. Then if compaction, which is quite nice. Then if compaction, which is quite nice. Then if it crashed, the work would be lost. it crashed, the work would be lost. it crashed, the work would be lost. What? What does that even mean? It's a What? What does that even mean? It's a What? What does that even mean? It's a real machine on my network. It knows real machine on my network. It knows real machine on my network. It knows it's a real machine. It knows it's a it's a real machine. It knows it's a it's a real machine. It knows it's a MacBook. It chose this computer because MacBook. It chose this computer because MacBook. It chose this computer because it's a Mac similar to mine that's on my it's a Mac similar to mine that's on my it's a Mac similar to mine that's on my fleet. It has no reason to be concerned fleet. It has no reason to be concerned fleet. It has no reason to be concerned about a crash losing work. These little about a crash losing work. These little about a crash losing work. These little moments haven't happened too often. It's moments haven't happened too often. It's moments haven't happened too often. It's usually when I'm auditing why things are usually when I'm auditing why things are usually when I'm auditing why things are taking a while. So like, we're already taking a while. So like, we're already taking a while. So like, we're already in a bad state, so to speak, and then in a bad state, so to speak, and then in a bad state, so to speak, and then I'm following up to figure out how we I'm following up to figure out how we I'm following up to figure out how we got there. That's when this model's got there. That's when this model's got there. That's when this model's discernment seems to be less strong. And discernment seems to be less strong. And discernment seems to be less strong. And this is not a thing I think almost any this is not a thing I think almost any this is not a thing I think almost any benchmarks really cover, which is why benchmarks really cover, which is why benchmarks really cover, which is why it's hard to see in the numbers there. it's hard to see in the numbers there. it's hard to see in the numbers there. The vast majority of benchmarks don't The vast majority of benchmarks don't The vast majority of benchmarks don't even hit compaction by the time they're even hit compaction by the time they're even hit compaction by the time they're done running, which is worth noting done running, which is worth noting done running, which is worth noting here. But it still does give you that here. But it still does give you that here. But it still does give you that small model feel and smell when you hit small model feel and smell when you hit small model feel and smell when you hit these edges because I I would not expect these edges because I I would not expect these edges because I I would not expect Fable to say something nonsense like Fable to say something nonsense like Fable to say something nonsense like this. Astra perhaps, but not Fable. Oh this. Astra perhaps, but not Fable. Oh this. Astra perhaps, but not Fable. Oh man, the performance improved version man, the performance improved version man, the performance improved version that it was working on is actually that it was working on is actually that it was working on is actually flying though. The fact that it looks flying though. The fact that it looks flying though. The fact that it looks this good and runs at 120 FPS. This this good and runs at 120 FPS. This this good and runs at 120 FPS. This might be the first version I have to might be the first version I have to might be the first version I have to publish so people can actually play it publish so people can actually play it publish so people can actually play it and see it because it's hard to believe.

  28. and see it because it's hard to believe. and see it because it's hard to believe. It does have its weird issues here and It does have its weird issues here and It does have its weird issues here and there. Like I just noticed when I was there. Like I just noticed when I was there. Like I just noticed when I was low enough and I looked up the fish's low enough and I looked up the fish's low enough and I looked up the fish's transparency is kind of brokenly the transparency is kind of brokenly the transparency is kind of brokenly the jellyfish by the water treatment on the jellyfish by the water treatment on the jellyfish by the water treatment on the ceiling there. But other than that, ceiling there. But other than that, ceiling there. But other than that, very very good. very very good. very very good. I want to talk a little bit about the I want to talk a little bit about the I want to talk a little bit about the spikiness of the model which touches on spikiness of the model which touches on spikiness of the model which touches on those weird dumb spikes that we were those weird dumb spikes that we were those weird dumb spikes that we were showing there because this model might showing there because this model might showing there because this model might have dumb spikes still. I don't know how have dumb spikes still. I don't know how have dumb spikes still. I don't know how frequent or how dumb those spikes will frequent or how dumb those spikes will frequent or how dumb those spikes will be because it's still day one of testing be because it's still day one of testing be because it's still day one of testing for me. I didn't have early access and I for me. I didn't have early access and I for me. I didn't have early access and I haven't shipped much beyond like five or haven't shipped much beyond like five or haven't shipped much beyond like five or so PRs with this model. From my so PRs with this model. From my so PRs with this model. From my rudimentary numbers with those poll rudimentary numbers with those poll rudimentary numbers with those poll requests, it is showing similar to lower requests, it is showing similar to lower requests, it is showing similar to lower rates of egregious high severity issues rates of egregious high severity issues rates of egregious high severity issues in the code that it's putting up with a in the code that it's putting up with a in the code that it's putting up with a similar to slightly higher number of similar to slightly higher number of similar to slightly higher number of small issues and nitpicks compared to small issues and nitpicks compared to small issues and nitpicks compared to something like Fable 5.1. But that's something like Fable 5.1. But that's something like Fable 5.1. But that's still just a day of medium-sized poll still just a day of medium-sized poll still just a day of medium-sized poll requests. It's not enough data to be requests. It's not enough data to be requests. It's not enough data to be 100% sure here. But my gut feel if we 100% sure here. But my gut feel if we 100% sure here. But my gut feel if we were to look at my previous response were to look at my previous response were to look at my previous response quality diagram with 5.1 versus Astra is quality diagram with 5.1 versus Astra is quality diagram with 5.1 versus Astra is that it is very similar to Fable 5.1's that it is very similar to Fable 5.1's that it is very similar to Fable 5.1's performance here with slightly higher performance here with slightly higher performance here with slightly higher overall quality, especially if you care overall quality, especially if you care overall quality, especially if you care about the quality of the pros and like about the quality of the pros and like about the quality of the pros and like what it says when you work with it. It's what it says when you work with it. It's what it says when you work with it. It's so much more readable, but it spikes so much more readable, but it spikes so much more readable, but it spikes down definitely go a little lower than down definitely go a little lower than down definitely go a little lower than fables do too. How much lower? I don't fables do too. How much lower? I don't fables do too. How much lower? I don't know. How frequently? I'm not sure yet.

  29. know. How frequently? I'm not sure yet. know. How frequently? I'm not sure yet. I want to make sure we aren't getting I want to make sure we aren't getting I want to make sure we aren't getting too excited here in just outright too excited here in just outright too excited here in just outright throwing out Fable in favor of Opus throwing out Fable in favor of Opus throwing out Fable in favor of Opus because this model is smaller. It is because this model is smaller. It is because this model is smaller. It is dumber. It will make mistakes. So will dumber. It will make mistakes. So will dumber. It will make mistakes. So will Fable, just less so. And as we're doing Fable, just less so. And as we're doing Fable, just less so. And as we're doing longer and longer jobs with more and longer and longer jobs with more and longer and longer jobs with more and more work, the likelihood you hit one of more work, the likelihood you hit one of more work, the likelihood you hit one of those failures goes up, not down. And those failures goes up, not down. And those failures goes up, not down. And the result of that is you'll sometimes the result of that is you'll sometimes the result of that is you'll sometimes see Opus 5.5 get stuck repairing things see Opus 5.5 get stuck repairing things see Opus 5.5 get stuck repairing things it doesn't need to when you send it off it doesn't need to when you send it off it doesn't need to when you send it off on some performance work or trapping on some performance work or trapping on some performance work or trapping itself thinking its context window is itself thinking its context window is itself thinking its context window is about to close and it's going to die about to close and it's going to die about to close and it's going to die when it's not. Opus will have these when it's not. Opus will have these when it's not. Opus will have these quirks. It has them a lot less than Opus quirks. It has them a lot less than Opus quirks. It has them a lot less than Opus 5 did but it has them enough to consider 5 did but it has them enough to consider 5 did but it has them enough to consider and we should all be realistic about and we should all be realistic about and we should all be realistic about what this model can do. That all said, what this model can do. That all said, what this model can do. That all said, it is unbelievable just how much more it is unbelievable just how much more it is unbelievable just how much more usage you can get, not only compared to usage you can get, not only compared to usage you can get, not only compared to Fable inside of your Cloud Code plan, Fable inside of your Cloud Code plan, Fable inside of your Cloud Code plan, but compared to all of the other models but compared to all of the other models but compared to all of the other models you can use with your codeex plans. It you can use with your codeex plans. It you can use with your codeex plans. It is kind of crazy to say that Anthropic is kind of crazy to say that Anthropic is kind of crazy to say that Anthropic put out a really good value here, but put out a really good value here, but put out a really good value here, but they did. Maybe not in the pure API they did. Maybe not in the pure API they did. Maybe not in the pure API prices if you're paying by the token, prices if you're paying by the token, prices if you're paying by the token, but at the absolute least, the amount of but at the absolute least, the amount of but at the absolute least, the amount of value you get out of your $200 sub in value you get out of your $200 sub in value you get out of your $200 sub in Claude, it is running laps around our Claude, it is running laps around our Claude, it is running laps around our friends at OpenAI. They did also bump friends at OpenAI. They did also bump friends at OpenAI. They did also bump the five hour limit, which is worth the five hour limit, which is worth the five hour limit, which is worth noting. It's still not as generous as noting. It's still not as generous as noting. It's still not as generous as OpenAI's entire lack of a weekly limit OpenAI's entire lack of a weekly limit OpenAI's entire lack of a weekly limit on the $1200 plans, but at least the on the $1200 plans, but at least the on the $1200 plans, but at least the five hour limit is much more generous five hour limit is much more generous five hour limit is much more generous now. I haven't come close to hitting it now. I haven't come close to hitting it now. I haven't come close to hitting it for what it's worth. That all said, I am for what it's worth. That all said, I am for what it's worth. That all said, I am still quite excited to dive deeper into still quite excited to dive deeper into still quite excited to dive deeper into GBD6 Soul and see what it is capable of.

  30. GBD6 Soul and see what it is capable of. GBD6 Soul and see what it is capable of. It is even cheaper than Opus and seems It is even cheaper than Opus and seems It is even cheaper than Opus and seems to be performing quite well as well. And to be performing quite well as well. And to be performing quite well as well. And historically, OpenAI has been better at historically, OpenAI has been better at historically, OpenAI has been better at taking the capability of the high-end taking the capability of the high-end taking the capability of the high-end expensive model and forcing it into the expensive model and forcing it into the expensive model and forcing it into the cheaper, smaller models. Anthropic's cheaper, smaller models. Anthropic's cheaper, smaller models. Anthropic's never been the best at that. They're the never been the best at that. They're the never been the best at that. They're the pre-training kings and OpenAI is the pre-training kings and OpenAI is the pre-training kings and OpenAI is the post-training kings. But it seems like post-training kings. But it seems like post-training kings. But it seems like Anthropic caught up a lot here on the Anthropic caught up a lot here on the Anthropic caught up a lot here on the post-training side. So, I cannot wait to post-training side. So, I cannot wait to post-training side. So, I cannot wait to dive deeper into what OpenAI has been dive deeper into what OpenAI has been dive deeper into what OpenAI has been doing. Make sure you're subbed if you doing. Make sure you're subbed if you doing. Make sure you're subbed if you haven't yet in order to see that video haven't yet in order to see that video haven't yet in order to see that video right when it drops. I can't believe right when it drops. I can't believe right when it drops. I can't believe this week's been so chaotic and it's this week's been so chaotic and it's this week's been so chaotic and it's only Tuesday. I have a feeling it's only Tuesday. I have a feeling it's only Tuesday. I have a feeling it's going to get worse. And next week has going to get worse. And next week has going to get worse. And next week has OpenAI's dev day, so it's probably going OpenAI's dev day, so it's probably going OpenAI's dev day, so it's probably going to get even crazier from there. So, make to get even crazier from there. So, make to get even crazier from there. So, make sure you stay tuned because there's sure you stay tuned because there's sure you stay tuned because there's going to be a lot of stuff to cover and going to be a lot of stuff to cover and going to be a lot of stuff to cover and I have a feeling I will not be sleeping I have a feeling I will not be sleeping I have a feeling I will not be sleeping enough this week. Let me know how y'all enough this week. Let me know how y'all enough this week. Let me know how y'all feel about this model. I think it is a feel about this model. I think it is a feel about this model. I think it is a way better solution than Opus 5 was and way better solution than Opus 5 was and way better solution than Opus 5 was and it's getting close enough to Fable that it's getting close enough to Fable that it's getting close enough to Fable that it makes a lot more sense in your cloud it makes a lot more sense in your cloud it makes a lot more sense in your cloud subs. So, congrats to everybody who's subs. So, congrats to everybody who's subs. So, congrats to everybody who's working at a company that restricts what working at a company that restricts what working at a company that restricts what models you can use. Opus 55 is going to models you can use. Opus 55 is going to models you can use. Opus 55 is going to be a lifechanging experience for y'all. be a lifechanging experience for y'all. be a lifechanging experience for y'all. And to everybody else trying to maximize And to everybody else trying to maximize And to everybody else trying to maximize the usage of your subs, congrats. You the usage of your subs, congrats. You the usage of your subs, congrats. You can now do that in a way that makes can now do that in a way that makes can now do that in a way that makes actual sense with code that I'm not actual sense with code that I'm not actual sense with code that I'm not scared to look at, much less merge. I scared to look at, much less merge. I scared to look at, much less merge. I want to go play with these models more.

  31. want to go play with these models more. want to go play with these models more. So, uh yeah. Bards.

No summary available yet.

View original episode ↗