AI Engineer Paris 2026 Main Stage: Google DeepMind, ElevenLabs, Hugging Face & Stripe | Day 2
Read full transcript 239 segments
-
Ladies and gentlemen, please join me in Ladies and gentlemen, please join me in welcoming to the stage your MC for the welcoming to the stage your MC for the welcoming to the stage your MC for the AI engineer Paris 2026, AI engineer Paris 2026, AI engineer Paris 2026, developer relations engineer at Replet, developer relations engineer at Replet, developer relations engineer at Replet, Rou Chevrey. Good morning. M check. Good morning Good morning. M check. Good morning everybody. How's it going? everybody. How's it going? everybody. How's it going? Hey, this feels like day two energy, so Hey, this feels like day two energy, so Hey, this feels like day two energy, so we need to bring it up. Okay. How's it we need to bring it up. Okay. How's it we need to bring it up. Okay. How's it going everybody? going everybody? going everybody? >> Yeah. Nice, nice, nice. Well, welcome to >> Yeah. Nice, nice, nice. Well, welcome to >> Yeah. Nice, nice, nice. Well, welcome to AI Engineer Paris day two. All righty. AI Engineer Paris day two. All righty. AI Engineer Paris day two. All righty. So, who was with us yesterday? Show So, who was with us yesterday? Show So, who was with us yesterday? Show hands. Okay, most of you guys. That's hands. Okay, most of you guys. That's hands. Okay, most of you guys. That's awesome. So yesterday was all about how awesome. So yesterday was all about how awesome. So yesterday was all about how to think about the evolution of to think about the evolution of to think about the evolution of technology to think about AI today and technology to think about AI today and technology to think about AI today and then we talked about the future of AI then we talked about the future of AI then we talked about the future of AI right the future of AI um software right the future of AI um software right the future of AI um software factories um and also uh the the future factories um and also uh the the future factories um and also uh the the future of frontier models but I'm particularly of frontier models but I'm particularly of frontier models but I'm particularly excited about today today we're going to excited about today today we're going to excited about today today we're going to talk about generative media about talk about generative media about talk about generative media about software factories again we're going to software factories again we're going to software factories again we're going to talk about uh graph rags and we're and talk about uh graph rags and we're and talk about uh graph rags and we're and we're going to cover so much more so So we're going to cover so much more so So we're going to cover so much more so So with that, I want you to uh think about with that, I want you to uh think about with that, I want you to uh think about what you want to see. Please pay what you want to see. Please pay what you want to see. Please pay attention to everything that is going to attention to everything that is going to attention to everything that is going to be said here. Our speakers are going to be said here. Our speakers are going to be said here. Our speakers are going to be hanging out. So also take the time be hanging out. So also take the time be hanging out. So also take the time and go to the expo and talk to everybody and go to the expo and talk to everybody and go to the expo and talk to everybody to learn as much as you can. All right.
-
to learn as much as you can. All right. to learn as much as you can. All right. So you know the drill by now if you were So you know the drill by now if you were So you know the drill by now if you were here, right? This event wouldn't be here, right? This event wouldn't be here, right? This event wouldn't be possible without Mrol. So I would like possible without Mrol. So I would like possible without Mrol. So I would like to give give it up for them. So please to give give it up for them. So please to give give it up for them. So please let's give it up for M. Thanks for let's give it up for M. Thanks for let's give it up for M. Thanks for organizing and let's give it up for our organizing and let's give it up for our organizing and let's give it up for our sponsors, our platinum sponsor, Nvidia. sponsors, our platinum sponsor, Nvidia. sponsors, our platinum sponsor, Nvidia. Yes, please. Yeah, keep it going. So, to Yes, please. Yeah, keep it going. So, to Yes, please. Yeah, keep it going. So, to for all our uh gold sponsors, silver for all our uh gold sponsors, silver for all our uh gold sponsors, silver sponsors, sponsors, sponsors, bronze sponsors, and also our supporting bronze sponsors, and also our supporting bronze sponsors, and also our supporting sponsors. Let's give it up one more sponsors. Let's give it up one more sponsors. Let's give it up one more time. Yeah, thank you. time. Yeah, thank you. time. Yeah, thank you. Okay. All right. So, without further Okay. All right. So, without further Okay. All right. So, without further ado, our next speaker is from Black ado, our next speaker is from Black ado, our next speaker is from Black Forest Labs, and he's got some tendency. Forest Labs, and he's got some tendency. Forest Labs, and he's got some tendency. He loves going to the dentist. I don't He loves going to the dentist. I don't He loves going to the dentist. I don't know why. Maybe we can ask him later. know why. Maybe we can ask him later. know why. Maybe we can ask him later. But he's going to talk to us about um But he's going to talk to us about um But he's going to talk to us about um about how they went from image about how they went from image about how they went from image generation to robotics. So, please join generation to robotics. So, please join generation to robotics. So, please join me in welcoming to the stage Yakob me in welcoming to the stage Yakob me in welcoming to the stage Yakob Porchman.
-
Thank you very much. Super nice to meet Thank you very much. Super nice to meet you everyone. Just need my slides maybe you everyone. Just need my slides maybe you everyone. Just need my slides maybe in here. Awesome. Nice to meet you in here. Awesome. Nice to meet you in here. Awesome. Nice to meet you everyone. My name is Jakob. I lead the everyone. My name is Jakob. I lead the everyone. My name is Jakob. I lead the AI solutions team. Um that's the four AI solutions team. Um that's the four AI solutions team. Um that's the four deployed engineering team and solutions deployed engineering team and solutions deployed engineering team and solutions engineering team at Black Forest Labs. engineering team at Black Forest Labs. engineering team at Black Forest Labs. And I would like to start in with a And I would like to start in with a And I would like to start in with a hopefully trivial question. So give me a hopefully trivial question. So give me a hopefully trivial question. So give me a brief show of hands of who has generated brief show of hands of who has generated brief show of hands of who has generated an image or a video before using AI. an image or a video before using AI. an image or a video before using AI. Just give me a brief show of hands. This Just give me a brief show of hands. This Just give me a brief show of hands. This should be trivial. Come on. This is should be trivial. Come on. This is should be trivial. Come on. This is almost everyone hopefully. Um almost everyone hopefully. Um almost everyone hopefully. Um and um then hopefully if you generate an and um then hopefully if you generate an and um then hopefully if you generate an image or video before um you hopefully image or video before um you hopefully image or video before um you hopefully were in contact with Flux uh or were in contact with Flux uh or were in contact with Flux uh or Blackfire Labs before because Blackfire Blackfire Labs before because Blackfire Blackfire Labs before because Blackfire Labs at Blackfire Labs we are a research Labs at Blackfire Labs we are a research Labs at Blackfire Labs we are a research lab based in Germany and we build lab based in Germany and we build lab based in Germany and we build foundational image and video models for foundational image and video models for foundational image and video models for image generation. Um and you can see image generation. Um and you can see image generation. Um and you can see some of the samples um videos and images some of the samples um videos and images some of the samples um videos and images that were generated with our recent that were generated with our recent that were generated with our recent model flux 3 in the back uh back here. model flux 3 in the back uh back here. model flux 3 in the back uh back here. Now the second question that I have for Now the second question that I have for Now the second question that I have for you is who of you would believe me um if you is who of you would believe me um if you is who of you would believe me um if I would tell you that you can use that I would tell you that you can use that I would tell you that you can use that same video model to also control robots same video model to also control robots same video model to also control robots and play video games.
-
and play video games. and play video games. Okay, that's a few people, not really Okay, that's a few people, not really Okay, that's a few people, not really everyone. Um well, it's kind of a it's everyone. Um well, it's kind of a it's everyone. Um well, it's kind of a it's kind of a rhetorical question obviously. kind of a rhetorical question obviously. kind of a rhetorical question obviously. um otherwise I wouldn't talk about it um otherwise I wouldn't talk about it um otherwise I wouldn't talk about it because that's exactly the final thing because that's exactly the final thing because that's exactly the final thing that I'm going to close with today is that I'm going to close with today is that I'm going to close with today is how we're going to customize Flux um to how we're going to customize Flux um to how we're going to customize Flux um to not just generate media, generate not just generate media, generate not just generate media, generate images, generate videos and and edit um images, generate videos and and edit um images, generate videos and and edit um edit both um but we're going to talk edit both um but we're going to talk edit both um but we're going to talk about how to steer video games and steer about how to steer video games and steer about how to steer video games and steer robots with it. So you can see two robots with it. So you can see two robots with it. So you can see two different samples here that were also on different samples here that were also on different samples here that were also on the left generated with Flux. On the the left generated with Flux. On the the left generated with Flux. On the left you have Flux actually playing a left you have Flux actually playing a left you have Flux actually playing a video game. Um so this is not just a video game. Um so this is not just a video game. Um so this is not just a video of a video of a video of a generated but it's actually flux taking generated but it's actually flux taking generated but it's actually flux taking the control and and generating the the control and and generating the the control and and generating the frames that are happening. And on the frames that are happening. And on the frames that are happening. And on the right you have a Lero that is completing right you have a Lero that is completing right you have a Lero that is completing pick and place tasks um with Flux the pick and place tasks um with Flux the pick and place tasks um with Flux the video our video model flux 3 as a video our video model flux 3 as a video our video model flux 3 as a intelligence backbone and what is in intelligence backbone and what is in intelligence backbone and what is in between video generation and using flux between video generation and using flux between video generation and using flux as a backbone for a robot that's exactly as a backbone for a robot that's exactly as a backbone for a robot that's exactly what we're going to talk about today what we're going to talk about today what we're going to talk about today because the part that is in between because the part that is in between because the part that is in between that's basically bridging that gap is that's basically bridging that gap is that's basically bridging that gap is customization of this flux model right customization of this flux model right customization of this flux model right there is a super powerful model that we there is a super powerful model that we there is a super powerful model that we pre-trained trained mid-trained on pre-trained trained mid-trained on pre-trained trained mid-trained on essentially understanding the visual essentially understanding the visual essentially understanding the visual world having having an understanding for world having having an understanding for world having having an understanding for um how objects are behaving in the um how objects are behaving in the um how objects are behaving in the visual world and we are taking this visual world and we are taking this visual world and we are taking this model and customizing it for certain use model and customizing it for certain use model and customizing it for certain use cases and that's exactly what we do in cases and that's exactly what we do in cases and that's exactly what we do in our AI solutions team as well and this our AI solutions team as well and this our AI solutions team as well and this is what I want to talk to you about um
-
is what I want to talk to you about um is what I want to talk to you about um in a couple different steps to break in a couple different steps to break in a couple different steps to break this down a little bit before we go into this down a little bit before we go into this down a little bit before we go into robots that's going to be the final robots that's going to be the final robots that's going to be the final stage of what we're going to talk about stage of what we're going to talk about stage of what we're going to talk about um um um I want to talk a little bit about the I want to talk a little bit about the I want to talk a little bit about the anatomy of the inference pipeline anatomy of the inference pipeline anatomy of the inference pipeline because this is also going to give you because this is also going to give you because this is also going to give you an overview of sort of the customization an overview of sort of the customization an overview of sort of the customization vectors that we have when we walk work vectors that we have when we walk work vectors that we have when we walk work with flux and we talk to enterprise with flux and we talk to enterprise with flux and we talk to enterprise customers that want to use flux for very customers that want to use flux for very customers that want to use flux for very specific use cases. Um this breaks down specific use cases. Um this breaks down specific use cases. Um this breaks down in approximately six steps. First of in approximately six steps. First of in approximately six steps. First of course there is a user input uh that can course there is a user input uh that can course there is a user input uh that can be optimized. Um this is usually a text be optimized. Um this is usually a text be optimized. Um this is usually a text prompt, an image prompt, a video prompt, prompt, an image prompt, a video prompt, prompt, an image prompt, a video prompt, fairly straightforward, right? Um second fairly straightforward, right? Um second fairly straightforward, right? Um second step is the prompt upsampling. Prompt step is the prompt upsampling. Prompt step is the prompt upsampling. Prompt upsampling essentially means that we upsampling essentially means that we upsampling essentially means that we have an LLM or a VLM that sits between have an LLM or a VLM that sits between have an LLM or a VLM that sits between the user prompt that is sent to us and the user prompt that is sent to us and the user prompt that is sent to us and the flux model. And we essentially try the flux model. And we essentially try the flux model. And we essentially try to understand the intent. We try to to understand the intent. We try to to understand the intent. We try to understand what the user actually wants understand what the user actually wants understand what the user actually wants to do and we try to improve the prompt to do and we try to improve the prompt to do and we try to improve the prompt that the model generates something that the model generates something that the model generates something that's closer to what the user is that's closer to what the user is that's closer to what the user is actually intending to do. Then there's actually intending to do. Then there's actually intending to do. Then there's moderation. And you might be surprised moderation. And you might be surprised moderation. And you might be surprised in seeing moderation when it when it in seeing moderation when it when it in seeing moderation when it when it comes to when it comes to customization comes to when it comes to customization comes to when it comes to customization of a model. Well, actually moderation is of a model. Well, actually moderation is of a model. Well, actually moderation is one of the most important product one of the most important product one of the most important product surfaces. Um because moderation surfaces. Um because moderation surfaces. Um because moderation essentially decides what you're going to essentially decides what you're going to essentially decides what you're going to see or what you're not going to see.
-
see or what you're not going to see. see or what you're not going to see. It's not part of mostly not part of the It's not part of mostly not part of the It's not part of mostly not part of the model. Um but it is part of the user model. Um but it is part of the user model. Um but it is part of the user experience and we're going to talk a experience and we're going to talk a experience and we're going to talk a little bit about how we customize it and little bit about how we customize it and little bit about how we customize it and also how we actually embed moderation also how we actually embed moderation also how we actually embed moderation into the model weights. Fourth, there's into the model weights. Fourth, there's into the model weights. Fourth, there's of course the model. There are the model of course the model. There are the model of course the model. There are the model weights which we can fine-tune. We can weights which we can fine-tune. We can weights which we can fine-tune. We can run laws on it. We can fine-tune for run laws on it. We can fine-tune for run laws on it. We can fine-tune for style, for characters, for products. And style, for characters, for products. And style, for characters, for products. And we can fine-tune also for behavior and we can fine-tune also for behavior and we can fine-tune also for behavior and using fine-tunes, we can also add whole using fine-tunes, we can also add whole using fine-tunes, we can also add whole new modalities. And that's exactly what new modalities. And that's exactly what new modalities. And that's exactly what we did with flux flux 3 action, which we did with flux flux 3 action, which we did with flux flux 3 action, which I'm going to talk about in the end of my I'm going to talk about in the end of my I'm going to talk about in the end of my presentation. Then there's a presentation. Then there's a presentation. Then there's a post-processing step. Of course, this is post-processing step. Of course, this is post-processing step. Of course, this is kind of arbitrary. There can be kind of arbitrary. There can be kind of arbitrary. There can be upscaling, there can be restoration, upscaling, there can be restoration, upscaling, there can be restoration, there can be further moderation, right? there can be further moderation, right? there can be further moderation, right? Um, endless customization opportunities Um, endless customization opportunities Um, endless customization opportunities here. And then finally there is um we here. And then finally there is um we here. And then finally there is um we also work a lot about how we actually also work a lot about how we actually also work a lot about how we actually serve the model, how we deploy it, how serve the model, how we deploy it, how serve the model, how we deploy it, how we optimize it, how we distill it um etc we optimize it, how we distill it um etc we optimize it, how we distill it um etc etc. And what I want to do in the next etc. And what I want to do in the next etc. And what I want to do in the next 20-ish minutes is I want to go through 20-ish minutes is I want to go through 20-ish minutes is I want to go through essentially the four deployed engineers essentially the four deployed engineers essentially the four deployed engineers toolbox at least scratch the surface a toolbox at least scratch the surface a toolbox at least scratch the surface a little bit um of how we customize flux little bit um of how we customize flux little bit um of how we customize flux um in our day-to-day. So starting uh um in our day-to-day. So starting uh um in our day-to-day. So starting uh with the first major customization I with the first major customization I with the first major customization I want to focus on this prompt upsampling want to focus on this prompt upsampling want to focus on this prompt upsampling step. As I already mentioned prompt step. As I already mentioned prompt step. As I already mentioned prompt upsambling means that there is a VLM or upsambling means that there is a VLM or upsambling means that there is a VLM or an LLM that sits between the user and an LLM that sits between the user and an LLM that sits between the user and the model.
-
the model. the model. Um that looks in practice something like Um that looks in practice something like Um that looks in practice something like this. Um so on the left here you have a this. Um so on the left here you have a this. Um so on the left here you have a user prompt. So this might be super user prompt. So this might be super user prompt. So this might be super super straightforward. a cabin in the super straightforward. a cabin in the super straightforward. a cabin in the forest. And if you send that straight to forest. And if you send that straight to forest. And if you send that straight to the flux model, generate an image, the flux model, generate an image, the flux model, generate an image, you'll see something like the image on you'll see something like the image on you'll see something like the image on the left. It's a plain um cabin in the the left. It's a plain um cabin in the the left. It's a plain um cabin in the forest, right? The mission completed. forest, right? The mission completed. forest, right? The mission completed. However, it might not be optimized for However, it might not be optimized for However, it might not be optimized for user intent. It might not be optimized user intent. It might not be optimized user intent. It might not be optimized for user preferences. Um it highly for user preferences. Um it highly for user preferences. Um it highly depends, of course, what the user wants depends, of course, what the user wants depends, of course, what the user wants here. In this case, the prompt is fairly here. In this case, the prompt is fairly here. In this case, the prompt is fairly simple. Um but it demonstrates the simple. Um but it demonstrates the simple. Um but it demonstrates the purpose quite well. So essentially we purpose quite well. So essentially we purpose quite well. So essentially we take the cabin in the forest and we take the cabin in the forest and we take the cabin in the forest and we essentially extend the prompt. In this essentially extend the prompt. In this essentially extend the prompt. In this case the default upsampler just case the default upsampler just case the default upsampler just optimizes for user preference. Extends optimizes for user preference. Extends optimizes for user preference. Extends the prompt the user the cabin in the the prompt the user the cabin in the the prompt the user the cabin in the forest to something like a forest to something like a forest to something like a hyperrealistic right angle stone chimney hyperrealistic right angle stone chimney hyperrealistic right angle stone chimney softarm golden light etc etc. And then softarm golden light etc etc. And then softarm golden light etc etc. And then you see the final result um um on the you see the final result um um on the you see the final result um um on the right. Whether you like the left or the right. Whether you like the left or the right. Whether you like the left or the right better, that's up for up for right better, that's up for up for right better, that's up for up for discussion, of course. But the point discussion, of course. But the point discussion, of course. But the point being here is that the prompt upsling being here is that the prompt upsling being here is that the prompt upsling step gives us an incredibly powerful step gives us an incredibly powerful step gives us an incredibly powerful tool to actually steer the user intent tool to actually steer the user intent tool to actually steer the user intent and the generation that we're taking and the generation that we're taking and the generation that we're taking into any direction possible. So let's into any direction possible. So let's into any direction possible. So let's take a look at a concrete um concrete take a look at a concrete um concrete take a look at a concrete um concrete case. So this specific this was a real case. So this specific this was a real case. So this specific this was a real customer case. The customer essentially customer case. The customer essentially customer case. The customer essentially came to us and told us well actually we came to us and told us well actually we came to us and told us well actually we want to illustrate certain conversations want to illustrate certain conversations want to illustrate certain conversations that a user having and users are having that a user having and users are having that a user having and users are having in our application with flux images. So in our application with flux images. So in our application with flux images. So they want to create illustrations for they want to create illustrations for they want to create illustrations for topics. These topics could be completely topics. These topics could be completely topics. These topics could be completely random and they're machine generated. So
-
random and they're machine generated. So random and they're machine generated. So they're not hunt created and they're they're not hunt created and they're they're not hunt created and they're also not image prompts. So for example also not image prompts. So for example also not image prompts. So for example this could be familyfriendly travel this could be familyfriendly travel this could be familyfriendly travel destinations or draining water for my destinations or draining water for my destinations or draining water for my dishwasher. Right? Completely random dishwasher. Right? Completely random dishwasher. Right? Completely random topics not in the form of an image topics not in the form of an image topics not in the form of an image prompt machine generated. uh this could prompt machine generated. uh this could prompt machine generated. uh this could be a vast variety of topics. Now our be a vast variety of topics. Now our be a vast variety of topics. Now our task is to generate an illustration task is to generate an illustration task is to generate an illustration somehow illustrate this nicely. But the somehow illustrate this nicely. But the somehow illustrate this nicely. But the trick to this is that of course there trick to this is that of course there trick to this is that of course there are certain design requirements that we are certain design requirements that we are certain design requirements that we need to meet. So for example um in this need to meet. So for example um in this need to meet. So for example um in this specific case the subject is supposed to specific case the subject is supposed to specific case the subject is supposed to be on the right. It's supposed to be a be on the right. It's supposed to be a be on the right. It's supposed to be a very clean image. There's supposed to be very clean image. There's supposed to be very clean image. There's supposed to be a lot of space on the left for a text a lot of space on the left for a text a lot of space on the left for a text overlay. And there are certain design overlay. And there are certain design overlay. And there are certain design rules that also apply. For example, no rules that also apply. For example, no rules that also apply. For example, no people involved, no landmarks, and a people involved, no landmarks, and a people involved, no landmarks, and a couple others. So what we essentially do couple others. So what we essentially do couple others. So what we essentially do or what this actually comes down to is or what this actually comes down to is or what this actually comes down to is that the model the image model is that the model the image model is that the model the image model is actually not the bottleneck anymore actually not the bottleneck anymore actually not the bottleneck anymore because the model technically can because the model technically can because the model technically can generate all of these all of these generate all of these all of these generate all of these all of these different different things, right? No different different things, right? No different different things, right? No problem. But really the design problem. But really the design problem. But really the design requirements end up being the bottleneck requirements end up being the bottleneck requirements end up being the bottleneck and where prompt assembling comes in is and where prompt assembling comes in is and where prompt assembling comes in is matching and bridging the user prompt matching and bridging the user prompt matching and bridging the user prompt with the design requirements and with the design requirements and with the design requirements and matching these design requirements to matching these design requirements to matching these design requirements to the training distribution of flux. So the training distribution of flux. So the training distribution of flux. So finding the exact right spot that we finding the exact right spot that we finding the exact right spot that we need to send to the model um from the need to send to the model um from the need to send to the model um from the training distribution to generate the training distribution to generate the training distribution to generate the the exact image um that the exact image um that the exact image um that um um um slides are gone um here we go um to slides are gone um here we go um to slides are gone um here we go um to generate the exact image that actually generate the exact image that actually generate the exact image that actually follows the design requirements. So to follows the design requirements. So to follows the design requirements. So to give you two examples from the the
-
give you two examples from the the give you two examples from the the topics that you already saw, um first topics that you already saw, um first topics that you already saw, um first they're familyfriendly travel they're familyfriendly travel they're familyfriendly travel destinations. Um if you send that as a destinations. Um if you send that as a destinations. Um if you send that as a raw prompt to flux, completely raw prompt to flux, completely raw prompt to flux, completely unseampled, you're going to see complete unseampled, you're going to see complete unseampled, you're going to see complete gibberish, complete uh trash trash gibberish, complete uh trash trash gibberish, complete uh trash trash images essentially. For example, there's images essentially. For example, there's images essentially. For example, there's this this postcard style. The v variety this this postcard style. The v variety this this postcard style. The v variety between different seats also is between different seats also is between different seats also is extremely high when you just send this extremely high when you just send this extremely high when you just send this completely unupled. If we use our completely unupled. If we use our completely unupled. If we use our default upsampler, which is just default upsampler, which is just default upsampler, which is just optimized for general user user optimized for general user user optimized for general user user preference, then you you'll see preference, then you you'll see preference, then you you'll see something like in the middle, probably a something like in the middle, probably a something like in the middle, probably a little bit closer to somewhat of an little bit closer to somewhat of an little bit closer to somewhat of an illustration, but not at all matching illustration, but not at all matching illustration, but not at all matching the design requirements. But if we the design requirements. But if we the design requirements. But if we actually customize the upsampling step, actually customize the upsampling step, actually customize the upsampling step, and we embed the design requirements and we embed the design requirements and we embed the design requirements directly in the prompt upsampling, directly in the prompt upsampling, directly in the prompt upsampling, you'll see something on the right. So you'll see something on the right. So you'll see something on the right. So the subject is on the right. It's a very the subject is on the right. It's a very the subject is on the right. It's a very clean image, very sort of glossy. clean image, very sort of glossy. clean image, very sort of glossy. There's a lot of space on the left for There's a lot of space on the left for There's a lot of space on the left for text overlay. So meeting exactly the text overlay. So meeting exactly the text overlay. So meeting exactly the design requirements. Same goes for the design requirements. Same goes for the design requirements. Same goes for the second example. So draining water in a second example. So draining water in a second example. So draining water in a dishwasher without propped up sampling dishwasher without propped up sampling dishwasher without propped up sampling complete complete random image. Um complete complete random image. Um complete complete random image. Um getting a little bit better some sort of getting a little bit better some sort of getting a little bit better some sort of instruction video but still not at all instruction video but still not at all instruction video but still not at all what we want in the second case. But if what we want in the second case. But if what we want in the second case. But if we actually customize this upsampler we actually customize this upsampler we actually customize this upsampler which is essentially the system prompt which is essentially the system prompt which is essentially the system prompt that looks at the intent and tries to that looks at the intent and tries to that looks at the intent and tries to match that into the training match that into the training match that into the training distribution knowing the design distribution knowing the design distribution knowing the design requirements. we'll get a very nice, requirements. we'll get a very nice, requirements. we'll get a very nice, very clean image, subject on the right, very clean image, subject on the right, very clean image, subject on the right, a lot of space on the left, text overlay a lot of space on the left, text overlay a lot of space on the left, text overlay possible. Um, and this is essentially a possible. Um, and this is essentially a possible. Um, and this is essentially a pretty simple trick. Um, because the um pretty simple trick. Um, because the um pretty simple trick. Um, because the um the massive example and what we
-
the massive example and what we the massive example and what we essentially learned from this um the essentially learned from this um the essentially learned from this um the massive advantage is that it's super massive advantage is that it's super massive advantage is that it's super quickly to iterate, right? You quickly to iterate, right? You quickly to iterate, right? You essentially just iterate over a system essentially just iterate over a system essentially just iterate over a system prompt. There's no training needed. Um, prompt. There's no training needed. Um, prompt. There's no training needed. Um, and you have maximum use case and you have maximum use case and you have maximum use case flexibility. And this is sort of flexibility. And this is sort of flexibility. And this is sort of necessary as well to include the step necessary as well to include the step necessary as well to include the step because what we learn especially with because what we learn especially with because what we learn especially with when talking with enterprise customers when talking with enterprise customers when talking with enterprise customers is that users are just lazy. Users don't is that users are just lazy. Users don't is that users are just lazy. Users don't necessarily want to know exactly how our necessarily want to know exactly how our necessarily want to know exactly how our training distribution looks like right training distribution looks like right training distribution looks like right how to perfectly um perfectly prompt how to perfectly um perfectly prompt how to perfectly um perfectly prompt flux. That's in a lot of cases where we flux. That's in a lot of cases where we flux. That's in a lot of cases where we come in and we customize prompt come in and we customize prompt come in and we customize prompt upsampling to match exactly user intent upsampling to match exactly user intent upsampling to match exactly user intent design requirements and training design requirements and training design requirements and training distribution to find exactly the right distribution to find exactly the right distribution to find exactly the right sweet spot here. Um yeah, front up sling sweet spot here. Um yeah, front up sling sweet spot here. Um yeah, front up sling seems super simple but it's an seems super simple but it's an seems super simple but it's an incredibly powerful tool and because it incredibly powerful tool and because it incredibly powerful tool and because it is so simple to to iterate over and it is so simple to to iterate over and it is so simple to to iterate over and it doesn't need any training. Moving on to doesn't need any training. Moving on to doesn't need any training. Moving on to the next stop is moderation and safety. the next stop is moderation and safety. the next stop is moderation and safety. And you might be again a little bit And you might be again a little bit And you might be again a little bit surprised to see this in the surprised to see this in the surprised to see this in the customization surface here, but customization surface here, but customization surface here, but moderation is actually one of the most moderation is actually one of the most moderation is actually one of the most important product surfaces that we have important product surfaces that we have important product surfaces that we have um especially when offering a model to um especially when offering a model to um especially when offering a model to the through the API because essentially the through the API because essentially the through the API because essentially moderation decides on what the user sees moderation decides on what the user sees moderation decides on what the user sees and what not, what which prompts we and what not, what which prompts we and what not, what which prompts we allow and which not. Um so this is an allow and which not. Um so this is an allow and which not. Um so this is an incredibly important piece of the user incredibly important piece of the user incredibly important piece of the user experience that we offer. If you have experience that we offer. If you have experience that we offer. If you have tried some video models um tried some video models um tried some video models um state-of-the-art video models in the state-of-the-art video models in the state-of-the-art video models in the past, you might be very aware of error past, you might be very aware of error past, you might be very aware of error messages like this. You put in an image messages like this. You put in an image messages like this. You put in an image of yourself and you want to animate of yourself and you want to animate of yourself and you want to animate yourself, but you're getting something yourself, but you're getting something yourself, but you're getting something like the images of videos might find like the images of videos might find like the images of videos might find likenesses of real people, etc. This is
-
likenesses of real people, etc. This is likenesses of real people, etc. This is obviously not a model limitation. This obviously not a model limitation. This obviously not a model limitation. This is purely moderation service, right? is purely moderation service, right? is purely moderation service, right? That doesn't allow you to generate That doesn't allow you to generate That doesn't allow you to generate something. And when talking to larger something. And when talking to larger something. And when talking to larger customers that want to do very specific customers that want to do very specific customers that want to do very specific thing thinks this is a massive issue and thing thinks this is a massive issue and thing thinks this is a massive issue and we need to customize and solve for this. we need to customize and solve for this. we need to customize and solve for this. Um, so two ways that I would like to Um, so two ways that I would like to Um, so two ways that I would like to talk about moderation. First is the most talk about moderation. First is the most talk about moderation. First is the most simple way and it is essentially in the simple way and it is essentially in the simple way and it is essentially in the configs. Um we have a modular setup to configs. Um we have a modular setup to configs. Um we have a modular setup to essentially customize configs. There is essentially customize configs. There is essentially customize configs. There is a fairly straightforward safety a fairly straightforward safety a fairly straightforward safety parameter here um that allows you to set parameter here um that allows you to set parameter here um that allows you to set that via API. And when customizing uh that via API. And when customizing uh that via API. And when customizing uh when customizing moderation for specific when customizing moderation for specific when customizing moderation for specific customers, we have a very fine grained customers, we have a very fine grained customers, we have a very fine grained confict setup that essentially allows us confict setup that essentially allows us confict setup that essentially allows us to set safety tolerance but also allow to set safety tolerance but also allow to set safety tolerance but also allow or block certain um certain categories or block certain um certain categories or block certain um certain categories specifically. So violence, sexual specifically. So violence, sexual specifically. So violence, sexual prompts, self harm, IP and likeness. And prompts, self harm, IP and likeness. And prompts, self harm, IP and likeness. And we also combine that with certain block we also combine that with certain block we also combine that with certain block lists and allow lists that we can set lists and allow lists that we can set lists and allow lists that we can set specifically for every customer. So specifically for every customer. So specifically for every customer. So think about, for example, a kids think about, for example, a kids think about, for example, a kids platform being much more strict, platform being much more strict, platform being much more strict, blocking violent, sexual, self harm, IP blocking violent, sexual, self harm, IP blocking violent, sexual, self harm, IP likeness, obviously, and then having likeness, obviously, and then having likeness, obviously, and then having even a specific block list. Uh fashion even a specific block list. Uh fashion even a specific block list. Uh fashion retailer being a little bit more a retailer being a little bit more a retailer being a little bit more a little bit more open, maybe allowing a little bit more open, maybe allowing a little bit more open, maybe allowing a little bit more of sexual leaning little bit more of sexual leaning little bit more of sexual leaning prompts. Um but blocking specifically prompts. Um but blocking specifically prompts. Um but blocking specifically competitor brands and allowing the own competitor brands and allowing the own competitor brands and allowing the own portfolio essentially only trying to portfolio essentially only trying to portfolio essentially only trying to generate their own their own stuff, of generate their own their own stuff, of generate their own their own stuff, of course, right? And then there might be a course, right? And then there might be a course, right? And then there might be a media or a movie studio that uh is much
-
media or a movie studio that uh is much media or a movie studio that uh is much more flexible, allows even some violence more flexible, allows even some violence more flexible, allows even some violence and specifically wants to allow certain and specifically wants to allow certain and specifically wants to allow certain licensed characters even if they're AP licensed characters even if they're AP licensed characters even if they're AP protected and wants to block every other protected and wants to block every other protected and wants to block every other character that is not that they don't character that is not that they don't character that is not that they don't have a license for. So this is fairly have a license for. So this is fairly have a license for. So this is fairly straightforward. We just set the config straightforward. We just set the config straightforward. We just set the config super modular, super simple. However, super modular, super simple. However, super modular, super simple. However, this doesn't work for every case um this doesn't work for every case um this doesn't work for every case um because we also deploy our models because we also deploy our models because we also deploy our models directly on customers infrastructure. So directly on customers infrastructure. So directly on customers infrastructure. So in these cases, we don't ship the entire in these cases, we don't ship the entire in these cases, we don't ship the entire moderation stack. we somehow need to moderation stack. we somehow need to moderation stack. we somehow need to find a way to actually embed the find a way to actually embed the find a way to actually embed the moderation and the safety into the model moderation and the safety into the model moderation and the safety into the model rates. And this is where for example um rates. And this is where for example um rates. And this is where for example um a slider law comes in which is the next a slider law comes in which is the next a slider law comes in which is the next sort of tool in the toolbox of um our sort of tool in the toolbox of um our sort of tool in the toolbox of um our FTEEs. A slider law tries to learn an FTEEs. A slider law tries to learn an FTEEs. A slider law tries to learn an abstract concept or tries to learn the abstract concept or tries to learn the abstract concept or tries to learn the specific abstract con concept that we're specific abstract con concept that we're specific abstract con concept that we're trying to block and tries to then turn trying to block and tries to then turn trying to block and tries to then turn around the rates that it learned to around the rates that it learned to around the rates that it learned to block that specifically. So let's take block that specifically. So let's take block that specifically. So let's take the example of a desexualization the example of a desexualization the example of a desexualization slider lower um which you have slider lower um which you have slider lower um which you have essentially simplified here on the essentially simplified here on the essentially simplified here on the slide. Essentially you try to tweak the slide. Essentially you try to tweak the slide. Essentially you try to tweak the model tweak the lower adaptation matrix model tweak the lower adaptation matrix model tweak the lower adaptation matrix you try to tweak that into the d the you try to tweak that into the d the you try to tweak that into the d the rates into the direction of the of the rates into the direction of the of the rates into the direction of the of the pure concept of sexualization. How we do pure concept of sexualization. How we do pure concept of sexualization. How we do this is that you have a generic data set this is that you have a generic data set this is that you have a generic data set of images for example of people in this of images for example of people in this of images for example of people in this case one image of a woman in a black case one image of a woman in a black case one image of a woman in a black dress and you feed that to every um dress and you feed that to every um dress and you feed that to every um training loop tries once with a modest training loop tries once with a modest training loop tries once with a modest caption and once with a sexualized caption and once with a sexualized caption and once with a sexualized caption. So then you essentially train caption. So then you essentially train caption. So then you essentially train the model. This is the image and you see the model. This is the image and you see the model. This is the image and you see it tries one once modest and once
-
it tries one once modest and once it tries one once modest and once sexualized and what what what is left in sexualized and what what what is left in sexualized and what what what is left in the rates is essentially the delta the rates is essentially the delta the rates is essentially the delta between both learnings. And the delta between both learnings. And the delta between both learnings. And the delta between both learnings is both both between both learnings is both both between both learnings is both both captions essentially contain there's a captions essentially contain there's a captions essentially contain there's a woman, black hair, black death, woman, black hair, black death, woman, black hair, black death, whatever. But the delta between those is whatever. But the delta between those is whatever. But the delta between those is essentially the concept of sexualization essentially the concept of sexualization essentially the concept of sexualization because the sexualized caption contains because the sexualized caption contains because the sexualized caption contains that but the modest caption doesn't. that but the modest caption doesn't. that but the modest caption doesn't. It's much more descriptive. So you end It's much more descriptive. So you end It's much more descriptive. So you end up with a law matrix with an adaptation up with a law matrix with an adaptation up with a law matrix with an adaptation matrix that essentially describes the matrix that essentially describes the matrix that essentially describes the concept of sexualization, right? Um and concept of sexualization, right? Um and concept of sexualization, right? Um and in inference we can now go ahead and in inference we can now go ahead and in inference we can now go ahead and apply that to our model weights with a apply that to our model weights with a apply that to our model weights with a negative alpha. So we turn the weights negative alpha. So we turn the weights negative alpha. So we turn the weights exactly around to go in exactly the exactly around to go in exactly the exactly around to go in exactly the opposite direction which then turned out opposite direction which then turned out opposite direction which then turned out in this specific case for uh for minus in this specific case for uh for minus in this specific case for uh for minus 20 nudity detections and um 68% judged 20 nudity detections and um 68% judged 20 nudity detections and um 68% judged less less sexualized image quality. less less sexualized image quality. less less sexualized image quality. while at the same time the image quality while at the same time the image quality while at the same time the image quality generally is a complete tie because we generally is a complete tie because we generally is a complete tie because we literally only turned around the concept literally only turned around the concept literally only turned around the concept of sexualization. Um so this is a nice of sexualization. Um so this is a nice of sexualization. Um so this is a nice tweak to really bake in moderation into tweak to really bake in moderation into tweak to really bake in moderation into the weights and this is something that the weights and this is something that the weights and this is something that you can also ship directly to customer you can also ship directly to customer you can also ship directly to customer they can deploy it on their own they can deploy it on their own they can deploy it on their own infrastructure um to have a much safer infrastructure um to have a much safer infrastructure um to have a much safer model without any deterministic uh model without any deterministic uh model without any deterministic uh moderation around it. So what we learned moderation around it. So what we learned moderation around it. So what we learned from the moderation piece is that first from the moderation piece is that first from the moderation piece is that first of all safety is an extreme extremely of all safety is an extreme extremely of all safety is an extreme extremely important product product surface. Um important product product surface. Um important product product surface. Um it's essentially the first thing that it's essentially the first thing that it's essentially the first thing that users interact with in some sense. It users interact with in some sense. It users interact with in some sense. It decides what users see what users don't
-
decides what users see what users don't decides what users see what users don't see and it can be a USB. It can even be see and it can be a USB. It can even be see and it can be a USB. It can even be a business model in certain certain a business model in certain certain a business model in certain certain cases. Modularity wins. cases. Modularity wins. cases. Modularity wins. Baking in moderation into model rates is Baking in moderation into model rates is Baking in moderation into model rates is cool. It's fun. It's a nice concept to cool. It's fun. It's a nice concept to cool. It's fun. It's a nice concept to build some sliders. However, it's build some sliders. However, it's build some sliders. However, it's expensive to train. it's static, you expensive to train. it's static, you expensive to train. it's static, you need to update it somehow. So in many need to update it somehow. So in many need to update it somehow. So in many cases or in most cases, we just go with cases or in most cases, we just go with cases or in most cases, we just go with simple confict config customizations simple confict config customizations simple confict config customizations because it's just modular. It's easy. because it's just modular. It's easy. because it's just modular. It's easy. It's scalable, right? Um it's it's much It's scalable, right? Um it's it's much It's scalable, right? Um it's it's much cheaper to actually set up. Um and yeah, cheaper to actually set up. Um and yeah, cheaper to actually set up. Um and yeah, as as I already mentioned, motivation as as I already mentioned, motivation as as I already mentioned, motivation and weights is possible. However, it is and weights is possible. However, it is and weights is possible. However, it is expensive and static. All right, let's expensive and static. All right, let's expensive and static. All right, let's move on. So, we already looked at the move on. So, we already looked at the move on. So, we already looked at the fine tune that changes behavior in the fine tune that changes behavior in the fine tune that changes behavior in the moderation direction, but I promised you moderation direction, but I promised you moderation direction, but I promised you some robots and I promised you some um some robots and I promised you some um some robots and I promised you some um flux controlling some things. So, flux controlling some things. So, flux controlling some things. So, finally, we're going to look at finally, we're going to look at finally, we're going to look at customization of model weights um and customization of model weights um and customization of model weights um and how to fine-tune flux for specifically how to fine-tune flux for specifically how to fine-tune flux for specifically action prediction. Um just to give you action prediction. Um just to give you action prediction. Um just to give you an overview, generally if you fine-tune an overview, generally if you fine-tune an overview, generally if you fine-tune a model, usually we come with a certain a model, usually we come with a certain a model, usually we come with a certain goal in mind. If we're talking about goal in mind. If we're talking about goal in mind. If we're talking about video models or or image models, we video models or or image models, we video models or or image models, we either fine-tune for a certain style, we either fine-tune for a certain style, we either fine-tune for a certain style, we fine-tune for a certain object that we fine-tune for a certain object that we fine-tune for a certain object that we want to present or a set of objects. We want to present or a set of objects. We want to present or a set of objects. We fine tune for a certain character that fine tune for a certain character that fine tune for a certain character that needs to be um extremely accurately um needs to be um extremely accurately um needs to be um extremely accurately um displayed or we fine tune for behavior.
-
displayed or we fine tune for behavior. displayed or we fine tune for behavior. We also have some models that we find in We also have some models that we find in We also have some models that we find in for example outpainting, specific local for example outpainting, specific local for example outpainting, specific local editing, desexualization, right? All of editing, desexualization, right? All of editing, desexualization, right? All of these are behavioral finetunes. And then these are behavioral finetunes. And then these are behavioral finetunes. And then if you go even further, you can even if you go even further, you can even if you go even further, you can even think about fine-tuning for certain to think about fine-tuning for certain to think about fine-tuning for certain to to add certain modalities if we even go to add certain modalities if we even go to add certain modalities if we even go a little bit into mid-training and post- a little bit into mid-training and post- a little bit into mid-training and post- training. Again, that's exactly what we training. Again, that's exactly what we training. Again, that's exactly what we did with Flux action. Flux action is a did with Flux action. Flux action is a did with Flux action. Flux action is a new um version of Flux uh Flux Backbone new um version of Flux uh Flux Backbone new um version of Flux uh Flux Backbone that we um continued mid-training and that we um continued mid-training and that we um continued mid-training and postrained and fine-tuned for postrained and fine-tuned for postrained and fine-tuned for specifically action prediction. This was specifically action prediction. This was specifically action prediction. This was actually released into open weights actually released into open weights actually released into open weights yesterday night. So check it out. It's yesterday night. So check it out. It's yesterday night. So check it out. It's brand new. It's in hugging phase. You brand new. It's in hugging phase. You brand new. It's in hugging phase. You can try it out. pull the rates and can try it out. pull the rates and can try it out. pull the rates and control your own layer robot with it. control your own layer robot with it. control your own layer robot with it. But just to give you a brief rundown on But just to give you a brief rundown on But just to give you a brief rundown on how it works. So we have an image model how it works. So we have an image model how it works. So we have an image model that essentially predicts frames um that essentially predicts frames um that essentially predicts frames um predicts video frames based on um based predicts video frames based on um based predicts video frames based on um based on a text input mostly, right? Um so how on a text input mostly, right? Um so how on a text input mostly, right? Um so how flux action works is that it has a text flux action works is that it has a text flux action works is that it has a text conditioning and it has a video conditioning and it has a video conditioning and it has a video conditioning. In this specific case, we conditioning. In this specific case, we conditioning. In this specific case, we work with a LE robot that completes pick work with a LE robot that completes pick work with a LE robot that completes pick and place task. The robot in this case and place task. The robot in this case and place task. The robot in this case has two camera sets. one is at the uh has two camera sets. one is at the uh has two camera sets. one is at the uh grabber hand, one is above um the top grabber hand, one is above um the top grabber hand, one is above um the top camera essentially. At the same time, we camera essentially. At the same time, we camera essentially. At the same time, we also fine-tuned flux action for the um also fine-tuned flux action for the um also fine-tuned flux action for the um for the embodiment of of the SL101, the for the embodiment of of the SL101, the for the embodiment of of the SL101, the Lobot robot arm specifically. So we Lobot robot arm specifically. So we Lobot robot arm specifically. So we actually teach flux that this specific actually teach flux that this specific actually teach flux that this specific robot has six axes and this the the the robot has six axes and this the the the robot has six axes and this the the the movement of these axes essentially movement of these axes essentially movement of these axes essentially represented by a vector that in every
-
represented by a vector that in every represented by a vector that in every frame defines the direction that every frame defines the direction that every frame defines the direction that every axis needs to move into. So essentially axis needs to move into. So essentially axis needs to move into. So essentially is simplified a sixdimensional vector is simplified a sixdimensional vector is simplified a sixdimensional vector that we're trying to predict here. Now that we're trying to predict here. Now that we're trying to predict here. Now we feed the the the vector the the we feed the the the vector the the we feed the the the vector the the embodiment knowledge is already embedded embodiment knowledge is already embedded embodiment knowledge is already embedded in the model rates. We feed in the video in the model rates. We feed in the video in the model rates. We feed in the video frames and the text conditioning which frames and the text conditioning which frames and the text conditioning which defines the task and the output defines the task and the output defines the task and the output essentially is a video of the essentially is a video of the essentially is a video of the continuation of these video frames and continuation of these video frames and continuation of these video frames and also um the actual vector that defines also um the actual vector that defines also um the actual vector that defines the direction that the robot needs to the direction that the robot needs to the direction that the robot needs to move in to actually complete the task. move in to actually complete the task. move in to actually complete the task. And then you can on the right see how And then you can on the right see how And then you can on the right see how the layer robot um is actually the layer robot um is actually the layer robot um is actually controlled by flux. I mean you'll have controlled by flux. I mean you'll have controlled by flux. I mean you'll have to trust me that it's actually to trust me that it's actually to trust me that it's actually controlled by flux but you can check controlled by flux but you can check controlled by flux but you can check that out yourself. Um, and it actually that out yourself. Um, and it actually that out yourself. Um, and it actually completes exactly the task that it's uh completes exactly the task that it's uh completes exactly the task that it's uh supposed to. You actually can see the supposed to. You actually can see the supposed to. You actually can see the the text bombs there. But um, so that is the text bombs there. But um, so that is the text bombs there. But um, so that is pretty fun. How that works in the pretty fun. How that works in the pretty fun. How that works in the background. Um, let's take a little look background. Um, let's take a little look background. Um, let's take a little look at the architecture. Um, is that we have at the architecture. Um, is that we have at the architecture. Um, is that we have um a video model. Traditionally, it um a video model. Traditionally, it um a video model. Traditionally, it takes certain frames and it feeds them takes certain frames and it feeds them takes certain frames and it feeds them through a variable autoenccoder. um the through a variable autoenccoder. um the through a variable autoenccoder. um the autoenccoder projects my frames, my autoenccoder projects my frames, my autoenccoder projects my frames, my pixel space into a latent space which is pixel space into a latent space which is pixel space into a latent space which is an abstract representation um of of an abstract representation um of of an abstract representation um of of these exact frames and then we go these exact frames and then we go these exact frames and then we go through this den noising process that through this den noising process that through this den noising process that gradually reduces noise and tries to gradually reduces noise and tries to gradually reduces noise and tries to predict these next video frames. Um predict these next video frames. Um predict these next video frames. Um usually when we just predict an image or usually when we just predict an image or usually when we just predict an image or a video we then have a a decoder as well
-
a video we then have a a decoder as well a video we then have a a decoder as well that then takes these latent these that then takes these latent these that then takes these latent these abstract representation and brings them abstract representation and brings them abstract representation and brings them back into the pixel space. Um now flux back into the pixel space. Um now flux back into the pixel space. Um now flux action is a world action model. Um that action is a world action model. Um that action is a world action model. Um that means that we don't even send these means that we don't even send these means that we don't even send these action vectors um through the viral action vectors um through the viral action vectors um through the viral autoenccoder into the latent space autoenccoder into the latent space autoenccoder into the latent space anymore but we just normalize them into anymore but we just normalize them into anymore but we just normalize them into the dimensionality of the um of the the dimensionality of the um of the the dimensionality of the um of the latent space and we directly latent space and we directly latent space and we directly dise the uh directly dise essentially dise the uh directly dise essentially dise the uh directly dise essentially the action vector. Um that is a huge the action vector. Um that is a huge the action vector. Um that is a huge example because we essentially dise both example because we essentially dise both example because we essentially dise both in the same transformer. Um in the end in the same transformer. Um in the end in the same transformer. Um in the end for the action case we actually don't for the action case we actually don't for the action case we actually don't even use the um decoder for the um for even use the um decoder for the um for even use the um decoder for the um for the frames and video part. We only the frames and video part. We only the frames and video part. We only denormalize the action part. Uh we have denormalize the action part. Uh we have denormalize the action part. Uh we have the robot actually apply it. Um and the robot actually apply it. Um and the robot actually apply it. Um and that's how we complete the task. that's how we complete the task. that's how we complete the task. Essentially the video needs encoder and Essentially the video needs encoder and Essentially the video needs encoder and decoder. Actions are already a vector. decoder. Actions are already a vector. decoder. Actions are already a vector. So we can actually naturally feed them So we can actually naturally feed them So we can actually naturally feed them through the through the model that has through the through the model that has through the through the model that has been mittrained and fine-tuned on been mittrained and fine-tuned on been mittrained and fine-tuned on exactly solving these tasks.
-
exactly solving these tasks. exactly solving these tasks. So flux flux to action uh flux 3 action So flux flux to action uh flux 3 action So flux flux to action uh flux 3 action um you can check out the research. You um you can check out the research. You um you can check out the research. You can download the rates here. It has can download the rates here. It has can download the rates here. It has dropped literally last night. It's right dropped literally last night. It's right dropped literally last night. It's right now topping the leaderboards and being now topping the leaderboards and being now topping the leaderboards and being being state-of-the-art in exactly um being state-of-the-art in exactly um being state-of-the-art in exactly um this action prediction space. So check this action prediction space. So check this action prediction space. So check it out. It's pretty exciting research. it out. It's pretty exciting research. it out. It's pretty exciting research. Um and we're pretty excited about Um and we're pretty excited about Um and we're pretty excited about venturing in this in this domain uh as a venturing in this in this domain uh as a venturing in this in this domain uh as a new modality. Um and this is also new modality. Um and this is also new modality. Um and this is also already where I'd like to wrap up. Um already where I'd like to wrap up. Um already where I'd like to wrap up. Um what we basically looked at is what we basically looked at is what we basically looked at is customizing flux for certain use cases. customizing flux for certain use cases. customizing flux for certain use cases. We looked at prompt upsampling and how We looked at prompt upsampling and how We looked at prompt upsampling and how prompt upsampling can be an incredibly prompt upsampling can be an incredibly prompt upsampling can be an incredibly efficient method to actually match user efficient method to actually match user efficient method to actually match user intent with the design requirements into intent with the design requirements into intent with the design requirements into the training distribution of flux to the training distribution of flux to the training distribution of flux to actually get out an image that at scale actually get out an image that at scale actually get out an image that at scale uh conforms certain design requirements. uh conforms certain design requirements. uh conforms certain design requirements. We also looked at moderation. We looked We also looked at moderation. We looked We also looked at moderation. We looked at sort of static deterministic or more at sort of static deterministic or more at sort of static deterministic or more or less deterministic moderation uh that or less deterministic moderation uh that or less deterministic moderation uh that modularly um is embedded in config. Um modularly um is embedded in config. Um modularly um is embedded in config. Um we also looked at how to embed we also looked at how to embed we also looked at how to embed moderation preferences into model moderation preferences into model moderation preferences into model weights directly using a using a slider weights directly using a using a slider weights directly using a using a slider law. And then finally finally we looked law. And then finally finally we looked law. And then finally finally we looked at how to um add a new modality to flux at how to um add a new modality to flux at how to um add a new modality to flux specifically looking at flux action and specifically looking at flux action and specifically looking at flux action and how to tweak flux to actually control a how to tweak flux to actually control a how to tweak flux to actually control a robot and predict the next action. If robot and predict the next action. If robot and predict the next action. If you want to take away one single thing you want to take away one single thing you want to take away one single thing across all these methods and this is across all these methods and this is across all these methods and this is obviously just scratching the surface.
-
obviously just scratching the surface. obviously just scratching the surface. If if you want to take away one single If if you want to take away one single If if you want to take away one single thing um about good customizations of AI thing um about good customizations of AI thing um about good customizations of AI models in general and also flux models in general and also flux models in general and also flux specifically it is that great custom specifically it is that great custom specifically it is that great custom customizations in all of these cases are customizations in all of these cases are customizations in all of these cases are as specific as needed and as general as as specific as needed and as general as as specific as needed and as general as possible. So they solve one specific possible. So they solve one specific possible. So they solve one specific problem to exactly the extent that it problem to exactly the extent that it problem to exactly the extent that it needs to solve them. But it also needs to solve them. But it also needs to solve them. But it also generalizes to actually maybe solve an generalizes to actually maybe solve an generalizes to actually maybe solve an entire domain, maybe be a solution that entire domain, maybe be a solution that entire domain, maybe be a solution that you can also roll out further. So with you can also roll out further. So with you can also roll out further. So with that, I'll wrap it up. I have another that, I'll wrap it up. I have another that, I'll wrap it up. I have another advertising block here. We are hiring advertising block here. We are hiring advertising block here. We are hiring for the product engineers, solution for the product engineers, solution for the product engineers, solution engineers. Um if you want, check that engineers. Um if you want, check that engineers. Um if you want, check that out. You can also find me um find me at out. You can also find me um find me at out. You can also find me um find me at the conference, find me online. And the conference, find me online. And the conference, find me online. And yeah, thank you very much for attending. Thank you, Yak. Thank you, Yak. Good job. Good job. Good job. All right. Okay. All right. Okay. All right. Okay. Yeah. Let's give it up once more for Yeah. Let's give it up once more for Yeah. Let's give it up once more for Yaku, please.
-
Thank you guys. I'm so excited for Black Thank you guys. I'm so excited for Black Forest Labs. And uh the nice thing about Forest Labs. And uh the nice thing about Forest Labs. And uh the nice thing about them as well is that they're based in them as well is that they're based in them as well is that they're based in Europe. They're based in Germany. And Europe. They're based in Germany. And Europe. They're based in Germany. And they're the at the front of everything they're the at the front of everything they're the at the front of everything generative media related. and now they generative media related. and now they generative media related. and now they do robotics which is so cool. But you do robotics which is so cool. But you do robotics which is so cool. But you know another company that is European know another company that is European know another company that is European based and also at the the frontier of AI based and also at the the frontier of AI based and also at the the frontier of AI is 11 Labs. So our next speaker is is 11 Labs. So our next speaker is is 11 Labs. So our next speaker is actually is going to talk to you about actually is going to talk to you about actually is going to talk to you about generative identity at scale with orbs generative identity at scale with orbs generative identity at scale with orbs and basically uh who who here has used and basically uh who who here has used and basically uh who who here has used the chat GPT voice mode or 11 labs. the chat GPT voice mode or 11 labs. the chat GPT voice mode or 11 labs. So you've seen that little circle, that So you've seen that little circle, that So you've seen that little circle, that orb. If you use that, you probably have orb. If you use that, you probably have orb. If you use that, you probably have used something that our next speaker has used something that our next speaker has used something that our next speaker has built. So please join me in welcoming to built. So please join me in welcoming to built. So please join me in welcoming to the stage creative experience engineer the stage creative experience engineer the stage creative experience engineer at 11 Labs, Dorian Lods.
-
Seems that internet's not working. Seems that internet's not working. That's a good start. Hi everyone. I'm Dorian. I'm a creative Hi everyone. I'm Dorian. I'm a creative engineer at 11 Labs but currently based engineer at 11 Labs but currently based engineer at 11 Labs but currently based in Paris. Um if you want to know more in Paris. Um if you want to know more in Paris. Um if you want to know more about what we do at 11 Labs or if you about what we do at 11 Labs or if you about what we do at 11 Labs or if you have any questions later uh please feel have any questions later uh please feel have any questions later uh please feel free to note the my socials. So free to note the my socials. So free to note the my socials. So I'm here today to talk to you about our I'm here today to talk to you about our I'm here today to talk to you about our how in a creative per perspective things how in a creative per perspective things how in a creative per perspective things have changed a little bit in the past have changed a little bit in the past have changed a little bit in the past few years and how how we deal with all few years and how how we deal with all few years and how how we deal with all these new tools set of tools AI tools these new tools set of tools AI tools these new tools set of tools AI tools enabled and so on. So uh just let's take enabled and so on. So uh just let's take enabled and so on. So uh just let's take uh let's let me just take a take a step uh let's let me just take a take a step uh let's let me just take a take a step back and back and back and tell you how I used to make the thing tell you how I used to make the thing tell you how I used to make the thing and now how I make the thing that makes and now how I make the thing that makes and now how I make the thing that makes the thing which is a bit of a change. So the thing which is a bit of a change. So the thing which is a bit of a change. So I came from an industry where used to I came from an industry where used to I came from an industry where used to make immersive creative experience make immersive creative experience make immersive creative experience websites are mainly based in WebGL or websites are mainly based in WebGL or websites are mainly based in WebGL or immersive installation or whatnot. They immersive installation or whatnot. They immersive installation or whatnot. They were all designed around a feeling uh were all designed around a feeling uh were all designed around a feeling uh refined over thousands of hours. Each refined over thousands of hours. Each refined over thousands of hours. Each one has a specific craft and and one has a specific craft and and one has a specific craft and and purpose.
-
purpose. purpose. Um during that journey um I interacted a Um during that journey um I interacted a Um during that journey um I interacted a lot with uh how you can streamline stuff lot with uh how you can streamline stuff lot with uh how you can streamline stuff with a system, how you can produce more with a system, how you can produce more with a system, how you can produce more or just expand your capabilities or just expand your capabilities or just expand your capabilities um with new tools for example comfy UI um with new tools for example comfy UI um with new tools for example comfy UI which is an amazing tool which is an amazing tool which is an amazing tool that helped creative new immersive like that helped creative new immersive like that helped creative new immersive like new deeper uh AI systems but still made new deeper uh AI systems but still made new deeper uh AI systems but still made for a single outcome and you will see for a single outcome and you will see for a single outcome and you will see that's important for that's important for that's important for And all of this background led me to a And all of this background led me to a And all of this background led me to a specific project I work on which is this specific project I work on which is this specific project I work on which is this guy. guy. guy. Um it's the OpenAI voice the face of Um it's the OpenAI voice the face of Um it's the OpenAI voice the face of OpenAI when they released the voice OpenAI when they released the voice OpenAI when they released the voice model. Um few notes but it's was the model. Um few notes but it's was the model. Um few notes but it's was the first iconic orbs that I had to do. uh first iconic orbs that I had to do. uh first iconic orbs that I had to do. uh everything was handcrafted, had to fit everything was handcrafted, had to fit everything was handcrafted, had to fit on brand and was the goal was to to find on brand and was the goal was to to find on brand and was the goal was to to find uh a way to visually represent the voice uh a way to visually represent the voice uh a way to visually represent the voice that you talk to uh while representing that you talk to uh while representing that you talk to uh while representing the brand. So pretty much the same stack the brand. So pretty much the same stack the brand. So pretty much the same stack than before. It was hand tuned, than before. It was hand tuned, than before. It was hand tuned, deterministic, everything you see on deterministic, everything you see on deterministic, everything you see on screen is made with a shader. Um so math screen is made with a shader. Um so math screen is made with a shader. Um so math actually makes the visuals, all the actually makes the visuals, all the actually makes the visuals, all the inputs, all the colors, all the way the inputs, all the colors, all the way the inputs, all the colors, all the way the thing behaves were carefully tuned. But thing behaves were carefully tuned. But thing behaves were carefully tuned. But it was just made for a single purpose.
-
it was just made for a single purpose. it was just made for a single purpose. And I spent few weeks designing one And I spent few weeks designing one And I spent few weeks designing one sphere. sphere. sphere. Then I had to design for 40 million Then I had to design for 40 million Then I had to design for 40 million voices which is a big change of scale if voices which is a big change of scale if voices which is a big change of scale if you ask me. So I joined 11 Labs. Uh for you ask me. So I joined 11 Labs. Uh for you ask me. So I joined 11 Labs. Uh for the one who don't know us, 11 Labs is a the one who don't know us, 11 Labs is a the one who don't know us, 11 Labs is a European firstbased company. Uh we are European firstbased company. Uh we are European firstbased company. Uh we are leading the way on we have our own leading the way on we have our own leading the way on we have our own frontier in the research team dedicated frontier in the research team dedicated frontier in the research team dedicated to any voice surface any voice to any voice surface any voice to any voice surface any voice interaction music any sounds to be interaction music any sounds to be interaction music any sounds to be honest but we do much more than that. We honest but we do much more than that. We honest but we do much more than that. We have our own creative platform where you have our own creative platform where you have our own creative platform where you can um dub make new clone your voice can um dub make new clone your voice can um dub make new clone your voice prompt another voice um and so on. So prompt another voice um and so on. So prompt another voice um and so on. So many of our or many many creators uh many of our or many many creators uh many of our or many many creators uh already use the platform. We have the already use the platform. We have the already use the platform. We have the web one, we have the app, we also have web one, we have the app, we also have web one, we have the app, we also have uh 11 agents now. So we handle our uh 11 agents now. So we handle our uh 11 agents now. So we handle our customer support call, we do our own customer support call, we do our own customer support call, we do our own orchestration, orchestration, orchestration, many many many things.
-
many many many things. many many many things. Um here are some key numbers because Um here are some key numbers because Um here are some key numbers because everyone loves numbers, right? Uh we everyone loves numbers, right? Uh we everyone loves numbers, right? Uh we support over 90 languages. Uh we already support over 90 languages. Uh we already support over 90 languages. Uh we already have millions of business developers, have millions of business developers, have millions of business developers, creators on the platform. Uh we crossed creators on the platform. Uh we crossed creators on the platform. Uh we crossed 600 millions of RR by end of June. uh 600 millions of RR by end of June. uh 600 millions of RR by end of June. uh current valuation is at 11 billion as of current valuation is at 11 billion as of current valuation is at 11 billion as of February and now we actually expanded February and now we actually expanded February and now we actually expanded the team to over 700 people across 50 the team to over 700 people across 50 the team to over 700 people across 50 countries. countries. countries. We have more than 10 million agent on We have more than 10 million agent on We have more than 10 million agent on the platform. It's a lot of numbers I the platform. It's a lot of numbers I the platform. It's a lot of numbers I know. Uh but some here are some fun know. Uh but some here are some fun know. Uh but some here are some fun ones. So every day we generate more than ones. So every day we generate more than ones. So every day we generate more than 34 years of audio with our speech to 34 years of audio with our speech to 34 years of audio with our speech to text model. Um so the scale factor is text model. Um so the scale factor is text model. Um so the scale factor is definitely here and definitely here and definitely here and uh one that uh we we made 25 million uh one that uh we we made 25 million uh one that uh we we made 25 million tracks with our music our music models tracks with our music our music models tracks with our music our music models and different music touch points and one and different music touch points and one and different music touch points and one that I really really u that I really that I really really u that I really that I really really u that I really love I think it's my favorite number love I think it's my favorite number love I think it's my favorite number here is that we so far helped to restore here is that we so far helped to restore here is that we so far helped to restore more than 11,000 voices for people who more than 11,000 voices for people who more than 11,000 voices for people who have lost them due to illness or injury.
-
have lost them due to illness or injury. have lost them due to illness or injury. We're super committed towards that goal. We're super committed towards that goal. We're super committed towards that goal. Our impact team is doing an exterior job Our impact team is doing an exterior job Our impact team is doing an exterior job uh to bring voice capabilities back to uh to bring voice capabilities back to uh to bring voice capabilities back to the people who actually need it for the people who actually need it for the people who actually need it for free. free. free. Um but more specifically I joined 11 Um but more specifically I joined 11 Um but more specifically I joined 11 Labs design. So the design team is Labs design. So the design team is Labs design. So the design team is pretty much all us here. Um pretty much all us here. Um pretty much all us here. Um we try to push the frontier of both the we try to push the frontier of both the we try to push the frontier of both the brand, the product, the way you interact brand, the product, the way you interact brand, the product, the way you interact with AI with all our stuff that we have with AI with all our stuff that we have with AI with all our stuff that we have on the platform and how we can on the platform and how we can on the platform and how we can streamline that. streamline that. streamline that. Um I join with my first uh challenge Um I join with my first uh challenge Um I join with my first uh challenge actually uh was to explore how voices actually uh was to explore how voices actually uh was to explore how voices which is our main interaction which is our main interaction which is our main interaction uh could have a visual identity and how uh could have a visual identity and how uh could have a visual identity and how that identity could work across the that identity could work across the that identity could work across the product. So how you bring more soul product. So how you bring more soul product. So how you bring more soul visual souls to our voices and how those visual souls to our voices and how those visual souls to our voices and how those souls should reflect the brand. souls should reflect the brand. souls should reflect the brand. Uh here's what we've been cooking in the Uh here's what we've been cooking in the Uh here's what we've been cooking in the past few months.
-
I might have spoiled a few things with I might have spoiled a few things with that video. that video. that video. Uh so um the pitch was simple but yet Uh so um the pitch was simple but yet Uh so um the pitch was simple but yet not so much. We believe that every voice not so much. We believe that every voice not so much. We believe that every voice is every face. Um and then is every face. Um and then is every face. Um and then we had 40 million voices on the we had 40 million voices on the we had 40 million voices on the platform. So uh users they can prom platform. So uh users they can prom platform. So uh users they can prom voice, they can clone their voice, you voice, they can clone their voice, you voice, they can clone their voice, you have we serve many type of different have we serve many type of different have we serve many type of different voices and so they all uh they all voices and so they all uh they all voices and so they all uh they all represent now more than 40 million represent now more than 40 million represent now more than 40 million voices. And how do you design something voices. And how do you design something voices. And how do you design something that is iconic enough that still that is iconic enough that still that is iconic enough that still translate the voice characteristic but translate the voice characteristic but translate the voice characteristic but keep it in one visual system that each keep it in one visual system that each keep it in one visual system that each voice needs to bring the identity of 11 voice needs to bring the identity of 11 voice needs to bring the identity of 11 labs and it still need to feel like 11 labs and it still need to feel like 11 labs and it still need to feel like 11 Labs. So it's all about feeling and Labs. So it's all about feeling and Labs. So it's all about feeling and developing it at scale. developing it at scale. developing it at scale. Uh first thing first it needed a Uh first thing first it needed a Uh first thing first it needed a reframe. So it could not just be me reframe. So it could not just be me reframe. So it could not just be me making the outputs anymore. It was me making the outputs anymore. It was me making the outputs anymore. It was me making the system that makes an output. making the system that makes an output. making the system that makes an output. very close to what we've been hearing in very close to what we've been hearing in very close to what we've been hearing in the past yesterday basically and today the past yesterday basically and today the past yesterday basically and today this morning uh with this morning uh with this morning uh with code factories and so on but here more code factories and so on but here more code factories and so on but here more on the design side uh the core concept on the design side uh the core concept on the design side uh the core concept of the or is it's actually a stack of of the or is it's actually a stack of of the or is it's actually a stack of two layers you have the image layer that two layers you have the image layer that two layers you have the image layer that carries the brand characteristic the carries the brand characteristic the carries the brand characteristic the voice that it's actually linked to the voice that it's actually linked to the voice that it's actually linked to the voice and then you have the realtime voice and then you have the realtime voice and then you have the realtime layer on the top which brings that to layer on the top which brings that to layer on the top which brings that to life at the interactivity had the living
-
life at the interactivity had the living life at the interactivity had the living piece to to the one in the back. So, a piece to to the one in the back. So, a piece to to the one in the back. So, a bit of a change compared to the first bit of a change compared to the first bit of a change compared to the first one I worked on at OpenAI. one I worked on at OpenAI. one I worked on at OpenAI. Uh, and that touch point then is diffuse Uh, and that touch point then is diffuse Uh, and that touch point then is diffuse everywhere across the board. Let's focus everywhere across the board. Let's focus everywhere across the board. Let's focus first on the image layer. Um, first on the image layer. Um, first on the image layer. Um, it's as it core it's linked to the it's as it core it's linked to the it's as it core it's linked to the voice. So, one voice equals one face. If voice. So, one voice equals one face. If voice. So, one voice equals one face. If you make a new voice, you make a new a you make a new voice, you make a new a you make a new voice, you make a new a new or So we need to we went down to the new or So we need to we went down to the new or So we need to we went down to the source and we we analyze the voice source and we we analyze the voice source and we we analyze the voice output for air like the audio output for output for air like the audio output for output for air like the audio output for every voice. We extract key metrics, key every voice. We extract key metrics, key every voice. We extract key metrics, key numbers um that that then fall under numbers um that that then fall under numbers um that that then fall under different category of pitch, energy, different category of pitch, energy, different category of pitch, energy, expressiveness that that drive the rest expressiveness that that drive the rest expressiveness that that drive the rest of the pipeline where at this core it's of the pipeline where at this core it's of the pipeline where at this core it's very much linked to the voice. Uh in the very much linked to the voice. Uh in the very much linked to the voice. Uh in the first place we wanted to use embeddings first place we wanted to use embeddings first place we wanted to use embeddings directly but canary directly but canary directly but canary became a bit complicated later. Um became a bit complicated later. Um became a bit complicated later. Um to work to design the process and to to work to design the process and to to work to design the process and to really uh expand or to explore how much really uh expand or to explore how much really uh expand or to explore how much range do we need to work with. We range do we need to work with. We range do we need to work with. We extracted a set of 200,000 voices the extracted a set of 200,000 voices the extracted a set of 200,000 voices the most popular we have on the platform.
-
most popular we have on the platform. most popular we have on the platform. And we realized for example that they And we realized for example that they And we realized for example that they are not distributed evenly. you have are not distributed evenly. you have are not distributed evenly. you have much more deep voices mainly for much more deep voices mainly for much more deep voices mainly for narration purposes. So we really had to narration purposes. So we really had to narration purposes. So we really had to focus in those kind of categories where focus in those kind of categories where focus in those kind of categories where the other ones might fall under like the other ones might fall under like the other ones might fall under like niche where it's super high super niche where it's super high super niche where it's super high super energetic like maybe cartoonish style of energetic like maybe cartoonish style of energetic like maybe cartoonish style of voices. But uh that really helped us a voices. But uh that really helped us a voices. But uh that really helped us a lot during the process to to explore the lot during the process to to explore the lot during the process to to explore the thing at scale. Uh first thing first we thing at scale. Uh first thing first we thing at scale. Uh first thing first we wanted to um the voices to feel like the wanted to um the voices to feel like the wanted to um the voices to feel like the brand. So we needed to teach a model to brand. So we needed to teach a model to brand. So we needed to teach a model to speak 11 laps. Each voice is unique. The speak 11 laps. Each voice is unique. The speak 11 laps. Each voice is unique. The goal is to create a generate new voice goal is to create a generate new voice goal is to create a generate new voice per new visual new or per voice. So we per new visual new or per voice. So we per new visual new or per voice. So we needed to find a way so a model can needed to find a way so a model can needed to find a way so a model can actually learn our design language and actually learn our design language and actually learn our design language and how that translate. Uh so of the chef how that translate. Uh so of the chef how that translate. Uh so of the chef they don't really know your brand. Um so they don't really know your brand. Um so they don't really know your brand. Um so we built a synthetic data set to train a we built a synthetic data set to train a we built a synthetic data set to train a specific lura on the top and actually specific lura on the top and actually specific lura on the top and actually there's pretty fun thing that we there's pretty fun thing that we there's pretty fun thing that we explored is that we steer the inference explored is that we steer the inference explored is that we steer the inference uh with the caption. So on top of the of uh with the caption. So on top of the of uh with the caption. So on top of the of the action words we hand selected the action words we hand selected the action words we hand selected multiple type of colors and things that multiple type of colors and things that multiple type of colors and things that we wanted to attribute to different kind we wanted to attribute to different kind we wanted to attribute to different kind of voices. Let's say high pitch, mid of voices. Let's say high pitch, mid of voices. Let's say high pitch, mid pitch, medium pitch. Um and then we we pitch, medium pitch. Um and then we we pitch, medium pitch. Um and then we we put that in the caption. So during the put that in the caption. So during the put that in the caption. So during the training um the model could actually training um the model could actually training um the model could actually steer a bit and like develop a bit more steer a bit and like develop a bit more steer a bit and like develop a bit more uh the varity and all those keywords
-
uh the varity and all those keywords uh the varity and all those keywords were actually were put back into were actually were put back into were actually were put back into inventory. So at inference level when inventory. So at inference level when inventory. So at inference level when you rebuild the prompt the model can you rebuild the prompt the model can you rebuild the prompt the model can choose exactly the same keywords like choose exactly the same keywords like choose exactly the same keywords like sub keyword that has been in the sub keyword that has been in the sub keyword that has been in the training data. That was really fun. Um training data. That was really fun. Um training data. That was really fun. Um we explore different type of model we explore different type of model we explore different type of model uh one from our friends. uh one from our friends. uh one from our friends. So it's a good segue actually. Um so we So it's a good segue actually. Um so we So it's a good segue actually. Um so we we try to yeah we explored different we try to yeah we explored different we try to yeah we explored different kind of open weight model. The goal was kind of open weight model. The goal was kind of open weight model. The goal was to f to find a good sweet spot between to f to find a good sweet spot between to f to find a good sweet spot between uh between like inferencing cost and uh between like inferencing cost and uh between like inferencing cost and development at large scale while still development at large scale while still development at large scale while still maintaining the quality. maintaining the quality. maintaining the quality. um post inference after you generate the um post inference after you generate the um post inference after you generate the base image the idea is that basically base image the idea is that basically base image the idea is that basically you blur it and you crop it inside the you blur it and you crop it inside the you blur it and you crop it inside the inside the base image. So each variation inside the base image. So each variation inside the base image. So each variation each or has a really really good each or has a really really good each or has a really really good variation into it. Um that's what we variation into it. Um that's what we variation into it. Um that's what we call post inference. Um and that call post inference. Um and that call post inference. Um and that prepares for the rest of the shader. So prepares for the rest of the shader. So prepares for the rest of the shader. So let's now focus on the real time layer let's now focus on the real time layer let's now focus on the real time layer of those orbs. Um during the design of those orbs. Um during the design of those orbs. Um during the design phase we we actually used code as a core phase we we actually used code as a core phase we we actually used code as a core design medium. So it lives as a shader.
-
design medium. So it lives as a shader. design medium. So it lives as a shader. Uh it takes the image in the back and Uh it takes the image in the back and Uh it takes the image in the back and has multiple layers on top of it. has multiple layers on top of it. has multiple layers on top of it. Everything was uh made through an Everything was uh made through an Everything was uh made through an interface where all the design team was interface where all the design team was interface where all the design team was able to tweak and like see how we can able to tweak and like see how we can able to tweak and like see how we can feel things. Uh some parameters are feel things. Uh some parameters are feel things. Uh some parameters are tweaked with the sound itself with the tweaked with the sound itself with the tweaked with the sound itself with the audio feedback and so on. Um, WebJ is audio feedback and so on. Um, WebJ is audio feedback and so on. Um, WebJ is the interactive experience on the the interactive experience on the the interactive experience on the platform on the on the marketing website platform on the on the marketing website platform on the on the marketing website but uh uh but we converted this to but uh uh but we converted this to but uh uh but we converted this to OpenGL support for server server side OpenGL support for server server side OpenGL support for server server side rendering for many previews or rendering for many previews or rendering for many previews or other touch points that cannot actually other touch points that cannot actually other touch points that cannot actually run the shader themsel. run the shader themsel. run the shader themsel. Uh the shader itself is composed of Uh the shader itself is composed of Uh the shader itself is composed of basically six six stacks together. You basically six six stacks together. You basically six six stacks together. You have a realtime fleet simulation that have a realtime fleet simulation that have a realtime fleet simulation that has um the reaper effect that's linked has um the reaper effect that's linked has um the reaper effect that's linked to the audio. It's really bring the to the audio. It's really bring the to the audio. It's really bring the organicness into it. The goal I might organicness into it. The goal I might organicness into it. The goal I might have skipped that but the goal was to have skipped that but the goal was to have skipped that but the goal was to anchor the orbs into nature. Nature is anchor the orbs into nature. Nature is anchor the orbs into nature. Nature is the main characteristics of the orbs.
-
the main characteristics of the orbs. the main characteristics of the orbs. You wanted to fill them to feel organic. You wanted to fill them to feel organic. You wanted to fill them to feel organic. Then you have multiple phase of of noise Then you have multiple phase of of noise Then you have multiple phase of of noise on the top to bring it to life. FBM and on the top to bring it to life. FBM and on the top to bring it to life. FBM and so on. And we have a a further sound so on. And we have a a further sound so on. And we have a a further sound wave that is actually a link with the wave that is actually a link with the wave that is actually a link with the one we had before and the the the brand one we had before and the the the brand one we had before and the the the brand grain that we put on many of our of our grain that we put on many of our of our grain that we put on many of our of our touch points. Um so all of that led to touch points. Um so all of that led to touch points. Um so all of that led to the shader playground like I said before the shader playground like I said before the shader playground like I said before it was the place where most of the craft it was the place where most of the craft it was the place where most of the craft was happening early on. Um but behind was happening early on. Um but behind was happening early on. Um but behind that we actually developed a reviewing that we actually developed a reviewing that we actually developed a reviewing tool to debug everything at scale. So tool to debug everything at scale. So tool to debug everything at scale. So one voice equal one output. So you also one voice equal one output. So you also one voice equal one output. So you also mean that you cannot choose the output mean that you cannot choose the output mean that you cannot choose the output and every output had to be right. So we and every output had to be right. So we and every output had to be right. So we quickly realized that in order to be quickly realized that in order to be quickly realized that in order to be able to maintain quality across the able to maintain quality across the able to maintain quality across the board, we needed to redo it at scale. Uh board, we needed to redo it at scale. Uh board, we needed to redo it at scale. Uh and that was really a key factor because and that was really a key factor because and that was really a key factor because you might end up with your debug ones you might end up with your debug ones you might end up with your debug ones that are super nice and then as soon as that are super nice and then as soon as that are super nice and then as soon as you run it at scale, it's actually fall you run it at scale, it's actually fall you run it at scale, it's actually fall down. So building internal tools to down. So building internal tools to down. So building internal tools to review things at scale really really review things at scale really really review things at scale really really make the make a change.
-
make the make a change. make the make a change. Um that's the less fun part. Uh so I had Um that's the less fun part. Uh so I had Um that's the less fun part. Uh so I had an we had an amazing an amazing op an we had an amazing an amazing op an we had an amazing an amazing op process and amazing workflows. Uh and process and amazing workflows. Uh and process and amazing workflows. Uh and then I run through uh finance. Then they then I run through uh finance. Then they then I run through uh finance. Then they told me that okay D is super nice but told me that okay D is super nice but told me that okay D is super nice but for millions times 10 cents it's way too for millions times 10 cents it's way too for millions times 10 cents it's way too much. uh so you need to find a way to much. uh so you need to find a way to much. uh so you need to find a way to reduce both latency and the cost. So we reduce both latency and the cost. So we reduce both latency and the cost. So we went back to the uh to the chart that I went back to the uh to the chart that I went back to the uh to the chart that I shared before where we have all the shared before where we have all the shared before where we have all the distribution of voices in the platform distribution of voices in the platform distribution of voices in the platform and we actually um took this as the core and we actually um took this as the core and we actually um took this as the core source. So instead of generating a new source. So instead of generating a new source. So instead of generating a new image every time, we generated a image every time, we generated a image every time, we generated a gigantic atlas at once. And then from gigantic atlas at once. And then from gigantic atlas at once. And then from this atlas, we actually crop a bit this atlas, we actually crop a bit this atlas, we actually crop a bit deeper in each of the image. And so uh deeper in each of the image. And so uh deeper in each of the image. And so uh that help us to not generate every time that help us to not generate every time that help us to not generate every time we we make a new voices but still we we make a new voices but still we we make a new voices but still maintaining a uniqueness. So every maintaining a uniqueness. So every maintaining a uniqueness. So every voices is unique because we rotate it, voices is unique because we rotate it, voices is unique because we rotate it, we crop it differently, we change a bit we crop it differently, we change a bit we crop it differently, we change a bit to U and so on. uh and they are still to U and so on. uh and they are still to U and so on. uh and they are still really much um locked into the audio really much um locked into the audio really much um locked into the audio characteristics. That helped to save characteristics. That helped to save characteristics. That helped to save more than 3.6 million. So pretty good.
-
more than 3.6 million. So pretty good. more than 3.6 million. So pretty good. Uh run it 10 to 15 times faster and have Uh run it 10 to 15 times faster and have Uh run it 10 to 15 times faster and have less than three seconds per voice to less than three seconds per voice to less than three seconds per voice to generate the visual. So it have it's generate the visual. So it have it's generate the visual. So it have it's happening on the back as soon as you as happening on the back as soon as you as happening on the back as soon as you as you make a new voice and then it run you make a new voice and then it run you make a new voice and then it run through all the the process audio through all the the process audio through all the the process audio process. It spin up it spin up the task process. It spin up it spin up the task process. It spin up it spin up the task on the infra and then generate the one. on the infra and then generate the one. on the infra and then generate the one. So less than 3 seconds actually make it So less than 3 seconds actually make it So less than 3 seconds actually make it not visible for the user. not visible for the user. not visible for the user. Um and so here's a recap of the pipeline Um and so here's a recap of the pipeline Um and so here's a recap of the pipeline for the voices. You have the voice for the voices. You have the voice for the voices. You have the voice analysis in the first place that maps analysis in the first place that maps analysis in the first place that maps that map key metrics to uh visual that map key metrics to uh visual that map key metrics to uh visual mapping like a medium pen and energy mapping like a medium pen and energy mapping like a medium pen and energy level texture and so on. Then we have level texture and so on. Then we have level texture and so on. Then we have the entire generative pipeline I just the entire generative pipeline I just the entire generative pipeline I just showed you that prepared the assets, showed you that prepared the assets, showed you that prepared the assets, prepared the shader base, um the prepared the shader base, um the prepared the shader base, um the texture, the preview, the thumbnail, texture, the preview, the thumbnail, texture, the preview, the thumbnail, animated MP4, all of these diffused CDN animated MP4, all of these diffused CDN animated MP4, all of these diffused CDN and so on. And at the end you have 40 and so on. And at the end you have 40 and so on. And at the end you have 40 million plus outputs uh that serve million plus outputs uh that serve million plus outputs uh that serve pretty much everywhere as you can as pretty much everywhere as you can as pretty much everywhere as you can as you're going to see in a few.
-
you're going to see in a few. you're going to see in a few. Um quality development was really really Um quality development was really really Um quality development was really really important. Uh the concept the core important. Uh the concept the core important. Uh the concept the core design concept was was cool but we design concept was was cool but we design concept was was cool but we wanted it to feel alive everywhere to wanted it to feel alive everywhere to wanted it to feel alive everywhere to have to to to follow all the pipelines have to to to follow all the pipelines have to to to follow all the pipelines of the orbs that we have on the of the orbs that we have on the of the orbs that we have on the platform. Uh here you see on the left platform. Uh here you see on the left platform. Uh here you see on the left you have the the voice cloning um the you have the the voice cloning um the you have the the voice cloning um the voice cloning interface. So the orbs voice cloning interface. So the orbs voice cloning interface. So the orbs change depending on your voice. Uh we change depending on your voice. Uh we change depending on your voice. Uh we wanted to maintain quality across iOS, wanted to maintain quality across iOS, wanted to maintain quality across iOS, Android, uh it's the one you see in the Android, uh it's the one you see in the Android, uh it's the one you see in the middle. Every shader had to look the middle. Every shader had to look the middle. Every shader had to look the same. Super important for the brand. Uh same. Super important for the brand. Uh same. Super important for the brand. Uh and we needed to review it at scale. So and we needed to review it at scale. So and we needed to review it at scale. So on the right side, you have like the on the right side, you have like the on the right side, you have like the admin interface where you can review admin interface where you can review admin interface where you can review every voice that we have. every voice that we have. every voice that we have. Um all of that started to grow. Uh we Um all of that started to grow. Uh we Um all of that started to grow. Uh we made a dedicated uh service. So we have made a dedicated uh service. So we have made a dedicated uh service. So we have a dedicated a dedicated a dedicated um cloud service for making our orb um cloud service for making our orb um cloud service for making our orb visuals. Now it's called Picasso because visuals. Now it's called Picasso because visuals. Now it's called Picasso because it's painter and and I'm French.
-
it's painter and and I'm French. it's painter and and I'm French. Uh it's a shared API across all our Uh it's a shared API across all our Uh it's a shared API across all our touch points. We we push those orbs touch points. We we push those orbs touch points. We we push those orbs everywhere. So in the core app the everywhere. So in the core app the everywhere. So in the core app the marketing website iOS and Android app uh marketing website iOS and Android app uh marketing website iOS and Android app uh 11 music which are which is our music 11 music which are which is our music 11 music which are which is our music brand branding and expanded into the brand branding and expanded into the brand branding and expanded into the brand core itself. So we made brand brand core itself. So we made brand brand core itself. So we made brand tools motion designs everything now used tools motion designs everything now used tools motion designs everything now used orbs the carrier brand the carrier of orbs the carrier brand the carrier of orbs the carrier brand the carrier of voice and the carrier of face. Uh it was voice and the carrier of face. Uh it was voice and the carrier of face. Uh it was backward compatible because we had backward compatible because we had backward compatible because we had already previous type of orbs in the already previous type of orbs in the already previous type of orbs in the platform. So we were able to run a back platform. So we were able to run a back platform. So we were able to run a back field of 40 million voices. Uh that back field of 40 million voices. Uh that back field of 40 million voices. Uh that back that back field run for a few days. Um that back field run for a few days. Um that back field run for a few days. Um really stressful moment but it worked at really stressful moment but it worked at really stressful moment but it worked at the end. Uh and now they are everywhere. the end. Uh and now they are everywhere. the end. Uh and now they are everywhere. They we cannot hold them anymore. They They we cannot hold them anymore. They They we cannot hold them anymore. They they are part of our um homepage. Uh the they are part of our um homepage. Uh the they are part of our um homepage. Uh the home page the first thing you see on the home page the first thing you see on the home page the first thing you see on the homepage is the orbs. Um they are part homepage is the orbs. Um they are part homepage is the orbs. Um they are part of the app itself. So you see for of the app itself. So you see for of the app itself. So you see for example here you have deeper voices has example here you have deeper voices has example here you have deeper voices has a tendency to be a bit darker. if they a tendency to be a bit darker. if they a tendency to be a bit darker. if they have more energy, they will have like have more energy, they will have like have more energy, they will have like blue hints on and so on. So, I really blue hints on and so on. So, I really blue hints on and so on. So, I really invite you to play with the platform and invite you to play with the platform and invite you to play with the platform and to try to filter voices to try to filter voices to try to filter voices um to to see how each characteristic of um to to see how each characteristic of um to to see how each characteristic of voice translate into visuals.
-
voice translate into visuals. voice translate into visuals. Uh more renders Uh more renders Uh more renders and they actually came out of the screen and they actually came out of the screen and they actually came out of the screen which is really really cool and which is really really cool and which is really really cool and something I'm super proud of. uh we use something I'm super proud of. uh we use something I'm super proud of. uh we use them as a core brand uh elements toward them as a core brand uh elements toward them as a core brand uh elements toward everything we do. So on the left here everything we do. So on the left here everything we do. So on the left here it's a autoome campaign we made in char it's a autoome campaign we made in char it's a autoome campaign we made in char uh in France uh where you could have uh in France uh where you could have uh in France uh where you could have spot orbs based on the on the on the spot orbs based on the on the on the spot orbs based on the on the on the branding of some of our clients. Here is branding of some of our clients. Here is branding of some of our clients. Here is Alan for example or on the right side Alan for example or on the right side Alan for example or on the right side you have London uh where we took over um you have London uh where we took over um you have London uh where we took over um one of the main square of London with one of the main square of London with one of the main square of London with train line train line train line color DNA and so on. So once again color DNA and so on. So once again color DNA and so on. So once again another clients of us and then we the another clients of us and then we the another clients of us and then we the orbs as itself doesn't carry identity orbs as itself doesn't carry identity orbs as itself doesn't carry identity but it's a shape shifter so we can also but it's a shape shifter so we can also but it's a shape shifter so we can also use it for expressing use it for expressing use it for expressing um so illustrating the the needs of our um so illustrating the the needs of our um so illustrating the the needs of our clients towards the thing we do. clients towards the thing we do. clients towards the thing we do. And here another example in the Paris And here another example in the Paris And here another example in the Paris metro. It was really fun to go to work metro. It was really fun to go to work metro. It was really fun to go to work and then see that on the metro.
-
and then see that on the metro. and then see that on the metro. Um so here is what we've learned uh Um so here is what we've learned uh Um so here is what we've learned uh during this little process. Uh using during this little process. Uh using during this little process. Uh using machine learning as a core creative machine learning as a core creative machine learning as a core creative medium change how you design. It changed medium change how you design. It changed medium change how you design. It changed the way you think. You don't think about the way you think. You don't think about the way you think. You don't think about only the outputs but you need to think only the outputs but you need to think only the outputs but you need to think about how you create that. um machine about how you create that. um machine about how you create that. um machine learning AI tools are basically learning AI tools are basically learning AI tools are basically superpowers but they need to be used superpowers but they need to be used superpowers but they need to be used carefully uh if every design thing that carefully uh if every design thing that carefully uh if every design thing that you do don't overlook it now the cost of you do don't overlook it now the cost of you do don't overlook it now the cost of making the thing has become so much making the thing has become so much making the thing has become so much lower that we have tendency to ship fast lower that we have tendency to ship fast lower that we have tendency to ship fast and fast and fast and we believe as a and fast and fast and we believe as a and fast and fast and we believe as a design team that the the edge is like design team that the the edge is like design team that the the edge is like the quality edge is going to make you the quality edge is going to make you the quality edge is going to make you actually um want more of brand awareness actually um want more of brand awareness actually um want more of brand awareness brand quality and so on So uh you need brand quality and so on So uh you need brand quality and so on So uh you need to own the pipeline. You need to own all to own the pipeline. You need to own all to own the pipeline. You need to own all the outputs like if they were your own. the outputs like if they were your own. the outputs like if they were your own. Uh the crash shifted into the data and Uh the crash shifted into the data and Uh the crash shifted into the data and rules. Uh this is how you trying to rules. Uh this is how you trying to rules. Uh this is how you trying to maintain all the AI stuff that we have maintain all the AI stuff that we have maintain all the AI stuff that we have and to be sure that they follow a and to be sure that they follow a and to be sure that they follow a branding rule, a feeling rule and so on.
-
branding rule, a feeling rule and so on. branding rule, a feeling rule and so on. Uh every stage had to preserve the brand Uh every stage had to preserve the brand Uh every stage had to preserve the brand vision. Like I said, branding is getting vision. Like I said, branding is getting vision. Like I said, branding is getting more and more key. Um and the patterns more and more key. Um and the patterns more and more key. Um and the patterns revealed what a single output hit means revealed what a single output hit means revealed what a single output hit means that we need to debug at scale. If that we need to debug at scale. If that we need to debug at scale. If you're making a if you're making a a you're making a if you're making a a you're making a if you're making a a pipeline a process that it's meant to to pipeline a process that it's meant to to pipeline a process that it's meant to to be deployed 40 million times you need be deployed 40 million times you need be deployed 40 million times you need for example to to debug it at scale. So for example to to debug it at scale. So for example to to debug it at scale. So don't you don't hesitate to run multiple don't you don't hesitate to run multiple don't you don't hesitate to run multiple run of of generation at a smaller scale run of of generation at a smaller scale run of of generation at a smaller scale to be sure that every you can own every to be sure that every you can own every to be sure that every you can own every single output. single output. single output. So I still care about every output like So I still care about every output like So I still care about every output like I said. Um they are all my they are all I said. Um they are all my they are all I said. Um they are all my they are all I'm all proud of them. There is not any I'm all proud of them. There is not any I'm all proud of them. There is not any one that I don't like. Uh but instead of one that I don't like. Uh but instead of one that I don't like. Uh but instead of designing them now I designed the designing them now I designed the designing them now I designed the condition that produce it and that's condition that produce it and that's condition that produce it and that's really really important and it's really really important and it's really really important and it's something that we've seen on the code something that we've seen on the code something that we've seen on the code side and we also see it now on the side and we also see it now on the side and we also see it now on the design side. design side. design side. That's pretty much it for me. Uh thank That's pretty much it for me. Uh thank That's pretty much it for me. Uh thank you for your attention and if you have you for your attention and if you have you for your attention and if you have any questions please don't hesitate to any questions please don't hesitate to any questions please don't hesitate to reach out. uh check out our website or reach out. uh check out our website or reach out. uh check out our website or apps or everything and uh thank you very apps or everything and uh thank you very apps or everything and uh thank you very much.
-
All right, thank you so much, Dorian. All right, thank you so much, Dorian. Let's give it up one more time for Let's give it up one more time for Let's give it up one more time for Dorian, please. Thank you. All right. Dorian, please. Thank you. All right. Dorian, please. Thank you. All right. Okay, so show of hands. Who's here has Okay, so show of hands. Who's here has Okay, so show of hands. Who's here has uh used open weights models before? uh used open weights models before? uh used open weights models before? Okay, cool. So if you did, you probably Okay, cool. So if you did, you probably Okay, cool. So if you did, you probably have used the tools from our next have used the tools from our next have used the tools from our next speaker company. So up next is someone speaker company. So up next is someone speaker company. So up next is someone from hugging face. He's going to dive from hugging face. He's going to dive from hugging face. He's going to dive deep into how a transformers model loads deep into how a transformers model loads deep into how a transformers model loads in VLM. So yeah, hold on to your seats in VLM. So yeah, hold on to your seats in VLM. So yeah, hold on to your seats guys. That's going it's going to go guys. That's going it's going to go guys. That's going it's going to go deep. So please join me in welcoming to deep. So please join me in welcoming to deep. So please join me in welcoming to machine learning engineer at Hogenface, machine learning engineer at Hogenface, machine learning engineer at Hogenface, Harry Miller. Harry Miller. Harry Miller. All right.
-
Okay. So, hello everyone. I'm Harry and Okay. So, hello everyone. I'm Harry and as you just heard I'm a machine learning as you just heard I'm a machine learning as you just heard I'm a machine learning engineer at HuggingFace and I'm also a engineer at HuggingFace and I'm also a engineer at HuggingFace and I'm also a maintainer of VLM maintainer of VLM maintainer of VLM and in today's talk I'm going to talk to and in today's talk I'm going to talk to and in today's talk I'm going to talk to you about how a hugging face you about how a hugging face you about how a hugging face transformers model can be loaded transformers model can be loaded transformers model can be loaded directly into VLM and why because of directly into VLM and why because of directly into VLM and why because of what we do to it will run just as fast what we do to it will run just as fast what we do to it will run just as fast as if the model had been dedicated had a as if the model had been dedicated had a as if the model had been dedicated had a dedicated implementation in VLM itself. dedicated implementation in VLM itself. dedicated implementation in VLM itself. But first a little bit of context in But first a little bit of context in But first a little bit of context in case you haven't heard of VLM before. case you haven't heard of VLM before. case you haven't heard of VLM before. Uh, VLM is an open-source LLM inference Uh, VLM is an open-source LLM inference Uh, VLM is an open-source LLM inference engine. It came from the UC Berkeley engine. It came from the UC Berkeley engine. It came from the UC Berkeley Skylab and is now a PyTorch Foundation Skylab and is now a PyTorch Foundation Skylab and is now a PyTorch Foundation project. Um, and the key things that it project. Um, and the key things that it project. Um, and the key things that it brings is page detention which allows brings is page detention which allows brings is page detention which allows you to KV cache efficiently. Um, you to KV cache efficiently. Um, you to KV cache efficiently. Um, continuous batching which allows you to continuous batching which allows you to continuous batching which allows you to process a request as soon as it arrives process a request as soon as it arrives process a request as soon as it arrives rather than waiting for the next batch rather than waiting for the next batch rather than waiting for the next batch to begin. Fast kernels scale out through to begin. Fast kernels scale out through to begin. Fast kernels scale out through many different axes of parallelism, many different axes of parallelism, many different axes of parallelism, portability by running on any hardware portability by running on any hardware portability by running on any hardware um in any environment and then support um in any environment and then support um in any environment and then support broad support for hundreds of models broad support for hundreds of models broad support for hundreds of models both through dedicated implementations both through dedicated implementations both through dedicated implementations and the transformers modeling back end and the transformers modeling back end and the transformers modeling back end which you'll hear more about imminently.
-
which you'll hear more about imminently. which you'll hear more about imminently. And so then the transformers modeling And so then the transformers modeling And so then the transformers modeling back end um just again more context uh back end um just again more context uh back end um just again more context uh is what lets you take this transformers is what lets you take this transformers is what lets you take this transformers modeling code from hugging face modeling code from hugging face modeling code from hugging face transformers and run it directly in VLM. transformers and run it directly in VLM. transformers and run it directly in VLM. So the model definition the config and So the model definition the config and So the model definition the config and the weights all come from transformers the weights all come from transformers the weights all come from transformers and then the runtime and the all the and then the runtime and the all the and then the runtime and the all the performance critical modules come from performance critical modules come from performance critical modules come from VLM. So you kind of get a best of both. VLM. So you kind of get a best of both. VLM. So you kind of get a best of both. And so why might you end up using the And so why might you end up using the And so why might you end up using the transformers modeling back end? Well, transformers modeling back end? Well, transformers modeling back end? Well, there are three main reasons. One of there are three main reasons. One of there are three main reasons. One of them might be that you're already using them might be that you're already using them might be that you're already using the transformers ecosystem. And so you the transformers ecosystem. And so you the transformers ecosystem. And so you don't want to reimplement your model don't want to reimplement your model don't want to reimplement your model again for each inference engine that you again for each inference engine that you again for each inference engine that you might want to support. So through might want to support. So through might want to support. So through backends like this one for VLM and then backends like this one for VLM and then backends like this one for VLM and then also SG lang um you can take this code also SG lang um you can take this code also SG lang um you can take this code that you wrote once for your that you wrote once for your that you wrote once for your transformers training loop and run it transformers training loop and run it transformers training loop and run it straight in VLM or SG lang. straight in VLM or SG lang. straight in VLM or SG lang. uh you might be using it for RL in which uh you might be using it for RL in which uh you might be using it for RL in which case you would be using traditionally uh case you would be using traditionally uh case you would be using traditionally uh potentially transformers for your potentially transformers for your potentially transformers for your training and then for rollouts you would training and then for rollouts you would training and then for rollouts you would use VLM uh both with different modeling use VLM uh both with different modeling use VLM uh both with different modeling implementations because previously that implementations because previously that implementations because previously that was what you had to do. Uh but now with was what you had to do. Uh but now with was what you had to do. Uh but now with this you can use the same code for this you can use the same code for this you can use the same code for training and inference during your RL training and inference during your RL training and inference during your RL loops. uh meaning that there's no loops. uh meaning that there's no loops. uh meaning that there's no potential drift in the two potential drift in the two potential drift in the two implementations and the the training implementations and the the training implementations and the the training should go a little better.
-
should go a little better. should go a little better. And then finally, it might just be the And then finally, it might just be the And then finally, it might just be the only way that your model is supported in only way that your model is supported in only way that your model is supported in VLM. So today there are 14 architectures VLM. So today there are 14 architectures VLM. So today there are 14 architectures which only which are officially which only which are officially which only which are officially supported via the transformers modeling supported via the transformers modeling supported via the transformers modeling back end. Um and that list is only back end. Um and that list is only back end. Um and that list is only growing as time goes on. growing as time goes on. growing as time goes on. So now we start getting into the into So now we start getting into the into So now we start getting into the into the weeds and talking about how we get the weeds and talking about how we get the weeds and talking about how we get into the transformers modeling back end into the transformers modeling back end into the transformers modeling back end and what it does. Uh so obviously first and what it does. Uh so obviously first and what it does. Uh so obviously first we start with starting up a VLM engine. we start with starting up a VLM engine. we start with starting up a VLM engine. Um and here we have a whole load of uh Um and here we have a whole load of uh Um and here we have a whole load of uh command line arguments. One of them is command line arguments. One of them is command line arguments. One of them is to force the transformers modeling back to force the transformers modeling back to force the transformers modeling back end because for this uh quen 3VL model end because for this uh quen 3VL model end because for this uh quen 3VL model that we're looking at uh there is a that we're looking at uh there is a that we're looking at uh there is a dedicated BLM implementation. So we have dedicated BLM implementation. So we have dedicated BLM implementation. So we have to tell BLLM that transformers is what to tell BLLM that transformers is what to tell BLLM that transformers is what we want. Um and then we've used a a not we want. Um and then we've used a a not we want. Um and then we've used a a not recommended but possible uh parallelism recommended but possible uh parallelism recommended but possible uh parallelism scheme where we're using as many as we scheme where we're using as many as we scheme where we're using as many as we can across eight GPUs. We've got tensor can across eight GPUs. We've got tensor can across eight GPUs. We've got tensor pipeline data and expert parallel all at pipeline data and expert parallel all at pipeline data and expert parallel all at once. Um and then finally we've enabled once. Um and then finally we've enabled once. Um and then finally we've enabled torch compile on the multimodel encoder torch compile on the multimodel encoder torch compile on the multimodel encoder which not all um VLM models will do by which not all um VLM models will do by which not all um VLM models will do by default but because of the generality of default but because of the generality of default but because of the generality of the transformers modeling back end. If the transformers modeling back end. If the transformers modeling back end. If the transformers code uh has a the transformers code uh has a the transformers code uh has a compilable multimodel encoder you can compilable multimodel encoder you can compilable multimodel encoder you can just enable this flag and if your model just enable this flag and if your model just enable this flag and if your model works it should also work via VLM.
-
works it should also work via VLM. works it should also work via VLM. So we first cross the boundary into So we first cross the boundary into So we first cross the boundary into transformers from VLM by looking into transformers from VLM by looking into transformers from VLM by looking into the config where we're looking at what the config where we're looking at what the config where we're looking at what is this model, what shape is it and then is this model, what shape is it and then is this model, what shape is it and then any other features like what any other features like what any other features like what quantization scheme is it using quantization scheme is it using quantization scheme is it using and we then once we've chosen the and we then once we've chosen the and we then once we've chosen the transformers modeling back end had a transformers modeling back end had a transformers modeling back end had a look inside the config we need to choose look inside the config we need to choose look inside the config we need to choose which transformers modeling backend which transformers modeling backend which transformers modeling backend class to load our transformers model class to load our transformers model class to load our transformers model into. Uh and on the right you can see into. Uh and on the right you can see into. Uh and on the right you can see the selection of classes available. Um the selection of classes available. Um the selection of classes available. Um and the way that we select them is by and the way that we select them is by and the way that we select them is by looking into the config and establishing looking into the config and establishing looking into the config and establishing which characteristics the model has which characteristics the model has which characteristics the model has which will determine the correct class which will determine the correct class which will determine the correct class to use. So first they all start with to use. So first they all start with to use. So first they all start with transformers as you can see. Um and then transformers as you can see. Um and then transformers as you can see. Um and then we detect this one is multimodal because we detect this one is multimodal because we detect this one is multimodal because it has both a text and a vision config it has both a text and a vision config it has both a text and a vision config uh in that uh config JSON you saw on the uh in that uh config JSON you saw on the uh in that uh config JSON you saw on the previous slide. It's because it has more than zero experts because it has more than zero experts and then forcal lm is the default uh if and then forcal lm is the default uh if and then forcal lm is the default uh if your model is a causal model or if you your model is a causal model or if you your model is a causal model or if you don't specify that you want to do a don't specify that you want to do a don't specify that you want to do a pooling task with your causal model.
-
pooling task with your causal model. pooling task with your causal model. And so we choose the transformers And so we choose the transformers And so we choose the transformers multimodal for cos lm class. multimodal for cos lm class. multimodal for cos lm class. Uh and what are these classes made of? Uh and what are these classes made of? Uh and what are these classes made of? Because there were quite a lot of them Because there were quite a lot of them Because there were quite a lot of them and you'd think that if we didn't do and you'd think that if we didn't do and you'd think that if we didn't do anything clever, we would end up with anything clever, we would end up with anything clever, we would end up with lots of duplicated code um and lots of lots of duplicated code um and lots of lots of duplicated code um and lots of drift and it would be really hard to drift and it would be really hard to drift and it would be really hard to maintain. Um but all of these are made maintain. Um but all of these are made maintain. Um but all of these are made of mixins which enable each individual of mixins which enable each individual of mixins which enable each individual uh capability. uh capability. uh capability. So for this one we first start with the So for this one we first start with the So for this one we first start with the MOE mixin which provides uh all of the MOE mixin which provides uh all of the MOE mixin which provides uh all of the necessary protocols for expert parallel necessary protocols for expert parallel necessary protocols for expert parallel and expert parallel load balancing. Then and expert parallel load balancing. Then and expert parallel load balancing. Then we have the multimodal mixin which we have the multimodal mixin which we have the multimodal mixin which handles all of the vision encoding the handles all of the vision encoding the handles all of the vision encoding the encoding the vision encoding caching um encoding the vision encoding caching um encoding the vision encoding caching um mrop if your model has it which this one mrop if your model has it which this one mrop if your model has it which this one does. Um and then merging the vision does. Um and then merging the vision does. Um and then merging the vision embeddings into the text embeddings embeddings into the text embeddings embeddings into the text embeddings before passing it over to transformers. before passing it over to transformers. before passing it over to transformers. Then the causal mixin which is for the Then the causal mixin which is for the Then the causal mixin which is for the lm head and the logics computation lm head and the logics computation lm head and the logics computation and then finally the base class which and then finally the base class which and then finally the base class which does most of the heavy lifting and all does most of the heavy lifting and all does most of the heavy lifting and all of the mixins actually now uh inherit of the mixins actually now uh inherit of the mixins actually now uh inherit this and it ends up at the bottom of the this and it ends up at the bottom of the this and it ends up at the bottom of the inheritance stack here. Um and that inheritance stack here. Um and that inheritance stack here. Um and that handles the the weight loading, the handles the the weight loading, the handles the the weight loading, the weight mapping, um all of the weight mapping, um all of the weight mapping, um all of the modifications that enable this to run modifications that enable this to run modifications that enable this to run quickly um just as fast as a BLM model quickly um just as fast as a BLM model quickly um just as fast as a BLM model would. Um and then there is a whole host would. Um and then there is a whole host would. Um and then there is a whole host of supports this supports that uh of supports this supports that uh of supports this supports that uh protocols for things like quantization, protocols for things like quantization, protocols for things like quantization, Laura, pipeline parallel, eagle and Laura, pipeline parallel, eagle and Laura, pipeline parallel, eagle and eagle 3.
-
eagle 3. eagle 3. So now we're going to look at how the So now we're going to look at how the So now we're going to look at how the transformers class is instantiated and transformers class is instantiated and transformers class is instantiated and what we do to it uh during what we do to it uh during what we do to it uh during instantiation. So it all looks a bit instantiation. So it all looks a bit instantiation. So it all looks a bit faint now, but that's because it has not faint now, but that's because it has not faint now, but that's because it has not yet been instantiated. Um first we yet been instantiated. Um first we yet been instantiated. Um first we actually before doing any of that we actually before doing any of that we actually before doing any of that we register two attention shins with register two attention shins with register two attention shins with transformers attention registry. Um and transformers attention registry. Um and transformers attention registry. Um and that is the VLM attention and VLM MLA that is the VLM attention and VLM MLA that is the VLM attention and VLM MLA attention forward methods. Um then we attention forward methods. Um then we attention forward methods. Um then we patch the config so that when the patch the config so that when the patch the config so that when the forward pass of the transformers model forward pass of the transformers model forward pass of the transformers model happens uh one of those shims depending happens uh one of those shims depending happens uh one of those shims depending on which is correct for your model is on which is correct for your model is on which is correct for your model is what is actually selected rather than uh what is actually selected rather than uh what is actually selected rather than uh flash attention flex attention or or flash attention flex attention or or flash attention flex attention or or whichever other options the model uh whichever other options the model uh whichever other options the model uh claims that it can use. claims that it can use. claims that it can use. Then we decorate uh the language model Then we decorate uh the language model Then we decorate uh the language model always for torch compile from the LM's always for torch compile from the LM's always for torch compile from the LM's perspective. Um, and then because of perspective. Um, and then because of perspective. Um, and then because of that command line argument I mentioned that command line argument I mentioned that command line argument I mentioned earlier, we also decorate the vision earlier, we also decorate the vision earlier, we also decorate the vision tower. Um, which tells VLM that you can tower. Um, which tells VLM that you can tower. Um, which tells VLM that you can compile both of these as whole graphs compile both of these as whole graphs compile both of these as whole graphs independently of each other before uh before torch compile gets a before uh before torch compile gets a chance to kind of wing it and figure out chance to kind of wing it and figure out chance to kind of wing it and figure out on its own. And now we instantiate the on its own. And now we instantiate the on its own. And now we instantiate the model uh using the automodel from config model uh using the automodel from config model uh using the automodel from config method on the meta device. Um and for method on the meta device. Um and for method on the meta device. Um and for automodel is slightly different to um
-
automodel is slightly different to um automodel is slightly different to um auto mode for cos lm which might be a auto mode for cos lm which might be a auto mode for cos lm which might be a class you're more familiar with if you class you're more familiar with if you class you're more familiar with if you use the transformers ecosystem. Uh and use the transformers ecosystem. Uh and use the transformers ecosystem. Uh and the auto model just loads what the auto model just loads what the auto model just loads what transformers refers to as the base class transformers refers to as the base class transformers refers to as the base class which is just the language backbone and which is just the language backbone and which is just the language backbone and the vision tower but without the the vision tower but without the the vision tower but without the language modeling head. I might have to step back to the laptop. I might have to step back to the laptop. This appears to have stopped working. This appears to have stopped working. This appears to have stopped working. Um, Um, Um, then we create the HF to VLLM mapper then we create the HF to VLLM mapper then we create the HF to VLLM mapper which is how we specify how checkpoint which is how we specify how checkpoint which is how we specify how checkpoint weights might be translated into how weights might be translated into how weights might be translated into how they now look in the Python class that they now look in the Python class that they now look in the Python class that gets created. uh Transformers has a few gets created. uh Transformers has a few gets created. uh Transformers has a few of these uh built in if they have moved of these uh built in if they have moved of these uh built in if they have moved from an old style naive mixture of from an old style naive mixture of from an old style naive mixture of experts to a fused mixture of experts experts to a fused mixture of experts experts to a fused mixture of experts style. Um so at this stage this is a style. Um so at this stage this is a style. Um so at this stage this is a pretty minimal mapper but as we move pretty minimal mapper but as we move pretty minimal mapper but as we move through this process we'll be through this process we'll be through this process we'll be contributing to it as we modify the contributing to it as we modify the contributing to it as we modify the model.
-
model. model. So after that we then delete all of the So after that we then delete all of the So after that we then delete all of the layers that the current pipeline stage layers that the current pipeline stage layers that the current pipeline stage is not using. And in this example, we're is not using. And in this example, we're is not using. And in this example, we're on rank zero. So we delete the second on rank zero. So we delete the second on rank zero. So we delete the second half of the layer stack and the final half of the layer stack and the final half of the layer stack and the final layer norm. layer norm. layer norm. And now we begin recursive replace which And now we begin recursive replace which And now we begin recursive replace which is where the bulk of the work happens to is where the bulk of the work happens to is where the bulk of the work happens to bring the performance and access to all bring the performance and access to all bring the performance and access to all the features that BLM has. Um fore the features that BLM has. Um fore the features that BLM has. Um fore models it happens in two passes because models it happens in two passes because models it happens in two passes because thee is a little bit special and there thee is a little bit special and there thee is a little bit special and there are two different ways which we can do are two different ways which we can do are two different ways which we can do it. Uh for this model it uses the it. Uh for this model it uses the it. Uh for this model it uses the internal routting method which means internal routting method which means internal routting method which means that we by using torch FX uh detect the that we by using torch FX uh detect the that we by using torch FX uh detect the structure of the sparsee block. Um and structure of the sparsee block. Um and structure of the sparsee block. Um and if we determine that it is entirely if we determine that it is entirely if we determine that it is entirely representable by VLM's runner um then we representable by VLM's runner um then we representable by VLM's runner um then we replace it wholesale. So we replace the replace it wholesale. So we replace the replace it wholesale. So we replace the experts with this MOE runner. Um and experts with this MOE runner. Um and experts with this MOE runner. Um and that contains the rooted experts and that contains the rooted experts and that contains the rooted experts and then the routter gets replaced with a then the routter gets replaced with a then the routter gets replaced with a gate linear and that gets passed into gate linear and that gets passed into gate linear and that gets passed into the runner. Um and that handles the runner. Um and that handles the runner. Um and that handles everything. And so that lets us do uh everything. And so that lets us do uh everything. And so that lets us do uh things like use monolithic kernels for things like use monolithic kernels for things like use monolithic kernels for routing and experts. Uh if there were routing and experts. Uh if there were routing and experts. Uh if there were shared experts in this model, it would shared experts in this model, it would shared experts in this model, it would allow us to do asynchronous shared allow us to do asynchronous shared allow us to do asynchronous shared expert computation. Um and all of that expert computation. Um and all of that expert computation. Um and all of that is managed by thee runner.
-
And then once that is done, we go to the And then once that is done, we go to the second pass which is the base class second pass which is the base class second pass which is the base class doing a depth first search through all doing a depth first search through all doing a depth first search through all of the children of the base model. So of the children of the base model. So of the children of the base model. So first we walk through the vision model first we walk through the vision model first we walk through the vision model replacing all of the layers that we know replacing all of the layers that we know replacing all of the layers that we know how. Um to its features. So all of VLM's to its features. So all of VLM's parallelization is built into its parallelization is built into its parallelization is built into its layers. So by using VLM's linear layers layers. So by using VLM's linear layers layers. So by using VLM's linear layers and VLM's experts layers, that's how we and VLM's experts layers, that's how we and VLM's experts layers, that's how we gain access to tensor parallel expert gain access to tensor parallel expert gain access to tensor parallel expert parallel etc. parallel etc. parallel etc. Then we do the vocab parallel embedding Then we do the vocab parallel embedding Then we do the vocab parallel embedding swap um by surgically replacing just the swap um by surgically replacing just the swap um by surgically replacing just the torch nn embeddings module. So if you torch nn embeddings module. So if you torch nn embeddings module. So if you have a model like Gemma where you do have a model like Gemma where you do have a model like Gemma where you do some scaling during your token embedding some scaling during your token embedding some scaling during your token embedding um that is all still preserved and we um that is all still preserved and we um that is all still preserved and we don't have to special case it on the VLM don't have to special case it on the VLM don't have to special case it on the VLM side.
-
side. side. Now the next fuser uh is used which is Now the next fuser uh is used which is Now the next fuser uh is used which is the QKV fuser which again uses um FX to the QKV fuser which again uses um FX to the QKV fuser which again uses um FX to detect the QKV and output projections detect the QKV and output projections detect the QKV and output projections um and swap in BLM's QKV parallel linear um and swap in BLM's QKV parallel linear um and swap in BLM's QKV parallel linear and a corresponding row parallel linear and a corresponding row parallel linear and a corresponding row parallel linear um using a so so this is all it's using um using a so so this is all it's using um using a so so this is all it's using torch FX to detect what we need to swap torch FX to detect what we need to swap torch FX to detect what we need to swap but then it's using Python's as module but then it's using Python's as module but then it's using Python's as module to actually make the modifications. So to actually make the modifications. So to actually make the modifications. So we are yet to torch compile anything and we are yet to torch compile anything and we are yet to torch compile anything and all of that is still left up to later all of that is still left up to later all of that is still left up to later stages in BLM. stages in BLM. stages in BLM. Then similarly we replace the RMS norm Then similarly we replace the RMS norm Then similarly we replace the RMS norm which is deceptively difficult for how which is deceptively difficult for how which is deceptively difficult for how simple uh it is but we do it because VLM simple uh it is but we do it because VLM simple uh it is but we do it because VLM uses um some global torch compile passes uses um some global torch compile passes uses um some global torch compile passes to detect where it can fuse layers and to detect where it can fuse layers and to detect where it can fuse layers and it does it by detecting its own modules.
-
it does it by detecting its own modules. it does it by detecting its own modules. Um and so there are some um RMS norm Um and so there are some um RMS norm Um and so there are some um RMS norm passes or some compilation passes that passes or some compilation passes that passes or some compilation passes that the RMS norm is involved in that will the RMS norm is involved in that will the RMS norm is involved in that will only be detected if we're using VLM's only be detected if we're using VLM's only be detected if we're using VLM's class. class. class. Then we create VLM's attention instances Then we create VLM's attention instances Then we create VLM's attention instances and as before this gives us access to and as before this gives us access to and as before this gives us access to the rest of VLM. Uh and the key thing the rest of VLM. Uh and the key thing the rest of VLM. Uh and the key thing that this gives us access to is VLM's KB that this gives us access to is VLM's KB that this gives us access to is VLM's KB cache. Um so we don't need to use any KB cache. Um so we don't need to use any KB cache. Um so we don't need to use any KB caching that has been implemented on caching that has been implemented on caching that has been implemented on transformers. Um, and VLM gets to do all transformers. Um, and VLM gets to do all transformers. Um, and VLM gets to do all of the clever stuff that it can do with of the clever stuff that it can do with of the clever stuff that it can do with the KB cache like offloading, um, the KB cache like offloading, um, the KB cache like offloading, um, disagregated prefill between nodes, um, disagregated prefill between nodes, um, disagregated prefill between nodes, um, KB cache reuse or all of that fun stuff. KB cache reuse or all of that fun stuff. KB cache reuse or all of that fun stuff. Then anything that we haven't touched Then anything that we haven't touched Then anything that we haven't touched yet, we make sure that it is on the yet, we make sure that it is on the yet, we make sure that it is on the device. Uh, so the only thing that device. Uh, so the only thing that device. Uh, so the only thing that changed there was those two norms in the changed there was those two norms in the changed there was those two norms in the vision model. And then lastly, we just vision model. And then lastly, we just vision model. And then lastly, we just make sure that anything that make sure that anything that make sure that anything that Transformers insists should be in FP32 Transformers insists should be in FP32 Transformers insists should be in FP32 stays in FP32 even if the rest of the stays in FP32 even if the rest of the stays in FP32 even if the rest of the model is FP8.
-
model is FP8. model is FP8. So we've built the model um and now a So we've built the model um and now a So we've built the model um and now a request has been received. We have had request has been received. We have had request has been received. We have had it scheduled through VLN and we've it scheduled through VLN and we've it scheduled through VLN and we've reached the base classes forward method. reached the base classes forward method. reached the base classes forward method. Um and this is the first hop back into Um and this is the first hop back into Um and this is the first hop back into transformers. transformers. transformers. Uh so we simply call the forward method Uh so we simply call the forward method Uh so we simply call the forward method on the original Quen 3VLOE model class on the original Quen 3VLOE model class on the original Quen 3VLOE model class um and pass in the input ids or the um and pass in the input ids or the um and pass in the input ids or the input embeds um in this case it's the input embeds um in this case it's the input embeds um in this case it's the inputs embeds because it's a multimodal inputs embeds because it's a multimodal inputs embeds because it's a multimodal model and VLM pre-mbbeds the vision part model and VLM pre-mbbeds the vision part model and VLM pre-mbbeds the vision part uh of the prompt and uh of the prompt and uh of the prompt and inserts it into the text embeddings and inserts it into the text embeddings and inserts it into the text embeddings and passes that all in one go rather than passes that all in one go rather than passes that all in one go rather than relying on the model itself to to do relying on the model itself to to do relying on the model itself to to do that we explicitly disable transformers that we explicitly disable transformers that we explicitly disable transformers is um KV cache because VM owns it is um KV cache because VM owns it is um KV cache because VM owns it through those attention modules. Um and through those attention modules. Um and through those attention modules. Um and then we just forward through the then we just forward through the then we just forward through the position ids with a slight modification position ids with a slight modification position ids with a slight modification for MRO models um which happens in the for MRO models um which happens in the for MRO models um which happens in the multimodal mixin for this particular multimodal mixin for this particular multimodal mixin for this particular type of model.
-
And so while we're going through the And so while we're going through the forward pass of the transformers model, forward pass of the transformers model, forward pass of the transformers model, there are two seams where we return back there are two seams where we return back there are two seams where we return back to VLM um in kind of performance to VLM um in kind of performance to VLM um in kind of performance critical ways. Uh and so the first and critical ways. Uh and so the first and critical ways. Uh and so the first and most interesting one is the attention most interesting one is the attention most interesting one is the attention module. Um so on the transformers side, module. Um so on the transformers side, module. Um so on the transformers side, we're starting again in the the we're starting again in the the we're starting again in the the transformers native class. It goes transformers native class. It goes transformers native class. It goes through its registry to choose which through its registry to choose which through its registry to choose which attention method to use which we attention method to use which we attention method to use which we pre-registered earlier as the VLM pre-registered earlier as the VLM pre-registered earlier as the VLM um attention shim and then selected it um attention shim and then selected it um attention shim and then selected it by modifying the config by modifying the config by modifying the config and then we enter that and transition and then we enter that and transition and then we enter that and transition back to VLM where VLM does some minor back to VLM where VLM does some minor back to VLM where VLM does some minor shape manipulation shape manipulation shape manipulation um gets the attention kernel to use um gets the attention kernel to use um gets the attention kernel to use which is any attention kernel supported which is any attention kernel supported which is any attention kernel supported by your device for this particular model by your device for this particular model by your device for this particular model in VLM it gets the KV cache using uh the in VLM it gets the KV cache using uh the in VLM it gets the KV cache using uh the attention context and then it actually attention context and then it actually attention context and then it actually calls the attention and passes the calls the attention and passes the calls the attention and passes the result back to transformers to continue result back to transformers to continue result back to transformers to continue with the rest of the model and the with the rest of the model and the with the rest of the model and the second is thee through internal second is thee through internal second is thee through internal routting. Um so as I mentioned thee has routting. Um so as I mentioned thee has routting. Um so as I mentioned thee has a special case where if we're able to a special case where if we're able to a special case where if we're able to detect an entirely representable sparsee detect an entirely representable sparsee detect an entirely representable sparsee block then we can replace it entirely block then we can replace it entirely block then we can replace it entirely and actually the forward method that we and actually the forward method that we and actually the forward method that we use is just experts forward and use is just experts forward and use is just experts forward and everything happens inside that runner everything happens inside that runner everything happens inside that runner class. So we go straight into the VLM
-
class. So we go straight into the VLM class. So we go straight into the VLM side. We gate the hidden states. We get side. We gate the hidden states. We get side. We gate the hidden states. We get the top K ids and logits. Perform the the top K ids and logits. Perform the the top K ids and logits. Perform the rooted experts for computation and pass rooted experts for computation and pass rooted experts for computation and pass the results back to transformers. And after this point the transformers And after this point the transformers model is done. It has provided all of model is done. It has provided all of model is done. It has provided all of the information about what the graph of the information about what the graph of the information about what the graph of the model is and all the computation the model is and all the computation the model is and all the computation for well and all the journey through the for well and all the journey through the for well and all the journey through the computation rather because as we saw the computation rather because as we saw the computation rather because as we saw the modules are all VLM modules um and then modules are all VLM modules um and then modules are all VLM modules um and then after that the logic computation and LM after that the logic computation and LM after that the logic computation and LM head evaluation is all done by VLM um head evaluation is all done by VLM um head evaluation is all done by VLM um and the sampling is also handled by VLM. and the sampling is also handled by VLM. and the sampling is also handled by VLM. So there are a few features that we So there are a few features that we So there are a few features that we haven't used with this model that I want haven't used with this model that I want haven't used with this model that I want to talk about. The first is external to talk about. The first is external to talk about. The first is external routting. Uh so if your block isn't the routting. Uh so if your block isn't the routting. Uh so if your block isn't the same as something that we've same as something that we've same as something that we've anticipated, uh that doesn't mean that anticipated, uh that doesn't mean that anticipated, uh that doesn't mean that it won't work. It just means that we it won't work. It just means that we it won't work. It just means that we will only replace the experts will only replace the experts will only replace the experts themselves. Um so that gives you the themselves. Um so that gives you the themselves. Um so that gives you the crucial performance because that's where crucial performance because that's where crucial performance because that's where most of the work is done. Um, but it most of the work is done. Um, but it most of the work is done. Um, but it gives you the flexibility to do whatever gives you the flexibility to do whatever gives you the flexibility to do whatever you might come up with in the routing or you might come up with in the routing or you might come up with in the routing or anything around it.
-
anything around it. anything around it. And there's also shared experts, which I And there's also shared experts, which I And there's also shared experts, which I alluded to earlier. If they had existed alluded to earlier. If they had existed alluded to earlier. If they had existed and they were detected in the internal and they were detected in the internal and they were detected in the internal routing, they would also have been routing, they would also have been routing, they would also have been included in the MOE runner. And included in the MOE runner. And included in the MOE runner. And depending on your configuration of BLM, depending on your configuration of BLM, depending on your configuration of BLM, um, it would be computed in a way that um, it would be computed in a way that um, it would be computed in a way that is more efficient than could be done is more efficient than could be done is more efficient than could be done with external routting. with external routting. with external routting. Then there's Laura and Eagle and Eagle Then there's Laura and Eagle and Eagle Then there's Laura and Eagle and Eagle 3, which we just didn't touch on because 3, which we just didn't touch on because 3, which we just didn't touch on because these are kind of determined by the these are kind of determined by the these are kind of determined by the checkpoint you use and then the requests checkpoint you use and then the requests checkpoint you use and then the requests that you pass. that you pass. that you pass. Then there's a a whole load of different Then there's a a whole load of different Then there's a a whole load of different attention flavors. So there's encoder attention flavors. So there's encoder attention flavors. So there's encoder only attention which would be used by only attention which would be used by only attention which would be used by models like BERT or sliding window models like BERT or sliding window models like BERT or sliding window attention that you might find in Gemma attention that you might find in Gemma attention that you might find in Gemma models. Um and for encoder we support models. Um and for encoder we support models. Um and for encoder we support encoder only models um not no encoder encoder only models um not no encoder encoder only models um not no encoder decoder models but for sliding window decoder models but for sliding window decoder models but for sliding window attention it can be interled um similar attention it can be interled um similar attention it can be interled um similar to how they started doing it with Gemma to how they started doing it with Gemma to how they started doing it with Gemma where you would alternate between full where you would alternate between full where you would alternate between full and sliding attention. Um and then also and sliding attention. Um and then also and sliding attention. Um and then also MLA attention which comes through a MLA attention which comes through a MLA attention which comes through a separate separate separate um fuser um fuser um fuser and that actually gives us the KB cache and that actually gives us the KB cache and that actually gives us the KB cache compression benefits that we didn't used compression benefits that we didn't used compression benefits that we didn't used to get when we used the we just used to to get when we used the we just used to to get when we used the we just used to pad to full attention when you loaded a pad to full attention when you loaded a pad to full attention when you loaded a a MLA attention model.
-
a MLA attention model. a MLA attention model. Um and then finally pooling and Um and then finally pooling and Um and then finally pooling and classification tasks um which would have classification tasks um which would have classification tasks um which would have been automatically used had we used a uh been automatically used had we used a uh been automatically used had we used a uh a pooling model. Um or we could have a pooling model. Um or we could have a pooling model. Um or we could have manually selected it and BLM would have manually selected it and BLM would have manually selected it and BLM would have created a classification head using the created a classification head using the created a classification head using the uh the LM head or it could have given uh the LM head or it could have given uh the LM head or it could have given you back the the raw embeddings out the you back the the raw embeddings out the you back the the raw embeddings out the end of the model if that is what you end of the model if that is what you end of the model if that is what you were after. were after. were after. And so to recap, we looked into the And so to recap, we looked into the And so to recap, we looked into the config, determined which hugging face config, determined which hugging face config, determined which hugging face architecture it was and then which architecture it was and then which architecture it was and then which transformers modeling backend class we transformers modeling backend class we transformers modeling backend class we should use. should use. should use. Um we loaded it into this generic class Um we loaded it into this generic class Um we loaded it into this generic class in BLM in BLM in BLM and we manipulated it such that all of and we manipulated it such that all of and we manipulated it such that all of the advanced features of VLM became the advanced features of VLM became the advanced features of VLM became accessible. Um and all of the accessible. Um and all of the accessible. Um and all of the performance that comes with them came performance that comes with them came performance that comes with them came with them. Um then we went through the with them. Um then we went through the with them. Um then we went through the forward pass of the base model with two forward pass of the base model with two forward pass of the base model with two important seams um where we go back into important seams um where we go back into important seams um where we go back into VLM and do some non-trivial VLM and do some non-trivial VLM and do some non-trivial um work and then we pass back to VLM to um work and then we pass back to VLM to um work and then we pass back to VLM to do the logic computation and the do the logic computation and the do the logic computation and the sampling sampling sampling and that was my talk and the slides are and that was my talk and the slides are and that was my talk and the slides are available on the QR code there um if available on the QR code there um if available on the QR code there um if anyone wants to look at them again anyone wants to look at them again anyone wants to look at them again Thank you.
-
That's what happens. All right, I think That's what happens. All right, I think we're on a break now. Thank you so much we're on a break now. Thank you so much we're on a break now. Thank you so much everybody. So, I think we were going to everybody. So, I think we were going to everybody. So, I think we were going to take a little break and we're back in take a little break and we're back in take a little break and we're back in half an hour, I believe. All right, half an hour, I believe. All right, half an hour, I believe. All right, thank you so much.
-
Hello, welcome back. All right, let's Hello, welcome back. All right, let's get it for you guys. Okay. So get it for you guys. Okay. So get it for you guys. Okay. So um so far we have spoken about you know um so far we have spoken about you know um so far we have spoken about you know the AI economy like we talked about the AI economy like we talked about the AI economy like we talked about harnesses, software factories, we talked harnesses, software factories, we talked harnesses, software factories, we talked about generative media and all those about generative media and all those about generative media and all those tools are for you guys to build agents tools are for you guys to build agents tools are for you guys to build agents to build applications that use AI. Who to build applications that use AI. Who to build applications that use AI. Who who here built actually with AI? who here built actually with AI? who here built actually with AI? Okay. Yeah, most of you guys do, right? Okay. Yeah, most of you guys do, right? Okay. Yeah, most of you guys do, right? And the purpose of that is to create And the purpose of that is to create And the purpose of that is to create value to for people, right? Whatever value to for people, right? Whatever value to for people, right? Whatever agent you're building, whatever tool agent you're building, whatever tool agent you're building, whatever tool you're making, ultimately is to bring you're making, ultimately is to bring you're making, ultimately is to bring value to your users. Our next speaker is value to your users. Our next speaker is value to your users. Our next speaker is actually deep into uh the the the data actually deep into uh the the the data actually deep into uh the the the data and um she works at Stripe and what they and um she works at Stripe and what they and um she works at Stripe and what they see is that the entire economy is see is that the entire economy is see is that the entire economy is changing thanks to AI and uh she's going changing thanks to AI and uh she's going changing thanks to AI and uh she's going to come to to tell us about and give us to come to to tell us about and give us to come to to tell us about and give us some insights about what Stripe uh what some insights about what Stripe uh what some insights about what Stripe uh what stripes data is revealing about AI. So stripes data is revealing about AI. So stripes data is revealing about AI. So without further ado, please join me in without further ado, please join me in without further ado, please join me in welcoming to the stage the head of welcoming to the stage the head of welcoming to the stage the head of product of AI and startups, Ariel Lubai.
-
Hi everyone. Hi everyone. Thank you. Thank you so much for being Thank you. Thank you so much for being Thank you. Thank you so much for being here. Um, I'm super happy to be here and here. Um, I'm super happy to be here and here. Um, I'm super happy to be here and I hope I will get some slides at some I hope I will get some slides at some I hope I will get some slides at some point on the screen. point on the screen. point on the screen. Uh, Uh, Uh, can I get the slides in the please? Okay, perfect. Okay, let's come back Okay, perfect. Okay, let's come back there. Okay, perfect. So yeah uh really there. Okay, perfect. So yeah uh really there. Okay, perfect. So yeah uh really happy to be here. Uh so at Stripe we happy to be here. Uh so at Stripe we happy to be here. Uh so at Stripe we work with the fastest AI growing work with the fastest AI growing work with the fastest AI growing companies. We we have this luck and many companies. We we have this luck and many companies. We we have this luck and many of them are actually in this room. So of them are actually in this room. So of them are actually in this room. So really happy to be here and we have a really happy to be here and we have a really happy to be here and we have a front seat of what's working, what's not front seat of what's working, what's not front seat of what's working, what's not and like how the best companies scale and like how the best companies scale and like how the best companies scale today in this new world of AI. So today today in this new world of AI. So today today in this new world of AI. So today I want to share some of those insights I want to share some of those insights I want to share some of those insights with you guys and you know have some with you guys and you know have some with you guys and you know have some thoughts on the playbook because this thoughts on the playbook because this thoughts on the playbook because this has been completely rewritten in the has been completely rewritten in the has been completely rewritten in the past few years and even months. So first past few years and even months. So first past few years and even months. So first like let's start with uh figures sorry like let's start with uh figures sorry like let's start with uh figures sorry um first we are working with um 1.6% um first we are working with um 1.6% um first we are working with um 1.6% six% uh of uh of the GDP globally that six% uh of uh of the GDP globally that six% uh of uh of the GDP globally that is processed by Stripe and that is is processed by Stripe and that is is processed by Stripe and that is really something that puts us in that really something that puts us in that really something that puts us in that front row seat at the five patterns that front row seat at the five patterns that front row seat at the five patterns that we're seeing. When the top AI companies we're seeing. When the top AI companies we're seeing. When the top AI companies build faster, I think this is something build faster, I think this is something build faster, I think this is something we're all feeling. Two, they sell
-
we're all feeling. Two, they sell we're all feeling. Two, they sell globally by default. When you're globally by default. When you're globally by default. When you're building today, your day one market is building today, your day one market is building today, your day one market is now the whole world. Um, three, pricing now the whole world. Um, three, pricing now the whole world. Um, three, pricing is evolving is evolving is evolving completely. Like there are a lot of completely. Like there are a lot of completely. Like there are a lot of workshops on pricing today by the way workshops on pricing today by the way workshops on pricing today by the way because how you price has completely because how you price has completely because how you price has completely changed in the past few years. Um, and changed in the past few years. Um, and changed in the past few years. Um, and four they adapt their go to market four they adapt their go to market four they adapt their go to market motions as well. We have like channel motions as well. We have like channel motions as well. We have like channel sales that are becoming increasingly sales that are becoming increasingly sales that are becoming increasingly important. Uh, we'll dig into that important. Uh, we'll dig into that important. Uh, we'll dig into that afterwards. And five, they monitor fraud afterwards. And five, they monitor fraud afterwards. And five, they monitor fraud early and often. Why? is like what's early and often. Why? is like what's early and often. Why? is like what's creating potentially viral growth with creating potentially viral growth with creating potentially viral growth with agents is also creating just a whole new agents is also creating just a whole new agents is also creating just a whole new ground for fraud attacks for your ground for fraud attacks for your ground for fraud attacks for your business. So we'll dig into it of course business. So we'll dig into it of course business. So we'll dig into it of course but first let's look at like big picture but first let's look at like big picture but first let's look at like big picture what's happening um in the world on what's happening um in the world on what's happening um in the world on that. that. that. Um we saw the top AI companies grow by Um we saw the top AI companies grow by Um we saw the top AI companies grow by 145% in 2025 which was already 145% in 2025 which was already 145% in 2025 which was already impressive but like in 26 this number impressive but like in 26 this number impressive but like in 26 this number grew by 195%.
-
grew by 195%. grew by 195%. So tripling in a single year. Uh that So tripling in a single year. Uh that So tripling in a single year. Uh that growth isn't slowing down it's actually growth isn't slowing down it's actually growth isn't slowing down it's actually accelerating today. And the top ones accelerating today. And the top ones accelerating today. And the top ones that you know are even more mind-blowing that you know are even more mind-blowing that you know are even more mind-blowing of course like lovable they reported 100 of course like lovable they reported 100 of course like lovable they reported 100 million in eight months. uh eight months million in eight months. uh eight months million in eight months. uh eight months later they were actually at 400 million. later they were actually at 400 million. later they were actually at 400 million. Anthropic that you all know started in Anthropic that you all know started in Anthropic that you all know started in January 23 with zero and went to um to 1 January 23 with zero and went to um to 1 January 23 with zero and went to um to 1 billion in in just two years and now billion in in just two years and now billion in in just two years and now they're at 30 billion. So that's just they're at 30 billion. So that's just they're at 30 billion. So that's just completely mind-blowing. completely mind-blowing. completely mind-blowing. And this isn't just B2B like we saw this And this isn't just B2B like we saw this And this isn't just B2B like we saw this in in our link data like consumer in in our link data like consumer in in our link data like consumer adoption is also increasing a lot. like adoption is also increasing a lot. like adoption is also increasing a lot. like it doubled just from 6 million to over it doubled just from 6 million to over it doubled just from 6 million to over 14 million just in one year. 14 million just in one year. 14 million just in one year. Consumers are also spending more. What's Consumers are also spending more. What's Consumers are also spending more. What's actually interesting is that they're actually interesting is that they're actually interesting is that they're spending more on AI tools. Like today, spending more on AI tools. Like today, spending more on AI tools. Like today, we're around like $360 we're around like $360 we're around like $360 uh per person per month on AI, which is uh per person per month on AI, which is uh per person per month on AI, which is basically double what it was three basically double what it was three basically double what it was three months ago. It was around $180.
-
months ago. It was around $180. months ago. It was around $180. And this is basically And this is basically And this is basically like basically what the average European like basically what the average European like basically what the average European spends for like phone service um uh spends for like phone service um uh spends for like phone service um uh internet and streaming services internet and streaming services internet and streaming services combined. Like people aren't treating AI combined. Like people aren't treating AI combined. Like people aren't treating AI today like some streaming service that today like some streaming service that today like some streaming service that you can cancel uh but they're treating you can cancel uh but they're treating you can cancel uh but they're treating that like a critical utility that they that like a critical utility that they that like a critical utility that they absolutely need. absolutely need. absolutely need. So what's behind all this? We're going So what's behind all this? We're going So what's behind all this? We're going to deep dive into these different to deep dive into these different to deep dive into these different patterns starting with speed. patterns starting with speed. patterns starting with speed. So just like you no longer need months So just like you no longer need months So just like you no longer need months of engineering to uh you know v code and of engineering to uh you know v code and of engineering to uh you know v code and accept the payments now you no longer uh accept the payments now you no longer uh accept the payments now you no longer uh need month of engineering to ship need month of engineering to ship need month of engineering to ship software. So what's actually interesting software. So what's actually interesting software. So what's actually interesting is that this chart shows when it is that this chart shows when it is that this chart shows when it clicked. Like if you take a look at iOS clicked. Like if you take a look at iOS clicked. Like if you take a look at iOS app releases, they were actually app releases, they were actually app releases, they were actually declining at the end of uh 24 and then declining at the end of uh 24 and then declining at the end of uh 24 and then the agent codic tools hit. And so this the agent codic tools hit. And so this the agent codic tools hit. And so this is where we saw the iOS app launches is where we saw the iOS app launches is where we saw the iOS app launches just grow again and not like small grow just grow again and not like small grow just grow again and not like small grow by 24% month over month.
-
by 24% month over month. by 24% month over month. And so now even like now that we embed And so now even like now that we embed And so now even like now that we embed payments into developer platforms like payments into developer platforms like payments into developer platforms like Lovable, Riplet, Versal, etc. Um it's Lovable, Riplet, Versal, etc. Um it's Lovable, Riplet, Versal, etc. Um it's actually even faster to go from your actually even faster to go from your actually even faster to go from your idea to first charge like even faster. idea to first charge like even faster. idea to first charge like even faster. So what you can see here is that it's So what you can see here is that it's So what you can see here is that it's actually taking less than six weeks from actually taking less than six weeks from actually taking less than six weeks from first idea and like sandbox to your first idea and like sandbox to your first idea and like sandbox to your first charge. Like only six weeks. first charge. Like only six weeks. first charge. Like only six weeks. That's really crazy. That's really crazy. That's really crazy. And that strategy is faster for And that strategy is faster for And that strategy is faster for technical folks. So I I've played a bit technical folks. So I I've played a bit technical folks. So I I've played a bit with a prompt to show you how it looks with a prompt to show you how it looks with a prompt to show you how it looks like today to build with cloud and like today to build with cloud and like today to build with cloud and stripe. Uh because we've we've made a stripe. Uh because we've we've made a stripe. Uh because we've we've made a lot of work to make sure that we have an lot of work to make sure that we have an lot of work to make sure that we have an agentic developer experience that agentic developer experience that agentic developer experience that enables you to build without having uh enables you to build without having uh enables you to build without having uh to leave where you're building and to leave where you're building and to leave where you're building and having to go on Stripe dashboard and and having to go on Stripe dashboard and and having to go on Stripe dashboard and and build stuff. So um let's say you are a build stuff. So um let's say you are a build stuff. So um let's say you are a developer today and that you've built an developer today and that you've built an developer today and that you've built an app and now you're ready to set to set app and now you're ready to set to set app and now you're ready to set to set up payments.
-
up payments. up payments. So let's jump into cloud and see what it So let's jump into cloud and see what it So let's jump into cloud and see what it looks like. So I have my prompt here looks like. So I have my prompt here looks like. So I have my prompt here where I've asked you know to set up where I've asked you know to set up where I've asked you know to set up payments. I want um invoicing and payments. I want um invoicing and payments. I want um invoicing and subscription subscription subscription and so and so and so yeah. Okay. So here you can see that yeah. Okay. So here you can see that yeah. Okay. So here you can see that Claude correctly recommends Stripe which Claude correctly recommends Stripe which Claude correctly recommends Stripe which is of course what we want and that they is of course what we want and that they is of course what we want and that they will be more powerful because they are will be more powerful because they are will be more powerful because they are connecting to Stripe MCP and the skills connecting to Stripe MCP and the skills connecting to Stripe MCP and the skills that we created with this. So Claude that we created with this. So Claude that we created with this. So Claude uses the Stripe integration recommener uses the Stripe integration recommener uses the Stripe integration recommener um MCP tool along with those skills uh um MCP tool along with those skills uh um MCP tool along with those skills uh to know what to use and what to ask to to know what to use and what to ask to to know what to use and what to ask to the user. So it's basically asking like the user. So it's basically asking like the user. So it's basically asking like where are your customers located? The where are your customers located? The where are your customers located? The reason it's asking is because like cloud reason it's asking is because like cloud reason it's asking is because like cloud knows now that they need to have this knows now that they need to have this knows now that they need to have this answer to provide the best answer to provide the best answer to provide the best recommendation for the integration. recommendation for the integration. recommendation for the integration. Um what's interesting is that they Um what's interesting is that they Um what's interesting is that they recommend checkout which is a stripe recommend checkout which is a stripe recommend checkout which is a stripe integration and not like cloud element integration and not like cloud element integration and not like cloud element for example because it knows that um for for example because it knows that um for for example because it knows that um for for a better subscription integration we for a better subscription integration we for a better subscription integration we need checkout. So that's again like need checkout. So that's again like need checkout. So that's again like showing that uh it knows the right showing that uh it knows the right showing that uh it knows the right thing.
-
thing. thing. So then I responded like my customers So then I responded like my customers So then I responded like my customers are split between the US and the EU uh are split between the US and the EU uh are split between the US and the EU uh with a particular concentration in with a particular concentration in with a particular concentration in Ireland. So what's interesting now is Ireland. So what's interesting now is Ireland. So what's interesting now is that I get feedback on the type of tax that I get feedback on the type of tax that I get feedback on the type of tax knowledge I need to have. So basically knowledge I need to have. So basically knowledge I need to have. So basically uh I need to know that digital uh I need to know that digital uh I need to know that digital subscriptions um are taxable here. And subscriptions um are taxable here. And subscriptions um are taxable here. And it's actually interesting because it it it's actually interesting because it it it's actually interesting because it it recommends me to use stripe tax um recommends me to use stripe tax um recommends me to use stripe tax um because I don't want to have to know all because I don't want to have to know all because I don't want to have to know all the regional tax rules for all the the regional tax rules for all the the regional tax rules for all the countries that I'm going to launch to. countries that I'm going to launch to. countries that I'm going to launch to. And so Stripe Tax is going to help me And so Stripe Tax is going to help me And so Stripe Tax is going to help me without uh without having to know those without uh without having to know those without uh without having to know those rules. rules. rules. Um I can see that it's provisioning a Um I can see that it's provisioning a Um I can see that it's provisioning a Stripe sandbox for me. Um, what's Stripe sandbox for me. Um, what's Stripe sandbox for me. Um, what's actually interesting is that you don't actually interesting is that you don't actually interesting is that you don't have to, you know, go outside of the have to, you know, go outside of the have to, you know, go outside of the flow here to to get a sandbox. Um, the flow here to to get a sandbox. Um, the flow here to to get a sandbox. Um, the agent uses the new CLI command to create agent uses the new CLI command to create agent uses the new CLI command to create the sandbox without needing any stripe the sandbox without needing any stripe the sandbox without needing any stripe credentials, which was not the case credentials, which was not the case credentials, which was not the case before. Ed the command also fetches the before. Ed the command also fetches the before. Ed the command also fetches the API keys and securely stores them. And API keys and securely stores them. And API keys and securely stores them. And that way I basically don't have to break that way I basically don't have to break that way I basically don't have to break my flow and I can build my integration my flow and I can build my integration my flow and I can build my integration like without having to to go away. And like without having to to go away. And like without having to to go away. And so with the key sorted I can see that so with the key sorted I can see that so with the key sorted I can see that now Claude is checking all the different now Claude is checking all the different now Claude is checking all the different skills um that we stripe provide to make skills um that we stripe provide to make skills um that we stripe provide to make sure that they're um it's building the sure that they're um it's building the sure that they're um it's building the right thing.
-
right thing. right thing. So, it's saying that it has everything So, it's saying that it has everything So, it's saying that it has everything it needs and it's asking me like just do it needs and it's asking me like just do it needs and it's asking me like just do you want to proceed with MCP Stripe you want to proceed with MCP Stripe you want to proceed with MCP Stripe create product and and if I want to create product and and if I want to create product and and if I want to continue with that and create my continue with that and create my continue with that and create my checkout. checkout. checkout. So, I'm pressing yes here and we're So, I'm pressing yes here and we're So, I'm pressing yes here and we're going to fast forward and try to see the going to fast forward and try to see the going to fast forward and try to see the checkout. checkout. checkout. Um, so this is what it looks like. Um Um, so this is what it looks like. Um Um, so this is what it looks like. Um I'm going to test if the checkout works I'm going to test if the checkout works I'm going to test if the checkout works properly properly properly and so and so and so yeah that's how uh it would look like. yeah that's how uh it would look like. yeah that's how uh it would look like. Okay so it's done like strap and claude Okay so it's done like strap and claude Okay so it's done like strap and claude are basically making it easy for you to are basically making it easy for you to are basically making it easy for you to for you and like all the developers to for you and like all the developers to for you and like all the developers to build set up payments and start build set up payments and start build set up payments and start monetizing without going away uh from monetizing without going away uh from monetizing without going away uh from your uh building interface. your uh building interface. your uh building interface. So what does this mean for you? Like So what does this mean for you? Like So what does this mean for you? Like many of you are probably already doing many of you are probably already doing many of you are probably already doing this but I want to reemphasize that this but I want to reemphasize that this but I want to reemphasize that you know in the playbook today it's you know in the playbook today it's you know in the playbook today it's really important to obsess over really important to obsess over really important to obsess over developer productivity like the best developer productivity like the best developer productivity like the best companies today are really working on companies today are really working on companies today are really working on speed and just making sure that they speed and just making sure that they speed and just making sure that they completely collapse build times.
-
completely collapse build times. completely collapse build times. Second like build, sell and iterate at Second like build, sell and iterate at Second like build, sell and iterate at the same time. the linear playbook that the same time. the linear playbook that the same time. the linear playbook that we had before like you build your we had before like you build your we had before like you build your product then you go find your customers product then you go find your customers product then you go find your customers isn't just working anymore because of of isn't just working anymore because of of isn't just working anymore because of of the fact that too many people are the fact that too many people are the fact that too many people are building at the same time. So you building at the same time. So you building at the same time. So you basically need to do all three at the basically need to do all three at the basically need to do all three at the same time um in parallel. same time um in parallel. same time um in parallel. And then on the build build versus buy And then on the build build versus buy And then on the build build versus buy question, we're actually quite convinced question, we're actually quite convinced question, we're actually quite convinced from what we're seeing that the best from what we're seeing that the best from what we're seeing that the best companies do both like they built what's companies do both like they built what's companies do both like they built what's making them differentiated and they making them differentiated and they making them differentiated and they actually also buy you know the actually also buy you know the actually also buy you know the infrastructure and building blocks that infrastructure and building blocks that infrastructure and building blocks that they need for their products. they need for their products. they need for their products. So two like getting to market is one So two like getting to market is one So two like getting to market is one thing but it is also a big question of thing but it is also a big question of thing but it is also a big question of like where am I going to sell and how am like where am I going to sell and how am like where am I going to sell and how am I going to get those customers abroad. I going to get those customers abroad. I going to get those customers abroad. So the data confirms it completely. Um a So the data confirms it completely. Um a So the data confirms it completely. Um a few years ago the fastest growing SAS few years ago the fastest growing SAS few years ago the fastest growing SAS companies were basically reaching 25 companies were basically reaching 25 companies were basically reaching 25 countries in year one and 51 by year countries in year one and 51 by year countries in year one and 51 by year three. Um and today the top AI companies three. Um and today the top AI companies three. Um and today the top AI companies they're already at 42 countries in year they're already at 42 countries in year they're already at 42 countries in year one and at 120 countries in year three.
-
one and at 120 countries in year three. one and at 120 countries in year three. So this is like what I was saying about So this is like what I was saying about So this is like what I was saying about like day one market is basically the like day one market is basically the like day one market is basically the whole world like it's not about having a whole world like it's not about having a whole world like it's not about having a presence in in those countries. It's presence in in those countries. It's presence in in those countries. It's actually getting revenue uh from these actually getting revenue uh from these actually getting revenue uh from these countries from day one. countries from day one. countries from day one. So Gamma is a great example. you So Gamma is a great example. you So Gamma is a great example. you probably know them. Like they're based probably know them. Like they're based probably know them. Like they're based in SF. It's an AI powered like slides in SF. It's an AI powered like slides in SF. It's an AI powered like slides and website builder. Um and they and website builder. Um and they and website builder. Um and they reported 100 million in revenue in year reported 100 million in revenue in year reported 100 million in revenue in year um in year one and the majority of their um in year one and the majority of their um in year one and the majority of their revenue comes from outside the US. So revenue comes from outside the US. So revenue comes from outside the US. So that's definitely showing how um the the that's definitely showing how um the the that's definitely showing how um the the trend is going. And today if you take a trend is going. And today if you take a trend is going. And today if you take a look at overall like all the top AI look at overall like all the top AI look at overall like all the top AI companies 48% of their revenue comes companies 48% of their revenue comes companies 48% of their revenue comes from outside their home market. Like from outside their home market. Like from outside their home market. Like nearly half of a dollar that you get or nearly half of a dollar that you get or nearly half of a dollar that you get or a euro is actually from a foreign a euro is actually from a foreign a euro is actually from a foreign country. And this figure was actually country. And this figure was actually country. And this figure was actually 33% 33% 33% um a few months ago.
-
um a few months ago. um a few months ago. So let's take a look at which markets So let's take a look at which markets So let's take a look at which markets are actually spending a lot on AI tools. are actually spending a lot on AI tools. are actually spending a lot on AI tools. Um so not surprisingly we can see that Um so not surprisingly we can see that Um so not surprisingly we can see that the countries with the high GDP growth the countries with the high GDP growth the countries with the high GDP growth are here. So you can see France, the UK, are here. So you can see France, the UK, are here. So you can see France, the UK, South Korea. Um but we also have like South Korea. Um but we also have like South Korea. Um but we also have like new emerging uh countries that are new emerging uh countries that are new emerging uh countries that are spending a lot on AI. So for example, spending a lot on AI. So for example, spending a lot on AI. So for example, Switzerland, Poland, Turkey Switzerland, Poland, Turkey Switzerland, Poland, Turkey and the demand is there. So now um as and the demand is there. So now um as and the demand is there. So now um as you're building your company, you need you're building your company, you need you're building your company, you need to think about like how am I going to to think about like how am I going to to think about like how am I going to get that demand? how am I going to get get that demand? how am I going to get get that demand? how am I going to get those customers? those customers? those customers? And like what we're saying what we're And like what we're saying what we're And like what we're saying what we're seeing is that you really need to meet seeing is that you really need to meet seeing is that you really need to meet buyers expectations and meet them where buyers expectations and meet them where buyers expectations and meet them where they are. So the way we're doing this is they are. So the way we're doing this is they are. So the way we're doing this is that you offer local currencies and that you offer local currencies and that you offer local currencies and local payment methods. That sounds easy local payment methods. That sounds easy local payment methods. That sounds easy and that is actually easy uh to do with and that is actually easy uh to do with and that is actually easy uh to do with St. But like basically it's really St. But like basically it's really St. But like basically it's really meeting them where they are and with meeting them where they are and with meeting them where they are and with their habits. So localized pricing for their habits. So localized pricing for their habits. So localized pricing for example its own sounds like completely example its own sounds like completely example its own sounds like completely simple but basically it increases 18% simple but basically it increases 18% simple but basically it increases 18% higher crossber revenue like it's crazy higher crossber revenue like it's crazy higher crossber revenue like it's crazy the impact that it has. Uh and basically the impact that it has. Uh and basically the impact that it has. Uh and basically when you add one just one local payment when you add one just one local payment when you add one just one local payment method for a country you have at least method for a country you have at least method for a country you have at least 7% conversion uplift uh in this country.
-
7% conversion uplift uh in this country. 7% conversion uplift uh in this country. So, three checks to pressure test your So, three checks to pressure test your So, three checks to pressure test your global readiness as an AI company today. global readiness as an AI company today. global readiness as an AI company today. First, you localize your prices and your First, you localize your prices and your First, you localize your prices and your payment methods. payment methods. payment methods. Two, automate tax collection because Two, automate tax collection because Two, automate tax collection because like no one wants to get in trouble like no one wants to get in trouble like no one wants to get in trouble either. And three, track performance by either. And three, track performance by either. And three, track performance by country like performance um conversion country like performance um conversion country like performance um conversion and obsess over optimizing those. and obsess over optimizing those. and obsess over optimizing those. So now pricing um because the old models So now pricing um because the old models So now pricing um because the old models are are not working anymore. Uh now we are are not working anymore. Uh now we are are not working anymore. Uh now we have to think about like new pricing have to think about like new pricing have to think about like new pricing models. models. models. The interesting thing is that like value The interesting thing is that like value The interesting thing is that like value is elastic. You probably know that if is elastic. You probably know that if is elastic. You probably know that if you've taken a look a bit at pricing and you've taken a look a bit at pricing and you've taken a look a bit at pricing and so everyone might be using the same AI so everyone might be using the same AI so everyone might be using the same AI tool, but they might not using the same tool, but they might not using the same tool, but they might not using the same way. Um, so take two users. One is an way. Um, so take two users. One is an way. Um, so take two users. One is an engineer. They're going to bed. Next day engineer. They're going to bed. Next day engineer. They're going to bed. Next day they're having a code review. Um, and they're having a code review. Um, and they're having a code review. Um, and the output is really engineering output.
-
the output is really engineering output. the output is really engineering output. The other user is just an everyday user. The other user is just an everyday user. The other user is just an everyday user. They replace uh Safari but GPT and the They replace uh Safari but GPT and the They replace uh Safari but GPT and the output is just an upgraded search output is just an upgraded search output is just an upgraded search experience. experience. experience. So for this like the cost is not the So for this like the cost is not the So for this like the cost is not the same, the value is not the same and so same, the value is not the same and so same, the value is not the same and so the pricing should definitely not be the the pricing should definitely not be the the pricing should definitely not be the same. So this change in value isn't new. same. So this change in value isn't new. same. So this change in value isn't new. We had this with all the technology We had this with all the technology We had this with all the technology shifts that happened like for example we shifts that happened like for example we shifts that happened like for example we had local host hosting at a time. So it had local host hosting at a time. So it had local host hosting at a time. So it was normal to download the software for was normal to download the software for was normal to download the software for a yearly license. Then cloud computing a yearly license. Then cloud computing a yearly license. Then cloud computing arrived and so we had auto updates of arrived and so we had auto updates of arrived and so we had auto updates of softwares. So it was actually uh normal softwares. So it was actually uh normal softwares. So it was actually uh normal to move to subscriptions to move to subscriptions to move to subscriptions and today like what do we have with AI and today like what do we have with AI and today like what do we have with AI so that's basically where you'd expect so that's basically where you'd expect so that's basically where you'd expect me to tell you what the the the best me to tell you what the the the best me to tell you what the the the best answer is. Um so good news bad news bad answer is. Um so good news bad news bad answer is. Um so good news bad news bad news first we it's still early like we news first we it's still early like we news first we it's still early like we don't have a gold standard of what all don't have a gold standard of what all don't have a gold standard of what all companies are doing and it is absolutely companies are doing and it is absolutely companies are doing and it is absolutely the best but good news is that we're the best but good news is that we're the best but good news is that we're clearly seeing um emerging a pattern clearly seeing um emerging a pattern clearly seeing um emerging a pattern um take for example poke to our MC today um take for example poke to our MC today um take for example poke to our MC today um a developer tools company that you um a developer tools company that you um a developer tools company that you probably know that's been around for probably know that's been around for probably know that's been around for around a decade basically and when AI around a decade basically and when AI around a decade basically and when AI tools arrived they pivoted their product
-
tools arrived they pivoted their product tools arrived they pivoted their product they started adding AI um AI tools AI they started adding AI um AI tools AI they started adding AI um AI tools AI features in the product and they also features in the product and they also features in the product and they also started to adapt the pricing like started to adapt the pricing like started to adapt the pricing like layering credits etc. So they started layering credits etc. So they started layering credits etc. So they started with flat subscriptions which is how you with flat subscriptions which is how you with flat subscriptions which is how you stick and create the relationship with stick and create the relationship with stick and create the relationship with the user and then added credits as AI the user and then added credits as AI the user and then added credits as AI became core to the value of the product. became core to the value of the product. became core to the value of the product. So the result is that today they're So the result is that today they're So the result is that today they're targeting 1 billion ARR uh this year and targeting 1 billion ARR uh this year and targeting 1 billion ARR uh this year and they're not the only ones to do this. they're not the only ones to do this. they're not the only ones to do this. This is just one of the examples of what This is just one of the examples of what This is just one of the examples of what we can see. we can see. we can see. So today we have like two out of three So today we have like two out of three So today we have like two out of three companies of the Forbes AI50 companies of the Forbes AI50 companies of the Forbes AI50 um listing that are using usage based um listing that are using usage based um listing that are using usage based pricing. So a majority of these pricing. So a majority of these pricing. So a majority of these companies are actually using a hybrid companies are actually using a hybrid companies are actually using a hybrid model. They're using subscriptions still model. They're using subscriptions still model. They're using subscriptions still for the the stickiness of the for the the stickiness of the for the the stickiness of the relationship, but they're adding a lot relationship, but they're adding a lot relationship, but they're adding a lot of usage based pricing to make sure that of usage based pricing to make sure that of usage based pricing to make sure that the value is actually reflected in the the value is actually reflected in the the value is actually reflected in the pricing.
-
pricing. pricing. So it's good to know that you have to do So it's good to know that you have to do So it's good to know that you have to do some form of usage based pricing but some form of usage based pricing but some form of usage based pricing but getting it right is actually another getting it right is actually another getting it right is actually another story. So in our view there are three story. So in our view there are three story. So in our view there are three things you need to do to add uh to your things you need to do to add uh to your things you need to do to add uh to your playbook for pricing. playbook for pricing. playbook for pricing. First price in your customers language First price in your customers language First price in your customers language like developer think in tokens like developer think in tokens like developer think in tokens enterprises most of the time think in enterprises most of the time think in enterprises most of the time think in seats like you have to think about these seats like you have to think about these seats like you have to think about these uh like who are your customers and make uh like who are your customers and make uh like who are your customers and make sure that you match their pricing with sure that you match their pricing with sure that you match their pricing with the way that they experience value in the way that they experience value in the way that they experience value in your product. your product. your product. Second, you want to show like as close Second, you want to show like as close Second, you want to show like as close as real-time visibility on what they're as real-time visibility on what they're as real-time visibility on what they're consuming because like you know this is consuming because like you know this is consuming because like you know this is this is not a nice to have and this is this is not a nice to have and this is this is not a nice to have and this is basically how you prevent churn as a basically how you prevent churn as a basically how you prevent churn as a company. Like this means finding ways to company. Like this means finding ways to company. Like this means finding ways to make sure they're seeing real time what make sure they're seeing real time what make sure they're seeing real time what they're consuming. And three, we really they're consuming. And three, we really they're consuming. And three, we really see that it's it's strong to sell see that it's it's strong to sell see that it's it's strong to sell credits and not costs because like at credits and not costs because like at credits and not costs because like at one point your customers will exchange one point your customers will exchange one point your customers will exchange money for credits, but then the way they money for credits, but then the way they money for credits, but then the way they will interact with your products will will interact with your products will will interact with your products will not be with like each euro or each cent.
-
not be with like each euro or each cent. not be with like each euro or each cent. It will be with credits and it will be It will be with credits and it will be It will be with credits and it will be related to value mostly, not costs. related to value mostly, not costs. related to value mostly, not costs. So So So now we're going into go to market now we're going into go to market now we're going into go to market motions changing uh which is also a big motions changing uh which is also a big motions changing uh which is also a big part even if like a lot of people are part even if like a lot of people are part even if like a lot of people are actually taking that bit of the side um actually taking that bit of the side um actually taking that bit of the side um but that's actually a big motion but that's actually a big motion but that's actually a big motion changing as well. Um actually entirely changing as well. Um actually entirely changing as well. Um actually entirely new motions are emerging. You might have new motions are emerging. You might have new motions are emerging. You might have seen that and the biggest one is channel seen that and the biggest one is channel seen that and the biggest one is channel sales like it means where your product sales like it means where your product sales like it means where your product is going to be to be sold and how. And is going to be to be sold and how. And is going to be to be sold and how. And like today in the AI go to market motion like today in the AI go to market motion like today in the AI go to market motion for example a lot of um companies are for example a lot of um companies are for example a lot of um companies are actually buying open AI or anthropic via actually buying open AI or anthropic via actually buying open AI or anthropic via their cloud provider. So you need to their cloud provider. So you need to their cloud provider. So you need to think about these different channel think about these different channel think about these different channel sales that are arriving and a lot of AI sales that are arriving and a lot of AI sales that are arriving and a lot of AI companies are actually building their companies are actually building their companies are actually building their own marketplaces where you are going to own marketplaces where you are going to own marketplaces where you are going to be able to be a reference and to be to be able to be a reference and to be to be able to be a reference and to be to be bought as well.
-
be bought as well. be bought as well. And of course like the billion or maybe And of course like the billion or maybe And of course like the billion or maybe trillion dollar question like how do you trillion dollar question like how do you trillion dollar question like how do you sell your products to your newest buyer sell your products to your newest buyer sell your products to your newest buyer which is you know the agent like we've which is you know the agent like we've which is you know the agent like we've spent decades thinking about what is the spent decades thinking about what is the spent decades thinking about what is the best way you know human psychology I best way you know human psychology I best way you know human psychology I will price my product 999 and not€10 but will price my product 999 and not€10 but will price my product 999 and not€10 but like agents absolutely don't care about like agents absolutely don't care about like agents absolutely don't care about any of this any of this any of this and today I think this chart is one of and today I think this chart is one of and today I think this chart is one of the clearest signal of like I need to the clearest signal of like I need to the clearest signal of like I need to think about my agent as a buyer because think about my agent as a buyer because think about my agent as a buyer because like for example this is trade docks uh like for example this is trade docks uh like for example this is trade docks uh so our public documentation the pink so our public documentation the pink so our public documentation the pink line is uh how it's read by agent and line is uh how it's read by agent and line is uh how it's read by agent and the purple line is by humans and as you the purple line is by humans and as you the purple line is by humans and as you can see like this very summer in 26 we can see like this very summer in 26 we can see like this very summer in 26 we had the two lines crossed which means had the two lines crossed which means had the two lines crossed which means that our public documentation today is that our public documentation today is that our public documentation today is more read by agents than by humans. So more read by agents than by humans. So more read by agents than by humans. So we we live it every day at ripe like we we we live it every day at ripe like we we we live it every day at ripe like we have to think about okay how is my have to think about okay how is my have to think about okay how is my product how is my documentation read and product how is my documentation read and product how is my documentation read and understood by agents. So you have to understood by agents. So you have to understood by agents. So you have to change the way you you do things. You change the way you you do things. You change the way you you do things. You have to change the way you are going to have to change the way you are going to have to change the way you are going to showcase your products and make them showcase your products and make them showcase your products and make them discoverable by agents.
-
discoverable by agents. discoverable by agents. And go to market is like the tip of the And go to market is like the tip of the And go to market is like the tip of the iceberg because if you think about like iceberg because if you think about like iceberg because if you think about like changing uh channel motions then you changing uh channel motions then you changing uh channel motions then you also think about like adapting your also think about like adapting your also think about like adapting your product to that making sure that your product to that making sure that your product to that making sure that your onboarding is going to reflect that new onboarding is going to reflect that new onboarding is going to reflect that new uh onboarding motion. Pricing and uh onboarding motion. Pricing and uh onboarding motion. Pricing and revenue of course is going to have to revenue of course is going to have to revenue of course is going to have to reflect that as well your organization. reflect that as well your organization. reflect that as well your organization. Um so for example if you have enterprise Um so for example if you have enterprise Um so for example if you have enterprise sales you are going to have to adapt sales you are going to have to adapt sales you are going to have to adapt this to the new uh new go to mac and this to the new uh new go to mac and this to the new uh new go to mac and motions as well. So this is how I would motions as well. So this is how I would motions as well. So this is how I would think about the playbook. Like basically think about the playbook. Like basically think about the playbook. Like basically today the products are channel agnostic today the products are channel agnostic today the products are channel agnostic and the customers themselves are channel and the customers themselves are channel and the customers themselves are channel agnostic. They don't really care about agnostic. They don't really care about agnostic. They don't really care about where they buy. So they may be self-s where they buy. So they may be self-s where they buy. So they may be self-s served at some at some point graduate to served at some at some point graduate to served at some at some point graduate to enterprise and then like have an agent enterprise and then like have an agent enterprise and then like have an agent implement a new product. So you need to implement a new product. So you need to implement a new product. So you need to think about all of this journey uh for think about all of this journey uh for think about all of this journey uh for your product. your product. your product. So have a clear graduation project pro So have a clear graduation project pro So have a clear graduation project pro progress from productled growth to progress from productled growth to progress from productled growth to salesled growth. Uh then make sure that salesled growth. Uh then make sure that salesled growth. Uh then make sure that you have unified system you have like you have unified system you have like you have unified system you have like one product catalog that you have one one product catalog that you have one one product catalog that you have one customer object um one data model and customer object um one data model and customer object um one data model and that you don't go away in all that you don't go away in all that you don't go away in all directions.
-
directions. directions. And third, like as I was saying, please And third, like as I was saying, please And third, like as I was saying, please design for agent readiness. Like make design for agent readiness. Like make design for agent readiness. Like make sure that your product is ready to be sure that your product is ready to be sure that your product is ready to be discovered and bought by agents without discovered and bought by agents without discovered and bought by agents without a human in the loop. a human in the loop. a human in the loop. Okay. So last pattern, fraud. Um this is Okay. So last pattern, fraud. Um this is Okay. So last pattern, fraud. Um this is actually something very important that actually something very important that actually something very important that you need to think about. Monitor fraud you need to think about. Monitor fraud you need to think about. Monitor fraud um and risk like early and often. um and risk like early and often. um and risk like early and often. So what makes it hard basically is that So what makes it hard basically is that So what makes it hard basically is that like it doesn't show up in one place like it doesn't show up in one place like it doesn't show up in one place like at all at all time of your product like at all at all time of your product like at all at all time of your product life cycle they're going to be a life cycle they're going to be a life cycle they're going to be a potential for an attack. So if you think potential for an attack. So if you think potential for an attack. So if you think about account creation, payment method about account creation, payment method about account creation, payment method added, payment processing and then after added, payment processing and then after added, payment processing and then after that like post usage all of these places that like post usage all of these places that like post usage all of these places are potential um places for fraud. are potential um places for fraud. are potential um places for fraud. So let's take first like sign up top of So let's take first like sign up top of So let's take first like sign up top of the funnel. The thing is, you're going the funnel. The thing is, you're going the funnel. The thing is, you're going to have someone that like a bad actor to have someone that like a bad actor to have someone that like a bad actor that signs up again and again, and so that signs up again and again, and so that signs up again and again, and so they're basically just stealing all of they're basically just stealing all of they're basically just stealing all of your user credits um at at once.
-
your user credits um at at once. your user credits um at at once. One in six signups, to give you an idea, One in six signups, to give you an idea, One in six signups, to give you an idea, today uh involves multi- accounts abuse. today uh involves multi- accounts abuse. today uh involves multi- accounts abuse. So, they're basically farming you free So, they're basically farming you free So, they're basically farming you free trial um without your evering a dollar trial um without your evering a dollar trial um without your evering a dollar from them. And the problem also is that from them. And the problem also is that from them. And the problem also is that JI today has made it very hard to, you JI today has made it very hard to, you JI today has made it very hard to, you know, find fake accounts because they know, find fake accounts because they know, find fake accounts because they actually look real. actually look real. actually look real. And what's interesting is that multi- And what's interesting is that multi- And what's interesting is that multi- account abuse is often like exactly how account abuse is often like exactly how account abuse is often like exactly how free trial um abuse happens at scale. So free trial um abuse happens at scale. So free trial um abuse happens at scale. So you can see that it increases you can see that it increases you can see that it increases dramatically like this is how it dramatically like this is how it dramatically like this is how it accelerated in the straight network. So accelerated in the straight network. So accelerated in the straight network. So you can see that in 2026 we have this you can see that in 2026 we have this you can see that in 2026 we have this crazy increase of free trial views crazy increase of free trial views crazy increase of free trial views because like if you're having success because like if you're having success because like if you're having success like you're actually also an attractive like you're actually also an attractive like you're actually also an attractive target on the internet. target on the internet. target on the internet. Um what's good news though is that it's Um what's good news though is that it's Um what's good news though is that it's a solvable problem. Like just as you can a solvable problem. Like just as you can a solvable problem. Like just as you can see an exponential growth of the see an exponential growth of the see an exponential growth of the attacks, you can also see a very strong attacks, you can also see a very strong attacks, you can also see a very strong and steep decline of the attacks if you and steep decline of the attacks if you and steep decline of the attacks if you have the right uh solution. Like just in have the right uh solution. Like just in have the right uh solution. Like just in four months we blocked like three four months we blocked like three four months we blocked like three million attacks like this for free million attacks like this for free million attacks like this for free trial. So that's definitely part of top trial. So that's definitely part of top trial. So that's definitely part of top of mind for us.
-
of mind for us. of mind for us. And then like of course there's another And then like of course there's another And then like of course there's another vector you you might be familiar with vector you you might be familiar with vector you you might be familiar with pay as you go pricing. So pay as you go pay as you go pricing. So pay as you go pay as you go pricing. So pay as you go abuse goes with that. abuse goes with that. abuse goes with that. So what happens is that a consumer would So what happens is that a consumer would So what happens is that a consumer would basically consume let's say a thou a basically consume let's say a thou a basically consume let's say a thou a thousand or thousands of dollars of thousand or thousands of dollars of thousand or thousands of dollars of tokens in a month and then at the end tokens in a month and then at the end tokens in a month and then at the end you try to get the money and the card is you try to get the money and the card is you try to get the money and the card is just invalid. just invalid. just invalid. So it's kind of like you know Michelin So it's kind of like you know Michelin So it's kind of like you know Michelin start start start like we say in French uh but for tokens like we say in French uh but for tokens like we say in French uh but for tokens right. So by the time you know the right. So by the time you know the right. So by the time you know the compute is gone and you can't charge compute is gone and you can't charge compute is gone and you can't charge anything to that user. So pay as you go anything to that user. So pay as you go anything to that user. So pay as you go is an amazing tool for pricing but you is an amazing tool for pricing but you is an amazing tool for pricing but you have to think about how to manage fraud have to think about how to manage fraud have to think about how to manage fraud in that case. So the companies handling in that case. So the companies handling in that case. So the companies handling that well are monitoring usage signals that well are monitoring usage signals that well are monitoring usage signals in real time and make sure that they see in real time and make sure that they see in real time and make sure that they see all the different signals uh and and all the different signals uh and and all the different signals uh and and avoid that from happening. avoid that from happening. avoid that from happening. Okay. So what do you do about this like Okay. So what do you do about this like Okay. So what do you do about this like three things again?
-
three things again? three things again? uh please monitor this at every stage of uh please monitor this at every stage of uh please monitor this at every stage of the life cycle from account creation to the life cycle from account creation to the life cycle from account creation to post usage. Um so defend all of this for post usage. Um so defend all of this for post usage. Um so defend all of this for your customers. Then I would recommend your customers. Then I would recommend your customers. Then I would recommend to close the gap between usage and to close the gap between usage and to close the gap between usage and billing. Like you don't want to wait a billing. Like you don't want to wait a billing. Like you don't want to wait a month to bill. You want to make sure month to bill. You want to make sure month to bill. You want to make sure that you close that gap as much as that you close that gap as much as that you close that gap as much as possible. possible. possible. And then like please build fraud into And then like please build fraud into And then like please build fraud into your launch plan. like it's not your launch plan. like it's not your launch plan. like it's not necessarily the best or more most necessarily the best or more most necessarily the best or more most exciting part about building but that's exciting part about building but that's exciting part about building but that's actually incredibly important like at actually incredibly important like at actually incredibly important like at every step of your product you need to every step of your product you need to every step of your product you need to think about okay if I am a back actor think about okay if I am a back actor think about okay if I am a back actor now and that I want to take advantage now and that I want to take advantage now and that I want to take advantage what does the attack looks like and so what does the attack looks like and so what does the attack looks like and so in my view if you haven't thought about in my view if you haven't thought about in my view if you haven't thought about this or you don't have the answer yet this or you don't have the answer yet this or you don't have the answer yet like you're not ready to ship it like you're not ready to ship it like you're not ready to ship it so so so the playbook is being rewritten We all the playbook is being rewritten We all the playbook is being rewritten We all know that. But like how you build, how know that. But like how you build, how know that. But like how you build, how you go global, how you price, how you you go global, how you price, how you you go global, how you price, how you adapt your go to market motions and how adapt your go to market motions and how adapt your go to market motions and how you manage fraud, that's what's changed.
-
you manage fraud, that's what's changed. you manage fraud, that's what's changed. Not just the tools, but actually the Not just the tools, but actually the Not just the tools, but actually the whole starting line of building a whole starting line of building a whole starting line of building a company. And you you know that you can company. And you you know that you can company. And you you know that you can build in hours, you can be present in 42 build in hours, you can be present in 42 build in hours, you can be present in 42 countries on day one. As we can see, you countries on day one. As we can see, you countries on day one. As we can see, you can price for value. Um, so that's can price for value. Um, so that's can price for value. Um, so that's actually in our view, you know, the actually in our view, you know, the actually in our view, you know, the funniest and maybe craziest time of like funniest and maybe craziest time of like funniest and maybe craziest time of like building uh for for building a company building uh for for building a company building uh for for building a company in the world. Um, so we we hope that in the world. Um, so we we hope that in the world. Um, so we we hope that these insights give you some some idea these insights give you some some idea these insights give you some some idea of like how to build this and yeah, like of like how to build this and yeah, like of like how to build this and yeah, like our job at Stripe is to be there with our job at Stripe is to be there with our job at Stripe is to be there with you to help you scale. So um I'm really you to help you scale. So um I'm really you to help you scale. So um I'm really excited to see what you build. Uh and excited to see what you build. Uh and excited to see what you build. Uh and just please know that we are there if just please know that we are there if just please know that we are there if you need to scale as well. and drop me a you need to scale as well. and drop me a you need to scale as well. and drop me a message if you want to chat more. Thank message if you want to chat more. Thank message if you want to chat more. Thank you. >> All right. >> All right. Thank you, Ariel. Let's give it up for Thank you, Ariel. Let's give it up for Thank you, Ariel. Let's give it up for Ariel once again, please. Ariel once again, please. Ariel once again, please. All right. Okay. I think what uh the All right. Okay. I think what uh the All right. Okay. I think what uh the fact that um or the information the fact that um or the information the fact that um or the information the piece of information that uh stayed with piece of information that uh stayed with piece of information that uh stayed with me through this presentation is that you me through this presentation is that you me through this presentation is that you need to design for humans but you also need to design for humans but you also need to design for humans but you also need to design your products for agents need to design your products for agents need to design your products for agents because they might you know be because they might you know be because they might you know be purchasing your products as well which purchasing your products as well which purchasing your products as well which is I think is awesome. Um I asked is I think is awesome. Um I asked is I think is awesome. Um I asked earlier a question. I said how many earlier a question. I said how many earlier a question. I said how many people are builders here? And many of people are builders here? And many of people are builders here? And many of you raised their hands, right? Like you raised their hands, right? Like you raised their hands, right? Like yeah, most of you are builders, right?
-
yeah, most of you are builders, right? yeah, most of you are builders, right? Like you're building products. Um and Like you're building products. Um and Like you're building products. Um and you're uh maybe chatting with agents. Um you're uh maybe chatting with agents. Um you're uh maybe chatting with agents. Um my personal experience is that I kind of my personal experience is that I kind of my personal experience is that I kind of stopped uh typing to my agent. Um so I stopped uh typing to my agent. Um so I stopped uh typing to my agent. Um so I only chat mostly with voice and I find only chat mostly with voice and I find only chat mostly with voice and I find it like a a more natural way to interact it like a a more natural way to interact it like a a more natural way to interact with agent. I read faster than I than I with agent. I read faster than I than I with agent. I read faster than I than I hear. But I also I I talk a lot faster hear. But I also I I talk a lot faster hear. But I also I I talk a lot faster than I than I than I type. Uh I don't than I than I than I type. Uh I don't than I than I than I type. Uh I don't know if anybody shares that experience know if anybody shares that experience know if anybody shares that experience here. Who talks to their agent? Yeah. here. Who talks to their agent? Yeah. here. Who talks to their agent? Yeah. Okay. So it's a actually a good number. Okay. So it's a actually a good number. Okay. So it's a actually a good number. Well, our next speaker is here to talk Well, our next speaker is here to talk Well, our next speaker is here to talk to us actually about voice AI and real to us actually about voice AI and real to us actually about voice AI and real time. Um so without further ado, please time. Um so without further ado, please time. Um so without further ado, please join me in welcoming to the stage chief join me in welcoming to the stage chief join me in welcoming to the stage chief scientist and co-founder of Piano AI, scientist and co-founder of Piano AI, scientist and co-founder of Piano AI, Erve Ba.
-
Okay, let me try this thing. Yeah, it's Okay, let me try this thing. Yeah, it's working. Um, hi everyone. Uh, thank you working. Um, hi everyone. Uh, thank you working. Um, hi everyone. Uh, thank you Rau for the introduction. Uh uh it's the Rau for the introduction. Uh uh it's the Rau for the introduction. Uh uh it's the first time I've been ever announced on a first time I've been ever announced on a first time I've been ever announced on a stage by an MC. I feel like a bit like stage by an MC. I feel like a bit like stage by an MC. I feel like a bit like Celindio in Paris uh tonight uh or today Celindio in Paris uh tonight uh or today Celindio in Paris uh tonight uh or today at least. Don't worry, I won't sing at least. Don't worry, I won't sing at least. Don't worry, I won't sing anything but I will talk to you about anything but I will talk to you about anything but I will talk to you about Voice AI and in particular about uh Voice AI and in particular about uh Voice AI and in particular about uh speaker dization. Um so uh quick speaker dization. Um so uh quick speaker dization. Um so uh quick introduction about myself. So I'm the introduction about myself. So I'm the introduction about myself. So I'm the co-founder and chief science officer at co-founder and chief science officer at co-founder and chief science officer at Pianoai. I've been an academic Pianoai. I've been an academic Pianoai. I've been an academic researcher most of my career up until researcher most of my career up until researcher most of my career up until two years ago when we uh created this two years ago when we uh created this two years ago when we uh created this company and um so yeah I've been mostly company and um so yeah I've been mostly company and um so yeah I've been mostly focusing on speakerization for the last focusing on speakerization for the last focusing on speakerization for the last 15 years and uh it looks like people 15 years and uh it looks like people 15 years and uh it looks like people think that uh uh I know a few things think that uh uh I know a few things think that uh uh I know a few things about speakerization and I hope today about speakerization and I hope today about speakerization and I hope today that uh I'll tell you what that uh I'll tell you what that uh I'll tell you what speakerization here is and that you'll speakerization here is and that you'll speakerization here is and that you'll learn a bit about how that works and uh learn a bit about how that works and uh learn a bit about how that works and uh yeah that at least you you take yeah that at least you you take yeah that at least you you take something out of this talk.
-
something out of this talk. something out of this talk. So what is speakerization? So what is speakerization? So what is speakerization? So speakerization u basically uh is the So speakerization u basically uh is the So speakerization u basically uh is the task of taking a recording of a task of taking a recording of a task of taking a recording of a conversation between multiple speakers conversation between multiple speakers conversation between multiple speakers uh and outputting some metadata about uh and outputting some metadata about uh and outputting some metadata about it. What you can do starting from a it. What you can do starting from a it. What you can do starting from a conversation is basically do voice conversation is basically do voice conversation is basically do voice activity detection first where you would activity detection first where you would activity detection first where you would uh detect uh the parts where uh someone uh detect uh the parts where uh someone uh detect uh the parts where uh someone is speaking and when nobody's speaking. is speaking and when nobody's speaking. is speaking and when nobody's speaking. That's voice activity detection. The That's voice activity detection. The That's voice activity detection. The second step that you can do on top of second step that you can do on top of second step that you can do on top of that is uh splitting those uh very long that is uh splitting those uh very long that is uh splitting those uh very long uh speech turn regions into uh single uh speech turn regions into uh single uh speech turn regions into uh single speaker regions. Uh that's this task is speaker regions. Uh that's this task is speaker regions. Uh that's this task is called segmentation. So it includes both called segmentation. So it includes both called segmentation. So it includes both uh speaker change detection as well as uh speaker change detection as well as uh speaker change detection as well as overlapping speech detection. So in this overlapping speech detection. So in this overlapping speech detection. So in this example here uh in in the middle you see example here uh in in the middle you see example here uh in in the middle you see that there are it looks like there might that there are it looks like there might that there are it looks like there might be uh u at least two speakers because uh be uh u at least two speakers because uh be uh u at least two speakers because uh there's a uh someone interrupting um uh there's a uh someone interrupting um uh there's a uh someone interrupting um uh another one or starting speaking after another one or starting speaking after another one or starting speaking after the end of the another speech turn and a the end of the another speech turn and a the end of the another speech turn and a bit later in the conversation there is a bit later in the conversation there is a bit later in the conversation there is a very small speech turn that overlap with very small speech turn that overlap with very small speech turn that overlap with the other one. This these are usually the other one. This these are usually the other one. This these are usually what we call back channels like mhm okay what we call back channels like mhm okay what we call back channels like mhm okay yes that you say when you are nodding at yes that you say when you are nodding at yes that you say when you are nodding at someone uh uh talking to you and those someone uh uh talking to you and those someone uh uh talking to you and those events are very important to uh uh for events are very important to uh uh for events are very important to uh uh for voice AI agents to uh basically voice AI agents to uh basically voice AI agents to uh basically understand really the core of the understand really the core of the understand really the core of the conversation the words are usually not conversation the words are usually not conversation the words are usually not enough knowing that when someone is enough knowing that when someone is enough knowing that when someone is interrupting someone else or nodding at interrupting someone else or nodding at interrupting someone else or nodding at what they are saying it gives a lot more
-
what they are saying it gives a lot more what they are saying it gives a lot more information about the the conversation information about the the conversation information about the the conversation and the next step really what we call and the next step really what we call and the next step really what we call speakerization is then once the the speakerization is then once the the speakerization is then once the the conversation is segmented into speech conversation is segmented into speech conversation is segmented into speech turn is to actually assign each speech turn is to actually assign each speech turn is to actually assign each speech turn to the right speaker and this is turn to the right speaker and this is turn to the right speaker and this is done without knowing in advance the done without knowing in advance the done without knowing in advance the number of speakers. So that makes the number of speakers. So that makes the number of speakers. So that makes the task more difficult than usual uh task more difficult than usual uh task more difficult than usual uh supervised machine learning task and we supervised machine learning task and we supervised machine learning task and we don't know the identity as well of the don't know the identity as well of the don't know the identity as well of the speaker. We have to uh basically speaker. We have to uh basically speaker. We have to uh basically generalize to any speaker even those generalize to any speaker even those generalize to any speaker even those that are not part of any training sets. that are not part of any training sets. that are not part of any training sets. So that's what makes speakerization So that's what makes speakerization So that's what makes speakerization difficult and so there are many tasks uh difficult and so there are many tasks uh difficult and so there are many tasks uh where speakerization can be useful uh where speakerization can be useful uh where speakerization can be useful uh and lots of those task can fall actually and lots of those task can fall actually and lots of those task can fall actually under the knowing who said what is just under the knowing who said what is just under the knowing who said what is just as important as what was said. So for as important as what was said. So for as important as what was said. So for instance uh for video dubbing instance uh for video dubbing instance uh for video dubbing application when you want to start from application when you want to start from application when you want to start from a a video in French for instance and you a a video in French for instance and you a a video in French for instance and you want uh basically to put it on the want uh basically to put it on the want uh basically to put it on the internet and reach a larger audience an internet and reach a larger audience an internet and reach a larger audience an international audience uh then you can international audience uh then you can international audience uh then you can do use video dubbing uh tools to uh do use video dubbing uh tools to uh do use video dubbing uh tools to uh basically translate from French to other basically translate from French to other basically translate from French to other languages and making sure that the you languages and making sure that the you languages and making sure that the you know the conversation remains natural.
-
know the conversation remains natural. know the conversation remains natural. So you have to put the right uh uh So you have to put the right uh uh So you have to put the right uh uh speaker clone uh voice to on the right speaker clone uh voice to on the right speaker clone uh voice to on the right speaker and for that you need to know speaker and for that you need to know speaker and for that you need to know exactly who speaks when in the original exactly who speaks when in the original exactly who speaks when in the original audio. Other task uh obvious task are audio. Other task uh obvious task are audio. Other task uh obvious task are meeting note takers uh especially in meeting note takers uh especially in meeting note takers uh especially in hybrid meetings where uh some people are hybrid meetings where uh some people are hybrid meetings where uh some people are joining from the same meeting room. when joining from the same meeting room. when joining from the same meeting room. when there are multiple streams, there are multiple streams, there are multiple streams, speakerization not necessarily uh is speakerization not necessarily uh is speakerization not necessarily uh is useful, but when people are in the same useful, but when people are in the same useful, but when people are in the same room, that's where it's it's really room, that's where it's it's really room, that's where it's it's really useful. And this allows to to make some useful. And this allows to to make some useful. And this allows to to make some nice uh uh summaries of the conversation nice uh uh summaries of the conversation nice uh uh summaries of the conversation of who's supposed to do what after the of who's supposed to do what after the of who's supposed to do what after the meeting, for instance. And other meeting, for instance. And other meeting, for instance. And other applications like podcast intelligence applications like podcast intelligence applications like podcast intelligence where you want to track one particular where you want to track one particular where you want to track one particular speaker in multiple podcasts or track speaker in multiple podcasts or track speaker in multiple podcasts or track the host of a podcast across episodes. the host of a podcast across episodes. the host of a podcast across episodes. And um so we have this uh open source And um so we have this uh open source And um so we have this uh open source toolkit called uh piano to audio uh that toolkit called uh piano to audio uh that toolkit called uh piano to audio uh that uh allows you actually to get from raw uh allows you actually to get from raw uh allows you actually to get from raw audio recording to uh uh an actual audio recording to uh uh an actual audio recording to uh uh an actual speaker deration. So it's it's become speaker deration. So it's it's become speaker deration. So it's it's become very popular over the years. Uh I have very popular over the years. Uh I have very popular over the years. Uh I have some vanity metrics here uh uh thanks to some vanity metrics here uh uh thanks to some vanity metrics here uh uh thanks to our friends at hugging face. It's also our friends at hugging face. It's also our friends at hugging face. It's also very visible there. Uh it has uh on very visible there. Uh it has uh on very visible there. Uh it has uh on average 10 million downloads every every average 10 million downloads every every average 10 million downloads every every month. And basically in a few lines of month. And basically in a few lines of month. And basically in a few lines of code you can get from the audio to the code you can get from the audio to the code you can get from the audio to the actual speakerization. So here in this actual speakerization. So here in this actual speakerization. So here in this example we start for the first block of example we start for the first block of example we start for the first block of the code as basically downloading uh the the code as basically downloading uh the the code as basically downloading uh the community one model the open source community one model the open source community one model the open source community one model from hugging phase.
-
community one model from hugging phase. community one model from hugging phase. You get this uh object called community You get this uh object called community You get this uh object called community one here in the P Python code and then one here in the P Python code and then one here in the P Python code and then you simply apply uh this uh uh pipeline you simply apply uh this uh uh pipeline you simply apply uh this uh uh pipeline on a recording of a conversation and you on a recording of a conversation and you on a recording of a conversation and you get the prediction that you can then get the prediction that you can then get the prediction that you can then iterate over to know exactly start time, iterate over to know exactly start time, iterate over to know exactly start time, end time and the actual speaker time and end time and the actual speaker time and end time and the actual speaker time and this really uh in in three lines of this really uh in in three lines of this really uh in in three lines of code. code. code. Um Um Um yeah so this uh that I described was yeah so this uh that I described was yeah so this uh that I described was batch in the sense that uh the whole batch in the sense that uh the whole batch in the sense that uh the whole processing happened after the processing happened after the processing happened after the conversation took place. So why uh go conversation took place. So why uh go conversation took place. So why uh go real time? So there are many uh reasons real time? So there are many uh reasons real time? So there are many uh reasons why we we would want to go real time. Uh why we we would want to go real time. Uh why we we would want to go real time. Uh for voice agent for voice agent for voice agent it's not that obvious because usually it's not that obvious because usually it's not that obvious because usually there's one person talking on an or to there's one person talking on an or to there's one person talking on an or to an agent and uh you you don't really an agent and uh you you don't really an agent and uh you you don't really care about uh speakerization because care about uh speakerization because care about uh speakerization because there's only one speaker. But very soon there's only one speaker. But very soon there's only one speaker. But very soon and it's already there actually the and it's already there actually the and it's already there actually the agent will listen to to the conversation agent will listen to to the conversation agent will listen to to the conversation while we're having a meeting for while we're having a meeting for while we're having a meeting for instance and to a use case would be to instance and to a use case would be to instance and to a use case would be to flag some speakers interrupting another flag some speakers interrupting another flag some speakers interrupting another speaker too often or to flag a speaker speaker too often or to flag a speaker speaker too often or to flag a speaker talking too much. Uh and doing that in talking too much. Uh and doing that in talking too much. Uh and doing that in real time is is necessary.
-
real time is is necessary. real time is is necessary. And uh the next step is also that we And uh the next step is also that we And uh the next step is also that we soon are going to have robots in our soon are going to have robots in our soon are going to have robots in our house. And uh to basically answer to uh house. And uh to basically answer to uh house. And uh to basically answer to uh a child or to the adult in the in the a child or to the adult in the in the a child or to the adult in the in the household, you would uh basically need household, you would uh basically need household, you would uh basically need to know who's who. And this is really to know who's who. And this is really to know who's who. And this is really where uh real time voice darization where uh real time voice darization where uh real time voice darization speakerization is useful. speakerization is useful. speakerization is useful. So now that we know that it's actually So now that we know that it's actually So now that we know that it's actually useful, how uh did we at PenAI uh manage useful, how uh did we at PenAI uh manage useful, how uh did we at PenAI uh manage to go real time? to go real time? to go real time? Uh I need first to take a step back to Uh I need first to take a step back to Uh I need first to take a step back to explain how the batch speakerization explain how the batch speakerization explain how the batch speakerization that I just described earlier in three that I just described earlier in three that I just described earlier in three lines of code called community one works lines of code called community one works lines of code called community one works internally. internally. internally. So uh I'm going to spend some time on So uh I'm going to spend some time on So uh I'm going to spend some time on this slide a bit more. So um basically this slide a bit more. So um basically this slide a bit more. So um basically given a let's say a 1 hour conversation given a let's say a 1 hour conversation given a let's say a 1 hour conversation uh a recording of a 1 hour conversation uh a recording of a 1 hour conversation uh a recording of a 1 hour conversation how it works how community one works how it works how community one works how it works how community one works today is we start by splitting the today is we start by splitting the today is we start by splitting the conversation into pieces let's say 20 conversation into pieces let's say 20 conversation into pieces let's say 20 seconds or 30 seconds pieces. So each uh seconds or 30 seconds pieces. So each uh seconds or 30 seconds pieces. So each uh green rectangle here in in this slide uh green rectangle here in in this slide uh green rectangle here in in this slide uh actually uh represent 30 seconds of actually uh represent 30 seconds of actually uh represent 30 seconds of speech let's say speech let's say speech let's say you see that they are overlapping a bit.
-
you see that they are overlapping a bit. you see that they are overlapping a bit. So we we move that in a in a sliding uh So we we move that in a in a sliding uh So we we move that in a in a sliding uh manner uh uh in a sliding window manner manner uh uh in a sliding window manner manner uh uh in a sliding window manner and then we apply the first uh neural and then we apply the first uh neural and then we apply the first uh neural network that basically performs uh network that basically performs uh network that basically performs uh speakerization on each chunk separately. speakerization on each chunk separately. speakerization on each chunk separately. So why do we do that? Why why do we uh So why do we do that? Why why do we uh So why do we do that? Why why do we uh process the conversation pieces by process the conversation pieces by process the conversation pieces by pieces rather than uh trying to do that pieces rather than uh trying to do that pieces rather than uh trying to do that globally? It's because uh working on globally? It's because uh working on globally? It's because uh working on chunks allows to uh have a limited chunks allows to uh have a limited chunks allows to uh have a limited number of speakers as output. So the the number of speakers as output. So the the number of speakers as output. So the the the machine learning task is actually the machine learning task is actually the machine learning task is actually easier and safer to to solve like that. easier and safer to to solve like that. easier and safer to to solve like that. So basically the model is trained to So basically the model is trained to So basically the model is trained to take 30 seconds of audio as input and take 30 seconds of audio as input and take 30 seconds of audio as input and returns uh basically a matrix of let's returns uh basically a matrix of let's returns uh basically a matrix of let's say uh the time and the number of say uh the time and the number of say uh the time and the number of speakers uh and um it will output zero speakers uh and um it will output zero speakers uh and um it will output zero and once whether each speaker is active and once whether each speaker is active and once whether each speaker is active over time. And the problem with this over time. And the problem with this over time. And the problem with this kind of a first approach is that at that kind of a first approach is that at that kind of a first approach is that at that point uh the same speaker on the first point uh the same speaker on the first point uh the same speaker on the first chunk might have a different uh label on chunk might have a different uh label on chunk might have a different uh label on the same for the in a different chunk.
-
the same for the in a different chunk. the same for the in a different chunk. So we need to rely on a second step uh So we need to rely on a second step uh So we need to rely on a second step uh where each speaker in each chunk here where each speaker in each chunk here where each speaker in each chunk here represented by symbols the triangles the represented by symbols the triangles the represented by symbols the triangles the circles and the stars are projected into circles and the stars are projected into circles and the stars are projected into a highdimensional space using speaker a highdimensional space using speaker a highdimensional space using speaker embeddings. I'm pretty sure you're all embeddings. I'm pretty sure you're all embeddings. I'm pretty sure you're all familiar with what embeddings are, but familiar with what embeddings are, but familiar with what embeddings are, but basically those are uh highdimensional basically those are uh highdimensional basically those are uh highdimensional representation of a thing. Here the representation of a thing. Here the representation of a thing. Here the thing is a speech turn. So each speech thing is a speech turn. So each speech thing is a speech turn. So each speech turn is uh projected in into the space. turn is uh projected in into the space. turn is uh projected in into the space. Uh here it's represented as 3D but in Uh here it's represented as 3D but in Uh here it's represented as 3D but in practice it's uh Oops. Uh in practice practice it's uh Oops. Uh in practice practice it's uh Oops. Uh in practice it's um it's um it's um it's coming back. Is it? It is. I should it's coming back. Is it? It is. I should it's coming back. Is it? It is. I should probably stay here. Um and uh this probably stay here. Um and uh this probably stay here. Um and uh this speakering model, it's a neural network speakering model, it's a neural network speakering model, it's a neural network as well. It's trained to have two uh as well. It's trained to have two uh as well. It's trained to have two uh speech turns of the same speaker close speech turns of the same speaker close speech turns of the same speaker close to each other and two speech turns from to each other and two speech turns from to each other and two speech turns from two different speakers far away from two different speakers far away from two different speakers far away from each other. And this way we can actually each other. And this way we can actually each other. And this way we can actually find uh clusters in this space. And this find uh clusters in this space. And this find uh clusters in this space. And this is how basically we managed to reconcile is how basically we managed to reconcile is how basically we managed to reconcile all the uh the chunks finally and be all the uh the chunks finally and be all the uh the chunks finally and be able to basically stitch uh uh all those able to basically stitch uh uh all those able to basically stitch uh uh all those small speech turn together and and get small speech turn together and and get small speech turn together and and get the final results. So that's the the final results. So that's the the final results. So that's the starting point of um um of speaker starting point of um um of speaker starting point of um um of speaker dization in a batch manner.
-
dization in a batch manner. dization in a batch manner. So the question is how did we get from So the question is how did we get from So the question is how did we get from uh this uh batch diorization to an uh this uh batch diorization to an uh this uh batch diorization to an actual uh realtime implementation. So actual uh realtime implementation. So actual uh realtime implementation. So the first thing that we uh worked on is the first thing that we uh worked on is the first thing that we uh worked on is basically the clustering part that I basically the clustering part that I basically the clustering part that I just mentioned and instead of having it just mentioned and instead of having it just mentioned and instead of having it uh global having it incremental. So uh uh global having it incremental. So uh uh global having it incremental. So uh this uh um animation will will say it this uh um animation will will say it this uh um animation will will say it all. So uh so basically starting with a all. So uh so basically starting with a all. So uh so basically starting with a conversation uh between two speakers conversation uh between two speakers conversation uh between two speakers here in batch dorization we start by here in batch dorization we start by here in batch dorization we start by extracting speaker embeddings from all extracting speaker embeddings from all extracting speaker embeddings from all those speech turns. So you get them all those speech turns. So you get them all those speech turns. So you get them all those points are actually speaker those points are actually speaker those points are actually speaker embeddings and you simply do uh embeddings and you simply do uh embeddings and you simply do uh clustering uh into that. But if you turn clustering uh into that. But if you turn clustering uh into that. But if you turn to streaming you don't have access to to streaming you don't have access to to streaming you don't have access to the future because the conversation did the future because the conversation did the future because the conversation did not happen yet. So the idea is that uh not happen yet. So the idea is that uh not happen yet. So the idea is that uh the clustering algorithm needs to work the clustering algorithm needs to work the clustering algorithm needs to work incrementally and that makes its job way incrementally and that makes its job way incrementally and that makes its job way harder because you only have a limited harder because you only have a limited harder because you only have a limited view of the conversation. So uh you see view of the conversation. So uh you see view of the conversation. So uh you see here uh this is a fake example but it here uh this is a fake example but it here uh this is a fake example but it takes some time for the clustering takes some time for the clustering takes some time for the clustering algorithm to realize that there is algorithm to realize that there is algorithm to realize that there is actually two speakers and that's why it actually two speakers and that's why it actually two speakers and that's why it makes um you know switching from batch makes um you know switching from batch makes um you know switching from batch to real time quite difficult if we if we to real time quite difficult if we if we to real time quite difficult if we if we were to stay with the you know chunk were to stay with the you know chunk were to stay with the you know chunk wise approach that I just mentioned.
-
wise approach that I just mentioned. wise approach that I just mentioned. it still managed in the end uh to find it still managed in the end uh to find it still managed in the end uh to find the right clustering but you saw that the right clustering but you saw that the right clustering but you saw that during uh the actual uh conversation it during uh the actual uh conversation it during uh the actual uh conversation it made some mistakes made some mistakes made some mistakes and so with this kind of approach uh and so with this kind of approach uh and so with this kind of approach uh this is actually implemented in the this is actually implemented in the this is actually implemented in the dieart open source toolkit which is uh a dieart open source toolkit which is uh a dieart open source toolkit which is uh a creation of our CTO Juan Manuel Ka and creation of our CTO Juan Manuel Ka and creation of our CTO Juan Manuel Ka and we managed to uh basically get slightly we managed to uh basically get slightly we managed to uh basically get slightly uh higher diorization error rate uh at uh higher diorization error rate uh at uh higher diorization error rate uh at around 5-second latency because the around 5-second latency because the around 5-second latency because the chunks are around 5 seconds long or at chunks are around 5 seconds long or at chunks are around 5 seconds long or at least the steps between the chunks are 5 least the steps between the chunks are 5 least the steps between the chunks are 5 seconds. seconds. seconds. So we said okay uh that's great but the So we said okay uh that's great but the So we said okay uh that's great but the next step would be to manage to lower next step would be to manage to lower next step would be to manage to lower this diation error rate at least to be this diation error rate at least to be this diation error rate at least to be as good as a batch speaker derization as good as a batch speaker derization as good as a batch speaker derization and also this 5-second latency might be and also this 5-second latency might be and also this 5-second latency might be a bit too much. So um what we worked on a bit too much. So um what we worked on a bit too much. So um what we worked on next is basically uh turning this uh next is basically uh turning this uh next is basically uh turning this uh chunk wise segmentation into something a chunk wise segmentation into something a chunk wise segmentation into something a bit more streaming that we call causal bit more streaming that we call causal bit more streaming that we call causal segmentation where instead of waiting segmentation where instead of waiting segmentation where instead of waiting for the whole chunk to be available to for the whole chunk to be available to for the whole chunk to be available to process it basically pro processing in process it basically pro processing in process it basically pro processing in in streaming.
-
in streaming. in streaming. So that that brings uh a lot more uh of So that that brings uh a lot more uh of So that that brings uh a lot more uh of other difficulties. So uh let's focus other difficulties. So uh let's focus other difficulties. So uh let's focus for instance on speaker chain detection for instance on speaker chain detection for instance on speaker chain detection uh in batch mode. uh what we you we uh in batch mode. uh what we you we uh in batch mode. uh what we you we usually have access to the whole usually have access to the whole usually have access to the whole conversation. So speaker chain detection conversation. So speaker chain detection conversation. So speaker chain detection can be done like that. You slide a can be done like that. You slide a can be done like that. You slide a window uh around the course of the window uh around the course of the window uh around the course of the conversation. You compare what happens conversation. You compare what happens conversation. You compare what happens to the left of the the segment and to to the left of the the segment and to to the left of the the segment and to the right and you compute the difference the right and you compute the difference the right and you compute the difference and you get some nice peak where there and you get some nice peak where there and you get some nice peak where there is a speaker change. But in streaming uh is a speaker change. But in streaming uh is a speaker change. But in streaming uh the problem is that you don't have the problem is that you don't have the problem is that you don't have access to the future. So it takes some access to the future. So it takes some access to the future. So it takes some time for the model to actually reach uh time for the model to actually reach uh time for the model to actually reach uh you know confidence enough uh to um you know confidence enough uh to um you know confidence enough uh to um enough confidence to basically take a enough confidence to basically take a enough confidence to basically take a decision and this uh small amount of decision and this uh small amount of decision and this uh small amount of time uh that it needs in the future to time uh that it needs in the future to time uh that it needs in the future to to decide whether there was a speaker to decide whether there was a speaker to decide whether there was a speaker change or not. We call that the look change or not. We call that the look change or not. We call that the look ahead and that's basically a a a latency ahead and that's basically a a a latency ahead and that's basically a a a latency that we can't get rid of because if we that we can't get rid of because if we that we can't get rid of because if we at one particular point of time in the at one particular point of time in the at one particular point of time in the conversation you need to decide whether conversation you need to decide whether conversation you need to decide whether there is a speaker change uh unless you there is a speaker change uh unless you there is a speaker change uh unless you have a very strong modeling of the pro have a very strong modeling of the pro have a very strong modeling of the pro or something like that. You really have or something like that. You really have or something like that. You really have to wait whether someone else is starting to wait whether someone else is starting to wait whether someone else is starting speaking to to actually take the speaking to to actually take the speaking to to actually take the decision. So that brings this uh decision. So that brings this uh decision. So that brings this uh incompressible uh latency that is added incompressible uh latency that is added incompressible uh latency that is added to the to the thing.
-
to the to the thing. to the to the thing. And uh so we we went ahead and uh and And uh so we we went ahead and uh and And uh so we we went ahead and uh and implemented that and we uh with with implemented that and we uh with with implemented that and we uh with with basically lots of research on the on the basically lots of research on the on the basically lots of research on the on the research team on this particular research team on this particular research team on this particular segmentation streaming model uh we segmentation streaming model uh we segmentation streaming model uh we managed to actually using much larger managed to actually using much larger managed to actually using much larger model than the chunk wise one and also model than the chunk wise one and also model than the chunk wise one and also some uh some uh implementation of loss some uh some uh implementation of loss some uh some uh implementation of loss function. I won't go into the detail. function. I won't go into the detail. function. I won't go into the detail. I'm happy to discuss that after that. I'm happy to discuss that after that. I'm happy to discuss that after that. But we managed to basically reach the But we managed to basically reach the But we managed to basically reach the same level of accuracy as the batch same level of accuracy as the batch same level of accuracy as the batch speaker deration with a a very low speaker deration with a a very low speaker deration with a a very low latency of 400 milliseconds. And this latency of 400 milliseconds. And this latency of 400 milliseconds. And this 400 milliseconds actually um involves 400 milliseconds actually um involves 400 milliseconds actually um involves many things. It it involves also uh many things. It it involves also uh many things. It it involves also uh round trip from uh uh the client to the round trip from uh uh the client to the round trip from uh uh the client to the servers and back. So I'd like to spend servers and back. So I'd like to spend servers and back. So I'd like to spend some time on the actual 40 400 millconds some time on the actual 40 400 millconds some time on the actual 40 400 millconds latency. So what's in those 400 latency. So what's in those 400 latency. So what's in those 400 millisecond latency? millisecond latency? millisecond latency? So um we start by a client uh having a So um we start by a client uh having a So um we start by a client uh having a conversation locally. The the first conversation locally. The the first conversation locally. The the first latency that you get is basically that latency that you get is basically that latency that you get is basically that the conversation is split into 100 the conversation is split into 100 the conversation is split into 100 millisecond blocks. So you already have millisecond blocks. So you already have millisecond blocks. So you already have to wait 100 millisecond. Uh let's say at to wait 100 millisecond. Uh let's say at to wait 100 millisecond. Uh let's say at time t you have to wait 100 milliseconds time t you have to wait 100 milliseconds time t you have to wait 100 milliseconds to build this block. So already at time to build this block. So already at time to build this block. So already at time t we already have 100 millcond latency.
-
t we already have 100 millcond latency. t we already have 100 millcond latency. This latency that I put here uh is part This latency that I put here uh is part This latency that I put here uh is part of the 400 millisecond latency and it's of the 400 millisecond latency and it's of the 400 millisecond latency and it's not something that is often uh counted not something that is often uh counted not something that is often uh counted in the numbers that you can see here and in the numbers that you can see here and in the numbers that you can see here and there. But it's actually also something there. But it's actually also something there. But it's actually also something uh that is important to take into uh that is important to take into uh that is important to take into account. account. account. Then uh the next step basically is we Then uh the next step basically is we Then uh the next step basically is we are relying on the Cloudflare um durable are relying on the Cloudflare um durable are relying on the Cloudflare um durable object uh uh technology to basically object uh uh technology to basically object uh uh technology to basically send this block size to the uh closest send this block size to the uh closest send this block size to the uh closest uh Cloudflare server. So this on average uh Cloudflare server. So this on average uh Cloudflare server. So this on average takes 20 millisecond but it really takes 20 millisecond but it really takes 20 millisecond but it really depends on where the client is and uh depends on where the client is and uh depends on where the client is and uh how good cloudflare is basically. Um and how good cloudflare is basically. Um and how good cloudflare is basically. Um and on this server what we do is we receive on this server what we do is we receive on this server what we do is we receive many streams and we batch them. So we many streams and we batch them. So we many streams and we batch them. So we batch lots of 100 millisecond blocks so batch lots of 100 millisecond blocks so batch lots of 100 millisecond blocks so that uh we can process multiple streams that uh we can process multiple streams that uh we can process multiple streams in parallel. So this is very fast and in parallel. So this is very fast and in parallel. So this is very fast and then once we've uh built those batches then once we've uh built those batches then once we've uh built those batches we send them to model where uh the GPU we send them to model where uh the GPU we send them to model where uh the GPU horsework lives. horsework lives. horsework lives. So basically for segmenting those 100 So basically for segmenting those 100 So basically for segmenting those 100 millisecond latency uh a batch size millisecond latency uh a batch size millisecond latency uh a batch size let's say of 32 32 streams let's say let's say of 32 32 streams let's say let's say of 32 32 streams let's say that's just an example here but it takes that's just an example here but it takes that's just an example here but it takes around uh 50 milliseconds just for the around uh 50 milliseconds just for the around uh 50 milliseconds just for the inference and also you have to take into inference and also you have to take into inference and also you have to take into account this 200 milliseconds look ahead account this 200 milliseconds look ahead account this 200 milliseconds look ahead that I mentioned earlier. So basically that I mentioned earlier. So basically that I mentioned earlier. So basically when you reach time t what the model the when you reach time t what the model the when you reach time t what the model the neural network does it predicts what neural network does it predicts what neural network does it predicts what happened 200 milliseconds in the past.
-
happened 200 milliseconds in the past. happened 200 milliseconds in the past. So this is uh this adds up and then in So this is uh this adds up and then in So this is uh this adds up and then in parallel uh we we find this trick of parallel uh we we find this trick of parallel uh we we find this trick of actually running the embedding actually running the embedding actually running the embedding extraction and the clustering in a extraction and the clustering in a extraction and the clustering in a totally async manner. So this does not totally async manner. So this does not totally async manner. So this does not delay the prediction whatsoever. So delay the prediction whatsoever. So delay the prediction whatsoever. So that's why I I put it uh zero that's why I I put it uh zero that's why I I put it uh zero millisecond here and then we do the the millisecond here and then we do the the millisecond here and then we do the the run trip back uh so five more run trip back uh so five more run trip back uh so five more milliseconds here. Then uh we basically milliseconds here. Then uh we basically milliseconds here. Then uh we basically uh using the state that is stored in the uh using the state that is stored in the uh using the state that is stored in the in the cloudfare we basically stitch the in the cloudfare we basically stitch the in the cloudfare we basically stitch the pieces together and ser serialize it to pieces together and ser serialize it to pieces together and ser serialize it to send it to the client. So in total you send it to the client. So in total you send it to the client. So in total you get those f four u uh 400 uh get those f four u uh 400 uh get those f four u uh 400 uh milliseconds latency milliseconds latency milliseconds latency and if we go into more details so and if we go into more details so and if we go into more details so basically there are just not just one basically there are just not just one basically there are just not just one latency basically there are many latency basically there are many latency basically there are many latencies here. There is the network latencies here. There is the network latencies here. There is the network latency that basically really depends on latency that basically really depends on latency that basically really depends on the client how close uh the client is to the client how close uh the client is to the client how close uh the client is to the networks how fast their connection the networks how fast their connection the networks how fast their connection is. So that accounts for like 50 is. So that accounts for like 50 is. So that accounts for like 50 millconds on average but this can be millconds on average but this can be millconds on average but this can be moving parts. Uh then there's the um uh moving parts. Uh then there's the um uh moving parts. Uh then there's the um uh inference latency that I mentioned the inference latency that I mentioned the inference latency that I mentioned the 50 millisecond. So same uh this 50 millisecond. So same uh this 50 millisecond. So same uh this inference latency you can tweak it a inference latency you can tweak it a inference latency you can tweak it a bit. Obviously if you make a a higher bit. Obviously if you make a a higher bit. Obviously if you make a a higher batch a larger batch it will take more batch a larger batch it will take more batch a larger batch it will take more time but then uh your cost will be lower time but then uh your cost will be lower time but then uh your cost will be lower because you can process more streams uh because you can process more streams uh because you can process more streams uh with the same u with the same GPU but if
-
with the same u with the same GPU but if with the same u with the same GPU but if you want to be faster then you have to you want to be faster then you have to you want to be faster then you have to uh lower uh the the b the batch size and uh lower uh the the b the batch size and uh lower uh the the b the batch size and uh uh your cost will go up. So really uh uh your cost will go up. So really uh uh your cost will go up. So really this inference latency this is where you this inference latency this is where you this inference latency this is where you control a cost uh latency trade-off and control a cost uh latency trade-off and control a cost uh latency trade-off and then there is this algorithmic latency then there is this algorithmic latency then there is this algorithmic latency that takes basically most of the uh uh that takes basically most of the uh uh that takes basically most of the uh uh the latency here and same uh it controls the latency here and same uh it controls the latency here and same uh it controls a new trade-off. This is a trade-off a new trade-off. This is a trade-off a new trade-off. This is a trade-off between accuracy and latency. The more between accuracy and latency. The more between accuracy and latency. The more you can uh wait in this look ahead the you can uh wait in this look ahead the you can uh wait in this look ahead the better the prediction you can make. So better the prediction you can make. So better the prediction you can make. So this is really uh like a as I say a a this is really uh like a as I say a a this is really uh like a as I say a a trade-off between the two and an easy trade-off between the two and an easy trade-off between the two and an easy way for to to to reduce the latency way for to to to reduce the latency way for to to to reduce the latency would be to reduce the block size would be to reduce the block size would be to reduce the block size obviously uh here it's 100 millisecond obviously uh here it's 100 millisecond obviously uh here it's 100 millisecond but that has repercussion on basically but that has repercussion on basically but that has repercussion on basically the inference latency because you get the inference latency because you get the inference latency because you get way more blocks that you need to process way more blocks that you need to process way more blocks that you need to process and uh so um yeah that makes the the the and uh so um yeah that makes the the the and uh so um yeah that makes the the the whole uh trade-off a bit more difficult whole uh trade-off a bit more difficult whole uh trade-off a bit more difficult to uh to achieve. to uh to achieve. to uh to achieve. Uh so um yeah now I've now reached the Uh so um yeah now I've now reached the Uh so um yeah now I've now reached the point where I think it would be good to point where I think it would be good to point where I think it would be good to make a a demo. So uh I'm going to ask my make a a demo. So uh I'm going to ask my make a a demo. So uh I'm going to ask my co-founder Vanson to to join uh stage co-founder Vanson to to join uh stage co-founder Vanson to to join uh stage just so that we have a chat quickly and just so that we have a chat quickly and just so that we have a chat quickly and you can take a mic here and uh I'll you can take a mic here and uh I'll you can take a mic here and uh I'll switch uh to a demo so that you you get switch uh to a demo so that you you get switch uh to a demo so that you you get really what 400 millconds really means really what 400 millconds really means really what 400 millconds really means in practice. Um so just about the setup in practice. Um so just about the setup in practice. Um so just about the setup uh even though we have two microphones uh even though we have two microphones uh even though we have two microphones here it's really my laptop that captures here it's really my laptop that captures here it's really my laptop that captures uh uh that captures the uh the uh uh that captures the uh the uh uh that captures the uh the conversation. So um
-
conversation. So um conversation. So um >> so you you're reing the tricks you see >> so you you're reing the tricks you see >> so you you're reing the tricks you see everything. everything. everything. >> Yeah. >> Yeah. >> Yeah. >> Hi everyone. Nice to meet you. >> Hi everyone. Nice to meet you. >> Hi everyone. Nice to meet you. >> Uh so how did my talk go? >> Uh so how did my talk go? >> Uh so how did my talk go? >> Quite good. I hope you share and people >> Quite good. I hope you share and people >> Quite good. I hope you share and people have been learning things about what we have been learning things about what we have been learning things about what we do. I'm always amazed how difficult it do. I'm always amazed how difficult it do. I'm always amazed how difficult it is. It sounds easy when you use it but is. It sounds easy when you use it but is. It sounds easy when you use it but uh uh uh >> yeah but we can even uh do overlapping >> yeah but we can even uh do overlapping >> yeah but we can even uh do overlapping speed. Absolutely. That's what we speed. Absolutely. That's what we speed. Absolutely. That's what we usually do. usually do. usually do. >> 20% of the speech usually is >> 20% of the speech usually is >> 20% of the speech usually is overlapping, which is kind of crazy to overlapping, which is kind of crazy to overlapping, which is kind of crazy to detect. detect. detect. >> Yeah, he keeps interpreting me in every >> Yeah, he keeps interpreting me in every >> Yeah, he keeps interpreting me in every meeting we have. It's an ideal effect. meeting we have. It's an ideal effect. meeting we have. It's an ideal effect. >> For some reason, Ralph says something in >> For some reason, Ralph says something in >> For some reason, Ralph says something in the in in the backside because we found the in in the backside because we found the in in the backside because we found speaker 02 here. speaker 02 here. speaker 02 here. >> Please don't back channel, please, Raul. >> Please don't back channel, please, Raul. >> Please don't back channel, please, Raul. >> Okay. Uh I leave you there. Thanks. >> Okay. Uh I leave you there. Thanks. >> Okay. Uh I leave you there. Thanks. >> H So, so you you get the idea of what uh >> H So, so you you get the idea of what uh >> H So, so you you get the idea of what uh basically 400 millconds feels like. And basically 400 millconds feels like. And basically 400 millconds feels like. And uh yeah, so I'm I I've reached uh almost uh yeah, so I'm I I've reached uh almost uh yeah, so I'm I I've reached uh almost the end of my talk. I would like I hope the end of my talk. I would like I hope the end of my talk. I would like I hope you you you learned a few things in this you you you learned a few things in this you you you learned a few things in this uh in this talk. And uh at least the uh in this talk. And uh at least the uh in this talk. And uh at least the three things that are I think three things that are I think three things that are I think interesting you'll tell me about this interesting you'll tell me about this interesting you'll tell me about this talk is that uh I I try to convince you talk is that uh I I try to convince you talk is that uh I I try to convince you that uh batch and real time derization that uh batch and real time derization that uh batch and real time derization are actually two different machine are actually two different machine are actually two different machine learning problems. They have very learning problems. They have very learning problems. They have very different constraints. Even though the different constraints. Even though the different constraints. Even though the input and the output are the same uh input and the output are the same uh input and the output are the same uh solving it uh uh needs different solving it uh uh needs different solving it uh uh needs different approaches.
-
approaches. approaches. Uh wanted to highlight as well that Uh wanted to highlight as well that Uh wanted to highlight as well that latency is just not one number. It latency is just not one number. It latency is just not one number. It really depends on many things. I mean really depends on many things. I mean really depends on many things. I mean you all all knew about that but in you all all knew about that but in you all all knew about that but in particular that for speakerization you particular that for speakerization you particular that for speakerization you see that the main latency is coming from see that the main latency is coming from see that the main latency is coming from the algorithmic part because because of the algorithmic part because because of the algorithmic part because because of the way this speakerization task is uh the way this speakerization task is uh the way this speakerization task is uh is set up and uh yeah very proud of the is set up and uh yeah very proud of the is set up and uh yeah very proud of the work that uh the research and tech team work that uh the research and tech team work that uh the research and tech team did for to to handle this and uh yeah did for to to handle this and uh yeah did for to to handle this and uh yeah there are many ways of trading uh there are many ways of trading uh there are many ways of trading uh between latency accuracy and cost and uh between latency accuracy and cost and uh between latency accuracy and cost and uh yeah so uh thank you very much that's yeah so uh thank you very much that's yeah so uh thank you very much that's the end of my talk And um yeah, happy to the end of my talk And um yeah, happy to the end of my talk And um yeah, happy to have a chat after if if you want. Thank you. Thank you so much, Alve. Thank you. Thank you so much, Alve. Let's give it up one more time for Alve, Let's give it up one more time for Alve, Let's give it up one more time for Alve, please. Nice. please. Nice. please. Nice. All right. All right. All right. Okay. So, this is a good moment because Okay. So, this is a good moment because Okay. So, this is a good moment because this is this was the last talk of this this is this was the last talk of this this is this was the last talk of this morning. Okay. Uh so, please uh you can morning. Okay. Uh so, please uh you can morning. Okay. Uh so, please uh you can go to the expo. We have a launch that is go to the expo. We have a launch that is go to the expo. We have a launch that is uh being sponsored by Sierra. So, shout uh being sponsored by Sierra. So, shout uh being sponsored by Sierra. So, shout out for to Sierra, please. Yeah, thank out for to Sierra, please. Yeah, thank out for to Sierra, please. Yeah, thank you for you for you for offering us some fuel, you know, to uh offering us some fuel, you know, to uh offering us some fuel, you know, to uh to to get going with this uh with with to to get going with this uh with with to to get going with this uh with with this event and to come back stronger in this event and to come back stronger in this event and to come back stronger in the afternoon because we still have so the afternoon because we still have so the afternoon because we still have so many things to talk about. We're going many things to talk about. We're going many things to talk about. We're going to be back here at 2 p.m. We're going to to be back here at 2 p.m. We're going to to be back here at 2 p.m. We're going to have a fireside chat between Mistrol and have a fireside chat between Mistrol and have a fireside chat between Mistrol and Nvidia. You should not miss this at all.
-
Nvidia. You should not miss this at all. Nvidia. You should not miss this at all. Uh so, I'll see you here at 2 p.m. All Uh so, I'll see you here at 2 p.m. All Uh so, I'll see you here at 2 p.m. All right, enjoy your lunch.
-
Ladies and gentlemen, Please join me in Ladies and gentlemen, Please join me in welcoming to the stage your MC for the welcoming to the stage your MC for the welcoming to the stage your MC for the AI engineer Paris 2026, AI engineer Paris 2026, AI engineer Paris 2026, developer relations engineer at Replet, developer relations engineer at Replet, developer relations engineer at Replet, Raul Chevrey. Raul Chevrey. Raul Chevrey. >> All right, and we're back. >> All right, and we're back. >> All right, and we're back. Okay. Well, congratulations to to to us Okay. Well, congratulations to to to us Okay. Well, congratulations to to to us all, right? You guys are making it for all, right? You guys are making it for all, right? You guys are making it for the final stretch of today's conference, the final stretch of today's conference, the final stretch of today's conference, day two of AI engineering Paris. Thank day two of AI engineering Paris. Thank day two of AI engineering Paris. Thank you so much for coming back here. Uh we you so much for coming back here. Uh we you so much for coming back here. Uh we actually have a great great panel that actually have a great great panel that actually have a great great panel that we're going to uh we're going to start we're going to uh we're going to start we're going to uh we're going to start in a second. But how are you feeling? in a second. But how are you feeling? in a second. But how are you feeling? >> Yeah. No, no. Okay. You you know the >> Yeah. No, no. Okay. You you know the >> Yeah. No, no. Okay. You you know the drill by now. Come on. We've been doing drill by now. Come on. We've been doing drill by now. Come on. We've been doing it since yesterday. So, we got to bring it since yesterday. So, we got to bring it since yesterday. So, we got to bring up the energy. Okay. So uh we have two up the energy. Okay. So uh we have two up the energy. Okay. So uh we have two gentlemen here that um are going to come gentlemen here that um are going to come gentlemen here that um are going to come uh on stage. So please I would like you uh on stage. So please I would like you uh on stage. So please I would like you to uh join me in welcoming to the stage to uh join me in welcoming to the stage to uh join me in welcoming to the stage uh VP of compute at Mistral Yan Leger uh VP of compute at Mistral Yan Leger uh VP of compute at Mistral Yan Leger and also uh senior manager solution and also uh senior manager solution and also uh senior manager solution architect at Nvidia Adolf Hall. So architect at Nvidia Adolf Hall. So architect at Nvidia Adolf Hall. So please let's give it up for them.
-
All right. And I kind of forgot to say All right. And I kind of forgot to say that. Um, yeah, we're going to be that. Um, yeah, we're going to be that. Um, yeah, we're going to be addressing like building the addressing like building the addressing like building the infrastructure behind the next infrastructure behind the next infrastructure behind the next generation of AI. So, we're going to generation of AI. So, we're going to generation of AI. So, we're going to have like a little bit of a futuristic have like a little bit of a futuristic have like a little bit of a futuristic view of uh what what's needed and uh to view of uh what what's needed and uh to view of uh what what's needed and uh to to build the next generation of u of AI. to build the next generation of u of AI. to build the next generation of u of AI. So, thank you so much for being here, So, thank you so much for being here, So, thank you so much for being here, Adolf uh Yan. Um Adolf uh Yan. Um Adolf uh Yan. Um Jan, you were here last year at AI Jan, you were here last year at AI Jan, you were here last year at AI engineer. Do you remember? It feels like engineer. Do you remember? It feels like engineer. Do you remember? It feels like it's been uh 10 years ago. it's been uh 10 years ago. it's been uh 10 years ago. >> Kind of a big year went by like our >> Kind of a big year went by like our >> Kind of a big year went by like our company was acquired in between. company was acquired in between. company was acquired in between. >> A lot has happened. >> A lot has happened. >> A lot has happened. >> Organizing like this conference together >> Organizing like this conference together >> Organizing like this conference together last year. Um and now we're making it last year. Um and now we're making it last year. Um and now we're making it again. again. again. >> Yeah, exactly. But it really feels like >> Yeah, exactly. But it really feels like >> Yeah, exactly. But it really feels like Yeah. One year in a in AI time. It's uh Yeah. One year in a in AI time. It's uh Yeah. One year in a in AI time. It's uh feels like ages ago. feels like ages ago. feels like ages ago. >> I wish I would have been here as well. >> I wish I would have been here as well. >> I wish I would have been here as well. >> Yeah. you invited next year. >> Yeah. you invited next year. >> Yeah. you invited next year. >> Yeah, that's nice. So, Jan, when you >> Yeah, that's nice. So, Jan, when you >> Yeah, that's nice. So, Jan, when you were here, you were here on stage, you were here, you were here on stage, you were here, you were here on stage, you were talking about building for the were talking about building for the were talking about building for the future of agentic era.
-
future of agentic era. future of agentic era. So, can you can you walk us through what So, can you can you walk us through what So, can you can you walk us through what do you think has changed in in the like do you think has changed in in the like do you think has changed in in the like in the interim? in the interim? in the interim? >> Yeah, I think like over the last two >> Yeah, I think like over the last two >> Yeah, I think like over the last two years I I've uh spoke about years I I've uh spoke about years I I've uh spoke about infrastructure for AI. I mean, and over infrastructure for AI. I mean, and over infrastructure for AI. I mean, and over actually like the more more than this actually like the more more than this actually like the more more than this now. um last year was specifically a now. um last year was specifically a now. um last year was specifically a focus on sandboxes and how basically focus on sandboxes and how basically focus on sandboxes and how basically it's creating a new category of it's creating a new category of it's creating a new category of infrastructure. Um and um we're seeing infrastructure. Um and um we're seeing infrastructure. Um and um we're seeing it in action now. I think that's what it in action now. I think that's what it in action now. I think that's what happened like the the one year happened like the the one year happened like the the one year we were in the beginning of this like we were in the beginning of this like we were in the beginning of this like motion of the agentic era and like uh motion of the agentic era and like uh motion of the agentic era and like uh the adoption kicking off uh and now um a the adoption kicking off uh and now um a the adoption kicking off uh and now um a lot of the software uh is written by lot of the software uh is written by lot of the software uh is written by agents uh and we have a huge penetration agents uh and we have a huge penetration agents uh and we have a huge penetration of the agentic workflows overall. Um the of the agentic workflows overall. Um the of the agentic workflows overall. Um the but the GPUs are still like the the but the GPUs are still like the the but the GPUs are still like the the essence uh after all like I mean like we essence uh after all like I mean like we essence uh after all like I mean like we we have in addition to the GPU we have in addition to the GPU we have in addition to the GPU infrastructure we have CPUs which play infrastructure we have CPUs which play infrastructure we have CPUs which play into the mix now. Um that's a key into the mix now. Um that's a key into the mix now. Um that's a key difference before between um the difference before between um the difference before between um the previous years and and now I think um so previous years and and now I think um so previous years and and now I think um so we're still going in the same direction we're still going in the same direction we're still going in the same direction like training large models it didn't like training large models it didn't like training large models it didn't stop uh at ML was still training like stop uh at ML was still training like stop uh at ML was still training like state-of-the-art uh models on one side state-of-the-art uh models on one side state-of-the-art uh models on one side uh we're doing large scale inferencing uh we're doing large scale inferencing uh we're doing large scale inferencing and the agent in KA is actually like and the agent in KA is actually like and the agent in KA is actually like kicking off even more the inference kicking off even more the inference kicking off even more the inference consumptions uh I think and which means
-
consumptions uh I think and which means consumptions uh I think and which means the stack is even more complex. than the stack is even more complex. than the stack is even more complex. than before. before. before. >> So you're saying yeah we were you were >> So you're saying yeah we were you were >> So you're saying yeah we were you were talking about send boxes and you've seen talking about send boxes and you've seen talking about send boxes and you've seen the rise of that type of instru the rise of that type of instru the rise of that type of instru infrastructure. So you're seeing like infrastructure. So you're seeing like infrastructure. So you're seeing like the progress in the load and I I think the progress in the load and I I think the progress in the load and I I think there's no doubt in this in this new world. What do you in this in this new world. What do you think um is standing between us and the think um is standing between us and the think um is standing between us and the next generation of AI from a hardware next generation of AI from a hardware next generation of AI from a hardware perspective and from Nvidia's perspective and from Nvidia's perspective and from Nvidia's perspective perspective perspective >> when you are longer in this business and >> when you are longer in this business and >> when you are longer in this business and then you used to speak with compute with then you used to speak with compute with then you used to speak with compute with networking and storage and actually networking and storage and actually networking and storage and actually since this agentic workload is so since this agentic workload is so since this agentic workload is so dominant well where where where is it dominant well where where where is it dominant well where where where is it going to uh be executed and um I think going to uh be executed and um I think going to uh be executed and um I think often these days in this equation it's often these days in this equation it's often these days in this equation it's forgotten where does the agenda workload forgotten where does the agenda workload forgotten where does the agenda workload go and that's a CPU question um and um go and that's a CPU question um and um go and that's a CPU question um and um that is um that is I would say pretty that is um that is I would say pretty that is um that is I would say pretty obvious if we speak in um model obvious if we speak in um model obvious if we speak in um model development so we have certain development so we have certain development so we have certain established architectures they have established architectures they have established architectures they have particular needs which drives basically particular needs which drives basically particular needs which drives basically the how the infrastructure and the the the how the infrastructure and the the the how the infrastructure and the the chips are designed for the for the next chips are designed for the for the next chips are designed for the for the next innovation cycle and uh they are innovation cycle and uh they are innovation cycle and uh they are certainly um it's influenced by mixture
-
certainly um it's influenced by mixture certainly um it's influenced by mixture of experts um and with its unique unique of experts um and with its unique unique of experts um and with its unique unique demands um of uh of communication. If we demands um of uh of communication. If we demands um of uh of communication. If we look into new modalities, I would say um look into new modalities, I would say um look into new modalities, I would say um we have I think the society and the we have I think the society and the we have I think the society and the industry has made very good use of data. industry has made very good use of data. industry has made very good use of data. We I think we we we trained on every We I think we we we trained on every We I think we we we trained on every book in the world. We likely we soon book in the world. We likely we soon book in the world. We likely we soon train on every movie and every train on every movie and every train on every movie and every multimodel. But that's not the end of multimodel. But that's not the end of multimodel. But that's not the end of the uh the game. I think there are other the uh the game. I think there are other the uh the game. I think there are other things uh um coming um which need to be things uh um coming um which need to be things uh um coming um which need to be tried out from uh um let's say if if we tried out from uh um let's say if if we tried out from uh um let's say if if we go into this robotic domain um it needs go into this robotic domain um it needs go into this robotic domain um it needs to have feedback. So giving this to have feedback. So giving this to have feedback. So giving this feedback is a question of uh of feedback is a question of uh of feedback is a question of uh of simulation. So I think that would that simulation. So I think that would that simulation. So I think that would that will pose new challenges. So we're not will pose new challenges. So we're not will pose new challenges. So we're not there with it's a very exciting uh there with it's a very exciting uh there with it's a very exciting uh journey. So I I just want to yeah double journey. So I I just want to yeah double journey. So I I just want to yeah double click on something you mentioned mixer click on something you mentioned mixer click on something you mentioned mixer of experts and it's just a way for us to of experts and it's just a way for us to of experts and it's just a way for us to have very very large models but who can have very very large models but who can have very very large models but who can be efficiently uh being used at be efficiently uh being used at be efficiently uh being used at inference time right inference time right inference time right >> yes >> yes >> yes >> okay um all right okay um >> okay um all right okay um >> okay um all right okay um if gentlemen if you'd like to follow me if gentlemen if you'd like to follow me if gentlemen if you'd like to follow me I would like to talk a little bit about I would like to talk a little bit about I would like to talk a little bit about silicon hardware silicon hardware silicon hardware um and um so yeah we know about Reuben um and um so yeah we know about Reuben um and um so yeah we know about Reuben the the latest generation of in the the latest generation of in the the latest generation of in leaderships and it's now in full leaderships and it's now in full leaderships and it's now in full production, right? Um so production, right? Um so production, right? Um so yeah, some of the claims are are like yeah, some of the claims are are like yeah, some of the claims are are like that Reuben is 10 times um more
-
that Reuben is 10 times um more that Reuben is 10 times um more efficient and it's uh yeah, four times efficient and it's uh yeah, four times efficient and it's uh yeah, four times fewer GPUs also. So as uh as AI is fewer GPUs also. So as uh as AI is fewer GPUs also. So as uh as AI is evolving like we're getting more evolving like we're getting more evolving like we're getting more efficient systems um yeah to train a efficient systems um yeah to train a efficient systems um yeah to train a model of experts. So how should teams model of experts. So how should teams model of experts. So how should teams you know how should these engineers be you know how should these engineers be you know how should these engineers be building AI products thinking about the building AI products thinking about the building AI products thinking about the next generation of architecture given next generation of architecture given next generation of architecture given this cadence and this scale this cadence and this scale this cadence and this scale >> I would say try to follow as best >> I would say try to follow as best >> I would say try to follow as best because we on a the entire industry I because we on a the entire industry I because we on a the entire industry I would say is on a on a learning cycle so would say is on a on a learning cycle so would say is on a on a learning cycle so when I when I joined Nvidia there was an when I when I joined Nvidia there was an when I when I joined Nvidia there was an architecture called uh Pascal and architecture called uh Pascal and architecture called uh Pascal and Voltaar maybe you don't know about that Voltaar maybe you don't know about that Voltaar maybe you don't know about that one but uh um if uh if we talk about one but uh um if uh if we talk about one but uh um if uh if we talk about data center economics and efficiency data center economics and efficiency data center economics and efficiency what you also stress with uh more more what you also stress with uh more more what you also stress with uh more more tokens cheaper tokens etc. Then um we tokens cheaper tokens etc. Then um we tokens cheaper tokens etc. Then um we started a technological path uh in that started a technological path uh in that started a technological path uh in that generation and that is now uh seven or generation and that is now uh seven or generation and that is now uh seven or eight years ago and that continues. So eight years ago and that continues. So eight years ago and that continues. So um we need to work with that and the um we need to work with that and the um we need to work with that and the software follows these paradigms. So software follows these paradigms. So software follows these paradigms. So software and hardware goes hand in hand software and hardware goes hand in hand software and hardware goes hand in hand and uh in order to be of use you want to and uh in order to be of use you want to and uh in order to be of use you want to use it by means of the software and uh use it by means of the software and uh use it by means of the software and uh um so there are some dependencies when um so there are some dependencies when um so there are some dependencies when you say I want to uh make use of the you say I want to uh make use of the you say I want to uh make use of the latest innovation for someone who um latest innovation for someone who um latest innovation for someone who um provides inference at scale obviously provides inference at scale obviously provides inference at scale obviously those are economic arguments and they those are economic arguments and they those are economic arguments and they need to be leveraged uh at best overall need to be leveraged uh at best overall need to be leveraged uh at best overall um we need to optimize um all layers of um we need to optimize um all layers of um we need to optimize um all layers of the of the stack. Um what we see
-
the of the stack. Um what we see the of the stack. Um what we see recently is that uh um we used to think recently is that uh um we used to think recently is that uh um we used to think in in infrastructure for trading and in in infrastructure for trading and in in infrastructure for trading and inference and I would say that is that inference and I would say that is that inference and I would say that is that is going a bit more together also for an is going a bit more together also for an is going a bit more together also for an economic reason because uh um if you if economic reason because uh um if you if economic reason because uh um if you if you make a pool it's it's static so in you make a pool it's it's static so in you make a pool it's it's static so in each pool you have some some waste each pool you have some some waste each pool you have some some waste >> and uh so that is that is I I think >> and uh so that is that is I I think >> and uh so that is that is I I think something which is uh uh go away and the something which is uh uh go away and the something which is uh uh go away and the the infrastructures for both domains um the infrastructures for both domains um the infrastructures for both domains um I expect um goes more together hand in I expect um goes more together hand in I expect um goes more together hand in hand. hand. hand. >> Interesting. Okay. >> Interesting. Okay. >> Interesting. Okay. >> Yeah, I do agree with with it all on >> Yeah, I do agree with with it all on >> Yeah, I do agree with with it all on like the the use the usage on on of the like the the use the usage on on of the like the the use the usage on on of the inference like there is actually two inference like there is actually two inference like there is actually two kind of I mean in the past if you look kind of I mean in the past if you look kind of I mean in the past if you look at some of the talks I I did like two at some of the talks I I did like two at some of the talks I I did like two years ago I was already speaking about years ago I was already speaking about years ago I was already speaking about different kind of infrastructure for different kind of infrastructure for different kind of infrastructure for different purpose. um you leverage like different purpose. um you leverage like different purpose. um you leverage like uh training clusters don't have the the uh training clusters don't have the the uh training clusters don't have the the same uh physical requirements. Uh they same uh physical requirements. Uh they same uh physical requirements. Uh they need the highest performance uh on the need the highest performance uh on the need the highest performance uh on the training side. Um and inference are training side. Um and inference are training side. Um and inference are workloads are a bit more tolerant. So workloads are a bit more tolerant. So workloads are a bit more tolerant. So they don't necessarily need the latest they don't necessarily need the latest they don't necessarily need the latest generation of hardware. Um and we see generation of hardware. Um and we see generation of hardware. Um and we see this um in practice where for training this um in practice where for training this um in practice where for training we might use the latest generation of we might use the latest generation of we might use the latest generation of hardware. For inference, we can use hardware. For inference, we can use hardware. For inference, we can use older generation of hardware. Um, which older generation of hardware. Um, which older generation of hardware. Um, which also means that like um when we design also means that like um when we design also means that like um when we design infrastructure uh on the physical side,
-
infrastructure uh on the physical side, infrastructure uh on the physical side, we need to u keep in mind that like the we need to u keep in mind that like the we need to u keep in mind that like the usage on this kind of infrastructure is usage on this kind of infrastructure is usage on this kind of infrastructure is going to evolve over the time. uh you going to evolve over the time. uh you going to evolve over the time. uh you might like deploy a cluster uh typically might like deploy a cluster uh typically might like deploy a cluster uh typically and get the two first years used for one and get the two first years used for one and get the two first years used for one purpose and the three four year years purpose and the three four year years purpose and the three four year years used for another purpose as Adolf used for another purpose as Adolf used for another purpose as Adolf mentioned like in terms of flexibility mentioned like in terms of flexibility mentioned like in terms of flexibility and ideally you can even go be more and ideally you can even go be more and ideally you can even go be more dynamic and and get your uh training dynamic and and get your uh training dynamic and and get your uh training workloads to be uh your training workloads to be uh your training workloads to be uh your training infrastructure to be reused dynamically infrastructure to be reused dynamically infrastructure to be reused dynamically for inference and so this is something for inference and so this is something for inference and so this is something which is a lot of software at the end in which is a lot of software at the end in which is a lot of software at the end in terms of orchestration to be able to terms of orchestration to be able to terms of orchestration to be able to reallocate dynamically the walkers. reallocate dynamically the walkers. reallocate dynamically the walkers. >> Okay, very interesting. So, we'll need >> Okay, very interesting. So, we'll need >> Okay, very interesting. So, we'll need eventually two different types of eventually two different types of eventually two different types of infrastructure for um for inference and infrastructure for um for inference and infrastructure for um for inference and for training down the line if I for training down the line if I for training down the line if I understand it correctly because the understand it correctly because the understand it correctly because the needs are different because we don't needs are different because we don't needs are different because we don't need necessarily the latest type of need necessarily the latest type of need necessarily the latest type of hardware. So we I mean we actually were hardware. So we I mean we actually were hardware. So we I mean we actually were more converging maybe I mean we have more converging maybe I mean we have more converging maybe I mean we have specialized inference in infrastructure specialized inference in infrastructure specialized inference in infrastructure on one hand which uh if you build it on one hand which uh if you build it on one hand which uh if you build it like just for inference uh you can like like just for inference uh you can like like just for inference uh you can like tolerate like having not the latest tolerate like having not the latest tolerate like having not the latest generation of hardware it also depends generation of hardware it also depends generation of hardware it also depends on the kind of models you're training on the kind of models you're training on the kind of models you're training like models now are as we know are old like models now are as we know are old like models now are as we know are old of um different sizes and shapes. If you of um different sizes and shapes. If you of um different sizes and shapes. If you uh train a 1 trillion parameter model, uh train a 1 trillion parameter model, uh train a 1 trillion parameter model, it's not the same as like an 8 billion it's not the same as like an 8 billion it's not the same as like an 8 billion billion parameter model, right? You
-
billion parameter model, right? You billion parameter model, right? You don't need the same training don't need the same training don't need the same training infrastructure in both cases. Um and infrastructure in both cases. Um and infrastructure in both cases. Um and actually we're more I think like actually we're more I think like actually we're more I think like converging on the long term in a way. I converging on the long term in a way. I converging on the long term in a way. I mean on having a global pool that we mean on having a global pool that we mean on having a global pool that we reallocate uh between the kind of reallocate uh between the kind of reallocate uh between the kind of trainings uh and inferencing. trainings uh and inferencing. trainings uh and inferencing. >> Okay. Interesting. Um okay. So the next >> Okay. Interesting. Um okay. So the next >> Okay. Interesting. Um okay. So the next question is to for you Adolf. Um, I'm question is to for you Adolf. Um, I'm question is to for you Adolf. Um, I'm going to talk about uh NVL72, but before going to talk about uh NVL72, but before going to talk about uh NVL72, but before I I dive into that, can you please I I dive into that, can you please I I dive into that, can you please explain what NVL72 is to the audience explain what NVL72 is to the audience explain what NVL72 is to the audience maybe who are not familiar with? maybe who are not familiar with? maybe who are not familiar with? >> In case you're not familiar, um, AI is >> In case you're not familiar, um, AI is >> In case you're not familiar, um, AI is in the meantime, when let's put it this in the meantime, when let's put it this in the meantime, when let's put it this way, when I started, my entry cost to AI way, when I started, my entry cost to AI way, when I started, my entry cost to AI was, I think, $150 was, I think, $150 was, I think, $150 with a GPU, I had the workbench, I could with a GPU, I had the workbench, I could with a GPU, I had the workbench, I could participate. Now, everything grew. Um, participate. Now, everything grew. Um, participate. Now, everything grew. Um, fascinating speed. So uh I would say fascinating speed. So uh I would say fascinating speed. So uh I would say that the smallest problem what Mistral that the smallest problem what Mistral that the smallest problem what Mistral has I think it's it's it's not going to has I think it's it's it's not going to has I think it's it's it's not going to fit on a single server anymore and uh um fit on a single server anymore and uh um fit on a single server anymore and uh um so at the same time uh the hardware pays so at the same time uh the hardware pays so at the same time uh the hardware pays attention to what is requested by the attention to what is requested by the attention to what is requested by the software and so uh multiple components software and so uh multiple components software and so uh multiple components need to work together in a most um need to work together in a most um need to work together in a most um economic and efficient way. Efficient economic and efficient way. Efficient economic and efficient way. Efficient means communication needs to be very means communication needs to be very means communication needs to be very very fast. And for that purpose um we very fast. And for that purpose um we very fast. And for that purpose um we designed NVL72 which is a a system in designed NVL72 which is a a system in designed NVL72 which is a a system in Iraq. And u um the the difference to a Iraq. And u um the the difference to a Iraq. And u um the the difference to a classic system is um that NVL72 is an classic system is um that NVL72 is an classic system is um that NVL72 is an interconnect between them much closer interconnect between them much closer interconnect between them much closer than than anything of TCP IP. And this
-
than than anything of TCP IP. And this than than anything of TCP IP. And this is uh necessary if you want to run the is uh necessary if you want to run the is uh necessary if you want to run the modern workloads. Um well and uh um I modern workloads. Um well and uh um I modern workloads. Um well and uh um I would say it's the most economic way to would say it's the most economic way to would say it's the most economic way to address uh highly parallel uh trainings address uh highly parallel uh trainings address uh highly parallel uh trainings at this point. at this point. at this point. >> Okay. So yeah. So uh Nv72 is is a rack >> Okay. So yeah. So uh Nv72 is is a rack >> Okay. So yeah. So uh Nv72 is is a rack system and um system and um system and um so yeah, what does that mean for the so yeah, what does that mean for the so yeah, what does that mean for the rest of the stack? Does it mean that uh rest of the stack? Does it mean that uh rest of the stack? Does it mean that uh yeah, we're headed towards the rack yeah, we're headed towards the rack yeah, we're headed towards the rack being like the new server? being like the new server? being like the new server? >> Well, there's a saying from Nvidia that >> Well, there's a saying from Nvidia that >> Well, there's a saying from Nvidia that the data center is the computer. So but the data center is the computer. So but the data center is the computer. So but uh thinking even bigger um uh but uh uh thinking even bigger um uh but uh uh thinking even bigger um uh but uh when you think about the workload you when you think about the workload you when you think about the workload you have to schedule various uh jobs on the have to schedule various uh jobs on the have to schedule various uh jobs on the units then um what is the most economic units then um what is the most economic units then um what is the most economic way of doing it and the NVL72 in itself way of doing it and the NVL72 in itself way of doing it and the NVL72 in itself is a unit of assignment and um so u is a unit of assignment and um so u is a unit of assignment and um so u economically there's no need to say I'm economically there's no need to say I'm economically there's no need to say I'm cut it in the in the in the middle you cut it in the in the in the middle you cut it in the in the in the middle you can do that um but you're definitely can do that um but you're definitely can do that um but you're definitely leaving the sweet spot um of that uh of leaving the sweet spot um of that uh of leaving the sweet spot um of that uh of that system so in terms of I I would uh that system so in terms of I I would uh that system so in terms of I I would uh assign your your um your summary that it assign your your um your summary that it assign your your um your summary that it the computer lifts basically from a the computer lifts basically from a the computer lifts basically from a single system to a rack scale.
-
single system to a rack scale. single system to a rack scale. >> Yes, it's it's going to be just uh like >> Yes, it's it's going to be just uh like >> Yes, it's it's going to be just uh like yeah just a bigger systems to serve our yeah just a bigger systems to serve our yeah just a bigger systems to serve our needs moving forward. Um, okay. So, needs moving forward. Um, okay. So, needs moving forward. Um, okay. So, yeah, I have a question for you, Yan. yeah, I have a question for you, Yan. yeah, I have a question for you, Yan. Uh, since I mean like Uh, since I mean like Uh, since I mean like you're you're at the compute level, so you're you're at the compute level, so you're you're at the compute level, so you serve a lot of uh inference. And you serve a lot of uh inference. And you serve a lot of uh inference. And what do you think um that you will need what do you think um that you will need what do you think um that you will need moving forward from maybe from the chips moving forward from maybe from the chips moving forward from maybe from the chips or from a hardware perspective in order or from a hardware perspective in order or from a hardware perspective in order to maybe better serve your customers in to maybe better serve your customers in to maybe better serve your customers in let's say in the near let's say in the near let's say in the near yeah in the near future? yeah in the near future? yeah in the near future? It's uh I think it's well it's a it's a It's uh I think it's well it's a it's a It's uh I think it's well it's a it's a tricky question like uh we I mean we tricky question like uh we I mean we tricky question like uh we I mean we have a way of consuming hardware which have a way of consuming hardware which have a way of consuming hardware which is like we take the latest generation of is like we take the latest generation of is like we take the latest generation of hardware for training frontier models. hardware for training frontier models. hardware for training frontier models. So which is like that's uh what we need So which is like that's uh what we need So which is like that's uh what we need in general they they provide significant in general they they provide significant in general they they provide significant gain of efficiency and like to be at the gain of efficiency and like to be at the gain of efficiency and like to be at the frontier you actually still need it's frontier you actually still need it's frontier you actually still need it's still like uh you still want more power.
-
still like uh you still want more power. still like uh you still want more power. Um so in general we have this this mix Um so in general we have this this mix Um so in general we have this this mix and and we don't only deploy NVL72 and and we don't only deploy NVL72 and and we don't only deploy NVL72 systems actually we do have a mix of systems actually we do have a mix of systems actually we do have a mix of both um system which is which with this both um system which is which with this both um system which is which with this high performance interconnect for high performance interconnect for high performance interconnect for training uh but we do also some training uh but we do also some training uh but we do also some deployments dedicated to inference and deployments dedicated to inference and deployments dedicated to inference and they might not be equipped uh with they might not be equipped uh with they might not be equipped uh with exactly the same system. Um so I think exactly the same system. Um so I think exactly the same system. Um so I think like we have um the the reality of the like we have um the the reality of the like we have um the the reality of the field is like our the diversity of infra field is like our the diversity of infra field is like our the diversity of infra infrastructure is increasing because um infrastructure is increasing because um infrastructure is increasing because um the AI wave has started like several the AI wave has started like several the AI wave has started like several years ago. So we are starting to have years ago. So we are starting to have years ago. So we are starting to have like different generation and getting to like different generation and getting to like different generation and getting to actually a cycle of infrastructure which actually a cycle of infrastructure which actually a cycle of infrastructure which might actually resemble CPU in a way. Um might actually resemble CPU in a way. Um might actually resemble CPU in a way. Um the designs are completely different. uh the designs are completely different. uh the designs are completely different. uh but you do have a long-term like but you do have a long-term like but you do have a long-term like rotation in terms of hardware and you rotation in terms of hardware and you rotation in terms of hardware and you need to deal with this heterogeneous need to deal with this heterogeneous need to deal with this heterogeneous hardware. So um yeah I think in general hardware. So um yeah I think in general hardware. So um yeah I think in general like u um the chip manufacturers and like u um the chip manufacturers and like u um the chip manufacturers and Nvidia are designing what is like coming Nvidia are designing what is like coming Nvidia are designing what is like coming next and uh we're going to leverage this next and uh we're going to leverage this next and uh we're going to leverage this to to to train even more performant to to to train even more performant to to to train even more performant models.
-
models. models. >> Okay. Yeah. Um okay I think uh I want to >> Okay. Yeah. Um okay I think uh I want to >> Okay. Yeah. Um okay I think uh I want to move on since we're talking about move on since we're talking about move on since we're talking about inference I would like to talk a little inference I would like to talk a little inference I would like to talk a little bit about inference economics if uh if bit about inference economics if uh if bit about inference economics if uh if that's okay with you guys. Okay. So, um, that's okay with you guys. Okay. So, um, that's okay with you guys. Okay. So, um, uh, yes. So, when Nvidia basically talks uh, yes. So, when Nvidia basically talks uh, yes. So, when Nvidia basically talks about, you know, AI factories, you about, you know, AI factories, you about, you know, AI factories, you always measure with tokens per second, always measure with tokens per second, always measure with tokens per second, per dollar, per watt, right? Like we per dollar, per watt, right? Like we per dollar, per watt, right? Like we want the fastest way to serve inference. want the fastest way to serve inference. want the fastest way to serve inference. We want it at the cheapest cost and the We want it at the cheapest cost and the We want it at the cheapest cost and the lowest energy possible, right? Um, but lowest energy possible, right? Um, but lowest energy possible, right? Um, but yeah, you mentioned it's different yeah, you mentioned it's different yeah, you mentioned it's different infrastructure for uh, yeah, for infrastructure for uh, yeah, for infrastructure for uh, yeah, for training and inference. But what do you training and inference. But what do you training and inference. But what do you think think think at what point we're going to be able to at what point we're going to be able to at what point we're going to be able to justify the the the revenue of inference justify the the the revenue of inference justify the the the revenue of inference rather than just saying like um so we're rather than just saying like um so we're rather than just saying like um so we're building basically for for training building basically for for training building basically for for training mostly. But when when are we going to uh mostly. But when when are we going to uh mostly. But when when are we going to uh to be able to justify the the revenue to be able to justify the the revenue to be able to justify the the revenue for inference and say that yeah all that for inference and say that yeah all that for inference and say that yeah all that build out is not going to build out is not going to build out is not going to is going to materialize at some point.
-
is going to materialize at some point. is going to materialize at some point. looking looking looking >> you're looking at me. >> Yes, I'm looking at you. Yes. Sorry, it >> Yes, I'm looking at you. Yes. Sorry, it was a question for if uh nobody was a question for if uh nobody was a question for if uh nobody understood, but yes, it's an Nvidia understood, but yes, it's an Nvidia understood, but yes, it's an Nvidia related question and related question and related question and >> well, everything what you what you what >> well, everything what you what you what >> well, everything what you what you what you use for for producing and making you use for for producing and making you use for for producing and making money that there needs to be some some money that there needs to be some some money that there needs to be some some it the economics will not change for it the economics will not change for it the economics will not change for this one. So um it might have been uh a this one. So um it might have been uh a this one. So um it might have been uh a pattern in the past that you say I'm pattern in the past that you say I'm pattern in the past that you say I'm using training uh systems and uh when using training uh systems and uh when using training uh systems and uh when once the next generation is available once the next generation is available once the next generation is available I'm using that one for for inference for I'm using that one for for inference for I'm using that one for for inference for a less demanding workload. Um but still a less demanding workload. Um but still a less demanding workload. Um but still um the economics must be taken into um the economics must be taken into um the economics must be taken into account. Um and uh um now um I would say account. Um and uh um now um I would say account. Um and uh um now um I would say the inferencing workloads with the um the inferencing workloads with the um the inferencing workloads with the um it's going through the roof. Um and um it's going through the roof. Um and um it's going through the roof. Um and um there is there is much more demand. I there is there is much more demand. I there is there is much more demand. I myself I'm using it. I love to use it. I myself I'm using it. I love to use it. I myself I'm using it. I love to use it. I use it all the time. It it I don't I use it all the time. It it I don't I use it all the time. It it I don't I don't search in a classic way anymore don't search in a classic way anymore don't search in a classic way anymore likely many of you don't do this.
-
likely many of you don't do this. likely many of you don't do this. >> Um and I don't want to miss it. Uh and >> Um and I don't want to miss it. Uh and >> Um and I don't want to miss it. Uh and so I'm leaving a trace um of this one so I'm leaving a trace um of this one so I'm leaving a trace um of this one and uh um so inferencing is uh um is and uh um so inferencing is uh um is and uh um so inferencing is uh um is also very important uh for us because we also very important uh for us because we also very important uh for us because we spend a lot of uh engineering cycles in spend a lot of uh engineering cycles in spend a lot of uh engineering cycles in making the software better all the time. making the software better all the time. making the software better all the time. Um so it has a um it is really at a top Um so it has a um it is really at a top Um so it has a um it is really at a top workload and so I would say um these workload and so I would say um these workload and so I would say um these days um you need to take it into account days um you need to take it into account days um you need to take it into account um right from the first day. um right from the first day. um right from the first day. >> Okay. >> Okay. >> Okay. >> Yeah. And so if you go in this root like >> Yeah. And so if you go in this root like >> Yeah. And so if you go in this root like I think there was a risk which was I think there was a risk which was I think there was a risk which was before which was um a fear because the before which was um a fear because the before which was um a fear because the inference side didn't materialize at inference side didn't materialize at inference side didn't materialize at scale and I think in the last year with scale and I think in the last year with scale and I think in the last year with the emergence of actually actually the emergence of actually actually the emergence of actually actually agentic workflows agentic workflows agentic workflows >> uh and workloads um we see inference >> uh and workloads um we see inference >> uh and workloads um we see inference taking off like I think like one one of taking off like I think like one one of taking off like I think like one one of the the thing which is discussed is like the the thing which is discussed is like the the thing which is discussed is like how much like for instance if you look how much like for instance if you look how much like for instance if you look at software engineering how much do you at software engineering how much do you at software engineering how much do you spend on like AI tools to accelerate spend on like AI tools to accelerate spend on like AI tools to accelerate your software engineering efforts your software engineering efforts your software engineering efforts >> and so obviously there are discussion >> and so obviously there are discussion >> and so obviously there are discussion around token maxing or whatever like to around token maxing or whatever like to around token maxing or whatever like to to on this angle but um there is also to on this angle but um there is also to on this angle but um there is also just a reality that it's a daily tool uh just a reality that it's a daily tool uh just a reality that it's a daily tool uh for anybody who is doing software for anybody who is doing software for anybody who is doing software engineering these days so um and we see engineering these days so um and we see engineering these days so um and we see the the revenue uh generation like uh the the revenue uh generation like uh the the revenue uh generation like uh from the the result of the previous from the the result of the previous from the the result of the previous investments are paying off now investments are paying off now investments are paying off now >> yeah so basically the inference is is
-
>> yeah so basically the inference is is >> yeah so basically the inference is is taking off and we're going to see it's taking off and we're going to see it's taking off and we're going to see it's just going to be uh increasing moving just going to be uh increasing moving just going to be uh increasing moving forward. So all this buildout is forward. So all this buildout is forward. So all this buildout is actually going to be used at some point actually going to be used at some point actually going to be used at some point by inference even though if it's for for by inference even though if it's for for by inference even though if it's for for training at any time but uh yeah you training at any time but uh yeah you training at any time but uh yeah you spoke about token maxing uh and um yeah spoke about token maxing uh and um yeah spoke about token maxing uh and um yeah I reflect like we we don't believe so I reflect like we we don't believe so I reflect like we we don't believe so much at uh at token maxing but we much at uh at token maxing but we much at uh at token maxing but we believe more um about um outcome maxing. believe more um about um outcome maxing. believe more um about um outcome maxing. So we want to provide value for the So we want to provide value for the So we want to provide value for the customer and I think this is a good customer and I think this is a good customer and I think this is a good segue for my next question actually segue for my next question actually segue for my next question actually because yeah despite or because yeah despite or because yeah despite or when you quantify you know the the uh when you quantify you know the the uh when you quantify you know the the uh like we were talking about cost etc but like we were talking about cost etc but like we were talking about cost etc but at some point it's not just a matter of at some point it's not just a matter of at some point it's not just a matter of cost it's a matter of value to to the cost it's a matter of value to to the cost it's a matter of value to to the people who use it right so h how do you people who use it right so h how do you people who use it right so h how do you guys actually quantify that so I mean guys actually quantify that so I mean guys actually quantify that so I mean beyond just beyond just beyond just raw bruter commodities ities and you raw bruter commodities ities and you raw bruter commodities ities and you know economics in this case. So how do know economics in this case. So how do know economics in this case. So how do we quantify intelligence? Uh >> so is a question for me or is it for >> so is a question for me or is it for >> this is a question for you?
-
>> this is a question for you? >> this is a question for you? >> How do we intelligence? I think it's a >> How do we intelligence? I think it's a >> How do we intelligence? I think it's a tough one like how do you quantify tough one like how do you quantify tough one like how do you quantify intelligence is more about yeah it's intelligence is more about yeah it's intelligence is more about yeah it's about business outcome. So it's uh and about business outcome. So it's uh and about business outcome. So it's uh and it's highly viable. you can have like um it's highly viable. you can have like um it's highly viable. you can have like um with a low amount of token you can with a low amount of token you can with a low amount of token you can generate tremendous amount of of value generate tremendous amount of of value generate tremendous amount of of value if you apply AI right I think I mean if you apply AI right I think I mean if you apply AI right I think I mean it's one of the thing that we are doing it's one of the thing that we are doing it's one of the thing that we are doing at Mistro it's not only about volume at Mistro it's not only about volume at Mistro it's not only about volume there are some activities which are um there are some activities which are um there are some activities which are um consuming a lot of tokens uh consuming a lot of tokens uh consuming a lot of tokens uh proportionally and there is a question proportionally and there is a question proportionally and there is a question there is all the efficiency and all the there is all the efficiency and all the there is all the efficiency and all the actually knowledge we still knew to to actually knowledge we still knew to to actually knowledge we still knew to to apply to optimize this processes because apply to optimize this processes because apply to optimize this processes because some of the workloads might be actually some of the workloads might be actually some of the workloads might be actually like using an abnormal amount of tokens like using an abnormal amount of tokens like using an abnormal amount of tokens uh for the outcome. Um and so I think uh for the outcome. Um and so I think uh for the outcome. Um and so I think that's where um all of the ecosystem can that's where um all of the ecosystem can that's where um all of the ecosystem can deliver a lot of value. If we look at deliver a lot of value. If we look at deliver a lot of value. If we look at like the the situation and like it's not like the the situation and like it's not like the the situation and like it's not only about the M it's not only about only about the M it's not only about only about the M it's not only about training it's about like the complete training it's about like the complete training it's about like the complete harness and how do you actually like harness and how do you actually like harness and how do you actually like leverage the tokens properly uh to have leverage the tokens properly uh to have leverage the tokens properly uh to have a ratio of like token cost versus uh a ratio of like token cost versus uh a ratio of like token cost versus uh value. Then on the infrastructure layer, value. Then on the infrastructure layer, value. Then on the infrastructure layer, we're just optimizing each layer. Uh but we're just optimizing each layer. Uh but we're just optimizing each layer. Uh but we are just providing like the tokens.
-
we are just providing like the tokens. we are just providing like the tokens. Uh the value is really in the uh it's Uh the value is really in the uh it's Uh the value is really in the uh it's captured by the higher levels of the captured by the higher levels of the captured by the higher levels of the stack I think. stack I think. stack I think. >> Okay. Um well since I think since you >> Okay. Um well since I think since you >> Okay. Um well since I think since you are talking about you know uh the are talking about you know uh the are talking about you know uh the consumption of tokens and the ever consumption of tokens and the ever consumption of tokens and the ever rising the the consumption of tokens and rising the the consumption of tokens and rising the the consumption of tokens and harnesses etc. So I think like um it's a harnesses etc. So I think like um it's a harnesses etc. So I think like um it's a good segue for me for my next question good segue for me for my next question good segue for me for my next question for Adolf about reasoning. So um yeah, for Adolf about reasoning. So um yeah, for Adolf about reasoning. So um yeah, one way to increase the number of tokens one way to increase the number of tokens one way to increase the number of tokens that we consume is reasoning, but it's that we consume is reasoning, but it's that we consume is reasoning, but it's also a good way also to to have a better also a good way also to to have a better also a good way also to to have a better uh a better outcome, right? So which uh a better outcome, right? So which uh a better outcome, right? So which part you think of the stack that is part you think of the stack that is part you think of the stack that is maybe the least prepared uh for all that maybe the least prepared uh for all that maybe the least prepared uh for all that uh you know tokens and that is generated uh you know tokens and that is generated uh you know tokens and that is generated by by reasoning agents. I would say I love it. It comes agents. I would say I love it. It comes at a price. price is um we we consume at a price. price is um we we consume at a price. price is um we we consume more tokens. Um that means for everyone more tokens. Um that means for everyone more tokens. Um that means for everyone who provides infrastructure um I need who provides infrastructure um I need who provides infrastructure um I need the uh the memory to keep uh the the the uh the memory to keep uh the the the uh the memory to keep uh the the long contexts uh and I I need uh the long contexts uh and I I need uh the long contexts uh and I I need uh the caching um to take the best out of uh caching um to take the best out of uh caching um to take the best out of uh reuse of these conversations and uh um reuse of these conversations and uh um reuse of these conversations and uh um so so so I could imagine that uh um we haven't uh I could imagine that uh um we haven't uh I could imagine that uh um we haven't uh um deployed the latest uh or the out of um deployed the latest uh or the out of um deployed the latest uh or the out of possible everywhere. Um I see we are possible everywhere. Um I see we are possible everywhere. Um I see we are basically in in a phase where uh where basically in in a phase where uh where basically in in a phase where uh where caching um gets gets more important caching um gets gets more important caching um gets gets more important caching across um um in in a distributed
-
caching across um um in in a distributed caching across um um in in a distributed systems um there are clear uh clear systems um there are clear uh clear systems um there are clear uh clear signs that uh it goes into u um signs that uh it goes into u um signs that uh it goes into u um distributed serving um and this caching distributed serving um and this caching distributed serving um and this caching paradigm is important to this one and um paradigm is important to this one and um paradigm is important to this one and um well I think uh um we if we sit here well I think uh um we if we sit here well I think uh um we if we sit here next year I think uh we will we will next year I think uh we will we will next year I think uh we will we will review um the value of caching in that review um the value of caching in that review um the value of caching in that equation. equation. equation. >> Ah, super interesting. Okay. Um, I have >> Ah, super interesting. Okay. Um, I have >> Ah, super interesting. Okay. Um, I have one final question for the two of you. one final question for the two of you. one final question for the two of you. Okay. Uh, we said, yeah, we're coming Okay. Uh, we said, yeah, we're coming Okay. Uh, we said, yeah, we're coming back next year. You're invited according back next year. You're invited according back next year. You're invited according to to Yan, but we'll probably be come to to Yan, but we'll probably be come to to Yan, but we'll probably be come back coming back hopefully in five years back coming back hopefully in five years back coming back hopefully in five years from now. And, uh, um, so yeah. Do you from now. And, uh, um, so yeah. Do you from now. And, uh, um, so yeah. Do you think that we're going to have any think that we're going to have any think that we're going to have any frontier models going that are going to frontier models going that are going to frontier models going that are going to be trained in uh, European AI factory? be trained in uh, European AI factory? be trained in uh, European AI factory? This is you. You can start with this This is you. You can start with this This is you. You can start with this one. uh Yan and uh one. uh Yan and uh one. uh Yan and uh >> yeah yeah I think I mean we're acting on >> yeah yeah I think I mean we're acting on >> yeah yeah I think I mean we're acting on it at Mistro like we're bringing up like it at Mistro like we're bringing up like it at Mistro like we're bringing up like 200 megawatt of capacity next year and 1 200 megawatt of capacity next year and 1 200 megawatt of capacity next year and 1 gawatt uh until 2030 so and um um and a gawatt uh until 2030 so and um um and a gawatt uh until 2030 so and um um and a lot of it is in Europe so um and we are lot of it is in Europe so um and we are lot of it is in Europe so um and we are training on this infrastructure and training on this infrastructure and training on this infrastructure and actually like the current uh MR models actually like the current uh MR models actually like the current uh MR models are already trained on our are already trained on our are already trained on our infrastructure so I think in five years infrastructure so I think in five years infrastructure so I think in five years for sure we'll have frontier models for sure we'll have frontier models for sure we'll have frontier models Uh, Uh, Uh, >> we're announcing already.
-
>> we're announcing already. >> we're announcing already. >> Okay, great. Okay, so what's the most >> Okay, great. Okay, so what's the most >> Okay, great. Okay, so what's the most overrated infrastructure trend in 2026? overrated infrastructure trend in 2026? overrated infrastructure trend in 2026? >> That's my love. >> That's my love. >> That's my love. That's for you. That's for you as well. Out of the blue, self-healing Out of the blue, self-healing everything. What? everything. What? everything. What? Okay. And what's the most underrated Okay. And what's the most underrated Okay. And what's the most underrated one? I think like uh in in infrastructure I think like uh in in infrastructure like it depends on which uh which uh like it depends on which uh which uh like it depends on which uh which uh layer we're we're looking at. It's a layer we're we're looking at. It's a layer we're we're looking at. It's a a tricky question but like software a tricky question but like software a tricky question but like software uh in a way. Um yeah I mean if you look uh in a way. Um yeah I mean if you look uh in a way. Um yeah I mean if you look at the purely infrastructure level uh at the purely infrastructure level uh at the purely infrastructure level uh infrastructure is software is probably infrastructure is software is probably infrastructure is software is probably like the the one of the underrated one like the the one of the underrated one like the the one of the underrated one uh software enabling infrastructure.
-
uh software enabling infrastructure. uh software enabling infrastructure. >> All right that's a good answer. I'm >> All right that's a good answer. I'm >> All right that's a good answer. I'm happy with that. So with this thank you happy with that. So with this thank you happy with that. So with this thank you so much gentlemen. Uh this is it for for so much gentlemen. Uh this is it for for so much gentlemen. Uh this is it for for for this chat and uh yeah thanks Adolf for this chat and uh yeah thanks Adolf for this chat and uh yeah thanks Adolf and thanks Yan. Please let's give it up and thanks Yan. Please let's give it up and thanks Yan. Please let's give it up for Adolf and Yan. Oops. for Adolf and Yan. Oops. for Adolf and Yan. Oops. All right. All right. All right. >> Well, the audience is part of the five >> Well, the audience is part of the five >> Well, the audience is part of the five years so mortal exercise, right? years so mortal exercise, right? years so mortal exercise, right? >> Absolutely. Yes. We're all going to be >> Absolutely. Yes. We're all going to be >> Absolutely. Yes. We're all going to be coming back in five years. Thank you so coming back in five years. Thank you so coming back in five years. Thank you so much. much. much. >> All right. Let's give it up for Yan and >> All right. Let's give it up for Yan and >> All right. Let's give it up for Yan and head off once more, please. Okay. And head off once more, please. Okay. And head off once more, please. Okay. And with that, we're ready for our next with that, we're ready for our next with that, we're ready for our next speaker. Our next speaker comes from speaker. Our next speaker comes from speaker. Our next speaker comes from Moto. Please, ladies and gentlemen, Moto. Please, ladies and gentlemen, Moto. Please, ladies and gentlemen, please welcome to the stage please welcome to the stage please welcome to the stage Charles Fry. All right. Charles Fry. All right. Charles Fry. All right. How's it going?
-
All right. Um, that was great intro All right. Um, that was great intro music. I'm very pumped. I'm ready to do music. I'm very pumped. I'm ready to do music. I'm very pumped. I'm ready to do like a jazzer size or something. Um, and like a jazzer size or something. Um, and like a jazzer size or something. Um, and you need a lot of energy for low latency you need a lot of energy for low latency you need a lot of energy for low latency inference. Um, not just wattage, but inference. Um, not just wattage, but inference. Um, not just wattage, but also uh grit. So, it's an appropriate also uh grit. So, it's an appropriate also uh grit. So, it's an appropriate start. Um, so I'm Charles. I work on start. Um, so I'm Charles. I work on start. Um, so I'm Charles. I work on inference engineering at Modal. We're a inference engineering at Modal. We're a inference engineering at Modal. We're a infrastructure platform. Um, and what infrastructure platform. Um, and what infrastructure platform. Um, and what I'm gonna tell you about today is our I'm gonna tell you about today is our I'm gonna tell you about today is our playbook that we've developed for playbook that we've developed for playbook that we've developed for optimizing uh low latency LLM inference. optimizing uh low latency LLM inference. optimizing uh low latency LLM inference. Um, so but before telling you how to Um, so but before telling you how to Um, so but before telling you how to make it good, uh, I want to talk about make it good, uh, I want to talk about make it good, uh, I want to talk about like what low latency inf low latency like what low latency inf low latency like what low latency inf low latency inference is and why it's so important. inference is and why it's so important. inference is and why it's so important. Um, so there's this thing that got Um, so there's this thing that got Um, so there's this thing that got released recently called Jev. Um, that released recently called Jev. Um, that released recently called Jev. Um, that probably hasn't been brought up every probably hasn't been brought up every probably hasn't been brought up every five minutes uh, today. Um, so, uh, Jev five minutes uh, today. Um, so, uh, Jev five minutes uh, today. Um, so, uh, Jev is a decision model, a a generic zeroot is a decision model, a a generic zeroot is a decision model, a a generic zeroot classifier that works really fast. Um, classifier that works really fast. Um, classifier that works really fast. Um, and people have gotten very excited and people have gotten very excited and people have gotten very excited about it. People have built these really about it. People have built these really about it. People have built these really cool demos. One of my favorite from cool demos. One of my favorite from cool demos. One of my favorite from Twitter is you like type a word and you Twitter is you like type a word and you Twitter is you like type a word and you get a pallet associated with it. Um, and get a pallet associated with it. Um, and get a pallet associated with it. Um, and the um, the person who created this the um, the person who created this the um, the person who created this demo, Matt Deloreier, said, "Oh, this is demo, Matt Deloreier, said, "Oh, this is demo, Matt Deloreier, said, "Oh, this is very cheap and fast. It feels like a very cheap and fast. It feels like a very cheap and fast. It feels like a leap forward for creating new UX and UI leap forward for creating new UX and UI leap forward for creating new UX and UI paradigms." Um, and the response, I paradigms." Um, and the response, I paradigms." Um, and the response, I think, for a lot of people who've been think, for a lot of people who've been think, for a lot of people who've been in this field in machine learning and in this field in machine learning and in this field in machine learning and natural language processing for a long natural language processing for a long natural language processing for a long time was actually kind of different. It time was actually kind of different. It time was actually kind of different. It was like, "Oh, I'm surprised that was like, "Oh, I'm surprised that was like, "Oh, I'm surprised that everyone's so excited about this." Um,
-
everyone's so excited about this." Um, everyone's so excited about this." Um, and the sort of like basic answer here and the sort of like basic answer here and the sort of like basic answer here is that they made this intelligence that is that they made this intelligence that is that they made this intelligence that was always there cheaper and faster at was always there cheaper and faster at was always there cheaper and faster at the same time. and all of a sudden that the same time. and all of a sudden that the same time. and all of a sudden that unlocked new use cases if you aren't unlocked new use cases if you aren't unlocked new use cases if you aren't sweating that like you might run up to a sweating that like you might run up to a sweating that like you might run up to a usage limit or you might like spend usage limit or you might like spend usage limit or you might like spend $1,000 just trying to make an icon spin $1,000 just trying to make an icon spin $1,000 just trying to make an icon spin um and you won't wait 30 minutes for um and you won't wait 30 minutes for um and you won't wait 30 minutes for that icon to start spinning then like that icon to start spinning then like that icon to start spinning then like all of a sudden there are more all of a sudden there are more all of a sudden there are more opportunities available to you um so opportunities available to you um so opportunities available to you um so um oh I forgot my corporate sponsorship um oh I forgot my corporate sponsorship um oh I forgot my corporate sponsorship statement yeah this is uh application uh statement yeah this is uh application uh statement yeah this is uh application uh Jeff and Typesafe building on top of our Jeff and Typesafe building on top of our Jeff and Typesafe building on top of our cloud platform So this is the kind of cloud platform So this is the kind of cloud platform So this is the kind of thing u that I'd like to help you figure thing u that I'd like to help you figure thing u that I'd like to help you figure out how to build yourself. Um and so Jev out how to build yourself. Um and so Jev out how to build yourself. Um and so Jev sort of sits in like one of three basic sort of sits in like one of three basic sort of sits in like one of three basic categories for the kinds of things that categories for the kinds of things that categories for the kinds of things that people build with uh with language model people build with uh with language model people build with uh with language model inference with sequence inference. Um inference with sequence inference. Um inference with sequence inference. Um and my sort of like infrared take and my sort of like infrared take and my sort of like infrared take thinking on the back end here um is that thinking on the back end here um is that thinking on the back end here um is that you kind of want to divide these out you kind of want to divide these out you kind of want to divide these out into what is the kind of like figure of into what is the kind of like figure of into what is the kind of like figure of merit for their performance. The merit for their performance. The merit for their performance. The performance is so critical here. It's performance is so critical here. It's performance is so critical here. It's what unlocks these UI UX u paradigms what unlocks these UI UX u paradigms what unlocks these UI UX u paradigms that something like you know Jev for the that something like you know Jev for the that something like you know Jev for the original release of chat GPT was able to original release of chat GPT was able to original release of chat GPT was able to just like open up that was not there just like open up that was not there just like open up that was not there before. Um so this performance is super before. Um so this performance is super before. Um so this performance is super critical and for chat bots the sort of critical and for chat bots the sort of critical and for chat bots the sort of figure of merit is how many tokens per
-
figure of merit is how many tokens per figure of merit is how many tokens per second can you get back to a user. So second can you get back to a user. So second can you get back to a user. So tokens per second per user um for a tokens per second per user um for a tokens per second per user um for a decision model or a zero. Um, and then decision model or a zero. Um, and then decision model or a zero. Um, and then the sort of final category, the sort of the sort of final category, the sort of the sort of final category, the sort of high latency inference category, these high latency inference category, these high latency inference category, these like data processors that might operate like data processors that might operate like data processors that might operate on a whole bunch of data from your on a whole bunch of data from your on a whole bunch of data from your backend or might consume in Reduct's backend or might consume in Reduct's backend or might consume in Reduct's case like you know thousands, millions case like you know thousands, millions case like you know thousands, millions of PDFs, they have a different sort of of PDFs, they have a different sort of of PDFs, they have a different sort of figure of merit, a different performance figure of merit, a different performance figure of merit, a different performance target, the number of millions of tokens target, the number of millions of tokens target, the number of millions of tokens that you can that you can process for a that you can that you can process for a that you can that you can process for a dollar or maybe billions of tokens. Um, dollar or maybe billions of tokens. Um, dollar or maybe billions of tokens. Um, so this is higher latency inference, so so this is higher latency inference, so so this is higher latency inference, so we won't talk about it. Um we're going we won't talk about it. Um we're going we won't talk about it. Um we're going to focus on the low latency inference. to focus on the low latency inference. to focus on the low latency inference. Um so in the sort of agent category um Um so in the sort of agent category um Um so in the sort of agent category um we've been working on serving the uh we've been working on serving the uh we've been working on serving the uh Kimmy models from Moonshot AI and the Kimmy models from Moonshot AI and the Kimmy models from Moonshot AI and the like the difference between what you get like the difference between what you get like the difference between what you get from a sort of like offtheshelf let me from a sort of like offtheshelf let me from a sort of like offtheshelf let me take the description in SG Lang's take the description in SG Lang's take the description in SG Lang's cookbook or VLM's cookbook and serve it cookbook or VLM's cookbook and serve it cookbook or VLM's cookbook and serve it versus what you can get if you sit down versus what you can get if you sit down versus what you can get if you sit down and look at the workload and optimize um and look at the workload and optimize um and look at the workload and optimize um do a lot of configuration tuning do a do a lot of configuration tuning do a do a lot of configuration tuning do a little bit of uh um uh like performance little bit of uh um uh like performance little bit of uh um uh like performance engineering you can get um in our case engineering you can get um in our case engineering you can get um in our case we were able to get about you know five we were able to get about you know five we were able to get about you know five and a half times more uh efficiency and a half times more uh efficiency and a half times more uh efficiency tokens per minute per GPU but also three tokens per minute per GPU but also three tokens per minute per GPU but also three about three times faster per user. So um about three times faster per user. So um about three times faster per user. So um about like 400 tokens per second per about like 400 tokens per second per about like 400 tokens per second per user instead of a little over uh like user instead of a little over uh like user instead of a little over uh like 130 or so. Um and this has allowed us to
-
130 or so. Um and this has allowed us to 130 or so. Um and this has allowed us to like serve this at large scale through a like serve this at large scale through a like serve this at large scale through a variety of platforms like open router variety of platforms like open router variety of platforms like open router and in partnership with our customers. and in partnership with our customers. and in partnership with our customers. Um and then on the low latency side, Um and then on the low latency side, Um and then on the low latency side, obviously you mentioned that Jev is obviously you mentioned that Jev is obviously you mentioned that Jev is running on modal, but also um we've uh running on modal, but also um we've uh running on modal, but also um we've uh written about um uh how we worked with written about um uh how we worked with written about um uh how we worked with Decagon to get the like lowest possible Decagon to get the like lowest possible Decagon to get the like lowest possible latency for their like uh for their latency for their like uh for their latency for their like uh for their agent operating procedures that sort of agent operating procedures that sort of agent operating procedures that sort of do um this decision work um inside of do um this decision work um inside of do um this decision work um inside of their voice agent pipelines. their voice agent pipelines. their voice agent pipelines. Um so let's talk about how that's Um so let's talk about how that's Um so let's talk about how that's actually done. actually done. actually done. Um so the sort of like playbook we've Um so the sort of like playbook we've Um so the sort of like playbook we've developed. Um so the first step is that developed. Um so the first step is that developed. Um so the first step is that you need to understand the workload that you need to understand the workload that you need to understand the workload that you are serving and the hardware that it you are serving and the hardware that it you are serving and the hardware that it runs on. Uh and then once you have a runs on. Uh and then once you have a runs on. Uh and then once you have a clear picture what that is then you can clear picture what that is then you can clear picture what that is then you can start optimizing the performance. The start optimizing the performance. The start optimizing the performance. The workload tells you what you have to do workload tells you what you have to do workload tells you what you have to do and the hardware is what sort of does it and the hardware is what sort of does it and the hardware is what sort of does it for you and it it says this thing will for you and it it says this thing will for you and it it says this thing will be easy. This thing will be hard. This be easy. This thing will be hard. This be easy. This thing will be hard. This thing will be 10 times as fast as this thing will be 10 times as fast as this thing will be 10 times as fast as this other thing.
-
other thing. other thing. So in understanding the workload just to So in understanding the workload just to So in understanding the workload just to make sure we like I think a lot of make sure we like I think a lot of make sure we like I think a lot of people have been substantively consuming people have been substantively consuming people have been substantively consuming AI through APIs and not necessarily AI through APIs and not necessarily AI through APIs and not necessarily thinking about what's going on inside. thinking about what's going on inside. thinking about what's going on inside. So just to make sure we're all on the So just to make sure we're all on the So just to make sure we're all on the same page. Um we like you send in same page. Um we like you send in same page. Um we like you send in something like this thou shalt not something like this thou shalt not something like this thou shalt not create uh into a language model create uh into a language model create uh into a language model inference server. Um this uh the uh inference server. Um this uh the uh inference server. Um this uh the uh internal representation of the model is internal representation of the model is internal representation of the model is calculated on this input in the prefill calculated on this input in the prefill calculated on this input in the prefill phase. Um that generates a few tokens phase. Um that generates a few tokens phase. Um that generates a few tokens and then you iterate that over and over and then you iterate that over and over and then you iterate that over and over again in the decode phase to generate again in the decode phase to generate again in the decode phase to generate the outputs and complete the commandment the outputs and complete the commandment the outputs and complete the commandment uh from the orange Catholic Bible that uh from the orange Catholic Bible that uh from the orange Catholic Bible that thou shalt not create a machine in the thou shalt not create a machine in the thou shalt not create a machine in the likeness of a human mind. And along the likeness of a human mind. And along the likeness of a human mind. And along the way, you store those internal way, you store those internal way, you store those internal representations in something called the representations in something called the representations in something called the KV cache so that you can uh retrieve it KV cache so that you can uh retrieve it KV cache so that you can uh retrieve it later instead of having to recalculate later instead of having to recalculate later instead of having to recalculate it. it. it. Um, and that's a pretty important piece Um, and that's a pretty important piece Um, and that's a pretty important piece especially if you want to serve low especially if you want to serve low especially if you want to serve low latency agent inference. If you're latency agent inference. If you're latency agent inference. If you're serving something that looks a little serving something that looks a little serving something that looks a little bit more like Jev, this is less bit more like Jev, this is less bit more like Jev, this is less important because of the sort of important because of the sort of important because of the sort of structure of the workload, the structure structure of the workload, the structure structure of the workload, the structure of the requests that you're going to of the requests that you're going to of the requests that you're going to receive. Uh so with an agent you see receive. Uh so with an agent you see receive. Uh so with an agent you see something like this where you have the something like this where you have the something like this where you have the the first user input comes in the LM the first user input comes in the LM the first user input comes in the LM responds and then later you get another responds and then later you get another responds and then later you get another request from the same user that has that request from the same user that has that request from the same user that has that LM response and their first input in it LM response and their first input in it LM response and their first input in it and then you get another maybe a tool and then you get another maybe a tool and then you get another maybe a tool call from the LM and once that tool call call from the LM and once that tool call call from the LM and once that tool call comes back that also is going to have
-
comes back that also is going to have comes back that also is going to have the whole prior context in it and so you the whole prior context in it and so you the whole prior context in it and so you have this iterative construction of the have this iterative construction of the have this iterative construction of the input sequence turn after turn. Um so input sequence turn after turn. Um so input sequence turn after turn. Um so some things about this um like this some things about this um like this some things about this um like this field and about the way LM and field and about the way LM and field and about the way LM and artificial intelligence work are kind of artificial intelligence work are kind of artificial intelligence work are kind of like flash in the pan. Um but this like flash in the pan. Um but this like flash in the pan. Um but this actually kind of feels very deep and actually kind of feels very deep and actually kind of feels very deep and seems like something that'll stick seems like something that'll stick seems like something that'll stick around for a while. Essentially the around for a while. Essentially the around for a while. Essentially the user, the model um and the external user, the model um and the external user, the model um and the external systems that it's accessing are systems that it's accessing are systems that it's accessing are iteratively constructing this context iteratively constructing this context iteratively constructing this context and they're creating something of and they're creating something of and they're creating something of increasing value over time. And so we increasing value over time. And so we increasing value over time. And so we should expect these kinds of interactive should expect these kinds of interactive should expect these kinds of interactive um like constructions to be part of how um like constructions to be part of how um like constructions to be part of how these sequence models end up working. these sequence models end up working. these sequence models end up working. Even once they become, you know, Even once they become, you know, Even once they become, you know, something inside of a robot moving something inside of a robot moving something inside of a robot moving around your house or uh once they start around your house or uh once they start around your house or uh once they start taking in videos and outputting videos, taking in videos and outputting videos, taking in videos and outputting videos, they'll still probably look quite a bit they'll still probably look quite a bit they'll still probably look quite a bit like this. like this. like this. Um yeah, now for a little something Um yeah, now for a little something Um yeah, now for a little something that's a little bit more narrow in scope that's a little bit more narrow in scope that's a little bit more narrow in scope or momentary. Uh the good news is that or momentary. Uh the good news is that or momentary. Uh the good news is that right now um if you're focused on high right now um if you're focused on high right now um if you're focused on high interactivity, low latency inference, interactivity, low latency inference, interactivity, low latency inference, you actually can usually run on no more you actually can usually run on no more you actually can usually run on no more than eight data center GPUs. So if you than eight data center GPUs. So if you than eight data center GPUs. So if you check out the really excellent check out the really excellent check out the really excellent benchmarking work from semi analysis and benchmarking work from semi analysis and benchmarking work from semi analysis and their inference X or inference max their inference X or inference max their inference X or inference max benchmarks, um this is maybe a little benchmarks, um this is maybe a little benchmarks, um this is maybe a little hard to see um because they're very hard to see um because they're very hard to see um because they're very dense charts. This is like the least dense charts. This is like the least dense charts. This is like the least dense version I could pull out. Um, you dense version I could pull out. Um, you dense version I could pull out. Um, you can see that if you want the highest can see that if you want the highest can see that if you want the highest throughput per chip, you want these big
-
throughput per chip, you want these big throughput per chip, you want these big like megawatt scale, not quite a like megawatt scale, not quite a like megawatt scale, not quite a megawatt, these really huge machines megawatt, these really huge machines megawatt, these really huge machines that have a bunch of GPUs in a in a that have a bunch of GPUs in a in a that have a bunch of GPUs in a in a tight domain with each other, an envy tight domain with each other, an envy tight domain with each other, an envy link domain. U, you need like what it link domain. U, you need like what it link domain. U, you need like what it says here, 40 of those GPUs, 64 of those says here, 40 of those GPUs, 64 of those says here, 40 of those GPUs, 64 of those GPUs. Um, but if you want to serve at GPUs. Um, but if you want to serve at GPUs. Um, but if you want to serve at the highest interactivity, if you want the highest interactivity, if you want the highest interactivity, if you want the highest tokens per second per user, the highest tokens per second per user, the highest tokens per second per user, you actually want to be all the way down you actually want to be all the way down you actually want to be all the way down here. Um, and if you look at the points here. Um, and if you look at the points here. Um, and if you look at the points that are down here, you can see that that are down here, you can see that that are down here, you can see that they have four or eight GPUs. Um, so you they have four or eight GPUs. Um, so you they have four or eight GPUs. Um, so you don't need um these fancy things, which don't need um these fancy things, which don't need um these fancy things, which is a huge operational win. is a huge operational win. is a huge operational win. Um, okay. So, we've understood like the Um, okay. So, we've understood like the Um, okay. So, we've understood like the workload that we're thinking about and workload that we're thinking about and workload that we're thinking about and the hardware that we want to run on. the hardware that we want to run on. the hardware that we want to run on. Now, let's think about how we're going Now, let's think about how we're going Now, let's think about how we're going to optimize the performance. to optimize the performance. to optimize the performance. Um, so roughly the way that things work, Um, so roughly the way that things work, Um, so roughly the way that things work, um, is you want to target a bunch of big um, is you want to target a bunch of big um, is you want to target a bunch of big wins first. So speculative decoding is wins first. So speculative decoding is wins first. So speculative decoding is one of the biggest ones. Um you can get one of the biggest ones. Um you can get one of the biggest ones. Um you can get a five time speed up or more. Uh a five time speed up or more. Uh a five time speed up or more. Uh quantization reducing the precision is quantization reducing the precision is quantization reducing the precision is another big speed up has quality another big speed up has quality another big speed up has quality implications. Um once you get those big implications. Um once you get those big implications. Um once you get those big wins out of the way, there's like kind wins out of the way, there's like kind wins out of the way, there's like kind of a grind of mediumsiz tasks on the of a grind of mediumsiz tasks on the of a grind of mediumsiz tasks on the host side. Um like sort of regular CPU host side. Um like sort of regular CPU host side. Um like sort of regular CPU systems performance engineering stuff systems performance engineering stuff systems performance engineering stuff you got to do. Um that you wins 50% you got to do. Um that you wins 50% you got to do. Um that you wins 50% maybe as large as 2x. um the weirder the maybe as large as 2x. um the weirder the maybe as large as 2x. um the weirder the thing you're doing is the more wins thing you're doing is the more wins thing you're doing is the more wins they're probably going to be on the host they're probably going to be on the host they're probably going to be on the host side. Um and then only then do you spend side. Um and then only then do you spend side. Um and then only then do you spend your time you know working on uh GPU your time you know working on uh GPU your time you know working on uh GPU kernels and the sort of thing that like kernels and the sort of thing that like kernels and the sort of thing that like everybody gets most excited about. Um so
-
everybody gets most excited about. Um so everybody gets most excited about. Um so let's go through that real quick at a let's go through that real quick at a let's go through that real quick at a high level like what do all these things high level like what do all these things high level like what do all these things look like and and when should you use look like and and when should you use look like and and when should you use them? So big wins first. Speculative them? So big wins first. Speculative them? So big wins first. Speculative decoding takes that decode process that decoding takes that decode process that decoding takes that decode process that we talked about where you generate um uh we talked about where you generate um uh we talked about where you generate um uh an output sentence like the LM is going an output sentence like the LM is going an output sentence like the LM is going to generate from modal is a to modal is to generate from modal is a to modal is to generate from modal is a to modal is a serverless computing platform. The a serverless computing platform. The a serverless computing platform. The thing you do instead of running this guy thing you do instead of running this guy thing you do instead of running this guy one step at a time is you take something one step at a time is you take something one step at a time is you take something much cheaper to run and you run it um much cheaper to run and you run it um much cheaper to run and you run it um and generate the output tokens and you and generate the output tokens and you and generate the output tokens and you take a guess at what those output tokens take a guess at what those output tokens take a guess at what those output tokens might be. you speculate and the idea might be. you speculate and the idea might be. you speculate and the idea here is the same as speculative here is the same as speculative here is the same as speculative execution inside processors. If you have execution inside processors. If you have execution inside processors. If you have some like slack in a system um which is some like slack in a system um which is some like slack in a system um which is what happens when you run looms one what happens when you run looms one what happens when you run looms one token at a time uh then you can uh sort token at a time uh then you can uh sort token at a time uh then you can uh sort of use some of that slack to do work of use some of that slack to do work of use some of that slack to do work that might not be accepted. So run that might not be accepted. So run that might not be accepted. So run instructions that might not actually be instructions that might not actually be instructions that might not actually be on the code path in speculative on the code path in speculative on the code path in speculative execution or process tokens that might execution or process tokens that might execution or process tokens that might not show up in the output like the word not show up in the output like the word not show up in the output like the word company here. Um, but you like it's such company here. Um, but you like it's such company here. Um, but you like it's such a big win to be able to run these things a big win to be able to run these things a big win to be able to run these things in parallel, especially when you're in parallel, especially when you're in parallel, especially when you're targeting high interactivity that this targeting high interactivity that this targeting high interactivity that this um is like very frequently a humongous um is like very frequently a humongous um is like very frequently a humongous win. Um, and it's sort of linear in how win. Um, and it's sort of linear in how win. Um, and it's sort of linear in how many tokens you can get the speculator many tokens you can get the speculator many tokens you can get the speculator to guess. This is from one of our blog to guess. This is from one of our blog to guess. This is from one of our blog posts about speculators. Um, basically posts about speculators. Um, basically posts about speculators. Um, basically as you increase the number of tokens as you increase the number of tokens as you increase the number of tokens that are accepted on average, your speed that are accepted on average, your speed that are accepted on average, your speed up goes up as well. And these are these up goes up as well. And these are these up goes up as well. And these are these are integral numbers like 4x faster, 8x
-
are integral numbers like 4x faster, 8x are integral numbers like 4x faster, 8x faster. If you're you know do faster. If you're you know do faster. If you're you know do performance work, you're used to like performance work, you're used to like performance work, you're used to like you know sweating over 5% or 10% but you know sweating over 5% or 10% but you know sweating over 5% or 10% but these are like 500%. these are like 500%. these are like 500%. Um, and then one of my favorite things Um, and then one of my favorite things Um, and then one of my favorite things about this as somebody who worked on about this as somebody who worked on about this as somebody who worked on machine learning training from the machine learning training from the machine learning training from the beginning is that the difference between beginning is that the difference between beginning is that the difference between a point like this and as a speculator a point like this and as a speculator a point like this and as a speculator and model pair like this and one like and model pair like this and one like and model pair like this and one like this is that you um you customize to the this is that you um you customize to the this is that you um you customize to the specific workload that you want the the specific workload that you want the the specific workload that you want the the the speculator uh or the target model to the speculator uh or the target model to the speculator uh or the target model to work on. you train the speculator to work on. you train the speculator to work on. you train the speculator to know how that model behaves on that data know how that model behaves on that data know how that model behaves on that data and all of a sudden you go from six and all of a sudden you go from six and all of a sudden you go from six times faster with it three times faster times faster with it three times faster times faster with it three times faster to five times faster. Um and this is you to five times faster. Um and this is you to five times faster. Um and this is you the bitter lesson of machine learning the bitter lesson of machine learning the bitter lesson of machine learning but apply to performance engineering but apply to performance engineering but apply to performance engineering which is exciting. Um people often ask which is exciting. Um people often ask which is exciting. Um people often ask like wait how can some dumber model like wait how can some dumber model like wait how can some dumber model guess what the smarter model is going to guess what the smarter model is going to guess what the smarter model is going to say? There's a couple reasons why this say? There's a couple reasons why this say? There's a couple reasons why this works. Um the biggest one is that works. Um the biggest one is that works. Um the biggest one is that there's a lot of structure in outputs.
-
there's a lot of structure in outputs. there's a lot of structure in outputs. Um, and so if you take something from Um, and so if you take something from Um, and so if you take something from your coding agent sessions like the uh your coding agent sessions like the uh your coding agent sessions like the uh the thing above and then you translate the thing above and then you translate the thing above and then you translate it into what the model sees, this it into what the model sees, this it into what the model sees, this becomes even more obvious. You have chat becomes even more obvious. You have chat becomes even more obvious. You have chat templating that has all these special templating that has all these special templating that has all these special tokens which are usually very tokens which are usually very tokens which are usually very predictable. And then this right here is predictable. And then this right here is predictable. And then this right here is a quote from the user which probably a quote from the user which probably a quote from the user which probably appears earlier in the transcript. So appears earlier in the transcript. So appears earlier in the transcript. So that's easy to guess. There are only so that's easy to guess. There are only so that's easy to guess. There are only so many tool calls. um there's only a few many tool calls. um there's only a few many tool calls. um there's only a few of them that are something you would say of them that are something you would say of them that are something you would say you would do after saying you were going you would do after saying you were going you would do after saying you were going to read that file. And so this is like to read that file. And so this is like to read that file. And so this is like pretty easy to predict. Um but what the pretty easy to predict. Um but what the pretty easy to predict. Um but what the model the target model is doing is model the target model is doing is model the target model is doing is actually like processing these tokens actually like processing these tokens actually like processing these tokens and creating this rich internal and creating this rich internal and creating this rich internal representation that is going to be used representation that is going to be used representation that is going to be used for many tokens in the future. Um so you for many tokens in the future. Um so you for many tokens in the future. Um so you by being able to like guess what those by being able to like guess what those by being able to like guess what those specific values are with a good rate, specific values are with a good rate, specific values are with a good rate, you're able to parallelize that um you're able to parallelize that um you're able to parallelize that um construction of their representation. Um construction of their representation. Um construction of their representation. Um which is pretty cool. Um, and that which is pretty cool. Um, and that which is pretty cool. Um, and that representation actually gets reused representation actually gets reused representation actually gets reused inside of current speculator models. inside of current speculator models. inside of current speculator models. They don't just have to they're not like They don't just have to they're not like They don't just have to they're not like a marov model, an engram speculator or a marov model, an engram speculator or a marov model, an engram speculator or another language model from scratch.
-
another language model from scratch. another language model from scratch. They take those rich internal They take those rich internal They take those rich internal representations from the target model representations from the target model representations from the target model and they uh they use that they use and they uh they use that they use and they uh they use that they use what's what's going on inside the model. what's what's going on inside the model. what's what's going on inside the model. So they're able to kind of like uh So they're able to kind of like uh So they're able to kind of like uh bootstrap from that. So the specific bootstrap from that. So the specific bootstrap from that. So the specific architecture there was Dlash. um that's architecture there was Dlash. um that's architecture there was Dlash. um that's the the sort of speculator architecture the the sort of speculator architecture the the sort of speculator architecture that we've found uh works best um and that we've found uh works best um and that we've found uh works best um and the one that we've invested the most in the one that we've invested the most in the one that we've invested the most in and gives us like large speed ups over and gives us like large speed ups over and gives us like large speed ups over both built-in speculative decoding in both built-in speculative decoding in both built-in speculative decoding in certain models and of course over the certain models and of course over the certain models and of course over the baseline. baseline. baseline. Um after uh speculative decoding and Um after uh speculative decoding and Um after uh speculative decoding and getting that right, the next big win is getting that right, the next big win is getting that right, the next big win is quantization. So with quantization, you quantization. So with quantization, you quantization. So with quantization, you take the floatingoint representation of take the floatingoint representation of take the floatingoint representation of the weights or activations of the model the weights or activations of the model the weights or activations of the model and you decrease the bit width. So it and you decrease the bit width. So it and you decrease the bit width. So it goes from 16 bits to 8 bits to maybe goes from 16 bits to 8 bits to maybe goes from 16 bits to 8 bits to maybe even four bits. Um so this is a linear even four bits. Um so this is a linear even four bits. Um so this is a linear reduction in memory bandwidth demand and reduction in memory bandwidth demand and reduction in memory bandwidth demand and a linear increase in floatingoint a linear increase in floatingoint a linear increase in floatingoint operations per second because the the operations per second because the the operations per second because the the the bits are smaller or the the the the bits are smaller or the the the the bits are smaller or the the the numbers are smaller. Um so this I say numbers are smaller. Um so this I say numbers are smaller. Um so this I say this requires full stack coordination this requires full stack coordination this requires full stack coordination because you have to go down to the because you have to go down to the because you have to go down to the hardware layer and be like okay am I hardware layer and be like okay am I hardware layer and be like okay am I running on a GPU that can actually make running on a GPU that can actually make running on a GPU that can actually make use of these things? Do I have a H100 or use of these things? Do I have a H100 or use of these things? Do I have a H100 or later for FP8? do I have a B200 or later later for FP8? do I have a B200 or later later for FP8? do I have a B200 or later for uh FP4? Um but it also changes model for uh FP4? Um but it also changes model for uh FP4? Um but it also changes model behavior. The magic of spec of decoding behavior. The magic of spec of decoding behavior. The magic of spec of decoding is that it changes nothing about the is that it changes nothing about the is that it changes nothing about the probability distribution of the model's probability distribution of the model's probability distribution of the model's outputs. Um but quantization does and so outputs. Um but quantization does and so outputs. Um but quantization does and so you kind of have to go all the way back you kind of have to go all the way back you kind of have to go all the way back up to the application layer and ask is up to the application layer and ask is up to the application layer and ask is this still sufficient intelligence? Is this still sufficient intelligence? Is this still sufficient intelligence? Is this are the capabilities still what I this are the capabilities still what I this are the capabilities still what I need if I quantize the model? Um, so
-
need if I quantize the model? Um, so need if I quantize the model? Um, so with quantization, there's this nice with quantization, there's this nice with quantization, there's this nice little visualizer on our LM engineers little visualizer on our LM engineers little visualizer on our LM engineers almanac that shows you like the effect almanac that shows you like the effect almanac that shows you like the effect of quantization. In this case, 4bit of quantization. In this case, 4bit of quantization. In this case, 4bit quantization on this uh image of a quantization on this uh image of a quantization on this uh image of a parrot here versus the original. I think parrot here versus the original. I think parrot here versus the original. I think maybe as a demo of when quantization maybe as a demo of when quantization maybe as a demo of when quantization does work. I think the people in the does work. I think the people in the does work. I think the people in the front row can probably see that this is front row can probably see that this is front row can probably see that this is a little bit um a little jank, but the a little bit um a little jank, but the a little bit um a little jank, but the people in the back row and maybe um you people in the back row and maybe um you people in the back row and maybe um you know people squinting at their phones know people squinting at their phones know people squinting at their phones like maybe can't see that that's um that like maybe can't see that that's um that like maybe can't see that that's um that that's been quantized. And so that's the that's been quantized. And so that's the that's been quantized. And so that's the effect that you can rely on. Um sort of effect that you can rely on. Um sort of effect that you can rely on. Um sort of depends on what you're doing. So those depends on what you're doing. So those depends on what you're doing. So those are big wins. Let's talk about the host are big wins. Let's talk about the host are big wins. Let's talk about the host performance grind. So the key thing here performance grind. So the key thing here performance grind. So the key thing here is that the GPU does all the work, but is that the GPU does all the work, but is that the GPU does all the work, but the CPU can get in the way. the CPU is the CPU can get in the way. the CPU is the CPU can get in the way. the CPU is coordinating all the work that's going coordinating all the work that's going coordinating all the work that's going on in the GPU. And you want to make sure on in the GPU. And you want to make sure on in the GPU. And you want to make sure that the CPU work that is going on does that the CPU work that is going on does that the CPU work that is going on does not block the GPU work. Like this is not block the GPU work. Like this is not block the GPU work. Like this is finished. Oh, now I need to do some work finished. Oh, now I need to do some work finished. Oh, now I need to do some work to decide what to do next. And then oh, to decide what to do next. And then oh, to decide what to do next. And then oh, I'm going to start running GPU stuff I'm going to start running GPU stuff I'm going to start running GPU stuff again. That's really bad. It's a bubble again. That's really bad. It's a bubble again. That's really bad. It's a bubble or a stall. Um if you've worked with an or a stall. Um if you've worked with an or a stall. Um if you've worked with an event loop, then you're familiar with event loop, then you're familiar with event loop, then you're familiar with this idea like don't block the fast this idea like don't block the fast this idea like don't block the fast thing. Um there's a key technique for thing. Um there's a key technique for thing. Um there's a key technique for this called CUDA graph capture that this called CUDA graph capture that this called CUDA graph capture that allows you to take like instead of allows you to take like instead of allows you to take like instead of launching like one GPU kernel at a time, launching like one GPU kernel at a time, launching like one GPU kernel at a time, you can just record all the kernels that you can just record all the kernels that you can just record all the kernels that you ran one time and then just launch you ran one time and then just launch you ran one time and then just launch them all at once. Um which is a pretty them all at once. Um which is a pretty them all at once. Um which is a pretty effective way to like cut a lot of this effective way to like cut a lot of this effective way to like cut a lot of this host work out. Um one tool for keeping host work out. Um one tool for keeping host work out. Um one tool for keeping track of this is actually my my favorite track of this is actually my my favorite track of this is actually my my favorite is to just look at the power usage and
-
is to just look at the power usage and is to just look at the power usage and the temperature of the GPUs. if you the temperature of the GPUs. if you the temperature of the GPUs. if you aren't close to 100% power, you probably aren't close to 100% power, you probably aren't close to 100% power, you probably aren't using the whole GPU. They're very aren't using the whole GPU. They're very aren't using the whole GPU. They're very much power limited uh when they're much power limited uh when they're much power limited uh when they're achieving peak performance. Uh and so in achieving peak performance. Uh and so in achieving peak performance. Uh and so in this case, we were able like just by eye this case, we were able like just by eye this case, we were able like just by eye noticing that this thing was way below noticing that this thing was way below noticing that this thing was way below its total GPU power usage even at this its total GPU power usage even at this its total GPU power usage even at this uh uh between this one which is at well uh uh between this one which is at well uh uh between this one which is at well over 2,000 watts and this one which is over 2,000 watts and this one which is over 2,000 watts and this one which is uh below it or kind of capped at it. Um uh below it or kind of capped at it. Um uh below it or kind of capped at it. Um that doesn't tell you what the problem that doesn't tell you what the problem that doesn't tell you what the problem is. You want to use something like a is. You want to use something like a is. You want to use something like a profiler to get in there. You want these profiler to get in there. You want these profiler to get in there. You want these traces. It's a very empirical traces. It's a very empirical traces. It's a very empirical discipline. Um, and so the endsite discipline. Um, and so the endsite discipline. Um, and so the endsite systems or torch profiler, the tools systems or torch profiler, the tools systems or torch profiler, the tools that come from Nvidia for this are kind that come from Nvidia for this are kind that come from Nvidia for this are kind of your uh your go-to. And if you taking of your uh your go-to. And if you taking of your uh your go-to. And if you taking a look at that specific case, we're able a look at that specific case, we're able a look at that specific case, we're able to determine that one of these GPUs is to determine that one of these GPUs is to determine that one of these GPUs is actually falling behind. All these guys actually falling behind. All these guys actually falling behind. All these guys here are all doing the same thing. The here are all doing the same thing. The here are all doing the same thing. The the highlighted blocks, they're all the highlighted blocks, they're all the highlighted blocks, they're all running this deep gem. And these three running this deep gem. And these three running this deep gem. And these three GPUs are all waiting on this one that's GPUs are all waiting on this one that's GPUs are all waiting on this one that's slow. turned out to be like a NUMO slow. turned out to be like a NUMO slow. turned out to be like a NUMO region awareness thing. Um that was region awareness thing. Um that was region awareness thing. Um that was causing one of them to always be a causing one of them to always be a causing one of them to always be a little slower because it was on the little slower because it was on the little slower because it was on the wrong socket. Um there's a lot of host wrong socket. Um there's a lot of host wrong socket. Um there's a lot of host level things that you can use PI Spy level things that you can use PI Spy level things that you can use PI Spy like the things are running in Python.
-
like the things are running in Python. like the things are running in Python. So you could look at what's going on in So you could look at what's going on in So you could look at what's going on in Python and figure things out pretty Python and figure things out pretty Python and figure things out pretty well. Um we have a blog post that talks well. Um we have a blog post that talks well. Um we have a blog post that talks about one case in particular where we about one case in particular where we about one case in particular where we took this like repeated sort of took this like repeated sort of took this like repeated sort of allocation and turned it into just a allocation and turned it into just a allocation and turned it into just a cache, a Python dictionary. Um, but I cache, a Python dictionary. Um, but I cache, a Python dictionary. Um, but I guess you get to call it a CUDA IPC pool guess you get to call it a CUDA IPC pool guess you get to call it a CUDA IPC pool handle cache if it's uh AIG GPU stuff. handle cache if it's uh AIG GPU stuff. handle cache if it's uh AIG GPU stuff. Um, but a Python dictionary that just Um, but a Python dictionary that just Um, but a Python dictionary that just held on to handles for these and held on to handles for these and held on to handles for these and suddenly that made the thing run 10% suddenly that made the thing run 10% suddenly that made the thing run 10% faster. faster. faster. Um, so this was mostly focused on like a Um, so this was mostly focused on like a Um, so this was mostly focused on like a one single replica like a couple of GPUs one single replica like a couple of GPUs one single replica like a couple of GPUs at a time. When you go to scale up, um, at a time. When you go to scale up, um, at a time. When you go to scale up, um, you're going to run into routing you're going to run into routing you're going to run into routing problems. So you might predict problems. So you might predict problems. So you might predict performance before scaling that that performance before scaling that that performance before scaling that that just doesn't uh happen once you scale just doesn't uh happen once you scale just doesn't uh happen once you scale up. Um, so one particular thing that up. Um, so one particular thing that up. Um, so one particular thing that might happen is you might uh end up with might happen is you might uh end up with might happen is you might uh end up with like tail latencies that are much higher like tail latencies that are much higher like tail latencies that are much higher than you expect and that kind of show up than you expect and that kind of show up than you expect and that kind of show up seemingly at random. Um, and the answer seemingly at random. Um, and the answer seemingly at random. Um, and the answer is that they are kind of at random. They is that they are kind of at random. They is that they are kind of at random. They come from the way that you do your come from the way that you do your come from the way that you do your routing. Um, and this is especially routing. Um, and this is especially routing. Um, and this is especially important for agent workloads because of important for agent workloads because of important for agent workloads because of the KV cache. If you're not routing the KV cache. If you're not routing the KV cache. If you're not routing carefully like with respect to the KV carefully like with respect to the KV carefully like with respect to the KV cache, then the difference between like cache, then the difference between like cache, then the difference between like a cache hit and a cache missed is maybe a cache hit and a cache missed is maybe a cache hit and a cache missed is maybe 30 seconds versus 500 milliseconds. Um, 30 seconds versus 500 milliseconds. Um, 30 seconds versus 500 milliseconds. Um, and so uh measured this in one of our and so uh measured this in one of our and so uh measured this in one of our deployments and determined that there deployments and determined that there deployments and determined that there was variety of reasons causing instead was variety of reasons causing instead was variety of reasons causing instead of the you know four or five uh of the you know four or five uh of the you know four or five uh concurrent requests we were trying to concurrent requests we were trying to concurrent requests we were trying to aim for, we we saw this major overd aim for, we we saw this major overd aim for, we we saw this major overd dispersion of requests and the solution dispersion of requests and the solution dispersion of requests and the solution was just to make the router smarter. had
-
was just to make the router smarter. had was just to make the router smarter. had to become much more stateful um and to become much more stateful um and to become much more stateful um and aware application aware of things like aware application aware of things like aware application aware of things like the load and the KV cache state but that the load and the KV cache state but that the load and the KV cache state but that was able to sort of pull in this uh was able to sort of pull in this uh was able to sort of pull in this uh overd dispersion and uh massively you overd dispersion and uh massively you overd dispersion and uh massively you know get rid of these uh spikes and know get rid of these uh spikes and know get rid of these uh spikes and massively improve performance. massively improve performance. massively improve performance. Um so then and only then are you allowed Um so then and only then are you allowed Um so then and only then are you allowed to think about you know GPU performance to think about you know GPU performance to think about you know GPU performance and optimizing things on GPUs. In and optimizing things on GPUs. In and optimizing things on GPUs. In general kernels are relatively standard. general kernels are relatively standard. general kernels are relatively standard. They're close to the speed of light. They're close to the speed of light. They're close to the speed of light. architectures are not changing that architectures are not changing that architectures are not changing that much. Um, and so there are really good much. Um, and so there are really good much. Um, and so there are really good kernels out there from Nvidia and from kernels out there from Nvidia and from kernels out there from Nvidia and from like treehouse group and others. So like treehouse group and others. So like treehouse group and others. So there's only a few percentage points to there's only a few percentage points to there's only a few percentage points to gain there. Um, so that matters a lot at gain there. Um, so that matters a lot at gain there. Um, so that matters a lot at scale, but it's not something to worry scale, but it's not something to worry scale, but it's not something to worry about at the like, you know, at the about at the like, you know, at the about at the like, you know, at the beginning of serving your own inference. beginning of serving your own inference. beginning of serving your own inference. Um, the specific tool for this endsite Um, the specific tool for this endsite Um, the specific tool for this endsite compute allows you to like go all the compute allows you to like go all the compute allows you to like go all the way down to the assembler level and way down to the assembler level and way down to the assembler level and figure out where performance problems figure out where performance problems figure out where performance problems are coming from. Um, but before you pull are coming from. Um, but before you pull are coming from. Um, but before you pull out insight compute, I'd suggest you out insight compute, I'd suggest you out insight compute, I'd suggest you pull out um a marker and a whiteboard. pull out um a marker and a whiteboard. pull out um a marker and a whiteboard. Um, if you look at uh Tree Dow's uh and Um, if you look at uh Tree Dow's uh and Um, if you look at uh Tree Dow's uh and together's blog about their flash together's blog about their flash together's blog about their flash tension for kernel work. Um, the sort of tension for kernel work. Um, the sort of tension for kernel work. Um, the sort of the way they decided what to work on, the way they decided what to work on, the way they decided what to work on, the way they decided what problems the way they decided what problems the way they decided what problems mattered was just to write down the mattered was just to write down the mattered was just to write down the feeds and speeds like how much data is feeds and speeds like how much data is feeds and speeds like how much data is coming in, how many operations need to coming in, how many operations need to coming in, how many operations need to come uh happen on it and how quickly can come uh happen on it and how quickly can come uh happen on it and how quickly can we move it back out. And uh that is we move it back out. And uh that is we move it back out. And uh that is something that you can do you know something that you can do you know something that you can do you know before um you know from first principles before um you know from first principles before um you know from first principles and this sort of applies across the and this sort of applies across the and this sort of applies across the stack. If you understand these things stack. If you understand these things stack. If you understand these things about the like hardware that you're
-
about the like hardware that you're about the like hardware that you're operating with then uh you can really um operating with then uh you can really um operating with then uh you can really um you can get a lot further um without you can get a lot further um without you can get a lot further um without confusing yourself and wasting money. confusing yourself and wasting money. confusing yourself and wasting money. Um great. So I said that we weren't Um great. So I said that we weren't Um great. So I said that we weren't going to talk about high about high going to talk about high about high going to talk about high about high throughput but I actually lied. Um throughput but I actually lied. Um throughput but I actually lied. Um there's a little piece of work that there's a little piece of work that there's a little piece of work that actually is just coming out today and actually is just coming out today and actually is just coming out today and that I wanted to share with folks here that I wanted to share with folks here that I wanted to share with folks here at AI Engineer Paris. Um we've been at AI Engineer Paris. Um we've been at AI Engineer Paris. Um we've been thinking a lot about low latency thinking a lot about low latency thinking a lot about low latency inference at Modal because of how much inference at Modal because of how much inference at Modal because of how much people, you know, love coding agents, people, you know, love coding agents, people, you know, love coding agents, how much they love things like Jev. Um how much they love things like Jev. Um how much they love things like Jev. Um but there's actually quite a bit of but there's actually quite a bit of but there's actually quite a bit of demand for high latency inference for demand for high latency inference for demand for high latency inference for things that look very different. Um, so things that look very different. Um, so things that look very different. Um, so the sort of prototypical example of this the sort of prototypical example of this the sort of prototypical example of this and that we're going to talk about is AI and that we're going to talk about is AI and that we're going to talk about is AI SQL. So this is not the same as AI SQL. So this is not the same as AI SQL. So this is not the same as AI writing SQL, which my understanding is writing SQL, which my understanding is writing SQL, which my understanding is AI now writes all of the SQL in the AI now writes all of the SQL in the AI now writes all of the SQL in the world. Um, certainly I wrote this um, world. Um, certainly I wrote this um, world. Um, certainly I wrote this um, but you know, this is this is not but you know, this is this is not but you know, this is this is not something I run in production. Um, but something I run in production. Um, but something I run in production. Um, but yeah, so this is not AI writing SQL.
-
yeah, so this is not AI writing SQL. yeah, so this is not AI writing SQL. It's actual it's actually the other like It's actual it's actually the other like It's actual it's actually the other like direction, the opposite. It's using SQL direction, the opposite. It's using SQL direction, the opposite. It's using SQL to write prompts for AI. Um, so the idea to write prompts for AI. Um, so the idea to write prompts for AI. Um, so the idea is, you know, you have all this is, you know, you have all this is, you know, you have all this information about your customers and information about your customers and information about your customers and about your products. Have you ever about your products. Have you ever about your products. Have you ever wanted to just ask Chad GBT like which wanted to just ask Chad GBT like which wanted to just ask Chad GBT like which of my customers might buy which of my of my customers might buy which of my of my customers might buy which of my products? Um, and that can be expressed products? Um, and that can be expressed products? Um, and that can be expressed as like a relational join, the classic as like a relational join, the classic as like a relational join, the classic kind of thing that you would normally do kind of thing that you would normally do kind of thing that you would normally do in your database or your and in your in your database or your and in your in your database or your and in your transactional or analytic database. Um, transactional or analytic database. Um, transactional or analytic database. Um, but the problem here is that this guy but the problem here is that this guy but the problem here is that this guy here, this prompt is going to like here, this prompt is going to like here, this prompt is going to like bankrupt you if you're running against bankrupt you if you're running against bankrupt you if you're running against like um a proprietary API. So like like um a proprietary API. So like like um a proprietary API. So like running this on your database might running this on your database might running this on your database might produce like 8 million queries all at produce like 8 million queries all at produce like 8 million queries all at once. Um, and then you know fire them once. Um, and then you know fire them once. Um, and then you know fire them all off and there goes your token all off and there goes your token all off and there goes your token budget. Um, yeah, that's one way to budget. Um, yeah, that's one way to budget. Um, yeah, that's one way to achieve token maxing, but I think we now achieve token maxing, but I think we now achieve token maxing, but I think we now all agree that that's a bad idea. Um so all agree that that's a bad idea. Um so all agree that that's a bad idea. Um so uh the key thing that makes this a uh the key thing that makes this a uh the key thing that makes this a totally different optimization problem totally different optimization problem totally different optimization problem from the low latency inference is that from the low latency inference is that from the low latency inference is that you have this tremendous amount of you have this tremendous amount of you have this tremendous amount of structure. So this is a query plan for a structure. So this is a query plan for a structure. So this is a query plan for a little bit more complicated of a query little bit more complicated of a query little bit more complicated of a query that has a couple of joins in it. Um and that has a couple of joins in it. Um and that has a couple of joins in it. Um and this query plan says these are all the this query plan says these are all the this query plan says these are all the sort of documents I'm going to use to sort of documents I'm going to use to sort of documents I'm going to use to construct prompts. This is how I'm going construct prompts. This is how I'm going construct prompts. This is how I'm going to have them interact with each other.
-
to have them interact with each other. to have them interact with each other. Um, and this is all the 8 million or 10 Um, and this is all the 8 million or 10 Um, and this is all the 8 million or 10 million or whatever prompts I'm gonna million or whatever prompts I'm gonna million or whatever prompts I'm gonna I'm gonna generate and send to the I'm gonna generate and send to the I'm gonna generate and send to the inference engine. And the existing inference engine. And the existing inference engine. And the existing inference engines have no idea about inference engines have no idea about inference engines have no idea about SQL. Um, they like they only know sort SQL. Um, they like they only know sort SQL. Um, they like they only know sort of like request response maybe batch of like request response maybe batch of like request response maybe batch request kind of semantics. Um, so what request kind of semantics. Um, so what request kind of semantics. Um, so what if you took the sort of uh page out of if you took the sort of uh page out of if you took the sort of uh page out of the database book and you did like query the database book and you did like query the database book and you did like query planning and combine that with the planning and combine that with the planning and combine that with the inference engine. So we did that with um inference engine. So we did that with um inference engine. So we did that with um partnered with Treya Shankar at CMU uh partnered with Treya Shankar at CMU uh partnered with Treya Shankar at CMU uh to get about a 14 times speed up over to get about a 14 times speed up over to get about a 14 times speed up over the like baseline VLM on that specific the like baseline VLM on that specific the like baseline VLM on that specific query that I just showed you and achieve query that I just showed you and achieve query that I just showed you and achieve like over a billion tokens per minute on like over a billion tokens per minute on like over a billion tokens per minute on a single H100 and finish this query at a single H100 and finish this query at a single H100 and finish this query at just like $2 to execute instead of like just like $2 to execute instead of like just like $2 to execute instead of like uh almost $30 with VLM and like uh quite uh almost $30 with VLM and like uh quite uh almost $30 with VLM and like uh quite a bit more with um something like uh a bit more with um something like uh a bit more with um something like uh even the smallest OpenAI models. even the smallest OpenAI models. even the smallest OpenAI models. Um, so this is actually open source. We Um, so this is actually open source. We Um, so this is actually open source. We just released it today. Um, this is the just released it today. Um, this is the just released it today. Um, this is the first announcement of it. Um, you can first announcement of it. Um, you can first announcement of it. Um, you can find it at the full stack data labs find it at the full stack data labs find it at the full stack data labs website and their GitHub. Um, and you website and their GitHub. Um, and you website and their GitHub. Um, and you can read about it on our blog on uh, can read about it on our blog on uh, can read about it on our blog on uh, yeah, check it out here. Uh, all the yeah, check it out here. Uh, all the yeah, check it out here. Uh, all the information that I shared today is in information that I shared today is in information that I shared today is in greater detail on our blog and a bunch greater detail on our blog and a bunch greater detail on our blog and a bunch of resources about running your own of resources about running your own of resources about running your own inference on modal.com.
-
inference on modal.com. inference on modal.com. Um, I'll I'm gonna get pulled off the Um, I'll I'm gonna get pulled off the Um, I'll I'm gonna get pulled off the stage in about 30 seconds. So, I'll just stage in about 30 seconds. So, I'll just stage in about 30 seconds. So, I'll just say the reason why I'm telling you all say the reason why I'm telling you all say the reason why I'm telling you all about this is both because I want people about this is both because I want people about this is both because I want people to run their own inference and because I to run their own inference and because I to run their own inference and because I think the contain the uh um serverless think the contain the uh um serverless think the contain the uh um serverless infrastructure platform that we've built infrastructure platform that we've built infrastructure platform that we've built at Bodal is the best place to run this at Bodal is the best place to run this at Bodal is the best place to run this stuff. So, when you go to try and run stuff. So, when you go to try and run stuff. So, when you go to try and run your own inference, I think you'll your own inference, I think you'll your own inference, I think you'll realize that modal is the right place to realize that modal is the right place to realize that modal is the right place to do it. um whether that's uh by running do it. um whether that's uh by running do it. um whether that's uh by running code yourself or by taking advantage of code yourself or by taking advantage of code yourself or by taking advantage of modal auto endpoints um our automated modal auto endpoints um our automated modal auto endpoints um our automated system for producing custom inference system for producing custom inference system for producing custom inference deployments for uh for users. All right, deployments for uh for users. All right, deployments for uh for users. All right, thank you very much. Thank you so much, Charles. All right, Thank you so much, Charles. All right, let's give it up for Charles once more. let's give it up for Charles once more. let's give it up for Charles once more. All right. All right. All right. Yeah, I totally didn't don't want to run Yeah, I totally didn't don't want to run Yeah, I totally didn't don't want to run that query and bankrupt myself. That's that query and bankrupt myself. That's that query and bankrupt myself. That's for sure. Um yeah, our next speaker is for sure. Um yeah, our next speaker is for sure. Um yeah, our next speaker is actually going to talk to us about the actually going to talk to us about the actually going to talk to us about the exact opposite. Uh he's going to talk to exact opposite. Uh he's going to talk to exact opposite. Uh he's going to talk to us about how to run models in uh and in us about how to run models in uh and in us about how to run models in uh and in uh in smaller devices. And so he's going uh in smaller devices. And so he's going uh in smaller devices. And so he's going to talk to us about edi. Our next to talk to us about edi. Our next to talk to us about edi. Our next speaker is Dominic Payak uh and uh he's speaker is Dominic Payak uh and uh he's speaker is Dominic Payak uh and uh he's a senior director at ARM. So please join a senior director at ARM. So please join a senior director at ARM. So please join me in welcoming to the stage senior me in welcoming to the stage senior me in welcoming to the stage senior director at ARM, Dominic Payak.
-
director at ARM, Dominic Payak. director at ARM, Dominic Payak. All right. All right. Hi everyone. All right. Hi everyone. So yeah, my name is Dominic Pike. time So yeah, my name is Dominic Pike. time So yeah, my name is Dominic Pike. time at ARM and today we're going to talk at ARM and today we're going to talk at ARM and today we're going to talk about what happens when you run about what happens when you run about what happens when you run inference at the edge. So there's inference at the edge. So there's inference at the edge. So there's there's three things going on right now there's three things going on right now there's three things going on right now which I think are going to unlock some which I think are going to unlock some which I think are going to unlock some incredible opportunities. One is just incredible opportunities. One is just incredible opportunities. One is just the speed that small openweight capable the speed that small openweight capable the speed that small openweight capable models are becoming available. You models are becoming available. You models are becoming available. You combine that with the fact there's tons combine that with the fact there's tons combine that with the fact there's tons of accessible hardware you can then put of accessible hardware you can then put of accessible hardware you can then put these things on and there's this these things on and there's this these things on and there's this potential for devices that are a little potential for devices that are a little potential for devices that are a little bit more intelligent, a little bit more bit more intelligent, a little bit more bit more intelligent, a little bit more um aware of the intention of the user um aware of the intention of the user um aware of the intention of the user and deliver experiences that have not and deliver experiences that have not and deliver experiences that have not been seen before in all kinds of been seen before in all kinds of been seen before in all kinds of different environments. And so yeah, I'm different environments. And so yeah, I'm different environments. And so yeah, I'm I'm from ARM. And so a little bit about I'm from ARM. And so a little bit about I'm from ARM. And so a little bit about ARM. ARM is the compute platform at the ARM. ARM is the compute platform at the ARM. ARM is the compute platform at the heart of uh devices all around you heart of uh devices all around you heart of uh devices all around you everywhere from highly efficient data everywhere from highly efficient data everywhere from highly efficient data center infrastructure. There's physical center infrastructure. There's physical center infrastructure. There's physical AI with automotive and robotics, so hard AI with automotive and robotics, so hard AI with automotive and robotics, so hard real time safety critical applications real time safety critical applications real time safety critical applications all the way down to pretty much every all the way down to pretty much every all the way down to pretty much every smartphone you've ever owned, client smartphone you've ever owned, client smartphone you've ever owned, client computing and industrial IoT and computing and industrial IoT and computing and industrial IoT and emerging devices. And that's the the
-
emerging devices. And that's the the emerging devices. And that's the the stuff that me and my team have been stuff that me and my team have been stuff that me and my team have been working on enabling new classes of working on enabling new classes of working on enabling new classes of devices to be created and come to devices to be created and come to devices to be created and come to market. market. market. And so I am a computer history geek. I And so I am a computer history geek. I And so I am a computer history geek. I love the fact by the way that this uh love the fact by the way that this uh love the fact by the way that this uh station F is on Parve Allen Turing. I station F is on Parve Allen Turing. I station F is on Parve Allen Turing. I don't know if I pronounced PV correctly, don't know if I pronounced PV correctly, don't know if I pronounced PV correctly, but um it's it's fascinating to look at but um it's it's fascinating to look at but um it's it's fascinating to look at the history of computing and where it's the history of computing and where it's the history of computing and where it's taking us today. So if we look back the taking us today. So if we look back the taking us today. So if we look back the past 80 years of how people have past 80 years of how people have past 80 years of how people have interfaced with machines starting in the interfaced with machines starting in the interfaced with machines starting in the very first days, you know, maybe very first days, you know, maybe very first days, you know, maybe starting 80 or more years ago with punch starting 80 or more years ago with punch starting 80 or more years ago with punch cards and tape and then keyboards and cards and tape and then keyboards and cards and tape and then keyboards and then mice and pointers and touchcreens. then mice and pointers and touchcreens. then mice and pointers and touchcreens. Each generation has made computing Each generation has made computing Each generation has made computing easier for people to access and interact easier for people to access and interact easier for people to access and interact with. But irrespective with. But irrespective with. But irrespective every time it's still the case that the every time it's still the case that the every time it's still the case that the user is having to give their attention user is having to give their attention user is having to give their attention to that device and have the mental to that device and have the mental to that device and have the mental friction of converting the thing they friction of converting the thing they friction of converting the thing they want to get done into steps that the want to get done into steps that the want to get done into steps that the machine understands in a form that it machine understands in a form that it machine understands in a form that it can be inputed into that device. And the can be inputed into that device. And the can be inputed into that device. And the you know the amazing thing about the AI you know the amazing thing about the AI you know the amazing thing about the AI era is it inverts this relationship.
-
era is it inverts this relationship. era is it inverts this relationship. Suddenly the machines understand us. Not Suddenly the machines understand us. Not Suddenly the machines understand us. Not I'm not saying in a conscious way but I'm not saying in a conscious way but I'm not saying in a conscious way but you know they can understand natural you know they can understand natural you know they can understand natural language and they can convert our intent language and they can convert our intent language and they can convert our intent into actions and tool calls to get into actions and tool calls to get into actions and tool calls to get things done. And so um in some ways you things done. And so um in some ways you things done. And so um in some ways you know we're talking about edge inference know we're talking about edge inference know we're talking about edge inference in some ways this isn't new. So there's in some ways this isn't new. So there's in some ways this isn't new. So there's been object detection and CNN's running been object detection and CNN's running been object detection and CNN's running on edge devices for for many years. So on edge devices for for many years. So on edge devices for for many years. So you know envision applications in you know envision applications in you know envision applications in factory inspection or parking bay factory inspection or parking bay factory inspection or parking bay monitoring and all this kind of stuff. monitoring and all this kind of stuff. monitoring and all this kind of stuff. Um more recently embeddings LLMs and Um more recently embeddings LLMs and Um more recently embeddings LLMs and then agentic applications which I guess then agentic applications which I guess then agentic applications which I guess is LLMs that can reason strongly enough is LLMs that can reason strongly enough is LLMs that can reason strongly enough call tools and you know deal with call tools and you know deal with call tools and you know deal with context long enough to act autonomously context long enough to act autonomously context long enough to act autonomously on behalf of the users. These are just on behalf of the users. These are just on behalf of the users. These are just becoming possible today. I I think it's becoming possible today. I I think it's becoming possible today. I I think it's good to step back a little bit and just good to step back a little bit and just good to step back a little bit and just look how fast this has happened. And so look how fast this has happened. And so look how fast this has happened. And so I, you know, I graphed this out and I, you know, I graphed this out and I, you know, I graphed this out and really this is kind of, you know, look really this is kind of, you know, look really this is kind of, you know, look at um GPT3.5 at um GPT3.5 at um GPT3.5 and, you know, that was the first time I and, you know, that was the first time I and, you know, that was the first time I really experienced um what an LLM was really experienced um what an LLM was really experienced um what an LLM was and it was, you know, a mind-blowing and it was, you know, a mind-blowing and it was, you know, a mind-blowing experience. But if you look how quickly experience. But if you look how quickly experience. But if you look how quickly this capability is then being compressed this capability is then being compressed this capability is then being compressed into models with fewer and fewer into models with fewer and fewer into models with fewer and fewer parameters um up until the present day, parameters um up until the present day, parameters um up until the present day, you know, this can now fit in a billion you know, this can now fit in a billion you know, this can now fit in a billion parameters roundabout and then run on a parameters roundabout and then run on a parameters roundabout and then run on a smartphone or run on a Raspberry Pi in smartphone or run on a Raspberry Pi in smartphone or run on a Raspberry Pi in your hand. This is phenomenal. And I by your hand. This is phenomenal. And I by your hand. This is phenomenal. And I by the way, I know that benchmarks are not the way, I know that benchmarks are not the way, I know that benchmarks are not perfect and then an LU probably less so.
-
perfect and then an LU probably less so. perfect and then an LU probably less so. But it's indicative of a trend which is But it's indicative of a trend which is But it's indicative of a trend which is going to continue. You know, if I'm here going to continue. You know, if I'm here going to continue. You know, if I'm here next year, this is going to have gone next year, this is going to have gone next year, this is going to have gone down even further. And it's it's not down even further. And it's it's not down even further. And it's it's not just general knowledge question just general knowledge question just general knowledge question answering, it's reasoning too. Um and so answering, it's reasoning too. Um and so answering, it's reasoning too. Um and so here we're looking at a frontier of you here we're looking at a frontier of you here we're looking at a frontier of you know what can you do with four billion know what can you do with four billion know what can you do with four billion active parameters. So the kind of active parameters. So the kind of active parameters. So the kind of compute on uh today's edge devices and compute on uh today's edge devices and compute on uh today's edge devices and again you know this has gone uh from again you know this has gone uh from again you know this has gone uh from especially with mixture of experts especially with mixture of experts especially with mixture of experts coming into the picture gone from coming into the picture gone from coming into the picture gone from something which is now um exceeding what something which is now um exceeding what something which is now um exceeding what uh Claude Sonnet 3.5 was doing when it uh Claude Sonnet 3.5 was doing when it uh Claude Sonnet 3.5 was doing when it was first released with its frontier was first released with its frontier was first released with its frontier GPQA diamond score. I think this was GPQA diamond score. I think this was GPQA diamond score. I think this was like a year and a half ago. Um this like a year and a half ago. Um this like a year and a half ago. Um this stuff is just moving phenomenally fast. stuff is just moving phenomenally fast. stuff is just moving phenomenally fast. And of course, it's not just about And of course, it's not just about And of course, it's not just about capability. It's about latency when capability. It's about latency when capability. It's about latency when we're talking about human user we're talking about human user we're talking about human user interaction. Um, and there are also interaction. Um, and there are also interaction. Um, and there are also models now which can run comfortably on models now which can run comfortably on models now which can run comfortably on device which do uh speech to text in a device which do uh speech to text in a device which do uh speech to text in a latency which are useful for creating latency which are useful for creating latency which are useful for creating voice UI without any need to send this voice UI without any need to send this voice UI without any need to send this inference to the cloud. And I think it's inference to the cloud. And I think it's inference to the cloud. And I think it's very interesting to consider what drives very interesting to consider what drives very interesting to consider what drives you know these latency requirements. And you know these latency requirements. And you know these latency requirements. And it's really, you know, it's what people it's really, you know, it's what people it's really, you know, it's what people expect, what they're comfortable with is expect, what they're comfortable with is expect, what they're comfortable with is the way we should be designing these the way we should be designing these the way we should be designing these things. I find this graph really things. I find this graph really things. I find this graph really fascinating. So this is an fascinating. So this is an fascinating. So this is an anthropological study of human anthropological study of human anthropological study of human conversation turntaking and the latency conversation turntaking and the latency conversation turntaking and the latency between speakers. And it's interesting between speakers. And it's interesting between speakers. And it's interesting that it's culturally dependent and that it's culturally dependent and that it's culturally dependent and language dependent. So, uh, Japanese is
-
language dependent. So, uh, Japanese is language dependent. So, uh, Japanese is very quick and efficient in responding. very quick and efficient in responding. very quick and efficient in responding. Danish is very laid-back and very slow. Danish is very laid-back and very slow. Danish is very laid-back and very slow. I don't know if that's because the verb I don't know if that's because the verb I don't know if that's because the verb is at the end of the sentence, by the is at the end of the sentence, by the is at the end of the sentence, by the way. Um, and then English is kind of in way. Um, and then English is kind of in way. Um, and then English is kind of in the middle. And so that gives you a kind the middle. And so that gives you a kind the middle. And so that gives you a kind of gauge of of where these things need of gauge of of where these things need of gauge of of where these things need to land. to land. to land. And so, yeah, so I hope you're as And so, yeah, so I hope you're as And so, yeah, so I hope you're as excited as as I am about all of this excited as as I am about all of this excited as as I am about all of this inference being possible on edge inference being possible on edge inference being possible on edge devices. You may say, well, you know, we devices. You may say, well, you know, we devices. You may say, well, you know, we could do this on the data center could do this on the data center could do this on the data center already. Why, you know, why put it at already. Why, you know, why put it at already. Why, you know, why put it at the edge? And it's it's a good question. the edge? And it's it's a good question. the edge? And it's it's a good question. I think there's some, you know, very I think there's some, you know, very I think there's some, you know, very very good reasons for this. Number one very good reasons for this. Number one very good reasons for this. Number one is efficiency. And so you know I think is efficiency. And so you know I think is efficiency. And so you know I think many people are already looking at their many people are already looking at their many people are already looking at their token usage and thinking how how much of token usage and thinking how how much of token usage and thinking how how much of this could I offload onto a local device this could I offload onto a local device this could I offload onto a local device or a device on prem. Um there's the fact or a device on prem. Um there's the fact or a device on prem. Um there's the fact that you can customize this. So maybe that you can customize this. So maybe that you can customize this. So maybe you want to use personal data or you want to use personal data or you want to use personal data or proprietary data to your organization proprietary data to your organization proprietary data to your organization and you would like to keep this uh and you would like to keep this uh and you would like to keep this uh inhouse. Um, and then there's kind of inhouse. Um, and then there's kind of inhouse. Um, and then there's kind of this opportunity for attent a uh sorry, this opportunity for attent a uh sorry, this opportunity for attent a uh sorry, yeah, intentionaware devices where maybe yeah, intentionaware devices where maybe yeah, intentionaware devices where maybe they're observing your speech or have they're observing your speech or have they're observing your speech or have cameras um and are constantly looking cameras um and are constantly looking cameras um and are constantly looking and listening. And a lot of people may and listening. And a lot of people may and listening. And a lot of people may want to choose whether that stuff is want to choose whether that stuff is want to choose whether that stuff is shared off the device onto a data center shared off the device onto a data center shared off the device onto a data center or not. So I think these are some really or not. So I think these are some really or not. So I think these are some really compelling reasons why we would want to compelling reasons why we would want to compelling reasons why we would want to put the inference at the edge. And so I put the inference at the edge. And so I put the inference at the edge. And so I did say, you know, there's something did say, you know, there's something did say, you know, there's something else going on here. And I think it's else going on here. And I think it's else going on here. And I think it's this combination of you've got this this combination of you've got this this combination of you've got this accessible hardware. Um you've got these accessible hardware. Um you've got these accessible hardware. Um you've got these openw weight models putting this openw weight models putting this openw weight models putting this together I think is is a fascinating
-
together I think is is a fascinating together I think is is a fascinating opportunity. And so I don't know how opportunity. And so I don't know how opportunity. And so I don't know how many people are like hardware makers or many people are like hardware makers or many people are like hardware makers or into hardware here. Are there hardware into hardware here. Are there hardware into hardware here. Are there hardware people here? Can I have a show of hands? people here? Can I have a show of hands? people here? Can I have a show of hands? Good. Okay. I'm glad to see that. And Good. Okay. I'm glad to see that. And Good. Okay. I'm glad to see that. And so, you know, it's been notable that for so, you know, it's been notable that for so, you know, it's been notable that for many years now, people have been many years now, people have been many years now, people have been building pretty amazing um hardware building pretty amazing um hardware building pretty amazing um hardware innovations using these composable innovations using these composable innovations using these composable modules from like Raspberry Pi or modules from like Raspberry Pi or modules from like Raspberry Pi or Arduino maybe in the first instance and Arduino maybe in the first instance and Arduino maybe in the first instance and then uh creating totally new designs then uh creating totally new designs then uh creating totally new designs based on this stuff. And I think that based on this stuff. And I think that based on this stuff. And I think that combination of hardware innovation plus combination of hardware innovation plus combination of hardware innovation plus these open models is going to open up these open models is going to open up these open models is going to open up some really interesting potential. Um, I some really interesting potential. Um, I some really interesting potential. Um, I think the observant among you may notice think the observant among you may notice think the observant among you may notice that the micro duck is in the corner that the micro duck is in the corner that the micro duck is in the corner there. And you know, my my favorite there. And you know, my my favorite there. And you know, my my favorite example of this combination because I example of this combination because I example of this combination because I don't think it's really uh as widely um, don't think it's really uh as widely um, don't think it's really uh as widely um, you know, it's not happening yet as you know, it's not happening yet as you know, it's not happening yet as widely as it could do, but a really nice widely as it could do, but a really nice widely as it could do, but a really nice early example of this is the hugging early example of this is the hugging early example of this is the hugging face robotics. Oh, sorry, hugging face face robotics. Oh, sorry, hugging face face robotics. Oh, sorry, hugging face pollen robotics, Reichi Mini. And so I pollen robotics, Reichi Mini. And so I pollen robotics, Reichi Mini. And so I know Reichi Mini is from France. Um, I know Reichi Mini is from France. Um, I know Reichi Mini is from France. Um, I was very happy to be on the beta uh, was very happy to be on the beta uh, was very happy to be on the beta uh, tester program for Reichi Mini. So last tester program for Reichi Mini. So last tester program for Reichi Mini. So last Christmas with my son, I built Reachi Christmas with my son, I built Reachi Christmas with my son, I built Reachi Mini and I was really happy to see. So Mini and I was really happy to see. So Mini and I was really happy to see. So inside of this thing, it has a Stewart inside of this thing, it has a Stewart inside of this thing, it has a Stewart platform so it can express emotions with platform so it can express emotions with platform so it can express emotions with its head and it's got a bunch of servos its head and it's got a bunch of servos its head and it's got a bunch of servos that um move the head platform. So this that um move the head platform. So this that um move the head platform. So this has ARM Cortex M0 in there. They have has ARM Cortex M0 in there. They have has ARM Cortex M0 in there. They have actuators on the antenna. They have a actuators on the antenna. They have a actuators on the antenna. They have a Raspberry Pi compute module 4 for the Raspberry Pi compute module 4 for the Raspberry Pi compute module 4 for the Wi-Fi version. And so with that you can Wi-Fi version. And so with that you can Wi-Fi version. And so with that you can either run some applications on the
-
either run some applications on the either run some applications on the device itself or you can use that to device itself or you can use that to device itself or you can use that to then call inference services off device. then call inference services off device. then call inference services off device. Um and so that's super cool. But it's Um and so that's super cool. But it's Um and so that's super cool. But it's not just hardware. It's also an SDK not just hardware. It's also an SDK not just hardware. It's also an SDK which is open source and it's a which is open source and it's a which is open source and it's a community of developers creating really community of developers creating really community of developers creating really cool applications for this platform. And cool applications for this platform. And cool applications for this platform. And here's one application example I would here's one application example I would here's one application example I would like to highlight. So this is um by our like to highlight. So this is um by our like to highlight. So this is um by our very own Marco Domingo at ARM and he very own Marco Domingo at ARM and he very own Marco Domingo at ARM and he created an app which basically it sits created an app which basically it sits created an app which basically it sits on your desk and if you use your phone on your desk and if you use your phone on your desk and if you use your phone or like start doom scrolling or getting or like start doom scrolling or getting or like start doom scrolling or getting distracted it tells you not to. And this distracted it tells you not to. And this distracted it tells you not to. And this this won a prize. So this won him a VIP this won a prize. So this won him a VIP this won a prize. So this won him a VIP ticket to Nvidia GTC in uh San Jose this ticket to Nvidia GTC in uh San Jose this ticket to Nvidia GTC in uh San Jose this year and a DGX Spark. Um and it was my year and a DGX Spark. Um and it was my year and a DGX Spark. Um and it was my Reichi Mini. I have to say I'm not Reichi Mini. I have to say I'm not Reichi Mini. I have to say I'm not bitter about that but well done Marco. bitter about that but well done Marco. bitter about that but well done Marco. It's a cool app and actually if you go It's a cool app and actually if you go It's a cool app and actually if you go on the app store it is there today on on the app store it is there today on on the app store it is there today on the it's a hugging face space where they the it's a hugging face space where they the it's a hugging face space where they keep reaching many apps. So in the keep reaching many apps. So in the keep reaching many apps. So in the design of these types of ondevice design of these types of ondevice design of these types of ondevice applications you know there's a applications you know there's a applications you know there's a engineering challenge because you have engineering challenge because you have engineering challenge because you have constraints right so you have to think constraints right so you have to think constraints right so you have to think about the available compute and memory about the available compute and memory about the available compute and memory and oftent times you'll have to divide and oftent times you'll have to divide and oftent times you'll have to divide up this task to make best a use of the up this task to make best a use of the up this task to make best a use of the available hardware. So for a vision available hardware. So for a vision available hardware. So for a vision application, it's quite common to maybe application, it's quite common to maybe application, it's quite common to maybe have like a motion gate. Uh maybe this have like a motion gate. Uh maybe this have like a motion gate. Uh maybe this triggers object detection and if you triggers object detection and if you triggers object detection and if you classify things you're interested in, classify things you're interested in, classify things you're interested in, you grab a region of interest, maybe you you grab a region of interest, maybe you you grab a region of interest, maybe you do image uh embeddings and maybe only do image uh embeddings and maybe only do image uh embeddings and maybe only then do you go and then fire off a VLM.
-
then do you go and then fire off a VLM. then do you go and then fire off a VLM. And so these tasks may be carried out by And so these tasks may be carried out by And so these tasks may be carried out by different types of processor on device different types of processor on device different types of processor on device and in some cases may then go off to um and in some cases may then go off to um and in some cases may then go off to um other devices entirely or even to the other devices entirely or even to the other devices entirely or even to the data center to be run. Um and here's data center to be run. Um and here's data center to be run. Um and here's another example. So this is the app I another example. So this is the app I another example. So this is the app I made during the beta test of reachi many made during the beta test of reachi many made during the beta test of reachi many and this is running entirely on a and this is running entirely on a and this is running entirely on a Raspberry Pi 5. What I love about this Raspberry Pi 5. What I love about this Raspberry Pi 5. What I love about this is you you know you say something and is you you know you say something and is you you know you say something and reach will immediately turn his head and reach will immediately turn his head and reach will immediately turn his head and there's something magical about embodied there's something magical about embodied there's something magical about embodied AI in the way that you know a smart AI in the way that you know a smart AI in the way that you know a smart speaker may just bong or the LED lights speaker may just bong or the LED lights speaker may just bong or the LED lights up but having a character turn its head up but having a character turn its head up but having a character turn its head to face you and then respond is a really to face you and then respond is a really to face you and then respond is a really cool experience. And um I think in doing cool experience. And um I think in doing cool experience. And um I think in doing this, you know, I was using a Raspberry this, you know, I was using a Raspberry this, you know, I was using a Raspberry Pi which is definitely capable of Pi which is definitely capable of Pi which is definitely capable of running a bunch of different LLMs, but running a bunch of different LLMs, but running a bunch of different LLMs, but just because of the latency constraint, just because of the latency constraint, just because of the latency constraint, I chose to use embeddings and then do a I chose to use embeddings and then do a I chose to use embeddings and then do a semantic match and call tools against semantic match and call tools against semantic match and call tools against this. So you could still ask it the this. So you could still ask it the this. So you could still ask it the weather or the time and it could weather or the time and it could weather or the time and it could generate using text to speech on device generate using text to speech on device generate using text to speech on device also and reply. also and reply. also and reply. And here's Marco's demo. And by the way, And here's Marco's demo. And by the way, And here's Marco's demo. And by the way, I forgot to say, so in the main expo I forgot to say, so in the main expo I forgot to say, so in the main expo outside after this talk, Marco will be outside after this talk, Marco will be outside after this talk, Marco will be showing this um demo live and you can showing this um demo live and you can showing this um demo live and you can ask him questions about it.
-
ask him questions about it. ask him questions about it. Fundamentally, what we have here is he Fundamentally, what we have here is he Fundamentally, what we have here is he has his DJX Spark, so he's got a more has his DJX Spark, so he's got a more has his DJX Spark, so he's got a more capable device uh from his prize winning capable device uh from his prize winning capable device uh from his prize winning um and he's running um another um and he's running um another um and he's running um another interactive agent like demo. Um this interactive agent like demo. Um this interactive agent like demo. Um this time he's using Mistral Minstrol 3 on time he's using Mistral Minstrol 3 on time he's using Mistral Minstrol 3 on device and he's doing speech to text. device and he's doing speech to text. device and he's doing speech to text. Um, and then he's uh creating responses, Um, and then he's uh creating responses, Um, and then he's uh creating responses, synthesizing a response using text to synthesizing a response using text to synthesizing a response using text to speech and playing it back. In addition speech and playing it back. In addition speech and playing it back. In addition to that, he's doing a bunch of tool to that, he's doing a bunch of tool to that, he's doing a bunch of tool calling. And so he has this ability to, calling. And so he has this ability to, calling. And so he has this ability to, you know, you can ask it what it's you know, you can ask it what it's you know, you can ask it what it's saying. It'll tell you, you know, saying. It'll tell you, you know, saying. It'll tell you, you know, actually, maybe I'll play the video. actually, maybe I'll play the video. actually, maybe I'll play the video. That could be a better way of doing That could be a better way of doing That could be a better way of doing this. Let me just do that really this. Let me just do that really this. Let me just do that really quickly. quickly. quickly. >> Okay. >> Okay. >> Okay. >> Yes. >> Yes. >> Yes. >> Can you greet the room? >> Can you greet the room? >> Can you greet the room? Hi everyone, I'm Reachi. Nice to meet Hi everyone, I'm Reachi. Nice to meet Hi everyone, I'm Reachi. Nice to meet you all. you all. you all. >> I've got a badge here for AI engineer >> I've got a badge here for AI engineer >> I've got a badge here for AI engineer and let's see if Reachi can do anything and let's see if Reachi can do anything and let's see if Reachi can do anything with it. Hey Ree. with it. Hey Ree. with it. Hey Ree. >> Yes.
-
>> Yes. >> Yes. >> Can you meet the attendee? >> Can you meet the attendee? >> Can you meet the attendee? >> Show me your badge. >> Show me your badge. >> Show me your badge. Yes. Got it. Give me a second. Yes. Got it. Give me a second. Yes. Got it. Give me a second. Hi Dominic. Good to see you at the AI Hi Dominic. Good to see you at the AI Hi Dominic. Good to see you at the AI engineer conference. I see you're from engineer conference. I see you're from engineer conference. I see you're from ARM. They're known for designing risk ARM. They're known for designing risk ARM. They're known for designing risk based CPU architectures powering based CPU architectures powering based CPU architectures powering everything from smartphones to everything from smartphones to everything from smartphones to supercomputers. supercomputers. supercomputers. >> Now finally, can you dance? >> How was that? Too much? Want me to dance >> How was that? Too much? Want me to dance again? See you later, everyone. Reachi's again? See you later, everyone. Reachi's again? See you later, everyone. Reachi's off to recharge. Have a great day. off to recharge. Have a great day. off to recharge. Have a great day. >> All right. >> All right. >> All right. >> So, I love the character that um and >> So, I love the character that um and >> So, I love the character that um and personality that Reachi has through personality that Reachi has through personality that Reachi has through minstrel by the way. too much is a minstrel by the way. too much is a minstrel by the way. too much is a really cool uh demo. If you go outside, really cool uh demo. If you go outside, really cool uh demo. If you go outside, like I said, you can actually go and see like I said, you can actually go and see like I said, you can actually go and see if it can read your badge and tell your if it can read your badge and tell your if it can read your badge and tell your fortune, etc. It's really cool to fortune, etc. It's really cool to fortune, etc. It's really cool to actually interact with it in person. So, actually interact with it in person. So, actually interact with it in person. So, those are a couple of examples, but I those are a couple of examples, but I those are a couple of examples, but I did talk at the beginning about intent. did talk at the beginning about intent. did talk at the beginning about intent. And so, you probably heard about the And so, you probably heard about the And so, you probably heard about the intention economy, and you know, the intention economy, and you know, the intention economy, and you know, the attention economy is, you know, you get attention economy is, you know, you get attention economy is, you know, you get people looking at apps and there's people looking at apps and there's people looking at apps and there's adverts and and stuff like that. The adverts and and stuff like that. The adverts and and stuff like that. The intention economy is understanding intention economy is understanding intention economy is understanding um what people want and providing um what people want and providing um what people want and providing solutions to it without them having to solutions to it without them having to solutions to it without them having to spell out everything. I think that's the spell out everything. I think that's the spell out everything. I think that's the way I would describe it. And here's a way I would describe it. And here's a way I would describe it. And here's a little example of an interaction that little example of an interaction that little example of an interaction that could be possible through this intent.
-
could be possible through this intent. could be possible through this intent. So you know if you have for example So you know if you have for example So you know if you have for example reaches a bed light bed night lamp there reaches a bed light bed night lamp there reaches a bed light bed night lamp there um that's also connected to your home um that's also connected to your home um that's also connected to your home assistant system for example um if you assistant system for example um if you assistant system for example um if you ask it you know you tell it you're done ask it you know you tell it you're done ask it you know you tell it you're done for the day it may not be able to for the day it may not be able to for the day it may not be able to interpret that right if it understands interpret that right if it understands interpret that right if it understands your preferences in terms of uh you know your preferences in terms of uh you know your preferences in terms of uh you know AC temperature or lighting it can go and AC temperature or lighting it can go and AC temperature or lighting it can go and act on it if it's got ambient context act on it if it's got ambient context act on it if it's got ambient context and if reachi is maybe looking or and if reachi is maybe looking or and if reachi is maybe looking or listening and remembering has memory um listening and remembering has memory um listening and remembering has memory um about the context it's it's receiving. about the context it's it's receiving. about the context it's it's receiving. It can maybe see, oh, your partner's It can maybe see, oh, your partner's It can maybe see, oh, your partner's asleep. And so maybe Reichi doesn't say asleep. And so maybe Reichi doesn't say asleep. And so maybe Reichi doesn't say anything. It just nods and it kind of anything. It just nods and it kind of anything. It just nods and it kind of dims the lights and it goes to sleep dims the lights and it goes to sleep dims the lights and it goes to sleep itself. And I think, you know, this itself. And I think, you know, this itself. And I think, you know, this ability to deliver a better UX with ability to deliver a better UX with ability to deliver a better UX with fewer interactions with a device is fewer interactions with a device is fewer interactions with a device is really the key to this thing. This is really the key to this thing. This is really the key to this thing. This is just one small example and I think there just one small example and I think there just one small example and I think there are thousands of others. are thousands of others. are thousands of others. And inherent in this is the fact that And inherent in this is the fact that And inherent in this is the fact that you know there there is an agent running you know there there is an agent running you know there there is an agent running here. So we talked about vision and here. So we talked about vision and here. So we talked about vision and audio and sensor streams and maybe audio and sensor streams and maybe audio and sensor streams and maybe integration with with home assistant.
-
integration with with home assistant. integration with with home assistant. These are effectively events or maybe These are effectively events or maybe These are effectively events or maybe state that can be uh queried from an state that can be uh queried from an state that can be uh queried from an agentic system. And of course, you know, agentic system. And of course, you know, agentic system. And of course, you know, with with the agent, there's a harness with with the agent, there's a harness with with the agent, there's a harness which is running on a CPU and that CPU which is running on a CPU and that CPU which is running on a CPU and that CPU is doing a bunch of work managing that is doing a bunch of work managing that is doing a bunch of work managing that context and and the memory. And then I context and and the memory. And then I context and and the memory. And then I think the uh key thing here is how much think the uh key thing here is how much think the uh key thing here is how much of the inference can be done off or off of the inference can be done off or off of the inference can be done off or off of the device. And I think this is like of the device. And I think this is like of the device. And I think this is like a an orchestration problem which a lot a an orchestration problem which a lot a an orchestration problem which a lot of the industry are focusing on today. of the industry are focusing on today. of the industry are focusing on today. Certainly, you know, task aware routting Certainly, you know, task aware routting Certainly, you know, task aware routting and being able to judge whether a task and being able to judge whether a task and being able to judge whether a task is possible to run on a device is is a is possible to run on a device is is a is possible to run on a device is is a really key one. We'll look at that in a really key one. We'll look at that in a really key one. We'll look at that in a second. And finally, you know, we've second. And finally, you know, we've second. And finally, you know, we've been talking a lot about consumerbased been talking a lot about consumerbased been talking a lot about consumerbased applications and how people interact applications and how people interact applications and how people interact with devices. There are a ton of with devices. There are a ton of with devices. There are a ton of embedded systems out there that have embedded systems out there that have embedded systems out there that have computers that actually, you know, maybe computers that actually, you know, maybe computers that actually, you know, maybe they're in industry or smart green they're in industry or smart green they're in industry or smart green houses or, you know, all kinds of houses or, you know, all kinds of houses or, you know, all kinds of commercial building systems where um commercial building systems where um commercial building systems where um they're mainly monitoring sensors or they're mainly monitoring sensors or they're mainly monitoring sensors or maybe doing actuation but with limited maybe doing actuation but with limited maybe doing actuation but with limited uh interaction from a person. And so I uh interaction from a person. And so I uh interaction from a person. And so I really like to think about this idea of really like to think about this idea of really like to think about this idea of of those thousands of embedded systems of those thousands of embedded systems of those thousands of embedded systems that are out there. Is there a potential that are out there. Is there a potential that are out there. Is there a potential for, you know, an agentic embedded for, you know, an agentic embedded for, you know, an agentic embedded system? So I'm not saying that the agent system? So I'm not saying that the agent system? So I'm not saying that the agent carries out all of these tasks more that carries out all of these tasks more that carries out all of these tasks more that the agent takes the role of a supervisor the agent takes the role of a supervisor the agent takes the role of a supervisor and when things happen out of band in and when things happen out of band in and when things happen out of band in otherwise in other words you know if otherwise in other words you know if otherwise in other words you know if there's a fault that the designer of there's a fault that the designer of there's a fault that the designer of this system hadn't anticipated can these this system hadn't anticipated can these this system hadn't anticipated can these agents improvise to go and fix them and
-
agents improvise to go and fix them and agents improvise to go and fix them and the kind of thought experiment uh I the kind of thought experiment uh I the kind of thought experiment uh I think is really interesting is you know think is really interesting is you know think is really interesting is you know I I'm also a space geek right so Voyager I I'm also a space geek right so Voyager I I'm also a space geek right so Voyager one is like I think it's the furthest one is like I think it's the furthest one is like I think it's the furthest man-made object from the earth right now man-made object from the earth right now man-made object from the earth right now and it's left our solar system and it's and it's left our solar system and it's and it's left our solar system and it's um you know there was a fault a couple um you know there was a fault a couple um you know there was a fault a couple of years ago with one of the memory of years ago with one of the memory of years ago with one of the memory chips on this device and NASA were able chips on this device and NASA were able chips on this device and NASA were able to create a patch to re you know reroute to create a patch to re you know reroute to create a patch to re you know reroute this basically and and fix the error but this basically and and fix the error but this basically and and fix the error but they had a roundtrip delay of around 22 they had a roundtrip delay of around 22 they had a roundtrip delay of around 22 and a half hours and it was like and a half hours and it was like and a half hours and it was like something like a in the region of 100 something like a in the region of 100 something like a in the region of 100 bits per second communication to do this bits per second communication to do this bits per second communication to do this and so you know it's interesting to and so you know it's interesting to and so you know it's interesting to consider could agents be used to consider could agents be used to consider could agents be used to preserve these systems systems and preserve these systems systems and preserve these systems systems and monitor them um and keep them in good monitor them um and keep them in good monitor them um and keep them in good health and I think this is one of the health and I think this is one of the health and I think this is one of the really interesting avenues for really interesting avenues for really interesting avenues for exploration exploration exploration on the tip on the topic of hybrid AI um on the tip on the topic of hybrid AI um on the tip on the topic of hybrid AI um so you know there are a ton of so you know there are a ton of so you know there are a ton of interesting uh research papers in this interesting uh research papers in this interesting uh research papers in this area and there's some open source area and there's some open source area and there's some open source projects you know um light LLM or root projects you know um light LLM or root projects you know um light LLM or root LLM um that you can apply to this LLM um that you can apply to this LLM um that you can apply to this challenge of you know can inference run challenge of you know can inference run challenge of you know can inference run on this constrained device in a model on this constrained device in a model on this constrained device in a model that fits there do I need to send it that fits there do I need to send it that fits there do I need to send it externally internally to a more capable externally internally to a more capable externally internally to a more capable piece of hardware. And so I was running piece of hardware. And so I was running piece of hardware. And so I was running some experiments the other day that you some experiments the other day that you some experiments the other day that you know there's a couple of different know there's a couple of different know there's a couple of different strategies. I think that the embedding strategies. I think that the embedding strategies. I think that the embedding kernel is the the quickest because you kernel is the the quickest because you kernel is the the quickest because you don't actually have to run the prompt.
-
don't actually have to run the prompt. don't actually have to run the prompt. You're really just uh you know you have You're really just uh you know you have You're really just uh you know you have some kind of calibration set. You look some kind of calibration set. You look some kind of calibration set. You look at prompts that successfully ran locally at prompts that successfully ran locally at prompts that successfully ran locally and ones that didn't and you see which and ones that didn't and you see which and ones that didn't and you see which one is more similar to your you know one is more similar to your you know one is more similar to your you know prompting question. There's other prompting question. There's other prompting question. There's other solutions out there like calibrated solutions out there like calibrated solutions out there like calibrated uncertainty or even decoder state uncertainty or even decoder state uncertainty or even decoder state probing which are possible once you've probing which are possible once you've probing which are possible once you've actually run inference. So you've actually run inference. So you've actually run inference. So you've already drafted uh a response to the already drafted uh a response to the already drafted uh a response to the prompt and it maybe then you need to prompt and it maybe then you need to prompt and it maybe then you need to kind of defer this off device to be run kind of defer this off device to be run kind of defer this off device to be run um if the confidence is low or if it um if the confidence is low or if it um if the confidence is low or if it classifies as something which would classifies as something which would classifies as something which would would have failed. This is super would have failed. This is super would have failed. This is super interesting stuff I have to say. You interesting stuff I have to say. You interesting stuff I have to say. You know these evals are really just looking know these evals are really just looking know these evals are really just looking at single um single prompt uh eval sets. at single um single prompt uh eval sets. at single um single prompt uh eval sets. there are, you know, there's pinch bench there are, you know, there's pinch bench there are, you know, there's pinch bench uh chlor is it chloral? There's a bunch uh chlor is it chloral? There's a bunch uh chlor is it chloral? There's a bunch of evals looking at multi-step tasks. Um of evals looking at multi-step tasks. Um of evals looking at multi-step tasks. Um I tend to think, you know, it's really I tend to think, you know, it's really I tend to think, you know, it's really your own application which is the best your own application which is the best your own application which is the best thing to consider in these cases. thing to consider in these cases. thing to consider in these cases. So, I've talked a lot about what we can So, I've talked a lot about what we can So, I've talked a lot about what we can do with this combination of really do with this combination of really do with this combination of really efficient models and accessible hardware efficient models and accessible hardware efficient models and accessible hardware and just like the inspiration of and just like the inspiration of and just like the inspiration of creating devices that are a little bit creating devices that are a little bit creating devices that are a little bit more I think you know a more more uh more I think you know a more more uh more I think you know a more more uh considered and less obtrusive. I think considered and less obtrusive. I think considered and less obtrusive. I think this is something that would benefit this is something that would benefit this is something that would benefit everyone. Um but of course you know I everyone. Um but of course you know I everyone. Um but of course you know I should also talk about the hardware should also talk about the hardware should also talk about the hardware beneath a lot of this stuff. And I beneath a lot of this stuff. And I beneath a lot of this stuff. And I mentioned at the beginning I work for mentioned at the beginning I work for mentioned at the beginning I work for ARM and ARM has uh processor solutions ARM and ARM has uh processor solutions ARM and ARM has uh processor solutions that span from these really low power um that span from these really low power um that span from these really low power um wearable or you know sensor node type
-
wearable or you know sensor node type wearable or you know sensor node type devices with our CortexM and ethos up to devices with our CortexM and ethos up to devices with our CortexM and ethos up to Linux-based platforms with Cortex A Linux-based platforms with Cortex A Linux-based platforms with Cortex A which you might find in um in Jetson or which you might find in um in Jetson or which you might find in um in Jetson or in Raspberry Pi right the way up to data in Raspberry Pi right the way up to data in Raspberry Pi right the way up to data center infrastructure and our focus is center infrastructure and our focus is center infrastructure and our focus is providing a common software framework providing a common software framework providing a common software framework and uh we've recently announced an AI and uh we've recently announced an AI and uh we've recently announced an AI portal which has optimized models and portal which has optimized models and portal which has optimized models and also performance optimization tools. And also performance optimization tools. And also performance optimization tools. And one thing, you know, I would definitely one thing, you know, I would definitely one thing, you know, I would definitely encourage you um you know, I'll show encourage you um you know, I'll show encourage you um you know, I'll show this QR code at the end. Uh please this QR code at the end. Uh please this QR code at the end. Uh please register uh become an early access register uh become an early access register uh become an early access customer to this thing. We're interested customer to this thing. We're interested customer to this thing. We're interested to get your feedback and improve this to get your feedback and improve this to get your feedback and improve this for the applications you'd like to for the applications you'd like to for the applications you'd like to develop. And uh just a couple of develop. And uh just a couple of develop. And uh just a couple of examples. So you know, we've upstreamed examples. So you know, we've upstreamed examples. So you know, we've upstreamed uh kernels for Clyde AI into Llama CPP uh kernels for Clyde AI into Llama CPP uh kernels for Clyde AI into Llama CPP into Onyx runtime. So when Gemma 4 came into Onyx runtime. So when Gemma 4 came into Onyx runtime. So when Gemma 4 came out, it ran optimally on SME2. Um, but out, it ran optimally on SME2. Um, but out, it ran optimally on SME2. Um, but it's not only the new stuff. There's it's not only the new stuff. There's it's not only the new stuff. There's also consideration for uh really widely also consideration for uh really widely also consideration for uh really widely used kind of established models and used kind of established models and used kind of established models and hardware. So I really like this example hardware. So I really like this example hardware. So I really like this example which we have just released or due to which we have just released or due to which we have just released or due to release very shortly um which is release very shortly um which is release very shortly um which is Ultralytics YOLO 26 kind of optimized Ultralytics YOLO 26 kind of optimized Ultralytics YOLO 26 kind of optimized specifically for the Raspberry Pi 5 um specifically for the Raspberry Pi 5 um specifically for the Raspberry Pi 5 um ISA. And so they got 1.87 eight, seven ISA. And so they got 1.87 eight, seven ISA. And so they got 1.87 eight, seven times faster inference. I think the key times faster inference. I think the key times faster inference. I think the key thing to point out here is, you know, thing to point out here is, you know, thing to point out here is, you know, typically when you look at the FPS typically when you look at the FPS typically when you look at the FPS numbers for some of these models, numbers for some of these models, numbers for some of these models, they're occupying, if they're running on they're occupying, if they're running on they're occupying, if they're running on CPU, all four cores. And really in an CPU, all four cores. And really in an CPU, all four cores. And really in an application that is not ideal. You would
-
application that is not ideal. You would application that is not ideal. You would like some headroom to do other stuff. like some headroom to do other stuff. like some headroom to do other stuff. And this kind of optimization is the And this kind of optimization is the And this kind of optimization is the difference between, you know, having two difference between, you know, having two difference between, you know, having two cores free to run your application, run cores free to run your application, run cores free to run your application, run other kinds of algorithm on that output. other kinds of algorithm on that output. other kinds of algorithm on that output. So this is really interesting stuff, I So this is really interesting stuff, I So this is really interesting stuff, I think. think. think. Okay. So, if you would like to get more Okay. So, if you would like to get more Okay. So, if you would like to get more resources, there's that QR code again resources, there's that QR code again resources, there's that QR code again for developer.arm.comai. for developer.arm.comai. for developer.arm.comai. Please sign up. Uh, yeah, there's early Please sign up. Uh, yeah, there's early Please sign up. Uh, yeah, there's early access to our optimization tools which I access to our optimization tools which I access to our optimization tools which I would encourage you to get into. Um, would encourage you to get into. Um, would encourage you to get into. Um, also I'm really happy to connect with also I'm really happy to connect with also I'm really happy to connect with you either on LinkedIn or I'll be out by you either on LinkedIn or I'll be out by you either on LinkedIn or I'll be out by reachy later on. And so, yeah, just to reachy later on. And so, yeah, just to reachy later on. And so, yeah, just to say again, so Marco Domingo, I think say again, so Marco Domingo, I think say again, so Marco Domingo, I think he'll be there at 400 pm. So, not he'll be there at 400 pm. So, not he'll be there at 400 pm. So, not directly after this, but at the break. directly after this, but at the break. directly after this, but at the break. uh he'll be on the the expo stage with uh he'll be on the the expo stage with uh he'll be on the the expo stage with Richi Mini if you'd like to go and chat Richi Mini if you'd like to go and chat Richi Mini if you'd like to go and chat with him and interact with Richi. All with him and interact with Richi. All with him and interact with Richi. All right. Thank you very much. Cheers. How you feeling everybody?
-
How you feeling everybody? Yeah. Good. Yeah. Okay. You know what Yeah. Good. Yeah. Okay. You know what Yeah. Good. Yeah. Okay. You know what I'm going to do, right? How you feeling I'm going to do, right? How you feeling I'm going to do, right? How you feeling guys? Okay. I think the next one is guys? Okay. I think the next one is guys? Okay. I think the next one is going to be great. Yeah. Well, thank you going to be great. Yeah. Well, thank you going to be great. Yeah. Well, thank you so much uh uh thank you so much do so much uh uh thank you so much do so much uh uh thank you so much do Dominic for the presentation. Now we're Dominic for the presentation. Now we're Dominic for the presentation. Now we're gonna move on to our next speaker who's gonna move on to our next speaker who's gonna move on to our next speaker who's going to talk to us about agentic map going to talk to us about agentic map going to talk to us about agentic map produce for large scale coding agents. produce for large scale coding agents. produce for large scale coding agents. I'm I'm super curious about this one. I I'm I'm super curious about this one. I I'm I'm super curious about this one. I want to see what what this means. But want to see what what this means. But want to see what what this means. But please join me in welcoming to the stage please join me in welcoming to the stage please join me in welcoming to the stage VP of engineering at Cognition Yanis VP of engineering at Cognition Yanis VP of engineering at Cognition Yanis Sorakis. Okay, Okay, let's see. Oops. Am I connecting to the let's see. Oops. Am I connecting to the let's see. Oops. Am I connecting to the right projector? Classic.
-
Please give me a few seconds to sort it Please give me a few seconds to sort it out. out. out. Uhhuh. That's good. That's good. It's Uhhuh. That's good. That's good. It's Uhhuh. That's good. That's good. It's promising. promising. promising. And I need to see how I extend my And I need to see how I extend my And I need to see how I extend my screen. Ah, slideshow. slideshow. Excellent. Excellent. Excellent. Thank you very much for having me here. Thank you very much for having me here. Thank you very much for having me here. It's such a great pleasure to be It's such a great pleasure to be It's such a great pleasure to be speaking to AI engineering in Paris, speaking to AI engineering in Paris, speaking to AI engineering in Paris, this great audience. Um, very excited this great audience. Um, very excited this great audience. Um, very excited about what I have to show you today. Um, about what I have to show you today. Um, about what I have to show you today. Um, an approach on tackling very large an approach on tackling very large an approach on tackling very large codebasewide problems. codebasewide problems. codebasewide problems. And uh, yeah, let's get going. Um, a few And uh, yeah, let's get going. Um, a few And uh, yeah, let's get going. Um, a few words of intro for us. We are Cognition. words of intro for us. We are Cognition. words of intro for us. We are Cognition. We're an apply AI research lab. Our We're an apply AI research lab. Our We're an apply AI research lab. Our founding DNA is in um AI research, founding DNA is in um AI research, founding DNA is in um AI research, competitive programming, competitive competitive programming, competitive competitive programming, competitive maths. We're based in San Francisco, but maths. We're based in San Francisco, but maths. We're based in San Francisco, but um we are now all around the world um um we are now all around the world um um we are now all around the world um including London. This is our European including London. This is our European including London. This is our European base of operations where I'm based and base of operations where I'm based and base of operations where I'm based and we are the makers of Devon.
-
we are the makers of Devon. we are the makers of Devon. um few so Devon Devon is the first um AI um few so Devon Devon is the first um AI um few so Devon Devon is the first um AI coding agent in the cloud um which now coding agent in the cloud um which now coding agent in the cloud um which now has grown as a platform both in terms of has grown as a platform both in terms of has grown as a platform both in terms of kind of like different form factors um kind of like different form factors um kind of like different form factors um both local and desktop and also in depth both local and desktop and also in depth both local and desktop and also in depth in terms of the different integrations in terms of the different integrations in terms of the different integrations and agentic personas across the SDLC and agentic personas across the SDLC and agentic personas across the SDLC that we provide. that we provide. that we provide. Um we deploy Devon across the most like Um we deploy Devon across the most like Um we deploy Devon across the most like sophisticated um organizations in the sophisticated um organizations in the sophisticated um organizations in the world. They range from um systemically world. They range from um systemically world. They range from um systemically critical banks, Goldman City, Santandere critical banks, Goldman City, Santandere critical banks, Goldman City, Santandere all the way to the US Army, all the way to the US Army, all the way to the US Army, big tech um smaller like a fastmoving big tech um smaller like a fastmoving big tech um smaller like a fastmoving technative startups technative startups technative startups and uh we love tackling like a really and uh we love tackling like a really and uh we love tackling like a really really like difficult uh coding tasks really like difficult uh coding tasks really like difficult uh coding tasks and I want to reflect on our experiences and I want to reflect on our experiences and I want to reflect on our experiences on sort of like the classes of problems on sort of like the classes of problems on sort of like the classes of problems that one uh can See let us consider that one uh can See let us consider that one uh can See let us consider let's start by think two categories okay let's start by think two categories okay let's start by think two categories okay let's call them kind of like localized let's call them kind of like localized let's call them kind of like localized um agentic workloads which are basically um agentic workloads which are basically um agentic workloads which are basically kind of like take your median session kind of like take your median session kind of like take your median session with an agent I want to implement a a with an agent I want to implement a a with an agent I want to implement a a feature or given a particular area of feature or given a particular area of feature or given a particular area of the codebase I want to kind of like the codebase I want to kind of like the codebase I want to kind of like write some tests or augment the existing write some tests or augment the existing write some tests or augment the existing test coverage um given a particular like
-
test coverage um given a particular like test coverage um given a particular like slice of my DB proxy I want to optimize slice of my DB proxy I want to optimize slice of my DB proxy I want to optimize queries and so on and so forth this queries and so on and so forth this queries and so on and so forth this concerns work that's kind like a like a concerns work that's kind like a like a concerns work that's kind like a like a bounded problem it is like a set of bounded problem it is like a set of bounded problem it is like a set of relevant files in the codebase okay that relevant files in the codebase okay that relevant files in the codebase okay that basically modularize the whole logic and basically modularize the whole logic and basically modularize the whole logic and equally importantly if I were to tell an equally importantly if I were to tell an equally importantly if I were to tell an agent what what to do I will either agent what what to do I will either agent what what to do I will either point them explicitly to the files or if point them explicitly to the files or if point them explicitly to the files or if I don't do they can trivial I don't do they can trivial I don't do they can trivial go and find the relevant go and find the relevant go and find the relevant area neighborhood of the codebase to area neighborhood of the codebase to area neighborhood of the codebase to carry out the task. carry out the task. carry out the task. Now Now Now let us consider what I would kind of let us consider what I would kind of let us consider what I would kind of like simply call codebase wide task. And like simply call codebase wide task. And like simply call codebase wide task. And by the way codebase wide doesn't only by the way codebase wide doesn't only by the way codebase wide doesn't only mean like a single repo. You can have mean like a single repo. You can have mean like a single repo. You can have multiple repos. multiple repos. multiple repos. Let's take for example a case that we're Let's take for example a case that we're Let's take for example a case that we're going to revisit again again today like going to revisit again again today like going to revisit again again today like security scanning. what is the security security scanning. what is the security security scanning. what is the security profile of my application?
-
profile of my application? profile of my application? We know that there are like the We know that there are like the We know that there are like the individual security kind of like a individual security kind of like a individual security kind of like a quality levels of different modules of quality levels of different modules of quality levels of different modules of the codebase. But what is equally the codebase. But what is equally the codebase. But what is equally important is the different kind of like important is the different kind of like important is the different kind of like the ways are interplay and how those the ways are interplay and how those the ways are interplay and how those things can evolve in the real world, how things can evolve in the real world, how things can evolve in the real world, how they can connect in order to reveal they can connect in order to reveal they can connect in order to reveal vulnerabilities. vulnerabilities. vulnerabilities. similar in a similar vein extracting similar in a similar vein extracting similar in a similar vein extracting common logic. Um we work with clients common logic. Um we work with clients common logic. Um we work with clients that have like really really large code that have like really really large code that have like really really large code bases. They want to commonize ways they bases. They want to commonize ways they bases. They want to commonize ways they treat arithmetic um financial formulas treat arithmetic um financial formulas treat arithmetic um financial formulas or even kind of like infrastructure or even kind of like infrastructure or even kind of like infrastructure ways of interacting with the database. ways of interacting with the database. ways of interacting with the database. Um they're all sort of like first of all Um they're all sort of like first of all Um they're all sort of like first of all it is a search problem but it's also a it is a search problem but it's also a it is a search problem but it's also a problem of sort of like understanding problem of sort of like understanding problem of sort of like understanding and finding all the different nuances and finding all the different nuances and finding all the different nuances okay by which you need to a common okay by which you need to a common okay by which you need to a common library needs to be extracted need to library needs to be extracted need to library needs to be extracted need to conform to and so on and so forth conform to and so on and so forth conform to and so on and so forth um there's a complex interplay between um there's a complex interplay between um there's a complex interplay between the components take the security example the components take the security example the components take the security example the security quality of my individual the security quality of my individual the security quality of my individual let's say microservices is okay could be let's say microservices is okay could be let's say microservices is okay could be kind of like a you know okay I I have kind of like a you know okay I I have kind of like a you know okay I I have like if I were to look at them in like if I were to look at them in like if I were to look at them in isolation isolation isolation low um priority vulnerabilities they can low um priority vulnerabilities they can low um priority vulnerabilities they can be in the backlog but when chained be in the backlog but when chained be in the backlog but when chained together they can give rise to a together they can give rise to a together they can give rise to a critical exploit critical exploit critical exploit and then therefore like for those tasks and then therefore like for those tasks and then therefore like for those tasks the result is trustworthy when you
-
the result is trustworthy when you the result is trustworthy when you consider the totality of the codebase or consider the totality of the codebase or consider the totality of the codebase or the totality of the deployed application the totality of the deployed application the totality of the deployed application or sets of applications. All right. And or sets of applications. All right. And or sets of applications. All right. And of course, one can say kind of like, of course, one can say kind of like, of course, one can say kind of like, yeah, that's why we have end to end yeah, that's why we have end to end yeah, that's why we have end to end tests. Okay. That's why kind of like in tests. Okay. That's why kind of like in tests. Okay. That's why kind of like in our SDLC, we have the different quality our SDLC, we have the different quality our SDLC, we have the different quality gates and so on and so forth. gates and so on and so forth. gates and so on and so forth. Absolutely. Our goal is always to to Absolutely. Our goal is always to to Absolutely. Our goal is always to to shift left, right, to to catch those shift left, right, to to catch those shift left, right, to to catch those things and be aware of those things as things and be aware of those things as things and be aware of those things as um as soon as possible. um as soon as possible. um as soon as possible. Now if if I were to double click a Now if if I were to double click a Now if if I were to double click a little bit on sort of like uh what the little bit on sort of like uh what the little bit on sort of like uh what the challenges an agent could face in those challenges an agent could face in those challenges an agent could face in those codebase wide tasks codebase wide tasks codebase wide tasks I'm telling can you please find I want I'm telling can you please find I want I'm telling can you please find I want to find instances of a particular to find instances of a particular to find instances of a particular formula of a particular logic that formula of a particular logic that formula of a particular logic that appears in different variations across appears in different variations across appears in different variations across my codebase because I want to extract to my codebase because I want to extract to my codebase because I want to extract to a common library. Okay, there is a a common library. Okay, there is a a common library. Okay, there is a challenge of finding the work in the challenge of finding the work in the challenge of finding the work in the first place. Okay, it is a search first place. Okay, it is a search first place. Okay, it is a search problem. Okay, agents will spend a lot problem. Okay, agents will spend a lot problem. Okay, agents will spend a lot of time searching and backtracking like of time searching and backtracking like of time searching and backtracking like burning tokens. Okay, number two, burning tokens. Okay, number two, burning tokens. Okay, number two, polluting the context. There is a great polluting the context. There is a great polluting the context. There is a great greedy amount that's going to build up greedy amount that's going to build up greedy amount that's going to build up that is going to be competing for that is going to be competing for that is going to be competing for attention. Very very variable and attention. Very very variable and attention. Very very variable and diverse type of context. And number diverse type of context. And number diverse type of context. And number three, what is effectively our stopping three, what is effectively our stopping three, what is effectively our stopping criteria? Okay.
-
criteria? Okay. criteria? Okay. Again, to make this as strong argument Again, to make this as strong argument Again, to make this as strong argument as possible, one will say, "Yeah, of as possible, one will say, "Yeah, of as possible, one will say, "Yeah, of course, we're not going to solve those course, we're not going to solve those course, we're not going to solve those big problems, okay, in a kind of like in big problems, okay, in a kind of like in big problems, okay, in a kind of like in a single session. There will be some a single session. There will be some a single session. There will be some notion of agent orchestration. notion of agent orchestration. notion of agent orchestration. Okay, there will be some notion of okay, Okay, there will be some notion of okay, Okay, there will be some notion of okay, we're going to have kind of like an the we're going to have kind of like an the we're going to have kind of like an the agent outer loop. We're going to break agent outer loop. We're going to break agent outer loop. We're going to break down the problem into different tasks. down the problem into different tasks. down the problem into different tasks. Okay, we're going to define a dag how Okay, we're going to define a dag how Okay, we're going to define a dag how those things need to be completed and those things need to be completed and those things need to be completed and then we're going to have some notion of then we're going to have some notion of then we're going to have some notion of inner loop and the agent. We have to inner loop and the agent. We have to inner loop and the agent. We have to build all that stuff. Okay, because build all that stuff. Okay, because build all that stuff. Okay, because obviously no nobody is going to oneshot obviously no nobody is going to oneshot obviously no nobody is going to oneshot prompt hey kind of like please fix prompt hey kind of like please fix prompt hey kind of like please fix security vulnerabilities. Absolutely. security vulnerabilities. Absolutely. security vulnerabilities. Absolutely. My argument is to today is to present to My argument is to today is to present to My argument is to today is to present to you a framework okay how we think about you a framework okay how we think about you a framework okay how we think about the problem that gives the best um both the problem that gives the best um both the problem that gives the best um both in terms of outcome performance and cost in terms of outcome performance and cost in terms of outcome performance and cost and we make this okay as a first class and we make this okay as a first class and we make this okay as a first class notion in our product notion in our product notion in our product um I touched upon this before kind of um I touched upon this before kind of um I touched upon this before kind of like single agent reasoning I'm not like single agent reasoning I'm not like single agent reasoning I'm not going to labor too much on this point going to labor too much on this point going to labor too much on this point okay both in terms of kind of like time okay both in terms of kind of like time okay both in terms of kind of like time spent searching context pollution and so spent searching context pollution and so spent searching context pollution and so on and so forth on and so forth on and so forth single agent problem does not solve very single agent problem does not solve very single agent problem does not solve very very large codebasewide tasks.
-
very large codebasewide tasks. very large codebasewide tasks. How we approach this? Okay. So um we How we approach this? Okay. So um we How we approach this? Okay. So um we propose okay a variation or an propose okay a variation or an propose okay a variation or an augmentation of the map reduce framework augmentation of the map reduce framework augmentation of the map reduce framework an idea from um distributed systems. How an idea from um distributed systems. How an idea from um distributed systems. How many people here have worked with map many people here have worked with map many people here have worked with map reduce? reduce? reduce? Awesome. So super super quick example Awesome. So super super quick example Awesome. So super super quick example map reduce introduced by Google was a map reduce introduced by Google was a map reduce introduced by Google was a way to solve to perform analysis on like way to solve to perform analysis on like way to solve to perform analysis on like a really really large data sets that a really really large data sets that a really really large data sets that they are distributed that they live on a they are distributed that they live on a they are distributed that they live on a distributed file system. Okay. So you if distributed file system. Okay. So you if distributed file system. Okay. So you if the classic example that we give hey we the classic example that we give hey we the classic example that we give hey we have a huge textual corpus on the have a huge textual corpus on the have a huge textual corpus on the distributed file system um different distributed file system um different distributed file system um different files and we need to do some lexographic files and we need to do some lexographic files and we need to do some lexographic analysis I'll keep it simple count the analysis I'll keep it simple count the analysis I'll keep it simple count the number of words on each book okay will number of words on each book okay will number of words on each book okay will be a series of map functions these are be a series of map functions these are be a series of map functions these are simple stateless as in pure functions simple stateless as in pure functions simple stateless as in pure functions that they kind of like we count each that they kind of like we count each that they kind of like we count each book they can run in parallel and book they can run in parallel and book they can run in parallel and there's a reducer step that consolidates there's a reducer step that consolidates there's a reducer step that consolidates the results.
-
the results. the results. In a very similar vein with a In a very similar vein with a In a very similar vein with a adaptation, we want to port this into adaptation, we want to port this into adaptation, we want to port this into the agentic coding world. Okay, where the agentic coding world. Okay, where the agentic coding world. Okay, where instead of like huge data sets, consider instead of like huge data sets, consider instead of like huge data sets, consider huge code bases. Okay, an overview of huge code bases. Okay, an overview of huge code bases. Okay, an overview of how this thing would work. Okay, how this thing would work. Okay, how this thing would work. Okay, firstly, there are two steps that we firstly, there are two steps that we firstly, there are two steps that we introduce at the beginning. There is a introduce at the beginning. There is a introduce at the beginning. There is a main agent that creates a plan. The plan main agent that creates a plan. The plan main agent that creates a plan. The plan is takes your prompt or takes including is takes your prompt or takes including is takes your prompt or takes including taking skills, taking existing knowledge taking skills, taking existing knowledge taking skills, taking existing knowledge from the repo. Okay, whatever you from the repo. Okay, whatever you from the repo. Okay, whatever you provide and performs an anal and those provide and performs an anal and those provide and performs an anal and those all those considerations in order to all those considerations in order to all those considerations in order to create a series of selectors. Selectors create a series of selectors. Selectors create a series of selectors. Selectors are areas, okay, are deterministic are areas, okay, are deterministic are areas, okay, are deterministic functions, okay, when when you run them, functions, okay, when when you run them, functions, okay, when when you run them, okay, it will create a catalog of your okay, it will create a catalog of your okay, it will create a catalog of your codebase with candidate areas for codebase with candidate areas for codebase with candidate areas for investigation. Number three, for each of those Number three, for each of those candidate areas, we spin a series of candidate areas, we spin a series of candidate areas, we spin a series of parallel sub aents where they look these parallel sub aents where they look these parallel sub aents where they look these areas and only these areas and they areas and only these areas and they areas and only these areas and they perform and then the analysis. This is perform and then the analysis. This is perform and then the analysis. This is the map step and then there is the the map step and then there is the the map step and then there is the reduce step. The main agent again reduce step. The main agent again reduce step. The main agent again considers the results the combined considers the results the combined considers the results the combined results the distilled results from all results the distilled results from all results the distilled results from all the different sub aents. Okay. And the different sub aents. Okay. And the different sub aents. Okay. And performs some reconciliation some performs some reconciliation some performs some reconciliation some consolidation analysis. This this can be
-
consolidation analysis. This this can be consolidation analysis. This this can be duplication prioritization or some more duplication prioritization or some more duplication prioritization or some more richer analysis on the combined result. richer analysis on the combined result. richer analysis on the combined result. Let's kind of like double click a bit on Let's kind of like double click a bit on Let's kind of like double click a bit on those concepts. Okay. And I'm going to those concepts. Okay. And I'm going to those concepts. Okay. And I'm going to use I'm going to use security. Okay. As use I'm going to use security. Okay. As use I'm going to use security. Okay. As kind of like a motivating example. Okay. kind of like a motivating example. Okay. kind of like a motivating example. Okay. I have this large code base. And what is I have this large code base. And what is I have this large code base. And what is the security kind of like the security the security kind of like the security the security kind of like the security quality security level of this codebase. quality security level of this codebase. quality security level of this codebase. Okay. I will provide Okay. I will provide Okay. I will provide the threat model. I will provide any the threat model. I will provide any the threat model. I will provide any additional context on the codebase. additional context on the codebase. additional context on the codebase. Okay. What the pre-work step is going to Okay. What the pre-work step is going to Okay. What the pre-work step is going to do is okay, we're going to write this do is okay, we're going to write this do is okay, we're going to write this determin a deterministic search function determin a deterministic search function determin a deterministic search function that will identify areas of the codebase that will identify areas of the codebase that will identify areas of the codebase of interest in the context of security. of interest in the context of security. of interest in the context of security. You folks can imagine API boundaries um You folks can imagine API boundaries um You folks can imagine API boundaries um where tokens are being minted and be where tokens are being minted and be where tokens are being minted and be passed around um SQL entry uh code entry passed around um SQL entry uh code entry passed around um SQL entry uh code entry points and so on and so forth.
-
points and so on and so forth. points and so on and so forth. The outcome of that is to produce The outcome of that is to produce The outcome of that is to produce basically those deterministic selectors. basically those deterministic selectors. basically those deterministic selectors. They're going to run and they're going They're going to run and they're going They're going to run and they're going to produce like a series of batches, to produce like a series of batches, to produce like a series of batches, okay, or a series of shards of the okay, or a series of shards of the okay, or a series of shards of the codebase where we can then run each one codebase where we can then run each one codebase where we can then run each one of those sub agents. Okay, those habages of those sub agents. Okay, those habages of those sub agents. Okay, those habages then they have a very nice clean and then they have a very nice clean and then they have a very nice clean and bounded context which they can perform bounded context which they can perform bounded context which they can perform the investigations and reports and so on the investigations and reports and so on the investigations and reports and so on and so forth and produce kind of like a and so forth and produce kind of like a and so forth and produce kind of like a bounded outcome for that particular area bounded outcome for that particular area bounded outcome for that particular area of the codebase of the codebase of the codebase and then and then and then the results go back to the main agent the results go back to the main agent the results go back to the main agent where we do kind of like a holistic where we do kind of like a holistic where we do kind of like a holistic reasoning across all the different reasoning across all the different reasoning across all the different findings findings findings in the again in the context of security in the again in the context of security in the again in the context of security to make it like a little bit more to make it like a little bit more to make it like a little bit more concrete. concrete. concrete. probably we're going to duplicate number probably we're going to duplicate number probably we're going to duplicate number one. Number two, there is a one. Number two, there is a one. Number two, there is a prioritization okay of which prioritization okay of which prioritization okay of which vulnerabilities are the most critical vulnerabilities are the most critical vulnerabilities are the most critical and there is a ranking of them. Number and there is a ranking of them. Number and there is a ranking of them. Number three, but this is crucial again in the three, but this is crucial again in the three, but this is crucial again in the context of codebasewide tasks. Um context of codebasewide tasks. Um context of codebasewide tasks. Um presence of chained vulnerabilities.
-
presence of chained vulnerabilities. presence of chained vulnerabilities. Remember individual agents may tell you Remember individual agents may tell you Remember individual agents may tell you that their individual narrow part of the that their individual narrow part of the that their individual narrow part of the codebase has some low um um criticality codebase has some low um um criticality codebase has some low um um criticality finding. But what the reducer step can finding. But what the reducer step can finding. But what the reducer step can do is can assess combinations of those do is can assess combinations of those do is can assess combinations of those low vulnerabil low criticality findings. low vulnerabil low criticality findings. low vulnerabil low criticality findings. Can they produce something much more Can they produce something much more Can they produce something much more serious? serious? serious? And again with the right hardness, okay, And again with the right hardness, okay, And again with the right hardness, okay, um Devon can run, okay, can actually um Devon can run, okay, can actually um Devon can run, okay, can actually test, okay, and try to simulate those test, okay, and try to simulate those test, okay, and try to simulate those attack paths. And to do it kind of like attack paths. And to do it kind of like attack paths. And to do it kind of like a little bit more visually, we treat the a little bit more visually, we treat the a little bit more visually, we treat the tend to treat of code bases as basically tend to treat of code bases as basically tend to treat of code bases as basically a graph. Okay, all there's all those a graph. Okay, all there's all those a graph. Okay, all there's all those references between um the different references between um the different references between um the different parts of the codebase. selectors look parts of the codebase. selectors look parts of the codebase. selectors look for off areas, look for API entry for off areas, look for API entry for off areas, look for API entry points, uh data SQL entry points and points, uh data SQL entry points and points, uh data SQL entry points and then from those selectors that they have then from those selectors that they have then from those selectors that they have run cheaply and deterministic in the run cheaply and deterministic in the run cheaply and deterministic in the codebase, we deploy okay individual codebase, we deploy okay individual codebase, we deploy okay individual devons, individual agents that perform devons, individual agents that perform devons, individual agents that perform security analysis for this particular security analysis for this particular security analysis for this particular narrow bounded path of the codebase narrow bounded path of the codebase narrow bounded path of the codebase and then on the reduce step Devon will and then on the reduce step Devon will and then on the reduce step Devon will produce something like this. Okay, where produce something like this. Okay, where produce something like this. Okay, where it will sort of like number one triage, it will sort of like number one triage, it will sort of like number one triage, number two the duplicate, number three number two the duplicate, number three number two the duplicate, number three run and test and try to recreate those run and test and try to recreate those run and test and try to recreate those vulnerabilities including kind of like vulnerabilities including kind of like vulnerabilities including kind of like false positives as well.
-
false positives as well. false positives as well. Um so yeah mention about reproducing Um so yeah mention about reproducing Um so yeah mention about reproducing because obviously at the end of the day because obviously at the end of the day because obviously at the end of the day you don't want to ship kind of like you don't want to ship kind of like you don't want to ship kind of like tickets because we know whenever we have tickets because we know whenever we have tickets because we know whenever we have a release you know we run second a release you know we run second a release you know we run second vulnerability from tools there are like vulnerability from tools there are like vulnerability from tools there are like tons and tons of things to turn through tons and tons of things to turn through tons and tons of things to turn through we want to be shipping kind of like we want to be shipping kind of like we want to be shipping kind of like legit PRs against a known vulnerability legit PRs against a known vulnerability legit PRs against a known vulnerability so recreating is extremely extremely so recreating is extremely extremely so recreating is extremely extremely important now we're pushing the PRs and important now we're pushing the PRs and important now we're pushing the PRs and just want to comment Okay, a little bit just want to comment Okay, a little bit just want to comment Okay, a little bit on kind of like the relationship between on kind of like the relationship between on kind of like the relationship between cost and performance. Okay, of course cost and performance. Okay, of course cost and performance. Okay, of course kind of like you can run very kind of kind of like you can run very kind of kind of like you can run very kind of like expensive agents. You can run them like expensive agents. You can run them like expensive agents. You can run them against really large code bases. Okay, against really large code bases. Okay, against really large code bases. Okay, it's what is extremely extremely it's what is extremely extremely it's what is extremely extremely important as I mentioned earlier is to important as I mentioned earlier is to important as I mentioned earlier is to think hard and have the best possible think hard and have the best possible think hard and have the best possible ways okay to mitigate the cost of ways okay to mitigate the cost of ways okay to mitigate the cost of searching. Searching is a huge problem searching. Searching is a huge problem searching. Searching is a huge problem in large code bases and context in large code bases and context in large code bases and context pollution and context overflow. This is pollution and context overflow. This is pollution and context overflow. This is solved by the solved by the solved by the Devon and the main agent in the agentic Devon and the main agent in the agentic Devon and the main agent in the agentic map reduce writing those deterministic map reduce writing those deterministic map reduce writing those deterministic selectors and creating lots and lots of selectors and creating lots and lots of selectors and creating lots and lots of banded context which end up being much banded context which end up being much banded context which end up being much cheaper and end up having a better cheaper and end up having a better cheaper and end up having a better recall profile rather uh compared to recall profile rather uh compared to recall profile rather uh compared to competing models. Okay. And we've run competing models. Okay. And we've run competing models. Okay. And we've run this on 50 data sets, 50 vulnerabilities this on 50 data sets, 50 vulnerabilities this on 50 data sets, 50 vulnerabilities from kind of like a GitHub and they run from kind of like a GitHub and they run from kind of like a GitHub and they run across like a multitude of languages,
-
across like a multitude of languages, across like a multitude of languages, Golang, Python, Java and so on and so Golang, Python, Java and so on and so Golang, Python, Java and so on and so forth. Um, forth. Um, forth. Um, another thing that we should think another thing that we should think another thing that we should think about, I talked a lot about security about, I talked a lot about security about, I talked a lot about security agentic map reduce in the context of agentic map reduce in the context of agentic map reduce in the context of vulnerability detection and remediation. vulnerability detection and remediation. vulnerability detection and remediation. Um, think about it about it can be a Um, think about it about it can be a Um, think about it about it can be a generalized concept across all different generalized concept across all different generalized concept across all different codebasewide workloads. Okay, dead code codebasewide workloads. Okay, dead code codebasewide workloads. Okay, dead code removal I mention a lot extracting removal I mention a lot extracting removal I mention a lot extracting common common logic into some common common logic into some common common logic into some centralized library from a very large centralized library from a very large centralized library from a very large code base. Um, dead code detection and code base. Um, dead code detection and code base. Um, dead code detection and so on and so forth. Okay. And this is so on and so forth. Okay. And this is so on and so forth. Okay. And this is something that again we from our something that again we from our something that again we from our research we have kind of like moved it research we have kind of like moved it research we have kind of like moved it to security and we elevating it into a to security and we elevating it into a to security and we elevating it into a first class feature um in our product. first class feature um in our product. first class feature um in our product. Um we are as I mentioned kind of like Um we are as I mentioned kind of like Um we are as I mentioned kind of like earlier we are um applied AI research earlier we are um applied AI research earlier we are um applied AI research lab we doing work on areas like aentic lab we doing work on areas like aentic lab we doing work on areas like aentic map reduce building our own coding map reduce building our own coding map reduce building our own coding models 2.0 know research on model models 2.0 know research on model models 2.0 know research on model harnesses like dev infusion um agentic harnesses like dev infusion um agentic harnesses like dev infusion um agentic infrastructure um please check out our infrastructure um please check out our infrastructure um please check out our research blog and if you have any research blog and if you have any research blog and if you have any questions um this is my email I'm questions um this is my email I'm questions um this is my email I'm looking forward to uh your comments and looking forward to uh your comments and looking forward to uh your comments and see you around in the conference thank see you around in the conference thank see you around in the conference thank you very much thank you so much let's give it up for thank you so much let's give it up for Janice once more please
-
Janice once more please Janice once more please all All right. Okay. So, we all deserve all All right. Okay. So, we all deserve all All right. Okay. So, we all deserve a break right now. I know that it's been a break right now. I know that it's been a break right now. I know that it's been a long day for everybody, but we have a a long day for everybody, but we have a a long day for everybody, but we have a great line of of speakers still. Okay. great line of of speakers still. Okay. great line of of speakers still. Okay. So, we're going to be back here at 4:30, So, we're going to be back here at 4:30, So, we're going to be back here at 4:30, please. And uh enjoy your break. All please. And uh enjoy your break. All please. And uh enjoy your break. All right. See you in a bit.
-
Ladies and gentlemen, please join me in Ladies and gentlemen, please join me in welcoming to the stage your MC for the welcoming to the stage your MC for the welcoming to the stage your MC for the AI engineer Paris 2026 AI engineer Paris 2026 AI engineer Paris 2026 developer relations engineer at Replet. developer relations engineer at Replet. developer relations engineer at Replet. Rahul Chevrey. Okay. All good. How's it going, guys? Okay. All good. How's it going, guys? Hey, this is the last stretch. Okay. So, Hey, this is the last stretch. Okay. So, Hey, this is the last stretch. Okay. So, this is where we need to gather all the this is where we need to gather all the this is where we need to gather all the energy and we get going for the last energy and we get going for the last energy and we get going for the last stretch of this AI engineer uh at Paris. stretch of this AI engineer uh at Paris. stretch of this AI engineer uh at Paris. And uh I'm so excited about our next And uh I'm so excited about our next And uh I'm so excited about our next speaker. So our next speaker actually speaker. So our next speaker actually speaker. So our next speaker actually you don't need me. He doesn't need any you don't need me. He doesn't need any you don't need me. He doesn't need any introduction. He was known as the introduction. He was known as the introduction. He was known as the Typescript wizard and then uh you know Typescript wizard and then uh you know Typescript wizard and then uh you know he was also known for his work at he was also known for his work at he was also known for his work at Verscell before he started talking about Verscell before he started talking about Verscell before he started talking about skills and then it turns out that he has skills and then it turns out that he has skills and then it turns out that he has really strong opinions and good opinions really strong opinions and good opinions really strong opinions and good opinions about AI. So our next speaker is going about AI. So our next speaker is going about AI. So our next speaker is going to talk to you about how to fix your PI to talk to you about how to fix your PI to talk to you about how to fix your PI bottlenecks. And my god, I relate to bottlenecks. And my god, I relate to bottlenecks. And my god, I relate to that so much. Anybody relates to that?
-
that so much. Anybody relates to that? that so much. Anybody relates to that? Seriously, like who's who's having like Seriously, like who's who's having like Seriously, like who's who's having like so many PRs coming their way and then so many PRs coming their way and then so many PRs coming their way and then they they can't review any of them. So, they they can't review any of them. So, they they can't review any of them. So, um, well, without further ado, please um, well, without further ado, please um, well, without further ado, please join me in welcoming to the stage Matt join me in welcoming to the stage Matt join me in welcoming to the stage Matt PCO. Hello, folks. Having a good conference Hello, folks. Having a good conference so far? so far? so far? having a good conference so far. Okay, having a good conference so far. Okay, having a good conference so far. Okay, good. So, I'm here to talk about fixing good. So, I'm here to talk about fixing good. So, I'm here to talk about fixing the PR bottleneck. And this is kind of a the PR bottleneck. And this is kind of a the PR bottleneck. And this is kind of a grand title for trying to fix the thing grand title for trying to fix the thing grand title for trying to fix the thing that most organizations struggle with, I that most organizations struggle with, I that most organizations struggle with, I think, and have kind of historically think, and have kind of historically think, and have kind of historically struggled with before AI. You know, we struggled with before AI. You know, we struggled with before AI. You know, we have always had huge numbers of PRs just have always had huge numbers of PRs just have always had huge numbers of PRs just laying around that no one's bothered to laying around that no one's bothered to laying around that no one's bothered to review. And this has now increased review. And this has now increased review. And this has now increased massively because of the new strains on massively because of the new strains on massively because of the new strains on us because of AI imposing all these us because of AI imposing all these us because of AI imposing all these weird constraints. And to do this, I'm weird constraints. And to do this, I'm weird constraints. And to do this, I'm going to use the rubric of my skills, going to use the rubric of my skills, going to use the rubric of my skills, which is you've kind of heard about, which is you've kind of heard about, which is you've kind of heard about, maybe you've used them. And I have a maybe you've used them. And I have a maybe you've used them. And I have a couple of new skills to announce that couple of new skills to announce that couple of new skills to announce that are going to hopefully improve the way are going to hopefully improve the way are going to hopefully improve the way that you do PRs, improve the way that or that you do PRs, improve the way that or that you do PRs, improve the way that or improve the speed at which you can improve the speed at which you can improve the speed at which you can review them and do them. Speed. Now review them and do them. Speed. Now review them and do them. Speed. Now we're being pushed to do more with less we're being pushed to do more with less we're being pushed to do more with less essentially or more PRs, more work, more essentially or more PRs, more work, more essentially or more PRs, more work, more stuff. And this is kind of the central stuff. And this is kind of the central stuff. And this is kind of the central promise of AI that we're going to be promise of AI that we're going to be promise of AI that we're going to be able to use these agents to scale able to use these agents to scale able to use these agents to scale ourselves up to do more work. And this
-
ourselves up to do more work. And this ourselves up to do more work. And this has resulted in the software factory, has resulted in the software factory, has resulted in the software factory, probably the biggest buzzword of the probably the biggest buzzword of the probably the biggest buzzword of the day. Everyone's talking about software day. Everyone's talking about software day. Everyone's talking about software factory that I chat to. And I think of a factory that I chat to. And I think of a factory that I chat to. And I think of a software factory as primarily something software factory as primarily something software factory as primarily something where instead of the human initiating where instead of the human initiating where instead of the human initiating all of this work, we're going to pass all of this work, we're going to pass all of this work, we're going to pass some of that initiation, some of the some of that initiation, some of the some of that initiation, some of the initiation is going to be done by initiation is going to be done by initiation is going to be done by agents. And I think of their maybe from there you have a classifier maybe from there you have a classifier like Jev come in and Okay. Okay. Let's like Jev come in and Okay. Okay. Let's like Jev come in and Okay. Okay. Let's turn that into a fix or turn that into a turn that into a fix or turn that into a turn that into a fix or turn that into a reproduction or maybe I ping people reproduction or maybe I ping people reproduction or maybe I ping people straight away. And maybe you have other straight away. And maybe you have other straight away. And maybe you have other things. Maybe you have planet scale things. Maybe you have planet scale things. Maybe you have planet scale hooked up. So it gives you query reports hooked up. So it gives you query reports hooked up. So it gives you query reports on the slow queries on your database.
-
on the slow queries on your database. on the slow queries on your database. Maybe that then triggers a different um Maybe that then triggers a different um Maybe that then triggers a different um thing of your software factory. All of thing of your software factory. All of thing of your software factory. All of this is not humans triggering it. It's this is not humans triggering it. It's this is not humans triggering it. It's uh deterministic code triggering it. uh deterministic code triggering it. uh deterministic code triggering it. Right? And so these accelerate your Right? And so these accelerate your Right? And so these accelerate your software factory. They push more code software factory. They push more code software factory. They push more code through it. But then you need breaks, through it. But then you need breaks, through it. But then you need breaks, right? If you just have permanent right? If you just have permanent right? If you just have permanent acceleration pushing stuff through your acceleration pushing stuff through your acceleration pushing stuff through your factory, you're going to end up with a factory, you're going to end up with a factory, you're going to end up with a slop cannon, right? You're just going to slop cannon, right? You're just going to slop cannon, right? You're just going to end up with a ton of slop crappy PRs end up with a ton of slop crappy PRs end up with a ton of slop crappy PRs that you're not going to be able to that you're not going to be able to that you're not going to be able to touch or review or even freaking look touch or review or even freaking look touch or review or even freaking look at. So, you need breaks. These are at. So, you need breaks. These are at. So, you need breaks. These are mechanisms that slow down, that increase mechanisms that slow down, that increase mechanisms that slow down, that increase quality, that make sure that your quality, that make sure that your quality, that make sure that your codebase doesn't turn into a software codebase doesn't turn into a software codebase doesn't turn into a software entropy nightmare because code is the entropy nightmare because code is the entropy nightmare because code is the environment your agent operates in. And environment your agent operates in. And environment your agent operates in. And if you have bad code in your codebase, if you have bad code in your codebase, if you have bad code in your codebase, that is going to beget more bad code. that is going to beget more bad code. that is going to beget more bad code. And so I'm going to talk about these And so I'm going to talk about these And so I'm going to talk about these three breaks in this talk and talk about three breaks in this talk and talk about three breaks in this talk and talk about how we can use them to how we can use them to how we can use them to counterintuitively go faster. So counterintuitively go faster. So counterintuitively go faster. So automated checks, these are the automated checks, these are the automated checks, these are the deterministic checks in your repo that deterministic checks in your repo that deterministic checks in your repo that we've had for thousands of years or you we've had for thousands of years or you we've had for thousands of years or you know since the 50s. Deterministic checks know since the 50s. Deterministic checks know since the 50s. Deterministic checks where we have linting and tests and type where we have linting and tests and type where we have linting and tests and type checking and code quality metrics. All checking and code quality metrics. All checking and code quality metrics. All of these things going together and they of these things going together and they of these things going together and they all work the same every time. Layered on all work the same every time. Layered on all work the same every time. Layered on top of that we have automated review. So top of that we have automated review. So top of that we have automated review. So we have agents who look at our code and we have agents who look at our code and we have agents who look at our code and say okay you know these are for the say okay you know these are for the say okay you know these are for the things that the tests didn't catch or things that the tests didn't catch or things that the tests didn't catch or this is looking at the structure of the this is looking at the structure of the this is looking at the structure of the codebase in general. And then on top of codebase in general. And then on top of codebase in general. And then on top of that the third layer the final layer is
-
that the third layer the final layer is that the third layer the final layer is human review. So people looking at the human review. So people looking at the human review. So people looking at the PR. And these three layers form this PR. And these three layers form this PR. And these three layers form this kind of cake that we end up with when we kind of cake that we end up with when we kind of cake that we end up with when we get to human review. And so more speed get to human review. And so more speed get to human review. And so more speed of course means more PRs. And so the of course means more PRs. And so the of course means more PRs. And so the goal here is to make human review faster goal here is to make human review faster goal here is to make human review faster by leaning on those first two phases. by leaning on those first two phases. by leaning on those first two phases. And we got to stop the slop. That's the And we got to stop the slop. That's the And we got to stop the slop. That's the first principle here, which is if you first principle here, which is if you first principle here, which is if you raise the quality of the code that raise the quality of the code that raise the quality of the code that you're shipping, you're going to end up you're shipping, you're going to end up you're shipping, you're going to end up doing less human review because it's doing less human review because it's doing less human review because it's just going to be better work. And so just going to be better work. And so just going to be better work. And so you're going to end up needing to u make you're going to end up needing to u make you're going to end up needing to u make fewer interventions. fewer interventions. fewer interventions. So automated checks. Now automated So automated checks. Now automated So automated checks. Now automated checks are cheap. That's the cool thing checks are cheap. That's the cool thing checks are cheap. That's the cool thing about them is that they don't cost about them is that they don't cost about them is that they don't cost tokens like automated review does. They tokens like automated review does. They tokens like automated review does. They don't cost human effort. They just cost don't cost human effort. They just cost don't cost human effort. They just cost CPU cycles. So these are for instance, CPU cycles. So these are for instance, CPU cycles. So these are for instance, you know, you run your tests on every you know, you run your tests on every you know, you run your tests on every code change. Maybe those tests do incur code change. Maybe those tests do incur code change. Maybe those tests do incur some tokens because let's say an agent some tokens because let's say an agent some tokens because let's say an agent uh creates a bug and the tests catch it, uh creates a bug and the tests catch it, uh creates a bug and the tests catch it, then you need to spend some tokens to go then you need to spend some tokens to go then you need to spend some tokens to go and fix it, but those are tokens pretty and fix it, but those are tokens pretty and fix it, but those are tokens pretty well spent in my opinion. So checks are well spent in my opinion. So checks are well spent in my opinion. So checks are cheap. That means you can layer on loads cheap. That means you can layer on loads cheap. That means you can layer on loads and loads and loads of them on your and loads and loads of them on your and loads and loads of them on your repos. you're probably not using enough repos. you're probably not using enough repos. you're probably not using enough of them or not being creative enough of them or not being creative enough of them or not being creative enough with your use.
-
with your use. with your use. But checks can lie. But checks can lie. But checks can lie. Does a green CI mean that the code is Does a green CI mean that the code is Does a green CI mean that the code is ready for merge? No, it does not. And so ready for merge? No, it does not. And so ready for merge? No, it does not. And so we've always needed on top of these we've always needed on top of these we've always needed on top of these checks some extra layer to figure out if checks some extra layer to figure out if checks some extra layer to figure out if there's anything catastrophically wrong there's anything catastrophically wrong there's anything catastrophically wrong with the code before we ship it. And so with the code before we ship it. And so with the code before we ship it. And so all of the other phases, the human all of the other phases, the human all of the other phases, the human review and automated review, these are review and automated review, these are review and automated review, these are lie detectors. These are for finding lie detectors. These are for finding lie detectors. These are for finding lies in the automated checks. Now, I lies in the automated checks. Now, I lies in the automated checks. Now, I want to show you some of these lies want to show you some of these lies want to show you some of these lies first of all because this helps when first of all because this helps when first of all because this helps when we're thinking about code and thinking we're thinking about code and thinking we're thinking about code and thinking about automated checks to see how bad it about automated checks to see how bad it about automated checks to see how bad it is and how bad things can get. The first is and how bad things can get. The first is and how bad things can get. The first is tortological tests. A test that just is tortological tests. A test that just is tortological tests. A test that just reasserts the implementation. Opus 5 got reasserts the implementation. Opus 5 got reasserts the implementation. Opus 5 got addicted to these. I don't quite addicted to these. I don't quite addicted to these. I don't quite understand why it would have for understand why it would have for understand why it would have for instance x post character limit equals instance x post character limit equals instance x post character limit equals 280. Can anyone guess the test that was 280. Can anyone guess the test that was 280. Can anyone guess the test that was written to test this behavior? Right? written to test this behavior? Right? written to test this behavior? Right? You've probably seen this a thousand You've probably seen this a thousand You've probably seen this a thousand times. This is real code from agents or times. This is real code from agents or times. This is real code from agents or from stuff that I found my agents doing. from stuff that I found my agents doing. from stuff that I found my agents doing. It said expect x post character limit to It said expect x post character limit to It said expect x post character limit to be 280.
-
be 280. be 280. So the implementation looked like that So the implementation looked like that So the implementation looked like that and the test essentially reasserted the and the test essentially reasserted the and the test essentially reasserted the implementation. That is a tortological implementation. That is a tortological implementation. That is a tortological test. And tological tests are bad test. And tological tests are bad test. And tological tests are bad because they're extremely structure because they're extremely structure because they're extremely structure sensitive. They're very sensitive to the sensitive. They're very sensitive to the sensitive. They're very sensitive to the actual internal workings of the system. actual internal workings of the system. actual internal workings of the system. So it means I cannot change that So it means I cannot change that So it means I cannot change that constant without a test failing. But I constant without a test failing. But I constant without a test failing. But I cannot rename that constant without a cannot rename that constant without a cannot rename that constant without a test failing. Like I have to do is so test failing. Like I have to do is so test failing. Like I have to do is so tied into the structure of my system. tied into the structure of my system. tied into the structure of my system. And I found another one which is even And I found another one which is even And I found another one which is even more egregious I would say. This is more egregious I would say. This is more egregious I would say. This is incredible uh test. What it's doing here incredible uh test. What it's doing here incredible uh test. What it's doing here is it's essentially testing whether two is it's essentially testing whether two is it's essentially testing whether two things in the UI appear in the right things in the UI appear in the right things in the UI appear in the right order. So it's checking that the pitch order. So it's checking that the pitch order. So it's checking that the pitch detail page uh or rather the video detail page uh or rather the video detail page uh or rather the video section comes after the content plan. section comes after the content plan. section comes after the content plan. What it does is it doesn't render it to What it does is it doesn't render it to What it does is it doesn't render it to a screen. It just reads the actual file a screen. It just reads the actual file a screen. It just reads the actual file the module into its own memory and then the module into its own memory and then the module into its own memory and then it's finds the right thing. So, finds it's finds the right thing. So, finds it's finds the right thing. So, finds content plan, finds videos, and then it content plan, finds videos, and then it content plan, finds videos, and then it expects the videos to be after it in the expects the videos to be after it in the expects the videos to be after it in the source material, which is crazy if you source material, which is crazy if you source material, which is crazy if you think about it because I can just like think about it because I can just like think about it because I can just like change the way the source material looks change the way the source material looks change the way the source material looks and this test will fail. It's too and this test will fail. It's too and this test will fail. It's too sensitive to the structure of my sensitive to the structure of my sensitive to the structure of my codebase. So, that's another way that codebase. So, that's another way that codebase. So, that's another way that automated checks can fail. And there are automated checks can fail. And there are automated checks can fail. And there are also or sorry, automated checks can lie.
-
also or sorry, automated checks can lie. also or sorry, automated checks can lie. There are also tests that literally There are also tests that literally There are also tests that literally cannot fail. And this will feel familiar cannot fail. And this will feel familiar cannot fail. And this will feel familiar to you if you've abused mocking in the to you if you've abused mocking in the to you if you've abused mocking in the past or abused various things. For past or abused various things. For past or abused various things. For instance, here we have a use audio boost instance, here we have a use audio boost instance, here we have a use audio boost function. And this use audio boost function. And this use audio boost function. And this use audio boost function uh its internals use the audio function uh its internals use the audio function uh its internals use the audio context API in the DOM. Don't worry if context API in the DOM. Don't worry if context API in the DOM. Don't worry if you don't know any of this, but what you don't know any of this, but what you don't know any of this, but what we're doing here is we're just stubbing we're doing here is we're just stubbing we're doing here is we're just stubbing it out with some fake methods. And it it out with some fake methods. And it it out with some fake methods. And it turns out that audio context has some turns out that audio context has some turns out that audio context has some complicated error modes and it will fail complicated error modes and it will fail complicated error modes and it will fail if you use it under strange conditions. if you use it under strange conditions. if you use it under strange conditions. And so just doing this means our tests And so just doing this means our tests And so just doing this means our tests cannot fail using those modes and we're cannot fail using those modes and we're cannot fail using those modes and we're going to hit strange errors in going to hit strange errors in going to hit strange errors in production that our tests can't fix. And production that our tests can't fix. And production that our tests can't fix. And so the question is then if you can cheat so the question is then if you can cheat so the question is then if you can cheat on automated checks if you know and even on automated checks if you know and even on automated checks if you know and even in good faith ways as well. The the AI in good faith ways as well. The the AI in good faith ways as well. The the AI isn't trying to write bad tests here. isn't trying to write bad tests here. isn't trying to write bad tests here. It's just taking our instructions and It's just taking our instructions and It's just taking our instructions and writing tests that are too tied into the writing tests that are too tied into the writing tests that are too tied into the structure instead of actually executing structure instead of actually executing structure instead of actually executing code. So how do we make automated checks code. So how do we make automated checks code. So how do we make automated checks harder to cheat? If we can do that, then harder to cheat? If we can do that, then harder to cheat? If we can do that, then we can increase the quality of those we can increase the quality of those we can increase the quality of those checks, which means which increases our checks, which means which increases our checks, which means which increases our quality bar.
-
quality bar. quality bar. And the first thing I really like about And the first thing I really like about And the first thing I really like about this is codebased design. So you can this is codebased design. So you can this is codebased design. So you can actually design your way out of these actually design your way out of these actually design your way out of these bad checks. So what does good codebased bad checks. So what does good codebased bad checks. So what does good codebased design look like? I've talked about this design look like? I've talked about this design look like? I've talked about this before in previous AI engineer talks before in previous AI engineer talks before in previous AI engineer talks I've given, which are deep modules. I've given, which are deep modules. I've given, which are deep modules. These are modules that hide complex These are modules that hide complex These are modules that hide complex behavior behind simple interfaces. This behavior behind simple interfaces. This behavior behind simple interfaces. This is a John Asterout idea from the is a John Asterout idea from the is a John Asterout idea from the philosophy of software design. If we philosophy of software design. If we philosophy of software design. If we look at these two modules, we've got A look at these two modules, we've got A look at these two modules, we've got A which has a large implementation hiding which has a large implementation hiding which has a large implementation hiding behind a tiny little interface at the behind a tiny little interface at the behind a tiny little interface at the top. Okay? And then B is a large top. Okay? And then B is a large top. Okay? And then B is a large interface, lots of functions you can interface, lots of functions you can interface, lots of functions you can call and those functions individually call and those functions individually call and those functions individually don't do very much. Does that make don't do very much. Does that make don't do very much. Does that make sense? Yeah. Now, if you have a deep sense? Yeah. Now, if you have a deep sense? Yeah. Now, if you have a deep module like a here, you're going to have module like a here, you're going to have module like a here, you're going to have fewer structure sensitive tests because fewer structure sensitive tests because fewer structure sensitive tests because you're hiding more of the implementation you're hiding more of the implementation you're hiding more of the implementation behind that interface. If it's just behind that interface. If it's just behind that interface. If it's just testing at that interface, you're going testing at that interface, you're going testing at that interface, you're going to get better tests. And so, your job to get better tests. And so, your job to get better tests. And so, your job here is to force the agent to use that here is to force the agent to use that here is to force the agent to use that little interface instead of reaching little interface instead of reaching little interface instead of reaching into the implementation to test these into the implementation to test these into the implementation to test these weird implementation details.
-
weird implementation details. weird implementation details. So, I've got a skill for this. You can So, I've got a skill for this. You can So, I've got a skill for this. You can have like the weirdest vibecoded like have like the weirdest vibecoded like have like the weirdest vibecoded like codebase, the crappiest codebase that codebase, the crappiest codebase that codebase, the crappiest codebase that you've ever set your eyes on and you can you've ever set your eyes on and you can you've ever set your eyes on and you can run this skill on it and it will make it run this skill on it and it will make it run this skill on it and it will make it better. What this does is essentially better. What this does is essentially better. What this does is essentially gives you opportunities for deepening gives you opportunities for deepening gives you opportunities for deepening modules and kind of looks like this. modules and kind of looks like this. modules and kind of looks like this. Raise your hands if you've used this Raise your hands if you've used this Raise your hands if you've used this skill. By the way, not sure how many of skill. By the way, not sure how many of skill. By the way, not sure how many of my folks are in this room. Yeah. Okay. my folks are in this room. Yeah. Okay. my folks are in this room. Yeah. Okay. It's really freaking nice. Essentially It's really freaking nice. Essentially It's really freaking nice. Essentially gives you a HTML document. I'll get out gives you a HTML document. I'll get out gives you a HTML document. I'll get out of the way here of all of the different of the way here of all of the different of the way here of all of the different um potential opportunities. it sees. So um potential opportunities. it sees. So um potential opportunities. it sees. So we can see here we have a before and we we can see here we have a before and we we can see here we have a before and we have an after where we're sort of like have an after where we're sort of like have an after where we're sort of like reducing duplication. We're creating a reducing duplication. We're creating a reducing duplication. We're creating a nice deep testable module and then you nice deep testable module and then you nice deep testable module and then you can go ahead and implement that. And can go ahead and implement that. And can go ahead and implement that. And attached to this, there's also this kind attached to this, there's also this kind attached to this, there's also this kind of language that I've put together for of language that I've put together for of language that I've put together for describing modules. Because like if you describing modules. Because like if you describing modules. Because like if you ever try and read up about how to ever try and read up about how to ever try and read up about how to structure a codebase, you're going to structure a codebase, you're going to structure a codebase, you're going to find 20 different approaches and they're find 20 different approaches and they're find 20 different approaches and they're all going to be called DDD. And like all going to be called DDD. And like all going to be called DDD. And like what you need is a consistent language what you need is a consistent language what you need is a consistent language that you can use in your team to talk that you can use in your team to talk that you can use in your team to talk about this stuff. And so I have a little about this stuff. And so I have a little about this stuff. And so I have a little codebased design skill that defines what codebased design skill that defines what codebased design skill that defines what locality is, defines what leverage is, locality is, defines what leverage is, locality is, defines what leverage is, defines what a seam is. I was using defines what a seam is. I was using defines what a seam is. I was using seams before they were cool. And what seams before they were cool. And what seams before they were cool. And what locality means is kind of how uh well locality means is kind of how uh well locality means is kind of how uh well located together all of the code is. How located together all of the code is. How located together all of the code is. How can you change like a small change in can you change like a small change in can you change like a small change in one module and have it ripple out? And one module and have it ripple out? And one module and have it ripple out? And also leverage is what you get when you also leverage is what you get when you also leverage is what you get when you have a deep module because the caller, have a deep module because the caller, have a deep module because the caller, the person who's actually calling that
-
the person who's actually calling that the person who's actually calling that module gets a lot of value of calling a module gets a lot of value of calling a module gets a lot of value of calling a simple function. Both of those are very simple function. Both of those are very simple function. Both of those are very good in code bases and good for agents good in code bases and good for agents good in code bases and good for agents too, it turns out. But I'm sort of too, it turns out. But I'm sort of too, it turns out. But I'm sort of describing all of these highluting describing all of these highluting describing all of these highluting coding standards. But how do we actually coding standards. But how do we actually coding standards. But how do we actually make sure the agent does them right? How make sure the agent does them right? How make sure the agent does them right? How do you make sure the agent creates deep do you make sure the agent creates deep do you make sure the agent creates deep modules and creates good tests and modules and creates good tests and modules and creates good tests and doesn't write these crap tological ones doesn't write these crap tological ones doesn't write these crap tological ones or structure sensitive tests? or structure sensitive tests? or structure sensitive tests? Well, Well, Well, I do think most people get this wrong. I do think most people get this wrong. I do think most people get this wrong. And And And my first piece of advice is don't put my first piece of advice is don't put my first piece of advice is don't put coding standards in your implement coding standards in your implement coding standards in your implement agent. Okay, let me explain this. If you agent. Okay, let me explain this. If you agent. Okay, let me explain this. If you imagine the implement agent kind of imagine the implement agent kind of imagine the implement agent kind of looks like this where this is all of the looks like this where this is all of the looks like this where this is all of the things the agent needs to be able to do things the agent needs to be able to do things the agent needs to be able to do in its single context window. It needs in its single context window. It needs in its single context window. It needs to be able to explore like to look for to be able to explore like to look for to be able to explore like to look for the code that it's going to change. It the code that it's going to change. It the code that it's going to change. It then needs to actually change it. So then needs to actually change it. So then needs to actually change it. So make the updates to the files in the make the updates to the files in the make the updates to the files in the green and then it needs some budget for green and then it needs some budget for green and then it needs some budget for actually debugging the thing. So for actually debugging the thing. So for actually debugging the thing. So for running those uh automated checks, for running those uh automated checks, for running those uh automated checks, for actually checking and verifying that it actually checking and verifying that it actually checking and verifying that it works.
-
works. works. Now this is quite a lot of work it turns Now this is quite a lot of work it turns Now this is quite a lot of work it turns out. And if you try to impose your out. And if you try to impose your out. And if you try to impose your coding standards on it as well, it's coding standards on it as well, it's coding standards on it as well, it's going to perform worse. going to perform worse. going to perform worse. So implementation is overloaded. That's So implementation is overloaded. That's So implementation is overloaded. That's the mental model I want you to have. And the mental model I want you to have. And the mental model I want you to have. And so can we find a way to impose those so can we find a way to impose those so can we find a way to impose those coding standards in a way that isn't so coding standards in a way that isn't so coding standards in a way that isn't so overloaded? Well, this is my effort. overloaded? Well, this is my effort. overloaded? Well, this is my effort. This is my code review skill. And it This is my code review skill. And it This is my code review skill. And it receives a diff and it reads a file receives a diff and it reads a file receives a diff and it reads a file inside the repository called coding inside the repository called coding inside the repository called coding standards which you can write, you can standards which you can write, you can standards which you can write, you can customize and then it checks if the code customize and then it checks if the code customize and then it checks if the code follows those standards. follows those standards. follows those standards. And so if we look at the reviewer agent And so if we look at the reviewer agent And so if we look at the reviewer agent here, it also crucially runs it in a sub here, it also crucially runs it in a sub here, it also crucially runs it in a sub agent. So it's got its own context agent. So it's got its own context agent. So it's got its own context window to kind of handle here. It's got window to kind of handle here. It's got window to kind of handle here. It's got its own budget. This one it needs to do its own budget. This one it needs to do its own budget. This one it needs to do some exploration, right? Because sure it some exploration, right? Because sure it some exploration, right? Because sure it receives the diff so it knows exactly receives the diff so it knows exactly receives the diff so it knows exactly where it's located where the code is but where it's located where the code is but where it's located where the code is but it should probably do a bit of it should probably do a bit of it should probably do a bit of exploration just so it has the wider exploration just so it has the wider exploration just so it has the wider context understands the code but it context understands the code but it context understands the code but it doesn't need to do any implementation doesn't need to do any implementation doesn't need to do any implementation doesn't need to do any debugging. So doesn't need to do any debugging. So doesn't need to do any debugging. So while implementation is overloaded while implementation is overloaded while implementation is overloaded review is actually underloaded. So it review is actually underloaded. So it review is actually underloaded. So it doesn't have that many jobs to do. This doesn't have that many jobs to do. This doesn't have that many jobs to do. This means you can pile in a bunch of coding means you can pile in a bunch of coding means you can pile in a bunch of coding standards to it and it will do a much standards to it and it will do a much standards to it and it will do a much better job than if you try to do it with better job than if you try to do it with better job than if you try to do it with implement.
-
implement. implement. So, I think of this, and this is So, I think of this, and this is So, I think of this, and this is uncomfortable, right? Because we we all uncomfortable, right? Because we we all uncomfortable, right? Because we we all want to be able to just get good code want to be able to just get good code want to be able to just get good code out the first time, but I think of this out the first time, but I think of this out the first time, but I think of this as the twopart process for writing good as the twopart process for writing good as the twopart process for writing good code, which is implement, you make it code, which is implement, you make it code, which is implement, you make it work, and then code review, you actually work, and then code review, you actually work, and then code review, you actually make it good. You impose your coding make it good. You impose your coding make it good. You impose your coding standards. And for, you know, the retro standards. And for, you know, the retro standards. And for, you know, the retro developers among us, this is essentially developers among us, this is essentially developers among us, this is essentially a red green refactor approach. We use a red green refactor approach. We use a red green refactor approach. We use one context window to make it okay. Do one context window to make it okay. Do one context window to make it okay. Do the red green and then we do another the red green and then we do another the red green and then we do another context window to refactor it. That's at context window to refactor it. That's at context window to refactor it. That's at least how it works in my head and it's least how it works in my head and it's least how it works in my head and it's been very successful for me. This means been very successful for me. This means been very successful for me. This means that when you have coding standards, you that when you have coding standards, you that when you have coding standards, you don't put them in global scope. You don't put them in global scope. You don't put them in global scope. You don't put them in agents.mmd because don't put them in agents.mmd because don't put them in agents.mmd because then they sort of drown out your then they sort of drown out your then they sort of drown out your implement. It may read them, it may not. implement. It may read them, it may not. implement. It may read them, it may not. You put them in coding standards.mmd and You put them in coding standards.mmd and You put them in coding standards.mmd and that way just the code review agent does that way just the code review agent does that way just the code review agent does it. it. it. Now, I've got another idea here, which Now, I've got another idea here, which Now, I've got another idea here, which is I've been talking to lots of people is I've been talking to lots of people is I've been talking to lots of people today. Lots of people saying, you know, today. Lots of people saying, you know, today. Lots of people saying, you know, I've been talking about review and I've been talking about review and I've been talking about review and automated review increased, you know, automated review increased, you know, automated review increased, you know, fixing the PR bottleneck. So many folks fixing the PR bottleneck. So many folks fixing the PR bottleneck. So many folks say, "Oh, yeah, we just use a third say, "Oh, yeah, we just use a third say, "Oh, yeah, we just use a third party service. We use a cursor bug bot.
-
party service. We use a cursor bug bot. party service. We use a cursor bug bot. We use code rabbit or something like We use code rabbit or something like We use code rabbit or something like that." I think that I've I've really that." I think that I've I've really that." I think that I've I've really tried to make a generic code review tried to make a generic code review tried to make a generic code review skill in the past that finds all the skill in the past that finds all the skill in the past that finds all the bugs and does security review and that bugs and does security review and that bugs and does security review and that kind of thing. It turns out it's really kind of thing. It turns out it's really kind of thing. It turns out it's really really hard because you either make it really hard because you either make it really hard because you either make it too general and it just gives you false too general and it just gives you false too general and it just gives you false positives that aren't actually relevant positives that aren't actually relevant positives that aren't actually relevant to your use case or you make it too to your use case or you make it too to your use case or you make it too specific. You say, "Okay, find all the specific. You say, "Okay, find all the specific. You say, "Okay, find all the TypeScript stuff and then Rust people TypeScript stuff and then Rust people TypeScript stuff and then Rust people can't use it." So, I would say don't can't use it." So, I would say don't can't use it." So, I would say don't outsource automated review. Build your outsource automated review. Build your outsource automated review. Build your own. Build up your own coding standards own. Build up your own coding standards own. Build up your own coding standards over time. Share them across your team. over time. Share them across your team. over time. Share them across your team. And if you have an opportunity to impose And if you have an opportunity to impose And if you have an opportunity to impose coding standards, if you've got some coding standards, if you've got some coding standards, if you've got some docs sitting around that no one reads, docs sitting around that no one reads, docs sitting around that no one reads, this is the place to put them in. And this is the place to put them in. And this is the place to put them in. And also when you have this automated also when you have this automated also when you have this automated reviewer, a really natural inclination reviewer, a really natural inclination reviewer, a really natural inclination for lots of people is to say, "Oh yeah, for lots of people is to say, "Oh yeah, for lots of people is to say, "Oh yeah, my code review agent, what it does is it my code review agent, what it does is it my code review agent, what it does is it reads the code and then it comments on reads the code and then it comments on reads the code and then it comments on the PR." the PR." the PR." So what your code review agent is doing So what your code review agent is doing So what your code review agent is doing in that case is it's providing more work in that case is it's providing more work in that case is it's providing more work for the human reviewer. The human for the human reviewer. The human for the human reviewer. The human reviewer then has to read all of these reviewer then has to read all of these reviewer then has to read all of these verbose comments and figure out, okay, verbose comments and figure out, okay, verbose comments and figure out, okay, do we implement this? Do we implement do we implement this? Do we implement do we implement this? Do we implement that? The reviewer should commit. It that? The reviewer should commit. It that? The reviewer should commit. It should make fixes. So, it should should make fixes. So, it should should make fixes. So, it should actually change the things that it finds actually change the things that it finds actually change the things that it finds because then when the human comes because then when the human comes because then when the human comes around, you're reviewing a really nice around, you're reviewing a really nice around, you're reviewing a really nice artifact. If it finds anything that it artifact. If it finds anything that it artifact. If it finds anything that it has any questions over, then of course has any questions over, then of course has any questions over, then of course it can comment, but the default should it can comment, but the default should it can comment, but the default should be commits.
-
be commits. be commits. Stop trying to oneshot good code. Stop Stop trying to oneshot good code. Stop Stop trying to oneshot good code. Stop trying to make the implementer agent the trying to make the implementer agent the trying to make the implementer agent the only thing that you do and go, "Okay, only thing that you do and go, "Okay, only thing that you do and go, "Okay, I'm going to force it to be amazing." It I'm going to force it to be amazing." It I'm going to force it to be amazing." It takes a little bit of, you know, takes a little bit of, you know, takes a little bit of, you know, thinking your way out of there, but once thinking your way out of there, but once thinking your way out of there, but once you realize it, it is fabulous. So, you realize it, it is fabulous. So, you realize it, it is fabulous. So, okay, with all of that process, we've okay, with all of that process, we've okay, with all of that process, we've run our automated checks. We've now run run our automated checks. We've now run run our automated checks. We've now run our automated review to make sure the our automated review to make sure the our automated review to make sure the automated checks aren't lying. How do we automated checks aren't lying. How do we automated checks aren't lying. How do we then maximize the PR's quality in terms then maximize the PR's quality in terms then maximize the PR's quality in terms of human review? How do we get it like of human review? How do we get it like of human review? How do we get it like working the best it can? So, we need a working the best it can? So, we need a working the best it can? So, we need a human friendly PR. And this is a new human friendly PR. And this is a new human friendly PR. And this is a new skill uh coming into the repo which is skill uh coming into the repo which is skill uh coming into the repo which is currently in progress, but I'll be currently in progress, but I'll be currently in progress, but I'll be releasing it soon, which is the PR releasing it soon, which is the PR releasing it soon, which is the PR skill. This is one I've been mulling skill. This is one I've been mulling skill. This is one I've been mulling over for a long, long time. I haven't over for a long, long time. I haven't over for a long, long time. I haven't quite figured out what the quite figured out what the quite figured out what the state-of-the-art is. And I realized the state-of-the-art is. And I realized the state-of-the-art is. And I realized the best way to make a PR skill is just to best way to make a PR skill is just to best way to make a PR skill is just to steal all the best ideas that everyone's steal all the best ideas that everyone's steal all the best ideas that everyone's got. And it's very nice. Now, what does got. And it's very nice. Now, what does got. And it's very nice. Now, what does a good PR body look like? How would you a good PR body look like? How would you a good PR body look like? How would you recreate this skill on your own? Well, recreate this skill on your own? Well, recreate this skill on your own? Well, the first principle is some reviews are the first principle is some reviews are the first principle is some reviews are more important than others. Not every more important than others. Not every more important than others. Not every review is essential, right?
-
review is essential, right? review is essential, right? And once you understand this, you And once you understand this, you And once you understand this, you realize, okay, that means I can focus my realize, okay, that means I can focus my realize, okay, that means I can focus my energy on the really important reviews. energy on the really important reviews. energy on the really important reviews. But how do we categorize that? Well, you But how do we categorize that? Well, you But how do we categorize that? Well, you got to think about the PR as using this got to think about the PR as using this got to think about the PR as using this AWS terminology, which is is it a AWS terminology, which is is it a AWS terminology, which is is it a one-way door or is it a two-way door? one-way door or is it a two-way door? one-way door or is it a two-way door? Now, most PRs that you have will be Now, most PRs that you have will be Now, most PRs that you have will be two-way doors. You can merge the PR and two-way doors. You can merge the PR and two-way doors. You can merge the PR and then always pull it back later. It's the then always pull it back later. It's the then always pull it back later. It's the glorious benefit of being a software glorious benefit of being a software glorious benefit of being a software engineer as opposed to a civil engineer. engineer as opposed to a civil engineer. engineer as opposed to a civil engineer. Right? Mostly when you're a civil Right? Mostly when you're a civil Right? Mostly when you're a civil engineer, it's a it's a one-way door, engineer, it's a it's a one-way door, engineer, it's a it's a one-way door, right? If you get something wrong, that right? If you get something wrong, that right? If you get something wrong, that bridge is going down. But if you have a bridge is going down. But if you have a bridge is going down. But if you have a two-way door, it means that you can two-way door, it means that you can two-way door, it means that you can easily revert the change. Now, that easily revert the change. Now, that easily revert the change. Now, that might be more um might be more um might be more um bit more nuanced than you might expect. bit more nuanced than you might expect. bit more nuanced than you might expect. It might be that a very simple change It might be that a very simple change It might be that a very simple change accidentally blasts out an email to accidentally blasts out an email to accidentally blasts out an email to 60,000 people or something in which case 60,000 people or something in which case 60,000 people or something in which case that is a one-way door. You want to that is a one-way door. You want to that is a one-way door. You want to review that very very carefully. review that very very carefully. review that very very carefully. Involves expensive migrations or data Involves expensive migrations or data Involves expensive migrations or data loss. That is a one-way door. Review the loss. That is a one-way door. Review the loss. That is a one-way door. Review the hell out of that PR. But also tied on to hell out of that PR. But also tied on to hell out of that PR. But also tied on to that we need to understand the blast that we need to understand the blast that we need to understand the blast radius of this PR. What can go wrong?
-
radius of this PR. What can go wrong? radius of this PR. What can go wrong? And if things do go wrong, how bad is And if things do go wrong, how bad is And if things do go wrong, how bad is it? And this means I end up with a nice it? And this means I end up with a nice it? And this means I end up with a nice little sort of summary right at the little sort of summary right at the little sort of summary right at the bottom of all of my PRs, which is the bottom of all of my PRs, which is the bottom of all of my PRs, which is the merge danger. I can see this one is a merge danger. I can see this one is a merge danger. I can see this one is a two-way door. It blast radius is two-way door. It blast radius is two-way door. It blast radius is localized. And so I can see, fantastic. localized. And so I can see, fantastic. localized. And so I can see, fantastic. I don't need to pay that much attention I don't need to pay that much attention I don't need to pay that much attention to this. I'm just going to sort of to this. I'm just going to sort of to this. I'm just going to sort of review it a little. That's really review it a little. That's really review it a little. That's really important. important. important. Next, we need to understand what the PR Next, we need to understand what the PR Next, we need to understand what the PR is even doing, right? And I've tried is even doing, right? And I've tried is even doing, right? And I've tried lots of different ways of figuring this lots of different ways of figuring this lots of different ways of figuring this out, and the best thing I've come up out, and the best thing I've come up out, and the best thing I've come up with is using pseudo code. Now, a huge with is using pseudo code. Now, a huge with is using pseudo code. Now, a huge um point of gratitude here to the show um point of gratitude here to the show um point of gratitude here to the show me skill from the human layer skills me skill from the human layer skills me skill from the human layer skills repo by Dex Holley. This is a phenomenal repo by Dex Holley. This is a phenomenal repo by Dex Holley. This is a phenomenal skill that just essentially dispenses skill that just essentially dispenses skill that just essentially dispenses with most text and shows things to you with most text and shows things to you with most text and shows things to you in images and diagrams instead. This in images and diagrams instead. This in images and diagrams instead. This makes it a lot easier to grasp actually makes it a lot easier to grasp actually makes it a lot easier to grasp actually what's changing and why it's changing. what's changing and why it's changing. what's changing and why it's changing. So you get the kind of standard sort of So you get the kind of standard sort of So you get the kind of standard sort of set of like mermaid diagrams and UML for set of like mermaid diagrams and UML for set of like mermaid diagrams and UML for you know this is a sort of sequence of you know this is a sort of sequence of you know this is a sort of sequence of things that happened. You also just get things that happened. You also just get things that happened. You also just get these lovely simple ones like this for these lovely simple ones like this for these lovely simple ones like this for instance. We can look at this and go instance. We can look at this and go instance. We can look at this and go okay we're working in a CLI. We can see okay we're working in a CLI. We can see okay we're working in a CLI. We can see a new command has been added and we've a new command has been added and we've a new command has been added and we've got two new things little flags up the got two new things little flags up the got two new things little flags up the top here. Just little summaries like top here. Just little summaries like top here. Just little summaries like this. It really does make a massive this. It really does make a massive this. It really does make a massive difference. It's hard to overstate. So difference. It's hard to overstate. So difference. It's hard to overstate. So you're trying to like make understanding you're trying to like make understanding you're trying to like make understanding the why as fast as possible.
-
the why as fast as possible. the why as fast as possible. And I think a third principle here is And I think a third principle here is And I think a third principle here is something you should be thinking about something you should be thinking about something you should be thinking about whenever you do human review because whenever you do human review because whenever you do human review because because we're not doing like because we're not doing like because we're not doing like um because our processes now are so sort um because our processes now are so sort um because our processes now are so sort of streamlined and all we're all of streamlined and all we're all of streamlined and all we're all collaborating around these same skill collaborating around these same skill collaborating around these same skill files, these same steering files. We're files, these same steering files. We're files, these same steering files. We're all building an environment for our all building an environment for our all building an environment for our agents to operate in together. You agents to operate in together. You agents to operate in together. You should think of the process that should think of the process that should think of the process that produces your code as just as important produces your code as just as important produces your code as just as important as the code itself. In other words, when as the code itself. In other words, when as the code itself. In other words, when you do a human review, you're not just you do a human review, you're not just you do a human review, you're not just reviewing the code, you're reviewing the reviewing the code, you're reviewing the reviewing the code, you're reviewing the system that creates it. And the theory system that creates it. And the theory system that creates it. And the theory here is that you never want to write the here is that you never want to write the here is that you never want to write the same comment twice, right? You never same comment twice, right? You never same comment twice, right? You never want to catch the agent doing the same want to catch the agent doing the same want to catch the agent doing the same thing over two PRs. And so what's the thing over two PRs. And so what's the thing over two PRs. And so what's the mechanism by which you can make your mechanism by which you can make your mechanism by which you can make your human review matter? Well, this is a new human review matter? Well, this is a new human review matter? Well, this is a new skill. This is called retro. Retro for skill. This is called retro. Retro for skill. This is called retro. Retro for retrospective. You essentially take a retrospective. You essentially take a retrospective. You essentially take a session that you've done. It can either session that you've done. It can either session that you've done. It can either be like a single um agent session or it be like a single um agent session or it be like a single um agent session or it can be a PR plus the session or you can can be a PR plus the session or you can can be a PR plus the session or you can just get it to look at okay look at all just get it to look at okay look at all just get it to look at okay look at all of the PRs that we've done over the last of the PRs that we've done over the last of the PRs that we've done over the last week all of the reviews pull them all in week all of the reviews pull them all in week all of the reviews pull them all in let's do a retrospective on them and it let's do a retrospective on them and it let's do a retrospective on them and it will suggest automated checks and coding will suggest automated checks and coding will suggest automated checks and coding standards to make the next one better.
-
standards to make the next one better. standards to make the next one better. So this is the compounding effect where So this is the compounding effect where So this is the compounding effect where you essentially by doing human review you essentially by doing human review you essentially by doing human review you're making the quality of the next you're making the quality of the next you're making the quality of the next human review higher and you're sort of human review higher and you're sort of human review higher and you're sort of saving less work from yourself next saving less work from yourself next saving less work from yourself next time. And retro is a really smart skill. time. And retro is a really smart skill. time. And retro is a really smart skill. It adds a bunch of stuff. So obviously It adds a bunch of stuff. So obviously It adds a bunch of stuff. So obviously it suggests automated checks. It it suggests automated checks. It it suggests automated checks. It suggests updates to coding suggests updates to coding suggests updates to coding standards.mmd. It does other smart stuff standards.mmd. It does other smart stuff standards.mmd. It does other smart stuff too. So it looks at navigation pointers. too. So it looks at navigation pointers. too. So it looks at navigation pointers. How easily did the agent find its How easily did the agent find its How easily did the agent find its information? Can we provide a pointer information? Can we provide a pointer information? Can we provide a pointer inside agents.mmd to help it out next inside agents.mmd to help it out next inside agents.mmd to help it out next time? It looks at tool economy. Are time? It looks at tool economy. Are time? It looks at tool economy. Are there different tools that we're using there different tools that we're using there different tools that we're using in the session that can, you know, could in the session that can, you know, could in the session that can, you know, could be made more token efficient? It's be made more token efficient? It's be made more token efficient? It's amazing how many things this catches amazing how many things this catches amazing how many things this catches actually because those are often really actually because those are often really actually because those are often really hard to debug from the outside. It just hard to debug from the outside. It just hard to debug from the outside. It just looks at bloat as well. So, are there looks at bloat as well. So, are there looks at bloat as well. So, are there bloated steering files? Are there bloated steering files? Are there bloated steering files? Are there bloated skills that contribute to these bloated skills that contribute to these bloated skills that contribute to these bad results? Can we make them more bad results? Can we make them more bad results? Can we make them more organized? organized? organized? So that's the goal is to make human So that's the goal is to make human So that's the goal is to make human review faster. We do that by layering up review faster. We do that by layering up review faster. We do that by layering up automated checks. We're layering up automated checks. We're layering up automated checks. We're layering up automated review. And we make the human automated review. And we make the human automated review. And we make the human review as painless, as simple, and as review as painless, as simple, and as review as painless, as simple, and as kind of optional as we need it to. You kind of optional as we need it to. You kind of optional as we need it to. You really don't need to review every single really don't need to review every single really don't need to review every single two-way door. Every single one-way door two-way door. Every single one-way door two-way door. Every single one-way door you do.
-
you do. you do. And so these are my skills. AI And so these are my skills. AI And so these are my skills. AI her.dev/skills. her.dev/skills. her.dev/skills. I'm going to be shipping version 1.3 I'm going to be shipping version 1.3 I'm going to be shipping version 1.3 this week. It has been glorious hanging this week. It has been glorious hanging this week. It has been glorious hanging out with you. It's been a really nice out with you. It's been a really nice out with you. It's been a really nice conference. I'm going to be outside in conference. I'm going to be outside in conference. I'm going to be outside in the lobby if anyone wants to have a the lobby if anyone wants to have a the lobby if anyone wants to have a chat. Uh thank you so much for having chat. Uh thank you so much for having chat. Uh thank you so much for having me. Thank you, Paris. >> Thank you so much. >> Thank you so much. >> All right, let's give it up for Matt >> All right, let's give it up for Matt >> All right, let's give it up for Matt once again. I love that retro skill. I think I'm I love that retro skill. I think I'm going to use it very very soon. All going to use it very very soon. All going to use it very very soon. All right. Okay, so up next is somebody from right. Okay, so up next is somebody from right. Okay, so up next is somebody from Deep Mind U. So we have Martin who's Deep Mind U. So we have Martin who's Deep Mind U. So we have Martin who's going to talk to us about Gemma. He's going to talk to us about Gemma. He's going to talk to us about Gemma. He's going to talk to us about some going to talk to us about some going to talk to us about some architectural fundamentals, but also architectural fundamentals, but also architectural fundamentals, but also what makes uh Gemma so fast and unique. what makes uh Gemma so fast and unique. what makes uh Gemma so fast and unique. So please let's give it up for M for So please let's give it up for M for So please let's give it up for M for Martin.
-
Yay, it works. Yay, it works. Today I get to talk about something that Today I get to talk about something that Today I get to talk about something that I think is a very central component of I think is a very central component of I think is a very central component of releasing open models and that's about releasing open models and that's about releasing open models and that's about efficiency. Because the moment you efficiency. Because the moment you efficiency. Because the moment you release a model out into the wild and release a model out into the wild and release a model out into the wild and people can actually use them, uh that's people can actually use them, uh that's people can actually use them, uh that's the moment that they become constrained. the moment that they become constrained. the moment that they become constrained. Constrained because you know they're Constrained because you know they're Constrained because you know they're being run on your mobile devices on your being run on your mobile devices on your being run on your mobile devices on your laptops and I can kind of assume that laptops and I can kind of assume that laptops and I can kind of assume that there are not data centers that are there are not data centers that are there are not data centers that are running these models. So efficiency running these models. So efficiency running these models. So efficiency becomes a very big part of it. And I becomes a very big part of it. And I becomes a very big part of it. And I would like to focus on a couple of would like to focus on a couple of would like to focus on a couple of perspectives on efficiency. knowledge, perspectives on efficiency. knowledge, perspectives on efficiency. knowledge, size, speed, the complexity of the size, speed, the complexity of the size, speed, the complexity of the architecture, but also the architecture architecture, but also the architecture architecture, but also the architecture in itself. And I want to dive deep into in itself. And I want to dive deep into in itself. And I want to dive deep into some of these components and see what we some of these components and see what we some of these components and see what we can learn from it. One part is per layer can learn from it. One part is per layer can learn from it. One part is per layer embeddings. It's a very big thing in the embeddings. It's a very big thing in the embeddings. It's a very big thing in the smaller models of Gemma 4. There's these smaller models of Gemma 4. There's these smaller models of Gemma 4. There's these two billion and four billion parameter two billion and four billion parameter two billion and four billion parameter models and they use per layer embeddings models and they use per layer embeddings models and they use per layer embeddings heavily. Now they essentially explain heavily. Now they essentially explain heavily. Now they essentially explain what that E actually means in E2B and what that E actually means in E2B and what that E actually means in E2B and E4B. We know the A right in mix of E4B. We know the A right in mix of E4B. We know the A right in mix of experts models they stand for active.
-
experts models they stand for active. experts models they stand for active. But what does the E then stand for? But what does the E then stand for? But what does the E then stand for? Well, let me show you. And this is the Well, let me show you. And this is the Well, let me show you. And this is the architecture that Gemma 4 generally architecture that Gemma 4 generally architecture that Gemma 4 generally uses, right? It has local tension, has uses, right? It has local tension, has uses, right? It has local tension, has global tension, group query tension, global tension, group query tension, global tension, group query tension, there's some normalizations here and there's some normalizations here and there's some normalizations here and there. And then right at the end with there. And then right at the end with there. And then right at the end with the smaller models there are suddenly the smaller models there are suddenly the smaller models there are suddenly per layer embeddings there. And as the per layer embeddings there. And as the per layer embeddings there. And as the name implies these are embeddings per name implies these are embeddings per name implies these are embeddings per tokens but even per layer. So they're tokens but even per layer. So they're tokens but even per layer. So they're much like the token embeddings that you much like the token embeddings that you much like the token embeddings that you know we all know so well but this time know we all know so well but this time know we all know so well but this time they're being added with every layer. they're being added with every layer. they're being added with every layer. That means that the token high has a That means that the token high has a That means that the token high has a very different embedding on layer 1 than very different embedding on layer 1 than very different embedding on layer 1 than it has for instance on layer four or it has for instance on layer four or it has for instance on layer four or layer 10. It's a way to add knowledge to layer 10. It's a way to add knowledge to layer 10. It's a way to add knowledge to the model without having to, you know, the model without having to, you know, the model without having to, you know, store them in memory. store them in memory. store them in memory. During this processing, the model can During this processing, the model can During this processing, the model can actually weigh how these tokens uh are actually weigh how these tokens uh are actually weigh how these tokens uh are being used. So for a given context, the being used. So for a given context, the being used. So for a given context, the model might say, "Okay, I want to focus model might say, "Okay, I want to focus model might say, "Okay, I want to focus a little bit more on this aspect of the a little bit more on this aspect of the a little bit more on this aspect of the embedding rather than that aspect."
-
embedding rather than that aspect." embedding rather than that aspect." And now the great thing about using And now the great thing about using And now the great thing about using something like this is that it's a something like this is that it's a something like this is that it's a lookup table of information about the lookup table of information about the lookup table of information about the tokens. And when you have a lookup tokens. And when you have a lookup tokens. And when you have a lookup table, you don't need to store it in table, you don't need to store it in table, you don't need to store it in VRAM. You don't need to store it in RAM. VRAM. You don't need to store it in RAM. VRAM. You don't need to store it in RAM. You can just store it on whatever You can just store it on whatever You can just store it on whatever storage you have. It's a rackl like storage you have. It's a rackl like storage you have. It's a rackl like behavior that's happening here because behavior that's happening here because behavior that's happening here because during inference, we only need to load during inference, we only need to load during inference, we only need to load in the tokens that are actually needed. in the tokens that are actually needed. in the tokens that are actually needed. And so that explains that the E actually And so that explains that the E actually And so that explains that the E actually stands for effective because that per stands for effective because that per stands for effective because that per layer embedding lookup tables, billions layer embedding lookup tables, billions layer embedding lookup tables, billions of parameters, right? But they're not of parameters, right? But they're not of parameters, right? But they're not being used during inference. we can being used during inference. we can being used during inference. we can separate in a way part of the knowledge separate in a way part of the knowledge separate in a way part of the knowledge from actual model capabilities and this from actual model capabilities and this from actual model capabilities and this this this idea of separating them is this this idea of separating them is this this idea of separating them is becoming more and more pronounced in the becoming more and more pronounced in the becoming more and more pronounced in the field of open models. So we have field of open models. So we have field of open models. So we have knowledge you can do knowledge knowledge you can do knowledge knowledge you can do knowledge efficiently. What about size for efficiently. What about size for efficiently. What about size for instance uh you can create a very small instance uh you can create a very small instance uh you can create a very small model but you still want to have a lot model but you still want to have a lot model but you still want to have a lot of capabilities. So what you can do is of capabilities. So what you can do is of capabilities. So what you can do is you can do quantization aware training. you can do quantization aware training. you can do quantization aware training. you reduce the precision to something you reduce the precision to something you reduce the precision to something smaller while trying to maintain the smaller while trying to maintain the smaller while trying to maintain the capabilities as much as you possibly capabilities as much as you possibly capabilities as much as you possibly can. And that's a very nice way to can. And that's a very nice way to can. And that's a very nice way to approach this. But we can actually take approach this. But we can actually take approach this. But we can actually take it a step further with mobile quants.
-
it a step further with mobile quants. it a step further with mobile quants. And that's a very specific schema that And that's a very specific schema that And that's a very specific schema that was developed for these two smaller was developed for these two smaller was developed for these two smaller models because when you have smaller models because when you have smaller models because when you have smaller models and you're going to quantize them models and you're going to quantize them models and you're going to quantize them to even a smaller size, you have to make to even a smaller size, you have to make to even a smaller size, you have to make sure that you don't lose too much sure that you don't lose too much sure that you don't lose too much performance. So a very specific schema performance. So a very specific schema performance. So a very specific schema and looks a little bit like this. So we and looks a little bit like this. So we and looks a little bit like this. So we have the token embedding layer right you have the token embedding layer right you have the token embedding layer right you can actually reduce that all the way to can actually reduce that all the way to can actually reduce that all the way to two bits and retain a lot of the two bits and retain a lot of the two bits and retain a lot of the capabilities of the model still uh and capabilities of the model still uh and capabilities of the model still uh and now the reason I will come into a little now the reason I will come into a little now the reason I will come into a little bit later bit later bit later for global attention four bits 4bit is for global attention four bits 4bit is for global attention four bits 4bit is what we generally use right it's a it's what we generally use right it's a it's what we generally use right it's a it's a nice balance between the size of the a nice balance between the size of the a nice balance between the size of the model and capabilities model and capabilities model and capabilities but for local attention which you know but for local attention which you know but for local attention which you know zooms in on the thing that we want to zooms in on the thing that we want to zooms in on the thing that we want to know more about it helps to have a know more about it helps to have a know more about it helps to have a little bit of a higher precision little bit of a higher precision little bit of a higher precision Especially with the KV cache feed Especially with the KV cache feed Especially with the KV cache feed forward network is four bits but then forward network is four bits but then forward network is four bits but then the per layer embeddings that we just the per layer embeddings that we just the per layer embeddings that we just talked about those are two bits because talked about those are two bits because talked about those are two bits because they kind of compensate for the lower they kind of compensate for the lower they kind of compensate for the lower position of the token embeddings as position of the token embeddings as position of the token embeddings as well. And so when you combine all of well. And so when you combine all of well. And so when you combine all of this you have a nice schema for a this you have a nice schema for a this you have a nice schema for a specific model. And what you see is that specific model. And what you see is that specific model. And what you see is that making sure that it's efficient also making sure that it's efficient also making sure that it's efficient also means doing a lot of customizations to means doing a lot of customizations to means doing a lot of customizations to these models especially at sizes like these models especially at sizes like these models especially at sizes like two billion four billion parameters.
-
two billion four billion parameters. two billion four billion parameters. And with all of that you have knowledge, And with all of that you have knowledge, And with all of that you have knowledge, you have efficiency in size. What about you have efficiency in size. What about you have efficiency in size. What about speeds? Well, with speeds we use speeds? Well, with speeds we use speeds? Well, with speeds we use something that's becoming more and more something that's becoming more and more something that's becoming more and more popular fortunately these days popular fortunately these days popular fortunately these days speculative decoding with multi-token speculative decoding with multi-token speculative decoding with multi-token prediction. And it works by using a prediction. And it works by using a prediction. And it works by using a rather small model to kind of suggest rather small model to kind of suggest rather small model to kind of suggest tokens that the bigger model, the model tokens that the bigger model, the model tokens that the bigger model, the model you want to actually run, can use you want to actually run, can use you want to actually run, can use because that allows the bigger model to because that allows the bigger model to because that allows the bigger model to then only have to validate all of the then only have to validate all of the then only have to validate all of the tokens that it sees without having to tokens that it sees without having to tokens that it sees without having to process all of those tokens themselves process all of those tokens themselves process all of those tokens themselves at once. So we have a target model, at once. So we have a target model, at once. So we have a target model, right? It's a big one. It can be right? It's a big one. It can be right? It's a big one. It can be whatever architecture that you have whatever architecture that you have whatever architecture that you have and a draft model, a smaller one. And and a draft model, a smaller one. And and a draft model, a smaller one. And with Gemma 4, what we aim to do is do with Gemma 4, what we aim to do is do with Gemma 4, what we aim to do is do some KVK cache sharing. Why? Because some KVK cache sharing. Why? Because some KVK cache sharing. Why? Because that simplifies the process for the that simplifies the process for the that simplifies the process for the smaller model, the candidate model or smaller model, the candidate model or smaller model, the candidate model or the draft model much, much more. the draft model much, much more. the draft model much, much more. Because what happens is that the target Because what happens is that the target Because what happens is that the target model takes in and query, does a single model takes in and query, does a single model takes in and query, does a single round of processing, one single pass, round of processing, one single pass, round of processing, one single pass, generates its KV cache, and then shares generates its KV cache, and then shares generates its KV cache, and then shares it with the smaller model. And the it with the smaller model. And the it with the smaller model. And the smaller model can then use the output of smaller model can then use the output of smaller model can then use the output of that and continue on processing and that and continue on processing and that and continue on processing and generating tokens. It will take the generating tokens. It will take the generating tokens. It will take the input, generate one, two, three, perhaps input, generate one, two, three, perhaps input, generate one, two, three, perhaps eight tokens for the larger model then eight tokens for the larger model then eight tokens for the larger model then to validate. And so the larger model to validate. And so the larger model to validate. And so the larger model only has to look at all of these tokens only has to look at all of these tokens only has to look at all of these tokens at once and decide, okay, I like these at once and decide, okay, I like these at once and decide, okay, I like these tokens, but I don't like so much these
-
tokens, but I don't like so much these tokens, but I don't like so much these tokens. So I'm only going to accept tokens. So I'm only going to accept tokens. So I'm only going to accept these first two, for instance. these first two, for instance. these first two, for instance. But it's still a single pass, right? So But it's still a single pass, right? So But it's still a single pass, right? So all it needs to do is then add another all it needs to do is then add another all it needs to do is then add another token because it can it did the token because it can it did the token because it can it did the computation for the entire sequence computation for the entire sequence computation for the entire sequence anyway anyway anyway and then it generates three tokens and then it generates three tokens and then it generates three tokens almost at the cost of a single one and almost at the cost of a single one and almost at the cost of a single one and that speeds things up tremendously. that speeds things up tremendously. that speeds things up tremendously. Now obviously nothing is a free lunch Now obviously nothing is a free lunch Now obviously nothing is a free lunch right? So there are pros and cons to a right? So there are pros and cons to a right? So there are pros and cons to a method like this. You add a model it's a method like this. You add a model it's a method like this. You add a model it's a very small model but you add it anyway. very small model but you add it anyway. very small model but you add it anyway. It also differs greatly on the use case, It also differs greatly on the use case, It also differs greatly on the use case, the speed up that you get from something the speed up that you get from something the speed up that you get from something like this. The thing is some tokens are like this. The thing is some tokens are like this. The thing is some tokens are super easy to predict, especially in super easy to predict, especially in super easy to predict, especially in structured tasks like code for instance. structured tasks like code for instance. structured tasks like code for instance. But in creative tasks, it's very hard But in creative tasks, it's very hard But in creative tasks, it's very hard for the smaller model to really figure for the smaller model to really figure for the smaller model to really figure out okay, what kind of tokens am I going out okay, what kind of tokens am I going out okay, what kind of tokens am I going to suggest? So more of these tokens will to suggest? So more of these tokens will to suggest? So more of these tokens will then be rejected. Well fortunately there then be rejected. Well fortunately there then be rejected. Well fortunately there are more and more coding task more and are more and more coding task more and are more and more coding task more and more structured task that we see out more structured task that we see out more structured task that we see out there. So for those use cases there. So for those use cases there. So for those use cases multi-token prediction works multi-token prediction works multi-token prediction works exceptionally well exceptionally well exceptionally well and then you have efficiency on let's and then you have efficiency on let's and then you have efficiency on let's see knowledge size speed uh what about see knowledge size speed uh what about see knowledge size speed uh what about complexity? Oh did I see this one here? Let me remove that this one here? Let me remove that server. And then we have complexity
-
server. And then we have complexity server. And then we have complexity because that's also a thing, right? If because that's also a thing, right? If because that's also a thing, right? If you have a model that's not open, it you have a model that's not open, it you have a model that's not open, it really doesn't matter that much how really doesn't matter that much how really doesn't matter that much how complex it is for the end user. But when complex it is for the end user. But when complex it is for the end user. But when you release a model out into the wild you release a model out into the wild you release a model out into the wild for folks to use, uh, let's make sure for folks to use, uh, let's make sure for folks to use, uh, let's make sure it's not too difficult to actually, you it's not too difficult to actually, you it's not too difficult to actually, you know, understand, use, uh, and process. know, understand, use, uh, and process. know, understand, use, uh, and process. And so, as a consequence of that, there And so, as a consequence of that, there And so, as a consequence of that, there was this model that was released called was this model that was released called was this model that was released called an encoderree model. And the encoder an encoderree model. And the encoder an encoderree model. And the encoder free model kind of looked at the free model kind of looked at the free model kind of looked at the encoders for you know processing these encoders for you know processing these encoders for you know processing these multimodal inputs and said that's a lot multimodal inputs and said that's a lot multimodal inputs and said that's a lot of work. What if we remove them? For of work. What if we remove them? For of work. What if we remove them? For many of these models in the Gemma 4 line many of these models in the Gemma 4 line many of these models in the Gemma 4 line you have an audio tokenizer an audio you have an audio tokenizer an audio you have an audio tokenizer an audio encoder and the same for images. encoder and the same for images. encoder and the same for images. And that's great. That works really And that's great. That works really And that's great. That works really well. It's a staple in the field right well. It's a staple in the field right well. It's a staple in the field right to have these encoders for different to have these encoders for different to have these encoders for different multimodal entities but they still take multimodal entities but they still take multimodal entities but they still take up a reasonable size. It reduces time to up a reasonable size. It reduces time to up a reasonable size. It reduces time to first token and it's something else to first token and it's something else to first token and it's something else to consider when for example fine-tuning consider when for example fine-tuning consider when for example fine-tuning your model your model your model with the encoder free model. They try to with the encoder free model. They try to with the encoder free model. They try to just remove them entirely and put a lot just remove them entirely and put a lot just remove them entirely and put a lot of the burden of processing these of the burden of processing these of the burden of processing these multimodal inputs onto the model rather multimodal inputs onto the model rather multimodal inputs onto the model rather than these encoders. For audio, it works than these encoders. For audio, it works than these encoders. For audio, it works a little bit like this. You have an a little bit like this. You have an a little bit like this. You have an input, a sequence of amplitude values input, a sequence of amplitude values input, a sequence of amplitude values and the the only thing that happens is and the the only thing that happens is and the the only thing that happens is they cut it up into pieces, make it a they cut it up into pieces, make it a they cut it up into pieces, make it a sequence of sequence of sequence of well audio tokens and then project it well audio tokens and then project it well audio tokens and then project it onto the dimensionality of the model
-
onto the dimensionality of the model onto the dimensionality of the model itself and that's it. So the only itself and that's it. So the only itself and that's it. So the only parameters that you have are kind of in parameters that you have are kind of in parameters that you have are kind of in the linear projection, but that's about the linear projection, but that's about the linear projection, but that's about it. And then all of the contextual stuff it. And then all of the contextual stuff it. And then all of the contextual stuff that was happening in the audio encoder that was happening in the audio encoder that was happening in the audio encoder before is now being handled by the LLM before is now being handled by the LLM before is now being handled by the LLM itself. So you're kind of shifting all itself. So you're kind of shifting all itself. So you're kind of shifting all of these parameters from the encoder to of these parameters from the encoder to of these parameters from the encoder to the LLM. the LLM. the LLM. The same can be done with the image The same can be done with the image The same can be done with the image encoder. encoder. encoder. The one problem though is with images is The one problem though is with images is The one problem though is with images is 3D information. We can't just cut it up 3D information. We can't just cut it up 3D information. We can't just cut it up into pieces, make a sequence out of it, into pieces, make a sequence out of it, into pieces, make a sequence out of it, and just hope the model processes it and just hope the model processes it and just hope the model processes it well because it doesn't. I I can promise well because it doesn't. I I can promise well because it doesn't. I I can promise you that. So what we need for that are you that. So what we need for that are you that. So what we need for that are positional information. So instead of positional information. So instead of positional information. So instead of doing a a direct projection from the doing a a direct projection from the doing a a direct projection from the patches onto the model, we take those patches onto the model, we take those patches onto the model, we take those patches and provide it with additional patches and provide it with additional patches and provide it with additional positional information. So it knows that positional information. So it knows that positional information. So it knows that well let's say patch four that small well let's say patch four that small well let's say patch four that small image token um has actually some meaning image token um has actually some meaning image token um has actually some meaning to it because you can imagine if you to it because you can imagine if you to it because you can imagine if you have a very wide image position four have a very wide image position four have a very wide image position four means something very differently than if means something very differently than if means something very differently than if you have very tall image.
-
you have very tall image. you have very tall image. And so what happens is you just use the And so what happens is you just use the And so what happens is you just use the positional embeddings uh projected and positional embeddings uh projected and positional embeddings uh projected and what you have again is an encoder that's what you have again is an encoder that's what you have again is an encoder that's entirely removed. entirely removed. entirely removed. What's so nice about something like this What's so nice about something like this What's so nice about something like this is that it reduces time to first token is that it reduces time to first token is that it reduces time to first token that very much helps. you put all of the that very much helps. you put all of the that very much helps. you put all of the capabilities onto the model capabilities onto the model capabilities onto the model and you reduce uh a big portion of the and you reduce uh a big portion of the and you reduce uh a big portion of the parameters because especially at smaller parameters because especially at smaller parameters because especially at smaller sizes something like 300 million sizes something like 300 million sizes something like 300 million parameters or 500 million parameters it parameters or 500 million parameters it parameters or 500 million parameters it doesn't sound large when we're talking doesn't sound large when we're talking doesn't sound large when we're talking about hundreds of billions of mod of about hundreds of billions of mod of about hundreds of billions of mod of parameters for some of these models but parameters for some of these models but parameters for some of these models but at two billion sizes or 12 billion this at two billion sizes or 12 billion this at two billion sizes or 12 billion this makes a big big difference especially if makes a big big difference especially if makes a big big difference especially if you can use it for the capabilities of you can use it for the capabilities of you can use it for the capabilities of the the the And so we have knowledge, we have size, And so we have knowledge, we have size, And so we have knowledge, we have size, we have speed, we have complexity, all we have speed, we have complexity, all we have speed, we have complexity, all different ways to focus on efficiency. different ways to focus on efficiency. different ways to focus on efficiency. And there was one last thing that I And there was one last thing that I And there was one last thing that I really wanted to focus on and you know really wanted to focus on and you know really wanted to focus on and you know explain a little bit more about what it explain a little bit more about what it explain a little bit more about what it means and how it works and that's means and how it works and that's means and how it works and that's architecture.
-
architecture. architecture. Regular models, regular LLMs are all Regular models, regular LLMs are all Regular models, regular LLMs are all memory bound and those have advantages memory bound and those have advantages memory bound and those have advantages and disadvantages and disadvantages and disadvantages and they're great. They're used for many and they're great. They're used for many and they're great. They're used for many different purposes and and many different purposes and and many different purposes and and many different fields. different fields. different fields. What if we make them computebound What if we make them computebound What if we make them computebound instead? What if we flip the narrative instead? What if we flip the narrative instead? What if we flip the narrative 180 degrees and try to approach it from 180 degrees and try to approach it from 180 degrees and try to approach it from a very different perspective? And so a very different perspective? And so a very different perspective? And so what you get is not auto regression, what you get is not auto regression, what you get is not auto regression, it's something entirely different. it's something entirely different. it's something entirely different. Diffusion. Diffusion. Diffusion. Can we use diffusion for LLMs and still Can we use diffusion for LLMs and still Can we use diffusion for LLMs and still have it being meaningful? have it used have it being meaningful? have it used have it being meaningful? have it used in a way that you know a lot of people in a way that you know a lot of people in a way that you know a lot of people can still use it but for very different can still use it but for very different can still use it but for very different use cases. Now let me go through a use cases. Now let me go through a use cases. Now let me go through a little bit about about what diffusion little bit about about what diffusion little bit about about what diffusion means in large language models uh means in large language models uh means in large language models uh especially because we will come to a especially because we will come to a especially because we will come to a subject that I think a lot of folks will subject that I think a lot of folks will subject that I think a lot of folks will have seen in the last couple of weeks. have seen in the last couple of weeks. have seen in the last couple of weeks. You start with an input, right? We have You start with an input, right? We have You start with an input, right? We have a query, an input that's being a query, an input that's being a query, an input that's being processed. And normally one model would processed. And normally one model would processed. And normally one model would do that and now kind of the same model do that and now kind of the same model do that and now kind of the same model does it. But what happens in the fusion, does it. But what happens in the fusion, does it. But what happens in the fusion, we have something called an encoder. And we have something called an encoder. And we have something called an encoder. And the encoder in the fusion gemma at least the encoder in the fusion gemma at least the encoder in the fusion gemma at least is still just a decoder with causal is still just a decoder with causal is still just a decoder with causal attention. It's a fine-tuned Gemma 4 attention. It's a fine-tuned Gemma 4 attention. It's a fine-tuned Gemma 4 model. No crazy different architecture model. No crazy different architecture model. No crazy different architecture or what have you. It's, you know, one of or what have you. It's, you know, one of or what have you. It's, you know, one of the models you saw before. The reason the models you saw before. The reason the models you saw before. The reason why it's called an encoder though why it's called an encoder though why it's called an encoder though because it's meant for processing the because it's meant for processing the because it's meant for processing the prompt and so it generates a KV cache
-
prompt and so it generates a KV cache prompt and so it generates a KV cache the context that we need to then build the context that we need to then build the context that we need to then build upon and slowly and more surely do upon and slowly and more surely do upon and slowly and more surely do inference. inference. inference. So we start with the KV cache but do to So we start with the KV cache but do to So we start with the KV cache but do to do the actual processing to actually do the actual processing to actually do the actual processing to actually generate tokens in the fusion you don't generate tokens in the fusion you don't generate tokens in the fusion you don't start from scratch. You don't just you start from scratch. You don't just you start from scratch. You don't just you know here's one token and here's the know here's one token and here's the know here's one token and here's the second and here's the third. Now you second and here's the third. Now you second and here's the third. Now you start with a noisy con canvas. It's start with a noisy con canvas. It's start with a noisy con canvas. It's typically sized around 256 typically sized around 256 typically sized around 256 tokens and then the job of the den tokens and then the job of the den tokens and then the job of the den noiser will be to actually process that noiser will be to actually process that noiser will be to actually process that and iteratively fill it up with tokens and iteratively fill it up with tokens and iteratively fill it up with tokens that it's very confident about. that it's very confident about. that it's very confident about. It is the exact same model though. It It is the exact same model though. It It is the exact same model though. It uses the exact same tight weights gem uses the exact same tight weights gem uses the exact same tight weights gem for 26B model. There's one difference for 26B model. There's one difference for 26B model. There's one difference though. If you have a large canvas and though. If you have a large canvas and though. If you have a large canvas and you want to update the tokens you want to update the tokens you want to update the tokens simultaneously at the beginning and at simultaneously at the beginning and at simultaneously at the beginning and at the end, you kind of need a different the end, you kind of need a different the end, you kind of need a different type of attention. So when it's in the type of attention. So when it's in the type of attention. So when it's in the denoiser mode, it uses birectional denoiser mode, it uses birectional denoiser mode, it uses birectional attention instead. So it can lock both attention instead. So it can lock both attention instead. So it can lock both directions.
-
directions. directions. What happens after that is it gets the What happens after that is it gets the What happens after that is it gets the KV cache of the encoder because again it KV cache of the encoder because again it KV cache of the encoder because again it was meant to process the state and get was meant to process the state and get was meant to process the state and get it ready for the deninoiser to do its it ready for the deninoiser to do its it ready for the deninoiser to do its work. And then the deninoiser can well work. And then the deninoiser can well work. And then the deninoiser can well predict or suggest what tokens should go predict or suggest what tokens should go predict or suggest what tokens should go at which specific places because we have at which specific places because we have at which specific places because we have this initial canvas of tokens and all it this initial canvas of tokens and all it this initial canvas of tokens and all it needs to do is update them at the needs to do is update them at the needs to do is update them at the appropriate places. So there's this appropriate places. So there's this appropriate places. So there's this predicted canvas that it then has but predicted canvas that it then has but predicted canvas that it then has but you know it doesn't do stuff in uh you know it doesn't do stuff in uh you know it doesn't do stuff in uh perfectly at once. So it needs a couple perfectly at once. So it needs a couple perfectly at once. So it needs a couple of steps to do that. There are some of steps to do that. There are some of steps to do that. There are some tokens it's very confident about tokens it's very confident about tokens it's very confident about obviously and some not so much. So what obviously and some not so much. So what obviously and some not so much. So what you then get is in so-called accepted you then get is in so-called accepted you then get is in so-called accepted canvas some tokens are accepted and some canvas some tokens are accepted and some canvas some tokens are accepted and some tokens are not. The tokens that are not tokens are not. The tokens that are not tokens are not. The tokens that are not accepted though those are renoised accepted though those are renoised accepted though those are renoised because the entire purpose of this model because the entire purpose of this model because the entire purpose of this model is to dn noiseise it to remove the noise is to dn noiseise it to remove the noise is to dn noiseise it to remove the noise from the original canvas and find the from the original canvas and find the from the original canvas and find the actual tokens that should be there. And actual tokens that should be there. And actual tokens that should be there. And so we end up with a renoised canvas and so we end up with a renoised canvas and so we end up with a renoised canvas and it go can go ahead and do its second it go can go ahead and do its second it go can go ahead and do its second step. And now in the second step it step. And now in the second step it step. And now in the second step it again uses the KV cache and it can use again uses the KV cache and it can use again uses the KV cache and it can use this KV cache to again do a similar this KV cache to again do a similar this KV cache to again do a similar processing. But what we found if if you processing. But what we found if if you processing. But what we found if if you were to do just this, it's not a perfect were to do just this, it's not a perfect were to do just this, it's not a perfect way to approach this. Because in the way to approach this. Because in the way to approach this. Because in the very first step, it had an initial idea very first step, it had an initial idea very first step, it had an initial idea of what it wanted to do. Even though of what it wanted to do. Even though of what it wanted to do. Even though some of the tokens weren't accepted, it some of the tokens weren't accepted, it some of the tokens weren't accepted, it still had a direction it wanted to go still had a direction it wanted to go still had a direction it wanted to go into. And so there's kind of a skip
-
into. And so there's kind of a skip into. And so there's kind of a skip connection there, a self conditioning connection there, a self conditioning connection there, a self conditioning where you take the logit from the where you take the logit from the where you take the logit from the previous one and feed it to the input of previous one and feed it to the input of previous one and feed it to the input of the second step. And then you can do the second step. And then you can do the second step. And then you can do this process over and over again, this process over and over again, this process over and over again, typically eight steps or something where typically eight steps or something where typically eight steps or something where it updates this canvas. it updates this canvas. it updates this canvas. And then in those eight steps, sometimes And then in those eight steps, sometimes And then in those eight steps, sometimes less, sometimes more, you get 256 less, sometimes more, you get 256 less, sometimes more, you get 256 tokens. And it makes your model go tokens. And it makes your model go tokens. And it makes your model go incredibly fast when served for a single incredibly fast when served for a single incredibly fast when served for a single user on a larger GPU. user on a larger GPU. user on a larger GPU. Efficiency takes a very different Efficiency takes a very different Efficiency takes a very different meaning than when you compare meaning than when you compare meaning than when you compare computebound to memory bound. But computebound to memory bound. But computebound to memory bound. But there's still many use cases for that. there's still many use cases for that. there's still many use cases for that. One we didn't quite expect would blow up One we didn't quite expect would blow up One we didn't quite expect would blow up in the last week or so, which is in the last week or so, which is in the last week or so, which is obviously Jeff. obviously Jeff. obviously Jeff. because Jeff being a a a foundational because Jeff being a a a foundational because Jeff being a a a foundational decision model, you know, essentially a decision model, you know, essentially a decision model, you know, essentially a foundational classifier that's being foundational classifier that's being foundational classifier that's being used to make decisions happens to work used to make decisions happens to work used to make decisions happens to work quite well with diffusion.
-
quite well with diffusion. quite well with diffusion. And it works a little bit like this. So And it works a little bit like this. So And it works a little bit like this. So let's say you have a situation, the let's say you have a situation, the let's say you have a situation, the printer is on fire. anything that printer is on fire. anything that printer is on fire. anything that requires you to want to take an action. requires you to want to take an action. requires you to want to take an action. You describe the state, the situation, You describe the state, the situation, You describe the state, the situation, what have you, you again take that what have you, you again take that what have you, you again take that encoder and you process it for the encoder and you process it for the encoder and you process it for the purpose of creating a KV cache. What purpose of creating a KV cache. What purpose of creating a KV cache. What then happens is instead of having a then happens is instead of having a then happens is instead of having a fully noisy canvas, you kind of prefill fully noisy canvas, you kind of prefill fully noisy canvas, you kind of prefill some of these tokens and the prefill are some of these tokens and the prefill are some of these tokens and the prefill are then related to the actions that you then related to the actions that you then related to the actions that you want the model to make a suggestion on. want the model to make a suggestion on. want the model to make a suggestion on. Okay, the printer is on fire and your Okay, the printer is on fire and your Okay, the printer is on fire and your options are to either, you know, decide options are to either, you know, decide options are to either, you know, decide if I should evacuate or call it. Those if I should evacuate or call it. Those if I should evacuate or call it. Those tokens don't change, but the others are tokens don't change, but the others are tokens don't change, but the others are still noisy. That's the one it needs to still noisy. That's the one it needs to still noisy. That's the one it needs to fill in. And what happens is that when fill in. And what happens is that when fill in. And what happens is that when you do that, it can make a prediction you do that, it can make a prediction you do that, it can make a prediction for only those two tokens for only those two tokens for only those two tokens and using the things that we saw before, and using the things that we saw before, and using the things that we saw before, make a probability of whether it should make a probability of whether it should make a probability of whether it should do something or shouldn't do something.
-
do something or shouldn't do something. do something or shouldn't do something. just one den noising step. Technically just one den noising step. Technically just one den noising step. Technically it can do more but if you want to make it can do more but if you want to make it can do more but if you want to make fast quick decisions uh this is all it fast quick decisions uh this is all it fast quick decisions uh this is all it takes to create well apparently it's takes to create well apparently it's takes to create well apparently it's called diffusion gemma Jeff now but that called diffusion gemma Jeff now but that called diffusion gemma Jeff now but that works extremely well because nobody works extremely well because nobody works extremely well because nobody fine-tuned that model the only thing fine-tuned that model the only thing fine-tuned that model the only thing that happens is just limit it in a way that happens is just limit it in a way that happens is just limit it in a way but that showcases also this concept of but that showcases also this concept of but that showcases also this concept of efficiency because it's okay to limit efficiency because it's okay to limit efficiency because it's okay to limit certain things in certain ways because certain things in certain ways because certain things in certain ways because it opens up the way to make sure it can it opens up the way to make sure it can it opens up the way to make sure it can be use for things we didn't imagine be use for things we didn't imagine be use for things we didn't imagine before. before. before. When you combine all of that together, When you combine all of that together, When you combine all of that together, one thing that makes all of this happen one thing that makes all of this happen one thing that makes all of this happen is that it's all about the community, is that it's all about the community, is that it's all about the community, right? Efficiency is not for us, it's right? Efficiency is not for us, it's right? Efficiency is not for us, it's for you. It's for the folks that use for you. It's for the folks that use for you. It's for the folks that use actually these models. And whether actually these models. And whether actually these models. And whether that's on mobile, on a laptop, on bigger that's on mobile, on a laptop, on bigger that's on mobile, on a laptop, on bigger GPU, that doesn't matter. when you GPU, that doesn't matter. when you GPU, that doesn't matter. when you release these open models and you've release these open models and you've release these open models and you've seen a bunch of talks already today seen a bunch of talks already today seen a bunch of talks already today talking about that talking about that talking about that it needs to be useful for you. That's it needs to be useful for you. That's it needs to be useful for you. That's the main purpose of trying to give back the main purpose of trying to give back the main purpose of trying to give back to the community. But in the same way to the community. But in the same way to the community. But in the same way when you do that you can see the when you do that you can see the when you do that you can see the community giving back to you because a community giving back to you because a community giving back to you because a decision model based on a model that decision model based on a model that decision model based on a model that wasn't necessarily intended for that.
-
wasn't necessarily intended for that. wasn't necessarily intended for that. Well, that's very cool to see that gives Well, that's very cool to see that gives Well, that's very cool to see that gives a lot of innovation from a field uh that a lot of innovation from a field uh that a lot of innovation from a field uh that you know we didn't do you know we didn't do you know we didn't do but by giving people the option and the but by giving people the option and the but by giving people the option and the opportunity to play around with these opportunity to play around with these opportunity to play around with these models to open them up to see what's models to open them up to see what's models to open them up to see what's happening inside them to tweak it to happening inside them to tweak it to happening inside them to tweak it to change it you get these fun cool models change it you get these fun cool models change it you get these fun cool models that who knows where it's going. I have that who knows where it's going. I have that who knows where it's going. I have no clue. uh it's now a big hype. I think no clue. uh it's now a big hype. I think no clue. uh it's now a big hype. I think for a good reason, but the thing is with for a good reason, but the thing is with for a good reason, but the thing is with hypes, we'll see how far it gets. hypes, we'll see how far it gets. hypes, we'll see how far it gets. And when you then talk about efficiency, And when you then talk about efficiency, And when you then talk about efficiency, there's all these different types of there's all these different types of there's all these different types of fields and perspectives that you can fields and perspectives that you can fields and perspectives that you can focus on. And I could stand here and focus on. And I could stand here and focus on. And I could stand here and talk for hours upon hours on what talk for hours upon hours on what talk for hours upon hours on what efficiency means alongside performance. efficiency means alongside performance. efficiency means alongside performance. And that's not an easy thing to do, And that's not an easy thing to do, And that's not an easy thing to do, right? There's so many of of these right? There's so many of of these right? There's so many of of these incredible open models out there that incredible open models out there that incredible open models out there that try to focus on these things because try to focus on these things because try to focus on these things because it's not just about performance but also it's not just about performance but also it's not just about performance but also of on whether you can actually run the of on whether you can actually run the of on whether you can actually run the model because it doesn't matter how good model because it doesn't matter how good model because it doesn't matter how good a model is if you can't run it it's a model is if you can't run it it's a model is if you can't run it it's meaningless and so a lot of attention meaningless and so a lot of attention meaningless and so a lot of attention from different parties and you see a lot from different parties and you see a lot from different parties and you see a lot happening in Jama devices devices and hopefully this helps you get an idea and hopefully this helps you get an idea and hopefully this helps you get an idea of what it means to focus focus on these of what it means to focus focus on these of what it means to focus focus on these on these smaller models. It's a very on these smaller models. It's a very on these smaller models. It's a very different narrative than going bigger different narrative than going bigger different narrative than going bigger and bigger and bigger. That's also and bigger and bigger. That's also and bigger and bigger. That's also interesting. Gives a lot of very cool
-
interesting. Gives a lot of very cool interesting. Gives a lot of very cool capabilities. capabilities. capabilities. But these open models, these small But these open models, these small But these open models, these small models, especially when they're on models, especially when they're on models, especially when they're on device, on edge, your small devices, device, on edge, your small devices, device, on edge, your small devices, they need to be capable and they need to they need to be capable and they need to they need to be capable and they need to be quick. And efficiency from many be quick. And efficiency from many be quick. And efficiency from many different perspectives is a very big different perspectives is a very big different perspectives is a very big part of that. part of that. part of that. Thank you very much. All All right.
-
All right. Thank you guys. Let's give it up one Thank you guys. Let's give it up one Thank you guys. Let's give it up one more time for Martin. All right. All right. So So So our next speaker is going to come in a our next speaker is going to come in a our next speaker is going to come in a few minutes. In the meantime, I have so few minutes. In the meantime, I have so few minutes. In the meantime, I have so many questions for you guys. So many questions for you guys. So many questions for you guys. So um I just would like to know what was um I just would like to know what was um I just would like to know what was your favorite part of this conference? your favorite part of this conference? your favorite part of this conference? Like does anybody want to share? Okay, maybe two open-ended questions Okay, maybe two open-ended questions don't work in a crowd big crowd like don't work in a crowd big crowd like don't work in a crowd big crowd like this. Okay. So, um uh you guys I I don't this. Okay. So, um uh you guys I I don't this. Okay. So, um uh you guys I I don't know if you guys went to uh to the expo, know if you guys went to uh to the expo, know if you guys went to uh to the expo, you checked the swag, you checked like you checked the swag, you checked like you checked the swag, you checked like uh but but I like I just want to know uh but but I like I just want to know uh but but I like I just want to know for example for me what I like the most for example for me what I like the most for example for me what I like the most is uh seeing people at the expo is uh seeing people at the expo is uh seeing people at the expo discussing and and and making new discussing and and and making new discussing and and and making new friends. I think it's something that is friends. I think it's something that is friends. I think it's something that is uh quite obvious to say right like every uh quite obvious to say right like every uh quite obvious to say right like every time we go to a new place, every time we time we go to a new place, every time we time we go to a new place, every time we go to an event is to meet people. But I go to an event is to meet people. But I go to an event is to meet people. But I really like that it really becomes like really like that it really becomes like really like that it really becomes like a family. I've been um uh attending AI a family. I've been um uh attending AI a family. I've been um uh attending AI engineer for quite a while now uh around engineer for quite a while now uh around engineer for quite a while now uh around the world. So in San Francisco, in New the world. So in San Francisco, in New the world. So in San Francisco, in New York, London, and here. And it's so cool York, London, and here. And it's so cool York, London, and here. And it's so cool to be able to see like sometimes we see to be able to see like sometimes we see to be able to see like sometimes we see the same faces over and over and then we the same faces over and over and then we the same faces over and over and then we make new connections. So my my advice to make new connections. So my my advice to make new connections. So my my advice to be honest with you guys I know we we be honest with you guys I know we we be honest with you guys I know we we don't have a lot of time left but yeah don't have a lot of time left but yeah don't have a lot of time left but yeah just make sure that you connect make just make sure that you connect make just make sure that you connect make make sure that uh you you collaborate
-
make sure that uh you you collaborate make sure that uh you you collaborate because I think that the through this because I think that the through this because I think that the through this community like a lot can happen and you community like a lot can happen and you community like a lot can happen and you never know it doesn't have to happen now never know it doesn't have to happen now never know it doesn't have to happen now it can happen in the future. Um, another it can happen in the future. Um, another it can happen in the future. Um, another nice moment that I saw is uh so many nice moment that I saw is uh so many nice moment that I saw is uh so many people lining up for to talk to uh to people lining up for to talk to uh to people lining up for to talk to uh to our speakers like yeah so many people our speakers like yeah so many people our speakers like yeah so many people come to me and they ask me also about um come to me and they ask me also about um come to me and they ask me also about um some some of our speakers apparently some some of our speakers apparently some some of our speakers apparently like yeah um P.A.I. like yeah um P.A.I. like yeah um P.A.I. was quite popular. Uh ARM also is quite was quite popular. Uh ARM also is quite was quite popular. Uh ARM also is quite popular. So it's good to see that uh our popular. So it's good to see that uh our popular. So it's good to see that uh our audience is uh is quite diverse. Nobody audience is uh is quite diverse. Nobody audience is uh is quite diverse. Nobody wants to share what their favorite wants to share what their favorite wants to share what their favorite moment was. the food. All right. Yeah, I I guess so. the food. All right. Yeah, I I guess so. Yeah, I guess so. Yeah, thank you Sarah Yeah, I guess so. Yeah, thank you Sarah Yeah, I guess so. Yeah, thank you Sarah again for sponsoring for the food. So, again for sponsoring for the food. So, again for sponsoring for the food. So, apparently we nailed that. Okay. Any any apparently we nailed that. Okay. Any any apparently we nailed that. Okay. Any any talk favorite talk? Gemma 4 was great. Okay. Wow. Nice. What Gemma 4 was great. Okay. Wow. Nice. What else?
-
Okay, that was Gemma 4 and that's it. Okay, that was Gemma 4 and that's it. Like we should have just have gym four Like we should have just have gym four Like we should have just have gym four for the entire conference and that's it. for the entire conference and that's it. for the entire conference and that's it. The kids that was hilarious and I love The kids that was hilarious and I love The kids that was hilarious and I love Kitza. Yeah, shout out to KZA. But but Kitza. Yeah, shout out to KZA. But but Kitza. Yeah, shout out to KZA. But but yeah, he's uh he's hilarious. Yeah, I yeah, he's uh he's hilarious. Yeah, I yeah, he's uh he's hilarious. Yeah, I agree agree with that one. Um all right. agree agree with that one. Um all right. agree agree with that one. Um all right. Okay, so Okay, so Okay, so enough of me, but this is our last enough of me, but this is our last enough of me, but this is our last presentation of the day. Okay, so thank presentation of the day. Okay, so thank presentation of the day. Okay, so thank you so much for hanging in there. Um but you so much for hanging in there. Um but you so much for hanging in there. Um but before we get started before we get started before we get started I would like to introduce introduce uh I would like to introduce introduce uh I would like to introduce introduce uh Olivia Tibu from Gradium who is going to Olivia Tibu from Gradium who is going to Olivia Tibu from Gradium who is going to talk about the missing layer of talk about the missing layer of talk about the missing layer of conversational AI. Let's give it up for conversational AI. Let's give it up for conversational AI. Let's give it up for Olivia. It's not what I meant either.
-
It's not what I meant either. All right. All right. All right. should get started. Let me see if that should get started. Let me see if that should get started. Let me see if that is working. That is working. is working. That is working. is working. That is working. Cool. All right. Should I get started? Cool. All right. Should I get started? Cool. All right. Should I get started? Yes. All right. So, let's get started. Yes. All right. So, let's get started. Yes. All right. So, let's get started. Hi everyone. My name is Olivia Tubul. Uh Hi everyone. My name is Olivia Tubul. Uh Hi everyone. My name is Olivia Tubul. Uh I'm a CTO and co-founder at Gradium and I'm a CTO and co-founder at Gradium and I'm a CTO and co-founder at Gradium and today I'm going to talk about the today I'm going to talk about the today I'm going to talk about the missing layers of conversational AI. missing layers of conversational AI. missing layers of conversational AI. It's a mysterous title. So, we'll see It's a mysterous title. So, we'll see It's a mysterous title. So, we'll see what that mean. Uh maybe before we do, what that mean. Uh maybe before we do, what that mean. Uh maybe before we do, let me introduce you uh to who we are at let me introduce you uh to who we are at let me introduce you uh to who we are at Gradium. Gradium. Gradium. So we are a one-year-old startup So we are a one-year-old startup So we are a one-year-old startup uh focused on training uh voice models. uh focused on training uh voice models. uh focused on training uh voice models. So those acronyms you might have seen if So those acronyms you might have seen if So those acronyms you might have seen if you don't did not ST means speech to you don't did not ST means speech to you don't did not ST means speech to text, TTS text to speech and S2S speech text, TTS text to speech and S2S speech text, TTS text to speech and S2S speech to speech. So all these kind of flavors to speech. So all these kind of flavors to speech. So all these kind of flavors of models of models of models uh we train and serve at Gradium. uh we train and serve at Gradium. uh we train and serve at Gradium. Maybe before that if I go back to to the Maybe before that if I go back to to the Maybe before that if I go back to to the genesis of um of Gradium is QAI. So genesis of um of Gradium is QAI. So genesis of um of Gradium is QAI. So maybe you've heard of it. I don't know.
-
maybe you've heard of it. I don't know. maybe you've heard of it. I don't know. Yes, some of you. Great. So CQA is a a Yes, some of you. Great. So CQA is a a Yes, some of you. Great. So CQA is a a Paris-based open science lab that was Paris-based open science lab that was Paris-based open science lab that was founded in 2023 with quite some money founded in 2023 with quite some money founded in 2023 with quite some money and a lot of freedom and out of this and a lot of freedom and out of this and a lot of freedom and out of this freedom they you know made some great freedom they you know made some great freedom they you know made some great scientific breakthrough. uh Moshi is scientific breakthrough. uh Moshi is scientific breakthrough. uh Moshi is probably the most famous and I'll go probably the most famous and I'll go probably the most famous and I'll go back to that a bit later in the back to that a bit later in the back to that a bit later in the presentation uh and of obviously some presentation uh and of obviously some presentation uh and of obviously some others but that um I would say the others but that um I would say the others but that um I would say the success of those voice models were so success of those voice models were so success of those voice models were so big that somehow we thought in order to big that somehow we thought in order to big that somehow we thought in order to get the you know bigger reach of our get the you know bigger reach of our get the you know bigger reach of our research idea we should turn that into research idea we should turn that into research idea we should turn that into products. So the goal of Gradium was products. So the goal of Gradium was products. So the goal of Gradium was really to bridge this gap from research really to bridge this gap from research really to bridge this gap from research to product and and here we are one year to product and and here we are one year to product and and here we are one year later. later. later. So, So, So, conversational AI, right? Uh where where conversational AI, right? Uh where where conversational AI, right? Uh where where do we got that? Actually, today pretty do we got that? Actually, today pretty do we got that? Actually, today pretty much everywhere you've heard about voice much everywhere you've heard about voice much everywhere you've heard about voice agents, whether it's in gaming, live agents, whether it's in gaming, live agents, whether it's in gaming, live streaming, streaming, streaming, assistance, robotics, wherever you want assistance, robotics, wherever you want assistance, robotics, wherever you want to talk to a machine, there is a voice to talk to a machine, there is a voice to talk to a machine, there is a voice agent involved.
-
agent involved. agent involved. And I would like to explain you a bit of And I would like to explain you a bit of And I would like to explain you a bit of the history uh today, not only but part the history uh today, not only but part the history uh today, not only but part of it. And if I go back to well what I of it. And if I go back to well what I of it. And if I go back to well what I could call provocatively the prehistoric could call provocatively the prehistoric could call provocatively the prehistoric time of voice agents time of voice agents time of voice agents meaning before LLMs meaning before LLMs meaning before LLMs uh you probably remember that right I uh you probably remember that right I uh you probably remember that right I think that was the closest think that was the closest think that was the closest um I would say glance at AGI back in the um I would say glance at AGI back in the um I would say glance at AGI back in the days right it really felt like AGI right days right it really felt like AGI right days right it really felt like AGI right that was the iPhone and Siri that was the iPhone and Siri that was the iPhone and Siri and Siri was nothing really agi was A and Siri was nothing really agi was A and Siri was nothing really agi was A very nice uh product, engineering very nice uh product, engineering very nice uh product, engineering product made of different layers, lots product made of different layers, lots product made of different layers, lots of layers, six of them, sorry. of layers, six of them, sorry. of layers, six of them, sorry. First one is P uh the speech to text the First one is P uh the speech to text the First one is P uh the speech to text the speech recognition. After the audio is speech recognition. After the audio is speech recognition. After the audio is turned into text, it was turned into you turned into text, it was turned into you turned into text, it was turned into you know classification classifiers trying know classification classifiers trying know classification classifiers trying to say what is the intent, what is the to say what is the intent, what is the to say what is the intent, what is the object, what is the subject, what is the object, what is the subject, what is the object, what is the subject, what is the complement, stuff like that. a dialog complement, stuff like that. a dialog complement, stuff like that. a dialog manager, manager, manager, uh, some calls to APIs like, you know, uh, some calls to APIs like, you know, uh, some calls to APIs like, you know, check the weather, check the calendar on check the weather, check the calendar on check the weather, check the calendar on the phone, stuff like that. Um, probably the phone, stuff like that. Um, probably the phone, stuff like that. Um, probably the most, you know, I would imagine the most, you know, I would imagine the most, you know, I would imagine template and regex's intensive part was template and regex's intensive part was template and regex's intensive part was to generate some text uh for Siri to be to generate some text uh for Siri to be to generate some text uh for Siri to be spoken out and then the TTS would speak spoken out and then the TTS would speak spoken out and then the TTS would speak it out loud, right?
-
it out loud, right? it out loud, right? Then what came after was LLMs and I say Then what came after was LLMs and I say Then what came after was LLMs and I say maybe we don't need all of that, right? maybe we don't need all of that, right? maybe we don't need all of that, right? This natural language understanding, This natural language understanding, This natural language understanding, those templates those templates those templates all gone. LLM, there's a bit of a lion all gone. LLM, there's a bit of a lion all gone. LLM, there's a bit of a lion here, right? There is still a dialogue here, right? There is still a dialogue here, right? There is still a dialogue manager in modern uh voice agents. And manager in modern uh voice agents. And manager in modern uh voice agents. And those LLMs, they become stronger, right? those LLMs, they become stronger, right? those LLMs, they become stronger, right? And they are now able to connect to the And they are now able to connect to the And they are now able to connect to the external world. They can interact. They external world. They can interact. They external world. They can interact. They can talk to a database, a web search, can talk to a database, a web search, can talk to a database, a web search, MCPs, tool calls, retrieval, rag and MCPs, tool calls, retrieval, rag and MCPs, tool calls, retrieval, rag and they became voice agents and that's what they became voice agents and that's what they became voice agents and that's what you're seeing here is actually how I you're seeing here is actually how I you're seeing here is actually how I would say 99% of voice agent are would say 99% of voice agent are would say 99% of voice agent are working. working. working. It's called the cascaded model ASR or TT It's called the cascaded model ASR or TT It's called the cascaded model ASR or TT ST LLM TTS ST LLM TTS ST LLM TTS right they are not trained jointly. right they are not trained jointly. right they are not trained jointly. Those are three pieces of AI modules Those are three pieces of AI modules Those are three pieces of AI modules that we plug one into the other.
-
that we plug one into the other. that we plug one into the other. There is a an extension of this which is There is a an extension of this which is There is a an extension of this which is okay maybe we don't need three pieces. okay maybe we don't need three pieces. okay maybe we don't need three pieces. Maybe you can just one big model and Maybe you can just one big model and Maybe you can just one big model and that's we call speech to speech and then that's we call speech to speech and then that's we call speech to speech and then there is no turn taking to take into there is no turn taking to take into there is no turn taking to take into account. The model itself knows where to account. The model itself knows where to account. The model itself knows where to speak and when to listen. So this is to give you an overview of So this is to give you an overview of what is convers conversational AI. Uh so what is convers conversational AI. Uh so what is convers conversational AI. Uh so I can also give you a glimpse of what we I can also give you a glimpse of what we I can also give you a glimpse of what we do and then I'll talk about the do and then I'll talk about the do and then I'll talk about the engineering aspect and the AI aspect of engineering aspect and the AI aspect of engineering aspect and the AI aspect of what's under the hood. So what do we do what's under the hood. So what do we do what's under the hood. So what do we do at Gradium? We basically do all of this, at Gradium? We basically do all of this, at Gradium? We basically do all of this, right? We do the text to speech, the right? We do the text to speech, the right? We do the text to speech, the speech to text, we do the text to speech speech to text, we do the text to speech speech to text, we do the text to speech on device. Uh we do some speechtoech on device. Uh we do some speechtoech on device. Uh we do some speechtoech translation. We do some voice design translation. We do some voice design translation. We do some voice design because when you want to say something because when you want to say something because when you want to say something out loud, you want to say it in some out loud, you want to say it in some out loud, you want to say it in some voice. Can be a French accent voice for voice. Can be a French accent voice for voice. Can be a French accent voice for example or it can be a female voice, example or it can be a female voice, example or it can be a female voice, high pitch, low pitch, whatever. And high pitch, low pitch, whatever. And high pitch, low pitch, whatever. And coming soon is actually the speech to coming soon is actually the speech to coming soon is actually the speech to speech.
-
speech. speech. So this talk is called uh the missing So this talk is called uh the missing So this talk is called uh the missing layer of voice AI. So what is actually layer of voice AI. So what is actually layer of voice AI. So what is actually missing? It feels like it's it's missing? It feels like it's it's missing? It feels like it's it's complete, right? And some people will complete, right? And some people will complete, right? And some people will try to convince you this picture is try to convince you this picture is try to convince you this picture is complete but it's not. complete but it's not. complete but it's not. And those are just some examples. It's And those are just some examples. It's And those are just some examples. It's not everything. Uh but we're going to not everything. Uh but we're going to not everything. Uh but we're going to talk more about that. So if you put you talk more about that. So if you put you talk more about that. So if you put you know voice AI in the wild, well most know voice AI in the wild, well most know voice AI in the wild, well most likely it's going to go wild, right? likely it's going to go wild, right? likely it's going to go wild, right? Because most of the um the applications Because most of the um the applications Because most of the um the applications do there are more on the phone, one do there are more on the phone, one do there are more on the phone, one speaker uh you know one computer talking speaker uh you know one computer talking speaker uh you know one computer talking to one another. But put that in real to one another. But put that in real to one another. But put that in real life you get like many speakers. it life you get like many speakers. it life you get like many speakers. it confuses the agent. Noises confusing the confuses the agent. Noises confusing the confuses the agent. Noises confusing the agent long sessions you leave the lose agent long sessions you leave the lose agent long sessions you leave the lose the context and also the real global the context and also the real global the context and also the real global scale. Uh and what I mean real global scale. Uh and what I mean real global scale. Uh and what I mean real global scale is like not the ability for anyone scale is like not the ability for anyone scale is like not the ability for anyone to reach your your service but like the to reach your your service but like the to reach your your service but like the ability as that was Siri was doing like ability as that was Siri was doing like ability as that was Siri was doing like everyone on his phone can talk to it everyone on his phone can talk to it everyone on his phone can talk to it instead of typing.
-
instead of typing. instead of typing. So underlying those issues are some So underlying those issues are some So underlying those issues are some challenges and I'm going to pick two of challenges and I'm going to pick two of challenges and I'm going to pick two of them to explain to dig into how how we them to explain to dig into how how we them to explain to dig into how how we do and what's the challenge. do and what's the challenge. do and what's the challenge. So if you want to do um voice models So if you want to do um voice models So if you want to do um voice models you need to you need to you need to several things. Uh basically you want several things. Uh basically you want several things. Uh basically you want something that's fast uh and something something that's fast uh and something something that's fast uh and something you can serve at scale. And so you have you can serve at scale. And so you have you can serve at scale. And so you have to think about this fast and at scale to think about this fast and at scale to think about this fast and at scale this real time at a scale from from the this real time at a scale from from the this real time at a scale from from the ground up. So it means like designing a ground up. So it means like designing a ground up. So it means like designing a model that will be able to scale and model that will be able to scale and model that will be able to scale and that will be able to be to serve fast that will be able to be to serve fast that will be able to be to serve fast optimizing the inference and having a optimizing the inference and having a optimizing the inference and having a serving infrastructure that will support serving infrastructure that will support serving infrastructure that will support it. I'm going to talk about the first it. I'm going to talk about the first it. I'm going to talk about the first one and the last one. one and the last one. one and the last one. So how do those how does model work? So how do those how does model work? So how do those how does model work? Um, so when you want to design a new Um, so when you want to design a new Um, so when you want to design a new model, you don't start from scratch model, you don't start from scratch model, you don't start from scratch basically, right? You start from basically, right? You start from basically, right? You start from something that's similar and try to something that's similar and try to something that's similar and try to adapt. So what you really need for and adapt. So what you really need for and adapt. So what you really need for and I'm talking about the speechtoech uh I'm talking about the speechtoech uh I'm talking about the speechtoech uh system you need something that will take system you need something that will take system you need something that will take the speech as it comes and that will be the speech as it comes and that will be the speech as it comes and that will be able to take more and more speech as the able to take more and more speech as the able to take more and more speech as the conversation keep going and that will be conversation keep going and that will be conversation keep going and that will be able to as soon as possible start able to as soon as possible start able to as soon as possible start speaking in return right and of course speaking in return right and of course speaking in return right and of course there is a one of the requirements that there is a one of the requirements that there is a one of the requirements that it's speech in speech out right so even
-
it's speech in speech out right so even it's speech in speech out right so even if we don't those models. There are a if we don't those models. There are a if we don't those models. There are a family of models that is really similar family of models that is really similar family of models that is really similar to this and that satisfies two out of to this and that satisfies two out of to this and that satisfies two out of those three requirements which are LMS. those three requirements which are LMS. those three requirements which are LMS. LMS you take a text in and a text out. LMS you take a text in and a text out. LMS you take a text in and a text out. So how do they work? You have a context. So how do they work? You have a context. So how do they work? You have a context. At the beginning it's the prompt but At the beginning it's the prompt but At the beginning it's the prompt but then it's a context. Those are words then it's a context. Those are words then it's a context. Those are words and the LLM is trained to predict what's and the LLM is trained to predict what's and the LLM is trained to predict what's next. Okay. So in that situation next. Okay. So in that situation next. Okay. So in that situation gradient is an AI model company based gradient is an AI model company based gradient is an AI model company based what's next is it in off at blue bird what's next is it in off at blue bird what's next is it in off at blue bird you know it gives a probability to all you know it gives a probability to all you know it gives a probability to all of them so of course you know in seems of them so of course you know in seems of them so of course you know in seems to be a good candidate so if you to be a good candidate so if you to be a good candidate so if you represent it that way at the bottom line represent it that way at the bottom line represent it that way at the bottom line this is the context on the top line this this is the context on the top line this this is the context on the top line this is the prediction what's interesting is is the prediction what's interesting is is the prediction what's interesting is that and that's why it's called auto that and that's why it's called auto that and that's why it's called auto reggressive models reggressive models reggressive models the next prediction is going to be part the next prediction is going to be part the next prediction is going to be part of the context at the next round. So you of the context at the next round. So you of the context at the next round. So you say oop sorry um gradium is unpredicted say oop sorry um gradium is unpredicted say oop sorry um gradium is unpredicted AI AI becomes part of the context and AI AI becomes part of the context and AI AI becomes part of the context and now it can bring a new one and another now it can bring a new one and another now it can bring a new one and another one and another one and this way on the one and another one and this way on the one and another one and this way on the go you will uh predict and generate a go you will uh predict and generate a go you will uh predict and generate a full sentence. So can it work on audio?
-
full sentence. So can it work on audio? full sentence. So can it work on audio? Feels like it's the same right? Audio is Feels like it's the same right? Audio is Feels like it's the same right? Audio is a sequence sequence to sequence. Why a sequence sequence to sequence. Why a sequence sequence to sequence. Why could not it work? could not it work? could not it work? Well, there is a dimensionality issue Well, there is a dimensionality issue Well, there is a dimensionality issue here. Um, the problem with audio is that here. Um, the problem with audio is that here. Um, the problem with audio is that this text, Grandom is an IMO company this text, Grandom is an IMO company this text, Grandom is an IMO company based in Paris. It's nine words. So, based in Paris. It's nine words. So, based in Paris. It's nine words. So, like eight words for the context one like eight words for the context one like eight words for the context one you're going to predict. you're going to predict. you're going to predict. It's 3 seconds when it's spoken out It's 3 seconds when it's spoken out It's 3 seconds when it's spoken out loud. But 3 seconds of audio, if you loud. But 3 seconds of audio, if you loud. But 3 seconds of audio, if you represent it on a computer on a normal represent it on a computer on a normal represent it on a computer on a normal quality, it will be like 24,000 quality, it will be like 24,000 quality, it will be like 24,000 samples per second. those eight words samples per second. those eight words samples per second. those eight words becomes seven 72,000 becomes seven 72,000 becomes seven 72,000 sample uh values. So it's much longer of sample uh values. So it's much longer of sample uh values. So it's much longer of a sequence. So you have 10,000 times a sequence. So you have 10,000 times a sequence. So you have 10,000 times more sequ uh values to predict more sequ uh values to predict more sequ uh values to predict and since the transformer architecture and since the transformer architecture and since the transformer architecture which what's the LLM is relying upon is which what's the LLM is relying upon is which what's the LLM is relying upon is relying on self attention which is relying on self attention which is relying on self attention which is quadratic. then your uh attention matrix quadratic. then your uh attention matrix quadratic. then your uh attention matrix is 100 million times bigger. So it's not is 100 million times bigger. So it's not is 100 million times bigger. So it's not tractable. So the first trick is to turn tractable. So the first trick is to turn tractable. So the first trick is to turn the audio into tokens. And that's the the audio into tokens. And that's the the audio into tokens. And that's the neural codec. And actually it turns out neural codec. And actually it turns out neural codec. And actually it turns out you can do that. You can train a model you can do that. You can train a model you can do that. You can train a model to compress the audio in tokens from 24 to compress the audio in tokens from 24 to compress the audio in tokens from 24 kHz to 12 hertz 12.5.
-
kHz to 12 hertz 12.5. kHz to 12 hertz 12.5. So you divide by 2,00 then you have a So you divide by 2,00 then you have a So you divide by 2,00 then you have a tractable number of tokens that's a tractable number of tokens that's a tractable number of tokens that's a sequence and you can start predicting sequence and you can start predicting sequence and you can start predicting because at the end of the story will because at the end of the story will because at the end of the story will decode them into back to audio. decode them into back to audio. decode them into back to audio. So if you wanted to do that to train a So if you wanted to do that to train a So if you wanted to do that to train a dialogue system you could say me I'm the dialogue system you could say me I'm the dialogue system you could say me I'm the I'm the LLM okay I'm the gray one. So, I'm the LLM okay I'm the gray one. So, I'm the LLM okay I'm the gray one. So, we have a sequence of me talking, me we have a sequence of me talking, me we have a sequence of me talking, me talking, me talking, you talking, you talking, me talking, you talking, you talking, me talking, you talking, you talking, me talking, blah, blah, blah. talking, me talking, blah, blah, blah. talking, me talking, blah, blah, blah. That's a sequence and we're going to That's a sequence and we're going to That's a sequence and we're going to predict the next audio tokens, right? predict the next audio tokens, right? predict the next audio tokens, right? But this, sorry, sorry, sorry, sorry. But this, sorry, sorry, sorry, sorry. But this, sorry, sorry, sorry, sorry. This doesn't really work like this This doesn't really work like this This doesn't really work like this because in a conversation, it's not like because in a conversation, it's not like because in a conversation, it's not like I talk or you talk. We might be talking I talk or you talk. We might be talking I talk or you talk. We might be talking at the same time. Otherwise, it's just at the same time. Otherwise, it's just at the same time. Otherwise, it's just like walkie-talkie uh conversation, like walkie-talkie uh conversation, like walkie-talkie uh conversation, which which is lame. which which is lame. which which is lame. So in practice you have to invent So in practice you have to invent So in practice you have to invent something new that we don't have in LLMs something new that we don't have in LLMs something new that we don't have in LLMs is multiream hierarchical transformers. is multiream hierarchical transformers. is multiream hierarchical transformers. So like now you don't have a single line So like now you don't have a single line So like now you don't have a single line of time for the context. You have two of time for the context. You have two of time for the context. You have two you have me speaking and you speaking you have me speaking and you speaking you have me speaking and you speaking and those two things will attend to to and those two things will attend to to and those two things will attend to to be able to predict what's next. And so be able to predict what's next. And so be able to predict what's next. And so those models they can listen and talk at those models they can listen and talk at those models they can listen and talk at the same time.
-
the same time. the same time. Give you an example. So, the planet is Sirius 22. Can you So, the planet is Sirius 22. Can you plot a trajectory course to it, please? plot a trajectory course to it, please? plot a trajectory course to it, please? >> Yes, sir. >> Yes, sir. >> Yes, sir. >> Okay. How long is it going to take us to >> Okay. How long is it going to take us to >> Okay. How long is it going to take us to get there? get there? get there? >> I've mapped it out. It's approximately 5 >> I've mapped it out. It's approximately 5 >> I've mapped it out. It's approximately 5 months to get there. Okay, that's that's months to get there. Okay, that's that's months to get there. Okay, that's that's not too bad. Uh, do you think we have not too bad. Uh, do you think we have not too bad. Uh, do you think we have all we need on board the ship to start all we need on board the ship to start all we need on board the ship to start the mission? the mission? the mission? >> Yes, sir. We have everything we need. >> Yes, sir. We have everything we need. >> Yes, sir. We have everything we need. >> Okay. >> Okay. >> Okay. >> Today, >> Today, >> Today, >> okay, I'm going to skip the second one. >> okay, I'm going to skip the second one. >> okay, I'm going to skip the second one. So, you see like it's very natural that So, you see like it's very natural that So, you see like it's very natural that the model is trained to listen and talk the model is trained to listen and talk the model is trained to listen and talk at the same time. So, there is no at the same time. So, there is no at the same time. So, there is no latency whatsoever. And that was Alex, latency whatsoever. And that was Alex, latency whatsoever. And that was Alex, our chief scientific officer in 2024. our chief scientific officer in 2024. our chief scientific officer in 2024. So now you're there is something I So now you're there is something I So now you're there is something I didn't tell you is that in 2024 we had didn't tell you is that in 2024 we had didn't tell you is that in 2024 we had this technology at QI but we're not this technology at QI but we're not this technology at QI but we're not selling it. So what's the missing layer selling it. So what's the missing layer selling it. So what's the missing layer here? Well, there is something here is here? Well, there is something here is here? Well, there is something here is that this one model is just one model that this one model is just one model that this one model is just one model and it's capable of and it's capable of and it's capable of um listening, speaking and you know um listening, speaking and you know um listening, speaking and you know processing language. That's a lot of processing language. That's a lot of processing language. That's a lot of things for a very small model. So when things for a very small model. So when things for a very small model. So when you talk to it at some point you will you talk to it at some point you will you talk to it at some point you will think okay it's not as good as you know think okay it's not as good as you know think okay it's not as good as you know the limitation is even worse is that the limitation is even worse is that the limitation is even worse is that this model is kind of airgapped it
-
this model is kind of airgapped it this model is kind of airgapped it doesn't talk to the external world. So doesn't talk to the external world. So doesn't talk to the external world. So if it doesn't talk to the world if you if it doesn't talk to the world if you if it doesn't talk to the world if you cannot act on you know booking an cannot act on you know booking an cannot act on you know booking an appointment or something in production appointment or something in production appointment or something in production it's it's just a gadget. So the big it's it's just a gadget. So the big it's it's just a gadget. So the big challenge for speech to speech today is challenge for speech to speech today is challenge for speech to speech today is to be able to do seamless integration to be able to do seamless integration to be able to do seamless integration with tool calling. And if you do that, with tool calling. And if you do that, with tool calling. And if you do that, then you have the same experience while then you have the same experience while then you have the same experience while being able to do what a voice agent being able to do what a voice agent being able to do what a voice agent should do. Okay, should do. Okay, should do. Okay, so that was about how we build a model so that was about how we build a model so that was about how we build a model roughly, not all the recipe. Um, and now roughly, not all the recipe. Um, and now roughly, not all the recipe. Um, and now let's say you have the model, you want let's say you have the model, you want let's say you have the model, you want to scale it and serve it to the world. to scale it and serve it to the world. to scale it and serve it to the world. How do you do that? So here we got the How do you do that? So here we got the How do you do that? So here we got the latency problem. So you have one person latency problem. So you have one person latency problem. So you have one person with his computer. I'm going to take the with his computer. I'm going to take the with his computer. I'm going to take the text to speech example here sending some text to speech example here sending some text to speech example here sending some text to our server and we have to answer text to our server and we have to answer text to our server and we have to answer some audio back and if you want to plug some audio back and if you want to plug some audio back and if you want to plug that into a voice agent it has to be that into a voice agent it has to be that into a voice agent it has to be fast has to be real time and by real fast has to be real time and by real fast has to be real time and by real time we mean less than 300 milliseconds.
-
time we mean less than 300 milliseconds. time we mean less than 300 milliseconds. So that's the latency budget that you So that's the latency budget that you So that's the latency budget that you have for the whole thing. If you don't, have for the whole thing. If you don't, have for the whole thing. If you don't, then it it's going to be like a weird then it it's going to be like a weird then it it's going to be like a weird conversation where you wait for a while conversation where you wait for a while conversation where you wait for a while before getting an answer and you're before getting an answer and you're before getting an answer and you're going to be bored. So, it has to be going to be bored. So, it has to be going to be bored. So, it has to be fast. And during that trip, you have fast. And during that trip, you have fast. And during that trip, you have actually three layers that you have to actually three layers that you have to actually three layers that you have to go through. One is a transport layer, go through. One is a transport layer, go through. One is a transport layer, right? Network layer. The information right? Network layer. The information right? Network layer. The information has to travel from point A to point B. has to travel from point A to point B. has to travel from point A to point B. Then you have the product layer because Then you have the product layer because Then you have the product layer because like someone sends me a request. I like someone sends me a request. I like someone sends me a request. I should check this is that legitimate. should check this is that legitimate. should check this is that legitimate. Does this person uh does this person Does this person uh does this person Does this person uh does this person have the permission the credits? He have the permission the credits? He have the permission the credits? He wants this voice. Do we have this voice? wants this voice. Do we have this voice? wants this voice. Do we have this voice? Where is it? And so on and so forth. So Where is it? And so on and so forth. So Where is it? And so on and so forth. So you have to enhance the request or deny you have to enhance the request or deny you have to enhance the request or deny the request. And then there is a the request. And then there is a the request. And then there is a compute. compute. compute. So how do we do that fast? So first So how do we do that fast? So first So how do we do that fast? So first problem you have to solve is a latency problem you have to solve is a latency problem you have to solve is a latency problem. So let's say I've put my server problem. So let's say I've put my server problem. So let's say I've put my server somewhere in Europe and my client is somewhere in Europe and my client is somewhere in Europe and my client is somewhere in the US. Well, just by going somewhere in the US. Well, just by going somewhere in the US. Well, just by going round trip through the Atlantic Ocean, I round trip through the Atlantic Ocean, I round trip through the Atlantic Ocean, I kind of I'm wasting 100 milliseconds, kind of I'm wasting 100 milliseconds, kind of I'm wasting 100 milliseconds, onethird of my budget for just onethird of my budget for just onethird of my budget for just transporting the information. So how do transporting the information. So how do transporting the information. So how do I fix it? Okay, easy. I'm going to put I fix it? Okay, easy. I'm going to put I fix it? Okay, easy. I'm going to put servers everywhere, right? It's more servers everywhere, right? It's more servers everywhere, right? It's more expensive first, but I'm going to reduce expensive first, but I'm going to reduce expensive first, but I'm going to reduce the trip from the client to the server.
-
the trip from the client to the server. the trip from the client to the server. But now I got new issues, right? Because But now I got new issues, right? Because But now I got new issues, right? Because I had the system with no replication and I had the system with no replication and I had the system with no replication and no distribution. Now I have a no distribution. Now I have a no distribution. Now I have a distributed systems distributed systems distributed systems where the information has to be where the information has to be where the information has to be synchronized. Of course, if you want to synchronized. Of course, if you want to synchronized. Of course, if you want to go fast, you put caches, right? So I put go fast, you put caches, right? So I put go fast, you put caches, right? So I put I have a database. I have cache of my I have a database. I have cache of my I have a database. I have cache of my database. It's already an issue per se database. It's already an issue per se database. It's already an issue per se for you know all the cache invalidation. for you know all the cache invalidation. for you know all the cache invalidation. But now you have many caches that has But now you have many caches that has But now you have many caches that has all to be synchronized alto together. all to be synchronized alto together. all to be synchronized alto together. Otherwise, you create an API key and Otherwise, you create an API key and Otherwise, you create an API key and you're going to California. It doesn't you're going to California. It doesn't you're going to California. It doesn't know about the API key. That's weird. know about the API key. That's weird. know about the API key. That's weird. Okay. So, you're trading latency for Okay. So, you're trading latency for Okay. So, you're trading latency for complexity, engineering complexity. We complexity, engineering complexity. We complexity, engineering complexity. We can handle it, but still it's it's can handle it, but still it's it's can handle it, but still it's it's important to be aware of it. important to be aware of it. important to be aware of it. Then the second thing which I think is Then the second thing which I think is Then the second thing which I think is very very new. very very new. very very new. You know when we started a year ago, we You know when we started a year ago, we You know when we started a year ago, we were thinking there is no way if I need were thinking there is no way if I need were thinking there is no way if I need a new GPU to handle my request. There is a new GPU to handle my request. There is a new GPU to handle my request. There is no way I can ask my cloud provider for a no way I can ask my cloud provider for a no way I can ask my cloud provider for a GPU. He gives me a GPU or she give me a GPU. He gives me a GPU or she give me a GPU. He gives me a GPU or she give me a GPU. Uh the GPU is ready and I pull the GPU. Uh the GPU is ready and I pull the GPU. Uh the GPU is ready and I pull the weights. I pull the Docker image. I warm weights. I pull the Docker image. I warm weights. I pull the Docker image. I warm up the GPU and this going to happen in up the GPU and this going to happen in up the GPU and this going to happen in 300 millconds. Impossible, right? like 300 millconds. Impossible, right? like 300 millconds. Impossible, right? like if you try you you you can maybe make it if you try you you you can maybe make it if you try you you you can maybe make it minutes right minutes right minutes right today what's happening so we we never today what's happening so we we never today what's happening so we we never put that as a as an option you have to put that as a as an option you have to put that as a as an option you have to predict and to be ready for the load but predict and to be ready for the load but predict and to be ready for the load but now what's even worse is that the GPU
-
now what's even worse is that the GPU now what's even worse is that the GPU scarcity is so important that you may scarcity is so important that you may scarcity is so important that you may not be able to have it even like in next not be able to have it even like in next not be able to have it even like in next day so maybe you want one more GPU but day so maybe you want one more GPU but day so maybe you want one more GPU but you don't have this GPU because there is you don't have this GPU because there is you don't have this GPU because there is no GPU available in your region because no GPU available in your region because no GPU available in your region because one of the cloud provider are saying one of the cloud provider are saying one of the cloud provider are saying sorry there is everyone wants a GPU we sorry there is everyone wants a GPU we sorry there is everyone wants a GPU we don't have it so how do you handle that don't have it so how do you handle that don't have it so how do you handle that so again you're going to trade you're so again you're going to trade you're so again you're going to trade you're going to mitigate this risk with more going to mitigate this risk with more going to mitigate this risk with more complexity and there are two things you complexity and there are two things you complexity and there are two things you can do the first thing is I'm going to can do the first thing is I'm going to can do the first thing is I'm going to go to different providers there is more go to different providers there is more go to different providers there is more the odds the odds to find the GPU is is the odds the odds to find the GPU is is the odds the odds to find the GPU is is better or I'm going to try to uh have my better or I'm going to try to uh have my better or I'm going to try to uh have my model run on different architecture on model run on different architecture on model run on different architecture on different models on different chip different models on different chip different models on different chip providers, right? Not only Nvidia or not providers, right? Not only Nvidia or not providers, right? Not only Nvidia or not only H100 or stuff like that. So that I only H100 or stuff like that. So that I only H100 or stuff like that. So that I increases the odds of finding a GPU. So increases the odds of finding a GPU. So increases the odds of finding a GPU. So again to mitigate that risk, I'm again to mitigate that risk, I'm again to mitigate that risk, I'm increasing the complexity of my system. increasing the complexity of my system. increasing the complexity of my system. And how do you handle like burst, right?
-
And how do you handle like burst, right? And how do you handle like burst, right? Because you can uh sometimes everybody Because you can uh sometimes everybody Because you can uh sometimes everybody wants at the same time uh to be served, wants at the same time uh to be served, wants at the same time uh to be served, right? And there are different right? And there are different right? And there are different strategies you can do and we take two strategies you can do and we take two strategies you can do and we take two options. First one is to say now let's options. First one is to say now let's options. First one is to say now let's take advantage of all those clusters and take advantage of all those clusters and take advantage of all those clusters and reroute the traffic from one to the reroute the traffic from one to the reroute the traffic from one to the other the closest one. Or you can also other the closest one. Or you can also other the closest one. Or you can also do change the batch size of your model. do change the batch size of your model. do change the batch size of your model. Maybe you can trade a little bit of Maybe you can trade a little bit of Maybe you can trade a little bit of latency with a bigger batch size right latency with a bigger batch size right latency with a bigger batch size right for the time where you hit the wave. You for the time where you hit the wave. You for the time where you hit the wave. You increase the batch size and then you increase the batch size and then you increase the batch size and then you decrease it. decrease it. decrease it. In practice, if you put all that one In practice, if you put all that one In practice, if you put all that one after the other, a good model, fast after the other, a good model, fast after the other, a good model, fast inference, fastly served, that's what inference, fastly served, that's what inference, fastly served, that's what you get. So, we did um we're building you get. So, we did um we're building you get. So, we did um we're building this week a new beta model for text to this week a new beta model for text to this week a new beta model for text to speech and we did that race where we speech and we did that race where we speech and we did that race where we click somehow at the same time, but we click somehow at the same time, but we click somehow at the same time, but we measure the time for uh how much it measure the time for uh how much it measure the time for uh how much it takes to get the audio back for us and takes to get the audio back for us and takes to get the audio back for us and some providers. And this is what we we some providers. And this is what we we some providers. And this is what we we got.
-
>> Hi, I'm Gradium's newest model and I'm >> Hi, I'm Gradium's newest model and I'm the fastest text to speech we've ever the fastest text to speech we've ever the fastest text to speech we've ever built. Faster first audio, more natural built. Faster first audio, more natural built. Faster first audio, more natural conversations, no waiting, and conversations, no waiting, and conversations, no waiting, and guaranteed speed. Try it today. guaranteed speed. Try it today. guaranteed speed. Try it today. >> Yeah. So, we're happy that we're, you >> Yeah. So, we're happy that we're, you >> Yeah. So, we're happy that we're, you know, pushing the boundaries of of know, pushing the boundaries of of know, pushing the boundaries of of latency. you know, the best models were latency. you know, the best models were latency. you know, the best models were about like 150 millconds and we made it about like 150 millconds and we made it about like 150 millconds and we made it like three times faster. So, at some like three times faster. So, at some like three times faster. So, at some point there is no there is no limit, point there is no there is no limit, point there is no there is no limit, right? It's not never going to be, you right? It's not never going to be, you right? It's not never going to be, you know, one zero milliseconds, but we're know, one zero milliseconds, but we're know, one zero milliseconds, but we're happy that we're able to put all that happy that we're able to put all that happy that we're able to put all that together uh to make it happen. together uh to make it happen. together uh to make it happen. Now, to close this talk, I'm going to Now, to close this talk, I'm going to Now, to close this talk, I'm going to take talk about take talk about take talk about can we go further? Actually, I'm telling can we go further? Actually, I'm telling can we go further? Actually, I'm telling I told you we cannot go zero millisecond I told you we cannot go zero millisecond I told you we cannot go zero millisecond and that's true. But can we go zero and that's true. But can we go zero and that's true. But can we go zero milliseconds of travel? Can we go zero milliseconds of travel? Can we go zero milliseconds of travel? Can we go zero milliseconds of product and still have milliseconds of product and still have milliseconds of product and still have something fast that would scale even something fast that would scale even something fast that would scale even more? And that's the idea of moving the more? And that's the idea of moving the more? And that's the idea of moving the compute compute compute from the server to the client directly.
-
from the server to the client directly. from the server to the client directly. And for this you have to have very tiny And for this you have to have very tiny And for this you have to have very tiny models that can run on your device on models that can run on your device on models that can run on your device on your phone on your brother on a your phone on your brother on a your phone on your brother on a raspberry whatever. raspberry whatever. raspberry whatever. Uh so we also released like grad phon Uh so we also released like grad phon Uh so we also released like grad phon which is the smallest it's not exactly which is the smallest it's not exactly which is the smallest it's not exactly the smallest one of the smallest but the smallest one of the smallest but the smallest one of the smallest but with a very low u number of parameter with a very low u number of parameter with a very low u number of parameter 100 million parameter um only uh and a 100 million parameter um only uh and a 100 million parameter um only uh and a very good accuracy and this you can also very good accuracy and this you can also very good accuracy and this you can also try feel free to reach out if you want try feel free to reach out if you want try feel free to reach out if you want to try those models I'm going to give to try those models I'm going to give to try those models I'm going to give you a little bit of examples of what it you a little bit of examples of what it you a little bit of examples of what it sounds sounds sounds >> only 100 million parameters and I can >> only 100 million parameters and I can >> only 100 million parameters and I can speak in any voice you give me. speak in any voice you give me. speak in any voice you give me. >> And here's an example with a brand new >> And here's an example with a brand new >> And here's an example with a brand new voice cloned from a 10-second sample. voice cloned from a 10-second sample. voice cloned from a 10-second sample. >> Okay, we're going to skip. >> Okay, we're going to skip. >> Okay, we're going to skip. >> This voice didn't exist a minute ago. I >> This voice didn't exist a minute ago. I >> This voice didn't exist a minute ago. I generated it from a single 10-second generated it from a single 10-second generated it from a single 10-second recording. recording. recording. >> Same model, completely different voice. >> Same model, completely different voice. >> Same model, completely different voice. >> Cloned on the fly from a short clip.
-
>> Cloned on the fly from a short clip. >> Cloned on the fly from a short clip. >> All right, so that's that's pretty much >> All right, so that's that's pretty much >> All right, so that's that's pretty much all I wanted to uh tell you about. you all I wanted to uh tell you about. you all I wanted to uh tell you about. you know what's below the hood in in the know what's below the hood in in the know what's below the hood in in the voice uh AI how we build conversational voice uh AI how we build conversational voice uh AI how we build conversational agents how we train those model how we agents how we train those model how we agents how we train those model how we serve those models and how if we you serve those models and how if we you serve those models and how if we you know put together nice research and nice know put together nice research and nice know put together nice research and nice engineering we can really uh push the engineering we can really uh push the engineering we can really uh push the boundaries of of conversational AI thank boundaries of of conversational AI thank boundaries of of conversational AI thank you >> thank you so much Olivier let's give it >> thank you so much Olivier let's give it up for Olivia for and Gradium up for Olivia for and Gradium up for Olivia for and Gradium Nice. Oh, all right. >> Okay. So, if you follow the QR code that >> Okay. So, if you follow the QR code that was uh that was shared with you guys, was uh that was shared with you guys, was uh that was shared with you guys, then you can have uh credits, right, to then you can have uh credits, right, to then you can have uh credits, right, to to try out Gradium. All right. Thank to try out Gradium. All right. Thank to try out Gradium. All right. Thank you. Okay. So, uh we still have a few you. Okay. So, uh we still have a few you. Okay. So, uh we still have a few more things for you guys. So, don't more things for you guys. So, don't more things for you guys. So, don't don't go don't leave just yet. Um don't go don't leave just yet. Um don't go don't leave just yet. Um actually I've been working in AI for uh actually I've been working in AI for uh actually I've been working in AI for uh quite some time now. I can't say that it quite some time now. I can't say that it quite some time now. I can't say that it was a long time but it feels like a long was a long time but it feels like a long was a long time but it feels like a long time. And one of the questions that I time. And one of the questions that I time. And one of the questions that I get the most is what do you think is get the most is what do you think is get the most is what do you think is going to happen to the next generation?
-
going to happen to the next generation? going to happen to the next generation? And it's a tough question right because And it's a tough question right because And it's a tough question right because things are moving so fast and it's things are moving so fast and it's things are moving so fast and it's already quite hard for us as adults to already quite hard for us as adults to already quite hard for us as adults to understand the impact of all this understand the impact of all this understand the impact of all this technological change. Um, but luckily technological change. Um, but luckily technological change. Um, but luckily for us, there are people that are doing for us, there are people that are doing for us, there are people that are doing something about it. And uh, I I would something about it. And uh, I I would something about it. And uh, I I would like to welcome Cassandra Chin, who's like to welcome Cassandra Chin, who's like to welcome Cassandra Chin, who's going to tell you more about it. So, going to tell you more about it. So, going to tell you more about it. So, please let's give it up for Cassandra. >> Hi. So, my name is Cassandra Chin and I >> Hi. So, my name is Cassandra Chin and I am one of the kids workshop instructors am one of the kids workshop instructors am one of the kids workshop instructors for AI Engineer Paris. This weekend we for AI Engineer Paris. This weekend we for AI Engineer Paris. This weekend we ran a kids event where we educated kids ran a kids event where we educated kids ran a kids event where we educated kids and helped teach them to learn about and helped teach them to learn about and helped teach them to learn about technology to prepare them for this AI technology to prepare them for this AI technology to prepare them for this AI world. But before I get too much into world. But before I get too much into world. But before I get too much into it, I want to talk a little bit about it, I want to talk a little bit about it, I want to talk a little bit about myself and why I'm so passionate about myself and why I'm so passionate about myself and why I'm so passionate about teaching kids technology. teaching kids technology. teaching kids technology. I actually started teaching kids when I I actually started teaching kids when I I actually started teaching kids when I was a kid myself. I was 13 years old at was a kid myself. I was 13 years old at was a kid myself. I was 13 years old at the time and I would teach kids the time and I would teach kids the time and I would teach kids workshops over the weekend, Raspberry Pi workshops over the weekend, Raspberry Pi workshops over the weekend, Raspberry Pi workshops where kids would get to tinker workshops where kids would get to tinker workshops where kids would get to tinker with different technologies.
-
with different technologies. with different technologies. But even before I started teaching in my But even before I started teaching in my But even before I started teaching in my household, there was lots of technology. household, there was lots of technology. household, there was lots of technology. We had 3D printers, a lot of different We had 3D printers, a lot of different We had 3D printers, a lot of different technology. I remember building my own technology. I remember building my own technology. I remember building my own computers. I really believe in inspiring computers. I really believe in inspiring computers. I really believe in inspiring kids to enjoy technology rather than kids to enjoy technology rather than kids to enjoy technology rather than pushing them. Because if you show kids pushing them. Because if you show kids pushing them. Because if you show kids that technology is this very fun thing, that technology is this very fun thing, that technology is this very fun thing, not just dad or mom at the computer, not just dad or mom at the computer, not just dad or mom at the computer, like show the kids what technology can like show the kids what technology can like show the kids what technology can actually do. That's what we really tried actually do. That's what we really tried actually do. That's what we really tried to do in these kids workshops. to do in these kids workshops. to do in these kids workshops. And for the kids workshops this weekend, And for the kids workshops this weekend, And for the kids workshops this weekend, I taught a workshop teaching kids how to I taught a workshop teaching kids how to I taught a workshop teaching kids how to code with AI. The kids used Mistral Vibe code with AI. The kids used Mistral Vibe code with AI. The kids used Mistral Vibe to build their own games. And I was to build their own games. And I was to build their own games. And I was really surprised at what the kids could really surprised at what the kids could really surprised at what the kids could do because we have these preconceived do because we have these preconceived do because we have these preconceived notions about how to code, what you can notions about how to code, what you can notions about how to code, what you can or can't do, but the kids, they just or can't do, but the kids, they just or can't do, but the kids, they just have ideas.
-
have ideas. have ideas. Some ideas they have are like building Some ideas they have are like building Some ideas they have are like building an RPG game, a platformer, an RPG game, a platformer, an RPG game, a platformer, a banana game. a lot of creative ideas a banana game. a lot of creative ideas a banana game. a lot of creative ideas from kids but they just prompt and the from kids but they just prompt and the from kids but they just prompt and the AI could build it for them and with the AI could build it for them and with the AI could build it for them and with the we had an organization called AI tinkers we had an organization called AI tinkers we had an organization called AI tinkers they brought in a lot of underprivileged they brought in a lot of underprivileged they brought in a lot of underprivileged kids for uh to the workshop and these kids for uh to the workshop and these kids for uh to the workshop and these kids they they could learn really well kids they they could learn really well kids they they could learn really well like it didn't matter what background like it didn't matter what background like it didn't matter what background they have and I think it was really they have and I think it was really they have and I think it was really great that they also brought in a lot of great that they also brought in a lot of great that they also brought in a lot of girls another issue with technology is girls another issue with technology is girls another issue with technology is there's a lot of gender biases. As kids there's a lot of gender biases. As kids there's a lot of gender biases. As kids get older and older, they start thinking get older and older, they start thinking get older and older, they start thinking that technology is a guy's thing because that technology is a guy's thing because that technology is a guy's thing because they see more of the industry. But if they see more of the industry. But if they see more of the industry. But if you catch kids when they're very young, you catch kids when they're very young, you catch kids when they're very young, they don't have this idea that they don't have this idea that they don't have this idea that technology is a guys thing. And a lot of technology is a guys thing. And a lot of technology is a guys thing. And a lot of girls can be inspired to join technology girls can be inspired to join technology girls can be inspired to join technology like I have. So I think these kids like I have. So I think these kids like I have. So I think these kids workshops are really great for kids. The workshops are really great for kids. The workshops are really great for kids. The workshop which I taught, we had the kids workshop which I taught, we had the kids workshop which I taught, we had the kids build the games and we actually build the games and we actually build the games and we actually connected it to a GitHub repo which had connected it to a GitHub repo which had connected it to a GitHub repo which had a lot of art assets and a agents.mmd a lot of art assets and a agents.mmd a lot of art assets and a agents.mmd file. So this created a more kids safe file. So this created a more kids safe file. So this created a more kids safe way with AI because a huge concern is way with AI because a huge concern is way with AI because a huge concern is how is it safe to let kids use AI?
-
how is it safe to let kids use AI? how is it safe to let kids use AI? But with this agents.mmd file, we were But with this agents.mmd file, we were But with this agents.mmd file, we were able to instruct the AI to act more able to instruct the AI to act more able to instruct the AI to act more friendly, use less tech jargon, and be friendly, use less tech jargon, and be friendly, use less tech jargon, and be more helpful, run commands rather than more helpful, run commands rather than more helpful, run commands rather than asking you to run commands. A lot of asking you to run commands. A lot of asking you to run commands. A lot of instructions which make the workshops instructions which make the workshops instructions which make the workshops more kids friendlies. We have this video more kids friendlies. We have this video more kids friendlies. We have this video which we produced about the kids which we produced about the kids which we produced about the kids workshop. So if you would please watch workshop. So if you would please watch workshop. So if you would please watch it, I think it's really nice. We're here at AI Engineer Kids Day in We're here at AI Engineer Kids Day in Paris teaching the local kids about AI Paris teaching the local kids about AI Paris teaching the local kids about AI technology, about models, about neural technology, about models, about neural technology, about models, about neural networks. This is a great way to expose networks. This is a great way to expose networks. This is a great way to expose kids to AI in a fun way, in an kids to AI in a fun way, in an kids to AI in a fun way, in an approachable way, and something we love approachable way, and something we love approachable way, and something we love doing at Neo Forj because we're helping doing at Neo Forj because we're helping doing at Neo Forj because we're helping train the next generation of AI train the next generation of AI train the next generation of AI engineers. There's a whole mix of kids here uh from There's a whole mix of kids here uh from a bunch of different backgrounds. Uh a bunch of different backgrounds. Uh a bunch of different backgrounds. Uh some of them are already expert coders some of them are already expert coders some of them are already expert coders at 8 years old and some of them are in at 8 years old and some of them are in at 8 years old and some of them are in their teens and they need a little bit their teens and they need a little bit their teens and they need a little bit extra help. But that's kind of what extra help. But that's kind of what extra help. But that's kind of what excites me is we're putting them all in excites me is we're putting them all in excites me is we're putting them all in one room. There's peer-to-peer learning.
-
one room. There's peer-to-peer learning. one room. There's peer-to-peer learning. They're teaching each other and they're They're teaching each other and they're They're teaching each other and they're helping each other so that when they helping each other so that when they helping each other so that when they grow up and they start working in the grow up and they start working in the grow up and they start working in the industry, they'll have a handle on their industry, they'll have a handle on their industry, they'll have a handle on their own lives, their own careers, their own own lives, their own careers, their own own lives, their own careers, their own ideas. For my workshop, the kids use Mistro For my workshop, the kids use Mistro Libe and it is a coding agent which lets Libe and it is a coding agent which lets Libe and it is a coding agent which lets them create their own games. them create their own games. them create their own games. It's really interesting what kind of It's really interesting what kind of It's really interesting what kind of games the kids build. Some of them games the kids build. Some of them games the kids build. Some of them choose to build an RPG game. Some do choose to build an RPG game. Some do choose to build an RPG game. Some do platformers. Another kid did a balloon platformers. Another kid did a balloon platformers. Another kid did a balloon popping game. And what's interesting is popping game. And what's interesting is popping game. And what's interesting is the there's no language barrier. Even so the there's no language barrier. Even so the there's no language barrier. Even so the kids mainly speak French. They can the kids mainly speak French. They can the kids mainly speak French. They can type French to the agent and it fully type French to the agent and it fully type French to the agent and it fully understands French and still codes the understands French and still codes the understands French and still codes the games for them. So I think this was a games for them. So I think this was a games for them. So I think this was a really good experience for the kids really good experience for the kids really good experience for the kids because they got to build their own because they got to build their own because they got to build their own unique games.
-
So my organization AI tinkers organizes So my organization AI tinkers organizes a lot of events and having an event that a lot of events and having an event that a lot of events and having an event that will bring together the kids and their will bring together the kids and their will bring together the kids and their family members. So our usual community family members. So our usual community family members. So our usual community members that could be the parent or the members that could be the parent or the members that could be the parent or the big brother or big sister. That sounded big brother or big sister. That sounded big brother or big sister. That sounded really exciting so that they can build really exciting so that they can build really exciting so that they can build together and I believe that's part of together and I believe that's part of together and I believe that's part of community building. And there's a reason community building. And there's a reason community building. And there's a reason why we were really excited about doing why we were really excited about doing why we were really excited about doing this. We are running these kids workshops as a We are running these kids workshops as a part of the AI engineer conference part of the AI engineer conference part of the AI engineer conference series. So we will be running the next series. So we will be running the next series. So we will be running the next workshop at AI engineer uh New York and workshop at AI engineer uh New York and workshop at AI engineer uh New York and also AI engineer code in San Francisco also AI engineer code in San Francisco also AI engineer code in San Francisco and next year we'll be running it at and next year we'll be running it at and next year we'll be running it at London. So if you're at those events, London. So if you're at those events, London. So if you're at those events, please bring your kids. We'll welcome please bring your kids. We'll welcome please bring your kids. We'll welcome your kids and teach them.
-
>> Thank you. Yeah. Thank you, Cassandra. >> Thank you. Yeah. Thank you, Cassandra. Let's give it up one more time. That was Let's give it up one more time. That was Let's give it up one more time. That was awesome. Yeah. awesome. Yeah. awesome. Yeah. Thank you, Cassandra. Thank you, Neil Thank you, Cassandra. Thank you, Neil Thank you, Cassandra. Thank you, Neil Forj, uh for doing yeah, such a great Forj, uh for doing yeah, such a great Forj, uh for doing yeah, such a great job. Uh I'm a father myself and I can job. Uh I'm a father myself and I can job. Uh I'm a father myself and I can relate. This is super important work. Um relate. This is super important work. Um relate. This is super important work. Um yeah, so thank you one one more time. yeah, so thank you one one more time. yeah, so thank you one one more time. And uh let's hear it from you guys as And uh let's hear it from you guys as And uh let's hear it from you guys as well. Thank you. You you made it. Hey, well. Thank you. You you made it. Hey, well. Thank you. You you made it. Hey, you made it all the way here, right? I you made it all the way here, right? I you made it all the way here, right? I know somebody somebody told me earlier know somebody somebody told me earlier know somebody somebody told me earlier is like, "Hey, you clap too much. My is like, "Hey, you clap too much. My is like, "Hey, you clap too much. My hands are hurting." I'm like, "Hey, hands are hurting." I'm like, "Hey, hands are hurting." I'm like, "Hey, dude." Like we gotta find a Oh, you're dude." Like we gotta find a Oh, you're dude." Like we gotta find a Oh, you're there. You recognize yourself. Nice. there. You recognize yourself. Nice. there. You recognize yourself. Nice. Yeah. Um, but thank you so much for Yeah. Um, but thank you so much for Yeah. Um, but thank you so much for being with us and I would like to uh being with us and I would like to uh being with us and I would like to uh thank everybody that helped us organize thank everybody that helped us organize thank everybody that helped us organize this event, right? Um, I would like to this event, right? Um, I would like to this event, right? Um, I would like to uh uh uh sorry, I would like to thank our sorry, I would like to thank our sorry, I would like to thank our sponsors, Mrol first for organizing the sponsors, Mrol first for organizing the sponsors, Mrol first for organizing the event and then Nvidia as our uh platinum event and then Nvidia as our uh platinum event and then Nvidia as our uh platinum sponsor, but also all our gold and uh sponsor, but also all our gold and uh sponsor, but also all our gold and uh bronze and and silver sponsors. So, bronze and and silver sponsors. So, bronze and and silver sponsors. So, let's give it up for them, please.
-
And also I think you should be proud of And also I think you should be proud of yourselves. So I I heard that many of yourselves. So I I heard that many of yourselves. So I I heard that many of you have traveled all the way here from you have traveled all the way here from you have traveled all the way here from different places. We heard somebody from different places. We heard somebody from different places. We heard somebody from French, Polynia and then Australia to French, Polynia and then Australia to French, Polynia and then Australia to live and different parts of Europe as live and different parts of Europe as live and different parts of Europe as well. I think this is so cool that well. I think this is so cool that well. I think this is so cool that people come from different parts of the people come from different parts of the people come from different parts of the world to come and and and be with us world to come and and and be with us world to come and and and be with us here and uh we really appreciate it. So here and uh we really appreciate it. So here and uh we really appreciate it. So um yeah. So let's let's hear it from um yeah. So let's let's hear it from um yeah. So let's let's hear it from you. Let's give it up for you. I think you. Let's give it up for you. I think you. Let's give it up for you. I think you did a good job. Okay, so you did a good job. Okay, so you did a good job. Okay, so before we close it out, I uh first of before we close it out, I uh first of before we close it out, I uh first of all, I hope I see you next year, but we all, I hope I see you next year, but we all, I hope I see you next year, but we have one more thing from Patrick and our have one more thing from Patrick and our have one more thing from Patrick and our sponsor, Back Blaze. So, yeah, let's sponsor, Back Blaze. So, yeah, let's sponsor, Back Blaze. So, yeah, let's hear it for Patrick. All right. hear it for Patrick. All right. hear it for Patrick. All right. >> All right, guys. The first thing I'm >> All right, guys. The first thing I'm >> All right, guys. The first thing I'm going to promise you is this will be going to promise you is this will be going to promise you is this will be exactly two minutes. Um, and let's give exactly two minutes. Um, and let's give exactly two minutes. Um, and let's give it up for how much he claps. Um, because it up for how much he claps. Um, because it up for how much he claps. Um, because you got to clap one more time. Can you you got to clap one more time. Can you you got to clap one more time. Can you do it? All right. do it? All right. do it? All right. So, uh, I'm Patrick from Back Blaze. So, uh, I'm Patrick from Back Blaze. So, uh, I'm Patrick from Back Blaze. Great to meet you all. I am buying you Great to meet you all. I am buying you Great to meet you all. I am buying you drinks tonight. Uh, so I am drinks tonight. Uh, so I am drinks tonight. Uh, so I am contractually obligated to explain to contractually obligated to explain to contractually obligated to explain to you what is Back Blaze. I'm going to do you what is Back Blaze. I'm going to do you what is Back Blaze. I'm going to do it as simply as I can. Um, I'll give you it as simply as I can. Um, I'll give you it as simply as I can. Um, I'll give you one sentence to remember. AI runs on one sentence to remember. AI runs on one sentence to remember. AI runs on GPUs. GPUs need data. Back Blaze is the GPUs. GPUs need data. Back Blaze is the GPUs. GPUs need data. Back Blaze is the best capacity storage tier for AI data.
-
best capacity storage tier for AI data. best capacity storage tier for AI data. So, if you remember anything, remember So, if you remember anything, remember So, if you remember anything, remember that. Two examples. uh Decart uh the that. Two examples. uh Decart uh the that. Two examples. uh Decart uh the world model builder. Um when they were world model builder. Um when they were world model builder. Um when they were getting their start, they uh Dean called getting their start, they uh Dean called getting their start, they uh Dean called me up and he's like, "Hey, I've got 15 me up and he's like, "Hey, I've got 15 me up and he's like, "Hey, I've got 15 pabytes and I'm moving it in days and I pabytes and I'm moving it in days and I pabytes and I'm moving it in days and I haven't broken you guys yet. You're haven't broken you guys yet. You're haven't broken you guys yet. You're awesome." And I was like, "Awesome. Who awesome." And I was like, "Awesome. Who awesome." And I was like, "Awesome. Who are you? What are you doing? I have no are you? What are you doing? I have no are you? What are you doing? I have no idea." Um but they chose Back Blaze idea." Um but they chose Back Blaze idea." Um but they chose Back Blaze because we don't charge egress and we because we don't charge egress and we because we don't charge egress and we have really incredible throughput. So have really incredible throughput. So have really incredible throughput. So they were able to race different GPU they were able to race different GPU they were able to race different GPU clusters and create incredibly efficient clusters and create incredibly efficient clusters and create incredibly efficient AI training runs. Uh another customer, AI training runs. Uh another customer, AI training runs. Uh another customer, Cororeweave, we are their object storage Cororeweave, we are their object storage Cororeweave, we are their object storage tier. They chose us because they ran a tier. They chose us because they ran a tier. They chose us because they ran a very exhaustive PC and what they found very exhaustive PC and what they found very exhaustive PC and what they found was that um from exabytes upward um our was that um from exabytes upward um our was that um from exabytes upward um our economics and our performance did not economics and our performance did not economics and our performance did not change and so they selected us because change and so they selected us because change and so they selected us because they believed we could scale with them they believed we could scale with them they believed we could scale with them forever. So whether you're working with forever. So whether you're working with forever. So whether you're working with terabytes or exabytes um we are a great terabytes or exabytes um we are a great terabytes or exabytes um we are a great solution uh to ensure that your data is solution uh to ensure that your data is solution uh to ensure that your data is someplace where you control it you can someplace where you control it you can someplace where you control it you can use it how you want to you can use it use it how you want to you can use it use it how you want to you can use it with who you want to um and uh as an S3 with who you want to um and uh as an S3 with who you want to um and uh as an S3 compatible storage layer that is a compatible storage layer that is a compatible storage layer that is a quarter the cost of S3 and has free quarter the cost of S3 and has free quarter the cost of S3 and has free egress you're going to pay a lot less to egress you're going to pay a lot less to egress you're going to pay a lot less to do it um like I said our mission which do it um like I said our mission which do it um like I said our mission which we execute for lots of genai companies we execute for lots of genai companies we execute for lots of genai companies and neo clouds and data scrapers and and neo clouds and data scrapers and and neo clouds and data scrapers and physical AI companies um is use your physical AI companies um is use your physical AI companies um is use your data however you want, wherever you
-
data however you want, wherever you data however you want, wherever you want, with whoever you want. We are only want, with whoever you want. We are only want, with whoever you want. We are only going to focus on storage and we're going to focus on storage and we're going to focus on storage and we're going to make sure every bite is there going to make sure every bite is there going to make sure every bite is there as fast as you need it. And our primary as fast as you need it. And our primary as fast as you need it. And our primary goal is that you spend as little money goal is that you spend as little money goal is that you spend as little money as possible on storage because after as possible on storage because after as possible on storage because after listening to all the talks here for the listening to all the talks here for the listening to all the talks here for the last couple of days, what you guys are last couple of days, what you guys are last couple of days, what you guys are doing is so interesting and storage doing is so interesting and storage doing is so interesting and storage should be the last of your concerns. So should be the last of your concerns. So should be the last of your concerns. So definitely check us out. I will be definitely check us out. I will be definitely check us out. I will be upstairs. We have a signature uh upstairs. We have a signature uh upstairs. We have a signature uh cocktail, Lebe Blaze. Very original. Um cocktail, Lebe Blaze. Very original. Um cocktail, Lebe Blaze. Very original. Um try one, talk to me. And if you are try one, talk to me. And if you are try one, talk to me. And if you are early stage, we also have a startup early stage, we also have a startup early stage, we also have a startup program where we offer up to $100,000 in program where we offer up to $100,000 in program where we offer up to $100,000 in storage credits, which if you can do the storage credits, which if you can do the storage credits, which if you can do the math, we're a quarter of the cost of S3 math, we're a quarter of the cost of S3 math, we're a quarter of the cost of S3 is going to last uh a good amount of is going to last uh a good amount of is going to last uh a good amount of time depending on how fast you're time depending on how fast you're time depending on how fast you're growing. So, thank you everyone. Thank growing. So, thank you everyone. Thank growing. So, thank you everyone. Thank you again to AI Engineer and Mistl and you again to AI Engineer and Mistl and you again to AI Engineer and Mistl and all of the sponsors and especially all of the sponsors and especially all of the sponsors and especially thanks to you guys. I flew here from thanks to you guys. I flew here from thanks to you guys. I flew here from Minnesota two nights ago. I'm flying Minnesota two nights ago. I'm flying Minnesota two nights ago. I'm flying back tomorrow morning. I'm exhausted. I back tomorrow morning. I'm exhausted. I back tomorrow morning. I'm exhausted. I have to go take care of my kids. I know have to go take care of my kids. I know have to go take care of my kids. I know how hard this is, but I think it was how hard this is, but I think it was how hard this is, but I think it was totally worth it. So, thank you, and totally worth it. So, thank you, and totally worth it. So, thank you, and we'll see you upstairs.
No summary available yet.
View original episode ↗