← Back
Aaron Zisk October 1, 2026 20m

M5 Ultra vs 2 DGX Sparks… The Number You're Not Looking At

Read full transcript 17 segments
  1. Are you thinking what I'm thinking? Are you thinking what I'm thinking? These edges are way too sharp. You These edges are way too sharp. You These edges are way too sharp. You should not let your kids play with them. should not let your kids play with them. should not let your kids play with them. It even says so on the box. I've got the It even says so on the box. I've got the It even says so on the box. I've got the M5 Ultra Max Studio here. 256 GB and two M5 Ultra Max Studio here. 256 GB and two M5 Ultra Max Studio here. 256 GB and two DGX Sparks. And go. And there they go. DGX Sparks. And go. And there they go. DGX Sparks. And go. And there they go. Okay. Both wrote exactly 800 tokens. Okay. Both wrote exactly 800 tokens. Okay. Both wrote exactly 800 tokens. 38.7 tokens a second on the M5 Ultra and 38.7 tokens a second on the M5 Ultra and 38.7 tokens a second on the M5 Ultra and 34.3 on the dual DGX Spark cluster. 34.3 on the dual DGX Spark cluster. 34.3 on the dual DGX Spark cluster. That's basically neck and neck. But That's basically neck and neck. But That's basically neck and neck. But we'll discover that there are quite a we'll discover that there are quite a we'll discover that there are quite a number of differences here. And this is number of differences here. And this is number of differences here. And this is the closest these two are going to get the closest these two are going to get the closest these two are going to get for the rest of the video. So, here's for the rest of the video. So, here's for the rest of the video. So, here's what we've got on this side. This is the what we've got on this side. This is the what we've got on this side. This is the brand new M5 Ultra. Top of the line brand new M5 Ultra. Top of the line brand new M5 Ultra. Top of the line right now with 256 gigs of memory all in right now with 256 gigs of memory all in right now with 256 gigs of memory all in one box. And in this corner, we've got one box. And in this corner, we've got one box. And in this corner, we've got two DJX Sparks. Each one of them has 128 two DJX Sparks. Each one of them has 128 two DJX Sparks. Each one of them has 128 GB. That's also 256 when you add them GB. That's also 256 when you add them GB. That's also 256 when you add them together. and they're connected by one together. and they're connected by one together. and they're connected by one fat 200 GB cable. That cable runs fat 200 GB cable. That cable runs fat 200 GB cable. That cable runs something called Rocky. Not that Rocky. something called Rocky. Not that Rocky. something called Rocky. Not that Rocky. RDMMA over converged Ethernet. It's like RDMMA over converged Ethernet. It's like RDMMA over converged Ethernet. It's like an acronym within an acronym. R is for an acronym within an acronym. R is for an acronym within an acronym. R is for RDMA, which is remote direct memory RDMA, which is remote direct memory RDMA, which is remote direct memory access. So basically, they can write access. So basically, they can write access. So basically, they can write into each other's memory directly. The into each other's memory directly. The into each other's memory directly. The MAC as configured here is $14,000 MAC as configured here is $14,000 MAC as configured here is $14,000 because it's got the 8 TBTE drive in because it's got the 8 TBTE drive in because it's got the 8 TBTE drive in there. The Sparks have 4 TB each. So there. The Sparks have 4 TB each. So there. The Sparks have 4 TB each. So together they're also 8 TB. And the together they're also 8 TB. And the together they're also 8 TB. And the Sparks when they came out they were 4 Sparks when they came out they were 4 Sparks when they came out they were 4 grand each. Now they're almost 5 grand grand each. Now they're almost 5 grand grand each. Now they're almost 5 grand each. So that's 10. Still cheaper than each. So that's 10. Still cheaper than each. So that's 10. Still cheaper than this. Plus the expensive cable. If

  2. this. Plus the expensive cable. If this. Plus the expensive cable. If you're a developer running big models you're a developer running big models you're a developer running big models locally or you want to serve a small locally or you want to serve a small locally or you want to serve a small team, you should watch this video. And team, you should watch this video. And team, you should watch this video. And if you're not, you should also watch if you're not, you should also watch if you're not, you should also watch this video cuz it's pretty cool stuff. this video cuz it's pretty cool stuff. this video cuz it's pretty cool stuff. Merlin AI. It's an all-in-one AI tool Merlin AI. It's an all-in-one AI tool Merlin AI. It's an all-in-one AI tool and they gave my audience a big and they gave my audience a big and they gave my audience a big discount. I keep multiple AI tools discount. I keep multiple AI tools discount. I keep multiple AI tools around because each one is good at around because each one is good at around because each one is good at something, but it gets really expensive something, but it gets really expensive something, but it gets really expensive and bouncing between tabs breaks my and bouncing between tabs breaks my and bouncing between tabs breaks my focus. Merlin AI puts Chad GPT, Claw, focus. Merlin AI puts Chad GPT, Claw, focus. Merlin AI puts Chad GPT, Claw, Gemini, and more in one place so I can Gemini, and more in one place so I can Gemini, and more in one place so I can pick the best one for the moment, pick the best one for the moment, pick the best one for the moment, whether I'm coding, researching, or whether I'm coding, researching, or whether I'm coding, researching, or writing for a video. Watch this. I click writing for a video. Watch this. I click writing for a video. Watch this. I click the Merlin AI extension, chat with the the Merlin AI extension, chat with the the Merlin AI extension, chat with the web page to summarize what I'm reading, web page to summarize what I'm reading, web page to summarize what I'm reading, and pull out the important parts. And I and pull out the important parts. And I and pull out the important parts. And I even have my choice of models right at even have my choice of models right at even have my choice of models right at my fingertips. If I need something my fingertips. If I need something my fingertips. If I need something deeper, I turn on deep research. And it deeper, I turn on deep research. And it deeper, I turn on deep research. And it builds a clean, structured report from builds a clean, structured report from builds a clean, structured report from multiple sources. And it also has quick multiple sources. And it also has quick multiple sources. And it also has quick modes like web, academic, and Reddit modes like web, academic, and Reddit modes like web, academic, and Reddit search. If you pay separately, Chad GPT search. If you pay separately, Chad GPT search. If you pay separately, Chad GPT is $20. Claude is $20. Gemini is $20. is $20. Claude is $20. Gemini is $20. is $20. Claude is $20. Gemini is $20. And that adds up fast. Merlin AI is And that adds up fast. Merlin AI is And that adds up fast. Merlin AI is cheaper because they buy AI API access cheaper because they buy AI API access cheaper because they buy AI API access in bulk. APIs cost less than the $20 in bulk. APIs cost less than the $20 in bulk. APIs cost less than the $20 plans. And most people don't even use plans. And most people don't even use plans. And most people don't even use $20 worth of API in a month. And here's $20 worth of API in a month. And here's $20 worth of API in a month. And here's the discount. I click pricing, continue.

  3. the discount. I click pricing, continue. the discount. I click pricing, continue. It takes me to Stripe. I enter the promo It takes me to Stripe. I enter the promo It takes me to Stripe. I enter the promo code and the total drops to $60 for the code and the total drops to $60 for the code and the total drops to $60 for the year. That's basically five bucks a year. That's basically five bucks a year. That's basically five bucks a month. I don't know how long this deal month. I don't know how long this deal month. I don't know how long this deal will be available, so grab it soon. The will be available, so grab it soon. The will be available, so grab it soon. The link is in the description. Now, two sparks just don't turn into one Now, two sparks just don't turn into one big computer magically. VLM, which is big computer magically. VLM, which is big computer magically. VLM, which is this high throughput and memory this high throughput and memory this high throughput and memory efficient inference and serving engine efficient inference and serving engine efficient inference and serving engine for LLMs. This is the software that most for LLMs. This is the software that most for LLMs. This is the software that most NVIDIA setups run and it splits the NVIDIA setups run and it splits the NVIDIA setups run and it splits the model across both of them. In this case, model across both of them. In this case, model across both of them. In this case, it's called Tensor Parallel. There's it's called Tensor Parallel. There's it's called Tensor Parallel. There's other kinds of parallelism and I talk other kinds of parallelism and I talk other kinds of parallelism and I talk about that in other videos, but today about that in other videos, but today about that in other videos, but today we're doing tensor parallel. And what we're doing tensor parallel. And what we're doing tensor parallel. And what that means is that every layer of the that means is that every layer of the that means is that every layer of the model gets cut in half, one half on each model gets cut in half, one half on each model gets cut in half, one half on each spark. For every single token, both spark. For every single token, both spark. For every single token, both boxes do their half. Then they swap boxes do their half. Then they swap boxes do their half. Then they swap results over the cable before the next results over the cable before the next results over the cable before the next layer can even start. So that's a lot of layer can even start. So that's a lot of layer can even start. So that's a lot of back and forth over a cable like this. back and forth over a cable like this. back and forth over a cable like this. This is a QSFP cable it's called. That's This is a QSFP cable it's called. That's This is a QSFP cable it's called. That's just that standard right there. So why just that standard right there. So why just that standard right there. So why bother with two? Well, modern models bother with two? Well, modern models bother with two? Well, modern models today are perfect fit for two of them.

  4. today are perfect fit for two of them. today are perfect fit for two of them. Deepseek V4 Flash, for example, that's Deepseek V4 Flash, for example, that's Deepseek V4 Flash, for example, that's about 150 160 GB. It doesn't fit on one about 150 160 GB. It doesn't fit on one about 150 160 GB. It doesn't fit on one 128 gig Spark. And yeah, I I tried 128 gig Spark. And yeah, I I tried 128 gig Spark. And yeah, I I tried loading it once. It didn't work out so loading it once. It didn't work out so loading it once. It didn't work out so well. Plus, you need extra space for well. Plus, you need extra space for well. Plus, you need extra space for context. Once you load it across two of context. Once you load it across two of context. Once you load it across two of them, each node is holding just a little them, each node is holding just a little them, each node is holding just a little bit over 100 gigs. I got it loaded on bit over 100 gigs. I got it loaded on bit over 100 gigs. I got it loaded on both of them right now. We got 111 gigs. both of them right now. We got 111 gigs. both of them right now. We got 111 gigs. It's just showing one of them right now It's just showing one of them right now It's just showing one of them right now out of 128. And that's with extra system out of 128. And that's with extra system out of 128. And that's with extra system stuff going on, too. So, each one gets stuff going on, too. So, each one gets stuff going on, too. So, each one gets half the model plus a little room to half the model plus a little room to half the model plus a little room to work. And the Mac just loads the whole work. And the Mac just loads the whole work. And the Mac just loads the whole thing on one box. Just a quick note, I'm thing on one box. Just a quick note, I'm thing on one box. Just a quick note, I'm running the same model on both machines, running the same model on both machines, running the same model on both machines, but it's not the same file. Each file is but it's not the same file. Each file is but it's not the same file. Each file is rounded down to four bits or quantiz, so rounded down to four bits or quantiz, so rounded down to four bits or quantiz, so it's smaller and faster. The Sparks are it's smaller and faster. The Sparks are it's smaller and faster. The Sparks are running VLM. Like I mentioned, Llama running VLM. Like I mentioned, Llama running VLM. Like I mentioned, Llama CPP, a popular tool, works there, too. CPP, a popular tool, works there, too. CPP, a popular tool, works there, too. But VLM is Nvidia's go-to for this kind But VLM is Nvidia's go-to for this kind But VLM is Nvidia's go-to for this kind of split. On the Mac, I tested both of split. On the Mac, I tested both of split. On the Mac, I tested both Llama CPP and MLX. Sometimes one wins Llama CPP and MLX. Sometimes one wins Llama CPP and MLX. Sometimes one wins and sometimes the other one wins, and and sometimes the other one wins, and and sometimes the other one wins, and I'll point that out as we go.

  5. I'll point that out as we go. I'll point that out as we go. So, I'm going to be running DeepS V4 So, I'm going to be running DeepS V4 So, I'm going to be running DeepS V4 Flash, Quen 3.8 Flash Next. That's Flash, Quen 3.8 Flash Next. That's Flash, Quen 3.8 Flash Next. That's another pretty new one that's requires another pretty new one that's requires another pretty new one that's requires two sparks cuz it's large enough. Now, two sparks cuz it's large enough. Now, two sparks cuz it's large enough. Now, every time you generate tokens, it every time you generate tokens, it every time you generate tokens, it actually happens in two steps. And I actually happens in two steps. And I actually happens in two steps. And I know some of you already know all this, know some of you already know all this, know some of you already know all this, but this is for the newcomers. First, but this is for the newcomers. First, but this is for the newcomers. First, the model reads your whole prompt all at the model reads your whole prompt all at the model reads your whole prompt all at once. That's pure math. It's matrix once. That's pure math. It's matrix once. That's pure math. It's matrix multiplication. So, whichever box has multiplication. So, whichever box has multiplication. So, whichever box has more compute wins. Usually, that stuff more compute wins. Usually, that stuff more compute wins. Usually, that stuff happens on the GPU. So, the more happens on the GPU. So, the more happens on the GPU. So, the more powerful GPUs get the job done faster. powerful GPUs get the job done faster. powerful GPUs get the job done faster. And that time is your time to first And that time is your time to first And that time is your time to first token that waiting of the GPU number token that waiting of the GPU number token that waiting of the GPU number crunching. Then it takes that answer and crunching. Then it takes that answer and crunching. Then it takes that answer and writes one token at a time. And every writes one token at a time. And every writes one token at a time. And every token is another trip through memory for token is another trip through memory for token is another trip through memory for the model's weights. Weights is just the model's weights. Weights is just the model's weights. Weights is just basically a collection of numbers. A basically a collection of numbers. A basically a collection of numbers. A huge collection of numbers. Many many huge collection of numbers. Many many huge collection of numbers. Many many gigabytes of collections of numbers. gigabytes of collections of numbers. gigabytes of collections of numbers. That's the files that you download. So That's the files that you download. So That's the files that you download. So because for every token we use in the because for every token we use in the because for every token we use in the memory, the writing of the answer comes memory, the writing of the answer comes memory, the writing of the answer comes down to memory speed. In other words, down to memory speed. In other words, down to memory speed. In other words, the second part of inference is token the second part of inference is token the second part of inference is token generation and it's reliant on memory generation and it's reliant on memory generation and it's reliant on memory bandwidth. So let's take a look at bandwidth. So let's take a look at bandwidth. So let's take a look at writing. That's that second part. And writing. That's that second part. And writing. That's that second part. And this is the race from the start, the one this is the race from the start, the one this is the race from the start, the one I showed you in the beginning. A I showed you in the beginning. A I showed you in the beginning. A slightly different uh version of it, but slightly different uh version of it, but slightly different uh version of it, but very close numbers cuz I ran it multiple very close numbers cuz I ran it multiple very close numbers cuz I ran it multiple times. A short prompt with one user. I'm times. A short prompt with one user. I'm times. A short prompt with one user. I'm pointing over here cuz I have my charts pointing over here cuz I have my charts pointing over here cuz I have my charts over here. In Deep Seek, we're getting over here. In Deep Seek, we're getting over here. In Deep Seek, we're getting about 38 tokens per second on the M5 about 38 tokens per second on the M5 about 38 tokens per second on the M5 Ultra and also about 38 tokens per Ultra and also about 38 tokens per Ultra and also about 38 tokens per second on the uh Dual Sparks. That's second on the uh Dual Sparks. That's second on the uh Dual Sparks. That's basically a tie. Both are going a pretty

  6. basically a tie. Both are going a pretty basically a tie. Both are going a pretty decent speed. This is a big model, so decent speed. This is a big model, so decent speed. This is a big model, so that's not bad at all. With Quan, we that's not bad at all. With Quan, we that's not bad at all. With Quan, we have a little bit of a difference there. have a little bit of a difference there. have a little bit of a difference there. 45 tokens per second on the Mac and 38 45 tokens per second on the Mac and 38 45 tokens per second on the Mac and 38 tokens per second on the dual sparks. tokens per second on the dual sparks. tokens per second on the dual sparks. That one I'll give to the Mac. Why did That one I'll give to the Mac. Why did That one I'll give to the Mac. Why did that happen? Well, my best guess is that happen? Well, my best guess is that happen? Well, my best guess is memory. Apple says the M5 Ultra has 1.2 memory. Apple says the M5 Ultra has 1.2 memory. Apple says the M5 Ultra has 1.2 terab per second of memory bandwidth. terab per second of memory bandwidth. terab per second of memory bandwidth. That's a lot. Each Spark is rated at 273 That's a lot. Each Spark is rated at 273 That's a lot. Each Spark is rated at 273 GB a second, significantly less. And on GB a second, significantly less. And on GB a second, significantly less. And on top of that, they're syncing over a top of that, they're syncing over a top of that, they're syncing over a cable that I measured at about 111 GB. cable that I measured at about 111 GB. cable that I measured at about 111 GB. Not the 200 it's rated for, but still Not the 200 it's rated for, but still Not the 200 it's rated for, but still pretty good. But I didn't test how much pretty good. But I didn't test how much pretty good. But I didn't test how much that cable actually slows things down, that cable actually slows things down, that cable actually slows things down, so take this with a grain of salt. But so take this with a grain of salt. But so take this with a grain of salt. But the Mac has another trick. The engine the Mac has another trick. The engine the Mac has another trick. The engine you pick for the model matters a lot. you pick for the model matters a lot. you pick for the model matters a lot. For example, DeepSeek on Llama CPP For example, DeepSeek on Llama CPP For example, DeepSeek on Llama CPP writes at about 40 tokens per second, writes at about 40 tokens per second, writes at about 40 tokens per second, but you take the same model and use MLX but you take the same model and use MLX but you take the same model and use MLX instead and you're getting 53 tokens per instead and you're getting 53 tokens per instead and you're getting 53 tokens per second. That's 34% faster just from second. That's 34% faster just from second. That's 34% faster just from switching software. And on Llama CPP switching software. And on Llama CPP switching software. And on Llama CPP served exactly the same way. The MAC and served exactly the same way. The MAC and served exactly the same way. The MAC and the Sparks were within 2%. So yeah, the Sparks were within 2%. So yeah, the Sparks were within 2%. So yeah, that's faster than Spark's 38 tokens a that's faster than Spark's 38 tokens a that's faster than Spark's 38 tokens a second, but I only tested DeepSeek on second, but I only tested DeepSeek on second, but I only tested DeepSeek on MLX in process. That's the benchmark MLX in process. That's the benchmark MLX in process. That's the benchmark talking straight to the engine, not talking straight to the engine, not talking straight to the engine, not served over HTTP like a real chat served over HTTP like a real chat served over HTTP like a real chat server. Anyway, I'm collecting server. Anyway, I'm collecting server. Anyway, I'm collecting information. Okay, I'm trying my best to information. Okay, I'm trying my best to information. Okay, I'm trying my best to do all the tests that I can. There's a do all the tests that I can. There's a do all the tests that I can. There's a lot of tests that can be done, but so lot of tests that can be done, but so lot of tests that can be done, but so far the Mac is looking pretty good.

  7. far the Mac is looking pretty good. far the Mac is looking pretty good. Now, just a quick little primer on Now, just a quick little primer on Now, just a quick little primer on tokens. A token is about 3/4 of a word tokens. A token is about 3/4 of a word tokens. A token is about 3/4 of a word about there. So, a common thing you'll about there. So, a common thing you'll about there. So, a common thing you'll hear is 32K or 32,000 tokens, which is hear is 32K or 32,000 tokens, which is hear is 32K or 32,000 tokens, which is kind of like a typical modern starting kind of like a typical modern starting kind of like a typical modern starting point. You start from there and you go point. You start from there and you go point. You start from there and you go up for context size. And that's about up for context size. And that's about up for context size. And that's about 24,000 words, which is roughly a 100page 24,000 words, which is roughly a 100page 24,000 words, which is roughly a 100page document. How long is it going to take document. How long is it going to take document. How long is it going to take you to read a 100page document, huh? I you to read a 100page document, huh? I you to read a 100page document, huh? I bet it's not going to take 30 seconds or bet it's not going to take 30 seconds or bet it's not going to take 30 seconds or however many seconds. We'll find out. however many seconds. We'll find out. however many seconds. We'll find out. Now, Now, Now, reading or prompt processing, that's the reading or prompt processing, that's the reading or prompt processing, that's the GPU cranking away. Remember, I started GPU cranking away. Remember, I started GPU cranking away. Remember, I started small. At 2,000 tokens, both start fast, small. At 2,000 tokens, both start fast, small. At 2,000 tokens, both start fast, but the sparks turn out to be twice as but the sparks turn out to be twice as but the sparks turn out to be twice as quick. Quen 1.5 seconds on the Mac, 85 quick. Quen 1.5 seconds on the Mac, 85 quick. Quen 1.5 seconds on the Mac, 85 seconds on the dual sparks. Deepseek, 2 seconds on the dual sparks. Deepseek, 2 seconds on the dual sparks. Deepseek, 2 and 1/2 seconds on the Mac, one and a/4 and 1/2 seconds on the Mac, one and a/4 and 1/2 seconds on the Mac, one and a/4 seconds on the dual sparks. You won't seconds on the dual sparks. You won't seconds on the dual sparks. You won't care. It's only a 2,00 token prompt. You care. It's only a 2,00 token prompt. You care. It's only a 2,00 token prompt. You probably won't even notice if you probably won't even notice if you probably won't even notice if you sneeze. By the time you wipe your nose, sneeze. By the time you wipe your nose, sneeze. By the time you wipe your nose, it's done. But you might care later.

  8. it's done. But you might care later. it's done. But you might care later. We'll get to that. On Quen, the max We'll get to that. On Quen, the max We'll get to that. On Quen, the max faster writing even wins it back after a faster writing even wins it back after a faster writing even wins it back after a couple of hundred tokens. If we make couple of hundred tokens. If we make couple of hundred tokens. If we make that prompt a little longer, at 8,000 that prompt a little longer, at 8,000 that prompt a little longer, at 8,000 tokens, Deepseek takes about 4 seconds tokens, Deepseek takes about 4 seconds tokens, Deepseek takes about 4 seconds on the sparks and about 10 seconds on on the sparks and about 10 seconds on on the sparks and about 10 seconds on the Mac. We double it to 16,000 and the the Mac. We double it to 16,000 and the the Mac. We double it to 16,000 and the max weight doubles also to 21 seconds. max weight doubles also to 21 seconds. max weight doubles also to 21 seconds. Now, you're going to have to sneeze a Now, you're going to have to sneeze a Now, you're going to have to sneeze a few times and wipe your nose a few few times and wipe your nose a few few times and wipe your nose a few times, huh? Yeah, you're going to notice times, huh? Yeah, you're going to notice times, huh? Yeah, you're going to notice this one. So, let's go big. this one. So, let's go big. this one. So, let's go big. Okay, I know you some of you going to Okay, I know you some of you going to Okay, I know you some of you going to say like, "Oh, 32,000 is not big." Let's say like, "Oh, 32,000 is not big." Let's say like, "Oh, 32,000 is not big." Let's just take it one step at a time. Okay, just take it one step at a time. Okay, just take it one step at a time. Okay, this time I'm using a real codebase. this time I'm using a real codebase. this time I'm using a real codebase. 32,000. Boom. There's 14 Python files in 32,000. Boom. There's 14 Python files in 32,000. Boom. There's 14 Python files in here. Ah, the DJX Sparks started here. Ah, the DJX Sparks started here. Ah, the DJX Sparks started streaming already, which means they're streaming already, which means they're streaming already, which means they're generating tokens. Now we're past the generating tokens. Now we're past the generating tokens. Now we're past the calculation stage and the Mac is still calculation stage and the Mac is still calculation stage and the Mac is still reading the prompt. It's still reading the prompt. It's still reading the prompt. It's still processing. We're at 17 seconds for the processing. We're at 17 seconds for the processing. We're at 17 seconds for the two DJX Sparks. We're done. And yeah, two DJX Sparks. We're done. And yeah, two DJX Sparks. We're done. And yeah, the Mac is still thinking. Still reading the Mac is still thinking. Still reading the Mac is still thinking. Still reading the prompt. Yeah, 50 seconds. I can hear the prompt. Yeah, 50 seconds. I can hear the prompt. Yeah, 50 seconds. I can hear it generating. Yeah. But yeah, that's a it generating. Yeah. But yeah, that's a it generating. Yeah. But yeah, that's a big difference there. 300 tokens big difference there. 300 tokens big difference there. 300 tokens generated. 32,000 prompt tokens. Now, if generated. 32,000 prompt tokens. Now, if generated. 32,000 prompt tokens. Now, if you look down here, Spark 17 seconds to you look down here, Spark 17 seconds to you look down here, Spark 17 seconds to Mac Studios 50 seconds. That's about Mac Studios 50 seconds. That's about Mac Studios 50 seconds. That's about three times the weight. But hold on, three times the weight. But hold on, three times the weight. But hold on, don't run away yet buying the Sparks.

  9. don't run away yet buying the Sparks. don't run away yet buying the Sparks. Some of you might still want to pick up Some of you might still want to pick up Some of you might still want to pick up that Mac. I'll let you know why. Now, that Mac. I'll let you know why. Now, that Mac. I'll let you know why. Now, here on this chart, this is where I used here on this chart, this is where I used here on this chart, this is where I used a slightly shorter prompt, but the ratio a slightly shorter prompt, but the ratio a slightly shorter prompt, but the ratio is about the same. So, we got 44 seconds is about the same. So, we got 44 seconds is about the same. So, we got 44 seconds on the M5 Ultra and 16.5 on the dual on the M5 Ultra and 16.5 on the dual on the M5 Ultra and 16.5 on the dual sparks. Quen 26 seconds on the Ultra, sparks. Quen 26 seconds on the Ultra, sparks. Quen 26 seconds on the Ultra, 11.2 on the Sparks. So, 2.4 times 11.2 on the Sparks. So, 2.4 times 11.2 on the Sparks. So, 2.4 times faster. Not even close. And after that, faster. Not even close. And after that, faster. Not even close. And after that, it kind of keeps going. Double the it kind of keeps going. Double the it kind of keeps going. Double the prompt, double the weight. Now, with the prompt, double the weight. Now, with the prompt, double the weight. Now, with the sparks, I did push them all the way up sparks, I did push them all the way up sparks, I did push them all the way up to 128,000 tokens. I didn't run the Mac to 128,000 tokens. I didn't run the Mac to 128,000 tokens. I didn't run the Mac at 128,000 tokens because I kind of saw at 128,000 tokens because I kind of saw at 128,000 tokens because I kind of saw a pattern here. And if it kept its 32k a pattern here. And if it kept its 32k a pattern here. And if it kept its 32k pace, it would have been over 3 minutes pace, it would have been over 3 minutes pace, it would have been over 3 minutes of waiting. So, it only slows down as of waiting. So, it only slows down as of waiting. So, it only slows down as the prompt gets longer. The Sparks did the prompt gets longer. The Sparks did the prompt gets longer. The Sparks did it in 72 seconds. Still way faster than it in 72 seconds. Still way faster than it in 72 seconds. Still way faster than the M3 Ultra. Okay. All right. Not the M3 Ultra. Okay. All right. Not the M3 Ultra. Okay. All right. Not trying to make excuses. I'm just saying trying to make excuses. I'm just saying trying to make excuses. I'm just saying we we still have improvements overall. Well, let's go back to that codebased Well, let's go back to that codebased race. Deepse seek with a 32k prompt. The race. Deepse seek with a 32k prompt. The race. Deepse seek with a 32k prompt. The spark gets about a 30-cond head start.

  10. spark gets about a 30-cond head start. spark gets about a 30-cond head start. After that, on llama CPP, the writing After that, on llama CPP, the writing After that, on llama CPP, the writing speed is a tie. Quen's different. The speed is a tie. Quen's different. The speed is a tie. Quen's different. The sparks start about 15 seconds ahead, but sparks start about 15 seconds ahead, but sparks start about 15 seconds ahead, but the Mac gains a few thousand of a second the Mac gains a few thousand of a second the Mac gains a few thousand of a second in every token. divide one by the other in every token. divide one by the other in every token. divide one by the other and by my math, the Mac needs a few and by my math, the Mac needs a few and by my math, the Mac needs a few thousand tokens in order to catch up. I thousand tokens in order to catch up. I thousand tokens in order to catch up. I didn't time that one. It's an estimate, didn't time that one. It's an estimate, didn't time that one. It's an estimate, but that's quite a lot. It's a whole new but that's quite a lot. It's a whole new but that's quite a lot. It's a whole new file basically. So big inputs favor the file basically. So big inputs favor the file basically. So big inputs favor the Spark. Big output favors the MAC, at Spark. Big output favors the MAC, at Spark. Big output favors the MAC, at least on Quen. This is one of the least on Quen. This is one of the least on Quen. This is one of the reasons we have a lot of interest in reasons we have a lot of interest in reasons we have a lot of interest in doing this kind of disagregated prefill doing this kind of disagregated prefill doing this kind of disagregated prefill decode using both the Sparks and a Mac. decode using both the Sparks and a Mac. decode using both the Sparks and a Mac. Sparks will do the prefill, Max will do Sparks will do the prefill, Max will do Sparks will do the prefill, Max will do the decode. I did a video about this, an the decode. I did a video about this, an the decode. I did a video about this, an early early prototype a couple months early early prototype a couple months early early prototype a couple months ago. I'll link to it down below. You can ago. I'll link to it down below. You can ago. I'll link to it down below. You can check it out. It's interesting. But this check it out. It's interesting. But this check it out. It's interesting. But this is a new project by this guy Ash Hart. is a new project by this guy Ash Hart. is a new project by this guy Ash Hart. You can check it out. He's on Twitter. You can check it out. He's on Twitter. You can check it out. He's on Twitter. He's posting a lot about this stuff now. He's posting a lot about this stuff now. He's posting a lot about this stuff now. But there's one thing that takes the But there's one thing that takes the But there's one thing that takes the edge off. You always see demos of, oh, edge off. You always see demos of, oh, edge off. You always see demos of, oh, uh, you know, these things when they're uh, you know, these things when they're uh, you know, these things when they're kicked off fresh, the prompt processing kicked off fresh, the prompt processing kicked off fresh, the prompt processing speed is so different. That's not how it speed is so different. That's not how it speed is so different. That's not how it happens in the real world. In a real happens in the real world. In a real happens in the real world. In a real scenario, in the same session, the scenario, in the same session, the scenario, in the same session, the servers keep what they've already read.

  11. servers keep what they've already read. servers keep what they've already read. There's a cache with 16,000 tokens of There's a cache with 16,000 tokens of There's a cache with 16,000 tokens of context. The first ask took 21 seconds context. The first ask took 21 seconds context. The first ask took 21 seconds on the Mac and 8.3 seconds on the on the Mac and 8.3 seconds on the on the Mac and 8.3 seconds on the Sparks. All right, we've already seen Sparks. All right, we've already seen Sparks. All right, we've already seen this part, but the follow-up question, 3 this part, but the follow-up question, 3 this part, but the follow-up question, 3 seconds on the Mac, 1.4 seconds on the seconds on the Mac, 1.4 seconds on the seconds on the Mac, 1.4 seconds on the Sparks. Yes, the sparks are still a Sparks. Yes, the sparks are still a Sparks. Yes, the sparks are still a little bit faster, but you pay for that little bit faster, but you pay for that little bit faster, but you pay for that long read only once per session, not long read only once per session, not long read only once per session, not every time. And I recently made a video every time. And I recently made a video every time. And I recently made a video for members of the channel uh detailing for members of the channel uh detailing for members of the channel uh detailing different techniques for how to speed different techniques for how to speed different techniques for how to speed things up and includes prefix caching. things up and includes prefix caching. things up and includes prefix caching. Thanks to the members of the channel, by Thanks to the members of the channel, by Thanks to the members of the channel, by the way. Really appreciate you. the way. Really appreciate you. the way. Really appreciate you. Sometimes they get extra videos uh when Sometimes they get extra videos uh when Sometimes they get extra videos uh when I get a chance to record them. I get a chance to record them. I get a chance to record them. Appreciate you all. Anyway, and that's Appreciate you all. Anyway, and that's Appreciate you all. Anyway, and that's how coding tools actually work. Your how coding tools actually work. Your how coding tools actually work. Your agent sends the code base once, the agent sends the code base once, the agent sends the code base once, the server keeps it, and every new question server keeps it, and every new question server keeps it, and every new question only adds a little bit on top. So that only adds a little bit on top. So that only adds a little bit on top. So that long wait is mostly an initial first long wait is mostly an initial first long wait is mostly an initial first question cost, not an every question question cost, not an every question question cost, not an every question cost. Where it hurts the most is when cost. Where it hurts the most is when cost. Where it hurts the most is when the context keeps changing. Obviously, a the context keeps changing. Obviously, a the context keeps changing. Obviously, a new repo, big new files, or a long new repo, big new files, or a long new repo, big new files, or a long session that outgrows the cache. Those session that outgrows the cache. Those session that outgrows the cache. Those are all possibilities. Another thing I are all possibilities. Another thing I are all possibilities. Another thing I found when running Frontier models is found when running Frontier models is found when running Frontier models is when you're changing a model like from when you're changing a model like from when you're changing a model like from Opus 5 to Opus 5.5 or Fable 1.1, you Opus 5 to Opus 5.5 or Fable 1.1, you Opus 5 to Opus 5.5 or Fable 1.1, you have to recalculate all that and the have to recalculate all that and the have to recalculate all that and the initial hit is longer usually. But we're initial hit is longer usually. But we're initial hit is longer usually. But we're not talking about Frontier models now.

  12. not talking about Frontier models now. not talking about Frontier models now. We're talking about local. All right. We're talking about local. All right. We're talking about local. All right. Same ideas though apply. Same ideas though apply. Same ideas though apply. Can MLX rescue the Mac on reading speed? Can MLX rescue the Mac on reading speed? Can MLX rescue the Mac on reading speed? Well, on Quen, it's kind of a split. Well, on Quen, it's kind of a split. Well, on Quen, it's kind of a split. Llama CPP writes faster and MLX reads Llama CPP writes faster and MLX reads Llama CPP writes faster and MLX reads faster. MLX gets through 32K in about 19 faster. MLX gets through 32K in about 19 faster. MLX gets through 32K in about 19 seconds instead of 26. That's still seconds instead of 26. That's still seconds instead of 26. That's still behind the Spark's 11 seconds. And for behind the Spark's 11 seconds. And for behind the Spark's 11 seconds. And for DeepSeek, even at MLX's best reading DeepSeek, even at MLX's best reading DeepSeek, even at MLX's best reading speed, 32,000 tokens would take at least speed, 32,000 tokens would take at least speed, 32,000 tokens would take at least 22 seconds, probably closer to 30. Yes, 22 seconds, probably closer to 30. Yes, 22 seconds, probably closer to 30. Yes, that's also slower than the Sparks. that's also slower than the Sparks. that's also slower than the Sparks. Whether MLX lets the Mac win that race Whether MLX lets the Mac win that race Whether MLX lets the Mac win that race back, I don't know yet. I haven't run back, I don't know yet. I haven't run back, I don't know yet. I haven't run deep yet on MLX with longer prompts or deep yet on MLX with longer prompts or deep yet on MLX with longer prompts or multiple users. told you there's a lot multiple users. told you there's a lot multiple users. told you there's a lot of tests to do and I'm crunching through of tests to do and I'm crunching through of tests to do and I'm crunching through them. them. them. More than one user or the number of More than one user or the number of More than one user or the number of concurrencies sometimes it's called concurrencies sometimes it's called concurrencies sometimes it's called you'll see a bigger number quoted you'll see a bigger number quoted you'll see a bigger number quoted usually and it's the total speed of the usually and it's the total speed of the usually and it's the total speed of the throughput of all the tokens generated throughput of all the tokens generated throughput of all the tokens generated for all the users. But careful with that for all the users. But careful with that for all the users. But careful with that number though because it shows number though because it shows number though because it shows everybody's number. It's everyone's everybody's number. It's everyone's everybody's number. It's everyone's tokens added up. Each person only gets a tokens added up. Each person only gets a tokens added up. Each person only gets a slice. The server writes for everyone slice. The server writes for everyone slice. The server writes for everyone together. And when somebody new shows together. And when somebody new shows together. And when somebody new shows up, well, the server has to stop and up, well, the server has to stop and up, well, the server has to stop and read their prompt, too, which uh is read their prompt, too, which uh is read their prompt, too, which uh is going to add to that initial going to add to that initial going to add to that initial calculation. On the Mac, that reading is calculation. On the Mac, that reading is calculation. On the Mac, that reading is slower past about four users. It's slower past about four users. It's slower past about four users. It's spending more time reading than writing.

  13. spending more time reading than writing. spending more time reading than writing. So, at that point, everyone's slice So, at that point, everyone's slice So, at that point, everyone's slice shrinks a little bit. And you can see it shrinks a little bit. And you can see it shrinks a little bit. And you can see it on the Mac, Deepseek peaks at four on the Mac, Deepseek peaks at four on the Mac, Deepseek peaks at four users, 66 tokens per second there. users, 66 tokens per second there. users, 66 tokens per second there. That's total. Then it drops to 46 tokens That's total. Then it drops to 46 tokens That's total. Then it drops to 46 tokens per second for eight users. But the per second for eight users. But the per second for eight users. But the sparks, they keep climbing to 70 tokens sparks, they keep climbing to 70 tokens sparks, they keep climbing to 70 tokens per second, and I haven't done 16, so I per second, and I haven't done 16, so I per second, and I haven't done 16, so I don't know where it drops off uh TBD. don't know where it drops off uh TBD. don't know where it drops off uh TBD. Now, give everybody 8,000 tokens of chat Now, give everybody 8,000 tokens of chat Now, give everybody 8,000 tokens of chat history, and at eight users, the Mac history, and at eight users, the Mac history, and at eight users, the Mac does about 11 tokens per second. The does about 11 tokens per second. The does about 11 tokens per second. The Sparks, 25, and that's total. But the Sparks, 25, and that's total. But the Sparks, 25, and that's total. But the real pain is the wait. Each person on real pain is the wait. Each person on real pain is the wait. Each person on the Mac waits over a minute for the the Mac waits over a minute for the the Mac waits over a minute for the first word and then gets about five first word and then gets about five first word and then gets about five tokens a second. on a Sparks it's about tokens a second. on a Sparks it's about tokens a second. on a Sparks it's about 24 second wait and then about seven 24 second wait and then about seven 24 second wait and then about seven tokens per second. So yeah, sharing tokens per second. So yeah, sharing tokens per second. So yeah, sharing these machines, these machines, these machines, keep it to yourself. All right, you can keep it to yourself. All right, you can keep it to yourself. All right, you can do it, but use smaller models maybe or do it, but use smaller models maybe or do it, but use smaller models maybe or yeah, there's there's different ways of yeah, there's there's different ways of yeah, there's there's different ways of splitting these machines up, but just splitting these machines up, but just splitting these machines up, but just don't use big models for a lot of don't use big models for a lot of don't use big models for a lot of people. It's not going to turn out so people. It's not going to turn out so people. It's not going to turn out so well. Looking at Quen here at eight well. Looking at Quen here at eight well. Looking at Quen here at eight users, the Sparks put out 120 tokens a users, the Sparks put out 120 tokens a users, the Sparks put out 120 tokens a second total. That's pretty good. Yeah, second total. That's pretty good. Yeah, second total. That's pretty good. Yeah, that's way better than DeepSeek. The Mac that's way better than DeepSeek. The Mac that's way better than DeepSeek. The Mac does 66 in Llama CVP and 70 in MLX. And does 66 in Llama CVP and 70 in MLX. And does 66 in Llama CVP and 70 in MLX. And for that, I used OMLX, which is a new for that, I used OMLX, which is a new for that, I used OMLX, which is a new tool, newish. There it is. No more tool, newish. There it is. No more tool, newish. There it is. No more waiting on your Mac. Well, there is some waiting on your Mac. Well, there is some waiting on your Mac. Well, there is some waiting. Okay. Really, the Mac is a waiting. Okay. Really, the Mac is a waiting. Okay. Really, the Mac is a machine for one person, maybe two. The machine for one person, maybe two. The machine for one person, maybe two. The Sparks are a little bit better for a Sparks are a little bit better for a Sparks are a little bit better for a small team, maybe four people. Eight.

  14. small team, maybe four people. Eight. small team, maybe four people. Eight. You're pushing it. Now, I didn't start You're pushing it. Now, I didn't start You're pushing it. Now, I didn't start out with two sparks. I went to four and out with two sparks. I went to four and out with two sparks. I went to four and then eight. That was a much bigger then eight. That was a much bigger then eight. That was a much bigger project. It was actually experimental, project. It was actually experimental, project. It was actually experimental, more like uh with breakout cables and more like uh with breakout cables and more like uh with breakout cables and extra switches that I had to buy. I made extra switches that I had to buy. I made extra switches that I had to buy. I made a whole video about it. Compared to a whole video about it. Compared to a whole video about it. Compared to that, two sparks is kind of a perfect that, two sparks is kind of a perfect that, two sparks is kind of a perfect setup. It's just one cable. And Nvidia's setup. It's just one cable. And Nvidia's setup. It's just one cable. And Nvidia's developer site build.envidia.com/spark developer site build.envidia.com/spark developer site build.envidia.com/spark has some really interesting recipes that has some really interesting recipes that has some really interesting recipes that are super easy to follow and it just are super easy to follow and it just are super easy to follow and it just works most of the time. Shows you how to works most of the time. Shows you how to works most of the time. Shows you how to connect two of them, how to run multiple connect two of them, how to run multiple connect two of them, how to run multiple workloads, vlm, everything. Pretty easy. workloads, vlm, everything. Pretty easy. workloads, vlm, everything. Pretty easy. You still have to line things up. I You still have to line things up. I You still have to line things up. I match the OS, the kernel, the driver, match the OS, the kernel, the driver, match the OS, the kernel, the driver, the firmware on both of the machines. the firmware on both of the machines. the firmware on both of the machines. Then you have to set up VLM and Then you have to set up VLM and Then you have to set up VLM and multi-node with a pile of nickel and multi-node with a pile of nickel and multi-node with a pile of nickel and rocky environment variables. After that, rocky environment variables. After that, rocky environment variables. After that, you have to check the traffic going over you have to check the traffic going over you have to check the traffic going over the RDMA connection. What you get for the RDMA connection. What you get for the RDMA connection. What you get for all that is CUDA and VLM. So basically, all that is CUDA and VLM. So basically, all that is CUDA and VLM. So basically, it's the same kind of stack that you'd it's the same kind of stack that you'd it's the same kind of stack that you'd run in the cloud. But on the Mac side, run in the cloud. But on the Mac side, run in the cloud. But on the Mac side, this machine isn't just for AI. It's this machine isn't just for AI. It's this machine isn't just for AI. It's also good at AI, but for example, my also good at AI, but for example, my also good at AI, but for example, my daily driver is an M5 Max, MacBook Pro.

  15. daily driver is an M5 Max, MacBook Pro. daily driver is an M5 Max, MacBook Pro. Everything that I do on that machine, I Everything that I do on that machine, I Everything that I do on that machine, I can do it on the M5 Ultra, but faster. can do it on the M5 Ultra, but faster. can do it on the M5 Ultra, but faster. That includes running large models and That includes running large models and That includes running large models and for everyday work. It's zero setup. for everyday work. It's zero setup. for everyday work. It's zero setup. Basically, I can offload a lot of stuff Basically, I can offload a lot of stuff Basically, I can offload a lot of stuff to it, like rendering my videos, for to it, like rendering my videos, for to it, like rendering my videos, for example, for this channel. And um example, for this channel. And um example, for this channel. And um sometimes the Mac gets a little bogged sometimes the Mac gets a little bogged sometimes the Mac gets a little bogged down. I run a lot of stuff on it. And down. I run a lot of stuff on it. And down. I run a lot of stuff on it. And right now I'm using 88 GB of memory. right now I'm using 88 GB of memory. right now I'm using 88 GB of memory. Yeah, it gets bugged down a little bit Yeah, it gets bugged down a little bit Yeah, it gets bugged down a little bit even with 128 gigs of memory. So, it's even with 128 gigs of memory. So, it's even with 128 gigs of memory. So, it's nice to be able to offload stuff very nice to be able to offload stuff very nice to be able to offload stuff very easily. And here is another thing. There easily. And here is another thing. There easily. And here is another thing. There they go. Full GPU utilization on both they go. Full GPU utilization on both they go. Full GPU utilization on both machines. Or yeah, you only see one machines. Or yeah, you only see one machines. Or yeah, you only see one here, but they're both working on the here, but they're both working on the here, but they're both working on the Sparks. And the Mac is using 100% of GPU Sparks. And the Mac is using 100% of GPU Sparks. And the Mac is using 100% of GPU as well. Oh boy. as well. Oh boy. as well. Oh boy. Now, during my M5 Ultra first look, I Now, during my M5 Ultra first look, I Now, during my M5 Ultra first look, I noticed that the machine was pretty noticed that the machine was pretty noticed that the machine was pretty toasty. And yeah, it is. But sparks also toasty. And yeah, it is. But sparks also toasty. And yeah, it is. But sparks also have been known to be toasty. You know, have been known to be toasty. You know, have been known to be toasty. You know, we're not alone here. Right now, I'm we're not alone here. Right now, I'm we're not alone here. Right now, I'm doing a very heavy load on the cluster doing a very heavy load on the cluster doing a very heavy load on the cluster here and the Mac Studio. And this is here and the Mac Studio. And this is here and the Mac Studio. And this is what I'm seeing. This is nuts. All what I'm seeing. This is nuts. All what I'm seeing. This is nuts. All right. Uh I don't want to pop a breaker right. Uh I don't want to pop a breaker right. Uh I don't want to pop a breaker here, but I might. 434 watts being used here, but I might. 434 watts being used here, but I might. 434 watts being used by the M5 Ultra and 410 by the Spark by the M5 Ultra and 410 by the Spark by the M5 Ultra and 410 by the Spark Cluster. So, we're very close. I should Cluster. So, we're very close. I should Cluster. So, we're very close. I should say at idle, it's a very different say at idle, it's a very different say at idle, it's a very different story. They are pretty warm now. And I'm story. They are pretty warm now. And I'm story. They are pretty warm now. And I'm hearing noise coming out of everywhere.

  16. hearing noise coming out of everywhere. hearing noise coming out of everywhere. Just like noise throughout. Just like noise throughout. Just like noise throughout. They feel about the same, actually. They feel about the same, actually. They feel about the same, actually. Yeah, they're both pretty orange. About Yeah, they're both pretty orange. About Yeah, they're both pretty orange. About 49° 49° 49° to 50° on the hottest part of the Mac to 50° on the hottest part of the Mac to 50° on the hottest part of the Mac Studio. About 48° Studio. About 48° Studio. About 48° on the sparks. Let's take a look at the on the sparks. Let's take a look at the on the sparks. Let's take a look at the back. Ooh, 56° on the grill, the back of back. Ooh, 56° on the grill, the back of back. Ooh, 56° on the grill, the back of the Mac Studio. And wow, 5860 the Mac Studio. And wow, 5860 the Mac Studio. And wow, 5860 63 I saw in there on the Sparks. So, 63 I saw in there on the Sparks. So, 63 I saw in there on the Sparks. So, yeah, both get pretty toasty. For AI on yeah, both get pretty toasty. For AI on yeah, both get pretty toasty. For AI on the Mac, there's a little setup the Mac, there's a little setup the Mac, there's a little setup depending on which stack you want. Llama depending on which stack you want. Llama depending on which stack you want. Llama CPP or MLX. OMLX is pretty easy. There's CPP or MLX. OMLX is pretty easy. There's CPP or MLX. OMLX is pretty easy. There's also newer ones like native without the also newer ones like native without the also newer ones like native without the E. And there's a splash. Some of these E. And there's a splash. Some of these E. And there's a splash. Some of these are so new I haven't even tried them are so new I haven't even tried them are so new I haven't even tried them yet. These are basically homebrew yet. These are basically homebrew yet. These are basically homebrew installs. So, is 256 gigs across two installs. So, is 256 gigs across two installs. So, is 256 gigs across two boxes the same as 256 in one? Well, it boxes the same as 256 in one? Well, it boxes the same as 256 in one? Well, it holds the same model. It just reads a holds the same model. It just reads a holds the same model. It just reads a lot faster. But before you go out and lot faster. But before you go out and lot faster. But before you go out and drop 10 grand or more on one of these drop 10 grand or more on one of these drop 10 grand or more on one of these setups, look at your own work. How big setups, look at your own work. How big setups, look at your own work. How big are your prompts? How long are the are your prompts? How long are the are your prompts? How long are the answers to your prompts? You want to answers to your prompts? You want to answers to your prompts? You want to learn more about clustering the sparks?

  17. learn more about clustering the sparks? learn more about clustering the sparks? Watch this video here. Clustering Mac Watch this video here. Clustering Mac Watch this video here. Clustering Mac Studios, watch this video here. Thanks Studios, watch this video here. Thanks Studios, watch this video here. Thanks for watching and I'll see you next time.

Summary

This analysis compares the performance of Apple's M5 Ultra and a dual DGX Spark cluster for AI tasks, highlighting their similar initial speeds but emphasizing their differing price points and architectures. The practical takeaway is that for developers running large models or serving small teams, understanding these hardware differences is crucial, and tools like Merlin AI can streamline AI workflows by consolidating multiple models into one interface.

View original episode ↗