← Back
AI Engineer October 10, 2026 15m

How MiniMax M3 Was Built: Sparse Attention and Native Multimodality — Olive Song

Read full transcript 12 segments
  1. So, okay. So, So, okay. So, hello everyone. My hello everyone. My hello everyone. My name is Olive. I work at name is Olive. I work at name is Olive. I work at Minimax on Minimax on reinforcement learning research. Today I'll talk a little Today I'll talk a little about our about our about our new Minimax M3 model, new Minimax M3 model, new Minimax M3 model, how we trained how we trained how we trained it as natively it as natively it as natively multimodal from the multimodal from the multimodal from the start, and how start, and how start, and how we used MSA we used MSA , i.e. Minimax sparse attention. Let , i.e. Minimax sparse attention. Let me tell you a little about me tell you a little about our team. We are one our team. We are one our team. We are one of the few of the few of the few laboratories in the world laboratories in the world laboratories in the world that works with that works with that works with multimodality, multimodality, multimodality, which means that we have already which means that we have already which means that we have already published our LLM published our LLM models before. We models before. We models before. We worked on our worked on our worked on our speech speech speech models. We have our models. We have our models. We have our own own own video generation models and video generation models and video generation models and models for agents models for agents models for agents writing code. writing code. We recently We recently released the Minimax M3—a released the Minimax M3—a model with open model with open model with open scales and advanced scales and advanced scales and advanced performance. This is performance. This is performance. This is actually the first actually the first actually the first open- open- open- scale model that combines the scale model that combines the scale model that combines the following three following three following three advanced features.

  2. advanced features. The first is what we The first is what we call the best call the best call the best coding coding coding and agent capabilities. The model and agent capabilities. The model and agent capabilities. The model is capable of performing is capable of performing is capable of performing many agent many agent many agent tasks you tasks you tasks you can imagine. can imagine. can imagine. It can It can It can automatically automatically automatically break down tasks break down tasks break down tasks into parts. It can into parts. It can into parts. It can work in your work in your work in your coding agents. coding agents. coding agents. It can replace 90% of It can replace 90% of your current your current your current code work. code work. The second pillar is the The second pillar is the context window of the context window of the context window of the 1 million 1 million 1 million token model. It is very token model. It is very token model. It is very large, which opens up large, which opens up large, which opens up the possibility for almost the possibility for almost the possibility for almost all basic all basic all basic long-term long-term long-term user interactions, user interactions, user interactions, such as processing such as processing such as processing entire code bases. We entire code bases. We entire code bases. We trained it trained it trained it using MSA, i.e. using MSA, i.e. using MSA, i.e. Minimax sparse attention. Minimax sparse attention. Minimax sparse attention. This makes This makes This makes pre-filling pre-filling pre-filling 9 9 9 times faster and times faster and times faster and decoding 15 decoding 15 decoding 15 times faster times faster times faster compared to full compared to full compared to full attention. Third, the model attention. Third, the model attention. Third, the model was trained as was trained as was trained as natively natively natively multimodal from multimodal from multimodal from scratch. This scratch. This scratch. This means that we have means that we have combined visual combined visual understanding with understanding with understanding with language generation from the very beginning. Therefore, language generation from the very beginning. Therefore, language generation from the very beginning. Therefore, the model the model understands both of these understands both of these possibilities very well from the very beginning. And this possibilities very well from the very beginning. And this possibilities very well from the very beginning. And this allows us allows us allows us to build to build to build even more even more even more applications based on it.

  3. applications based on it. applications based on it. For example, agents For example, agents For example, agents for controlling for controlling for controlling a computer, right? a computer, right? So, first I'll So, first I'll briefly explain what briefly explain what sparse attention does. Our sparse attention does. Our basic basic basic design principle is that a design principle is that a design principle is that a simpler simpler simpler architecture always architecture always architecture always leads to better leads to better scaling properties. scaling properties. We adhered to this principle We adhered to this principle throughout the entire throughout the entire design process design process . Our goal is to always . Our goal is to always . Our goal is to always strive for cleaner strive for cleaner strive for cleaner and simpler and simpler and simpler architecture. architecture. Sparse attention Sparse attention consists of two consists of two consists of two branches. We call the first one the branches. We call the first one the branches. We call the first one the index index index branch. It performs branch. It performs branch. It performs lightweight lightweight lightweight calculations with full calculations with full calculations with full attention. It attention. It attention. It uses four uses four uses four query query query heads and one heads and one heads and one key head. It is key head. It is key head. It is trained trained trained using KL- using KL- divergence to divergence to divergence to match the match the match the attention metrics of the attention metrics of the attention metrics of the main branch. The main branch. The index branch index branch is used to is used to is used to identify tokens with the identify tokens with the identify tokens with the highest highest highest attention scores. These attention scores. These attention scores. These selected tokens selected tokens selected tokens are then processed by the are then processed by the are then processed by the main branch. main branch. In other words, a In other words, a sparse branch does not sparse branch does not sparse branch does not pay attention to all pay attention to all pay attention to all tokens. But only on a tokens. But only on a tokens. But only on a subset of them. This subset of them. This subset of them. This subset subset subset is selected is selected is selected using easy using easy using easy index index index branch calculations. The branch calculations. The branch calculations. The sparse branch then sparse branch then sparse branch then uses the blocks uses the blocks uses the blocks selected by the index selected by the index selected by the index branch to determine branch to determine which keys and which keys and which keys and values ​​are needed values ​​are needed values ​​are needed to compute the attention.

  4. to compute the attention. OK. And I want OK. And I want to emphasize today to emphasize today that our design is not that our design is not that our design is not purely algorithmically purely algorithmically purely algorithmically oriented. oriented. Instead, we Instead, we took a took a took a co- co- co- design approach to design approach to design approach to algorithms and algorithms and algorithms and infrastructure. Um, a infrastructure. Um, a infrastructure. Um, a good example good example good example to start with is DSA. DSA to start with is DSA. DSA to start with is DSA. DSA was a very powerful was a very powerful was a very powerful solution, but it was solution, but it was solution, but it was also heavily also heavily also heavily optimized for optimized for Deep Seek's own architecture, especially the Deep Seek's own architecture, especially the combination of MQA and combination of MQA and combination of MQA and high high high head dimensionality. However, as we head dimensionality. However, as we head dimensionality. However, as we move to the more move to the more move to the more common common common GQA architecture, the direct GQA architecture, the direct GQA architecture, the direct application of the same application of the same application of the same design becomes less design becomes less design becomes less obvious. This poses obvious. This poses obvious. This poses several practical several practical several practical challenges, especially from an challenges, especially from an infrastructure and infrastructure and implementation perspective. Therefore, this implementation perspective. Therefore, this implementation perspective. Therefore, this determines our determines our determines our design choices. Rather than design choices. Rather than simply simply simply borrowing an existing borrowing an existing borrowing an existing sparse sparse sparse attention design, we attention design, we attention design, we reimagined reimagined reimagined the architecture within the the architecture within the the architecture within the constraints of GQA and constraints of GQA and real-world system efficiency.

  5. So, what are the problems So, what are the problems with with with directly applying directly applying directly applying existing existing existing sparse attention architectures? sparse attention architectures? First, from an First, from an First, from an algorithmic algorithmic algorithmic perspective, GQA has multiple KV perspective, GQA has multiple KV goals, which ideally goals, which ideally goals, which ideally should should should provide greater provide greater provide greater diversity between diversity between diversity between different KV groups. different KV groups. different KV groups. However, if we However, if we However, if we directly directly reuse the same token selection strategy, multiple KV goals will eventually goals will eventually goals will eventually pay attention pay attention pay attention to the same to the same to the same selected tokens. This selected tokens. This selected tokens. This essentially negates essentially negates essentially negates the diversity that the diversity that the diversity that GQA is supposed to provide. GQA is supposed to provide. Secondly, from an Secondly, from an infrastructure perspective, this infrastructure perspective, this infrastructure perspective, this design is not very design is not very design is not very friendly to friendly to friendly to GPU memory access. GPUs are much more GPU memory access. GPUs are much more GPU memory access. GPUs are much more efficient at efficient at efficient at reading continuous reading continuous reading continuous data. In data. In data. In MQA settings, we MQA settings, we MQA settings, we can consider a can consider a can consider a KV layout as a single KV KV layout as a single KV head with a large head with a large head with a large dimension, dimension, dimension, such as 1 by 512. This is such as 1 by 512. This is such as 1 by 512. This is relatively continuous relatively continuous relatively continuous and efficient to and efficient to and efficient to read. However, in GQA, the read. However, in GQA, the read. However, in GQA, the same overall same overall same overall dimension can dimension can dimension can be divided between be divided between be divided between multiple KV heads.

  6. multiple KV heads. multiple KV heads. For example, this is 4 by 128. For example, this is 4 by 128. For example, this is 4 by 128. This is no longer one continuous This is no longer one continuous This is no longer one continuous continuous continuous continuous reading. Instead, reading. Instead, reading. Instead, the GPU has to the GPU has to the GPU has to perform multiple perform multiple perform multiple independent reads independent reads independent reads for different KV heads. This for different KV heads. This for different KV heads. This increases the overhead of increases the overhead of memory access and reduces memory access and reduces the ratio the ratio the ratio of computation to of computation to of computation to memory access. And another memory access. And another memory access. And another problem here is the problem here is the problem here is the calculation of the top K. To calculation of the top K. To calculation of the top K. To select the most important select the most important select the most important tokens, the indexer tokens, the indexer tokens, the indexer must calculate the top K must calculate the top K must calculate the top K scores for a large scores for a large scores for a large number of number of number of token-level candidates. token-level candidates. However, top K has However, top K has internal internal internal data dependencies. Therefore, it is data dependencies. Therefore, it is data dependencies. Therefore, it is difficult to fully difficult to fully difficult to fully parallelize it on GPUs. As parallelize it on GPUs. As parallelize it on GPUs. As a result, a result, a result, token-level selection token-level selection token-level selection introduces significant introduces significant introduces significant computational computational computational overhead and overhead and overhead and association costs association costs . So how did we solve . So how did we solve . So how did we solve these problems at MSA? From these problems at MSA? From these problems at MSA? From the algorithm side, we the algorithm side, we the algorithm side, we remove the aggregation of remove the aggregation of remove the aggregation of index index index branch outputs. In the original branch outputs. In the original branch outputs. In the original design, the results of design, the results of design, the results of all index all index all index branch heads were combined into a branch heads were combined into a branch heads were combined into a single set of single set of single set of selected tokens. But selected tokens. But , as we discussed , as we discussed , as we discussed earlier, this would destroy earlier, this would destroy earlier, this would destroy the diversity that the diversity that the diversity that GQA provides.

  7. GQA provides. Instead, we make Instead, we make the choice different for the choice different for the choice different for different KV goals. different KV goals. different KV goals. According to our According to our According to our sparse sparse sparse attention design, when attention design, when attention design, when using GQA, different using GQA, different using GQA, different KV heads can KV heads can KV heads can focus on focus on focus on different sets of different sets of different sets of tokens. This helps tokens. This helps tokens. This helps maintain maintain maintain the diversity of GQA, the diversity of GQA, the diversity of GQA, instead of instead of instead of forcing all heads forcing all heads forcing all heads to share the same to share the same to share the same sparse attention. For the sparse attention. For the sparse attention. For the infrastructure infrastructure infrastructure part, we moved part, we moved part, we moved from token-level selection from token-level selection from token-level selection to to to block-level search. block-level search. Compared to Compared to token-level selection, token-level selection, token-level selection, block-level search block-level search block-level search significantly reduces significantly reduces significantly reduces the number of candidates the number of candidates the number of candidates and reduces the cost of and reduces the cost of and reduces the cost of computing top-k and computing top-k and computing top-k and merging merging merging results. results. At the same time, At the same time, block-level search is more block-level search is more hardware-friendly. It hardware-friendly. It improves improves memory access efficiency and memory access efficiency and increases the ratio of increases the ratio of increases the ratio of computations to computations to computations to memory accesses. memory accesses. memory accesses. Importantly, we also Importantly, we also Importantly, we also show that such a show that such a show that such a design does not result design does not result design does not result in a significant in a significant in a significant performance drop. performance drop. Ultimately, we developed Ultimately, we developed efficient kernels for both efficient kernels for both efficient kernels for both training and training and training and inference. I won't inference. I won't inference. I won't go into the details of how go into the details of how go into the details of how kernels work today, kernels work today, kernels work today, but it is a system but it is a system but it is a system component that component that component that makes makes makes sparse attention design sparse attention design sparse attention design practical and practical and practical and efficient. So, with efficient. So, with efficient. So, with this design, we this design, we this design, we were able to very were able to very were able to very effectively scale effectively scale effectively scale the training of our the training of our the training of our M3 model to 1 million M3 model to 1 million M3 model to 1 million contexts. And, as we

  8. contexts. And, as we contexts. And, as we said before, the said before, the said before, the third pillar is third pillar is third pillar is native native native multimodality. multimodality. multimodality. Our model, Minimax M3, Our model, Minimax M3, Our model, Minimax M3, was trained from scratch on was trained from scratch on was trained from scratch on multimodal multimodal multimodal data. What does this mean? data. What does this mean? data. What does this mean? Our model Our model Our model was trained was trained was trained to understand text and to understand text and to understand text and images simultaneously from the images simultaneously from the images simultaneously from the start, which is a completely start, which is a completely start, which is a completely different approach than different approach than different approach than current current current open-ended models. Our open-ended models. Our open-ended models. Our starting point is very starting point is very starting point is very clear. At the current clear. At the current clear. At the current stage, a separate VL model stage, a separate VL model stage, a separate VL model has relatively has relatively has relatively limited limited limited application scenarios. But application scenarios. But application scenarios. But combining it with combining it with combining it with understanding the text understanding the text understanding the text while studying can while studying can while studying can be a be a be a challenge. Therefore, the challenge. Therefore, the challenge. Therefore, the key question and key question and key question and challenge was how to challenge was how to challenge was how to implement implement implement multimodal multimodal multimodal capabilities without capabilities without capabilities without damaging the core damaging the core text productivity? To text productivity? To answer this, we answer this, we answer this, we explored several explored several large-scale learning strategies. The first and first and most common most common most common approach is to add approach is to add approach is to add multimodal multimodal multimodal training after the training after the training after the pre- pre- pre- training stage, which is essentially a training stage, which is essentially a training stage, which is essentially a CPT-style approach.

  9. CPT-style approach. However, by that time However, by that time the model had already largely the model had already largely the model had already largely converged. converged. Adding Adding multimodal multimodal multimodal data at this stage data at this stage data at this stage usually interferes with usually interferes with usually interferes with the learned textual the learned textual the learned textual capabilities and capabilities and capabilities and causes degradation causes degradation . Another approach . Another approach . Another approach is to move to is to move to is to move to multimodal multimodal multimodal learning before the learning before the learning before the pre- pre- pre- learning stage. For example, learning stage. For example, learning stage. For example, after the model has after the model has after the model has already processed a certain already processed a certain already processed a certain number of text number of text number of text tokens. We also tokens. We also tokens. We also tested this tested this tested this approach, but found approach, but found approach, but found that the results that the results that the results depended heavily on the mix of depended heavily on the mix of depended heavily on the mix of data and data and training hyperparameters, such as the training speed. In some training speed. In some settings, a slower settings, a slower settings, a slower learning rate learning rate learning rate may reduce the may reduce the performance drop, but performance drop, but this conclusion is not this conclusion is not this conclusion is not necessarily necessarily necessarily true for true for true for another run. More another run. More another run. More importantly, during importantly, during importantly, during large-scale large-scale large-scale pre- pre- pre- training, these training, these training, these parameters cannot be parameters cannot be parameters cannot be arbitrarily changed arbitrarily changed arbitrarily changed just for the sake of just for the sake of just for the sake of multimodal multimodal multimodal training. And that training. And that training. And that led us to our led us to our led us to our final strategy: final strategy: final strategy: native native native multimodal multimodal multimodal learning from scratch. With learning from scratch. With learning from scratch. With this approach, this approach, this approach, multimodal multimodal multimodal capabilities capabilities capabilities are implemented from the are implemented from the are implemented from the very beginning, rather than very beginning, rather than very beginning, rather than added after the added after the added after the text model has text model has text model has already formed its already formed its already formed its structure. With structure. With structure. With the same the same the same text token budget, we text token budget, we text token budget, we found that this found that this found that this approach does little approach does little approach does little to harm text to harm text to harm text capabilities.

  10. capabilities. In fact, during these In fact, during these experiments, we experiments, we experiments, we conducted many conducted many conducted many interesting analyses. interesting analyses. One of them is an One of them is an interesting observation interesting observation interesting observation based on attention maps. based on attention maps. As you can see from this As you can see from this image, in image, in image, in multimodal multimodal multimodal models like CPT, models like CPT, models like CPT, text tokens text tokens text tokens often often often pay little attention to pay little attention to pay little attention to visual tokens. visual tokens. The areas of visual The areas of visual tokens on the tokens on the tokens on the attention map are mostly attention map are mostly attention map are mostly dark. This suggests dark. This suggests dark. This suggests that that that multimodalities are not multimodalities are not multimodalities are not fully integrated fully integrated . In contrast, with . In contrast, with . In contrast, with native native native multimodal multimodal multimodal learning, the attention map learning, the attention map learning, the attention map appears much appears much appears much more integrated. more integrated. We don't see a We don't see a clear separation clear separation clear separation between text and between text and between text and visual tokens here visual tokens here . This suggests . This suggests . This suggests that the model is learning to that the model is learning to combine visual combine visual and textual and textual and textual information more naturally from the information more naturally from the information more naturally from the start. So, start. So, start. So, overall, our overall, our overall, our learning strategy learning strategy learning strategy is not just about is not just about is not just about adding adding adding visual capabilities.

  11. visual capabilities. It is about It is about integrating integrating integrating multimodality into the multimodality into the multimodality into the model while model while model while preserving its preserving its preserving its strongest strongest strongest textual textual textual capabilities. And, capabilities. And, capabilities. And, of course, we did of course, we did of course, we did a lot of experiments a lot of experiments a lot of experiments with choosing a with choosing a with choosing a multimodal multimodal multimodal architecture. Our architecture. Our architecture. Our overall architecture overall architecture overall architecture is relatively is relatively is relatively standard. standard. standard. The image or video The image or video The image or video is processed is processed is processed using ViT, which using ViT, which using ViT, which generates visual generates visual generates visual tokens. These tokens. These tokens. These visual tokens visual tokens visual tokens are then inserted into the are then inserted into the are then inserted into the input of the language input of the language input of the language model, replacing the model, replacing the model, replacing the original original image placeholders in the image placeholders in the text text text sequence. After sequence. After sequence. After that, the visual and that, the visual and that, the visual and text embeddings are text embeddings are text embeddings are fed together into the fed together into the fed together into the language model. One language model. One language model. One important important important finding is that finding is that finding is that applying 3D attention applying 3D attention applying 3D attention within ViT is very within ViT is very within ViT is very beneficial for beneficial for beneficial for visual visual visual comprehension. In the comprehension. In the comprehension. In the default default default ViT setting, attention ViT setting, attention ViT setting, attention is applied is applied is applied within each within each within each individual image individual image individual image or frame. For or frame. For or frame. For video inputs, this video inputs, this video inputs, this means that visual means that visual means that visual fragments from different fragments from different fragments from different frames do not frames do not frames do not interact interact interact directly at the directly at the directly at the ViT stage. But thanks to ViT stage. But thanks to ViT stage. But thanks to 3D attention, fragments from 3D attention, fragments from 3D attention, fragments from different frames can different frames can different frames can interact with each interact with each interact with each other directly.

  12. other directly. This gives the model This gives the model better capabilities better capabilities better capabilities for temporal and for temporal and for temporal and interframe interframe interframe visual visual visual modeling. And we modeling. And we modeling. And we found that it found that it found that it improves visual improves visual improves visual comprehension, and the benefit comprehension, and the benefit comprehension, and the benefit extends to extends to extends to some some some imaging benchmarks as well. This imaging benchmarks as well. This imaging benchmarks as well. This suggests that suggests that suggests that it enhances the it enhances the it enhances the overall visual overall visual overall visual presentation, rather than presentation, rather than presentation, rather than just improving just improving just improving specific specific specific video tasks. And one video tasks. And one video tasks. And one practical detail practical detail practical detail is that we is that we is that we can't simply can't simply can't simply extend 3D attention to an extend 3D attention to an extend 3D attention to an unlimited number of unlimited number of unlimited number of frames, as that would frames, as that would frames, as that would create additional create additional create additional system calls and system calls and load balancing issues. In load balancing issues. In our study, our study, our study, using using using four frames four frames four frames provided the provided the provided the best balance between best balance between best balance between performance and performance and performance and efficiency. So, efficiency. So, efficiency. So, thanks to the thanks to the thanks to the architectural architectural architectural design, the design, the design, the training strategies we training strategies we training strategies we tested at tested at tested at different stages, and different stages, and different stages, and our our our ViT analysis and research, ViT analysis and research, ViT analysis and research, we were able to train a we were able to train a we were able to train a model from scratch with model from scratch with model from scratch with the ability to the ability to understand multimodally. This understand multimodally. This concludes concludes concludes today's today's today's presentation. Thank you.

Summary

This tech transcript introduces the Minimax M3 model, a natively multimodal AI trained from scratch. Key features highlighted are its advanced coding and agent capabilities, a 1 million token context window, and its use of Minimax Sparse Attention (MSA) for significantly faster processing. The practical takeaway is that M3 enables building more sophisticated applications by seamlessly integrating visual and language understanding.

View original episode ↗