{"notes":[{"content":{"submission_length":{"value":"Long submission (more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"code":{"value":"https://github.com/BICLab/SpikingBrain-7B"},"supplementary_material":{"value":"/attachment/67f070fcc827ac32ac4614eeaa968673de8b8856.zip"},"abstract":{"value":"Mainstream Transformer-based large language models (LLMs) face significant efficiency bottlenecks: training computation scales quadratically with sequence length, and inference memory grows linearly. These constraints limit their ability to process long sequences effectively. In addition, building large models on non-NVIDIA computing platforms poses major challenges in achieving stable and efficient training and deployment. To address these issues, we introduce SpikingBrain, a new family of brain-inspired models designed for efficient long-context training and inference. SpikingBrain leverages the MetaX GPU cluster and focuses on three core aspects: (1) Model Architecture: linear and hybrid-linear attention architectures with adaptive spiking neurons; (2) Algorithmic Optimizations: an efficient, conversion-based training pipeline compatible with existing LLMs, along with a dedicated spike coding framework; (3) System Engineering: customized training frameworks, operator libraries, and parallelism strategies tailored to the MetaX hardware. Using these techniques, we develop two models: SpikingBrain-7B, a linear LLM, and SpikingBrain-76B, a hybrid-linear MoE LLM. These models demonstrate the feasibility of large-scale LLM development on non-NVIDIA platforms, and our training framework supports weeks of stable training on hundreds of MetaX GPUs with Model FLOPs Utilization (MFU) at expected levels. SpikingBrain achieves performance comparable to open-source Transformer baselines while using exceptionally low data resources (continual pre-training of approximately 150B tokens). Our models also significantly improve long-context efficiency and deliver inference with (partially) constant memory and event-driven spiking behavior. For example, SpikingBrain-7B achieves more than 100× speedup in Time to First Token (TTFT) for 4M-token sequences. Furthermore, the proposed spiking scheme achieves 69.15% sparsity, enabling low-power operation. Overall, this work demonstrates the potential of brain-inspired mechanisms to drive the next generation of efficient and scalable large model design."},"_bibtex":{"value":"@article{\npan2026spikingbrain,\ntitle={SpikingBrain: Spiking Brain-inspired Large Models},\nauthor={Yuqi Pan and Yupeng Feng and JingHao Zhuang and siyu ding and Han Xu and Zehao Liu and Bohan Sun and Yuhong Chou and Xuerui Qiu and Anlin Deng and Anjie Hu and Shurong Wang and Peng Zhou and Man Yao and Jibin Wu and jian yang and 孙国梁 and Bo XU and Guoqi Li},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2026},\nurl={https://openreview.net/forum?id=PNLShc0C6q},\nnote={}\n}"},"title":{"value":"SpikingBrain: Spiking Brain-inspired Large Models"},"pdf":{"value":"/pdf/5d8957628e08647be20fb705242fa90fd8e89a35.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"pan|spikingbrain_spiking_braininspired_large_models"},"authorids":{"value":["~Yuqi_Pan2","~Yupeng_Feng1","~JingHao_Zhuang1","~siyu_ding1","~Han_Xu12","~Zehao_Liu4","~Bohan_Sun2","~Yuhong_Chou1","~Xuerui_Qiu1","~Anlin_Deng1","~Anjie_Hu1","~Shurong_Wang1","~Peng_Zhou10","~Man_Yao1","~Jibin_Wu1","~jian_yang38","~孙国梁2","~Bo_XU10","~Guoqi_Li1"]},"assigned_action_editor":{"value":"~Binhang_Yuan1"},"authors":{"value":["Yuqi Pan","Yupeng Feng","JingHao Zhuang","siyu ding","Han Xu","Zehao Liu","Bohan Sun","Yuhong Chou","Xuerui Qiu","Anlin Deng","Anjie Hu","Shurong Wang","Peng Zhou","Man Yao","Jibin Wu","jian yang","孙国梁","Bo XU","Guoqi Li"]}},"tmdate":1774441182935,"pdate":1774441182818,"tcdate":1763804612134,"writers":["TMLR"],"signatures":["TMLR/Paper6607/Authors"],"forum":"PNLShc0C6q","license":"CC BY 4.0","number":6607,"cdate":1763804612134,"readers":["everyone"],"invitations":["TMLR/-/Submission","TMLR/Paper6607/-/Revision","TMLR/-/Edit","TMLR/-/Under_Review","TMLR/Paper6607/-/Camera_Ready_Revision","TMLR/-/Accepted"],"mdate":1774441182935,"odate":1764120076639,"domain":"TMLR","id":"PNLShc0C6q","version":2},{"content":{"comment":{"value":">  **Challenges 4: Experiment setups of long-context evaluations.**\n\nWe thank the reviewer for the careful examination of our experimental setups. To ensure transparency and reproducibility, we provide detailed hardware configurations and interconnect information for all long-context benchmarking experiments.\n\n1. 7B model TTFT comparison under SP up to 4M tokens.  \nThese experiments were conducted on **NVIDIA H100 80GB GPUs**. This hardware platform was selected because support for the ZeCO operator within the MetaX hardware ecosystem is still under active development.  \nWithin each GPU node, devices are interconnected via **NVIDIA 4th-generation NVLink**, while inter-node communication is enabled through **InfiniBand (400G)**.\n\n2. 7B and 76B model TTFT comparison without SP up to 128k tokens.    \nThese experiments were conducted on **MetaX C550 GPUs**. We adapted both SpikingBrain models to the HuggingFace and vLLM frameworks within the MetaX environment.\nWithin each GPU node, devices are connected via **MetaX MetaLink**, and inter-node communication is provided by **InfiniBand (400G)**.\n\n3. 7B model training TGS comparison under SP.  \nThe hardware configuration is identical to that in (2), and experiments were also conducted on **MetaX C550 GPUs**.\n\nRegarding the hardware used for the 4M-token benchmarking, this information was originally stated in the caption of Table 6 (later moved to the Appendix). **We will ensure that all such hardware and configuration details are explicitly documented in the main text of the revised manuscript**.\n\n\n>  **Requested Changes 1: Table 1 is dangerously misleading.**\n\nWe apologize for any confusion. In Tables 1, bolded numbers is used solely to highlight the results of the SpikingBrain models and to distinguish them from the baselines. To avoid ambiguity, we will replace bolded numbers with a gray background for SpikingBrain entries in the revised manuscript. Additionally, as the reviewer suggested, we use the boldface font for the best performing model (which is mostly Qwen-2.5).\n\n>  **Requested Changes 2: Evaluation of recall capability.**\n\nplease see our response to the *Claim Evidence Challenge 2*.\n\n>  **Requested Changes 3: Quantitative comparisons of MFU.**\n\nThank you for the rigorous suggestion. In the abstract, our reference to “expected levels” was intended to indicate that achieving MFU above 20% in multi-node Megatron training is generally considered an acceptable practical standard. We will revise the wording to avoid any potential misunderstanding.\n\nRegarding the requested quantitative comparisons, we conducted the following additional experiments:\n\n1. **MFU evaluation on NVIDIA A800 80GB under the same parallel strategy.**  \nUnder identical parallel configurations, the measured MFU on **NVIDIA A800 80GB** reaches **25.8% (TGS 1913)**, which is slightly higher than the result on **MetaX C550**, i.e., **23.4% (TGS 1558)**. This indicates that the NVIDIA ecosystem currently benefits from more mature optimization, particularly in communication and kernel-level efficiency, while there remains room for further optimization in the MetaX stack.  \nHowever, this comparison is not strictly apples-to-apples. Due to the current state of Megatron support on MetaX hardware, we performed substantial interface adaptations and relied on **older software versions (torch 2.1.2 and triton 2.1.0)**, which likely contributed to suboptimal MFU. In contrast, on NVIDIA A800, we followed the recommended Megatron setup and used **newer versions (torch 2.5.0 and triton 3.0.0)**.\n\n2. **Overall model speed and operator-level comparison on MetaX C550 and NVIDIA A100.**  \nWe further compared SpikingBrain-7B on MetaX C550 and NVIDIA A100, including both **end-to-end model speed** and **operator-level performance**, while ensuring a fully fair comparison covering environment, model, and framework. Details can be found in our response to *Reviewer gn98, Requested Changes 4*.  \nOverall, we verify that **SpikingBrain runs efficiently on both NVIDIA A100 and MetaX C550** using the same implementation at the operator level, and the **end-to-end latency remains within ±2%** across a wide range of tensor shapes. We also observe that **linear-attention kernel optimization** represents the most direct path to further narrowing the remaining overall performance gap.\n\n\n>  **Requested Changes 4: Writing changes.**\n\nWe will revise the manuscript accordingly to address the suggested writing improvements."},"title":{"value":"[4/6] Response to Reviewer WxpP (Claim Evidence Challenges 4 and Requested Changes 1~4)"}},"parentInvitations":"TMLR/-/Official_Comment","tmdate":1770809327223,"tcdate":1770809327223,"writers":["TMLR","TMLR/Paper6607/Authors"],"signatures":["TMLR/Paper6607/Authors"],"forum":"PNLShc0C6q","number":15,"license":"CC BY 4.0","cdate":1770809327223,"readers":["everyone"],"invitations":["TMLR/Paper6607/-/Official_Comment"],"mdate":1770809327223,"domain":"TMLR","replyto":"MwHdMjWBjW","id":"ZnDksAV5Lo","forumContent":{"submission_length":{"value":"Long submission (more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"code":{"value":"https://github.com/BICLab/SpikingBrain-7B"},"supplementary_material":{"value":"/attachment/67f070fcc827ac32ac4614eeaa968673de8b8856.zip"},"abstract":{"value":"Mainstream Transformer-based large language models (LLMs) face significant efficiency bottlenecks: training computation scales quadratically with sequence length, and inference memory grows linearly. These constraints limit their ability to process long sequences effectively. In addition, building large models on non-NVIDIA computing platforms poses major challenges in achieving stable and efficient training and deployment. To address these issues, we introduce SpikingBrain, a new family of brain-inspired models designed for efficient long-context training and inference. SpikingBrain leverages the MetaX GPU cluster and focuses on three core aspects: (1) Model Architecture: linear and hybrid-linear attention architectures with adaptive spiking neurons; (2) Algorithmic Optimizations: an efficient, conversion-based training pipeline compatible with existing LLMs, along with a dedicated spike coding framework; (3) System Engineering: customized training frameworks, operator libraries, and parallelism strategies tailored to the MetaX hardware. Using these techniques, we develop two models: SpikingBrain-7B, a linear LLM, and SpikingBrain-76B, a hybrid-linear MoE LLM. These models demonstrate the feasibility of large-scale LLM development on non-NVIDIA platforms, and our training framework supports weeks of stable training on hundreds of MetaX GPUs with Model FLOPs Utilization (MFU) at expected levels. SpikingBrain achieves performance comparable to open-source Transformer baselines while using exceptionally low data resources (continual pre-training of approximately 150B tokens). Our models also significantly improve long-context efficiency and deliver inference with (partially) constant memory and event-driven spiking behavior. For example, SpikingBrain-7B achieves more than 100× speedup in Time to First Token (TTFT) for 4M-token sequences. Furthermore, the proposed spiking scheme achieves 69.15% sparsity, enabling low-power operation. Overall, this work demonstrates the potential of brain-inspired mechanisms to drive the next generation of efficient and scalable large model design."},"_bibtex":{"value":"@article{\npan2026spikingbrain,\ntitle={SpikingBrain: Spiking Brain-inspired Large Models},\nauthor={Yuqi Pan and Yupeng Feng and JingHao Zhuang and siyu ding and Han Xu and Zehao Liu and Bohan Sun and Yuhong Chou and Xuerui Qiu and Anlin Deng and Anjie Hu and Shurong Wang and Peng Zhou and Man Yao and Jibin Wu and jian yang and 孙国梁 and Bo XU and Guoqi Li},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2026},\nurl={https://openreview.net/forum?id=PNLShc0C6q},\nnote={}\n}"},"title":{"value":"SpikingBrain: Spiking Brain-inspired Large Models"},"pdf":{"value":"/pdf/5d8957628e08647be20fb705242fa90fd8e89a35.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"pan|spikingbrain_spiking_braininspired_large_models"},"authorids":{"value":["~Yuqi_Pan2","~Yupeng_Feng1","~JingHao_Zhuang1","~siyu_ding1","~Han_Xu12","~Zehao_Liu4","~Bohan_Sun2","~Yuhong_Chou1","~Xuerui_Qiu1","~Anlin_Deng1","~Anjie_Hu1","~Shurong_Wang1","~Peng_Zhou10","~Man_Yao1","~Jibin_Wu1","~jian_yang38","~孙国梁2","~Bo_XU10","~Guoqi_Li1"]},"assigned_action_editor":{"value":"~Binhang_Yuan1"},"authors":{"value":["Yuqi Pan","Yupeng Feng","JingHao Zhuang","siyu ding","Han Xu","Zehao Liu","Bohan Sun","Yuhong Chou","Xuerui Qiu","Anlin Deng","Anjie Hu","Shurong Wang","Peng Zhou","Man Yao","Jibin Wu","jian yang","孙国梁","Bo XU","Guoqi Li"]}},"version":2},{"content":{"comment":{"value":"- In addition, recent work has already scaled ANN-to-SNN distillation to a 76B-parameter foundation model (SpikingBrain, arXiv, 2025). At such scale, applying NAS becomes unrealistic: (1) there is no diversified population of SNN architectures with comparable parameter budgets to even form a meaningful search space; and (2) the computational cost is prohibitive. If a method cannot scale to the regime where distillation truly matters, then NAS-based ANN-to-SNN distillation loses much of its scientific and practical significance, since it will never be able to bridge the gap between SNNs and large-scale ANN models.\n\nWe thank you for this thoughtful comment. We also appreciate your pointing us to the impressive recent work (SpikingBrain, arXiv, 2025), which scales ANN-to-SNN distillation to a 76B-parameter foundation model and highlights the great potential of SNNs at foundation-model scale.  To better address your concerns, we summarize your concerns as two points:\n\n**(1) Applying NAS to LLM is unrealistic, especially for SNN** \n\nWe agree that scaling NAS to spiking LLM is indeed highly challenging. The two points you mentioned, i.e., the lack of a meaningful search space for SNN and the high computational cost, are indeed key challenges for large-scale SNN NAS. However, previous studies [4-7] have shown that using NAS for LLMs is not unrealistic and can be a good way to enhance the designs and performance of LLMs. In our future work, we will explore NAS for large-scale SNNs to overcome the challenges you mentioned.\n\n**(2) NAS-based ANN-to-SNN distillation lacks significance if it cannot scale to large models**\n\nWhile large-scale distillation is an important motivation, we would like to highlight that ANN-to-SNN distillation also has significant scientific and practical value for small- to medium-scale models. This is because one of the original goals of both NAS and distillation is to enable efficient model deployment in resource-constrained scenarios, where SNNs can offer advantages in this scenario. Therefore, even in smaller models, NAS-based ANN-to-SNN distillation remains scientifically and practically meaningful.\n\n**References:**\n\n[1] Relational knowledge distillation[C]. CVPR, 2019.\n\n[2] DisWOT: Student architecture search for distillation without training[C]. CVPR. 2023.\n\n[3] A new similarity-based relational knowledge distillation method[C]. ICASSP, 2024.\n\n[4] Neural architecture search for parameter-efficient fine-tuning of large pre-trained language models[C]. Findings of ACL, 2023.\n\n[5] LLaMA-NAS: Efficient neural architecture search for large language models[C]. ECCV, 2024.\n\n[6] Large language model compression with neural architecture search[C]. Workshop of NeurIPS, 2024.\n\n[7] ZeroLM: Data-Free Transformer Architecture Search for Language Models[J]. arXiv, 2025."},"title":{"value":"Response to Reviewer 8Fkb [Part 3/3]"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1764651426439,"tcdate":1764651426439,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18211/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission18211/Authors"],"forum":"vhfR0LTB7x","number":19,"license":"CC BY 4.0","cdate":1764651426439,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18211/-/Official_Comment"],"mdate":1764651426439,"domain":"ICLR.cc/2026/Conference","replyto":"ErMQ6rc3JS","id":"EkZamr2cSm","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Neural architecture search","spiking neural network","knowledge distillation"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"abstract":{"value":"Bridging the performance gap between Spiking Neural Networks (SNNs) and Artificial Neural Networks (ANNs) under low timesteps remains a critical challenge in the SNN community. Recent work uses either ANN-supervised training or automated architecture design to narrow the gap. However, the combination of ANN-supervised training and SNN architecture search remains unexplored, leaving room for further improvement of SNN performance. To address this, we propose Distilling SNN Students from ANN Teachers via Spiking Neural Architecture Search (DSAS) method. It is a training-free spiking neural architecture search method that leverages pre-trained ANN teachers to discover efficient and high-performance SNNs with few timesteps. Specifically, DSAS employs an evolutionary neural architecture search guided by two novel metrics, i.e., Multi-layer Activation Similarity (MAS) and Threshold-guided Gradient Similarity (TGS). MAS aligns ANN and SNN feature maps, yet TGS ensures gradient alignment while tuning spiking activation thresholds. Experiments demonstrate that DSAS achieves state-of-the-art accuracy with four timesteps on both convolution-based and transformer-based search space, effectively narrowing the performance gap of ANN and SNN. For example, DSAS discovers architectures that achieve 65.50% top-1 accuracy on Tiny-ImageNet and 81.97% on CIFAR-100. Available code: https://anonymous.4open.science/r/DSAS-5764"},"_bibtex":{"value":"@misc{\nsong2026distilling,\ntitle={Distilling {SNN} Students from {ANN} Teachers via Spiking Neural Architecture Search},\nauthor={Xiaotian Song and Yanan Sun},\nyear={2026},\nurl={https://openreview.net/forum?id=vhfR0LTB7x}\n}"},"title":{"value":"Distilling SNN Students from ANN Teachers via Spiking Neural Architecture Search"},"pdf":{"value":"/pdf/97efe771d3ee2ee5871d013137187b9c801b7b6b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"song|distilling_snn_students_from_ann_teachers_via_spiking_neural_architecture_search"},"authorids":{"value":["~Xiaotian_Song1","~Yanan_Sun4"]},"authors":{"value":["Xiaotian Song","Yanan Sun"]}},"version":2},{"content":{"comment":{"value":">  **Challenges 1 & Requested Changes 6: Data-matched evaluations.**\n\n### Clarification of Table 1\n\nThank you for raising this concern. We first clarify the intent of Table 1.  \nBased on Table 1, we draw the following conclusions (see Section 5.1, Pre-trained Models):  \n1. SpikingBrain-7B recovers nearly 90% of the base model’s performance.\n2. SpikingBrain-7B achieves performance comparable to advanced Transformer-based models.\n3. With appropriate training strategies, linear attention architectures can preserve strong general-purpose modeling capability.\n4. A performance gap remains compared with Qwen2.5-7B, due to substantial architectural modifications.\n\nWe believe that **Table 1 sufficiently supports these conclusions**, even though the included open-source models are not strictly data-matched. Importantly, we explicitly report the training data scale for all models in the “Tokens” column.\n\n**The inclusion of Falcon-Mamba results is intended to better support conclusion (3)**: starting from a strong open-source base model and applying a limited amount of data for conversion can outperform same-size models trained from scratch with substantially larger pre-training corpora.\n\n### Additional Data-Matched Experiments\n\nAlthough our work does not aim to introduce a fundamentally new architecture—rigorous comparisons among GLA, Transformer, and Mamba have already been established in prior literature—we fully agree with the reviewer that **data-matched evaluations are essential for drawing reliable takeaways**.\n\nFor more controlled and strict evidence, **we conducted strict data-matched conversion ablation experiments**, using CPT from Qwen3-1.7B with 5B tokens.\n\n**(1) Attention comparison under identical data budgets**  \nWe compare the inter-layer hybrid attention paradigm used in SpikingBrain-7B (LA:SWA = 1:1), pure linear attention (LA, similar to Falcon-Mamba), and pure sliding-window attention (SWA, similar to Mistral) under identical conversion settings.  \nFor SWA training, we use a 4k window size, and report the equivalent window size for all models. For fair comparison, we additionally evaluate a pure SWA variant with the window size fixed to 128.\n\n| Models   | equivalent window size |   HellaSwag ($\\uparrow$) |  MMLU ($\\uparrow$) | SNIAH1-4k ($\\uparrow$) |\n| -------- | ----: | ----: | -------: | --------: |\n| FA Baseline | ∞  | 60.21 | 61.78 |   100.0 | \n| Pure SWA    |4096| 60.12 | 61.74 |   100.0 |\n| Pure SWA    | 128| -     | -     |   1.87  |\n| Pure LA     | 64 | 54.72 | 39.87 |   36.16 |\n| Hybrid      |2080| 57.88 | 56.08 |   96.83 |\n\n**Key takeaways:**\n1.  **General modeling performance is primarily determined by the degree of architectural deviation between the target and base models.**   \nOn HellaSwag and MMLU, pure SWA outperforms Hybrid and pure LA, indicating that under data-limited settings, smaller architectural changes enable faster recovery of language modeling capability.\n2. **Recall performance is primarily governed by effective state size, and LA exhibits stronger recall than SWA with comparable state sizes.**  \nDue to faster capability recovery, SWA-4k achieves better recall within its attention window. Although LA requires more tokens to fully recover overall performance, it already demonstrates substantially stronger recall than SWA-128, despite having a similar effective state size.\n3. **Hybrid attention provides the best overall trade-off.**   \nBy combining SWA and LA, hybrid attention achieves a balanced combination of reasonable recovery speed, strong recall capability, and reduced memory cost, making it the most effective architecture choice under efficient conversion settings.\n\n**(2) Isolated Impact of H, M, and S Components**  \nUnder the same experimental setup, we further evaluate the isolated impact of hybrid attention (H), MoE (M), and spike coding (S) during conversion. Details are provided in our response to *Reviewer gn98, Requested Changes 1*. The conclusions are summarized below.\n\n**Key takeaways:**\n1.  **Hybrid efficient attention (H)** is the dominant contributor to long-context efficiency gains, albeit with some performance degradation.\n2.  **MoE modules (M)** improve model performance with minimal impact on efficiency.   \n3.  **Spike coding (S)** consistently incurs only limited performance degradation across configurations, while offering potential energy-efficiency benefits under appropriate hardware assumptions.\n\n### Clarification of \"2% of the Training Data\"\n\nThank you for pointing out this ambiguity. We will revise the wording to clearly state that **“2% of the data” refers to the data required for CPT-based conversion, in comparison to training from scratch**.\n\nWe have also carefully reviewed related statements in the Abstract, Introduction, Figure 1, Section 3, and other relevant parts of the manuscript, to ensure that readers clearly understand that our emphasis is on the efficiency advantage of the conversion pipeline over pre-training from scratch."},"title":{"value":"[2/6] Response to Reviewer WxpP (Claim Evidence Challenges 1)"}},"parentInvitations":"TMLR/-/Official_Comment","tmdate":1770986827757,"tcdate":1770809217941,"writers":["TMLR","TMLR/Paper6607/Authors"],"signatures":["TMLR/Paper6607/Authors"],"forum":"PNLShc0C6q","number":13,"license":"CC BY 4.0","cdate":1770809217941,"readers":["everyone"],"invitations":["TMLR/Paper6607/-/Official_Comment"],"mdate":1770986827757,"domain":"TMLR","replyto":"MwHdMjWBjW","id":"PeKbXmoohD","forumContent":{"submission_length":{"value":"Long submission (more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"code":{"value":"https://github.com/BICLab/SpikingBrain-7B"},"supplementary_material":{"value":"/attachment/67f070fcc827ac32ac4614eeaa968673de8b8856.zip"},"abstract":{"value":"Mainstream Transformer-based large language models (LLMs) face significant efficiency bottlenecks: training computation scales quadratically with sequence length, and inference memory grows linearly. These constraints limit their ability to process long sequences effectively. In addition, building large models on non-NVIDIA computing platforms poses major challenges in achieving stable and efficient training and deployment. To address these issues, we introduce SpikingBrain, a new family of brain-inspired models designed for efficient long-context training and inference. SpikingBrain leverages the MetaX GPU cluster and focuses on three core aspects: (1) Model Architecture: linear and hybrid-linear attention architectures with adaptive spiking neurons; (2) Algorithmic Optimizations: an efficient, conversion-based training pipeline compatible with existing LLMs, along with a dedicated spike coding framework; (3) System Engineering: customized training frameworks, operator libraries, and parallelism strategies tailored to the MetaX hardware. Using these techniques, we develop two models: SpikingBrain-7B, a linear LLM, and SpikingBrain-76B, a hybrid-linear MoE LLM. These models demonstrate the feasibility of large-scale LLM development on non-NVIDIA platforms, and our training framework supports weeks of stable training on hundreds of MetaX GPUs with Model FLOPs Utilization (MFU) at expected levels. SpikingBrain achieves performance comparable to open-source Transformer baselines while using exceptionally low data resources (continual pre-training of approximately 150B tokens). Our models also significantly improve long-context efficiency and deliver inference with (partially) constant memory and event-driven spiking behavior. For example, SpikingBrain-7B achieves more than 100× speedup in Time to First Token (TTFT) for 4M-token sequences. Furthermore, the proposed spiking scheme achieves 69.15% sparsity, enabling low-power operation. Overall, this work demonstrates the potential of brain-inspired mechanisms to drive the next generation of efficient and scalable large model design."},"_bibtex":{"value":"@article{\npan2026spikingbrain,\ntitle={SpikingBrain: Spiking Brain-inspired Large Models},\nauthor={Yuqi Pan and Yupeng Feng and JingHao Zhuang and siyu ding and Han Xu and Zehao Liu and Bohan Sun and Yuhong Chou and Xuerui Qiu and Anlin Deng and Anjie Hu and Shurong Wang and Peng Zhou and Man Yao and Jibin Wu and jian yang and 孙国梁 and Bo XU and Guoqi Li},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2026},\nurl={https://openreview.net/forum?id=PNLShc0C6q},\nnote={}\n}"},"title":{"value":"SpikingBrain: Spiking Brain-inspired Large Models"},"pdf":{"value":"/pdf/5d8957628e08647be20fb705242fa90fd8e89a35.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"pan|spikingbrain_spiking_braininspired_large_models"},"authorids":{"value":["~Yuqi_Pan2","~Yupeng_Feng1","~JingHao_Zhuang1","~siyu_ding1","~Han_Xu12","~Zehao_Liu4","~Bohan_Sun2","~Yuhong_Chou1","~Xuerui_Qiu1","~Anlin_Deng1","~Anjie_Hu1","~Shurong_Wang1","~Peng_Zhou10","~Man_Yao1","~Jibin_Wu1","~jian_yang38","~孙国梁2","~Bo_XU10","~Guoqi_Li1"]},"assigned_action_editor":{"value":"~Binhang_Yuan1"},"authors":{"value":["Yuqi Pan","Yupeng Feng","JingHao Zhuang","siyu ding","Han Xu","Zehao Liu","Bohan Sun","Yuhong Chou","Xuerui Qiu","Anlin Deng","Anjie Hu","Shurong Wang","Peng Zhou","Man Yao","Jibin Wu","jian yang","孙国梁","Bo XU","Guoqi Li"]}},"version":2},{"content":{"comment":{"value":">  **Requested Changes 5: Typos and presentation issues.**\n\n1. Typos: Thank you for pointing this out. Some typographical errors were introduced during figure refinement, and we will correct them accordingly.\n2. Nit: We followed abbreviations used in some prior works (“HS” for HellaSwag) [1–2]. Based on the reviewer’s suggestion, we will use the full name HellaSwag consistently in the revised manuscript.\n\n---\n\n[1] Hierarchically Gated Recurrent Neural Network for Sequence Modeling\n, 2023.  \n[2] MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map, 2024. \n\n>  **Requested Changes 6: Data-matched evaluations.**\n\nplease see our response to the *Claim Evidence Challenge 1*.\n\n>  **Requested Changes 7: Implementation efforts of adaptive spiking mechanism.**\n\nThank you for your interest in our proposed adaptive spiking mechanism. We first briefly recap the mechanism and then clarify the corresponding implementation efforts under synchronous GPU-based computation and asynchronous neuromorphic execution paradigms. **We will incorporate this discussion into the revised manuscript**.\n\nAs described in Section 3.3, we propose a spiking scheme that converts continuous-valued activations into two equivalent representations: **integer spike counts** and **expanded sparse spike trains**. The integer formulation enables efficient execution on GPUs, while the spiking formulation naturally supports event-driven execution.\n\nIf the mechanism is used purely for numerical simulation, no special implementation effort is required; however, such usage only serves algorithmic validation and does not yield practical acceleration. To achieve actual performance and energy benefits across different platforms, the following implementation considerations are necessary.\n\n**(1) GPU-based implementation**\n\nOn GPUs, acceleration can be achieved by leveraging **INT-type quantized kernels**. In SpikingBrain, this is enabled by `INT8` quantization of weights (Section 5.5), allowing spike-count-based integer operations to be executed efficiently using existing GPU kernel optimizations.\n\n**(2) Asynchronous neuromorphic hardware implementation**\n\nAchieving memory-access and computation energy efficiency on asynchronous hardware requires more specialized, platform-aware design:\n\n- **Algorithm–hardware co-design for memory access**: Conduct offline profiling to identify sparsity pattern of SpikingBrain. On the software side, algorithm-level weight rearrangement and lightweight compression (e.g., run-length encoding) could be performed; on the hardware side, a channel-aware memory controller could be designed and channel-level memory partitioning could be implemented to isolate weights of active and inactive channels into separate memory blocks, enabling effective skipping of inactive channels.\n\n- **Low-overhead asynchronous control**: Adopt lightweight asynchronous protocols to reduce control overhead. For instance, click controllers and two-phase handshake protocols can minimize unnecessary signal switching. Additionally, an adaptive scheduler could be implemented to monitor spike sparsity in real time. The scheduler dynamically decides whether to merge low-activity channels or expand computing resources, further optimizing energy consumption.\n\n- **More comprehensive comparison and end-to-end modeling**: In hardware development, the theoretical estimation model is gradually revised by incorporating previously omitted overheads derived from hardware simulation into the model. The revised model can be used to guide hardware design and algorithm optimization, ensuring that the balance of power, performance and area is achieved, energy efficiency claims are consistent with practical implementation results, and a complete estimation of end-to-end energy consumption is realized."},"title":{"value":"[5/6] Response to Reviewer WxpP (Requested Changes 5~7)"}},"parentInvitations":"TMLR/-/Official_Comment","tmdate":1770809360721,"tcdate":1770809360721,"writers":["TMLR","TMLR/Paper6607/Authors"],"signatures":["TMLR/Paper6607/Authors"],"forum":"PNLShc0C6q","number":16,"license":"CC BY 4.0","cdate":1770809360721,"readers":["everyone"],"invitations":["TMLR/Paper6607/-/Official_Comment"],"mdate":1770809360721,"domain":"TMLR","replyto":"MwHdMjWBjW","id":"8YEaDqegCU","forumContent":{"submission_length":{"value":"Long submission (more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"code":{"value":"https://github.com/BICLab/SpikingBrain-7B"},"supplementary_material":{"value":"/attachment/67f070fcc827ac32ac4614eeaa968673de8b8856.zip"},"abstract":{"value":"Mainstream Transformer-based large language models (LLMs) face significant efficiency bottlenecks: training computation scales quadratically with sequence length, and inference memory grows linearly. These constraints limit their ability to process long sequences effectively. In addition, building large models on non-NVIDIA computing platforms poses major challenges in achieving stable and efficient training and deployment. To address these issues, we introduce SpikingBrain, a new family of brain-inspired models designed for efficient long-context training and inference. SpikingBrain leverages the MetaX GPU cluster and focuses on three core aspects: (1) Model Architecture: linear and hybrid-linear attention architectures with adaptive spiking neurons; (2) Algorithmic Optimizations: an efficient, conversion-based training pipeline compatible with existing LLMs, along with a dedicated spike coding framework; (3) System Engineering: customized training frameworks, operator libraries, and parallelism strategies tailored to the MetaX hardware. Using these techniques, we develop two models: SpikingBrain-7B, a linear LLM, and SpikingBrain-76B, a hybrid-linear MoE LLM. These models demonstrate the feasibility of large-scale LLM development on non-NVIDIA platforms, and our training framework supports weeks of stable training on hundreds of MetaX GPUs with Model FLOPs Utilization (MFU) at expected levels. SpikingBrain achieves performance comparable to open-source Transformer baselines while using exceptionally low data resources (continual pre-training of approximately 150B tokens). Our models also significantly improve long-context efficiency and deliver inference with (partially) constant memory and event-driven spiking behavior. For example, SpikingBrain-7B achieves more than 100× speedup in Time to First Token (TTFT) for 4M-token sequences. Furthermore, the proposed spiking scheme achieves 69.15% sparsity, enabling low-power operation. Overall, this work demonstrates the potential of brain-inspired mechanisms to drive the next generation of efficient and scalable large model design."},"_bibtex":{"value":"@article{\npan2026spikingbrain,\ntitle={SpikingBrain: Spiking Brain-inspired Large Models},\nauthor={Yuqi Pan and Yupeng Feng and JingHao Zhuang and siyu ding and Han Xu and Zehao Liu and Bohan Sun and Yuhong Chou and Xuerui Qiu and Anlin Deng and Anjie Hu and Shurong Wang and Peng Zhou and Man Yao and Jibin Wu and jian yang and 孙国梁 and Bo XU and Guoqi Li},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2026},\nurl={https://openreview.net/forum?id=PNLShc0C6q},\nnote={}\n}"},"title":{"value":"SpikingBrain: Spiking Brain-inspired Large Models"},"pdf":{"value":"/pdf/5d8957628e08647be20fb705242fa90fd8e89a35.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"pan|spikingbrain_spiking_braininspired_large_models"},"authorids":{"value":["~Yuqi_Pan2","~Yupeng_Feng1","~JingHao_Zhuang1","~siyu_ding1","~Han_Xu12","~Zehao_Liu4","~Bohan_Sun2","~Yuhong_Chou1","~Xuerui_Qiu1","~Anlin_Deng1","~Anjie_Hu1","~Shurong_Wang1","~Peng_Zhou10","~Man_Yao1","~Jibin_Wu1","~jian_yang38","~孙国梁2","~Bo_XU10","~Guoqi_Li1"]},"assigned_action_editor":{"value":"~Binhang_Yuan1"},"authors":{"value":["Yuqi Pan","Yupeng Feng","JingHao Zhuang","siyu ding","Han Xu","Zehao Liu","Bohan Sun","Yuhong Chou","Xuerui Qiu","Anlin Deng","Anjie Hu","Shurong Wang","Peng Zhou","Man Yao","Jibin Wu","jian yang","孙国梁","Bo XU","Guoqi Li"]}},"version":2},{"content":{"comment":{"value":"> One shortcoming is that the ablation study demonstrates that increasing the range of integer values for spikes improves performance; however, only two ranges are examined, so it would be good to explore this trend a little more thoroughly. It would also be useful to establish the energy/accuracy tradeoff a little more rigorously, perhaps as a scatterplot with energy and accuracy as axes.\n\n**Reply 1:** Thank you for the constructive suggestion. We have expanded the ablation study by including additional integer ranges and added a scatter plot illustrating the energy/accuracy tradeoff. We extended the range to $\\pm8$ and observed a performance gain. However, this also reduced energy efficiency and doubled the inference latency. \n\nThe new figure and analysis have been incorporated into the ablation study section and highlighted in blue in the revised manuscript on Page 10.\n\n---\n\n> Please keep the usage of the notation $T$ and $D$ consistent and define them both before they are used in Table 1.\n\n\n**Reply 2:** Thank you for the helpful suggestion. We have ensured consistent usage of the notations $T$ and $D$ throughout the paper. \n\nThe definitions have been added to the caption of Table 1, and the corresponding revisions are highlighted in blue in the updated manuscript.\n\n---\n\n> The evaluation protocol should be defined more clearly in the main text. What exactly does zero-shot accuracy mean here? What data are used to estimate it?\n\n> It would be good to provide a comprehensive summary of resources used in generating the experimental data.\n\n**Reply 3:** We sincerely thank you for this constructive comment. We have clarified the evaluation protocol in the main text. The relevant descriptions and dataset details have been moved from the Appendix to Section 5.1 and are highlighted in blue in the revised manuscripts.\n\nOur evaluation strictly follows the standard LM-Eval-Harness protocol, which has been widely adopted in prior work on large language models (e.g., GPT, LLaMA, Mamba) for reproducible zero-shot assessment across benchmarks such as LAMBADA, PIQA, ARC and so on.\n\nSpecifically, zero-shot accuracy measures a model’s performance on downstream tasks without any task-specific fine-tuning or gradient updates, i.e., the model is directly evaluated in an inference-only manner using task instructions and prompts.\n\n---\n\n> It is stated that SGC is only used for three layers: the \"first\", \"last\", and \"middle\". I'm not sure which these are.\n\n**Reply 4:** Thank you for this valuable comment. For the Mamba2-130M model, which contains 24 layers, the “first,” “middle,” and “last” layers correspond to the 1st, 12th, and 24th layers, respectively. For the 1.3B model, which contains 48 layers, they correspond to the 1st, 24th, and 48th layers.\n\nWe have clarified this in the revised manuscript on Page 7, and the corresponding update is highlighted in blue.\n\n---\n\n> The spiking neural networks community has surely explored neurons similar to the TI-LIF neurons examined here, right? It would be useful to mention this literature or else comment upon the gap.\n\n**Reply 5:** Thank you for the insightful comment. Indeed, neurons similar to the TI-LIF/I-LIF family have been explored in the SNN community. Expect discussed in section 3.1, concurrent work SpikingBrain [1] mainly focuses on integer-based quantized inference and does not verify end-to-end gradient-based training for the spiking model, as their quantization mechanisms are designed for forward pass efficiency rather than backward optimization.\n\nTo clarify this connection and distinguish our SI-LIF neuron from prior quantized LIF variants, we have added a brief discussion after Section 3.1, with the new text highlighted in blue in the revised manuscript on Page 4.\n\n---\n**Reference:**\n\n[1] Pan Y, Feng Y, Zhuang J, et al. SpikingBrain: Spiking Brain-inspired Large Models[J]. arXiv preprint arXiv:2509.05276, 2025.\n\n---\nFinally, we sincerely thank you again for your meticulous review. The constructive comments you provided have been highly valuable in strengthening the manuscript. Should there be any further questions, we would be more than willing to respond and make additional improvements."}},"parentInvitations":"TMLR/-/Official_Comment","tmdate":1763574339161,"tcdate":1763574339161,"writers":["TMLR","TMLR/Paper6077/Authors"],"signatures":["TMLR/Paper6077/Authors"],"forum":"uxb2jcCLxt","number":7,"license":"CC BY 4.0","cdate":1763574339161,"readers":["everyone"],"invitations":["TMLR/Paper6077/-/Official_Comment"],"mdate":1763574339161,"domain":"TMLR","replyto":"f6nhNICY5t","id":"syLBOA7p4d","forumContent":{"submission_length":{"value":"Regular submission (no more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"abstract":{"value":"Large Language Models (LLMs) have achieved remarkable performance across tasks but remain energy-intensive due to dense matrix operations. Spiking neural networks (SNNs) improve energy efficiency by replacing dense matrix multiplications with sparse accumulations. Their sparse spike activity enables efficient LLMs deployment on edge devices. However, prior SNN-based LLMs often sacrifice performance for efficiency, and recovering accuracy typically requires full pretraining, which is costly and impractical. To address this, we propose SpikingMamba, an energy-efficient SNN-based LLMs distilled from Mamba that improves energy efficiency with minimal accuracy sacrifice. SpikingMamba integrates two key components: (a) SI-LIF, a signed-integer spiking neuron that preserves semantic polarity through signed multi-level spike representations. (b) A training-exclusive Smoothed Gradient Compensation (SGC) path mitigating quantization loss while preserving spike-driven efficiency. We employ a single-stage distillation strategy to transfer the zero-shot ability of pretrained Mamba and further enhance it via reinforcement learning (RL). Experiments show that SpikingMamba-1.3B achieves a 4.76$\\times$ energy benefit, with only a 4.78\\% zero-shot accuracy gap compared to the original Mamba. The model achieves a further 2.55\\% accuracy improvement after RL, narrowing the performance gap from 4.78\\% to 2.23\\%."},"_bibtex":{"value":"@article{\nhuang2026spikingmamba,\ntitle={SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba},\nauthor={Yulong Huang and Jianxiong Tang and Chao Wang and Ziyi Wang and Jianguo Zhang and Zhichao Lu and Bojun Cheng and Luziwei Leng},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2026},\nurl={https://openreview.net/forum?id=uxb2jcCLxt},\nnote={}\n}"},"title":{"value":"SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba"},"pdf":{"value":"/pdf/6ce7e89a04f5887bd950a944ed8078bbb994d1b7.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"huang|spikingmamba_towards_energyefficient_large_language_models_via_knowledge_distillation_from_mamba"},"authorids":{"value":["~Yulong_Huang2","~Jianxiong_Tang1","~Chao_Wang50","~Ziyi_Wang19","~Jianguo_Zhang2","~Zhichao_Lu1","~Bojun_Cheng1","~Luziwei_Leng1"]},"assigned_action_editor":{"value":"~John_Timothy_Halloran1"},"authors":{"value":["Yulong Huang","Jianxiong Tang","Chao Wang","Ziyi Wang","Jianguo Zhang","Zhichao Lu","Bojun Cheng","Luziwei Leng"]}},"version":2},{"content":{"comment":{"value":">  **Challenges 2 & Requested Changes 2: Evaluation of recall capability.**\n\nThank you for raising this concern. In the original manuscript, we evaluated and discussed **long-context modeling and understanding capabilities** (as reported in Section 5.1), which are not designed for recall-intensive scenarios. \n\nWe agree with the reviewer that recall is a critical issue for linear and subquadratic models. To answer this, we **conducted additional data-matched conversion experiments** and will **add a dedicated paragraph with relevant citations** in the revised manuscript to explicitly discuss recall behavior.\n\n**(1) Experiments:**\n\nFollowing the setup described in our response to Challenges 1, we compare **Hybrid, pure SWA, and pure LA** architectures under identical CPT data budgets, and evaluate their **Single-NIAH** performance using the OpenCompass framework on RULER tasks. We additionally include:\n(i) CPT results of a full-attention (FA) Transformer, and\n(ii) an inference-only variant where SWA layers in the Hybrid model are replaced with FA.\n\n| Models   | Equivalent Window Size | SNIAH1-4k / 8k ($\\uparrow$) | SNIAH2-4k / 8k ($\\uparrow$) | SNIAH3-4k / 8k ($\\uparrow$) |\n| -------- | ----: | ----: |  ----: |  ----: |\n| FA Baseline       | ∞  |   100.0 / 98.41 |   98.41 / 96.83 |   96.83 / 95.24 |\n| Pure SWA          |4096|   100.0 / 47.60 |   97.80 / 47.00 |   98.60 / 50.20  |\n| Pure SWA          | 128|   1.87 / 0.00   |   0.93 / 0.93   | 1.87 / 1.87  |\n| Pure LA           | 64 |   36.16 / 39.40 |   11.43 / 12.40 |   0.7 / 0.6 |\n| Hybrid (SWA + LA) |2080|   96.83 / 44.44 |   93.65 / 39.68 |  79.37 / 26.98 |\n| Hybrid (FA + LA)  | ∞  |   74.61 / 71.00 |   92.86 / 55.00 |   97.62 / 31.40 |\n\n**Table Note**: For `96.83 / 44.44`, the first number `96.83` denotes model's accuracy on the SNIAH task with 4k context length, and the second number `44.44` denotes that of 8k context length\n\n**Our observations are as follows:**\n1. All converted models (SWA, LA, and Hybrid) exhibit degraded recall performance compared to the original FA model.\n2. A clear trade-off exists between effective state size and recall performance, with linear attention exhibiting stronger recall than SWA at comparable state sizes.\n3. Hybrid attention provides a more balanced combination of recall capability and state size compared with pure SWA and pure LA.\n4. Even when trained with SWA and LA, switching to a Hybrid of FA and LA at inference time (similar to SpikingBrain-76B) significantly reduces the long-context recall gap relative to FA.\n\n**(2) More Discussions:**\n\nWe further emphasize that **this gap can be further reduced by improving data quality**, for example by incorporating real long-context or recall-intensive training data. Downstream recall performance depends critically on both the quantity and quality of training data. The base Qwen3 model is trained with a 32k context length and includes rich retrieval-style data, whereas our current conversion experiments lack targeted data designed to explicitly elicit recall capabilities in subquadratic models.\n\nNevertheless, prior industrial studies have shown that **hybrid attention architectures do not significantly impair retrieval performance** relative to Transformers when trained with sufficient private data [1–2].\n\n---\n\n[1] MiniMax-01: Scaling Foundation Models with Lightning Attention, 2025.  \n[2] Qwen3-Next: Towards Ultimate Training & Inference Efficiency, 2025. \n\n>  **Challenges 3: Other hybrid models.**\n\nThank you for the suggestion. We will **add appropriate references** to acknowledge the strongest hybrid models currently developed in industry.\n\nWe would also like to respectfully clarify that, in the academic setting, implementing models of this scale and sequence length on open-source frameworks, particularly adapting distributed training to MetaX clusters, requires **a substantial amount of time and effort**. In fact, this project was initiated in **late 2024**, shortly after the release of Qwen2.5. As such, the central claim of SpikingBrain has consistently been that it nearly closes the performance gap with the base model Qwen2.5-7B, while simultaneously delivering long-context efficiency and potential event-driven energy efficiency.\n\n**We fully acknowledge that many strong open-source hybrid models have emerged over the past year, and we are encouraged to see growing industrial interest in linear attention architectures**. However, it is also important to note that, over the past years, the iteration speed of industrial large-model development especially at leading companies, has significantly outpaced that of academic research. As a result, it is difficult for us to conduct direct quantitative comparisons with very recent models such as Qwen3-Next or Kimi-Linear. Nevertheless, **we believe that the key findings validated in this work remain valuable to the research community, independent of the rapid pace of industrial model releases**."},"title":{"value":"[3/6] Response to Reviewer WxpP (Claim Evidence Challenges 2~3)"}},"parentInvitations":"TMLR/-/Official_Comment","tmdate":1770809264734,"tcdate":1770809264734,"writers":["TMLR","TMLR/Paper6607/Authors"],"signatures":["TMLR/Paper6607/Authors"],"forum":"PNLShc0C6q","number":14,"license":"CC BY 4.0","cdate":1770809264734,"readers":["everyone"],"invitations":["TMLR/Paper6607/-/Official_Comment"],"mdate":1770809264734,"domain":"TMLR","replyto":"MwHdMjWBjW","id":"Ai8YxWUcw5","forumContent":{"submission_length":{"value":"Long submission (more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"code":{"value":"https://github.com/BICLab/SpikingBrain-7B"},"supplementary_material":{"value":"/attachment/67f070fcc827ac32ac4614eeaa968673de8b8856.zip"},"abstract":{"value":"Mainstream Transformer-based large language models (LLMs) face significant efficiency bottlenecks: training computation scales quadratically with sequence length, and inference memory grows linearly. These constraints limit their ability to process long sequences effectively. In addition, building large models on non-NVIDIA computing platforms poses major challenges in achieving stable and efficient training and deployment. To address these issues, we introduce SpikingBrain, a new family of brain-inspired models designed for efficient long-context training and inference. SpikingBrain leverages the MetaX GPU cluster and focuses on three core aspects: (1) Model Architecture: linear and hybrid-linear attention architectures with adaptive spiking neurons; (2) Algorithmic Optimizations: an efficient, conversion-based training pipeline compatible with existing LLMs, along with a dedicated spike coding framework; (3) System Engineering: customized training frameworks, operator libraries, and parallelism strategies tailored to the MetaX hardware. Using these techniques, we develop two models: SpikingBrain-7B, a linear LLM, and SpikingBrain-76B, a hybrid-linear MoE LLM. These models demonstrate the feasibility of large-scale LLM development on non-NVIDIA platforms, and our training framework supports weeks of stable training on hundreds of MetaX GPUs with Model FLOPs Utilization (MFU) at expected levels. SpikingBrain achieves performance comparable to open-source Transformer baselines while using exceptionally low data resources (continual pre-training of approximately 150B tokens). Our models also significantly improve long-context efficiency and deliver inference with (partially) constant memory and event-driven spiking behavior. For example, SpikingBrain-7B achieves more than 100× speedup in Time to First Token (TTFT) for 4M-token sequences. Furthermore, the proposed spiking scheme achieves 69.15% sparsity, enabling low-power operation. Overall, this work demonstrates the potential of brain-inspired mechanisms to drive the next generation of efficient and scalable large model design."},"_bibtex":{"value":"@article{\npan2026spikingbrain,\ntitle={SpikingBrain: Spiking Brain-inspired Large Models},\nauthor={Yuqi Pan and Yupeng Feng and JingHao Zhuang and siyu ding and Han Xu and Zehao Liu and Bohan Sun and Yuhong Chou and Xuerui Qiu and Anlin Deng and Anjie Hu and Shurong Wang and Peng Zhou and Man Yao and Jibin Wu and jian yang and 孙国梁 and Bo XU and Guoqi Li},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2026},\nurl={https://openreview.net/forum?id=PNLShc0C6q},\nnote={}\n}"},"title":{"value":"SpikingBrain: Spiking Brain-inspired Large Models"},"pdf":{"value":"/pdf/5d8957628e08647be20fb705242fa90fd8e89a35.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"pan|spikingbrain_spiking_braininspired_large_models"},"authorids":{"value":["~Yuqi_Pan2","~Yupeng_Feng1","~JingHao_Zhuang1","~siyu_ding1","~Han_Xu12","~Zehao_Liu4","~Bohan_Sun2","~Yuhong_Chou1","~Xuerui_Qiu1","~Anlin_Deng1","~Anjie_Hu1","~Shurong_Wang1","~Peng_Zhou10","~Man_Yao1","~Jibin_Wu1","~jian_yang38","~孙国梁2","~Bo_XU10","~Guoqi_Li1"]},"assigned_action_editor":{"value":"~Binhang_Yuan1"},"authors":{"value":["Yuqi Pan","Yupeng Feng","JingHao Zhuang","siyu ding","Han Xu","Zehao Liu","Bohan Sun","Yuhong Chou","Xuerui Qiu","Anlin Deng","Anjie Hu","Shurong Wang","Peng Zhou","Man Yao","Jibin Wu","jian yang","孙国梁","Bo XU","Guoqi Li"]}},"version":2},{"content":{"comment":{"value":">  **Additional Comments 1: Performance gap.**\n\nThank you for raising this important point. We would like to respectfully clarify that training from scratch and conversion training represent **two fundamentally different comparison paradigms** when evaluating performance against Transformer baselines.\n\nIn the referenced studies, Transformers and hybrid models are typically **trained from scratch under identical data budgets**, often using trillions of tokens for PT. In this setting, all models benefit from extensive training to fully optimize their modeling capacity. As a result, when a sufficient number of full-attention layers are retained, hybrid models generally achieve language modeling performance (e.g., MMLU) **comparable to, or even exceeding**, that of Transformers.\n\nIn contrast, when **converting from an already trained and annealed Transformer checkpoint**, a natural performance gap is expected. The model must adapt to a modified architecture within a **limited additional training window**, which inevitably leads to relative under-training compared to the base Transformer. Consequently, a substantial body of prior work has focused on reducing this gap. For example, when converting from a Transformer to a pure linear-attention model, recovery on MMLU is often limited to around **60%** [1-2]. Other studies show that combining careful distillation with partial full-attention retention can improve recovery to approximately **80%** [3-4].\n\nNotably, **SpikingBrain does not employ any distillation techniques, yet it recovers approximately 89% of the MMLU performance of the base Transformer**. We therefore consider the remaining performance gap to be **acceptable** under the conversion-training paradigm.\n\n---\n[1] Finetuning Pretrained Transformers into RNNs, 2021.  \n[2] Gated Slot Attention for Efficient Linear-Time Sequence Modeling, 2024.  \n[3] The Mamba in the Llama: Distilling and Accelerating Hybrid Models, 2024.   \n[4] LoLCATs: On Low-Rank Linearizing of Large Language Models, 2024.\n\n>  **Additional Comments 2: What didn't work.**\n\nThank you for pointing this out. We will add a dedicated section in the revised manuscript to share practical insights about what did not work during the adaptation to MetaX hardware, including:\n- Low GPU utilization, and occasional out-of-memory (OOM) errors and deadlocks during multi-node sequence-parallel (SP) training.\n- The need to adapt Megatron framework interfaces to older versions of PyTorch and Triton, which limited access to newer optimization features.\n- Simplifications in Triton kernel design, such as reduced block sizes and constrained autotuning configurations, which were necessary for stability but resulted in suboptimal performance.\n\nWe believe that transparently documenting these limitations and failed attempts will provide useful guidance for future work on deploying large-scale hybrid and spiking models (especially on non-NVIDIA hardware platforms)."},"title":{"value":"[6/6] Response to Reviewer WxpP (Additional Comments)"}},"parentInvitations":"TMLR/-/Official_Comment","tmdate":1770809397421,"tcdate":1770809397421,"writers":["TMLR","TMLR/Paper6607/Authors"],"signatures":["TMLR/Paper6607/Authors"],"forum":"PNLShc0C6q","number":17,"license":"CC BY 4.0","cdate":1770809397421,"readers":["everyone"],"invitations":["TMLR/Paper6607/-/Official_Comment"],"mdate":1770809397421,"domain":"TMLR","replyto":"MwHdMjWBjW","id":"t5h99pqkG6","forumContent":{"submission_length":{"value":"Long submission (more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"code":{"value":"https://github.com/BICLab/SpikingBrain-7B"},"supplementary_material":{"value":"/attachment/67f070fcc827ac32ac4614eeaa968673de8b8856.zip"},"abstract":{"value":"Mainstream Transformer-based large language models (LLMs) face significant efficiency bottlenecks: training computation scales quadratically with sequence length, and inference memory grows linearly. These constraints limit their ability to process long sequences effectively. In addition, building large models on non-NVIDIA computing platforms poses major challenges in achieving stable and efficient training and deployment. To address these issues, we introduce SpikingBrain, a new family of brain-inspired models designed for efficient long-context training and inference. SpikingBrain leverages the MetaX GPU cluster and focuses on three core aspects: (1) Model Architecture: linear and hybrid-linear attention architectures with adaptive spiking neurons; (2) Algorithmic Optimizations: an efficient, conversion-based training pipeline compatible with existing LLMs, along with a dedicated spike coding framework; (3) System Engineering: customized training frameworks, operator libraries, and parallelism strategies tailored to the MetaX hardware. Using these techniques, we develop two models: SpikingBrain-7B, a linear LLM, and SpikingBrain-76B, a hybrid-linear MoE LLM. These models demonstrate the feasibility of large-scale LLM development on non-NVIDIA platforms, and our training framework supports weeks of stable training on hundreds of MetaX GPUs with Model FLOPs Utilization (MFU) at expected levels. SpikingBrain achieves performance comparable to open-source Transformer baselines while using exceptionally low data resources (continual pre-training of approximately 150B tokens). Our models also significantly improve long-context efficiency and deliver inference with (partially) constant memory and event-driven spiking behavior. For example, SpikingBrain-7B achieves more than 100× speedup in Time to First Token (TTFT) for 4M-token sequences. Furthermore, the proposed spiking scheme achieves 69.15% sparsity, enabling low-power operation. Overall, this work demonstrates the potential of brain-inspired mechanisms to drive the next generation of efficient and scalable large model design."},"_bibtex":{"value":"@article{\npan2026spikingbrain,\ntitle={SpikingBrain: Spiking Brain-inspired Large Models},\nauthor={Yuqi Pan and Yupeng Feng and JingHao Zhuang and siyu ding and Han Xu and Zehao Liu and Bohan Sun and Yuhong Chou and Xuerui Qiu and Anlin Deng and Anjie Hu and Shurong Wang and Peng Zhou and Man Yao and Jibin Wu and jian yang and 孙国梁 and Bo XU and Guoqi Li},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2026},\nurl={https://openreview.net/forum?id=PNLShc0C6q},\nnote={}\n}"},"title":{"value":"SpikingBrain: Spiking Brain-inspired Large Models"},"pdf":{"value":"/pdf/5d8957628e08647be20fb705242fa90fd8e89a35.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"pan|spikingbrain_spiking_braininspired_large_models"},"authorids":{"value":["~Yuqi_Pan2","~Yupeng_Feng1","~JingHao_Zhuang1","~siyu_ding1","~Han_Xu12","~Zehao_Liu4","~Bohan_Sun2","~Yuhong_Chou1","~Xuerui_Qiu1","~Anlin_Deng1","~Anjie_Hu1","~Shurong_Wang1","~Peng_Zhou10","~Man_Yao1","~Jibin_Wu1","~jian_yang38","~孙国梁2","~Bo_XU10","~Guoqi_Li1"]},"assigned_action_editor":{"value":"~Binhang_Yuan1"},"authors":{"value":["Yuqi Pan","Yupeng Feng","JingHao Zhuang","siyu ding","Han Xu","Zehao Liu","Bohan Sun","Yuhong Chou","Xuerui Qiu","Anlin Deng","Anjie Hu","Shurong Wang","Peng Zhou","Man Yao","Jibin Wu","jian yang","孙国梁","Bo XU","Guoqi Li"]}},"version":2},{"content":{"comment":{"value":"I appreciate the authors’ efforts in extending the experiments to Transformer architectures and in exploring additional spiking neuron types. However, I still have several questions and concerns:\n\n1. I requested deeper theoretical insights into the proposed metrics. Unfortunately, the current explanations are still insufficient and do not meaningfully clarify the underlying rationale or theoretical contribution.\n\n* From the definition in Paper Eq. (4)–(5), MAS normalizes the similarity matrices rather than the activations themselves. As a result, the metric is invariant to rescaling of activations and therefore does not enforce “similar activation norms,” contrary to the authors’ claim of \"(b) Similar activation norms through normalization and sparsity matching...\".\n* More broadly, the structure of $\\mathcal{M}_T, \\mathcal{M}_S \\in \\mathbb{R}^{B \\times B}$ essentially mirrors the representational similarity matrices used in Representational Similarity Analysis (RSA) [1]. Given this close conceptual alignment, the purported theoretical novelty of MAS appears limited, and the primary contribution seems to lie instead in the sparse alignment between SNN and ANN rather than in the metric itself.\n* Regarding TGS, the response does not address a fundamental issue: surrogate gradients for SNNs are known to introduce systematic approximation errors, which can degrade gradient direction alignment. Without an analysis of how TGS mitigates this effect, or whether it is robust to such errors, the claimed benefit of gradient alignment remains speculative.\n\n2. It's good to see that selecting neuron types offers additional accuracy improvements. Could the authors provide the actual search results, particularly which specific spiking neuron models were selected under different conditions?\n\n3. Before closing, I would again like to acknowledge the authors’ considerable effort. Overall, the paper is a relatively complete piece of work. However, I have a central concern regarding the necessity and value of applying NAS to the ANN-to-SNN distillation pipeline. \n* The purpose of distillation is fundamentally to transfer the teacher model’s representational structure to the student model. Yet the bottlenecks in SNNs, which stem from spike-based activations and errors induced by surrogate gradients, are structural and algorithmic, not architectural. I do not see how choosing a particular CNN/Transformer variant can meaningfully resolve these issues.\n* In addition, recent work has already scaled ANN-to-SNN distillation to a 76B-parameter foundation model [2]. At such scale, applying NAS becomes unrealistic: (1) there is no diversified population of SNN architectures with comparable parameter budgets to even form a meaningful search space; and (2) the computational cost is prohibitive. If a method cannot scale to the regime where distillation truly matters, then NAS-based ANN-to-SNN distillation loses much of its scientific and practical significance, since it will never be able to bridge the gap between SNNs and large-scale ANN models.\n\n[1] Kriegeskorte N, Mur M, Bandettini P A. Representational similarity analysis-connecting the branches of systems neuroscience[J]. Frontiers in systems neuroscience, 2008, 2: 249.\n\n[2] Pan Y, Feng Y, Zhuang J, et al. SpikingBrain: Spiking Brain-inspired Large Models[J]. arXiv preprint arXiv:2509.05276, 2025."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1764312351630,"tcdate":1764312351630,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18211/Reviewer_8Fkb"],"signatures":["ICLR.cc/2026/Conference/Submission18211/Reviewer_8Fkb"],"forum":"vhfR0LTB7x","number":16,"license":"CC BY 4.0","cdate":1764312351630,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18211/-/Official_Comment"],"mdate":1764312351630,"domain":"ICLR.cc/2026/Conference","replyto":"JKRq7FabzG","id":"ErMQ6rc3JS","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Neural architecture search","spiking neural network","knowledge distillation"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"abstract":{"value":"Bridging the performance gap between Spiking Neural Networks (SNNs) and Artificial Neural Networks (ANNs) under low timesteps remains a critical challenge in the SNN community. Recent work uses either ANN-supervised training or automated architecture design to narrow the gap. However, the combination of ANN-supervised training and SNN architecture search remains unexplored, leaving room for further improvement of SNN performance. To address this, we propose Distilling SNN Students from ANN Teachers via Spiking Neural Architecture Search (DSAS) method. It is a training-free spiking neural architecture search method that leverages pre-trained ANN teachers to discover efficient and high-performance SNNs with few timesteps. Specifically, DSAS employs an evolutionary neural architecture search guided by two novel metrics, i.e., Multi-layer Activation Similarity (MAS) and Threshold-guided Gradient Similarity (TGS). MAS aligns ANN and SNN feature maps, yet TGS ensures gradient alignment while tuning spiking activation thresholds. Experiments demonstrate that DSAS achieves state-of-the-art accuracy with four timesteps on both convolution-based and transformer-based search space, effectively narrowing the performance gap of ANN and SNN. For example, DSAS discovers architectures that achieve 65.50% top-1 accuracy on Tiny-ImageNet and 81.97% on CIFAR-100. Available code: https://anonymous.4open.science/r/DSAS-5764"},"_bibtex":{"value":"@misc{\nsong2026distilling,\ntitle={Distilling {SNN} Students from {ANN} Teachers via Spiking Neural Architecture Search},\nauthor={Xiaotian Song and Yanan Sun},\nyear={2026},\nurl={https://openreview.net/forum?id=vhfR0LTB7x}\n}"},"title":{"value":"Distilling SNN Students from ANN Teachers via Spiking Neural Architecture Search"},"pdf":{"value":"/pdf/97efe771d3ee2ee5871d013137187b9c801b7b6b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"song|distilling_snn_students_from_ann_teachers_via_spiking_neural_architecture_search"},"authorids":{"value":["~Xiaotian_Song1","~Yanan_Sun4"]},"authors":{"value":["Xiaotian Song","Yanan Sun"]}},"version":2}],"count":9}