Loading Now

Mixture-of-Experts: Powering Smarter, Safer, and More Efficient AI at Scale

Latest 82 papers on mixture-of-experts: Oct. 10, 2026

Mixture-of-Experts (MoE) models have rapidly become a cornerstone in the quest for highly capable yet efficient AI. By selectively activating a subset of specialized ‘experts’ for each input, MoEs promise to scale models to unprecedented sizes without a proportional increase in computational cost. However, realizing this potential comes with its own set of challenges, from optimizing routing and communication to ensuring safety, interpretability, and practical deployment. Recent research, as highlighted in a collection of groundbreaking papers, is pushing the boundaries of MoE capabilities, addressing these critical areas and paving the way for the next generation of intelligent systems.

The Big Idea(s) & Core Innovations

One of the central themes emerging from recent work is the intelligent orchestration and adaptation of experts across various tasks and architectures. For instance, DivMoE by Lou et al. from the National University of Singapore and Shanghai Jiao Tong University, tackles a critical pathology in fine-grained MoE upcycling: routing collapse. They demonstrate that combining domain-specialized expert initialization with diversity-constrained routing is crucial to prevent performance degradation, proving that cross-domain expert composition can surprisingly improve domain-specific tasks. This highlights the importance of thoughtful expert initialization and selection to harness diverse knowledge.

Meanwhile, several papers delve into making MoE models more robust and reliable. Zheng et al. from Shanghai AI Laboratory and co-authors introduce ReSI: Recursive Safety Improvement toward Resistant and Resilient AI [https://arxiv.org/pdf/2610.12233], a recursive framework that iteratively enhances AI safety. By applying diverse red-teaming methods and automated training recipe search, ReSI significantly reduces attack success rates on challenging benchmarks while preserving general capabilities. Complementing this, Teng et al. from New York University propose DRMoET (Distributionally Robust Mixture-of-Experts Training) [https://arxiv.org/pdf/2610.07207]. They address a hidden reliability problem where sparse routing can send tokens to undertrained experts, introducing an objective that optimizes high-loss routing outcomes to narrow expert-quality gaps and improve routing robustness, especially for mid-tier experts. Adding a crucial layer of interpretability and safety, Yang et al. from the University of Washington and NYU, in their paper “Reading, Not Manipulating: Leveraging Router Logits for Multimodal Safety in MoE Vision-Language Models” [https://arxiv.org/pdf/2610.07774], discover that router logits provide highly predictive signals for multimodal safety. Their lightweight router-logit detector identifies unsafe requests before generation, offering substantial safety improvements without altering model parameters or routing behavior. This aligns with Siddiky et al.’s work “Frequency Is Not Sensitivity: Identifying Safety-Sensitive Experts in Sparse MoE LLMs” [https://arxiv.org/pdf/2610.02910], which shows that router-gradient sensitivity, not mere activation frequency, is a superior proxy for identifying experts whose suppression can effectively reduce harmful refusals.

From a performance and efficiency standpoint, numerous advancements focus on optimizing MoE inference and training. Ma et al. from the University of Illinois Urbana-Champaign and co-authors, in “S3: Spectral Null-Space Swap Makes Reasoning Models Efficient” [https://arxiv.org/pdf/2609.37976], introduce a training-free method to combine components of ‘Non-thinking’ and ‘Thinking’ models, reducing token overhead by ~27% while improving accuracy, by identifying core reasoning in the null-space of weights. For specific tasks, such as automated program repair, Shah and Sharma from the National Institute of Technology, Calicut, demonstrate in “Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair” [https://arxiv.org/pdf/2610.08173] that MoE models achieve comparable correctness to much larger dense models with significantly fewer active parameters. Further efficiency is gained through novel routing and compression techniques. Chai et al. from Shanghai AI Laboratory and Shanghai Jiao Tong University propose Elastic Expert Routing in “Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts” [https://arxiv.org/pdf/2610.11575], which stochastically samples active expert budgets during training to smooth the rigid top-k selection threshold, yielding consistent performance gains without inference overhead. For compression, Cao et al. from Kyoto University and RIKEN AIP introduce Shared Low-rank Basis Factorization (SLBF) [https://arxiv.org/pdf/2610.09342], a data-free weight reconstruction method that avoids the structural errors of pruning/merging, demonstrating superior compression. Meanwhile, Sun et al. from Harbin Institute of Technology, Shenzhen, propose MoRA: MoE Pruning via Router Bias Learning and Expert Approximation [https://arxiv.org/pdf/2610.00367], which uses learnable router biases and a diversity regularizer to identify critical experts and approximate pruned ones, showing that learned biases correlate better with expert deletion sensitivity than activation frequency.

Beyond traditional language models, MoEs are expanding into diverse modalities. Cui et al. from AI at Meta and the University of Central Florida present MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models [https://arxiv.org/pdf/2609.08663], the first MoE-based Universal Multimodal Embedding model that scales capacity along the expert axis, achieving new state-of-the-art results on MMEB-V2. For robotics, Lee et al. from Yonsei University introduce WARP-VLA [https://arxiv.org/pdf/2610.11508], an MoE framework that adapts wrist-camera features to a canonical space, enabling vision-language-action policies to handle diverse camera configurations without explicit calibration. In video generation, Xu et al. from the University of Chinese Academy of Sciences and ByteDance propose SplitMoE [https://arxiv.org/pdf/2609.38140] to address the “uniformity trap” in video diffusion models, splitting experts into semantic and generic branches with prototype-guided routing to achieve coherent semantic grouping and improved video quality. Even medical applications are benefiting: Peng et al. from Northwestern University introduce an Anatomy-Aware Mixture-of-Experts model for bronchoscopy accessibility prediction from 3D CT [https://arxiv.org/pdf/2609.37386], outperforming human experts by integrating specialized modules for local morphology, anatomical priors, and airway geometry.

Further integrating MoEs with complex systems, Ma et al. and collaborators introduce “CipherGenome: Homomorphic Inference for Genomic Mixture-of-Experts” [https://arxiv.org/pdf/2609.35883], presenting a privacy-preserving protocol for genomic MoE models by outsourcing expert computations under module-LWE encryption, achieving exact homomorphic computation without bootstrapping. This is complemented by their “GenomeOcean Anywhere: Private WebGPU Inference for Genome MoEs” [https://arxiv.org/pdf/2609.35882], which runs a 15B genome MoE in web browsers using Lagrange coded computing to protect genomic privacy. These works demonstrate how MoE’s modularity can be leveraged for privacy-preserving distributed computing.

Under the Hood: Models, Datasets, & Benchmarks

Recent MoE research is heavily reliant on and contributes to a rich ecosystem of models, datasets, and benchmarks. The Qwen family of models (e.g., Qwen3-30B-A3B, Qwen3.5-35B-A3B, Qwen3-VL) appears frequently as a versatile backbone for various experiments, including safety, compression, multilingual alignment, and multimodal tasks. DeepSeek-MoE variants (e.g., DeepSeek-V2-Lite, DeepSeek-V3-671B) also serve as important testbeds, especially for efficiency and inference optimization.

Key resources enabling these innovations include:

  • Safety Benchmarks: WildJailbreak, HarmBench, X-Teaming (used in ReSI) are crucial for evaluating and improving AI model resistance and resilience to attacks.
  • Efficiency Benchmarks: MMLU, GSM8K, HumanEval, MBPP (for LLMs), VBench-2, T2V-CompBench (for video diffusion), and MMEB-V2, MRMR (for multimodal embeddings) are standard for assessing performance-compute trade-offs.
  • Specialized Datasets:
    • **ARC-TGI

Share this content:

mailbox@3x Mixture-of-Experts: Powering Smarter, Safer, and More Efficient AI at Scale
Hi there 👋

Get a roundup of the latest AI paper digests in a quick, clean weekly email.

Spread the love

Discover more from SciPapermill

Subscribe to get the latest posts sent to your email.

Post Comment

Discover more from SciPapermill

Subscribe now to keep reading and get access to the full archive.

Continue reading