Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking

Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert Masking

Mixture-of-Experts (MoE) architectures scale large language models (LLMs) while preserving computational efficiency through sparse activation. Despite their widespread adoption, the relative importance of individual MoE layers remains insufficiently characterized, particularly for model compression. This paper presents a systematic layer-wise sensitivity analysis of the Qwen3.6-35B-A3B model (40 MoE layers, 256 experts per layer, top-8 routing) using magnitude-based expert masking on the XLCoST cross-lingual code translation benchmark. We conduct a multi-phase study spanning 100, 300, and 500 prompt evaluation scales across three H100 GPU servers. Our central finding is that layer sensitivity is strongly depth-dependent: early layers (0-9) and middle layers (10-29) are highly fragile to expert masking, while late layers (30-39), and especially very-late layers (35-39), tolerate aggressive masking of low-magnitude experts.

Flat all-layer masking at 30% retains only 150/300 Good+Similar outputs at 300-prompt scale, whereas late-focused policies retain 249-255/300 while masking 640-1,145 experts. On a later 500-prompt held-out validation slice, the narrow very-late policy (layers 35-39 @ 50%) achieves the strongest quality/masked-expert tradeoff among tested candidates, retaining 419/500 Good+Similar outputs while masking only 640 of 10,240 total experts. We additionally characterize top-k routing width reduction from 8 to 6 active experts per token, which shows a large observed wall-clock reduction on a 100-prompt probe with no Good+Similar loss, though it does not yet compose cleanly with aggressive expert masking. These findings provide an empirical foundation for depth-aware MoE expert masking and establish a practical path toward physical weight surgery, activation-based expert scoring, and training-based recovery. ar Xiv Labs is a framework that allows collaborators to develop and share new ar Xiv features directly on our website.

Both individuals and organizations that work with ar Xiv Labs have embraced and accepted our values of openness, community, excellence, and user data privacy. ar Xiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for ar Xiv's community? Learn more about ar Xiv Labs .

Sources cited: ๐Ÿ“ฐ ArXiv AI โ†—

โšก Effects Interpreter

๐ŸŒWorld Economy

  • โ–ถCross-border money flows can gradually change direction after events like this.
  • โ–ถEconomies far from the headline can still catch the aftershocks.

๐Ÿ™๏ธLocal Economy

  • โ–ถHousehold budgets may notice a small ripple before long.
  • โ–ถLocal suppliers who import goods could pass on any change in costs.

๐ŸฆRates & Banks

  • โ–ถYour loan or mortgage rate is more likely to drift than to lurch here.
  • โ–ถBanks tend to wait and see before nudging the rates they offer.

โค๏ธHealth

  • โ–ถCommunity wellbeing could dip a little while people wait for clarity.
  • โ–ถLocal health services could get busier depending on how things develop.

๐Ÿ’ทWealth

  • โ–ถYour pension or investments might sway a touch as markets digest this.
  • โ–ถSavings and portfolios can see short-lived ups and downs after news like this.

๐Ÿ Housing

  • โ–ถHome costs usually respond later, once the bigger picture settles.
  • โ–ถFirst-time buyers might keep half an eye on mortgage rates after this.
Share: ๐• Twitter Facebook LinkedIn WhatsApp

Editorial note: This analysis was produced by the News Effects Interpreter, an AI editorial tool that cross-references 1 independent news sources and contextualises events in terms of their real-world impact on ordinary people. Original reporting is linked above. News Effects does not alter the facts of source reports.