UniPool: Learning Expert-to-Layer Ownership from Brief Global Access | Summary
07 May 2026 | Paper Review Mixture-of-Experts Expert Routing LLM PretrainingContents
- Summary
- 1 Introduction
- 2 Related WORK
- 3 Motivating Observation: Substitutable Experts Under Fixed OWNERSHIP
- 4 Method
- 5 Experiments
- 6 Conclusion
- AI Use Statement
- Reproducibility Statement
- Appendix
- Brief Thoughts
This article explains the key points of UniPool: Learning Expert-to-Layer Ownership from Brief Global Access.
- 2026-05-07 (arXiv)
- Huang, Minbin, Shi, Han, Zheng, Chuanyang, Wu, Yimeng, Chen, Guoxuan, Yu, Xintong, Yin, Yichun, Cheng, Hong.
- The Chinese University of Hong Kong, Huawei Technologies, The University of Hong Kong
- Paper
- Github
Summary
- UniPool: Learning Expert-to-Layer Ownership from Brief Global Access studies whether MoE expert ownership should be decided after early training has established layer-specific routing preferences. UNIPOOLlock briefly exposes every layer to a global expert pool, learns a balanced disjoint allocation, and then continues training with layer-private expert execution.
- With matched expert-FFN storage and routed expert FLOPs, UNIPOOLlock reduces test loss by 0.0244–0.0367 across five dense-equivalent scales from 182M to 1.5B. At 1.5B after about 60B tokens, loss falls from 1.6320 to 1.6073; measured post-lock step-time overhead ranges from −0.7% to +2.8% in four configurations at 182M and 469M.
- Matched 182M controls show contributions from both early global access and the allocation it produces: random ownership from initialization gives loss 1.9319, random locking after global access gives 1.9173, and learned locking gives 1.9040. The evidence is limited by one lock schedule, single-seed results at larger scales, and zero-shot downstream evaluation.
1 Introduction
Standard MoEs assign each expert to a private layer before training, when the experts are exchangeable and ownership is not informed by data. UNIPOOL learns ownership from early global competition, then removes cross-layer sharing to recover layer-private execution.
- The motivation combines prior reports of within-layer expert similarity and pruning tolerance with the paper’s routing-randomization probe.
- UNIPOOLfull retains global access throughout training as a quality and parameter-efficiency reference; UNIPOOLlock is the main method.
- Figure 1 separates the temporary global-access topology from the final disjoint ownership topology.
Global sharing is a temporary ownership-learning mechanism rather than the final execution architecture of UNIPOOLlock. Each colored expert in the post-lock pool belongs to one layer, so retaining a global parameter container does not imply continued cross-layer execution.
2 Related WORK
The paper distinguishes learned ownership from sparse routing, load-balancing alternatives, predetermined cross-layer sharing, and changes to MoE initialization or activation patterns. UNIPOOL changes which layer owns each expert while keeping expert activation sparsity fixed.
- Switch-style balancing acts within private expert banks; UNIPOOL balances aggregate usage across the shared pool.
- ReXMoE, MoUE, Universal Transformers, and related methods prescribe sharing patterns in advance, whereas UNIPOOL learns an allocation and subsequently eliminates sharing.
- Upcycling changes initialization; EvoMoE and dense-training/sparse-inference methods change activation patterns. These approaches do not learn a disjoint expert-to-layer allocation from brief global access.
3 Motivating Observation: Substitutable Experts Under Fixed OWNERSHIP
The motivating probe replaces the learned router in one deep-half MoE layer with uniform random assignment, repeats the intervention across deep-half layers, and averages five-task downstream accuracy. The drop is only 1.0–1.6 points in Qwen1.5-MoE, DeepSeek-V2-Lite, and Qwen3-30B-A3B, indicating limited sensitivity to expert identity under this intervention.
- Table 1 reports raw zero-shot accuracy: 67.92 to 66.29, 54.19 to 53.03, and 73.02 to 72.06, respectively.
- This observation motivates the ownership hypothesis but does not establish that fixed ownership causes redundancy.
- Later controls test allocation quality under identical expert budgets, and a matched probe examines the trained UNIPOOL models.
Replacing one deep-half layer’s router at a time produces small average raw-accuracy losses across all three production models. This establishes the motivating sensitivity observation, not a causal attribution of redundancy to fixed ownership.
4 Method
UNIPOOL combines global-pool routing, a pool-level auxiliary loss, and NormRouter with an ownership lock. UNIPOOLlock uses global access during the initial ownership-learning phase, whereas UNIPOOLfull keeps it throughout training.
4.1 Global EXPERT POOL
Each layer retains its own router but selects experts from a common pool. Its routed FFN is (\operatorname{FFN}l(x)=\sum{i\in\operatorname{Top-}k_{V_l}(r_l(x))}g_{l,i}(x)e_i(x)), where V_l is the layer’s visible expert set; global access exposes every expert, and the lock restricts visibility to the learned allocation.
- For L layers and E private experts per layer, matched-budget comparisons use M = EL global experts.
- Both architectures activate the same k expert FFNs per layer and token; reduced-pool experiments use M < EL.
- Layer-specific routers are retained because hidden-state distributions differ across depth.
4.2 POOL-Level AUXILIARY LOSS
Per-layer balancing can suppress depth-specific preferences when an expert unused by one layer can serve another. UNIPOOL instead minimizes L_pool = α_pool M ∑_{i=1}^{M} f_i P_i, with dispatch fractions and mean routing scores averaged across layers.
- The objective discourages dead experts at the pool level without requiring each layer to distribute traffic uniformly over every expert.
- The loss decomposes into per-layer terms, allowing each layer to compute its contribution independently.
- Aggregate dispatch fractions come from the previous micro-batch to avoid cross-layer tensor dependencies under activation checkpointing.
4.3 Normrouter
The method adopts NormRouter (Zheng et al., 2026), which scores experts using s_i = σc max(0, z_i/(∥z∥₂ + ϵ)), with z = W_l h. L2 normalization removes layer-dependent logit scale, ReLU zeroes negative candidates, σ is a learnable scalar initialized to 1, and c calibrates the initial selected-score scale.
- Auxiliary objectives use mean scores rather than mean softmax probabilities.
- After locking, the router still computes all M logits and normalizes over the full pool; only expert visibility changes.
- Retaining full-pool normalization preserves score calibration across the transition.
4.4 Learning OWNERSHIP and Locking It
The ownership-learning phase encourages each layer’s effective expert count toward E using L_card = (λ_card/L) ∑_{l=1}^{L}(H(p_l) − log K_tar)², with K_tar = E. At step 1K, routing counts from the held-out 3% validation split determine an assignment maximizing retained routing mass under exclusive ownership and per-layer cardinality constraints.
- Because M = EL and each layer receives at least E experts, every layer receives exactly E and every expert has one owner.
- Off-allocation scores are attenuated with a cosine schedule over steps 1K–2K, followed by a permanent visibility lock at step 2K.
- The same schedule is used at 1.5B; in a 60K-step run, the final 96.7% of training uses layer-private expert execution.
Global access splits the same routed computation among more, smaller grouped GEMMs, making full-pool steps 22%–36% slower in the reported measurements. After locking, compact dispatch gathers each layer’s selected global indices into E local expert groups, restoring vanilla-shaped token permutation, metadata, and grouped GEMM execution.
- A visibility mask alone is insufficient if the dispatcher still materializes M expert columns and groups.
- Compact groups reference the original expert parameters without copying them.
- Global scoring and pool statistics remain active after locking, accounting for the residual overhead.
5 Experiments
The experiments evaluate loss, zero-shot downstream accuracy, training throughput, ownership controls, reduced pools, component ablations, and routing sensitivity. Expert-FFN storage and routed expert FLOPs are matched within each architecture comparison, while wider UNIPOOL routers add parameter and scoring overhead.
5.1 Setup
LLaMA-style backbones are trained on the Pile at 182M, 469M, 650M, 830M, and 1.5B dense-equivalent active scales. The first four use 60K steps, global batch 512, sequence length 1,024, and about 30B tokens; the 1.5B model uses 36 layers, width 1,600, global batch 1,024, 57,220 steps, and about 60B tokens.
- Training uses AdamW, cosine learning-rate decay, bf16, and Megatron-LM.
- The default budget is eight experts per layer with top-1 routing; UNIPOOL uses M = 8L experts.
- Additional 16E/top-2 and 32E/top-4 configurations on the 182M backbone retain full-width experts, yielding about 268M and 438M active parameters for vanilla MoE.
- Wider routers add 0.10%–0.37% stored parameters and up to 2.2% active parameters in the reported four-scale configurations.
At 182M, vanilla MoE and UNIPOOLfull average three seeds with test-loss standard deviation 0.0023; UNIPOOLlock averages two same-recipe seeds, 1.9017 and 1.9040. Other scales use one seed, and the authors treat loss differences below about 0.005 as ties.
- Final losses use a held-out test split disjoint from the 3% validation split used to choose ownership.
- Limited replication at larger scales restricts conclusions about small differences between locked and persistent full-pool variants.
5.2 Main Results
UNIPOOLlock improves test loss over vanilla MoE at all five default scales: reductions are 0.0288, 0.0367, 0.0244, 0.0365, and 0.0247 from 182M through 1.5B. It also improves loss by 0.0298 and 0.0272 under the 16E/top-2 and 32E/top-4 budgets.
- Across the six settings with both variants, locked minus persistent full-pool loss ranges from −0.0021 to +0.0064.
- UNIPOOLfull is not trained at 1.5B.
- At the four default scales with matched dense baselines, all MoE variants achieve lower loss than the dense models.
Every locked-model column improves over its matched vanilla baseline, including the 1.5B run and the higher-activation expert budgets. Locked and persistent full-pool losses are close, but the 650M gap of 0.0064 exceeds the paper’s approximate 0.005 tie threshold.
The validation curves show that UNIPOOLfull’s advantage persists after warmup at 182M, 469M, and 650M. The sharing-scope trajectories broadly preserve their endpoint ordering, showing that the reported advantage is not confined to the final checkpoint.
- Figure 4 plots validation loss, whereas the headline endpoint comparisons report held-out test loss.
- These curves concern persistent full-pool training rather than a direct trajectory comparison of UNIPOOLlock.
The persistent full-pool curves remain below vanilla after warmup at three scales, showing that the advantage is not confined to the endpoint. The sharing-scope panel broadly reproduces endpoint ordering, and all curves measure validation rather than final test loss.
Zero-shot evaluation covers ARC-Easy, ARC-Challenge, PIQA, HellaSwag, WinoGrande, LAMBADA, and RACE using raw accuracy. UNIPOOLlock improves the unweighted average by 0.68–1.57 points across seven configurations and wins 43 of 49 scale–task cells.
- Its average accuracy is within 0.3 points of UNIPOOLfull wherever both are evaluated.
- At 1.5B, average accuracy rises from 47.50 to 48.45 and all seven tasks improve.
- Individual tasks do not improve uniformly at every smaller-scale configuration.
Average downstream improvements accompany the loss gains, while individual task cells reveal exceptions at smaller scales. The 1.5B locked model improves all seven tasks and raises the average from 47.50 to 48.45.
In four measured configurations, full-pool execution adds 22%–36% step time, while compact post-lock execution ranges from −0.7% to +2.8% relative to vanilla. Weighting 2K full-pool steps and 58K compact steps gives end-to-end overhead of +0.6% to +3.5%.
- Measurements use 8 GPUs and cover 182M with 8E, 16E, and 32E, plus 469M with 8E.
- These measurements do not establish the same overhead at 650M, 830M, or the EP = 8 configuration used at 1.5B.
The global-access phase occupies only 2K steps, so compact execution limits weighted end-to-end overhead to +0.6%–+3.5% in the four measured configurations. These measurements support near-vanilla execution cost in those settings rather than a throughput claim for every tested scale.
5.3 Where Does the Gain Come From?
Control A fixes a random balanced disjoint allocation at initialization while retaining UNIPOOL’s router and losses. Its loss of 1.9319 matches vanilla’s 1.9317, showing that this combination without early global access does not explain the improvement.
- Predetermined routed shared-core, adjacent-reuse, and windowed-universal adaptations yield losses between 1.9262 and 1.9411 at 182M.
- The routed shared core improves loss by 0.0055 at 182M and 0.0167 at 469M, less than half the corresponding persistent full-pool improvement.
- These comparisons use parameter-matched adaptations rather than unrestricted reproductions of the original sharing methods.
Controls B and C share the early global-access schedule but lock random or learned ownership, respectively. B reaches 1.9173 and C reaches 1.9040, so retaining the learned allocation adds 0.0133 improvement beyond the benefit that survives under random locking.
- Relative to A, random locking after global access improves loss by 0.0146, while learned locking improves it by 0.0279.
- C and the other UNIPOOLlock seed, 1.9017, use the same recipe; their mean is 1.9029, equal to the persistent full-pool result at reported precision.
- The B–C comparison identifies allocation value conditional on global training, but does not separately isolate how that phase changes expert weights; a matched warm-start control remains missing.
The ordering separates random initial allocation from early global access and learned locking. Random locking after global access improves over initialization-time locking, but learned ownership recovers substantially more of the persistent full-pool gain.
Persistent expert reuse is modest and depth-dependent: UNIPOOLfull touches 94.1%, 89.5%, and 83.4% distinct experts across 12, 24, and 36 layers. Independent uniform routing predicts 94.5%, 94.2%, and 94.2%, respectively, so reuse beyond chance is clearer in the deeper configurations.
- Vanilla MoE and post-lock UNIPOOLlock necessarily touch distinct expert tensors across layers because ownership is disjoint.
- Limited persistent reuse is consistent with, but does not explain conclusively, the small quality gap between locked and full-pool training.
5.4 What the Full POOL Reveals
Reduced-pool experiments keep top-1 expert compute fixed while shrinking UNIPOOLfull’s expert storage. A 66.7% pool beats vanilla at 182M, 50% pools beat vanilla by 0.007 and 0.011 at 469M and 650M, and a 41.7% pool matches vanilla within 0.002 at 830M.
- At 182M, the 50% shared pool gives 1.9479 versus 1.9664 for a parameter-matched baseline with four private experts per layer.
- The smallest competitive fraction decreases across jointly scaled configurations; the study does not isolate depth as the sole cause.
- These parameter-efficiency results apply to UNIPOOLfull, not reduced-budget UNIPOOLlock.
Points below zero indicate that a reduced shared pool beats the full-budget private baseline at unchanged top-1 expert compute. Competitive storage fractions become smaller across the jointly scaled configurations, although depth is not isolated from other model changes.
The full-pool benefit persists with more full-width experts and with two recurrent expert steps per layer. In the looped setting, UNIPOOLfull lowers loss from 1.9077 to 1.8727 at 182M and from 1.7768 to 1.7389 at 469M.
- Attention executes once per layer; the normalization, router, and expert path execute twice with shared parameters.
- The second step recomputes routing from the updated hidden state.
- Within each looped comparison, physical expert storage and per-token expert computation are matched.
Component ablations show that shared-pool performance depends on balancing and score calibration. Shared experts with per-layer balancing and softmax give loss 1.9480, the pool loss improves it to 1.9180, and adding NormRouter reaches 1.9029.
- Figure 3 contrasts aggregate usage concentrated on a few experts with broader aggregate utilization and layer-specific preferences under UNIPOOLfull.
- NormRouter alone worsens private-bank loss to 1.9375, indicating that its benefit in these experiments depends on the shared candidate set.
- UNIPOOLfull also beats the private sigmoid, auxiliary-loss-free baseline of 1.9239.
- Intermediate sharing scopes G = 6, 4, and 2 beat vanilla, but the global pool G = 1 is best; improvement is not monotonic across intermediate G values.
Sharing alone is insufficient: pool balancing improves the shared-softmax model, and NormRouter adds a further improvement in the shared setting but not in private banks. All intermediate pool counts beat vanilla, but their losses do not form a monotonic sequence.
The left aggregate histogram contains a few dominant usage spikes, whereas the right distributes usage more broadly while retaining different preferences across layer rows. Some usage spikes remain on the right; the comparison illustrates the combined pool-loss and NormRouter configuration rather than isolating either component.
At 469M, randomizing one deep-half layer reduces five-task average accuracy by 1.3 points for vanilla, 3.3 for UNIPOOLlock, and 4.1 for UNIPOOLfull under matched eight-expert choice sets. The larger drops are consistent with stronger expert differentiation after global competition.
- UNIPOOLlock samples from its eight allocated experts; UNIPOOLfull samples from each layer’s eight most-used experts.
- These evaluations use length-normalized accuracy for four tasks and raw accuracy for WinoGrande, so they are not directly comparable to Table 3 averages.
- Selected gates are re-normalized during intervention, making the probe a sensitivity test rather than a causal test of ownership.
With eight randomized candidates per layer, the locked model is more sensitive to expert replacement than vanilla despite identical private-bank cardinality. Gate re-normalization means the larger accuracy drop is evidence of routing sensitivity consistent with specialization, not a causal ownership test.
6 Conclusion
The paper concludes that brief global access can learn an expert allocation whose benefit persists after sharing is removed. The evidence comprises loss improvement across five scales, losses close to the persistent full-pool reference where both variants are tested, and near-vanilla post-lock step time in the measured configurations.
- Reduced pools provide a separate parameter-efficiency result for persistent global access.
- The experiments do not establish robustness to other lock schedules or downstream post-training procedures.
AI Use Statement
The authors report using generative AI for manuscript language editing and LaTeX format conversion, not for generating experimental results or altering empirical data. They state that they reviewed the assisted content against the underlying experiments and source material and take responsibility for the final manuscript.
Reproducibility Statement
The reproducibility statement directs readers to the setup, model and hyperparameter appendices, ownership-lock implementation, controls, sharing baselines, evaluation protocols, pool-loss derivation, and NormRouter initialization. The paper also provides the code URL https://github.com/Centaurus-Alpha/UniPool.
Appendix
- A ALTERNATIVE EXPERT-SHARING BASELINES specifies matched-budget routed shared-core, ReXMoE-R2, MoUE-stateless, and DeepSeekMoE-style controls. Added controls use single seed 1234; the MoUE adaptation omits the trajectory-state Universal Router, and the DeepSeekMoE-style control replaces 16 routed experts with 15 routed experts plus one always-active layer-local expert under the 16E/top-2 budget. The latter comparison matches global batch size but uses different micro-batch sizes for the control and vanilla baseline.
- B LIMITATIONS AND FUTURE WORK notes that allocation time, lock time, attenuation schedule, and cardinality coefficient are not ablated. Reduced pools and recurrent execution are evaluated only for UNIPOOLfull; combining reduced budgets with locking would require smaller or non-uniform layer allocations, and few-shot and post-training evaluation remain untested.
- C MODEL AND MOE CONFIGURATIONS lists the four 30B-token backbones: L/H values are 12/768, 24/1024, 36/1024, and 48/1024, with global pools of 96, 192, 288, and 384 experts. All use sequence length 1024, SwiGLU, and FFN intermediate size 4 × H; UNIPOOLfull and UNIPOOLlock have identical router parameter counts.
- D LOOPED-EXPERT COMPATIBILITY DETAILS defines two recurrent expert updates per layer with residual scaling 1/R and R = 2. Attention remains outside the recurrence; both expert steps reuse normalization, router, and expert parameters while recomputing top-1 selection, giving 2L expert evaluations per token in both compared architectures.
- E ADDITIONAL TRAINING CURVES documents the validation trajectories in Figure 4. UNIPOOLfull remains below vanilla after warmup at 182M, 469M, and 650M; the sharing-scope curves broadly interpolate between global and private endpoints, with global sharing lowest for most of training.
- F HYPERPARAMETER DETAILS reports RMSNorm, SwiGLU, RoPE with base 1M, four KV heads, and untied embeddings. The four 30B-token scales use learning rate 5 × 10−4, minimum learning rate 5 × 10−5, 1% warmup, gradient clipping 1.0, bf16, and initialization standard deviation 0.01; the 1.5B runs use EP = 8 and about 1.49B active and 9.23B stored vanilla-MoE parameters.
- G OWNERSHIP LOCK AND COMPACT DISPATCH DETAILS warms λ_card to 5 × 10−3 over the first 1K steps and keeps the pool loss active throughout training. Compact execution sorts visible global indices, enforces zero off-mask assignments, and references original parameters without copying; phase-median throughput measurements use 8 GPUs, dropless grouped GEMM, bfloat16, and EP = TP = PP = 1.
- H OWNERSHIP CONTROLS fixes the comparison architecture at 12 layers, 96 experts, top-1 routing, and eight disjoint experts per layer after locking. These constraints match the final allocation cardinality of the 182M UNIPOOLlock model.
- H.1 MATCHED CONTROLS FOR EARLY EXPERT ACCESS holds architecture, NormRouter, full-pool normalization, auxiliary coefficients, seed, batch size, sequence length, learning rate, and training schedule constant across A–C. A and B share the same random allocation; C chooses ownership from validation routing counts, and its final NLL improves over B by 0.0133.
- I DISTINCT-EXPERT ACCOUNTING compares observed distinct expert visits with the independent uniform expectation (M\left[1-\left(1-\frac{1}{M}\right)^L\right]). At 36 layers, the full pool gives 30.03 distinct experts per token versus 33.90 expected; at 12 layers, full-pool reuse is near chance and reduced pools touch more distinct experts than chance.
- J ADDITIONAL ROUTING-RANDOMIZATION DETAILS gives production-model per-task results and the matched top-8 protocol for UNIPOOLfull. Expert identities are randomized one deep-half layer at a time, selected gates are re-normalized, and results are averaged across layers; UNIPOOLlock is reported only through its five-task averages rather than per-task scores.
- K POOL AUXILIARY LOSS: DETAILED DERIVATION shows that expressing aggregate mean scores as their layer average yields independent per-layer pool-loss terms. Dispatch fractions are accumulated without gradients and lagged by one micro-batch; routing scores carry gradients through the objective, while expert FFN parameters receive no direct gradient from this loss path.
- L NORMROUTER: MONTE CARLO INITIALIZATION DETAILS calibrates c using Gaussian logits, L2 normalization, ReLU, and the largest k components. Algorithm 1 estimates the expected reciprocal L2 norm of those selected components using (N=10^5) samples at initialization.
Brief Thoughts
The ownership controls provide more direct evidence for allocation quality than the routing-sensitivity probe. Random locking after global access preserves part of the benefit, while learned locking adds another 0.0133 loss reduction under the same final expert cardinality, supporting data-informed allocation conditional on the shared phase.
The practical result is that improved loss survives conversion to private execution with low measured overhead. The main unresolved checks are lock-schedule sensitivity, replicated larger-scale results, and throughput at larger expert-parallel configurations; reduced-pool efficiency remains a separate result for UNIPOOLfull.