Gorio Tech Blog search

ERNIE 5.0 Technical Report | Summary

|

Contents

This article explains the key points of ERNIE 5.0 Technical Report.

  • 2026-02-04 (arXiv)
  • Wang, Haifeng, Wu, Hua, Wu, Tian, Sun, Yu, Liu, Jing, Yu, Dianhai, Ma, Yanjun, He, Jingzhou, He, Zhongjun, Hong, Dou, et al.
  • Baidu
  • Paper

Read this article in Korean


Summary

  • ERNIE 5.0 Technical Report describes a trillion-parameter foundation model for text, image, video, and audio understanding and generation. Its shared autoregressive backbone uses Next-Group-of-Tokens Prediction and modality-agnostic expert routing, with an activation rate below 3%.
  • Elastic training varies model depth, expert-pool width, and routing top-k during pre-training to support multiple deployment configurations. In the experimental ERNIE 5.0-Exp family, reducing routing top-k to 25% yields a reported decoding speed improvement of more than 15%; a jointly elastic variant uses 53.7% of activated parameters and 35.8% of total parameters and scores 75.17 versus 75.55 for the full model after additional mid-training and post-training.
  • Post-training combines supervised fine-tuning with unified multimodal reinforcement learning, supported by order-preserving rollout scheduling, mixed-granularity importance sampling, positive-sample masking, and adaptive reasoning hints. Results are strongest in factual knowledge, multi-instruction following, selected visual perception tasks, and selected audio tasks, with larger gaps on difficult reasoning and some video-understanding benchmarks. High-resolution visual generation also uses a separately trained diffusion refiner.

1 Introduction

The report addresses how to unify multimodal understanding and non-text generation without relying primarily on generators attached to a pre-trained language model. ERNIE 5.0 trains its backbone jointly across modalities from the beginning, routes tokens through a shared expert pool, and incorporates deployment flexibility into pre-training.

  • The introduction describes an “ability seesaw” in late-fusion systems: greater multimodal integration can compromise language performance. Native joint training is the proposed response, but the evaluations do not isolate this effect against a matched late-fusion control.
  • The contribution also includes multimodal RL stabilization and distributed infrastructure for training a trillion-parameter ultra-sparse MoE.

2 Architecture

Figure 1 organizes ERNIE 5.0 around a shared transformer backbone with text, vision, and audio input–output interfaces. Tokenization, embeddings, and output heads remain modality-specific, while sequence modeling, multimodal attention, and expert routing operate within the shared backbone under Next-Group-of-Tokens Prediction.

The shared router and expert pool sit inside the unified transformer layers, whereas embeddings, encoders, decoders, and output heads remain modality-specific. The diagram distinguishes backbone-level unification from eliminating specialized input–output components.

Shared multimodal transformer backbone with modality-specific text, vision, and audio interfaces.
Shared multimodal transformer backbone with modality-specific text, vision, and audio interfaces.

2.1 Unified Autoregressive Backbone with Ultra-Sparse Mixture-of-Experts

The backbone serializes heterogeneous inputs into a shared sequence and routes tokens according to their representations rather than explicit modality identifiers. Text uses Next-Token Prediction with Multi-Token Prediction (MTP), vision uses Next-Frame-and-Scale Prediction (NFSP), and audio uses Next-Codec Prediction (NCP).

  • The ultra-sparse, fine-grained MoE has an activation rate below 3%; auxiliary-loss-free load balancing stabilizes expert utilization.
  • The input–output interfaces accommodate different representation requirements: understanding emphasizes abstract semantics, whereas generation also requires fine-grained perceptual information.
  • The report discloses the trillion-parameter scale and activation-rate bound, but not a complete production-backbone specification with exact layer and expert counts.

2.2 Visual Modeling

2.2.1 Vision Tokenization treats images as single-frame videos and constructs a causal 3D convolutional tokenizer by inflating a pre-trained causal 2D multi-scale image tokenizer. Adversarial supervision and semantic regularization support reconstruction and representation quality; progressively switching from low-bit to higher-bit tokenizers introduces finer discrete representations during backbone training.

  • Bit-wise quantization represents visual latents as groups of bit-codes. The switching schedule targets early instability with large vocabularies, but exact bit counts and switching milestones are not disclosed.

2.2.2 Visual Understanding with Dual-Path Hybrid Representation avoids the compression associated with low-dimensional quantized features by combining pre-quantization CNN perceptual features with ViT semantic features. As shown in Figure 2, an Attention-based Patch Merger projects CNN features into the ViT feature space, concatenates the two paths along the patch dimension, applies self-attention, and mean-pools the result before projection into the backbone.

  • Image understanding groups K = 4 adjacent patches; video understanding groups K = 16 patches spanning 4 neighboring frames.
  • The authors report improvements over CNN-only, ViT-only, and naive MLP fusion, especially for documents and charts, without a numerical fusion-ablation table in this PDF.

2.2.3 Visual Generation with Next-Frame-and-Scale Prediction generates each frame from coarse to fine spatial scales, then predicts subsequent frames. Prediction is parallel within a scale, with local bidirectional attention and causal access to previous scales and historical frames.

  • Unified Spatiotemporal Rotary Positional Embedding (Uni-RoPE) assigns temporal and spatial coordinates to visual tokens. Text and audio use equal coordinates following their sequence index, while multi-scale visual coordinates are center-aligned.
  • Random bit flips in historical tokens train correction of accumulated errors. Early-scale loss reweighting, windowed temporal attention, and random historical frame masking further support visual generation.
  • A separately trained cascaded diffusion refiner enhances low-resolution autoregressive outputs. The shared backbone is therefore autoregressive, but the complete high-resolution visual generation pipeline is not exclusively autoregressive.

Understanding uses merged pre-quantization CNN–ViT features, whereas generation uses discrete multi-scale tokens. These different representations feed a common backbone, with video generation extending spatial next-scale prediction along the frame dimension.

Visual understanding through hybrid feature merging and visual generation through Next-Frame-and-Scale Prediction.
Visual understanding through hybrid feature merging and visual generation through Next-Frame-and-Scale Prediction.

2.3 Audio Modeling

2.3.1 Audio Tokenization uses Residual Vector Quantization (RVQ) at 12.5 Hz, assigning the first code to semantic information and subsequent codes to residual acoustic detail. The first-code representation is aligned with average-pooled Whisper encoder outputs, while residual codes preserve properties such as timbre and prosody.

  • 2.3.2 Audio Understanding and Generation with Next-Codec Prediction sums level-specific code embeddings for understanding rather than flattening all codes into a longer temporal sequence.
  • For generation, heads in the upper transformer layers predict semantic and residual codes successively across depth. Each code embedding is added back to the hidden state; training uses ground-truth feedback through teacher forcing, whereas inference uses predicted codes.
  • Figure 3 illustrates the depth-wise structure. A speaker embedding conditions speech timbre, and an audio decoder converts the completed code hierarchy into waveforms.

Code embeddings are summed at the input for understanding, but their codes are predicted successively across upper transformer layers for generation. Feedback between prediction levels avoids expanding the temporal sequence by flattening every codebook.

Depth-wise audio modeling with additive input embeddings and hierarchical Next-Codec Prediction.
Depth-wise audio modeling with additive input embeddings and hierarchical Next-Codec Prediction.

3 Pre-Training

Pre-training jointly develops representations across modalities, progressively extends context length, and optimizes an elastic family of sub-networks. The section provides data categories and selected optimization settings rather than a complete recipe with exact corpus sizes, modality proportions, and total training compute.

3.1 Pre-Training Data

Text Data includes multilingual web crawls, curated corpora, books, scientific publications, code repositories, and structured knowledge. Multimodal Data includes image–text, video–text, audio–text, and interleaved sequences, organized by input and output modalities through a standardized data platform.

  • The retrained text tokenizer uses UTF-16BE byte-level fallback and BPE dropout. For languages without whitespace boundaries, decomposable long unspaced phrases are filtered from the tokenizer vocabulary to reduce sparsity.
  • The report describes heuristic and model-based filtering, deduplication, and benchmark decontamination. It characterizes the corpus as comprising trillions of text tokens and multimodal instances without exact per-modality quantities or decontamination results.

3.2 Training Recipe

Stage 1: 8K Pre-Training uses a maximum context of 8K tokens, a 2,000-step linear warmup to 1 × 10−4, and a constant learning rate for the remainder of that stage. The global batch increases from 14M to 56M tokens, and the RoPE base is set to 1,000,000 from the outset.

  • Stage 2: 32K&128K Mid-Training extends context to 32K and then 128K tokens with the global batch unchanged, using cosine annealing from 1 × 10−4 to 1 × 10−5.
  • The auxiliary-loss-free balancing bias update speed decreases from 1 × 10−4 to 1 × 10−5, while the MTP loss weight decreases from 0.3 to 0.1.
  • Posterior-based weighting rescales modality-specific autoregressive losses to a common interval to reduce optimization imbalance.

3.3 Once-For-All with Elastic Training

Once-For-All with Elastic Training optimizes a super-network whose smaller configurations select subsets of layers and experts and vary the number of activated experts per token. Figure 4 distinguishes Elastic Depth, which changes active layers; Elastic Width, which changes the available expert pool; and Elastic Sparsity, which changes routing top-k without reducing stored model size.

  • Elastic Depth uses the full model with probability 75% and a reduced-depth sub-network with probability 25%.
  • Elastic Width retains all experts with probability 80% and restricts routing to a randomly sampled expert subset with probability 20%.
  • Elastic Sparsity retains the default routing configuration with probability 80% and samples a smaller top-k with probability 20%. Hidden-size elasticity is not implemented; extracted sub-networks can serve as starting points for mid-training or fine-tuning.

Layer removal and expert-pool restriction can reduce the parameters stored for a deployed model, whereas lowering top-k primarily reduces activated computation. The three controls address different deployment constraints.

Elastic depth, width, and routing sparsity within a shared MoE super-network.
Elastic depth, width, and routing sparsity within a shared MoE super-network.

4 Post-Training

Post-training follows supervised fine-tuning (SFT) with unified multimodal reinforcement learning (UM-RL). SFT develops instruction following and long-chain reasoning, while multi-stage RL combines reasoning, agent, and instruction-following tasks using an extended unified verifier system.

  • The main challenges are rollout cost, training–inference numerical mismatch amplified by sparse expert routing, and heterogeneous tasks with sparse rewards.
  • The report describes the following stabilization techniques but does not disclose a complete verifier specification or separate quantitative ablations for every component.

4.1 Enhancing Rollout Efficiency with Unbiased Replay Buffer

Rollout generation is reported to account for more than 90% of RL training time. U-RB: Unbiased Replay Buffer Generation uses spare inference capacity for future groups while requiring each training iteration to use its originally assigned query group, preventing shorter responses from displacing unfinished responses in that group.

  • The inference pool has capacity equal to the training batch size multiplied by the buffer size, while the training pool holds one training batch.
  • An iteration waits until the longest rollout in its assigned group reaches [EOS], then transfers that group’s trajectories for optimization. Trajectories may resume from earlier inference runs.
  • Figure 5 contrasts this ordering constraint with synchronous waiting and APRIL’s completion-count stopping rule. “Unbiased” here concerns query admission and completion-order bias, not a demonstrated guarantee that all resumed trajectories are sampled from the current policy.

U-RB keeps long query 11 in the current training group while using spare capacity to begin queries 16 and 17 for later iterations. APRIL instead admits query 16 before query 11, illustrating the completion-order bias targeted by the ordering constraint; the figure is a scheduling example rather than a throughput measurement.

Rollout scheduling under synchronous RL, APRIL, and the order-preserving Unbiased Replay Buffer.
Rollout scheduling under synchronous RL, APRIL, and the order-preserving Unbiased Replay Buffer.

4.2 Stabilizing Training with Mitigated Entropy Collapse

MISC: Multi-granularity Importance Sampling Clipping combines token-level training–inference correction with a length-normalized sequence-level policy importance ratio. Figure 6 shows a gradual entropy decrease and continued accuracy improvement, whereas directly applying sequence-level IcePop-calibrated GSPO produces an entropy spike and an accuracy decline.

  • The report attributes the unstable comparison to sequence-level truncation removing many low-entropy responses under training–inference mismatch. Its use of “entropy collapse” includes abrupt entropy increases as well as decreases.
  • WPSM: Well-learned Positive Sample Mask tracks query success rates and downweights responses when average accuracy exceeds τ and response entropy is below η. A masking strength α ∈ [0, 1] controls the reduction.
  • WPSM is intended to redirect gradient effort toward harder queries; the PDF supplies neither threshold values nor a standalone quantitative WPSM ablation.

The mixed-granularity run continues improving accuracy while entropy decreases gradually. The sequence-level comparison develops a sharp entropy increase and declining accuracy, supporting stabilization in this reported run without establishing generality across seeds or tasks.

Accuracy and policy-entropy trajectories for mixed-granularity and sequence-level IcePop-calibrated RL objectives.
Accuracy and policy-entropy trajectories for mixed-granularity and sequence-level IcePop-calibrated RL objectives.

4.3 Boosting Sample Efficiency with Hint-based Learning

AHRL: Adaptive Hint-based Reinforcement Learning addresses hard queries whose rollouts all receive zero reward, leaving group-relative optimization without a useful learning signal. It appends partial reasoning sketches to the query and gradually reduces the revealed fraction, as illustrated by the regular 24-gon problem in Figure 7.

  • The schedule is phint(xᵗ) = pinitial · exp(−γ · t · passinitialˣ), where t is the training iteration, γ is the decay rate, and passinitialˣ is the SFT model’s pass@k on the query.
  • Hints decay more slowly for queries with lower initial success. The mechanism is described without numerical sample-efficiency gains or a controlled assessment of its contribution to final unhinted performance.

The regular 24-gon example adds the beginning of a reasoning sketch rather than the final answer. It illustrates how AHRL changes exploration conditions for a hard query, but does not quantify learning gains.

A mathematical query augmented with a partial reasoning hint for Adaptive Hint-based Reinforcement Learning.
A mathematical query augmented with a partial reasoning hint for Adaptive Hint-based Reinforcement Learning.

5 Infrastructures

The infrastructure is built on PaddlePaddle and extends the ERNIE 4.5 training system. Its four stated challenges concern MoE memory and communication, heterogeneous tokenizer workloads, sample-dependent attention patterns, and scalable RL with consistent computation.

  • Tremendous Memory Pressure and Communication Overhead: large expert parameters and inter-node token dispatch motivate hybrid parallelism and fine-grained memory control.
  • Multimodal Training with Multiple Tokenizers: tokenizers are separated from the backbone so each component can use an appropriate parallelization strategy. Flexible Attention Patterns across Modalities: FlashMask supports causal text and audio processing alongside local bidirectional visual attention.
  • Scalable RL under Throughput and Consistent Constraints: disaggregated execution coordinates training, rollout, and environment interaction while addressing numerical mismatch, completion-order bias, and heterogeneous resource utilization.

5.1 Hybrid Parallelism for Training at Scale

The final configuration combines 4-way tensor parallelism, 12-way pipeline parallelism with virtual stages, 64-way expert parallelism, ZeRO-1 data parallelism, and context parallelism for long-context training. DeepEP handles inter-node communication, and training uses no token dropping despite occasional routing-induced out-of-memory events.

  • FP8 mixed-precision training and FP8 activation storage reduce peak memory demand.
  • Dynamic adaptive activation offloading is triggered when memory is insufficient; sub-batch computation reduces large allocation requests.
  • Automatic memory defragmentation uses CUDA virtual memory management (VMM). The report describes these measures as enabling training but does not quantify their individual overheads.

5.2 Disaggregation Architecture for Multimodal Training

Tokenizer-backbone disaggregation deploys modality-specific tokenizers as independent, horizontally scalable services on dedicated compute nodes under data parallelism. The backbone retrieves encoded representations through remote calls, allowing its parallelization strategy to differ from those of the tokenizers.

  • This separation targets load imbalance caused by variable sequence lengths and modality-dependent encoding costs. A matched throughput comparison with colocated tokenizers is not provided.

5.3 FlashMask for Flexible Multimodal Attention

FlashMask supports globally causal attention with locally bidirectional visual regions and masks that differ across samples. The report claims up to a 200% operator-level speedup over FlexAttention and more than 20% end-to-end training acceleration.

  • Kernel-level integration with context parallelism is reported to improve performance by 80% relative to the Megatron-LM solution.
  • The PDF does not specify hardware, input-shape distributions, or detailed measurement protocols for these comparisons.

5.4 Scalable and Disaggregated RL Infrastructure

The disaggregated RL infrastructure uses a centralized controller to coordinate training, inference, environment interaction, and reward evaluation asynchronously. A unified FP8 stack uses identical operators for training and rollout, together with Rollout Router Replay, to reduce numerical and routing mismatch.

  • The order-preserving replay buffer implements the query-admission constraint described in Section 4.1.
  • Elastic CPU pooling virtualizes idle cluster CPU capacity for environment interaction and verification. Wall-clock and cost-efficiency benefits are described qualitatively rather than measured in infrastructure ablations.

6 Evaluations

Evaluation uses the internal ERNIE-Eval framework across language, vision, and audio tasks, followed by routing analysis and elastic-training experiments. Code evaluations use LiveCodeBench and SandboxFusion for execution-based assessment.

  • Results are point estimates without confidence intervals. Language base-model comparisons specify few-shot settings, but prompting and sampling protocols are not comparably complete for all post-trained baselines.

6.1 Evaluation on Language Benchmarks

Language Benchmarks cover factual knowledge, general exam-style tasks, STEM, coding, multilingual understanding, reasoning, instruction following, and agents. Evaluation of Pre-trained Models. Table 1 shows ERNIE 5.0-Base leading most reported rows, including PreciseWikiQA10-shot at 74.48, ChineseSimpleQA10-shot at 90.09, MMLU-Pro5-shot at 75.58, HumanEval+0-shot at 80.86, and MMMLU5-shot at 78.94.

  • Kimi K2-Base is higher on C-Eval, CMMLU, and MBPP+. LiveCodeBench v6 covers 2408 to 2505 and uses a 1-shot setting for the base-model comparison.
  • Evaluation of Post-trained Models. Table 2 reports the highest listed scores on SimpleQA (74.01), ChineseSimpleQA (86.03), MultiChallenge (65.98), Multi-IF (85.56), ACEBench-en (87.70), and ACEBench-zh (89.60).
  • Difficult reasoning remains weaker: ERNIE 5.0 scores 25.81 on HLE, 79.58 on HMMT 2025, and 76.21 on LiveCodeBench v6, compared with Gemini 3-Pro’s 37.50, 93.33, and 86.34. Its BrowseComp-zh score of 64.71 also trails DS V3.2-Thinking’s 65.00, qualifying the report’s claim of leadership on that task.

Leading factual-QA and multi-instruction scores coexist with lower HLE, HMMT 2025, and LiveCodeBench v6 scores than Gemini 3-Pro. Agent performance is also task-dependent: ERNIE leads ACEBench but not TAU2-Bench, BrowseComp-zh, or SpreadSheetBench.

Post-trained language results across knowledge, reasoning, coding, instruction following, and agent tasks.
Post-trained language results across knowledge, reasoning, coding, instruction following, and agent tasks.

6.2 Evaluation on Vision Benchmarks

Visual Understanding & Generation Benchmarks cover visual STEM and reasoning, documents, general VQA, video understanding, and image and video generation. Evaluation of Visual Understanding: Table 3 shows strengths on VLMAreBlind (91.38) and DocVQAval (95.45), but ERNIE 5.0 trails Gemini 3-Pro on MMMU-Pro (68.63 versus 81.00) and VideoMME(w sub) (81.35 versus 88.40).

  • Table 4 reports only ERNIE 5.0-Base because comparable public visual base-model results are scarce. Scores include MathVista 84.40, ChartQA 87.68, and OCRBench 875; post-training improves several tasks but not all, with ChartXiv-DQ falling from 90.35 to 89.05.
  • Evaluation of Visual Generation: Table 5 reports GenEval increasing from 88.4 for the base model to 90.1 after post-training, compared with Qwen-Image at 91.0 and Nano Banana Pro at 89.0.
  • Table 6 reports VBench Quality 84.40, Semantic 83.40, and Overall 84.20. ERNIE 5.0 leads the listed models on Semantic, while Veo3 leads Quality (85.70) and Overall (85.06); these results do not establish superiority on every aesthetic or temporal-quality dimension.

VLMAreBlind and DocVQAval are standout ERNIE 5.0 results, whereas MMMU-Pro and several video-understanding benchmarks show gaps to stronger comparators. The results do not establish general leadership in visual reasoning.

Post-trained visual reasoning, document understanding, VQA, and video-understanding comparisons.
Post-trained visual reasoning, document understanding, VQA, and video-understanding comparisons.

The base model performs visual understanding before instruction tuning. Because the table includes only ERNIE 5.0-Base, it cannot establish an advantage over other visual base-model architectures.

ERNIE 5.0-Base results on visual understanding benchmarks.
ERNIE 5.0-Base results on visual understanding benchmarks.

GenEval improves from 88.4 to 90.1 after post-training, placing ERNIE 5.0 above Nano Banana Pro's 89.0 but below Qwen-Image's 91.0. This comparison assesses object-focused text-to-image alignment rather than comprehensive aesthetic quality.

Image generation comparison on GenEval for base, post-trained, and specialized models.
Image generation comparison on GenEval for base, post-trained, and specialized models.

ERNIE 5.0 obtains the highest listed semantic score, 83.40, while Veo3 leads quality and overall scores. Semantic alignment is therefore a relative strength without implying superior video quality on every dimension.

VBench quality, semantic, and overall scores for video generation.
VBench quality, semantic, and overall scores for video generation.

6.3 Evaluation on Audio Benchmarks

Audio Understanding & Generation Benchmarks cover ASR, VoiceBench speech-based interaction, general audio understanding, and SEED-TTS speech generation. Evaluation of Audio Understanding: Table 7 reports ASR WER of 0.31 on AISHELL-1, 1.16 on LibriSpeech clean, and 0.83 on Fleurs-zh, plus the highest listed scores on TUT2017 (68.09) and CochlScene (82.77).

  • Performance is uneven: Qwen3-Omni-Instruct is stronger on WenetSpeech, while ERNIE 5.0’s VoiceBench CommonEval and SD-QA scores decline from the base model’s 4.39 and 86.44 to 3.74 and 77.58. MMAU also decreases from 80.80 to 80.40.
  • Evaluation of Audio Generation: Table 8 reports SEED-TTS test-zh | test-en WER of 1.35 | 1.54 after post-training, versus 3.41 | 2.44 for the base model.
  • CosyVoice 3 scores 0.71 | 1.45 and Qwen3-Omni scores 1.07 | 1.39. SEED-TTS WER measures content consistency, not speech naturalness, prosody quality, or speaker similarity.

ERNIE leads selected ASR and acoustic-scene rows, but several VoiceBench scores weaken after post-training. Interpretation must respect metric direction: lower ASR WER is better, while higher dialogue and understanding scores are better.

ASR, VoiceBench, and general audio-understanding results for unified and specialized models.
ASR, VoiceBench, and general audio-understanding results for unified and specialized models.

6.4 Discussion

6.4.1 Modality-Agnostic Expert Routing examines utilization, collaboration, and load balance without explicit modality-specific routing rules. Figure 8 shows more concentrated expert activation for visual generation and audio tasks than for text, while Figure 9 shows stronger overlap within image/video understanding and within image/video generation than between understanding and generation.

  • Perspective of Expert Utilization and Specialization: specialization emerges despite shared routing, but activation frequencies identify usage patterns rather than independently establishing expert functions.
  • Perspective of Cross-Modality Expert Collaboration: IoU among the top 25% most frequently activated experts rises between text and audio from 0.12 in the first layer to 0.40 in the last. Image/video understanding overlap is 0.98, 0.52, and 0.64 in the representative first, middle, and last layers.
  • Perspective of Load Balancing: Figure 10 uses normalized entropy, NE = −Σᵢ₌₁ᴺ pᵢ log(pᵢ) / log N, where N is the expert count and pᵢ is the routing fraction. Text remains relatively uniform, visual generation is more concentrated, and visual understanding is better balanced in intermediate layers; these observations do not isolate a causal advantage over modality-partitioned routing.

Text activation is more diffuse, while visual generation and audio show sharper peaks in representative layers. Modality-agnostic routing thus permits task-dependent concentration rather than requiring uniform expert sharing.

Expert activation distributions by task in representative first, middle, and last layers.
Expert activation distributions by task in representative first, middle, and last layers.

Top-25% expert-set overlap is high within image/video understanding and within image/video generation, but lower between these task families. Text–audio overlap increases with depth; overlap alone does not establish semantic integration or a causal performance benefit.

Pairwise IoU of frequently activated expert sets across modalities and tasks.
Pairwise IoU of frequently activated expert sets across modalities and tasks.

Text retains high routing entropy across depth, while visual generation is substantially more concentrated. Visual understanding is relatively balanced in intermediate layers, showing why aggregate load balance can conceal task-specific differences.

Normalized expert-routing entropy across layers for text, visual, and audio tasks.
Normalized expert-routing entropy across layers for text, visual, and audio tasks.

6.4.2 Elastic Training: Controlled-Scale Experiments uses a MoE with 64 experts, 454M activated parameters, 3.2B total parameters, and default top-k = 8, trained on 250B tokens. Tables 9–11 separately test elastic depth, width, and sparsity using held-out validation loss.

  • Table 9: the 16-layer baseline loss is 1.945; elastic depth gives 1.941 at 16 layers and 2.137 at 12 layers. The controlled experiment specifies an 80% full-depth / 20% reduced-depth mixture, differing from the 75% / 25% recipe in Section 3.3 despite describing the configurations as consistent.
  • Table 10: the 64-expert baseline loss is 1.957; elastic width gives 1.964 with 64 experts and 2.218 with 32 experts.
  • Table 11: the baseline top-k = 8 loss is 1.945; elastic sparsity gives 1.969, 1.971, 2.003, and 2.175 at top-k = 8, 4, 2, and 1. The tables quantify reduction costs but do not compare reduced elastic configurations against equally reduced non-elastic checkpoints.

Elastic depth changes full-depth loss from the baseline's 1.945 to 1.941, while using 12 layers raises it to 2.137. The reduced-depth configuration incurs a measurable loss increase; the small full-depth improvement is not accompanied by uncertainty estimates.

Validation loss for full-depth and reduced-depth inference after elastic depth training.
Validation loss for full-depth and reduced-depth inference after elastic depth training.

The full 64-expert elastic configuration remains close to baseline loss, 1.964 versus 1.957, but restricting it to 32 experts increases loss to 2.218. This measures the capacity cost of reducing the stored expert pool.

Validation loss for 64-expert and 32-expert inference after elastic width training.
Validation loss for 64-expert and 32-expert inference after elastic width training.

Scaling to ERNIE 5.0 evaluates the experimental ERNIE 5.0-Exp family rather than the main benchmark model. Table 12 reports an average of 75.55 for ERNIE 5.0-Exp and 74.43 for ERNIE 5.0-Exp-ES25.0%, whose routing top-k is reduced to 25% with a reported decoding speed improvement of more than 15%.

  • ERNIE 5.0-Exp-EA35.8% uses 53.7% of activated parameters and 35.8% of total parameters. It scores 75.17 on the seven-benchmark average, including ZebraLogic 95.20 and VisualPuzzle 60.39 versus 95.00 and 59.93 for the full experimental model.
  • The jointly elastic compact model is extracted from ERNIE 5.0-Exp-Base, then receives mid-training and post-training using the same data and strategy as the full model. Its results are not a pure inference-time reduction without further training.
  • The large-scale comparison covers selected text and visual-understanding tasks, not audio or image/video generation. The decoding-speed claim lacks a hardware and workload breakdown.

The jointly elastic variant nearly matches the full experimental model's seven-task average, while the inference-only sparsity variant loses more accuracy. The compact variant also receives mid-training and post-training, so the rows differ in adaptation procedure as well as inference configuration.

Selected text and visual benchmark results for ERNIE 5.0-Exp and its elastic variants.
Selected text and visual benchmark results for ERNIE 5.0-Exp and its elastic variants.

7 Conclusion

The conclusion identifies native multimodal autoregressive modeling, shared ultra-sparse routing, elastic pre-training, and scalable RL as the central contributions. The results demonstrate broad multimodal capability and deployment trade-offs; the claim of the first unified trillion-level realization remains the authors’ qualified priority claim rather than an independently established finding.

8 Contributors

The report is authored by ERNIE Team, Baidu, and lists individual contributors across pages 27–28, beginning with Haifeng Wang, Hua Wu, Tian Wu, Yu Sun, Jing Liu, Dianhai Yu, and Yanjun Ma. The list does not assign individual technical roles or provide additional institutional affiliations.

Appendix

  • The supplied 36-page PDF contains no separate appendix. Pages 29–36 contain References after the contributor list, with no additional appendix experiments or implementation details.

Brief Thoughts

The elastic-training evidence is particularly concrete: controlled validation-loss measurements and a large-scale experimental comparison expose the costs of reducing layers, expert capacity, and routing top-k. The unified architecture is also technically substantive, but its claimed benefits for avoiding modality interference and improving capability through shared routing are not isolated by matched architectural controls. “Unified” most directly describes the shared autoregressive backbone and prediction framework, not the elimination of modality-specific interfaces or diffusion from the complete system. The tables support factual QA, multi-instruction following, selected perception tasks, and VBench semantic alignment as strengths, while difficult reasoning, several VoiceBench tasks, and some video-understanding evaluations remain weaker. Reproducibility is limited by the undisclosed production architecture, corpus mixture, total training compute, incomplete evaluation protocols, and missing standalone ablations for several RL and systems techniques.