DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation | Summary
06 Jul 2026 | Paper Review Speculative Decoding Semi-Autoregressive Generation Hardware-Aware Scheduling Confidence CalibrationContents
- Summary
- 3. Architecture
- 4.1. Experimental Setup
- 4.2. Experimental Results
- 4.3. Experimental Analysis
- 5. Real-World Deployment of DSpark
- Appendix
- Brief Thoughts
This article explains the key points of DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation.
- 2026-07-06 (arXiv)
- Cheng, Xin, Yu, Xingkai, Shao, Chenze, Li, Jiashi, Xiong, Yunfan, Qian, Yi, Zhu, Jiaqi, Ma, Shirong, Zhang, Xiaokang, Ye, Jiasheng, et al.
- Peking University, DeepSeek-AI
- Paper
Summary
- Speculative decoding proposes a token block with a draft model and verifies an accepted prefix with a target model through rejection sampling, preserving the target distribution. Autoregressive drafters incur sequential drafting latency; parallel drafters can propose longer blocks in one pass but lack dependencies among sampled tokens, while verifying low-value suffixes consumes batch capacity under load.
- DSpark pairs a parallel draft backbone with a lightweight sequential correction that supplies locally normalized per-token probabilities. A confidence head estimates conditional acceptance, and a scheduler uses calibrated prefix-survival estimates and a profiled engine step-rate curve to choose verification lengths; unlike globally normalized or latent-output alternatives discussed in the related work, the local correction permits standard rejection sampling.
- With confidence scheduling disabled in offline tests, DSpark improves macro-average accepted length over Eagle3 by 30.9%, 26.7%, and 30.0% on Qwen3-4B, 8B, and 14B, respectively. Under measured live DeepSeek-V4 traffic, per-user generation speed exceeds the former MTP-1 production baseline by 60%–85% on V4-Flash and 57%–78% on V4-Pro at matched aggregate throughput; these are configuration- and traffic-specific comparisons.
3. Architecture
DSpark targets per-generated-token latency L = (T_draft + T_verify)/τ, where τ is accepted length per decoding round, including the target-generated bonus token. The parallel backbone limits drafting passes, the sequential head addresses suffix acceptance, and confidence-scheduled verification avoids spending target capacity on selected low-survival suffixes.
3.1. Semi-Autoregressive Generation
DSpark modifies the DFlash backbone so γ inputs—an anchor and γ−1 masks—produce γ draft positions in one parallel pass. At position k, a lightweight sequential stage adds prefix-dependent bias B_k to backbone logits U_k and locally normalizes the distribution; left-to-right sampling supplies the conditional probabilities needed for rejection sampling.
- The default Markov head factors its transition matrix as B = W₁W₂, with rank r = 256 by default, and conditions on the immediately preceding token.
- An alternative RNN head maintains state across the block and incorporates the backbone hidden state; the Markov head is the default.
The correction addresses multi-modal collisions: independently chosen positions can combine fragments of different plausible continuations, whereas a sampled predecessor can favor a compatible successor. Its serial computation must remain small relative to the parallel pass for drafting latency to remain low.
3.2.1. Confidence Head
The confidence head predicts c_k from backbone state h_k and the preceding token’s Markov embedding. It targets the conditional chance of accepting position k after earlier positions have been accepted, using the analytical per-step acceptance rate c*_k = 1 − ½‖p^d_k − p^t_k‖₁ as a soft label.
- Sequential Temperature Scaling (STS) calibrates cumulative prefix survival on held-out data. A left-to-right one-dimensional grid search minimizes Expected Calibration Error (ECE) at each position while keeping earlier calibrated positions fixed.
3.2.2. Hardware-Aware Prefix Scheduler
For request r and position j, estimated prefix survival is a_r,j = ∏_{i≤j} c_r,i. For R requests with verification lengths ℓ_r, the scheduler estimates batch size B = Σ_r(1 + ℓ_r), accepted tokens τ = Σ_r(1 + Σ_{j=1}^{ℓ_r} a_r,j), and throughput Θ = τ · SPS(B), using an engine step-rate curve profiled at initialization; this model assumes step rate depends predominantly on verification batch size rather than context-length variation.
- Algorithm 1 ranks prefix extensions by survival probability. Non-increasing survival along each request makes this ordering prefix-respecting and optimal for expected accepts at a fixed batch size.
- The proposed causal search stops when estimated throughput fails to improve, preventing a retrospective choice about an earlier token from using its sampled value through successor confidence. Its stepwise stopping finds the global throughput maximum when Θ is unimodal along the admission path.
3.3. Training
Training samples multiple anchors per target sequence and freezes the target, shared embedding, and language-modeling head. The updated backbone, sequential module, and confidence head use position-weighted cross-entropy, draft–target total-variation matching, and binary cross-entropy against soft acceptance labels, with w_k = exp(−(k−1)/γ) and default coefficients α_ce = 0.1, α_tv = 0.9, α_conf = 1.0.
4.1. Experimental Setup
Offline evaluation uses Qwen3-{4B, 8B, 14B} and Gemma4-12B targets. Eagle3, DFlash, and DSpark are retrained in one framework using target-regenerated responses to prompts from the 1.3-million-sample Open-PerfectBlend dataset; the comparison aligns Eagle3’s TTT horizon and the DFlash/DSpark block size at 7, uses 1 versus 5 drafter layers, and evaluates non-thinking generation at sampling temperature 1.0 on math, code, and chat benchmarks, including AIME25 (Zhang and Math-AI, 2025).
- Reported accepted length includes the target-generated bonus token. All compared drafters use chain-based drafting.
4.2. Experimental Results
The offline comparison disables confidence scheduling to isolate draft quality. Table 1 reports the longest accepted length for DSpark in every listed target–benchmark combination: on Qwen3-4B, GSM8K gives 6.11 for DSpark, 5.40 for DFlash, and 5.14 for Eagle3; MT-Bench gives 3.64, 3.07, and 2.39.
- Relative to DFlash, DSpark’s Qwen3-4B, 8B, and 14B macro-average accepted-length gains are 16.3%, 18.4%, and 18.3%. Gemma4-12B also shows gains, but accepted length alone does not measure serving speed.
Every DSpark entry exceeds the corresponding Eagle3 and DFlash accepted length across the four targets and nine benchmarks. Scheduling is disabled in this comparison, so the table does not quantify production throughput.
4.3. Experimental Analysis
The analyses examine conditional acceptance by draft position, drafter depth and proposal length, serial-loop latency, and confidence-based pruning and calibration. These diagnostics distinguish draft quality from the throughput effects of scheduling, which are evaluated in deployment.
4.3.1. Why Can Parallel Generation Outperform Autoregression?
Figure 2 counts acceptance at position k only for rollouts whose earlier draft tokens were accepted. On Qwen3-4B, DFlash exceeds Eagle3 initially—about 0.88 versus 0.81 on math and 0.72 versus 0.53 on chat—but its later conditional acceptance declines; Eagle3’s chat acceptance rises from about 0.53 to 0.74.
- DSpark combines high initial conditional acceptance with less suffix decay in these measurements. The curves support the proposed account that the deeper parallel backbone benefits early positions while sequential correction helps later ones; they do not isolate backbone depth as the sole cause.
Conditional acceptance removes rollouts rejected before each plotted position. DFlash begins strongly but loses acceptance toward the suffix, Eagle3 is steadier or improves, and DSpark remains comparatively high across the block.
4.3.2. A Little Autoregression Goes a Long Way
At block size 7 on Qwen3-4B, accepted length rises as DSpark’s backbone grows from 1 to 5 layers, and the 2-layer variant exceeds 5-layer DFlash on the plotted benchmarks. With a 5-layer backbone, DSpark’s reported accepted-length gains over DFlash grow from 16%, 15%, and 18% on math, code, and chat at γ = 7 to 30%, 26%, and 22% at γ = 15.
- The RNN head offers marginal additional gains, chiefly for longer proposals. At batch size 128, with latency averaged over 512, 1024, 2048, and 4096 context tokens, the sequential head adds 0.2%–1.3% to full-round latency as proposal length increases from 4 to 16.
The math, code, and chat panels show increasing separation between DSpark and DFlash at longer proposals; the latency panel measures a small full-round overhead at batch size 128. The RNN head adds limited accepted-length improvement over the Markov head.
4.3.3. Verify Smarter, Not Longer: The Role of Confidence Head
A Qwen3-4B static-threshold sweep tests the confidence head separately from the production scheduler. Higher thresholds prune more rejected tokens and raise reported acceptance rates from 45.7% to 95.7% on chat, 76.9% to 92.5% on math, and 67.6% to 92.0% on code; they also shorten verified prefixes, so these rates do not establish throughput optimality.
- On Alpaca, Figure 6 reports ROC-AUC of 0.818, 0.812, 0.864, and 0.907 at positions 1, 3, 5, and 7. STS reduces the respective ECE values from 5.7%, 8.2%, 5.8%, and 3.3% to 2.0%, 1.7%, 0.8%, and 0.4%.
Increasing the threshold reduces the hashed rejected-token portions, particularly on chat, and increases the accepted fraction among verified tokens. The accompanying reduction in verified tokens means this sweep tests pruning behavior, not optimal serving throughput.
At Alpaca positions 1, 3, 5, and 7, post-calibration reliability curves approach the ideal diagonal more closely than raw predictions, with lower reported ECE. The ROC-AUC annotations measure discrimination separately from calibration.
5. Real-World Deployment of DSpark
DSpark is deployed with DeepSeek-V4-Flash (preview) and DeepSeek-V4-Pro (preview), where verification lengths must adapt to load and fit an asynchronous serving pipeline. The live-traffic evaluation compares the complete implementation against MTP-1, the former single-token production configuration, rather than isolating the effect of draft architecture.
5.1. Scalable and Flexible Training
The deployed parallel backbone has three MoE layers with mHC and sliding window attention of 128; its maximum draft block size is γ = 5, with a Markov head and subsequent STS calibration. To reduce training communication, workers transfer target hidden states instead of full-vocabulary logits and apply the language-modeling head locally at sampled positions; anchor-bounded sequence packing reduces draft-side work associated with full-document context.
- The paper specifies a target vocabulary of approximately 10⁵ and per-token hidden-state communication complexity of O(d), where d is hidden dimension.
5.2. Hardware-Aware Prefix Scheduler in Practice
The production scheduler adapts Algorithm 1 to jagged, discrete SPS(B) curves and Zero-Overhead Scheduling (ZOS). It uses confidence information from two steps earlier to set the upcoming capacity limit K while ranking current candidate tokens by current cumulative confidence, enabling dynamic top-K selection without waiting to schedule the GPU pipeline.
- The implementation omits the early-stopping break when searching historical capacity predictions to avoid local optima at hardware throughput cliffs. The paper argues that basing this capacity choice on historical information, rather than the current sampled token, retains the non-anticipating condition needed for exact target-distribution recovery.
5.3. High-Throughput and Low-Latency Inference
In the reported deployment, KV-cache limits or available traffic often keep effective batches below GPU compute saturation, allowing aggregate throughput and per-user speed to improve together. For variable verification lengths, the implementation flattens tokens across requests and passes sequence boundaries through a marker tensor in sparse attention, modifying the DeepSeek-V4 index-attention and compress kernels to avoid extensive padding.
5.4. Performance under Live User Traffic
Under measured live traffic, DSpark-5 improves the throughput–interactivity frontier relative to MTP-1. Aggregate throughput rises by 51% at 80 tok/s/user on V4-Flash and 52% at 35 tok/s/user on V4-Pro; at matched aggregate throughput, per-user speed rises by 60%–85% and 57%–78%, respectively.
- The nominal 661% Flash gain at 120 tok/s/user and 406% Pro gain at 50 tok/s/user occur where MTP-1 supports only low concurrency, so they are evidence of a wider feasible frontier rather than representative gains against a well-utilized baseline.
- Figure 8 shows a roughly 4–6-token per-request verification budget under moderate load, compared with MTP-1’s fixed 2 tokens, and a declining DSpark budget as concurrency grows. Difficult requests still incur the cost of generating an initial full draft block even when acceptance is low.
Live-traffic telemetry and fitted curves show higher aggregate output throughput for DSpark at comparable per-user speeds on both engines. The largest percentage gaps occur near MTP-1’s low-concurrency boundary; matched-throughput comparisons describe per-user speed gains at more practical operating points.
The top panels compare aggregate throughput as concurrency changes; the bottom panels show DSpark’s per-request verification budget falling with load while MTP-1 remains at 2 tokens. The plots connect the measured throughput advantage with load-dependent verification allocation.
Appendix
- Appendix A gives a counterexample to retrospective scheduling without early stopping. With one request, γ = 2, a₁ = 0.8, and SPS(1), SPS(2), SPS(3) equal to 1.0, 0.5, 0.45, the first two throughput estimates are Θ₀ = 1.0 and Θ₁ = 0.9; using the sampled first token to evaluate its successor makes c₂ = 0.9 select length 2 but c₂ = 0 select length 0.
- For first-position draft probabilities (0.5, 0.5) and target probabilities (0.7, 0.3) over {A, B}, the appendix’s retrospective rule produces output probabilities (0.85, 0.15), not the target distribution. Early stopping instead halts when Θ₁ falls below Θ₀, before evaluating continuation-dependent confidence.
Brief Thoughts
The scheduler-disabled comparisons establish longer accepted prefixes, while position-wise acceptance identifies where DSpark differs from the baselines. The production measurements establish gains for the reported DeepSeek-V4 engines and traffic against MTP-1, but do not compare the complete system with other adaptively scheduled multi-token production baselines; low-acceptance requests also retain the upfront draft-block cost.