AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing | Summary
08 Sep 2026 | Paper Review Speech Generation Speech Editing Flow Matching Reinforcement LearningContents
- Summary
- 2 Data Construction
- 3 Model Design
- 4 Model Training
- 5 Model Acceleration
- 7 Performance
- Appendix
- Brief Thoughts
This article explains the key points of AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing.
- 2026-09-08 (arXiv)
- Ma, Ziyang, Niu, Zhikang, Tu, Wenming, Wang, Tianrui, Yan, Ruiqi, Liu, Junxi, Huo, Yanru, Huang, Nickk, Liu, Yang, Xie, Qicong, et al.
- Paper
- Github
- Project Page
Summary
- AuK is an open-source foundational model that maps a natural-language instruction and optional audio context to a target waveform, unifying zero-shot and instruction-controlled TTS, acoustic and paralinguistic editing, content editing, and enhancement or separation. It is trained on approximately 3.03 billion instruction–audio instances providing 1.95 million hours of effective supervision; AuK-Flash uses four CFG-free sampling steps and is reported to provide a 4.5× wall-clock speedup under matched conditions.
This overview compares AuK and AuK-Flash with task-specific and unified baselines across generation, editing, enhancement, and separation. It illustrates the reported pattern that the full model emphasizes recognition and instruction accuracy, whereas Flash uses substantially fewer sampling steps and often has strong perceptual-quality estimates.
The diagram organizes the system around five capability families using a shared instruction-plus-audio interface. It depicts outputs that may be synthesized from text, locally edited, transformed in vocal delivery, or extracted and restored from mixtures.
2 Data Construction
The pre-training corpus represents every example as an instruction, optional input audio, and target waveform across five task families: speech generation, acoustic editing, paralinguistic editing, content editing, and enhancement and separation. This shared interface supplies supervision for synthesizing new speech, preserving untargeted attributes during edits, and selecting requested components from acoustic scenes.
The corpus map connects the five training families to representative operations, including generation, localized content edits, vocal-style edits, and source selection. Its sector widths are explicitly non-quantitative and therefore communicate task coverage rather than data proportions.
2.1 Speech Generation
Speech generation combines transcript-free zero-shot TTS with descriptive Instruct TTS. For each speaker, different utterances are paired bidirectionally so that prompt audio and target text are available without a prompt transcript; Instruct TTS uses Qwen3-Omni-generated captions and structured voice attributes without reference audio.
2.2 Acoustic Editing
Acoustic-editing supervision uses deterministic transformations intended to preserve content, speaker identity, and non-target attributes. Targets use speaking-rate multipliers of 0.5×, 0.75×, 1.25×, 1.5×, and 2.0×, loudness offsets of ±5, ±10, and ±15 dB, and pitch shifts of ±1, ±2, and ±3 semitones.
2.3 Paralinguistic Editing
Paralinguistic editing covers emotion, timbre, de-accenting, nonverbal vocalizations, and whisper-style conversion. The paired targets are constructed with existing speech systems, aligned recordings, or filtered reconstructions to alter delivery while retaining linguistic content and, where applicable, speaker characteristics.
2.4 Content Editing
For speech content editing, an LLM creates insertion, deletion, and substitution transcripts, while word-level alignments define localized masked regions for infilling. Lyric editing preserves Chinese character counts or English word counts, then conditions vocal synthesis on complete target lyrics and the original melody.
2.5 Enhancement and Separation
Enhancement and separation data include noise, reverberation, channel degradations, conversational mixtures, and music mixtures. Selective targets may remove only the component named by an instruction rather than always producing clean speech, requiring preservation of unrelated scene attributes.
3 Model Design
AuK comprises a Qwen2.5-Omni semantic encoder, an audio VAE, and a FLUX-style rectified-flow Transformer. The semantic encoder processes text alone or text jointly with reference audio; the VAE supplies reference latents and reconstructs waveforms; and the backbone predicts target-latent velocity.
3.1 Overall Architecture
The backbone begins with dual-stream MMDiT blocks that exchange information between semantic and acoustic streams while retaining distinct residual pathways, then concatenates the streams for single-stream DiT refinement. Reference-conditioned tasks provide the same audio to the MLLM audio encoder and VAE, whereas text-only tasks omit the reference branch.
The architecture traces how instructions and optional reference audio create semantic and acoustic conditions, which interact in dual-stream MMDiT layers before fused DiT processing and VAE decoding. It distinguishes frozen MLLM and VAE modules from the generative backbone.
4 Model Training
AuK-VAE encodes 24 kHz audio into 64-dimensional latents at 50 Hz and uses a causal decoder for waveform reconstruction. It is trained for 1.24 million updates on approximately 3 million hours of speech, music, and general audio with multi-scale log-mel reconstruction, adversarial and feature-matching losses, and KL regularization.
4.2 Unified Pre-Training
Unified pre-training freezes the MLLM and VAE, trains the approximately 1.5 billion-parameter Transformer for 50k generation-only warm-up updates, then jointly trains all task families for 600k updates. It uses masked rectified-flow MSE, hierarchical condition dropout, cropped zero-shot references, and a fixed mixture led by speech generation (28.10%), enhancement and separation (23.17%), and content editing (23.02%).
The table gives the fixed sampling mixture used in the second pre-training stage. Acoustic editing has the smallest probability at 3.96%, while each remaining family accounts for roughly one-fifth to just over one-quarter of batches.
4.3 Post-Training
Post-training first applies flow-based preference optimization to 818 informative human-rated editing groups comprising 9,080 candidates, using ordinal listwise feedback on completion quality. It then applies Flow-GRPO to generation: zero-shot TTS rewards content correctness and speaker similarity, while instruction TTS uses a style-consistency reward model trained from 30,000 balanced positive and negative samples.
5 Model Acceleration
AuK-Flash is distilled from the post-trained teacher through consistency initialization and task-routed Decoupled DMD. After observing separation regressions under DMD, the authors exclude separation examples from DMD and instead use supervised clean-latent regression; the student uses four function evaluations without classifier-free guidance, versus 32 evaluations and CFG scale 2.0 for AuK.
The configuration shows that AuK-Flash retains the same latent representation, decoder, prompt enhancement, and audio preparation as AuK, while reducing Euler evaluations from 32 to 4 and removing CFG. The 4.5× speedup is reported under matched hardware, output duration, and batch size.
7 Performance
Inference can use a Prompt Enhancer that routes a free-form request, rewrites it toward task-specific training formats, validates parameters, prepares input audio, and estimates duration. The report states that arbitrary free-form editing instructions are not yet handled robustly without this task routing and prompt rewriting; canonical supported instructions can bypass the enhancer.
This aggregate table supplies numerical comparisons across heterogeneous native metrics. Neither variant dominates every criterion: AuK commonly improves recognition error and edit fidelity, while Flash frequently obtains higher UTMOS scores or the best MMAE-Speech exact-match result.
The VAE comparison reports the best displayed PESQ, STOI, mel-distance, and STFT-distance values for AuK-VAE across the listed speech, general-audio, and music evaluations. These reconstruction results support latent-space fidelity but do not independently establish downstream editing quality.
7.2 Generation Ability
On Seed-TTS-Eval, AuK reports the lowest displayed average recognition error of 2.65% and the highest speaker similarity of 0.795; AuK-Flash reports 2.85% and 0.790. On InstructTTSEval, AuK reaches the highest displayed Chinese DSD accuracy of 83.37%, while AuK-Flash ties the best displayed English DSD score of 82.40%.
7.3 Editing Ability
On MMAE-Speech, AuK achieves the highest displayed instruction-following rate (48.23%) and consistency rate (88.11%), while Flash has the highest exact-match rate (13.85%). AuK reduces Full-split semantic-editing WER relative to Ming-UniAudio from 10.46% to 3.09% in Chinese and from 14.28% to 3.96% in English; on restoration benchmarks, the full model often yields stronger recognition preservation while Flash often attains higher predicted perceptual-quality scores.
Appendix
- The appendix provides expanded benchmark tables for zero-shot and instruction-following TTS, operation-level semantic and acoustic editing, enhancement, separation, and super-resolution. It specifies that MMAE-Speech uses the Prompt Enhancer and that generation results are means over three runs using benchmark-specific templates and duration estimation.
Brief Thoughts
The report combines broad instruction-conditioned speech tasks with preference optimization, generation reinforcement learning, and task-aware distillation. Its evidence is benchmark-specific, while unrestricted free-form use remains qualified by dependence on the Prompt Enhancer and by several data pipelines that use outputs from other speech models.