Gorio Tech Blog search

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing | Summary

|

Contents

This article explains the key points of AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing.

  • 2026-09-08 (arXiv)
  • Ma, Ziyang, Niu, Zhikang, Tu, Wenming, Wang, Tianrui, Yan, Ruiqi, Liu, Junxi, Huo, Yanru, Huang, Nickk, Liu, Yang, Xie, Qicong, et al.
  • Paper
  • Github
  • Project Page

Read this article in Korean


Summary

  • AuK is an open-source foundational model that maps a natural-language instruction and optional audio context to a target waveform, unifying zero-shot and instruction-controlled TTS, acoustic and paralinguistic editing, content editing, and enhancement or separation. It is trained on approximately 3.03 billion instruction–audio instances providing 1.95 million hours of effective supervision; AuK-Flash uses four CFG-free sampling steps and is reported to provide a 4.5× wall-clock speedup under matched conditions.

This overview compares AuK and AuK-Flash with task-specific and unified baselines across generation, editing, enhancement, and separation. It illustrates the reported pattern that the full model emphasizes recognition and instruction accuracy, whereas Flash uses substantially fewer sampling steps and often has strong perceptual-quality estimates.

Bar-chart comparison of speech generation, instruction-guided editing, enhancement, and separation metrics across AuK variants and baseline systems.
Bar-chart comparison of speech generation, instruction-guided editing, enhancement, and separation metrics across AuK variants and baseline systems.

The diagram organizes the system around five capability families using a shared instruction-plus-audio interface. It depicts outputs that may be synthesized from text, locally edited, transformed in vocal delivery, or extracted and restored from mixtures.

Illustration of instruction-based and zero-shot TTS, acoustic and paralinguistic editing, content editing, and speech or music enhancement and separation.
Illustration of instruction-based and zero-shot TTS, acoustic and paralinguistic editing, content editing, and speech or music enhancement and separation.

2 Data Construction

The pre-training corpus represents every example as an instruction, optional input audio, and target waveform across five task families: speech generation, acoustic editing, paralinguistic editing, content editing, and enhancement and separation. This shared interface supplies supervision for synthesizing new speech, preserving untargeted attributes during edits, and selecting requested components from acoustic scenes.

The corpus map connects the five training families to representative operations, including generation, localized content edits, vocal-style edits, and source selection. Its sector widths are explicitly non-quantitative and therefore communicate task coverage rather than data proportions.

Overview of the five-family pre-training corpus and representative instruction-level operations.
Overview of the five-family pre-training corpus and representative instruction-level operations.

2.1 Speech Generation

Speech generation combines transcript-free zero-shot TTS with descriptive Instruct TTS. For each speaker, different utterances are paired bidirectionally so that prompt audio and target text are available without a prompt transcript; Instruct TTS uses Qwen3-Omni-generated captions and structured voice attributes without reference audio.

2.2 Acoustic Editing

Acoustic-editing supervision uses deterministic transformations intended to preserve content, speaker identity, and non-target attributes. Targets use speaking-rate multipliers of 0.5×, 0.75×, 1.25×, 1.5×, and 2.0×, loudness offsets of ±5, ±10, and ±15 dB, and pitch shifts of ±1, ±2, and ±3 semitones.

2.3 Paralinguistic Editing

Paralinguistic editing covers emotion, timbre, de-accenting, nonverbal vocalizations, and whisper-style conversion. The paired targets are constructed with existing speech systems, aligned recordings, or filtered reconstructions to alter delivery while retaining linguistic content and, where applicable, speaker characteristics.

2.4 Content Editing

For speech content editing, an LLM creates insertion, deletion, and substitution transcripts, while word-level alignments define localized masked regions for infilling. Lyric editing preserves Chinese character counts or English word counts, then conditions vocal synthesis on complete target lyrics and the original melody.

2.5 Enhancement and Separation

Enhancement and separation data include noise, reverberation, channel degradations, conversational mixtures, and music mixtures. Selective targets may remove only the component named by an instruction rather than always producing clean speech, requiring preservation of unrelated scene attributes.

3 Model Design

AuK comprises a Qwen2.5-Omni semantic encoder, an audio VAE, and a FLUX-style rectified-flow Transformer. The semantic encoder processes text alone or text jointly with reference audio; the VAE supplies reference latents and reconstructs waveforms; and the backbone predicts target-latent velocity.

3.1 Overall Architecture

The backbone begins with dual-stream MMDiT blocks that exchange information between semantic and acoustic streams while retaining distinct residual pathways, then concatenates the streams for single-stream DiT refinement. Reference-conditioned tasks provide the same audio to the MLLM audio encoder and VAE, whereas text-only tasks omit the reference branch.

The architecture traces how instructions and optional reference audio create semantic and acoustic conditions, which interact in dual-stream MMDiT layers before fused DiT processing and VAE decoding. It distinguishes frozen MLLM and VAE modules from the generative backbone.

AuK architecture with multimodal semantic conditioning, VAE acoustic conditioning, dual-stream MMDiT fusion, and single-stream DiT generation.
AuK architecture with multimodal semantic conditioning, VAE acoustic conditioning, dual-stream MMDiT fusion, and single-stream DiT generation.

4 Model Training

AuK-VAE encodes 24 kHz audio into 64-dimensional latents at 50 Hz and uses a causal decoder for waveform reconstruction. It is trained for 1.24 million updates on approximately 3 million hours of speech, music, and general audio with multi-scale log-mel reconstruction, adversarial and feature-matching losses, and KL regularization.

4.2 Unified Pre-Training

Unified pre-training freezes the MLLM and VAE, trains the approximately 1.5 billion-parameter Transformer for 50k generation-only warm-up updates, then jointly trains all task families for 600k updates. It uses masked rectified-flow MSE, hierarchical condition dropout, cropped zero-shot references, and a fixed mixture led by speech generation (28.10%), enhancement and separation (23.17%), and content editing (23.02%).

The table gives the fixed sampling mixture used in the second pre-training stage. Acoustic editing has the smallest probability at 3.96%, while each remaining family accounts for roughly one-fifth to just over one-quarter of batches.

Fixed per-batch sampling probabilities for the five task families during joint pre-training.
Fixed per-batch sampling probabilities for the five task families during joint pre-training.

4.3 Post-Training

Post-training first applies flow-based preference optimization to 818 informative human-rated editing groups comprising 9,080 candidates, using ordinal listwise feedback on completion quality. It then applies Flow-GRPO to generation: zero-shot TTS rewards content correctness and speaker similarity, while instruction TTS uses a style-consistency reward model trained from 30,000 balanced positive and negative samples.

5 Model Acceleration

AuK-Flash is distilled from the post-trained teacher through consistency initialization and task-routed Decoupled DMD. After observing separation regressions under DMD, the authors exclude separation examples from DMD and instead use supervised clean-latent regression; the student uses four function evaluations without classifier-free guidance, versus 32 evaluations and CFG scale 2.0 for AuK.

The configuration shows that AuK-Flash retains the same latent representation, decoder, prompt enhancement, and audio preparation as AuK, while reducing Euler evaluations from 32 to 4 and removing CFG. The 4.5× speedup is reported under matched hardware, output duration, and batch size.

Inference configurations for the full AuK model and the task-routed distilled AuK-Flash variant.
Inference configurations for the full AuK model and the task-routed distilled AuK-Flash variant.

7 Performance

Inference can use a Prompt Enhancer that routes a free-form request, rewrites it toward task-specific training formats, validates parameters, prepares input audio, and estimates duration. The report states that arbitrary free-form editing instructions are not yet handled robustly without this task routing and prompt rewriting; canonical supported instructions can bypass the enhancer.

This aggregate table supplies numerical comparisons across heterogeneous native metrics. Neither variant dominates every criterion: AuK commonly improves recognition error and edit fidelity, while Flash frequently obtains higher UTMOS scores or the best MMAE-Speech exact-match result.

Selected generation, general-editing, enhancement, separation, and super-resolution benchmark results.
Selected generation, general-editing, enhancement, separation, and super-resolution benchmark results.

The VAE comparison reports the best displayed PESQ, STOI, mel-distance, and STFT-distance values for AuK-VAE across the listed speech, general-audio, and music evaluations. These reconstruction results support latent-space fidelity but do not independently establish downstream editing quality.

Reconstruction quality of AuK-VAE and comparison VAEs on speech, general audio, and music.
Reconstruction quality of AuK-VAE and comparison VAEs on speech, general audio, and music.

7.2 Generation Ability

On Seed-TTS-Eval, AuK reports the lowest displayed average recognition error of 2.65% and the highest speaker similarity of 0.795; AuK-Flash reports 2.85% and 0.790. On InstructTTSEval, AuK reaches the highest displayed Chinese DSD accuracy of 83.37%, while AuK-Flash ties the best displayed English DSD score of 82.40%.

7.3 Editing Ability

On MMAE-Speech, AuK achieves the highest displayed instruction-following rate (48.23%) and consistency rate (88.11%), while Flash has the highest exact-match rate (13.85%). AuK reduces Full-split semantic-editing WER relative to Ming-UniAudio from 10.46% to 3.09% in Chinese and from 14.28% to 3.96% in English; on restoration benchmarks, the full model often yields stronger recognition preservation while Flash often attains higher predicted perceptual-quality scores.

Appendix

  • The appendix provides expanded benchmark tables for zero-shot and instruction-following TTS, operation-level semantic and acoustic editing, enhancement, separation, and super-resolution. It specifies that MMAE-Speech uses the Prompt Enhancer and that generation results are means over three runs using benchmark-specific templates and duration estimation.

Brief Thoughts

The report combines broad instruction-conditioned speech tasks with preference optimization, generation reinforcement learning, and task-aware distillation. Its evidence is benchmark-specific, while unrestricted free-form use remains qualified by dependence on the Prompt Enhancer and by several data pipelines that use outputs from other speech models.