Gorio Tech Blog search

Qwen3-ASR Technical Report | Summary

|

Contents

This article explains the key points of Qwen3-ASR Technical Report.

  • 2026-01-29 (arXiv)
  • Shi, Xian, Wang, Xiong, Guo, Zhifang, Wang, Yongqi, Zhang, Pei, Zhang, Xinyu, Guo, Zishan, Hao, Hongkun, Xi, Yu, Yang, Baosong, et al.
  • Paper
  • Github

Read this article in Korean


Summary

  • Qwen3-ASR Technical Report introduces Qwen3-ASR-1.7B and Qwen3-ASR-0.6B for transcription and language identification across 30 languages and 22 Chinese dialects, alongside Qwen3-ForcedAligner-0.6B for timestamp alignment in 11 languages. The ASR models combine a pretrained AuT speech encoder with Qwen3 language models and use Qwen3-Omni as their foundation model.
  • The 1.7B model leads the evaluated systems on most public Mandarin and dialect benchmarks and every subset of the internal English–Chinese robustness suite. Its multilingual advantage is not universal: on the full 30-language Fleurs evaluation, its average error is 12.60 versus 8.16 for Whisper-large-v3.
  • The forced aligner predicts all requested timestamp slots non-autoregressively, using 80ms timestamp classes rather than next-token generation; its reported human-labeled average timestamp shift is 32.4ms. Qwen3-ASR-0.6B reaches 92ms average time-to-first-token at concurrency 1 and 2,000 audio seconds per second at concurrency 128, which are distinct operating points.

1 Introduction

The report motivates Large Audio-Language Models (LALMs) as a way to combine acoustic understanding with language modeling and world knowledge for noisy, multilingual, dialectal, and long-form transcription. It also treats timestamp prediction as a separate deployment requirement, proposing a multilingual LALM alternative to conventional CTC- or CIF-based post-processing with user-selected alignment granularities.

  • The proposed family includes two ASR models and one forced aligner, with weights released under the Apache 2.0 license and an open-source inference and fine-tuning framework.
  • Figure 1 illustrates the intended capabilities: noise robustness, singing and song transcription, multilingual recognition, streaming efficiency, and timestamp precision.
  • The authors supplement public benchmarks with internal evaluations because they argue that annotation errors limit discrimination on some established test sets.

The overview separates transcription capabilities from timestamp alignment. It illustrates the intended deployment scope rather than providing quantitative evidence for robustness, speed, or alignment accuracy.

Overview of multilingual transcription, noise robustness, singing and song recognition, streaming efficiency, and forced alignment in the Qwen3-ASR family.
Overview of multilingual transcription, noise robustness, singing and song recognition, streaming efficiency, and forced alignment in the Qwen3-ASR family.

2 Qwen3-ASR

Qwen3-ASR uses a speech-encoder–language-model architecture in two size configurations. Its training combines AuT and Omni pretraining with supervised and reinforcement-learning post-training, while the released models support both offline and streaming transcription.

2.1 Architecture

A separately pretrained AuT encoder converts 128-dimensional Fbank features into speech embeddings through 8 times downsampling, producing a 12.5Hz token rate. Dynamic FlashAttention windows ranging from 1s to 8s support short streaming chunks and longer offline inputs within the same architecture.

  • Qwen3-ASR-1.7B combines Qwen3-1.7B, a projector, and a 300M-parameter AuT encoder with hidden size 1024.
  • Qwen3-ASR-0.6B combines Qwen3-0.6B, a projector, and a 180M-parameter AuT encoder with hidden size 896.
  • Figure 2 distinguishes AuT’s attention-encoder-decoder pretraining from the released ASR pipeline, where encoded speech and optional entity context condition the Qwen3 LM.

The left panel shows AuT pretraining through an attention-encoder-decoder task. The right panel shows encoded speech and tokenized entity context conditioning Qwen3, clarifying how acoustic representations and contextual language modeling enter the transcription pipeline.

AuT pretraining architecture and the speech-encoder–language-model pipeline used by Qwen3-ASR.
AuT pretraining architecture and the speech-encoder–language-model pipeline used by Qwen3-ASR.

2.2 Training Strategies

Training proceeds through AuT pretraining, Omni pretraining, ASR supervised finetuning (SFT), and ASR reinforcement learning (RL). The first two stages follow Qwen3-Omni; post-training specializes the models for transcription, language identification, non-speech handling, streaming, and context biasing.

  • AuT pretraining uses approximately 40 million hours of pseudo-labeled ASR data, predominantly Chinese and English. Omni pretraining uses 3 trillion tokens for each model on multitask audio, vision, and text data.
  • SFT uses a substantially smaller multilingual dataset disjoint from the pretraining corpus, including non-speech, streaming-enhancement, and context-biasing examples. It trains ASR-only behavior that does not follow natural-language prompt instructions, with the stated aim of mitigating instruction injection, while allowing context tokens to supply background knowledge.
  • For recognizable speech, the example output is <|im_start|>assistant language English<asr_text>Today we release models including Qwen3-ASR-1.7B.<|im_end|>. For no detected speech, it is <|im_start|>assistant language None<asr_text><|im_end|>.
  • RL uses Group Sequence Policy Optimization (GSPO) on about 50k utterances: 35% Chinese and English data, 35% multilingual data, and 30% functional data. The authors attribute improved noise robustness and transcription stability to RL, but the report does not present a stage-specific ablation.

2.3 Features

Both ASR models support offline and streaming inference on individual audio inputs up to 1200s, or 20 minutes. Their stated coverage includes 30 languages, 22 Chinese dialects, ordinary speech, singing voice, and complete songs with background music.

  • Table 1 specifies the supported language and dialect inventories, inference modes, audio types, and maximum input durations.
  • Qwen3-ForcedAligner-0.6B has a different operating envelope: 11 languages, non-autoregressive inference, speech inputs, and a maximum single-input duration of 300s.
  • The aligner’s listed languages are Chinese, English, Cantonese, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish.

The feature comparison distinguishes the two ASR models' 1200s offline/streaming transcription limit from the aligner's 300s non-autoregressive speech-alignment limit. Explicit language and dialect inventories define the stated coverage, while the audio-type column distinguishes speech alignment from ASR on speech, singing, and songs with background music.

Supported languages, Chinese dialects, inference modes, maximum input durations, and audio types for the model family.
Supported languages, Chinese dialects, inference modes, maximum input durations, and audio types for the model family.

2.4 Inference Efficiency

ASR speed is measured using vLLM v0.14.0 with CUDA Graph enabled and bfloat16 precision, separately for offline batch generation and online asynchronous serving. Table 2 uses approximately 2-minute ASR inputs on a single computing resource, but the report does not identify the hardware model.

  • For Qwen3-ASR-0.6B, online concurrency 1 gives 92ms average TTFT, 105ms p95 TTFT, RTF 0.00940, and throughput 106.38 audio seconds per second.
  • At online concurrency 128, the same model reaches throughput 2000.00 and RTF 0.06400, while average TTFT rises to 3210ms and p95 TTFT to 6195ms. The highest throughput therefore does not retain the lowest first-token latency.
  • Qwen3-ASR-1.7B reaches online throughput 1219.51 at concurrency 128, with average TTFT 3392ms and RTF 0.10496. The smaller model has higher throughput at every matched concurrency in both inference modes, although its average TTFT is not lower in every online condition.

Increasing concurrency raises aggregate ASR throughput while also increasing first-token latency. For Qwen3-ASR-0.6B, 92ms average TTFT occurs at concurrency 1, whereas throughput 2000.00 occurs at concurrency 128 with 3210ms average TTFT; the alignment rows separately show offline throughput peaking at concurrency 8.

Offline and online inference measurements across concurrency levels, including RTF, throughput, and average and p95 TTFT.
Offline and online inference measurements across concurrency levels, including RTF, throughput, and average and p95 TTFT.

3 Qwen3-ForcedAligner

Qwen3-ForcedAligner-0.6B separates transcript generation from timestamp estimation. Given speech and its transcript, it predicts temporal boundaries through an LLM-based slot-filling task rather than generating a timestamp sequence autoregressively.

3.1 Overview

The transcript is augmented with [time] tokens that mark requested start and end boundaries for words or characters. The model directly predicts a discrete timestamp index for each slot, enabling selective alignment and word-, sentence-, or paragraph-level use through boundary placement.

  • The report claims a 67%–77% relative reduction in accumulated average shift on human-labeled datasets compared with competing alignment methods; the aggregate averages do not cover identical scenarios across methods.
  • The model supports speech in 11 languages, including cross-lingual scenarios, up to 300 seconds with a single set of weights.
  • Non-autoregressive prediction eliminates sequential next-token decoding for timestamp slots.

3.2 Model Design

A pretrained AuT encoder supplies speech embeddings, while the tokenizer supplies transcript embeddings containing [time] slots. Qwen3-0.6B processes their combined sequence, and a linear timestamp prediction layer classifies each requested boundary.

  • Training timestamps are discretized by dividing their values by the AuT encoder’s 80ms frame duration. The prediction layer has 3,750 classes, corresponding to the stated 300s input limit.
  • Neither the [time] token nor the prediction layer requires language-specific phoneme inventories or dictionaries; multilingual capability comes from the speech encoder and language model.
  • Figure 3 shows synchronized output and label positions, with cross-entropy supervision attached to timestamp slots rather than a shifted next-token objective.

Speech embeddings and a transcript containing dynamically inserted timestamp slots enter the same language-model backbone. The synchronized labels and outputs show how timestamp classification is supervised at each slot and computed in parallel without autoregressive timestamp generation.

Forced-alignment architecture with dynamic timestamp-slot insertion and synchronized cross-entropy supervision.
Forced-alignment architecture with dynamic timestamp-slot insertion and synchronized cross-entropy supervision.

3.3 Training Strategies

The aligner is trained with pseudo-timestamp labels generated by Montreal Forced Aligner (MFA), avoiding the cost of large-scale manual boundary annotation. The authors describe the learned model as distilling and smoothing noisy MFA labels; human-labeled evaluations test performance beyond agreement with its pseudo-label source.

  • Training retains causal processing but keeps output and label sequences non-shifted, so each timestamp slot is supervised at its own position and can use preceding context.
  • Cross-entropy loss is computed only at timestamp slots, focusing optimization on boundary prediction rather than transcript generation.
  • Dynamic slot insertion randomly determines whether start and end slots are added after each word or character, training the model to handle different boundary-selection patterns.

3.4 Inference and Usability

At inference, users insert boundary slots after selected words or characters, and the model predicts all slot indices simultaneously. Multiplying each predicted index by 80ms converts it to a timestamp.

  • Alignment speed is benchmarked with Transformers, FlashAttention, and bfloat16 on approximately 1-minute inputs; Table 2 identifies this as offline PyTorch inference only.
  • The highest listed alignment throughput is 961.54 audio seconds per second at concurrency 8, with RTF 0.00832. At concurrency 128, throughput is 649.35 and RTF is 0.19712.
  • The prose claim of RTF close to 0.001 under high concurrency does not match the per-request RTF values in Table 2. Near-1,000-fold aggregate throughput should not be conflated with per-request latency.

4 Experiments

The experiments evaluate transcription accuracy, language identification, streaming recognition, singing and song transcription, and forced-alignment precision. Public datasets are complemented by internal tests covering accents, dialects, speaker-age variation, difficult acoustics, and conversational speech.

4.1 Evaluation Details

ASR baselines include GPT-4o-Transcribe, Gemini-2.5-Pro, Doubao-ASR, Whisper-large-v3, FunASR-MLT-Nano, and, for multilingual evaluation, GLM-ASR-Nano. Recognition is measured with CER for character-based languages and WER for word-delimited languages; multilingual and dialect aggregates use unweighted macro-averages.

  • The internal robustness suite covers 16 English accent groups, 22 Chinese dialect varieties, elderly and children’s speech, extremely low-SNR audio, repetitive speech, and multi-speaker Mandarin conversation.
  • Language identification is evaluated by accuracy. Timestamp precision uses Accumulated Average Shift (AAS), the mean absolute difference between predicted and reference timestamps.
  • (\mathrm{AAS}=\frac{1}{N}\sum_{i=1}^{N}\left \hat{n}_i-n_i\right ), where N is the number of timestamp slots, n̂ᵢ is the predicted timestamp, and nᵢ is the MFA-derived or manually annotated reference.

4.2 English & Chinese ASR Performance

4.2.1 Opensource ASR Benchmarks: Table 3 shows strong Mandarin and dialect results, with mixed rankings on English. Qwen3-ASR-1.7B achieves WenetSpeech net/meeting errors of 4.97/5.88, compared with 9.86/19.11 for Whisper-large-v3, and WenetSpeech-Yue short/long errors of 5.82/8.85.

  • On English, the 1.7B model achieves 8.45 on GigaSpeech, 7.39 on CV-en, and 9.15 on VoxPopuli, the lowest listed results for those datasets.
  • It is not uniformly best: GPT-4o-Transcribe achieves 1.39 on LibriSpeech clean versus 1.63 for Qwen3-ASR-1.7B, and 2.40 on Fleurs-en versus 3.35. The smaller Qwen model also scores better on Tedlium, at 3.85 versus 4.50.
  • Doubao-ASR leads WenetSpeech-Chuan easy/hard at 11.40/20.20 versus 11.99/21.63 for Qwen3-ASR-1.7B, providing a dialect-specific exception.
  • Commercial API and Whisper results were obtained by the authors’ own inference runs; FunASR-MLT-Nano results were taken from its technical report. N/A denotes an API result the authors could not obtain reasonably, whereas a dash denotes an unreported result.

Public results support especially strong Mandarin and Cantonese performance, including large improvements over Whisper on WenetSpeech. English and Sichuanese rows also expose exceptions to a uniform superiority claim, with some commercial systems or the smaller Qwen model achieving lower errors.

English, Mandarin Chinese, and Chinese dialect recognition results for commercial APIs and open-source models.
English, Mandarin Chinese, and Chinese dialect recognition results for commercial APIs and open-source models.

4.2.2 Internal ASR Benchmarks: Qwen3-ASR-1.7B has the lowest error on all seven internal subsets in Table 4. Its accented-English WER is 16.07, compared with 16.62 for the 0.6B model and 19.96 for the strongest listed non-Qwen result.

  • Mandarin results are 3.81 for Elders&Kids, 16.17 for ExtremeNoise, 2.44 for TongueTwister, and 6.54 for Dialog-Mandarin.
  • Conversational Cantonese error is 4.12, and the macro-average over 22 Chinese dialects is 15.94, versus 18.24 for Qwen3-ASR-0.6B.
  • These results support robustness within the authors’ test suite, but the report does not specify its dataset sizes or provide a public release of the internal tests for independent replication.

Qwen3-ASR-1.7B has the lowest listed error in every internal subset, improving consistently over the 0.6B model. The rows extend evaluation beyond clean speech to accents, speaker-age variation, noise, repetition, conversation, and dialect diversity, although the report provides limited details for reproducing these tests.

Internal recognition results for accented English, challenging Mandarin speech, and conversational Chinese dialects.
Internal recognition results for accented English, challenging Mandarin speech, and conversational Chinese dialects.

4.3 Multilingual ASR and Language Identification

4.3.1 Multilingual ASR Performance: Qwen3-ASR-1.7B leads the listed open-source baselines on MLS, CommonVoice, MLC-SLM, the 12- and 20-language Fleurs subsets, and internal News-Multilingual. On the full 30-language Fleurs evaluation, however, Whisper-large-v3 has substantially lower average error.

  • The 1.7B model reports averages of 8.55 on MLS, 9.18 on CommonVoice, and 12.74 on MLC-SLM, compared with Whisper-large-v3 at 8.62, 10.77, and 15.68.
  • Fleurs, Fleurs†, and Fleurs†† cover progressively expanded inventories of 12, 20, and 30 languages. Qwen3-ASR-1.7B scores 4.90, 6.62, and 12.60; Whisper-large-v3 scores 5.27, 6.85, and 8.16.
  • On the 15-language News-Multilingual test, the 1.7B model scores 12.80 versus 14.80 for Whisper-large-v3 and 17.39 for Qwen3-ASR-0.6B.

The progressively expanded Fleurs inventories reveal a limitation hidden by narrower averages: Qwen3-ASR-1.7B leads on 12 and 20 languages but trails Whisper-large-v3 on 30 languages. Other rows show lower average errors than the listed baselines on MLS, CommonVoice, conversational multilingual speech, and internal news speech.

Macro-averaged multilingual recognition results on public datasets and internal multilingual news speech.
Macro-averaged multilingual recognition results on public datasets and internal multilingual news speech.

4.3.2 Language Identification Performance: Average language identification accuracy across four benchmarks is 97.9% for Qwen3-ASR-1.7B, 96.8% for Qwen3-ASR-0.6B, and 94.1% for Whisper-large-v3. The 1.7B model ties Whisper on MLS and exceeds it on CommonVoice, MLC-SLM, and Fleurs.

  • Qwen3-ASR-1.7B achieves 99.9% on MLS, 98.7% on CommonVoice, 94.1% on MLC-SLM, and 98.7% on 30-language Fleurs.
  • The authors identify Malay–Indonesian confusion as the main remaining source of Fleurs LID errors.
  • The prose describes MLS as covering 9 languages, whereas Table 5 lists 8 and Table 6 states that its inventories follow Table 5; this coverage inconsistency is unresolved in the report.

The broader 30-language transcription weakness is not accompanied by a corresponding language-identification weakness. Qwen3-ASR-1.7B achieves 98.7% identification accuracy on full Fleurs, compared with 94.6% for Whisper, despite its higher transcription error there.

Language identification accuracy on MLS, CommonVoice, MLC-SLM, and Fleurs.
Language identification accuracy on MLS, CommonVoice, MLC-SLM, and Fleurs.

4.4 Singing Voice & Songs Recognition Performance

Table 7 evaluates Qwen3-ASR-1.7B on singing-only recordings and entire songs with background music. It leads on M4Singer, MIR-1k-vocal, and Popcs, while FunASR-MLT-Nano narrowly leads on Opencpop.

  • The table reports WER (%) of 5.98 for M4Singer, 6.25 for MIR-1k-vocal, 3.08 for Opencpop, and 8.52 for Popcs. FunASR-MLT-Nano scores 2.98 on Opencpop.
  • For EntireSongs-zh, Qwen3-ASR-1.7B leads with 13.91 versus 18.68 for Gemini-2.5-Pro. For EntireSongs-en, Gemini leads with 12.18 versus 14.60 for Qwen.
  • Whisper-large-v3 and FunASR-MLT-Nano are marked N/A for entire-song recognition because of poor performance, not assigned numerical errors. The table does not evaluate the 0.6B Qwen model.

Separating singing-only recordings from entire songs tests transcription under different musical conditions. Qwen3-ASR-1.7B leads three singing datasets and Chinese entire songs, while FunASR leads Opencpop and Gemini leads English entire songs; N/A entries do not provide numerical baseline comparisons for entire-song recognition.

Singing-voice and entire-song transcription results, including songs with background music.
Singing-voice and entire-song transcription results, including songs with background music.

4.5 Streaming Speech Recognition

Streaming evaluation uses 2-second chunks, a 5-token fallback, and keeps the last four chunks unfixed. Table 8 evaluates a shared offline/streaming model, but streaming incurs a measurable accuracy penalty.

  • For Qwen3-ASR-1.7B, the reported average error increases from 2.69 offline to 3.33 streaming; LibriSpeech clean/other changes from 1.63/3.38 to 1.95/4.51.
  • For Qwen3-ASR-0.6B, average error increases from 3.48 to 4.40; LibriSpeech clean/other changes from 2.11/4.55 to 2.54/6.27.
  • The comparison reports recognition error rather than end-to-end streaming latency or the timing of transcript revisions.

4.6 Precision of Timestamps

Table 9 compares Qwen3-ForcedAligner-0.6B with Monotonic-Aligner, NFA, and WhisperX using AAS in milliseconds. The Qwen model maintains low error on both raw and concatenated inputs, including cross-lingual cases, whereas several baselines degrade sharply on long inputs.

  • On MFA-labeled data, Qwen’s average AAS is 42.9ms for raw inputs and 52.9ms for Concat-300s. The corresponding NFA averages are 129.8ms and 246.7ms; WhisperX averages are 133.2ms and 2708.4ms.
  • On human-labeled data, Qwen scores 27.8ms for Raw, 41.8ms for Raw-Noisy, 25.3ms for Concat-60s, 24.8ms for Concat-300s, and 42.5ms for Concat-Cross-lingual.
  • For human-labeled Concat-300s, Monotonic-Aligner and NFA score 410.8ms and 140.0ms, respectively. This matched condition provides direct evidence of Qwen’s long-input advantage.
  • Reported human-labeled averages are 32.4ms for Qwen, 141.3ms for Monotonic-Aligner, and 101.2ms for NFA. Missing baseline entries and different language or scenario coverage mean the aggregate columns are not fully matched comparisons.

5 Conclusion

The report concludes that large-scale speech pretraining, Qwen3-Omni initialization, and ASR post-training produce a broadly capable transcription family, with a separate non-autoregressive model supplying multilingual alignment. The results support strong Mandarin, dialect, internal robustness, and timestamp performance, but not uniform superiority across every English or multilingual benchmark.

  • Weights, inference tools, and fine-tuning support are released through https://huggingface.co/collections/Qwen/qwen3-asr, https://modelscope.cn/collections/Qwen/Qwen3-ASR, and https://github.com/QwenLM/Qwen3-ASR.
  • The stated single-input limits remain 1200s for ASR and 300s for forced alignment; longer recordings require processing beyond a single supported input.

6 Authors

The report is authored by Qwen Team and separates core contributors from a contributor list described as alphabetically ordered. Jin Xu and Junyang Lin are marked as corresponding authors; the PDF does not specify institutional affiliations.

Appendix

  • Table A.1 adds Qwen3-ASR-Flash-1208 API results for reference alongside the two released ASR models. Flash-1208 scores 1.33/2.40 on LibriSpeech clean/other and 4.60/5.80 on WenetSpeech net/meeting, but is not uniformly better: Qwen3-ASR-1.7B scores 8.45 versus 8.82 on GigaSpeech, and Qwen3-ASR-0.6B scores 3.85 versus 4.84 on Tedlium.
  • Table A.2 supplies per-language MLS, CommonVoice, MLC-SLM, and Fleurs results, including Flash-1208 as an API reference. The 1.7B model improves over the 0.6B model in every listed language row, but substantial absolute errors remain: Fleurs Hungarian is 34.22, Persian 29.90, Greek 28.08, and Finnish 25.23, compared with 2.41 for both Italian and Chinese.

Brief Thoughts

The forced aligner’s synchronized timestamp-slot objective is the report’s most distinctive methodological contribution: it uses a causal language-model backbone for parallel boundary classification without language-specific dictionaries. Human-labeled long-input results, particularly 24.8ms AAS on Concat-300s, provide stronger evidence for this design than agreement with MFA-generated labels alone.

The ASR family covers a broad set of tasks, but the report leaves attribution and reproducibility questions open. It provides no ablation isolating RL, context biasing, or dynamic attention windows, does not identify speed-benchmark hardware, and gives limited construction details for internal tests. Deployment judgments should retain the 30-language Fleurs gap and distinguish throughput from first-token and per-request latency.