Gorio Tech Blog search

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX | Summary

|

Contents

This article explains the key points of PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX.

  • 2026-08-18 (arXiv)
  • Zhang, Genghan, Dong, Yixin, Fan, Chengze, Zeng, Zhichen, Yuan, Yueming, Zhu, Shaowei, Olukotun, Kunle.
  • Stanford University, Carnegie Mellon University, RadixArk, Independent Researcher
  • Paper
  • Github
  • Project Page

Read this article in Korean


Summary

  • PTXBench evaluates whether LLMs can generate CUDA kernels with inline, architecture-specific PTX that are functionally correct, dynamically execute a requested instruction family, and improve latency relative to cuBLAS, cuDNN, or FlashInfer. It covers GEMM and attention workloads on NVIDIA H100 and B200 GPUs.
  • The results distinguish valid PTX generation, runtime instruction execution, and competitive performance. Repair-conditioned LoRA SFT improves Qwen3.6-27B on several tasks, but its transfer remains uneven and depends on coverage, balance, and the reasoning teacher.

1 Introduction

The paper motivates direct PTX programming as a means of exploiting GPU mechanisms that portable CUDA and higher-level kernel languages may not immediately expose. Unlike end-to-end kernel benchmarks, PTXBench fixes the target GPU and instruction family, prohibits cuDNN and cuBLAS headers in generated source, and tests whether the requested low-level mechanism actually runs.

2 PTXBench

PTXBench pairs a fixed workload and a frontier-library reference implementation with a target architecture and a required PTX instruction family. It uses FlashInfer-Trace to represent workloads, solutions, and checks, and MiniPTXAgent to generate, execute, and iteratively revise CUDA–PTX kernels.

The workflow separates the controlled capability probe from adaptation. Fixed architecture knowledge and a required instruction family feed a multi-turn CUDA–PTX agent; compilation and runtime feedback support revision, and successful repair trajectories can provide repair-and-rationale supervision for SFT.

PTXBench benchmark and adaptation workflow from controlled prompt construction through multi-turn generation, runtime measurement, and repair-conditioned adaptation.
PTXBench benchmark and adaptation workflow from controlled prompt construction through multi-turn generation, runtime measurement, and repair-conditioned adaptation.

2.1 Benchmark Tasks and Controlled Context

Every trajectory receives the same architecture-specific knowledge pack: architecture parameters, CUDA wrappers for PTX instructions, layout, synchronization, and memory-consistency contracts, and, for established operators, expert-validated scheduling principles. These packs contain 20k–30k tokens, so the benchmark controls for documentation retrieval rather than treating retrieval as a model capability.

2.2 Multi-turn Evaluation Protocol

MiniPTXAgent compiles candidates with nvcc, then sends successfully compiled candidates to an isolated profiling service for memory-safety checking, functional evaluation, diagnostics, and timing. Target-instruction correctness requires numerical correctness plus dynamic execution: static SASS inspection identifies an eligible family, and Nsight Compute must report a positive predicate-enabled thread count under the evaluated workload.

3 Adapting LLMs to Architecture-Specific PTX Programming

Fixit constructs SFT examples from failures generated by the pre-adaptation model. Given a failed kernel and execution feedback, a repair teacher supplies a correctness-filtered repair and a reasoning teacher produces a rationale; the student is conditioned on the original problem, failed kernel, and feedback and is supervised on the rationale and corrected kernel.

4 Benchmark Results

The benchmark reports turn-level functional correctness, target-instruction correctness, best speedup among correct kernels, and best speedup among correct kernels that execute the selected instruction. Results average three prompt variants with four eight-turn trajectories each (N = 32), using cuBLAS v13.1.0 for GEMM, cuDNN v9.20.0 for most attention tasks, and FlashInfer v0.6.14 for GQA.

4.1 Evaluation Metrics and Baselines

Fast_p is the fraction of all turns that produce a correct kernel faster than threshold p, whereas FastInst._p additionally requires runtime execution of the selected target instruction. Functional checking verifies output shape and dtype and uses torch.allclose with atol=rtol=1e-2; latency is the median of 50 timed iterations after 10 warmups.

4.2 Architecture-Specific PTX Capability Remains Uneven

Capability is strongly workload- and architecture-dependent. Claude Opus 4.8 reaches 1.012× cuBLAS on B200 GEMM with target-instruction execution, but no evaluated model consistently matches the frontier libraries; on B200, Claude’s target-instruction best speedups are 0.300× for MHA-Fwd, 0.269× for causal MHA-Fwd, and 0.149× for MHA-Bwd, with no qualifying causal-backward result.

The target-instruction FastInst._p curves show that performance-qualified success declines from GEMM and forward attention to backward attention, especially on B200. Claude Opus 4.8 has the strongest displayed B200 GEMM peak at 1.012×, while qualifying attention kernels remain uncommon and slower than their frontier baselines.

Target-instruction-qualified speedup distributions across GEMM and attention workloads on H100 and B200.
Target-instruction-qualified speedup distributions across GEMM and attention workloads on H100 and B200.

The table contextualizes the evaluated models by release date and disclosed knowledge cutoff relative to Hopper PTX ISA 8.0 and Blackwell PTX ISA 8.7. These dates support, but do not establish causally, the possibility that Blackwell-specific knowledge availability affects direct PTX capability.

Model release dates, knowledge cutoffs, and estimated lags after Hopper and Blackwell PTX ISA releases.
Model release dates, knowledge cutoffs, and estimated lags after Hopper and Blackwell PTX ISA releases.

4.3 High-Level Kernel Languages Remain Valuable on New Architectures

For Gemini 3.1 Pro, direct CUDA–PTX and Triton are relatively close on Hopper, including causal MHA-forward peaks of 0.768× and 0.759×, respectively. On B200 attention, Triton is substantially stronger: its backward peaks are 0.484× and 0.436×, versus 0.133× and 0.015× for CUDA–PTX; the authors suggest that Blackwell’s greater low-level coordination complexity and less available architecture-specific code may contribute to the gap.

For Gemini 3.1 Pro, Triton produces substantially higher B200 attention Fast_p distributions than CUDA–PTX, particularly in backward attention. The comparison supports the paper's conclusion that higher-level kernel languages remain useful on newly introduced architectures, despite occasional direct-PTX advantages on Hopper.

Unrestricted speedup distributions for Gemini 3.1 Pro generating Triton and CUDA-PTX kernels on H100 and B200.
Unrestricted speedup distributions for Gemini 3.1 Pro generating Triton and CUDA-PTX kernels on H100 and B200.

4.4 Explicit Architecture Knowledge Improves Target Instruction Execution

Explicit prompt knowledge substantially improves target-instruction use. For Gemini 3.1 Pro at eight turns, architecture parameters alone yield 26.0% unrestricted correctness but no target-instruction success; adding PTX template functions yields 19.8% target-instruction correctness, and adding the architecture contract raises this to 38.5%, without producing the highest peak speedup.

Adding PTX template functions changes Gemini 3.1 Pro from no target-instruction successes to 19.8% at eight turns, while adding architecture contracts raises that rate to 38.5%. The best-speedup columns show that more complete prompt knowledge most clearly improves reliable instruction-qualified correctness rather than monotonically maximizing peak performance.

Ablation of architecture parameters, PTX templates, and architecture contracts for Gemini 3.1 Pro.
Ablation of architecture parameters, PTX templates, and architecture contracts for Gemini 3.1 Pro.

5 Adaptation Results

The adaptation study applies LoRA SFT to Qwen3.6-27B because its 262K-token context can accommodate the long PTX prompts and kernel traces. It varies supervision format, task coverage and balance, reasoning teacher, transfer to held-out workloads, and inference-time alternatives to weight updates.

5.1 Training Format and Data Recipe Results

Fixit does not uniformly outperform direct-solution KernelGen supervision: with comparable coverage and record counts, the s3 Fixit recipe improves eight-turn correctness on GEMM, causal MHA-Fwd, and MHA-Bwd, but trails s0 on MHA-Fwd and causal MHA-Bwd. Among Fixit recipes, balanced coverage matters more than record count alone: s1 and s5 solve all five evaluation problems, whereas s6, using Qwen3.6-27B as the reasoning synthesizer for the same examples as s5, solves only GEMM.

The pie charts compare seven SFT recipes with different workload mixtures, record counts, supervision templates, and reasoning synthesizers. Their differing compositions provide the basis for comparing KernelGen with Fixit and for separating task distribution from raw record count.

Training data compositions, supervision formats, and reasoning teachers for the seven SFT recipes.
Training data compositions, supervision formats, and reasoning teachers for the seven SFT recipes.

The heatmaps compare correctness, target-instruction success, best speedup, and qualifying best speedup across Gemini and six Qwen adaptation recipes. They show the mixed KernelGen-versus-Fixit result and the broader task coverage of balanced s1 and s5, with no single recipe best across all workloads and measures.

Comparison of SFT data recipes across five kernel-generation problems and four evaluation measures.
Comparison of SFT data recipes across five kernel-generation problems and four evaluation measures.

5.2 Effects and Generalization of Fixit SFT

The selected s1 model, trained on four head-dimension-128 attention tasks, transfers to GEMM, all d64 MHA tasks, and two d96 forward tasks, but produces no correct d96 backward or GQA kernels. It reduces turn-level Triton correctness on every tested Hopper workload, although its best Triton speedup rises from 0.238× to 0.632× for causal MHA-Fwd and from 0.043× to 0.331× for causal MHA-Bwd.

After Fixit SFT, mean reasoning length increases across turns and the error distribution shifts away from compilation failures toward runtime and numerical errors. The adapted model therefore produces executable candidates more often, but execution validity and numerical correctness remain substantial bottlenecks.

Changes in reasoning-token length and error-type distributions after Fixit SFT.
Changes in reasoning-token length and error-type distributions after Fixit SFT.

The turn-level trajectories show the base Qwen3.6-27B remaining dominated by compilation errors and never reaching a correct GEMM kernel, whereas s1 produces correct kernels at turn 0 and at later turns. The remaining s1 failures shift toward runtime and numerical states.

Turn-by-turn transitions among compilation, runtime, numerical, timeout, extraction, other, and correct states for GEMM.
Turn-by-turn transitions among compilation, runtime, numerical, timeout, extraction, other, and correct states for GEMM.

5.3 SFT Versus In-context Supervision

Prompt-time guidance alone does not enable the base Qwen3.6-27B to generate correct MHA kernels, whereas s1 produces some correct MHA kernels without expert guidance. Retrieved repair notes alone also yield no correct kernels; retrieval becomes effective only when it includes the corrected kernel, a solution-bearing condition that places a near-solution in the prompt.

Neither the base model nor the base model supplemented with repair notes produces correct MHA kernels. S1 achieves nonzero target-instruction correctness in several conditions, while retrieval of both notes and a fixed kernel improves the base model because the prompt contains solution-bearing information.

Target-instruction-qualified and unrestricted MHA correctness under SFT, expert guidance, and retrieval-based prompt-time supervision.
Target-instruction-qualified and unrestricted MHA correctness under SFT, expert guidance, and retrieval-based prompt-time supervision.

The curves compare instruction-qualified speedups for s1 without expert guidance, s1 with guidance, and repair retrieval that includes a fixed kernel. They show that SFT yields some capability without solution-bearing retrieval, whereas including the corrected kernel creates much stronger opportunities for the base model to produce correct outputs.

Target-instruction-qualified speedup distributions under SFT and prompt-time supervision conditions.
Target-instruction-qualified speedup distributions under SFT and prompt-time supervision conditions.

Earlier kernel-generation benchmarks primarily assess functional correctness and overall speed, often permitting higher-level DSLs, ordinary CUDA, or library use. PTXBench instead isolates generation of from-scratch CUDA–PTX for a specified architecture and verifies at runtime that the required instruction family executes.

7 Limitations

The adaptation results use modest LoRA datasets and one 27B base model, so the effects of repair conditioning, data balance, and teacher quality may not transfer unchanged to industrial-scale post-training or other model families. PTXBench is also limited to BF16 GEMM and attention on H100 and B200 rather than the full range of GPU operators and hardware.

8 Conclusion

The authors present PTXBench as an auditable environment for measuring architecture-specific GPU-kernel capability and collecting targeted adaptation data. Current models sometimes solve forward workloads and execute requested instructions, but remain weak on backward attention and inconsistent relative to frontier libraries; Fixit helps selectively rather than uniformly.

Appendix

  • The appendix specifies evaluation on NVIDIA H100 80 GB and B200, nvcc -O3 compilation targeting sm 90a and sm 100a, and eight model calls per trajectory, with early stopping at 1.2× speedup or context exhaustion. It documents dynamic SASS and Nsight Compute measurement, primary BF16 workload shapes, LoRA rank 32 training for five epochs, prompt totals of 22,314 tokens for H100 and 28,815 for B200, profiling-service throughput measurements, corrected FlashAttention-3 pseudocode, and supplementary results.

Brief Thoughts

The benchmark’s central design choice is separating numerical correctness from dynamic instruction execution, which prevents dead-code and unlaunched-instruction claims from qualifying as architecture-specific success. The results also show that instruction execution is not a performance proxy: qualifying kernels often remain far below library baselines, while the adaptation evidence supports the importance of data quality and coverage more clearly than a broad claim of reliable repair-SFT transfer.