Gorio Tech Blog search

Program-as-Weights: A Programming Paradigm for Fuzzy Functions | Summary

|

Contents

This article explains the key points of Program-as-Weights: A Programming Paradigm for Fuzzy Functions.

  • 2026-07-02 (arXiv), Preprint
  • Zhang, Wentao, Hotsko, Liliana, Kim, Woojeong, Nie, Pengyu, Shieber, Stuart, Deng, Yuntian.
  • University of Waterloo, Cornell University, Harvard University
  • Paper
  • Github
  • Project Page

Read this article in Korean


Summary

  • Program-as-Weights (PAW) targets fuzzy functions, including log triage, malformed-JSON repair, and intent-based ranking, that resist precise symbolic rules. It compiles a natural-language specification into a reusable neural program that a small frozen interpreter can execute locally.
  • The program combines a text pseudo-program with a compiler-generated LoRA adapter. The trained 4B LoRA compiler learns from FuzzyBench, a released 10M-example synthetic dataset; the off-the-shelf pseudo compiler and default Qwen3 0.6B interpreter remain frozen.
  • On FuzzyBench’s model-verified test set, PAW with the Qwen3 0.6B interpreter achieves 73.78% exact match versus 68.70% for direct Qwen3-32B prompting, using approximately 50× less inference memory. Compiled programs can execute offline, but the evaluations concern single-step functions, and changing interpreter families requires retraining the compiler.

1 Introduction

Remote LLM calls make fuzzy functions convenient to implement but introduce per-input cost, network dependence, and reproducibility concerns when providers change models. PAW moves the larger model’s work to compilation: a developer describes a function, obtains a portable program, and runs subsequent inputs through a device-resident interpreter. The paper provides code at https://github.com/programasweights and a demo at https://programasweights.com.

The upper path sends an urgency-classification description through the compiler once; the lower path applies the resulting discrete text and continuous adapter to a new message on a local interpreter. The diagram separates compilation from work repeated for each input.

Compile-once and run-locally workflow for a hybrid neural program.
Compile-once and run-locally workflow for a hybrid neural program.

2 Programs as Weights

For a specification s and input x, the abstraction is p = Compiler(s) and ŷ = Interpreter(p, x) ≈ f(x). In this implementation, p contains a discrete pseudo-program and a continuous parameter-efficient component; the former describes the task with examples, while the latter adapts the frozen interpreter. A Q4_0 adapter for the 0.6B interpreter is approximately 23 MB, in addition to the shared base model.

3 The Compiler–Interpreter System

The pipeline separates specification rewriting, PEFT generation, and execution. The authors instantiate its continuous component as either generated LoRA weights or a prefix-tuning KV cache while retaining the compiler–interpreter abstraction.

3.1 Compiler–interpreter abstraction

An off-the-shelf Qwen3-4B-Instruct-2507 pseudo compiler rewrites the specification as a self-contained task description with representative input-output examples. A separate trained 4B PEFT compiler consumes the specification and pseudo-program; at execution, the frozen interpreter receives the pseudo-program and generated PEFT module.

3.2 Text-to-LoRA: our current best

The LoRA compiler processes the specification, pseudo-program, EOS, and T = 64 learned prefix tokens in one forward pass. Hidden states from L depth-aligned compiler layers are mean-pooled and mapped to coefficients that mix N = 64 shared bases into per-layer LoRA A and B matrices. With rank r = 64 across attention q/k/v/o and MLP gate/up/down modules, each function generates approximately 38.5M LoRA parameters for the frozen interpreter.

The panels connect compiler hidden states to shared LoRA bases through a mapper, then show the generated adapter attached to the frozen interpreter. They distinguish the trained compiler and mapper from the interpreter used for execution.

Text-to-LoRA compiler, LoRA mapper, and frozen-interpreter architecture.
Text-to-LoRA compiler, LoRA mapper, and frozen-interpreter architecture.

3.3 Prefix-tuning: a precursor instantiation

The precursor maps compiler hidden states position-wise to key–value pairs prepended to the interpreter’s attention cache. At the paper’s controlled, equal-training-compute comparison scale, prefix-tuning reaches 50.4% FuzzyBench exact match, text-to-LoRA reaches 56.5% at the size-matched r = 18 and 65.7% at r = 64, and direct prompting reaches 9.8%. The main text experiments therefore use the LoRA version.

Both compiled PEFT forms outperform the 0.098 prompting baseline in this controlled comparison. LoRA rises from 0.565 at r = 18 to 0.657 at r = 64, compared with 0.504 for prefix-tuning.

FuzzyBench accuracy for prompting, prefix-tuning, and two text-to-LoRA ranks.
FuzzyBench accuracy for prompting, prefix-tuning, and two text-to-LoRA ranks.

4 Training

Only the PEFT compiler and mapper are trained: for each (s, x, y), a pre-generated pseudo-program and compiler-produced adapter are supplied to the frozen interpreter, with a negative mean-token log-likelihood objective for y. Gradients pass through the interpreter without updating its weights. Appendix G specifies three epochs over 10M examples, effective batch size 48, and an approximately 72-hour Qwen3 0.6B run on 3 GPUs.

5 FuzzyBench: A 10M-Example Dataset of Fuzzy Functions

FuzzyBench contains 10M synthetic (specification, input, output) examples generated using gpt-5.2 across 29 thematic versions and more than 800 sub-categories. Generation produces specifications and then input-output pairs; an 80/10/10 split by specification keeps test specifications unseen during training. The verified evaluation set retains cases where gpt-5.2 and gpt-5-mini agree, filtering some ambiguous targets without providing independent human verification.

6 Main Results

On the verified FuzzyBench test set, PAW with Qwen3 0.6B reaches 73.78% exact match, compared with 68.70% for direct Qwen3-32B prompting, 85.45% for direct gpt-oss-20B prompting, and 96.09% for the data-generating gpt-5.2 API model. The result exceeds the specified Qwen3 prompting baselines, not every larger or API model. Four external text datasets use task-specific accuracy or F1 rather than FuzzyBench exact match.

The PAW Qwen3 0.6B row reports 73.78% on FuzzyBench and a 23 MB per-program shipping size, against 68.70% for direct Qwen3-32B prompting. The external columns use different metrics—SMS F1 and accuracy for the others—and relative performance varies by dataset.

Main text-task comparisons, containment, interpreter size, and per-program shipping size.
Main text-task comparisons, containment, interpreter size, and per-program shipping size.

For image-conditioned functions, the authors replace the text compiler with Qwen3-VL-4B while retaining a small text interpreter and the LoRA mapper. The Qwen3 0.6B PAW LoRA row exceeds the listed VLM baselines on Circuit, Chemical, and Music, but its Im2LaTeX score is 0.181 versus 0.391 for PAW prefix-tuning. The appendix’s prefix-tuning component test suggests that pseudo-program examples can crowd the interpreter’s context on long-form image-to-markup tasks.

With the compiler changed to Qwen3-VL-4B, PAW LoRA with a Qwen3 0.6B interpreter exceeds the listed VLM baselines on Circuit, Chemical, and Music, but scores below prefix-tuning on Im2LaTeX. The six tasks show both cross-modal application and task-dependent differences.

Results on six image-conditioned fuzzy-function tasks.
Results on six image-conditioned fuzzy-function tasks.

7 Ablations

In the mapper comparison, the default shared-basis design scores 0.6223, above per-position aggregation (0.5598), per-layer bases (0.6028), and a combined LoRA-plus-prefix pathway (0.6033); the authors offer no theoretical explanation. The reported same-base, same-data, same-budget comparisons give PAW 0.7378 versus 0.5840 for full fine-tuning and 0.5210 for the best fixed LoRA.

PAW reaches 0.7378 against 0.5840 for full fine-tuning and 0.5210 for the strongest fixed LoRA. The paper reports the same base, data, and training budget for these comparisons.

Same-base comparisons with fixed LoRAs and full fine-tuning.
Same-base comparisons with fixed LoRAs and full fine-tuning.

8 Robustness to Noisy Specifications

Noise variants alter specifications while leaving task inputs unchanged. At epoch 2, exact match changes from 0.6692 on clean specifications to 0.6326 under heavy combined noise; not every perturbation reduces accuracy. In a separate heavy-typo ablation, the rewritten pseudo-program scores 0.6108 as interpreter input versus 0.5662 for the raw specification, supporting a denoising role for the discrete component.

Heavy combined noise changes accuracy from 0.6692 clean to 0.6326. Grammar noise scores slightly above clean in this run, so the perturbations do not uniformly reduce accuracy.

Epoch-2 accuracy under heavily perturbed specifications.
Epoch-2 accuracy under heavily perturbed specifications.

Under heavy typos, the rewritten pseudo-program scores 0.6108 versus 0.5662 when the interpreter receives the raw specification. The clean-specification scores are 0.6443 versus 0.6285, making the observed gap larger under typos.

Interpreter-input ablation comparing rewritten pseudo-programs with raw specifications.
Interpreter-input ablation comparing rewritten pseudo-programs with raw specifications.

9 Local Execution

The interface uses paw.compile(prompt) to obtain a serializable program and paw.function(id_or_path) to load it as a local callable; compilation uses a compiler service, while execution after download needs no external API call. On a 4096-example quantization subset, a Q6_K base plus Q4_0 adapter scores 0.6575 versus 0.6580 for PyTorch bf16. A Q4_K_M base plus Q4_0 adapter occupies approximately 507 MB in total and scores 0.6453; the separate IQ4_XS configuration has an approximately 430 MB base plus a 23 MB adapter.

The paired Python listings show compilation with paw.compile and subsequent loading with paw.function before a function call. They illustrate the developer interface rather than an accuracy or latency result.

Minimal Python examples for compiling and locally calling a fuzzy function.
Minimal Python examples for compiling and locally calling a fuzzy function.

On the 4096-example subset, Q6_K plus Q4_0 nearly retains the bf16 accuracy with listed base and adapter sizes of 623 MB and 23 MB. Q4_K_M plus Q4_0 has a combined size of approximately 507 MB and scores 0.6453, compared with 0.6580 for bf16.

Qwen3 0.6B interpreter accuracy and file sizes across quantization configurations.
Qwen3 0.6B interpreter accuracy and file sizes across quantization configurations.

On a MacBook M3 with Metal acceleration, the Q5_K_M base plus Q4_0 adapter runs at 31.6 tokens/s with a 0.48 s cold load. Five case studies cover log monitoring, site navigation, search reranking, tool calling, and a multilingual word-guessing game; the 10-function tool-calling pipeline scores 93% on TOOLCALL-15. These demonstrations include surrounding application code and do not constitute a controlled evaluation of general long-horizon reasoning.

PAW builds on text-conditioned hypernetworks and PEFT methods that generate adapters or prefixes without task-specific gradient descent at deployment time. Its stated distinctions are the hybrid pseudo-program-plus-PEFT artifact, training on programmer-style fuzzy-function specifications, and a versionable developer-facing program interface. ALCHEmist [Huang et al., 2024b] instead emits executable Python labeling code, a symbolic alternative when the task can be expressed reliably as code.

11 Conclusion

The paper demonstrates compile-once, locally executable fuzzy functions with a fixed lightweight runtime and extends the abstraction to image-conditioned tasks by replacing the compiler. Its strongest quantitative text result is on the verified synthetic FuzzyBench test set; performance for arbitrary independently authored specifications or long-horizon programs is not established.

Appendix

  • Appendix A documents a hosted compile, test, and export workflow; B and C provide data-construction and compiler/interpreter prompts, including the off-the-shelf pseudo compiler’s request for 3-6 examples. D and D.1 give image-task prompts and a prefix-tuning component test: adding the discrete pseudo-program helps several short-answer tasks but reduces performance on Im2SMILES and Im2LaTeX relative to the continuous-only variant. E diagrams the prefix-tuning precursor, while F records the 29 dataset versions, including version 23’s 8.25M-example spaCy-superset/custom-NLP-pipeline theme.
  • Appendices G and H detail training and additional ablations; I reports an inconclusive, non-monotonic compiler-size and freezing comparison, and J provides per-noise-type results. K contains full quantization tables, while L’s inspection of 20 rollouts per interpreter describes observed strengths and failures without establishing population-level rates. M walks through the five case studies: log monitoring needs a separate stall timer, and the tool-call pipeline uses deterministic Python for JSON construction. N identifies interpreter coupling, opaque adapters, single-step evaluation, synthetic-data reliance, and task-dependent PEFT choice; O discusses offline benefits and the authors’ assessment of misuse risks.

Brief Thoughts

The same-base ablations and measured quantized-device results support amortizing compilation across repeated local calls. The principal limitation is evaluation scope: held-out synthetic specifications and five case studies do not measure reliability across independently authored production tasks, and the image results show that the preferred PEFT form depends on the task.