Gorio Tech Blog search

SPADE: Self-Play in Adaptive Synthetic Executable Environments | Summary

|

Contents

This article explains the key points of SPADE: Self-Play in Adaptive Synthetic Executable Environments.

  • 2026-08-19 (arXiv)
  • Liu, Bo, Yu, Simon, Jiang, Yiding, Qu, Ao, Zhao, Andrew, Liu, Zichen, Kim, Junsu, Zhou, Zijian, Kim, Seungone, Ren, Tongzheng, et al.
  • University of Washington, Stanford University, Northeastern University, Carnegie Mellon University, Massachusetts Institute of Technology, National University of Singapore, Seoul National University, Stevens Institute of Technology, University of Chicago
  • Paper

Read this article in Korean


Summary

  • SPADE trains a shared LLM to author executable, stateful training environments and solve them. A hint-based regret signal rewards environments where privileged guidance increases the Reasoning Agent’s return, while corpus grounding, memory, and a difficulty anchor are intended to maintain a feasible curriculum near the agent’s frontier.

Introduces the shared-policy, two-role training loop and plots reported improvement curves for games and tool use against fixed-environment baselines.

A single LLM designs executable environments and solves them with and without privileged hints; the return gap rewards the Designer, while task completion rewards the Agent.
A single LLM designs executable environments and solves them with and without privileged hints; the return gap rewards the Designer, while task completion rewards the Agent.

1 Introduction

The paper frames fixed human-curated, statically synthesized, and frozen-verifier environment pools as a bottleneck because their task distribution does not adapt as the learner improves. SPADE instead generates Python MDPs with a Gym-style reset()/step() interface, covering both reasoning games and multi-turn tool-use tasks.

Shows four environments from early through late training, illustrating the reported movement toward state-gated, multi-turn tasks as the Reasoning Agent improves.

Examples of generated environment observations, Python code, and Designer-written hints at training steps 0, 112, 200, and 384.
Examples of generated environment observations, Python code, and Designer-written hints at training steps 0, 112, 200, and 384.

SPADE relates its approach to LLM self-play, unsupervised environment design, and synthetic environment generation. Its stated distinction is online generation of complete executable MDPs, coupled with Designer optimization from measured Reasoning Agent returns rather than a fixed task generator or a fixed parameterized level space.

3 Preliminaries

The method combines Gym-style MDP environments with reinforcement learning from verifiable rewards optimized by GRPO. Generated Python code supplies state transitions, reward functions, and verification logic.

3.1 Markov Decision Processes and Gym-Style Environments

An MDP is defined as (S, A, T, R, ρ0). Generated environments expose reset(), which returns an initial observation and info dictionary, and step(a), which returns (s′, r, terminated, truncated, info); SPADE uses fully observed environments.

3.2 RLVR and GRPO

For each prompt, GRPO samples a group of responses and normalizes rewards relative to the group mean and standard deviation. The policy is optimized with a clipped objective and KL regularization against a reference policy.

4 SPADE: Self-Play in Adaptive Synthetic Executable Environments

SPADE alternates Environment Designer and Reasoning Agent trajectories within one shared policy πθ. The Designer produces validated environments and hints, while paired hinted and unhinted agent rollouts provide role-specific rewards for joint GRPO updates.

Depicts the complete data flow: corpus and memory condition the Designer, paired hinted and unhinted rollouts estimate regret, and Designer and Agent rewards update shared parameters.

SPADE framework with Environment Designer generation, Reasoning Agent rollouts, hint-based regret, and joint GRPO updates.
SPADE framework with Environment Designer generation, Reasoning Agent rollouts, hint-based regret, and joint GRPO updates.

4.1 Dual-Role Self-Play and Code-as-Environment

Role-specific prompts select Designer or Agent behavior while parameters remain shared. The Designer emits a Python program implementing transition and reward logic plus a privileged hint; syntax and executability checks reject invalid candidates before training.

4.2 Hint-Based Regret Reward

The basic Designer reward is the difference between mean hinted and unhinted Reasoning Agent returns. High regret identifies tasks solved with help but not without it, whereas low regret can reflect either mastery or infeasibility; deployment additionally blends floored regret with a difficulty anchor targeting a 0.4–0.6 win-rate band.

Provides rollout-level examples in which hints improve behavior. The displayed fiber pair has returns 0.00 and 1.00, while the audio pair displays 0.30 and 1.00; the caption clarifies that Designer reward uses arm means, including 0.30 to 0.65 for the audio task.

Two positive-regret examples comparing Reasoning Agent play with and without a privileged hint.
Two positive-regret examples comparing Reasoning Agent play with and without a privileged hint.

4.3 Environment Design

Fresh corpus documents provide topical material, while memory stores prior environments, regret scores, and skill tags for later variation. Games use 10k mathematics and 5k science documents from DCLM and MegaScience, whereas tool use uses 15k documents from the Nemotron pretraining code corpus.

5 Experimental Setup

Experiments use Qwen3-4B-Instruct-2507, Qwen3-8B, and Qwen3-30B-A3B-Instruct-2507. Held-out evaluation includes eight games-setting reasoning and code benchmarks, plus BFCL v4 multi-turn, τ²-bench, and ACEBench-Agent for tool use.

5.1 SPADE Training Recipe

The principal recipe uses GRPO, group size G=16, 24 environments per rollout, and 400 rollouts. Environments are regenerated every k rollouts, with k=4 for games and k=8 for tool use; delayed Designer updates use truncated importance sampling because their scoring is off-policy.

5.2 Domain: Games

For games, the Designer creates self-contained Python environments across six rotating cognitive skills: Mathematical Reasoning, Logical Deduction, Spatial Reasoning, Pattern Recognition, Optimization, and Causal Inference. The study compares SPADE with matched-budget Fixed-env GRPO and Fixed-env RLVE baselines.

5.3 Domain: Tool Use

For tool use, the Designer creates simulated OpenAI function-calling tools, mutable backend state, and three to five sequential user instructions with state-based criteria. Validation additionally checks deterministic resets and whether criteria are achievable by an available tool; benchmark tasks are not supplied to the generator.

6 Experimental Results

Across games and tool use, SPADE improves held-out scores over its untrained bases and the reported fixed-environment training sources. The evidence is presented as support for an adaptive curriculum built from online environment design, corpus grounding, and memory.

6.1 Quantitative Analysis

At Qwen3-30B-A3B, games-setting SPADE reaches an eight-benchmark average of 58.3, a +8.1-point gain over the 50.2 base and +5.3 over Fixed-env RLVE at 53.0. On tool use with the same backbone, it gains +5.7 on BFCL v4 multi-turn, +3.6 on τ²-bench, and +13.9 on ACEBench-Agent; cross-paper reference rows use different bases, data, and protocols.

Reports held-out games-setting results for three Qwen3 backbones. SPADE has the highest suite average within every backbone block, with its largest reported gain at 30B-A3B.

Held-out math, science, code, and Reasoning-Gym results for fixed-environment baselines and SPADE-generated games.
Held-out math, science, code, and Reasoning-Gym results for fixed-environment baselines and SPADE-generated games.

Shows checkpoint trajectories for eight games-setting evaluations at 30B-A3B. SPADE improves GPQA-Diamond, LiveCodeBench-v6, and procedural reasoning while competition-math performance remains near or above the untrained baseline at reported checkpoints.

Training trajectories on held-out math, science, code, and procedural-reasoning benchmarks for SPADE and fixed-environment baselines.
Training trajectories on held-out math, science, code, and procedural-reasoning benchmarks for SPADE and fixed-environment baselines.

Shows tool-use scores by backbone and benchmark; SPADE improves its Qwen3 bases on BFCL v4 multi-turn, τ²-bench, and ACEBench-Agent. The table explicitly notes that reference-system rows are not protocol-matched.

Tool-use benchmark results for synthetic-environment systems and SPADE post-training on Qwen3 backbones.
Tool-use benchmark results for synthetic-environment systems and SPADE post-training on Qwen3 backbones.

6.2 Qualitative Analysis

The reported environments become more state-gated and less likely to expose governing formulas directly: in Physics tasks, formula revelation falls from 25% early to 5% late. Qualitative trajectories shift from long up-front derivations toward short probes, evidence gathering, and targeted reasoning; without corpus grounding, the authors report repeated RotatingMazeEnv variants.

Tracks the fraction of generated environments with Reasoning Agent win rate in [0.2, 0.8]. Full SPADE rises to roughly one-third late in training, whereas displayed ablations decline or collapse.

Learnable-environment share over training for full SPADE and component ablations.
Learnable-environment share over training for full SPADE and component ablations.

Visualizes environment embeddings and reports a large diversity separation between corpus-grounded SPADE and the no-corpus ablation. The t-SNE is illustrative; the stated quantitative comparison is calibrated Vendi/n.

Corpus grounding comparison using t-SNE of SBERT-embedded environments and calibrated Vendi/n diversity scores.
Corpus grounding comparison using t-SNE of SBERT-embedded environments and calibrated Vendi/n diversity scores.

Shows representative Reasoning Agent behavior changing from invalid, front-loaded derivation to probing, revision, and evidence-based solution. It is qualitative evidence rather than a controlled causal measurement.

Reasoning Agent episodes at early, middle, and late training stages showing an evidence-first interaction shift.
Reasoning Agent episodes at early, middle, and late training stages showing an evidence-first interaction shift.

7 Ablations

Ablations test online Environment Designer adaptation and Designer reward design. The full configuration is strongest on the selected eight-benchmark average; removing both Designer training and memory yields 40.5, below the 50.2 untrained base, while a static GPT-5.5-generated pool reaches 53.0.

Quantifies the main configuration ablation at Qwen3-30B-A3B. The full system reaches 58.3, above all listed partial and non-adaptive variants, although some controls alter more than one factor.

Games-setting ablation results for Environment Designer training, corpus grounding, memory, and a fixed GPT-5.5 Designer.
Games-setting ablation results for Environment Designer training, corpus grounding, memory, and a fixed GPT-5.5 Designer.

7.1 Ablation: Environment Designer Adaptation

Removing memory lowers the selected suite average from 58.3 to 53.2, and removing corpus grounding lowers it to 53.5. These variants peak earlier than full SPADE, but the controls change multiple factors in the self-generation and GPT-5.5 comparisons, limiting component-wise causal attribution.

Displays per-benchmark trajectories for adaptation ablations, with full SPADE remaining strongest late while memory-, corpus-, or Designer-training removals deteriorate after earlier peaks.

Per-benchmark trajectories for full SPADE, partial ablations, and a fixed GPT-5.5 Environment Designer.
Per-benchmark trajectories for full SPADE, partial ablations, and a fixed GPT-5.5 Environment Designer.

7.2 Ablation: Environment Designer Reward

Hint-based regret raises the selected eight-benchmark average from 50.2 to 58.3, while the EMA learning-potential alternative reaches 55.9. The paper argues that unsigned deviation from a skill-level moving average can reward both mastered and hopeless environments and requires historical estimates before becoming informative.

Compares hint-based regret with EMA learning potential under trained-Designer settings. The regret variant reaches the highest reported suite average, while the no-training-and-memory control deteriorates.

Environment Designer reward ablation comparing hint-based regret, EMA learning potential, and no Designer training.
Environment Designer reward ablation comparing hint-based regret, EMA learning potential, and no Designer training.

8 Scaling Results

The games-setting gain over base rises from +5.2 at 4B to +5.7 at 8B and +8.1 at 30B-A3B, while Fixed-env GRPO remains near +1.2. A six-skill curriculum reaches 58.3 on the eight-benchmark suite, compared with 53.7 for the two-skill version.

Shows that the six-skill curriculum reaches a higher suite average than the restricted two-skill version.

Eight-benchmark suite-average trajectory for six-skill and two-skill SPADE curricula.
Eight-benchmark suite-average trajectory for six-skill and two-skill SPADE curricula.

Shows gains increasing with backbone scale and plots estimated hint regret. Only the 30B-A3B estimate remains positive for much of training; smaller models still improve despite noisy or negative finite-sample estimates.

Scaling of games-setting gain over base and hint-based regret trajectories for 4B, 8B, and 30B-A3B models.
Scaling of games-setting gain over base and hint-based regret trajectories for 4B, 8B, and 30B-A3B models.

9 Discussion

The results support adaptive environment design over the reported fixed pools and show that one code-as-environment interface can be used for games and multi-step tool use. Limitations include Designer complexity bounded by the base model and generation budget, a human-designed GRPO optimizer, no formal curriculum-optimality guarantee, and fixed-task rather than direct open-ended-growth evaluation.

10 Conclusion

SPADE makes environment design a learnable LLM post-training component through executable MDP generation, corpus and memory grounding, and hint-based regret. At 30B-A3B, the paper reports +8.1 average points over base in games, +13.9 on ACEBench-Agent, and +5.7 on BFCL v4 multi-turn.

Appendix

  • The appendix provides notation, an idealized equilibrium analysis, prompts, implementation details, reproducibility materials, extended ablations, extended related work, quantitative analyses, and generated-environment examples. Under assumptions of sound generation, articulated hints, and internalizability, its theoretical result states that every pure Nash equilibrium has zero hint regret and pointwise hint-free optimality on valid environments.

Brief Thoughts

The strongest empirical evidence concerns the joint role of corpus grounding and online adaptation: the paper reports Vendi/n of 0.68 with corpus grounding versus 0.04 without it. The conclusions remain bounded by generated-environment validity and non-matched external tool-use comparisons, making the released code, configurations, evaluation JSONs, and figure scripts important for independent verification.