Gorio Tech Blog search

Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization | Summary

|

Contents

This article explains the key points of Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization.

  • 2026-06-10 (arXiv), Preprint
  • Xiang, Hao, Tang, Qiaoyu, Yu, Le, Lu, Yaojie, Han, Xianpei, He, Ben, Sun, Le, Yu, Bowen, Wang, Peng, Lin, Hongyu, et al.
  • University of Chinese Academy of Sciences, Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences, Qwen Team, Alibaba Group
  • Paper

Read this article in Korean


Summary

  • RACES (Recursive Automated Composition for Environment Scaling) expands a finite pool of executable reasoning environments through recursive, type-compatible composition. Its four operators—SEQUENTIAL, PARALLEL, SORT, and SELECT—produce tasks involving state tracking, independent computations, operation-order inference, and subset selection.
  • Using 300 base environments and matched training-instance counts and RL steps, RACES improves DeepSeek-R1-Distill-Qwen-14B from 48.2 to 51.3 and Qwen3-14B from 58.8 to 61.1 on the reported benchmark average. Relative to RL on individual environments, the gains are 2.5 and 1.0 points, respectively, with improvements in every displayed benchmark column.
  • RACES with 50 base environments exceeds individual-environment RL with 300 environments on both backbones tested in the pool-size comparison. The depth analysis reveals a difficulty–trainability trade-off: deeper compositions are harder to optimize and do not consistently improve transfer. Sparse rewards on the smaller model, long context requirements, and unexplored conditional branching and bounded loops limit the demonstrated scope.

1 Introduction

The paper addresses the cost of expanding verifiable RL training distributions when environments are constructed individually. An environment can generate many input instances, but adding new transformation families requires further construction; RACES instead reuses executable transformations in different positions, combinations, and task presentations.

  • Models solve environment descriptions and inputs without external tools, while executable reference programs provide verifiable outputs.
  • Matching one environment’s output type to another’s input type enables recursive assembly, but execution and quality filtering remain necessary.
  • The proposed source of transfer is structural diversity, not merely a larger number of sampled inputs.

The related work covers RLVR, verifiable-environment synthesis, and problem composition. RACES moves composition to executable environment interfaces rather than connecting questions at the natural-language level.

  • RLVE demonstrates the value of adaptive verifiable environments; SCALER and RESYN automate synthesis but still construct environments individually.
  • MathFusion, H1, and Composition-RL combine problems or dependency chains. RACES supports programmatic recursive composition through standardized domain signatures.
  • Its distinction is reusable environment-level composition without hand-crafted adapters for each pair, subject to domain compatibility and runtime filtering.

3 Method

The method separates executable environment semantics from model-facing task presentation. It defines a standardized interface, discovers and filters compatible execution paths, and applies operators that specify the information shown to the model, required predictions, and rewards.

3.1 Preliminary

A verifiable environment is defined as (e=(G_e,f_e,D_e,V_e)). Its executable mapper supplies deterministic reference outputs, while the other components generate valid inputs, render natural-language problems, and check model answers.

  • G_e samples valid inputs from X_e; f_e: X_e → Y_e maps each valid input to a unique output.
  • The DOMAIN SIGNATURE τ_e = (X_e, Y_e) determines candidate composability.
  • D_e: X_e → Σ* renders problems, and V_e: Y_e × Y_e → {0, 1} compares predictions with reference answers.

3.2 Compositional Closure of Verifiable Environments

For environments whose signatures satisfy Y_ei = X_ej, the composite mapper is (f_ej ◦ f_ei)(x) = f_ej(f_ei(x)). Determinism is preserved, and the resulting interface maps X_ei to Y_ej, allowing further compatible extensions.

  • A length-t chain has mapper F_πt = f_et ◦ f_e(t−1) ◦ ··· ◦ f_e1 and signature (X_e1, Y_et).
  • Recursive closure expands the structures available from a finite pool; it does not establish that every type-compatible chain is executable or useful.
  • The implementation treats domain matching as a proposal mechanism and validates actual intermediate states.

3.3 Implementation of Recursive Composition

RACES uses randomized breadth-first search over executable paths: states are intermediate values, and outgoing edges are environments whose input domains match those states. Figure 1 relates this search-and-composition process to the standardized interface and four task operators.

  • At state y_t, candidates satisfy (C(y_t)={e\in\mathcal{E}:X_e=\operatorname{type}(y_t)}); an extension is retained only after successful execution and output-quality checks.
  • Search is bounded by maximum depth, per-execution limits, and a cap on extensions per frontier state. Remaining usage budgets weight candidate sampling to balance environment utilization.
  • Operator instantiation follows path discovery, allowing one executable composite to support distinct prediction tasks. Operator-specific size sampling prevents a narrow range of path lengths from dominating training.

The integer-to-integer, integer-to-string, and string-to-string examples show how matching output/input interfaces permits recursive assembly. The lower panels distinguish following an ordered pipeline, solving independent tasks, recovering an ordering, and selecting an ordered subset—different prediction tasks built from reusable executable components.

RACES environment components, type-compatible recursive assembly, and the SEQUENTIAL, PARALLEL, SORT, and SELECT operators.
RACES environment components, type-compatible recursive assembly, and the SEQUENTIAL, PARALLEL, SORT, and SELECT operators.

3.4 Composition Operators

SEQUENTIAL presents an initial input and ordered environment descriptors, then requires every intermediate output. Its reward is RSeq = K/t, where K is the length of the longest correct prefix; predictions after the first error receive no additional credit.

  • The prefix reward reflects causal dependence in the reference execution chain.
  • The task exercises transformation chaining and preservation of intermediate states.

PARALLEL packages n independent environment-input pairs into one context and requests all outputs. Its reward averages the individual verifier results, so a failure on one problem does not invalidate correct answers to the others.

  • This operator tests separation and maintenance of independent computational threads rather than sequential dependence.

SORT supplies the initial input, target terminal output, and shuffled environment descriptors, and asks for a permutation. The verifier executes the predicted ordering and awards a binary reward when it reaches the target, including valid non-canonical orderings.

  • The model must infer a working operation order rather than follow a disclosed chain.

SELECT adds executable, domain-compatible distractors to a discovered path and asks the model to select and order distinct candidates. A binary reward is awarded when the predicted sequence maps the initial input to the target output.

  • Distractors come from alternatives encountered during path discovery and are executable on intermediate states.
  • Verification checks the resulting computation rather than exact agreement with the original path.

4 Experiments

The experiments compare individual-environment RL and RACES using the same underlying pool, training-instance counts, and RL steps. Main results measure transfer to benchmarks not used to construct the training environments; smaller-model analyses examine training dynamics, pool size, and composition depth.

4.1 Experimental Setup

The main backbones are DeepSeek-R1-Distill-Qwen-14B and Qwen3-14B; analysis uses Qwen3-4B-Instruct-2507 to reduce computational cost. Evaluation covers LiveCodeBench, AIME 2024, AIME 2025, Enigmata, IFEval, and LongBench-v2.

  • The six benchmarks cover code generation, mathematics, logic, instruction following, and long-context understanding.
  • AIME 2024/2025 are evaluated 32 times each and other benchmarks 4 times each, with average scores reported.
  • The tables combine the two AIME benchmarks into one AIME column, leaving five displayed task columns.

Main training draws composition sizes uniformly from [2, 12] for SEQUENTIAL and PARALLEL, and [2, 6] for SORT and SELECT. Analysis reduces these ranges to [2, 6] for SEQUENTIAL and PARALLEL and [2, 3] for SORT; SELECT is excluded because its rewards are near zero on the smaller backbone.

  • Both configurations draw from the same pool of 300 environments; the baseline samples individual tasks, while RACES samples operators, sizes, and compatible environments.
  • An anonymized subset is available at https://anonymous.4open.science/r/Submission_of_NIPS2026_34776-B7FD.

Training uses VERL, vLLM rollouts, and GRPO on a 32 × NVIDIA A100 80GB cluster, with clip ratio 0.28 and no KL regularization to a reference policy. Main runs use 300 steps, 12,800 training instances, batch size 128, learning rate 2e-6, 8 rollouts per problem, and maximum sequence length 32K tokens.

  • Analysis runs use 200 steps, 6,400 instances, batch size 64, and maximum sequence length 16K tokens.
  • Matched instance counts and steps control data quantity and update count, but do not establish equal generated-token or total compute costs.

4.2 Main Results

Table 1 shows that RACES outperforms both the base models and individual-environment RL on every displayed task column. DeepSeek-R1-Distill-Qwen-14B reaches 51.3 versus 48.8 for RLindividual and 48.2 for Base; Qwen3-14B reaches 61.1 versus 60.1 and 58.8.

  • For DeepSeek-R1-Distill-Qwen-14B, RACES scores 48.8 on LiveCodeBench, 35.4 on Enigmata, 36.0 on LongBench-v2, 74.6 on IFEval, and 61.7 on AIME.
  • For Qwen3-14B, the corresponding scores are 57.0, 49.2, 35.5, 86.7, and 77.0.
  • The results support transfer beyond the constructed environments. Repeated evaluation reduces sampling noise, but Table 1 provides neither confidence intervals nor training-seed variability.

5 Analysis

The analysis examines transfer during training, reuse of smaller base pools, and the effect of composition depth on optimization. Most analyses use Qwen3-4B-Instruct-2507, while the pool-efficiency comparison also includes DeepSeek-R1-Distill-Qwen-14B.

5.1 Performance across Training Process

Figure 2 separates training reward from downstream transfer: RLindividual earns consistently higher rewards, whereas RACES achieves larger benchmark gains later in training. At step 200, the average scores are 51.9 for RACES and 50.4 for RLindividual.

  • Individual-environment training adapts faster to simpler tasks, while composite tasks remain harder to solve.
  • Higher training reward does not imply better transfer. The curves alone do not establish overfitting as the causal explanation, and neither benchmark trajectory is strictly increasing.
  • Rewards come from different task distributions and operator rules, so their magnitudes are not a common measure of reasoning ability.

Individual-environment RL maintains higher training rewards, but RACES achieves larger benchmark gains later in training. At step 200, RACES reaches 51.9 versus 50.4. The early crossings and fluctuations show that this is a later-stage transfer advantage, not a uniform advantage at every checkpoint.

Training rewards and average benchmark performance over 200 RL steps on Qwen3-4B-Instruct-2507.
Training rewards and average benchmark performance over 200 RL steps on Qwen3-4B-Instruct-2507.

5.2 Efficiency of Environment Utilization

Table 2 shows improved utilization of the base pool. On Qwen3-4B-Instruct-2507, RACES with 50 environments scores 50.8, exceeding RLindividual with 300 environments at 50.4; RACES with all 300 reaches 51.9.

  • The Qwen3-4B-Instruct-2507 base scores 49.2, and individual RL with 50 environments scores 50.2.
  • On DeepSeek-R1-Distill-Qwen-14B, RACES with 50 environments reaches 50.2 versus 48.8 for individual RL with 300.
  • These are average-score comparisons, not gains in every column for the 4B model, and they do not demonstrate a sixfold reduction in end-to-end training cost.

With 50 environments, RACES exceeds individual RL with 300 in average score on Qwen3-4B-Instruct-2507, 50.8 versus 50.4, and DeepSeek-R1-Distill-Qwen-14B, 50.2 versus 48.8. The 4B comparison does not favor RACES in every task column. These results support improved base-pool reuse, while total construction and training costs remain unmeasured.

Benchmark results for individual-environment RL and RACES with different initial environment-pool sizes.
Benchmark results for individual-environment RL and RACES with different initial environment-pool sizes.

5.3 Effect of Composition Size

With the pool fixed at 300 environments, the SEQUENTIAL analysis tests sizes 2–6 on Qwen3-4B-Instruct-2507. Figure 3 shows lower rewards and generally lower rollout-reward dispersion for deeper compositions, while Table 3 reports average scores of 50.8, 50.7, 51.0, 51.2, and 50.7 for sizes 2, 3, 4, 5, and 6.

  • Size 5 performs best in this experiment, while size 6 falls below size 2; depth should be tuned rather than maximized.
  • The prose describes steady improvement from size 2 to size 5, but the table includes a decrease at size 3. The supported relationship is non-monotonic, with an intermediate-depth optimum among the tested sizes.
  • Figure 3 labels its second metric reward variance, but the caption defines it as the mean per-problem rollout reward standard deviation.

Longer SEQUENTIAL compositions generally have lower training reward and rollout-reward dispersion, consistent with greater optimization difficulty. The second panel is labeled reward variance, but the caption defines the plotted quantity as the mean per-problem rollout reward standard deviation; it should therefore be interpreted as dispersion rather than a variance statistic.

Training reward and average rollout-reward dispersion for SEQUENTIAL composition lengths 2–6.
Training reward and average rollout-reward dispersion for SEQUENTIAL composition lengths 2–6.

Scores for lengths 2–6 are 50.8, 50.7, 51.0, 51.2, and 50.7. Length 5 performs best among the tested sizes, but the decreases at lengths 3 and 6 contradict a strictly increasing benefit from depth.

Average benchmark scores under different SEQUENTIAL composition sizes.
Average benchmark scores under different SEQUENTIAL composition sizes.

5.4 Pattern Analysis

Three qualitative cases associate RACES responses with reusable abstractions, explicit state tracking, hypothesis testing, and constraint propagation. They illustrate possible transfer behaviors but neither isolate operator-specific effects nor establish their frequency across the evaluation set.

  • In AIME grid coloring, Base makes variable-substitution errors and RLindividual inconsistently treats a red edge as blue; RACES uses a reusable counting function and two checks of the total.
  • In list transformation, the baselines return an incorrect list after unsuccessful hypothesis exploration. The RACES excerpt rejects a suffix hypothesis, tracks indices, and returns the ground-truth output, but does not supply a complete rule explaining all demonstrations.
  • In Sum Skyscraper, the baselines reach an incorrect no-solution conclusion through incomplete row-wise searches; RACES uses column-wise enumeration and constraint propagation, and the excerpt reports verification of all 16 clues.

6 Conclusion

The paper concludes that standardized executable interfaces enable recursive environment composition as an alternative to constructing every environment independently. The experiments support better benchmark transfer and base-pool utilization under the tested conditions, while the depth analysis shows an optimization trade-off in controlling task difficulty.

Appendix

  • A Limitations: The four operators cover representative patterns but leave conditional branching and bounded loops unexplored. Composite tasks require adequate initial reasoning competence and can produce sparse rewards for weaker models; extended traces also demand large context windows, with 32K tokens used in the main experiments.
  • B Declaration of LLM Usage: Claude-Sonnet-4.5 generates candidate environments, followed by static code self-consistency and multi-sample output-consistency filtering. The authors separately report limited writing assistance for grammar checking and minor sentence-level refinement.
  • C Broader Impacts: The authors argue that reusable verifiable environments can reduce annotation dependence and make RL data construction more accessible. They identify malicious adaptation and energy consumption as risks, state that model checkpoints are not released, and report that the 300 environments contain no content related to poison or dangerous material.
  • D Initial environments construction.: The 300 environments come from standardized algorithmic problem datasets, Claude-Sonnet-4.5 generation, and manual authoring; 45 are algorithm-derived. Quality review samples 20 inputs per environment and uses repeated Claude-Sonnet-4.5 solutions checked by the verifier. The supplied source truncates the pass-rate threshold after “exceeding 95,” so its complete value and unit cannot be verified.
  • D Initial environments construction. also specifies executable filtering: reject exceptions, missing outputs, execution exceeding a 2-second wall-clock timeout, and extensions requiring more than 400 atomic steps. Non-degeneracy filtering rejects repeated path states or sibling outputs, outputs equal to 0 or 1, integers greater than 500, string representations longer than 100, and strings consisting of one repeated character. Paths shorter than two steps are discarded.
  • E Case Studies for Pattern Analysis presents prompts and key response excerpts rather than complete responses. E.1 AIME: 2 × 2 Grid-Coloring asks for colorings of 12 unit segments such that each unit square has two red and two blue sides; “Problem (AIME 2025), ground truth:” gives 82. “Base, variable-confusion in case-table filling” produces 158, and “RLindividual, single inconsistent state update” produces 83. “RACES, functional abstraction with two-way verification” defines f(x, y) as 1 when x + y ∈ {0, 2} and 2 when x + y = 1, then checks the total 82 by grouping and sequential summation.
  • E.2 Enigmata: List-Transformation Rule Induction provides four demonstrations and test input [46, 94, 66, 98, 66, 66]; “Problem, ground truth:” gives [66, 66, 66]. “Base, broad hypothesis exploration without verification” and “RLindividual, contradictions noted but not used” both return [66, 66]. “RACES, index tracking and rule preservation across all demos” returns the correct list after rejecting a suffix hypothesis and tracking repeated values, but the excerpt does not state a complete rule explaining all demonstrations. It therefore supports a correct case-level answer more clearly than verified recovery of the underlying function.
  • E.3 Enigmata: Sum Skyscraper Logic Puzzle includes “Problem, Sum Skyscraper 4×4,” requiring row and column permutations of {1, 2, 3, 4} and visibility-sum clues on all four sides. “Base, visibility assumption error in row-wise casework” discards a valid row it had already found, while “RLindividual, incomplete row enumeration” misses valid candidates; both incorrectly conclude that no solution exists. “RACES, column-wise decomposition with cascading propagation” enumerates column candidates and propagates row-uniqueness constraints to obtain [[3, 2, 4, 1], [2, 3, 1, 4], [4, 1, 2, 3], [1, 4, 3, 2]], with the excerpt reporting that all 16 clues pass verification.

Brief Thoughts

The central contribution is an executable composition interface that produces structurally different training tasks without defining new semantics for every composite. Matched-pool comparisons and the 50-versus-300 environment results support its value, but equal instances and steps do not establish compute efficiency, and qualitative cases do not establish operator-specific causal mechanisms. Runtime filtering is as important as type matching: compatible types alone do not prevent invalid, degenerate, or excessively expensive intermediate computations. Near-zero SELECT rewards on the 4B model and the decline at SEQUENTIAL size 6 show that task complexity must be matched to model competence. An effectively unbounded composition space is a structural possibility, not an experimentally demonstrated supply of arbitrarily deep, useful training tasks.