Gorio Tech Blog search

Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models | Summary

|

Contents

This article explains the key points of Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models.

  • 2026-03-10 (arXiv)
  • Hennes, Daniel, Li, Zun, Schultz, John, Lanctot, Marc.
  • Google DeepMind
  • Paper

Read this article in Korean


Summary

  • Code-Space Response Oracles (CSRO) replaces the reinforcement learning best-response oracle in Policy-Space Response Oracles (PSRO) with an LLM that generates executable policies. The resulting strategies are stateful programs whose decision rules and opponent models can be inspected directly.
  • CSRO combines opponent-conditioned code generation with optional feedback-driven refinement through ZeroShot, LinearRefinement, or AlphaEvolve. Experiments use Gemini 2.5 Pro over 20 outer iterations in 1000-round Repeated Rock-Paper-Scissors and 100-hand repeated Leduc hold’em poker.
  • In Rock-Paper-Scissors, LinearRefinement with opponent code and Top 5 filtering achieves an Aggregate Score of 122.1 ± 9.8, compared with 126.0 for a turn-by-turn Gemma 3 27B agent. In Leduc, AlphaEvolve achieves 44.9 ± 4.1 versus 39.8 ± 0.3 for CFR+, but its population exploitability is 4.4 ± 0.6 rather than 0.0 ± 0.0. These results support competitive, inspectable policy synthesis on the tested populations, not general equilibrium convergence or robustness against all possible opponents.

1. Introduction

Standard PSRO expands a policy population by training approximate best responses to its current equilibrium mixture. Deep RL oracles produce neural policies that are difficult to inspect and can require extensive simulation. CSRO instead asks an LLM to synthesize a program from game rules, an environment API, and information about the opponent mixture.

  • Relative to LLM-PSRO (Bachrach et al., 2025), the proposed extensions are closed-loop candidate refinement and context abstraction through opponent descriptions and filtering.
  • The paper evaluates against external opponent populations and established baselines, extending the inter-population evaluation used in LLM-PSRO.

2. Code Space Response Oracles

CSRO retains PSRO’s population-based game-theoretic structure while changing the representation and construction of best responses. An empirical meta-game determines an equilibrium mixture over existing policies, and an LLM generates the next executable response to that mixture.

2.1. Preliminaries

The formal development focuses on two-player symmetric zero-sum games with a shared strategy space. PSRO maintains a policy set P, computes a symmetric equilibrium mixture σ in its empirical payoff matrix, and seeks a response π* ∈ arg maxπ Eπ′∼σ[u(π, π′)]. The new response is added to P before the process repeats.

  • The zero-sum condition is u(π′, π) = −u(π, π′).
  • The authors state that the methodology transfers to more general games, but the reported experiments remain within the symmetric zero-sum setting.

2.2. Code Policies

A code policy is a stateful program, such as a Python function or class, that maps observations to actions. Its prompt specifies the game, the executable interface, and the current opponent meta-strategy; opponent source code can be supplied directly, supplemented by summaries, or replaced by LLM-generated natural-language descriptions.

  • Comments and docstrings describe the intended strategy, while the executable implementation exposes the actual decision logic.
  • Descriptions and filtering reduce prompt size but also determine which opponent behaviors are represented in the oracle’s context.

The three heatmaps connect population growth to empirical payoffs, equilibrium weights, and external-bot returns in one RRPS run. Equilibrium support frequently concentrates on recently added policies, while returns differ across external opponents. The example motivates checking internal meta-game progress against an independent population.

An example CSRO run in Repeated Rock-Paper-Scissors, showing meta-game payoffs, equilibrium mixtures, and returns against competition bots across iterations.
An example CSRO run in Repeated Rock-Paper-Scissors, showing meta-game payoffs, equilibrium mixtures, and returns against competition bots across iterations.

2.3. CSRO Algorithm

Algorithm 1 initializes P with an initial policy, recomputes the payoff matrix and meta-equilibrium, constructs an opponent-conditioned prompt, and generates and evaluates a candidate. An optional inner loop updates the prompt using the candidate and its measured utility before adding the final program to P. The pseudocode returns the population and meta-strategy.

  • The outer loop changes the strategic target as the population grows; the inner loop refines a response to a fixed target mixture.

2.4. Oracle Refinement Mechanisms

2.4.1. Cross-Iteration Strategic Adaptation constructs strategic context from the current meta-equilibrium before each new oracle call. construct_prompt can include opponent code or LLM-generated summaries and limit the input using a minimum-support threshold or a top-k filter; the experiments include Top 5 filtering.

  • Filtering controls context length, but a concentrated equilibrium can leave a minimum-support prompt with very few opponent strategies.

2.4.2. Intra-Iteration Policy Refinement defines three oracle variants. ZeroShot generates one candidate without refinement; LinearRefinement starts revising only when the candidate has negative utility, retains score-improving changes, and stops at non-negative utility or budget M. AlphaEvolve uses distributed evolutionary program search scored by estimated expected utility against σ.

  • LinearRefinement performs its updates in a single thread.
  • AlphaEvolve maintains independently evolved subpopulations and samples past programs for LLM-generated mutations, exploring multiple revision paths (Novikov et al., 2025; Romera-Paredes et al., 2024).
  • Non-negative utility is a stopping condition, not evidence that the candidate is an exact best response.

3. Experiments

The experiments ask whether CSRO reaches low-exploitability strategies, whether its LLM oracle generates effective non-trivial policies with little task-specific feedback, and whether the programs expose strategic logic more clearly than neural policies. Both environments use OpenSpiel (Lanctot et al., 2019). The reported interpretability evidence consists of inspecting selected source programs rather than a controlled comparison of human understanding.

3.1. Environments

Repeated Rock-Paper-Scissors (RRPS): matches contain 1000 consecutive throws. Repeated Leduc hold’em poker: matches contain 100 hands, with dealer roles alternating and the first dealer selected randomly. These tasks test adaptation to recurring opponents under different information structures.

  • In RRPS, independent uniform random play is unexploitable but does not exploit predictable opponents.
  • Leduc uses six cards—two each of Jack, Queen, and King—with one private card per player, two betting rounds, and one public community card. Repetition allows policies to learn opponent tendencies across hands.

3.2. Evaluation Population

RRPS evaluation uses 43 hand-coded competition bots spanning fixed moves, periodic sequences, frequency analysis, Markov prediction, and adaptive predictor ensembles. Leduc evaluation uses AlwaysCall, AlwaysFold, and a single-hand strategy computed with CFR+ for 10 000 iterations and weighted averaging, repeated at every hand.

  • The RRPS population includes adaptive opponents such as iocainebot and greenberg alongside simple strategies such as rockbot and randbot.
  • AlwaysFold folds whenever possible and otherwise calls. The three-opponent Leduc population tests equilibrium-oriented play and exploitation of two conspicuous behavioral patterns.

Population Return is PopReturn(π) = Eπ′∼P[u(π, π′)], while Within Population Exploitability is PopExpl(π) = −minπ′∈P u(π, π′). Aggregate Score is AggScore(π) = PopReturn(π) − PopExpl(π), balancing average returns with the worst loss within the evaluation population.

  • A policy can have high average return while remaining vulnerable to one opponent; Aggregate Score penalizes that vulnerability.
  • PopExpl is population-restricted, not exploitability computed against an unrestricted best response.

3.3. Baselines

The principal baseline is PSRO-IMPALA, whose best-response oracle trains an LSTM policy with the Importance Weighted Actor-Learner Architecture. RRPS also includes pretrained Gemma 3 agents that receive textual action histories, predict subsequent actions, and counter the predicted opponent move.

  • The turn-by-turn LLM baseline performs inference during play, whereas CSRO generates a reusable program before deployment.
  • Additional RRPS comparisons use previously reported Q-learning with recall length 10, Contextual Regret Minimization (ContRM), and a Chinchilla 70B LLM agent (Lanctot et al., 2023).

3.4. Implementation Details

The code-generation oracle is Gemini 2.5 Pro. Experiments use K = 20 outer iterations and M = 10 for LinearRefinement, with five seeds unless otherwise specified; Leduc results use three seeds. Supplementary Table 7 specifies three seeds for RRPS AlphaEvolve and five for other CSRO variants, although Table 1’s general caption states five seeds.

  • The implementation details refer to the published AlphaEvolve procedure but do not specify a numerical evolutionary-search budget.
  • Supplementary materials provide the prompts, initial poker strategy, and PSRO-IMPALA hyperparameter sweep.

4. Results

The quantitative results evaluate final meta-equilibrium strategies after K = 20, followed by inspection of selected generated programs. In RRPS, the lowest population exploitability and highest Aggregate Score belong to different CSRO variants. In Leduc, AlphaEvolve has both the highest average return and lowest population exploitability among the three CSRO variants.

4.1. Repeated Rock-Paper-Scissors

In RRPS, AlphaEvolve has the lowest mean PopExpl among tested CSRO variants, 25.2 ± 20.3, but a PopReturn of only 50.5 ± 1.9. LinearRefinement with code input and Top 5 filtering provides the highest mean CSRO Aggregate Score, 122.1 ± 9.8, with PopReturn 159.8 ± 7.7 and PopExpl 37.7 ± 10.6.

  • Gemma 3 27B achieves an Aggregate Score of 126.0; the previously reported ContRM and Chinchilla 70B results are higher at 148.5 and 155.2.
  • The selected CSRO variants in Table 1 outperform PSRO-IMPALA, whose PopReturn is −108.9 ± 17.6, PopExpl is 423.2 ± 28.0, and Aggregate Score is −532.1 ± 41.5.
  • The oracle optimizes expected utility against the current equilibrium mixture, not external-population return or a direct worst-case objective. The authors attribute AlphaEvolve’s lower external average return to the limited equilibrium support assigned to weak opponents.

The three panels separate average exploitation from worst-opponent vulnerability. The no-opponent-input variant has substantial average return but extreme population exploitability, producing a strongly negative Aggregate Score. Input representation, refinement, and filtering interact; no input/filter choice improves every metric across all variants.

Population Return, Within Population Exploitability, and Aggregate Score for RRPS oracle variants, opponent-input formats, and filtering choices.
Population Return, Within Population Exploitability, and Aggregate Score for RRPS oracle variants, opponent-input formats, and filtering choices.

LinearRefinement with code input has the highest listed CSRO Aggregate Score at 122.1 ± 9.8, close to the Gemma 3 27B agent’s 126.0. AlphaEvolve has the lowest listed CSRO population exploitability at 25.2 ± 20.3, while previously reported ContRM and Chinchilla 70B results have higher Aggregate Scores. The general five-seed caption conflicts with Table 7’s three-seed specification for AlphaEvolve.

Final RRPS population metrics for selected CSRO variants and baselines, with asterisks marking results reported by Lanctot et al. (2023).
Final RRPS population metrics for selected CSRO variants and baselines, with asterisks marking results reported by Lanctot et al. (2023).

Removing opponent input yields PopReturn 135.3 ± 10.2 but PopExpl 614.2 ± 60.8 and Aggregate Score −478.9 ± 70.2, showing that strong average exploitation can coexist with severe vulnerability. Input representation also matters: unfiltered ZeroShot descriptions achieve Aggregate Score 63.5 ± 11.4, whereas unfiltered ZeroShot code achieves −54.3 ± 118.7.

  • The authors suggest that textual summaries make one-pass synthesis more tractable than processing many source programs, but do not directly test that explanation.
  • Top 5 filtering generally outperforms minimum-support filtering in the reported comparisons. The proposed explanation is that including several opponents reduces narrow adaptation to a dominant equilibrium policy.
  • The detailed ablations vary substantially across input and filtering choices; refinement does not improve robustness under every configuration.

4.2. Repeated Leduc Hold’em Poker

Leduc experiments use opponent descriptions without filtering for all CSRO oracles. AlphaEvolve achieves PopReturn 49.3 ± 3.7, PopExpl 4.4 ± 0.6, and Aggregate Score 44.9 ± 4.1; LinearRefinement achieves 43.8 ± 2.5, 9.8 ± 3.0, and 34.0 ± 4.3; ZeroShot achieves 40.4 ± 1.6, 19.6 ± 2.1, and 20.7 ± 3.0.

  • CFR+ achieves PopReturn and Aggregate Score 39.8 ± 0.3 with PopExpl 0.0 ± 0.0. AlphaEvolve has higher average exploitation but remains vulnerable to CFR+.
  • PSRO-IMPALA obtains Aggregate Score −45.0 ± 10.1 and PopExpl 58.4 ± 3.3.
  • The worst-population returns arise against CFR+, and the stronger CSRO oracles reduce those losses.

Across the three CSRO variants in Leduc, stronger refinement is associated with higher mean return and lower population exploitability. AlphaEvolve exceeds CFR+ on Aggregate Score, 44.9 ± 4.1 versus 39.8 ± 0.3, but retains a worst-population loss of 4.4 ± 0.6, whereas CFR+ has 0.0 ± 0.0. The comparison separates exploitation of the tested heuristics from robustness against the equilibrium-oriented opponent.

Repeated Leduc hold’em population metrics for CSRO, PSRO-IMPALA, and CFR+, averaged over three seeds.
Repeated Leduc hold’em population metrics for CSRO, PSRO-IMPALA, and CFR+, averaged over three seeds.

Opponent-specific returns show how AlphaEvolve exceeds CFR+ on Aggregate Score: it earns 110.3 ± 9.7 against AlwaysCall, compared with 62.1 ± 0.8 for CFR+ and 57.7 ± 3.3 for PSRO-IMPALA. LinearRefinement earns 57.3 ± 8.8 against AlwaysFold, close to CFR+ at 57.4 ± 0.2 and above AlphaEvolve at 42.0 ± 3.2.

  • The oracles produce different trade-offs rather than one policy dominating every opponent-specific comparison.
  • Higher returns than CFR+ against predictable heuristics demonstrate exploitation, not stronger worst-case guarantees than an equilibrium strategy.

4.3. Qualitative Analysis

4.3.1. Repeated Rock-Paper-Scissors presents an ensemble program associated with a selected result of PopReturn 238.3, rather than the mean performance of the final mixtures. Its 32 predictors combine Markov models, reactive and joint-history models, periodic detectors, and meta-predictors. Votes are weighted by the fifth power of each predictor’s score.

  • Listing 1 exposes Markov orders 1 through 8, score decay 0.985, randomized tie-breaking, and the meta_imitation component.
  • The meta-predictor assumes that the opponent uses a model resembling one of the agent’s successful predictors, estimates the opponent’s counter-move, and feeds that prediction into the ensemble. This is an inspectable heuristic, not independently validated evidence of a psychological theory-of-mind capability.
  • The qualitative text inconsistently labels the configuration as “LinearRefinement (no filter) - Top 5.” Table 7 associates the maximum PopReturn of 238.3 with LinearRefinement, code input, and no filter; it does not separately establish the standalone listing’s score.

4.3.2. Repeated Leduc hold’em poker examines a selected policy reported to achieve PopReturn 77.8 and Aggregate Score 69.1. The program estimates showdown equity, tracks opponent action frequencies by public betting context, and compares approximate expected values of legal actions. Its raise calculation combines profit when the opponent folds with estimated value when the opponent continues.

  • The intended adaptation is value betting against AlwaysCall and frequent bluffing against AlwaysFold, driven by estimated folding probabilities.
  • Listing 2 exposes the approximations: opponent hand weights use fixed action-based heuristics, and raise EV explicitly ignores opponent re-raises.
  • The opening docstring describes a continuous “bravery” parameter, but the executable decision logic uses equity and EV calculations. Comments alone therefore do not establish policy behavior.

CSRO changes the best-response oracle rather than the PSRO meta-solver. It also differs from game-theoretic steering of dialogue models and Game of Thoughts by producing persistent executable policies rather than strategic prompts or directly generated actions.

  • The paper contrasts its oracle change with meta-solver work such as alpha-rank and training schemes such as Pipeline PSRO.
  • The closest comparison is LLM-PSRO (Bachrach et al., 2025), which already synthesizes code policies within self-play.
  • The stated distinctions from LLM-PSRO are intra-iteration feedback, context abstraction, and external evaluation—not the first use of code policies in self-play.

6. Conclusion and Discussion

The conclusion emphasizes a deployment distinction: an action-generating LLM baseline requires 1000 model calls per RRPS match, while a synthesized CSRO policy can play subsequent matches without turn-level LLM inference. Policy generation still requires oracle calls and simulation, with additional calls for refinement and, when used, opponent summarization.

  • The paper does not report matched wall-clock, monetary, or total simulation budgets, so this deployment distinction is not a measured end-to-end cost advantage.
  • The authors acknowledge likely pretrained knowledge of classic game strategies. Dynamically generated opponent mixtures require contextual response synthesis, but the contribution of prior knowledge is not isolated experimentally.

The stated limitations are dependence on LLM capability and prompt quality, potentially invalid generated code, repeated API costs, and unresolved scaling to high-dimensional games such as Stratego or StarCraft. Representing complex states and opponent strategies within the model’s context remains an engineering challenge.

  • Interpretability is illustrated through source inspection of selected policies, not a controlled human study or formal verification.
  • The experiments establish performance on two compact domains and restricted evaluation populations; they do not establish general convergence or unrestricted low exploitability.

Appendix

  • A. Supplementary materials and A.1. Hyperparameters document A.1.1. PSRO-IMPALA, including sweeps over learning rates, hidden layer sizes, unroll length, entropy cost, batch size, and maximum gradient norm. Selected RRPS settings are learning rate 0.0001, hidden layers [256, 128], unroll length 20, entropy cost 0.001, batch size 16, and maximum gradient norm 40, using observation_tensor(). Selected Leduc settings are learning rate 0.001, hidden layers [512, 256], unroll length 40, entropy cost 0.01, batch size 64, and maximum gradient norm 40, using infostate_tensor().
  • A.2. Prompts and Initial Strategies includes A.2.1. Repeated Rock-Paper-Scissors, specifying Agent.act and previous-action observations, and A.2.2. Repeated Leduc Poker, specifying receive_outcome, restart, act, JSON observations, opponent summaries, and code-editing templates. A.2.3. Repeated Leduc Poker contains the initial non-learning heuristic poker bot, generated by zero-shot prompting. The supplement also supplies full selected-policy implementations, Gemma 3 model-size comparisons, CSRO input/filter ablations, and opponent-specific poker returns.

Brief Thoughts

The contribution is executable policy synthesis coupled to population-level feedback and external evaluation. The no-opponent-input ablation provides direct evidence that opponent context matters for robust synthesis: high average return coexists with severe worst-opponent losses when that context is removed.

The evidence supports inspectable policy construction more strongly than general equilibrium solving or end-to-end computational efficiency. Broader poker opponent populations, matched search and simulation budgets, and tests of human understanding would help assess the scope of the advantages. The supplied code already makes heuristic approximations and mismatches between documentation and execution identifiable.