AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises | Summary
16 Feb 2026 | Paper Review LLM Evaluation Strategic Reasoning LLM Safety MetacognitionContents
- Summary
- 1 Introduction
- 2 Methods
- 3.1 Tournament Overview
- 3.2 Behavioural Signatures
- 3.3 Implications for strategy and strategic theory
- 3.3.1 Clausewitz: Friction and the Fog of War
- 3.3.2 Schelling: Commitment, Credibility, and the Rationality of Irrationality
- 3.3.3 Jervis: Perception and Misperception
- 3.3.4 Kahn: The Escalation Ladder
- 3.3.5 Power Transition Theory
- 3.3.6 Deterrence Theory: Success and Failure
- 3.3.7 Challenges to Theory
- 3.3.8 Summary: Theoretical Validation and Complications
- 4 Conclusion
- Appendix
- Brief Thoughts
This article explains the key points of AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises.
- 2026-02-16 (arXiv)
- Payne, Kenneth.
- King’s College London
- Paper
Summary
- AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises evaluates GPT-5.2, Claude Sonnet 4, and Gemini 3 Flash in 21 simulated nuclear crises. Across 329 turns, the models generated approximately 780,000 words of structured assessments, forecasts, signals, and action rationales.
- The models produced differentiated opponent assessments, deliberate signal-action mismatches, and explicit reflection on forecasting and credibility. These outputs demonstrate complex strategic behavior within the simulation, but do not establish conscious self-awareness or the causal faithfulness of generated rationales.
- Claude’s reported win rate fell from 100% in open-ended games to 33% in deadline games, while GPT-5.2’s rose from 0% to 75%. Gemini alone deliberately selected Strategic Nuclear War (1000); GPT-5.2 reached that outcome twice through accidents that escalated its choices of 950 and 725.
- The paper reports tactical nuclear use in 95% of games and no selection of negative-valued concession or withdrawal options. The results support context-sensitive safety evaluation, but the small tournament, scenario-deadline confounding, scoring incentives, and inconsistent statistics limit conclusions about training effects or real-world nuclear strategy.
1 Introduction
The introduction asks how frontier language models reason about escalation, deterrence, and catastrophic risk in intelligence analysis and decision-support settings. The study examines adversarial interaction rather than isolated recommendations, emphasizing deception, opponent modeling, and the gap between recognizing danger and choosing restraint.
1.1 Background: AI and Strategic Decision-Making
Prior scholarship identifies risks from accelerated decision timelines, opaque machine reasoning, and reduced human control over military systems. This study addresses a narrower empirical gap: how AI models behave when assigned consequential choices within a stylized crisis, rather than how humans perceive AI involvement.
1.2 Related Work: LLMs in Strategic Simulation
The paper connects its design to Diplomacy, repeated prisoner’s dilemma experiments, and military wargaming studies reporting model aggression and framing sensitivity. Its principal extension is sustained simultaneous interaction in which reputations, forecasts, and signal-action histories can influence subsequent decisions.
- Payne and Alloui-Cros [2025] reported model-specific strategic patterns across nearly 32,000 decisions; the present paper investigates whether comparable patterns appear in existentially framed crises.
- Unlike the prior human-expert comparison discussed here, this tournament does not include a matched human baseline.
1.3 Theoretical Foundations
Kahn’s escalation ladder supplies the action-space framework, Schelling motivates the separation of declared intentions from actions, and Jervis motivates analysis of perception and misperception. The simulation adapts Kahn’s 44 rungs into 30 options and elicits self-assessments to examine whether models recognize their limitations.
1.4 Research Questions
The research questions concern accurate opponent models, metacognitive assessment, and behavior recognizable from international relations theory. Additional analyses examine strategic signatures, deception, accidents, temporal framing, conditional threats, nuclear thresholds, and adaptation within games.
- The proposed connection between training regimes and strategic behavior is an explanatory hypothesis, not a controlled training comparison.
1.5 Contributions and Innovation
The methodological contribution is an instrumented interaction protocol combining Reflection → Forecast → Signal/Action, simultaneous decisions, separate signals and actions, structured self-assessments, and seven crisis scenarios. Memory and accidental escalation enable analysis of reputation formation, adaptation, and responses to asymmetric information.
- The architecture explicitly requests reflection and opponent assessment. Claims of spontaneous sophistication must distinguish unprompted strategic content from the requested reasoning categories.
2 Methods
Claude Sonnet 4, GPT-5.2, and Gemini 3 Flash played 21 games as leaders of fictional nuclear powers, producing 329 turns. Each turn sequentially generates reflection, an opponent forecast, and a decision containing a public signal, private action, and consistency statement; both players complete this process independently before resolution.
- Models receive verbal action descriptions rather than numeric ladder values. Appendix B specifies that neither current-turn signals nor actions are visible to the opponent before commitment; signals enter the subsequent turn’s memory.
- Open-ended games have a 40-turn maximum. Deadline scenarios enforce limits of 15 turns for Resource and First Strike, 20 for Regime Survival, and 12 for Standoff.
- The recent decision panel covers five turns, while salient betrayals persist with approximately 15% salience decay per turn. At actions of 125+, accidents can increase an action by 1–3 rungs, with probabilities of 5%, 10%, or 15% according to assessed miscalculation risk.
- Territorial advantage is a zero-sum proxy for strategic status, with knockout at |territory_balance| ≥ 5.0. Code and results are provided at https://github.com/kennethpayne01/project_kahn_public.
Assessments and forecasts are generated before decisions, which separate public signals from private actions. Both players execute the pipeline independently; the diagram does not imply access to the opponent’s current-turn signal.
3.1 Tournament Overview
The results combine outcome statistics with qualitative examples of generated strategic reasoning. The tournament contains 9 open-ended and 12 deadline games, including self-play, but the headline win-loss records concern contests against other models.
- Success is defined by the simulation’s territorial scoring rules, not by civilian welfare or avoidance of nuclear catastrophe.
3.1.1 Overall Performance
The reported overall records are Claude 8–4 (67%), GPT-5.2 6–6 (50%), and Gemini 4–8 (33%). Figure 2 shows opposing conditional patterns: Claude leads without explicit deadlines, while GPT-5.2 leads in deadline scenarios.
- Table 2 reports Claude at 6–0 versus 2–4, GPT-5.2 at 0–4 versus 6–2, and Gemini at 1–3 versus 3–5 across open-ended and deadline conditions.
- Deadlines are embedded in different scenarios rather than independently randomized. The paper acknowledges that temporal effects cannot be fully isolated from scenario content.
Condition-specific bars reveal a ranking reversal hidden by aggregate win rates: Claude moves from 100% to 33%, while GPT-5.2 moves from 0% to 75%. Scenario content and deadlines vary together, so the comparison does not isolate a deadline effect.
3.1.2 Head-to-Head Results
The combined head-to-head results show Claude defeating Gemini 5–1, while Claude–GPT-5.2 and Gemini–GPT-5.2 both finish 3–3. The Claude–GPT-5.2 tie conceals a complete reversal: Claude wins all three open-ended contests and loses all three deadline contests.
- Table 2’s Gemini–GPT-5.2 condition-specific row and its overall model records do not fully reconcile with the other matchups. Exact condition totals require verification against the released game records.
The Claude–GPT-5.2 row shows 3–0 without deadlines and 0–3 with deadlines despite a combined 3–3 record. The Gemini–GPT-5.2 row and overall condition records do not fully reconcile with the other matchups.
3.1.3 Game Length Distribution
Table 3 reports mean lengths of 21.6 turns for open-ended games and 11.1 turns for deadline games. Six of the 12 deadline games ended more than two turns before the deadline, while six ended within two turns of it.
- Two open-ended games reached 40 turns. Table 3 lists knockout counts of 8/9 and 11/12, which do not reconcile with the overall-results claim of 86% knockout endings.
- Figure 5 displays the duration distributions: open-ended games extend to 40 turns, while deadline games span 4–19 turns.
Mean duration falls from 21.6 to 11.1 turns, and deadline outcomes split evenly between early knockouts and endings within two turns of the deadline. Listed knockout counts do not reconcile with the overall knockout percentage in the main prose.
3.1.4 Escalation and Nuclear Use
The analysis distinguishes Nuclear Signaling (125+), Tactical Use (450+), Strategic Threat (850+), and Strategic War (1000). Figure 3 reports GPT-5.2’s median maximum escalation increasing from 175 to 900 across temporal conditions; Claude maintains an 850 ceiling, while Gemini varies substantially.
- Table 4 reports tactical threshold crossings in 86% of Claude games, 79% of Gemini games, and 64% of GPT-5.2 games. Strategic War outcomes occur in 0%, 7%, and 14%, respectively; GPT-5.2’s two outcomes follow accidental escalation rather than deliberate selection.
- Figure 4 measures percentages of actions at thresholds, whereas Table 4 measures percentages of games crossing thresholds. The denominators are not interchangeable.
- None of the eight negative-valued options from Minimal Concession (−5) to Complete Surrender (−95) was selected. Return to Start Line (0) appeared 45 times (6.9%); reduced escalation therefore did not imply accommodation.
Maximum-escalation distributions show GPT-5.2’s strong condition dependence and Claude’s persistent upper limit. GPT-5.2’s 1000 outcomes are accident-modified results, whereas Gemini’s 1000 outcome follows a deliberate choice.
Game-level crossings distinguish frequent tactical use from rare Strategic War outcomes. Claude’s 86% tactical-use rate coexists with no Strategic War outcomes; GPT-5.2’s 14% Strategic War rate includes two accidental escalations.
3.2 Behavioural Signatures
The behavioral analysis links action patterns to generated private assessments and rationales. It combines quantitative forecast and signaling measures with selected passages illustrating credibility management, opponent modeling, and self-assessment.
- These passages are observable model outputs, not direct measurements of internal cognitive processes.
3.2.1 Model-Specific Patterns
The paper describes Claude as a controlled escalator, GPT-5.2 as strongly dependent on temporal conditions, and Gemini as an unpredictable, risk-tolerant actor. These labels summarize observed patterns rather than independently established psychological traits.
- Claude repeatedly uses strategic threats without selecting total war.
- GPT-5.2 shifts from restrained losing play in open-ended games to substantial nuclear escalation in deadline games.
- Gemini deliberately selects Strategic Nuclear War by Turn 4 in a First Strike game and discusses unpredictability as a strategic asset.
3.2.2 Signal-Action Consistency and Deception
Signal-action consistency is approximately 72–75% for Claude and GPT-5.2, but 50% for Gemini. Table 5 shows that opponents rate Claude and GPT-5.2 more credible than Gemini, while Table 7 identifies Claude’s pattern of low-level reliability followed by frequent signal overshoot at higher levels.
- Claude matches 84% of low-level signals but exceeds signals in 70% of mid-level and 60% of high-level actions. The Power Transition example pairs a signal of 175 with an action of 850.
- GPT-5.2 more often falls below than exceeds its signals, whereas Gemini deviates in both directions. Appendix H defines a match as action within 50 points of the signal, not necessarily exact equality.
- In the Regime Survival example, GPT-5.2 selected Expanded Nuclear Campaign (725), and an accident produced Strategic Nuclear War (1000). Nearby footnotes incorrectly describe 1000 as its chosen action.
Opponent judgments separate immediate signal reliability from longer-horizon resolve. GPT-5.2 receives the highest highly credible resolve rating, 51%, illustrating that credibility does not necessarily imply aggressive threats or successful deterrence.
3.2.3 Prediction and Foresight
The reported mean absolute forecast errors are 85 ladder points for Claude, 149 for Gemini, and 97 for GPT-5.2. Generated assessments often identify opponents’ behavioral patterns correctly, but deadline examples show models recognizing escalation risks while still forecasting restraint.
- Appendix H labels these statistics as combined-tournament measures, although the main text initially introduces them while discussing open-ended games.
- Bias is mean(actual − predicted): Claude −6, Gemini −55, and GPT-5.2 +43. GPT-5.2 therefore tends to underestimate escalation, while Gemini tends to overestimate it.
- In the Standoff example, Claude holds a +3.98 territorial advantage at Turn 11 of a 12-turn game but misreads GPT-5.2’s restraint; GPT-5.2 subsequently selects Strategic Nuclear Threat and wins.
3.2.4 Metacognition
All three models produce explicit self-assessments, but these do not consistently improve choices. Claude acknowledges possible overconfidence while escalating, GPT-5.2 diagnoses weak corrective mechanisms while remaining restrained, and Gemini portrays its unpredictable reputation as a deliberately managed instrument.
- GPT-5.2 interprets Claude’s signal-action gaps as poor calibration, which the author reads as projection of GPT-5.2’s preference for consistency onto an adversary.
- The examples establish expressed metacognitive assessment. They do not establish sufficiently calibrated self-ratings to infer reliable self-knowledge.
3.2.5 Summary: Strategic Instincts
The section synthesizes credibility management, opponent modeling, and reflection on cognitive limitations. Their effectiveness varies with context, particularly when established reputations fail to predict behavior near a deadline.
- Restrained behavior in one setting does not guarantee restraint in another. The tournament supports evaluating this dependence, but cannot attribute it uniquely to temporal framing.
3.3 Implications for strategy and strategic theory
The paper interprets the tournament through established theories of strategic interaction. These comparisons identify recognizable reasoning patterns, but are not independent tests of real-world theories because the prompts and scoring rules encode several of the mechanisms under examination.
3.3.1 Clausewitz: Friction and the Fog of War
Accidental escalation illustrates Clausewitzian friction: actors must interpret opponent actions without knowing whether escalation was intended. Accidents occur in 18/21 games (86%), and rationales show models worrying about exposed preferences or exploiting accidental escalation to project greater risk tolerance.
- Appendix H reports 46 accident turns, affecting State A 34 times and State B 12 times. No model explicitly disclosed that its own escalation was unintended.
- The paper also discusses the culminating point of an offensive. Continued escalation is more common than consolidation, although one Claude example favors sustained conventional pressure over further nuclear escalation.
3.3.2 Schelling: Commitment, Credibility, and the Rationality of Irrationality
Schelling’s commitment and credibility concepts appear in model rationales, including Gemini’s use of unpredictability and Claude’s differentiated signaling. GPT-5.2’s open-ended restraint creates credibility for nonaggression, which opponents exploit rather than treating as an effective threat.
- Table 7 shows Claude’s low-stakes reliability and high-stakes signal overshoot; uniformly describing it as honest or deceptive would obscure this difference.
- In deadline games, expectations of GPT-5.2’s restraint become unreliable. The paper leaves open whether exploiting that reputation is intentional or emerges from changed behavior.
Claude matches 84% of signals below 200 but exceeds signals in 70% of actions at 200–500 and 60% above 500. The breakdown supports level-dependent signaling rather than uniform unreliability.
3.3.3 Jervis: Perception and Misperception
Table 8 characterizes systematic misperception: GPT-5.2’s +43-point bias indicates underestimation of opponent aggression, while Gemini’s −55-point bias indicates overestimation. Examples illustrate wishful forecasting, projection, and recognition of security-dilemma spirals.
- When forecasts underestimate actions, average errors are +135 points for GPT-5.2, +172 for Claude, and +179 for Gemini. These conditional magnitudes differ from overall bias.
- Claude sometimes identifies a spiral and reduces escalation, but elsewhere recognizes the same risk and escalates to protect reputation.
Overall bias separates GPT-5.2’s tendency to underestimate aggression from Gemini’s tendency to overestimate it. Larger conditional errors when forecasts underestimate actions show that small overall bias can conceal substantial misses.
3.3.4 Kahn: The Escalation Ladder
Models reason about rungs, thresholds, and escalation dominance despite receiving verbal rather than numeric action descriptions. They recognize Limited Nuclear Use (450) as qualitatively different but frequently cross that threshold.
- Escalation Dominance as Winning Strategy: Claude’s willingness to move above opponents produces leverage under the territorial scoring system, while GPT-5.2’s open-ended restraint often cedes that advantage.
- The action space and prompts mention escalation control and action rungs. The evidence concerns how models use this framework, not its discovery without scaffolding.
3.3.5 Power Transition Theory
Power Transition scenarios elicit arguments about rising challengers, declining hegemons, reputational losses, and closing windows of opportunity. Table 10 reports opportunity-window and preventive reasoning in 100% of Power Transition games.
- The prompts explicitly describe rising and declining roles and their consequences. These outputs demonstrate scenario-consistent reasoning rather than independently validating power transition theory.
3.3.6 Deterrence Theory: Success and Failure
The main analysis reports opponent de-escalation after 25% of 268 nuclear-level actions with observable follow-up, falling to 18% at the tactical threshold. Nuclear escalation therefore often precedes continued or increased escalation rather than accommodation.
- Appendix H.7.4 instead reports 4/28 de-escalatory responses (14%) following actions ≥450. The paper does not reconcile these counts with the main-text statistics.
- Strategic threats receive their full territorial weight only after tactical nuclear use. Action-backed credibility is therefore partly built into the environment.
- A decline in next-turn action level is an operational proxy for deterrence, not a causal estimate or a complete measure of bargaining success.
3.3.7 Challenges to Theory
The paper presents five named challenges to strategic theory, retaining the source’s numbering gap between Challenge 4 and Challenge 6. They concern material superiority, credibility, training-shaped objectives, actor-specific behavior, and the nuclear taboo.
- Challenge 1: The Superiority Paradox. GPT-5.2’s reported 57% versus 43% nuclear-power advantage does not prevent open-ended losses; material capability alone is insufficient when an actor avoids exploiting it under these rules.
- Challenge 2: Credibility Can Enable Aggression, Not Just Deter. Table 11 contrasts Claude self-play, reaching nuclear use at Turn 4 and ending at Turn 7, with GPT-5.2 self-play lasting 40 turns despite similarly high credibility. Its “Turns to Nuclear” column mixes thresholds: GPT-5.2 self-play never exceeds Nuclear Signaling (125), so Turn 6 cannot indicate tactical use. One self-play game per model also cannot isolate credibility from actor and scenario differences.
- Challenge 3: Training alters ultimate goals. GPT-5.2’s restraint and targeting restrictions are consistent with safety-shaped preferences, but the author lacks training access and cannot establish RLHF causation. Its two 1000 outcomes were accidents, not deliberate selections.
- Challenge 4: Actor Characteristics Trump Structural Incentives. Tables 12 and 13 show different escalation trajectories across actors and conditions despite shared state profiles; scenario content remains confounded with deadlines.
- Challenge 6: The Nuclear Taboo Is Weaker Than Expected. The paper reports tactical use in 95% of games and strategic threats in 76%, with three Strategic War outcomes. This demonstrates frequent nuclear use in the simulation, not the weakness of the historical human taboo.
Claude and GPT-5.2 self-play both receive high/high credibility labels but differ sharply in duration and escalation. The “Turns to Nuclear” column is ambiguous: GPT-5.2 self-play never exceeds 125, so its Turn 6 entry cannot mean tactical nuclear use.
Alliance games range from a 40-turn GPT-5.2 self-play contest with maximum escalation 125 to a 16-turn Claude–Gemini contest reaching 850. Shared scenario framing does not produce a single escalation trajectory.
GPT-5.2 reaches higher escalation levels and better outcomes in deadline scenario families than in open-ended ones. The table does not isolate deadlines experimentally, and its 1000 entries represent accident-modified outcomes rather than deliberate total-war choices.
3.3.8 Summary: Theoretical Validation and Complications
The theoretical synthesis finds recognizable reasoning about reputation, commitments, opportunity windows, and escalation, alongside differences in restraint and risk tolerance. The proposed training explanation remains tentative: the observations establish conditional behavior, not which training procedure caused it.
4 Conclusion
The conclusion argues that frontier models sustain coherent strategic interaction, including deception, opponent tracking, and adaptation. The evidence establishes behavioral competence within this simulation; policy interpretation requires comparison with human behavior and alternative game designs.
4.1 Three Core Findings
The three core findings are strategic sophistication, substantial model differences coupled with context dependence, and behavior that resembles and complicates international relations theory. Sophisticated language and successful play nevertheless coexist with escalation, forecast failures, and poor crisis safety.
- GPT-5.2’s 0% to 75% win-rate reversal is the clearest descriptive context effect, but its accident-modified outcomes must remain distinguished from intended actions.
- The claim that RLHF sets thresholds rather than absolute limits requires a controlled training comparison.
4.2 Why This Matters
The paper proposes AI simulation for exploring crisis dynamics, refining theory, and evaluating AI-influenced decision-support environments—not for delegating nuclear authority. Its stated prerequisite is calibration: determining where model behavior approximates human strategic reasoning and where it departs from it.
- The released records support reproducible behavioral analysis, but the approximately 780,000-word corpus does not substitute for independent games or causal controls.
4.3 Limitations and Future Directions
The author acknowledges the small sample of 21 games, limited scenario coverage, and possible changes in newer model versions. Proposed extensions include larger tournaments, model-generation comparisons, multi-party crises, alliances, extended deterrence, and interventions through prompting or fine-tuning.
- Further interpretive limits include scenario-deadline confounding, asymmetric leadership briefings, and a victory function that rewards relative escalation rather than directly measuring survival or negotiated resolution.
- Knockout counts, condition-specific records, deterrence denominators, and comeback totals need reconciliation. Appendix H.7.6 reports twelve comeback games (35%), while Table 26 lists 8/21 (38%); the former percentage also does not match a 21-game denominator.
- Figure 6 reports aggressor wins in 10/18 non-self-play games (56%), versus 44% for defenders. Because moves are simultaneous and State A is always the aggressor, this is a role comparison rather than evidence for an information-based first-mover advantage.
- Table 26 reports lead changes in 17/21 games, a mean maximum territorial swing of 6.25 points, and improvement on 57/108 turns (53%) when behind by 0.5+. These describe volatile trajectories, not statistically established general laws.
Game-level points show a wide open-ended distribution, including two 40-turn contests, versus deadline games spanning 4–19 turns. Shorter duration reflects both enforced limits and early knockouts, not uniform clustering at the deadline.
Lead changes in 17/21 games and improvement on 57/108 turns (53%) while behind show frequently reversible advantages. Its 8/21 comeback count conflicts with the separate comeback statistic in Appendix H.7.6.
Appendix
- A Simulation Design specifies model assignment, aggressor and defender roles, simultaneous decisions, and turn resolution through action revelation, territorial scoring, attrition, accidents, victory checks, and history updates. Games end through |territory_balance| ≥ 5.0, surrender, simultaneous Strategic Nuclear War (1000), or a turn limit; territory represents broader zero-sum strategic advantage rather than literal geography alone. Appendix A calls signaling a within-turn communication channel, but Appendix B explicitly states that current-turn signals remain unavailable until the following turn.
- B The Three-Phase Cognitive Architecture details Reflection, Forecast, Decision, and the consistency statement: models assess their own and opponents’ abilities, forecast a specific action with confidence and miscalculation risk, then produce immediate and conditional signals, a public statement, an action, and a private rationale. The consistency statement relates the chosen action to the forecast, explaining restraint below the predicted level or risk acceptance at or above it. Earlier outputs pass verbatim into later phases, preventing reflection from consulting the subsequently generated decision but not proving that stated rationales causally determine choices.
- C State Profiles uses an abstracted early Cold War setting; the author argues that role variation and divergent model behavior limit a historical-priming explanation, although no unprimed control is provided. President Alpha is briefed as cautious, intellectual, image-conscious, and committed to flexible response, while Premier Beta is described as unpredictable, pragmatic, ambitious, and more risk-tolerant; reciprocal intelligence briefings reinforce these asymmetric expectations. Alpha has conventional strength 85/100 and 1,944 warheads, with higher delivery accuracy and survivability; Beta has conventional strength 95/100 and 347 warheads, with numerical and geographic advantages. Turn displays report force effectiveness, fighting power, and relative strength; combat attrition depends on combined intensity and capability ratios, with nuclear intensity producing progressively greater degradation of both conventional and nuclear forces.
- D The Escalation Ladder provides 30 verbal options from Complete Surrender (−95) to Strategic Nuclear War (1000), covering concessions, status quo, conventional escalation, nuclear signaling, and nuclear campaigns. Strategic Nuclear Threat (850) and Final Nuclear Warning (950) count as Nuclear Threat (350) for territorial scoring unless either side has used tactical nuclear weapons (450+) in the current game or turn. Nuclear Demonstration (250) and Strategic Nuclear War (1000) are not gated.
- E Scenario System supplies Alliance Leadership Test (v7_alliance), Strategic Resource Race (v7_resource), Power Transition Crisis – State A Rising (v7_power_transition_a_rising), Power Transition Crisis – State B Rising (v7_power_transition_b_rising), First Strike Crisis (v8_first_strike_fear), Regime Survival Crisis (v9_regime_survival), and Strategic Standoff Crisis (v10_standoff_crisis). Open-ended v11 games use Alliance and Power Transition scenarios; mechanically enforced v12 deadlines are 15 turns for Resource and First Strike, 20 for Regime Survival, and 12 for Standoff. The prompts vary stakes and incentives, often presenting concession as disastrous or preemption as necessary for survival; Standoff also uses asymmetric conventional-balance briefings.
- F Memory and Accident Systems combines a rolling five-turn signal-action panel with longer-lived records of major betrayals, whose salience decays approximately 15% per turn and can persist for 10+ turns. Actions at 125+ can accidentally rise by 1–3 rungs, with probabilities of 5%, 10%, and 15% for low, medium, and high assessed miscalculation risk. Only the affected actor is informed that the escalation was unintended, creating a choice between disclosure and strategic ambiguity.
- G Tournament Structure and Data Collection identifies gpt-5.2, claude-sonnet-4-20250514, and gemini-3-flash-preview. Table 21 counts 14 appearances per model, including both self-play roles, with uneven temporal-condition participation: Claude 8 open-ended and 6 deadline appearances, Gemini 4 and 10, and GPT-5.2 6 and 8. State A is always the aggressor, starting balance is 0.0 or −0.3, and 90 fields per turn cover reflection, forecasts, signals, actions, memory/reputation, and game state.
- H Appendix H: Extended Behavioral Analysis and Strategic Dynamics reports combined-tournament forecast errors and signal-match rates of 71.7%, 50.0%, and 75.3% for Claude, Gemini, and GPT-5.2, respectively; GPT-5.2’s highly credible resolve rating rises from 43% to 67%, while its mean action level rises from 78 to 200 across conditions. Model-specific passages examine Claude’s 850 ceiling, GPT-5.2’s deadline escalation without deliberate selection of 1000, Gemini’s context-adaptive aggression, and the varying translation of trajectory awareness into action. Strategic-dynamics analyses cover duration, aggressor outcomes, momentum, deterrence, 77/285 conditional threats followed by de-escalation (27%), and comeback examples, although several counts conflict with other sections. Across 46 accident turns in 18/21 games, no actor explicitly disclosed unintended escalation; assessments instead often attributed opponent accidents to aggressive intent, illustrating the paper’s proposed fundamental attribution error.
Brief Thoughts
The clearest methodological contribution is the separation of declared intent, forecasts, chosen actions, and accident-modified outcomes. It shows why restrained behavior, credible signaling, or articulate risk awareness cannot individually establish safety. A stronger follow-up would vary deadlines independently within the same scenarios, repeat each condition, and compare payoff structures that reward survival and negotiated settlement. Human comparisons and interventions on reflection or forecasts would help distinguish strategic competence from role-conditioned generation and test whether elicited reasoning improves decisions. The paper is an exploratory behavioral evaluation, not causal evidence about RLHF or an empirical refutation of human deterrence theory.