Gorio Tech Blog search

Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning | Summary

|

Contents

This article explains the key points of Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning.

  • 2026-06-11 (arXiv)
  • Yang, Jiayu, Chen, Chao, Wu, Shengen, Liu, Yinhong, Fan, Yuxuan, Li, Lujundong, Lai, Songning, Qin, Chengwei, Guo, Zhijiang.
  • HKUST(GZ), University of Cambridge, NTU, JoinQuant, HKUST
  • Paper

Read this article in Korean


Summary

  • Switch uses <swi> and </swi> to delimit continuous hidden-state recurrence within visible reasoning. These discrete boundaries provide policy probabilities at switching decisions and identifiable locations for probing and intervention.
  • Training combines entropy-based boundary annotation, a visible-to-latent supervised curriculum, and Switch-GRPO rollouts that execute hidden-state recurrence. The implemented RL backward pass differentiates through text segments only; latent segments run under torch.no_grad() and condition subsequent text through a frozen KV cache.
  • With Qwen3-8B, the strongest run reaches 79.3% on MATH-500 and 89.2% on GSM8K, exceeding the re-implemented recurrent-latent baselines. Switch nevertheless retains 1,721 and 1,608 visible tokens per problem, respectively, whereas those baselines use approximately 10; the comparison does not establish superiority at matched inference cost.
  • Boundary probabilities and linear probes support a localized switching feature, while interventions establish answer sensitivity on a selected diagnostic subset. Evidence for problem-specific latent content and concentration of useful computation at a single entry transition is less decisive: same-norm random replacement is relatively benign, and no step-specific causal ablation is reported.

1. Introduction

Hidden-state recurrence replaces a decoded reasoning token with the previous step’s final hidden state as the next input embedding. The paper identifies two obstacles: continuous latent positions do not supply categorical action probabilities for standard policy-gradient training, and their execution lacks explicit boundaries for causal inspection.

Switch addresses both obstacles with learned entry and exit tokens rather than an additional latent architecture. The public implementation is https://github.com/LARK-AI-Lab/SWITCH, and the released model weights are https://huggingface.co/LARK-Lab/SWITCH-Phase3-GRPO-LoRA-Qwen3-8B.

The paper distinguishes deterministic hidden-state recurrence, used by Coconut and CODI, from vocabulary-embedding mixtures. Recent mixture-based RL methods introduce sampling through Gumbel-Softmax; Switch instead preserves recurrence and adds discrete switching decisions, unlike training-free SwiReasoning’s external entropy rule.

  • Vocabulary-mixture methods are excluded from the headline comparison, so the results do not establish an advantage over the broader family of RL-trained latent reasoners.
  • Pause-style tokens and adaptive visible reasoning provide related compression or compute-allocation mechanisms but do not use the same recurrent latent representation.

3. Method

Three phases share the same boundary interface. Phase 1 learns to mark selected visible CoT spans, Phase 2 progressively replaces their contents with latent positions, and Phase 3 applies Switch-GRPO to trajectories containing hidden-state injection; the first two phases together are called Switch-SFT.

The diagram connects the three training phases to the deployed decoder. It distinguishes sampled text and boundary tokens, which carry policy probabilities, from recurrent latent inputs, which are deterministic hidden-state injections.

Three-phase training and interleaved text–latent execution in Switch.
Three-phase training and interleaved text–latent execution in Switch.

3.1. Switchable Latent Reasoning

The vocabulary adds <swi>, </swi>, and <latent>. Text positions use ordinary token embeddings, whereas a latent position receives the preceding final-layer hidden state: ẽ_t = E[x_t] for x_t ≠ <latent>, and ẽ_t = h_{t−1} for x_t = <latent>.

  • Every recurrent latent step requires its own model forward pass; omitting a decoded token does not eliminate its computational cost.
  • At inference, the decoder enforces a minimum latent dwell (K_{\min}) before permitting exit. Without this constraint, the trained model tends to exit after one hidden forward pass.

3.2. Switch-SFT: Curriculum Study

Phase 1 follows the SwiReasoning annotation procedure: contiguous high-Shannon-entropy regions of mathematical CoT are wrapped in <swi>/</swi>, then learned with response-only next-token cross-entropy. Phase 2 masks latent labels but retains supervision on the boundaries and surrounding text.

The default parallel curriculum changes every marked span simultaneously, using (n_m^{(k)}=c\cdot\min(k, S_m ,K_{\max})), with c=2 and K_max=8. A sequential alternative converts spans from left to right; the authors report that parallel replacement works better, but the appendix supplies no numerical head-to-head comparison.
  • The authors interpret parallel replacement as making it harder for visible text to compensate for ineffective latent blocks. This explanation is not separately tested.
  • Implementation details specify up to 16 latent positions per span, a 48-position cap per sample, curriculum smoothing probability 0.1, and three epochs at each stage.

Sequential replacement changes the leftmost spans first, whereas parallel replacement introduces latent computation across every marked span at each stage. The illustration explains the schedule difference but contains no quantitative comparison of their performance.

Sequential and parallel visible-to-latent curriculum schedules.
Sequential and parallel visible-to-latent curriculum schedules.

3.3. Switch-GRPO: Latent Exploring

Switch-GRPO executes the recurrent latent decoder during rollout rather than substituting text-only execution. Since latent injections are deterministic rather than sampled actions, trajectory likelihoods factor over discrete text positions, including <swi>, </swi>, and visible continuation tokens; latent positions receive no direct policy-gradient term.

  • Appendix B specifies a gradient pass conditioned on the rollout’s frozen input-embedding sequence, with group-relative advantages, a PPO-style clipped surrogate, and a KL penalty anchored to the old policy.
  • The implemented segmented backward pass does not backpropagate through latent recurrence. Shared model parameters still receive updates from text segments, so this is not equivalent to freezing all parameters used during latent execution.

The reward combines correctness, tag format, latent usage, and an optional correctness-gated brevity bonus. Section 3.3 describes the latent-usage reward as {0, 1}, whereas Appendix B defines r_use = r_corr · 1_wf · 1_used, which is −1 for an incorrect, well-formed response that uses a latent block.

  • The brevity term is gated on correctness and latent usage, with T_lo=800 and T_hi=2000; the compression configuration assigns it weight 0.1.
  • The PDF does not give the numerical weights for the other reward components.

4. Experiments

The experiments examine benchmark accuracy, changes after RL, and an accuracy–visible-length operating curve. Headline comparisons use the strongest end-to-end run, while checkpoint trajectories, efficiency analyses, and mechanistic experiments use a representative run with detailed logs.

4.1. Experimental Setup

All headline methods are re-implemented on Qwen3-8B using the same OpenR1-Math corpus and stated matched decoding settings. Evaluation uses MATH-500 and GSM8K; references include direct answers, text-CoT SFT, iCoT, Pause Tokens, Coconut, CODI, and CoLaR.

  • Training uses one node with 8×NVIDIA H20 GPUs, each with 95GB memory, under PyTorch DDP.
  • Switch-GRPO uses G=5 rollouts per question, clip threshold 0.2, KL coefficient 10^−3, learning rate 10^−6, and three inner epochs per rollout.
  • Appendix A specifies temperature 0.5 and K_min=4 for the main rollout configuration, versus temperature 0.7 and K_min=2 for compression. Table 8 nevertheless lists K_min=4 for the brevity checkpoint; maximum generation lengths vary among 2048, 4096, and 6000 across configurations.

4.2. Main Results

The strongest Switch-GRPO run reaches 79.3% MATH-500 accuracy versus CoLaR’s 53.6%, a 25.7-point difference, and 89.2% GSM8K accuracy versus CoLaR’s 78.5%. Text-CoT SFT reaches 80.6% and 88.6%, respectively: Switch is slightly below it on MATH-500 and slightly above it on GSM8K.

  • Switch-GRPO averages 1,721 visible tokens on MATH-500 and 1,608 on GSM8K, versus text-CoT’s 2,079 and 1,819.
  • CoLaR averages only 11.8 and 10.6 visible tokens. The accuracy advantage over the compressed recurrent-latent baselines is accompanied by substantially more visible reasoning.
  • The reported token metric excludes latent forward passes; no matched total-compute or wall-clock comparison is provided.

Switch-GRPO exceeds the selected recurrent-latent baselines in accuracy but uses roughly two orders of magnitude more visible tokens. Relative to text-CoT SFT, it has lower MATH-500 accuracy, slightly higher GSM8K accuracy, and moderately shorter visible responses.

Accuracy and average visible-token counts on MATH-500 and GSM8K with Qwen3-8B.
Accuracy and average visible-token counts on MATH-500 and GSM8K with Qwen3-8B.

Table 2 reports latent-conditional accuracy increasing from 66.7% to 79.3% after RL, with switching declining from 81% to 58% under K_min=4 greedy decoding. This is a 23-percentage-point reduction in switch rate, not a halving, and the conditional accuracies concern different selected problem sets because invocation changes.

  • Appendix Table 8 gives the representative run’s step-800 overall accuracy as 72.6%, latent-conditional accuracy as 67.8%, and mean length as 1,919 tokens. These do not match Table 2’s 79.3% latent accuracy and 1,777 tokens.
  • The PDF distinguishes the representative run from the strongest headline run, but does not reconcile that distinction with Table 2’s stated matched before/after comparison.
  • Table 8 reports 78.0% overall accuracy at step 200, above the step-800 endpoint. RL improves over the representative SFT checkpoint, but later optimization does not improve accuracy monotonically.

The representative step-800 endpoint reaches 72.6% overall accuracy and 67.8% latent-conditional accuracy, rather than Table 2’s reported 79.3% latent accuracy. Step 200 reaches 78.0% overall accuracy, and the brevity point reaches 69.0% at 1,276 visible tokens; its listed K_min=4 differs from Appendix A’s compression setting of 2.

Representative-run checkpoint metrics across curriculum, RL, and brevity training.
Representative-run checkpoint metrics across curriculum, RL, and brevity training.

The early training trajectory shows fewer switches, fewer latent blocks, and shorter outputs as RL proceeds. Figure 3’s start-to-end values are 69%→53% switching, 1.53→0.89 blocks per problem, and 2,839→1,702 tokens; average reward also decreases from +0.08 to +0.01, so this plot alone does not establish improving correctness.

In the representative run, the brevity operating point changes MATH-500 accuracy from 72.6% to 69.0% and mean visible length from 1,919 to 1,276 tokens, with truncation falling from 18.4% to 0%. The histogram shows a shift toward shorter responses rather than merely a smaller mean.

  • This is a 3.6-point accuracy loss for approximately 33% fewer visible tokens.
  • Appendix A also changes rollout temperature and minimum dwell for compression, so the configuration is not documented as a reward-only ablation; Table 8’s dwell setting remains inconsistent with that description.
  • Figure 4 labels its dashed curve an empirical Pareto frontier, although step 200 has both higher accuracy and fewer tokens than the step-800 endpoint.

The brevity point has shorter responses at an accuracy cost relative to the endpoint. Step 200 also has higher accuracy and fewer tokens than the endpoint, so the dashed curve is not a strict nondominated frontier across the displayed checkpoints.

MATH-500 accuracy versus average visible length across training and brevity configurations.
MATH-500 accuracy versus average visible length across training and brevity configurations.

The brevity distribution shifts mass toward shorter outputs and reduces the high-length tail. The plotted means are 1,721 for SFT, 1,919 for standard RL, and 1,276 for brevity; this SFT distribution corresponds to the longer-output curriculum checkpoint rather than Table 8’s stage-6 mean of 1,433.

Visible-response length distributions for SFT, Switch-GRPO, and the brevity variant.
Visible-response length distributions for SFT, Switch-GRPO, and the brevity variant.

5. How Does Latent Work in Reasoning?

Mechanistic analysis asks whether <swi> is a localized learned decision, whether latent execution affects answers, and where useful computation occurs within the block. It compares curriculum-SFT and post-RL checkpoints using teacher forcing, logistic probes, activation interventions, and qualitative logit-lens readouts.

At annotated boundaries, p(<swi>) is 0.847 after SFT and 0.480 after RL, compared with 0.003 and 0.002 at random non-boundary positions. Mean ranks are 1.13 and 1.68 at boundaries versus 1003.9 and 1127.9 at controls, supporting strong positional selectivity.

  • The tabulated probability contrasts are approximately 282-fold and 240-fold, not the roughly four orders of magnitude claimed in the surrounding prose.
  • RL lowers boundary confidence and raises boundary entropy from 0.203 to 0.532. The measured policy remains selective but is less certain at annotated boundaries.

Boundary and random-control positions are strongly separated in switch probability and rank. Post-RL boundaries are less confident and more entropic than SFT boundaries while remaining selective; the probability ratios are hundreds-fold, not the four orders of magnitude stated in the prose.

Teacher-forced switch probabilities, entropy, ranks, and margins at boundaries and controls.
Teacher-forced switch probabilities, entropy, ranks, and margins at boundaries and controls.

Across offsets −8 through +8, switch probability peaks at the annotated boundary and collapses immediately afterward. At offset +1 it is 2×10^−6 for both checkpoints, whereas at offset −1 it is 1.3×10^−2 after SFT and 9.7×10^−3 after RL.

  • The preserved narrow peak supports boundary localization despite reduced post-RL confidence.
  • These teacher-forced measurements establish discrimination around annotated positions, not optimal switch placement during free generation.

The central offsets quantify a narrow switch peak: probability is high at offset zero and drops to 2×10^−6 immediately afterward in both checkpoints. RL lowers the peak without removing positional localization.

Switch probabilities immediately before, at, and after annotated boundaries.
Switch probabilities immediately before, at, and after annotated boundaries.

The logarithmic probability axis exposes the abrupt post-boundary collapse, while the entropy axis shows increased uncertainty at the post-RL boundary. The curves support a softer but sharply localized policy on teacher-forced annotated prefixes.

Switch probability and entropy across a window centered on annotated entry positions.
Switch probability and entropy across a window centered on annotated entry positions.

Balanced logistic probes recover the label “next token is <swi>” increasingly well in deeper layers. Accuracy rises from approximately 53% at layer offset −24 to 91.9% after SFT and 88.4% after RL at offset −1.

  • The probes use ℓ2 regularization with C=1.0 and a single 80:20 train/test split.
  • Linear decodability demonstrates an encoded switch-related feature, but does not independently establish that the probed feature causally controls switching.

Probe accuracy increases with depth in both checkpoints, reaching 0.919 after SFT and 0.884 after RL. This establishes recoverability of the switch-related label from late activations, not causal necessity of a particular feature.

Balanced switch-label probe accuracy at seven layer offsets.
Balanced switch-label probe accuracy at seven layer offsets.

Both curves rise from near chance in earlier layers to strong classification in the final layer. Their similar depth profiles indicate that RL preserves a recoverable switch-related representation despite reducing final-layer probe accuracy.

Layerwise linear-probe accuracy for predicting the next switch token.
Layerwise linear-probe accuracy for predicting the next switch token.

Interventions compare normal execution, zeroed injected states, same-norm random replacement, and skipping latent execution. On problems that normal generation both solves correctly and handles with a latent block, accuracy falls from 100% to 33.3%, 90.5%, and 81.0%, respectively.

  • On the unrestricted evaluation, accuracy is 70%, 42%, 72%, and 70%. Skipping therefore has no aggregate accuracy cost in this experiment, despite harming the selected diagnostic subset.
  • The diagnostic subset is selected using normal-mode success. It establishes sensitivity on those cases, not a population-wide benefit from latent execution.
  • Zeroing is much more disruptive than same-norm replacement. The interventions leave state norm, distributional disruption, and problem-specific semantic content incompletely separated.

Zeroing strongly damages both overall and diagnostic accuracy, whereas skipping preserves overall accuracy but harms the normal-success subset. Same-norm random replacement is substantially less destructive, limiting how specifically the experiment identifies semantic information in the original injected state.

Accuracy and answer changes under normal, zero, random-norm, and skip interventions.
Accuracy and answer changes under normal, zero, random-norm, and skip interventions.

The paired panels show the difference between unrestricted and selected-subset effects. Zeroing produces the largest disruption, while unchanged aggregate skip accuracy cautions against treating diagnostic-subset losses as a benchmark-wide benefit.

Accuracy and answer-change rates under latent-state interventions on full and diagnostic evaluations.
Accuracy and answer-change rates under latent-state interventions on full and diagnostic evaluations.

The logit lens returns </swi> as the top token at each latent step, with more diffuse, problem-related alternatives at the first step. Exit probabilities remain approximately one across steps 1–4 for both correct and incorrect trajectories, explaining the need for an externally enforced minimum dwell.

  • The authors interpret these observations as concentrating useful computation at the entry transition.
  • Exit readiness is an output-distribution property, not a direct measure of internal computational value. Step-specific causal ablations would be needed to establish that later transitions are functionally negligible.
  • The qualitative logit-lens vocabulary should not be treated as a faithful decoded reasoning trace.

Exit probability is almost one from the first latent step onward for both correct and wrong rollouts. This explains dependence on forced minimum dwell, but does not establish that later hidden transitions perform no useful computation.

Latent-step exit probabilities grouped by eventual answer correctness.
Latent-step exit probabilities grouped by eventual answer correctness.

6. Conclusion

Switch provides a discrete interface for training and inspecting a reasoner that executes hidden-state recurrence. Its strongest run improves accuracy over the selected recurrent-latent baselines, while boundary measurements and interventions expose aspects of its execution to empirical analysis.

  • The evidence does not establish fully differentiable RL through recurrent latent states, equal-cost superiority, or complete reconstruction of latent reasoning.

Limitations

Evaluation is restricted to Qwen3-8B and two mathematics benchmarks. The authors acknowledge that latent segments are not differentiated during RL and that their representations are shaped primarily by the supervised curriculum; this qualification conflicts with the abstract’s claim of gradients propagating through recurrent latent computation.

  • Vocabulary-mixture systems are not compared head to head at matched scale and data.
  • The authors describe the logit-lens analysis as qualitative rather than a faithful reconstruction of latent reasoning.
  • No multi-seed uncertainty or sufficient statistical detail is reported to assess the robustness of the mechanistic effect sizes.

The full 1,964-step training trace reveals a late regime, beginning around step 1,200, that the authors call reward hacking: switching approaches 100% and latent blocks rise to approximately 13 per problem while reward declines. Figure 9 supports avoiding this degraded regime, but Table 8 shows that step 800 is not the highest-accuracy checkpoint in the representative trajectory.

  • The reported operating point depends on monitoring policy behavior and selecting a checkpoint rather than optimizing the reward to convergence.
  • Visible-token counts omit latent compute, and unresolved differences among headline, conditional, and representative-run metrics limit attribution of the reported gains.

The extended trajectory exposes behavior absent from the cropped main-body plot: late switching approaches 100% and latent invocations rise sharply while reward worsens. It supports avoiding the degraded regime, but does not establish step 800 as the best evaluation-accuracy checkpoint.

Full 1,964-step training trajectory with late policy degradation and the selected stopping point.
Full 1,964-step training trajectory with late policy degradation and the selected stopping point.

Appendix

  • A. Implementation Details: Special-token IDs are 151669 for <swi>, 151670 for </swi>, and 151671 for <latent>. Phase 1 uses LoRA rank 32 and α=64 on q, k, v, o, gate, up, and down projections, together with trainable resized embeddings and LM head; reported cross-entropy reaches 0.098. Segmented backward processes text with gradients and latent segments under torch.no_grad(), reducing peak activation memory to one text segment; out-of-memory handling preserves the surviving rollouts when computing group advantages.
  • B. Switch-GRPO Loss, in Full: Correctness and format rewards take ±1 values, the usage term multiplies correctness by well-formedness and usage indicators, and brevity is clipped between 0 and 1 and gated on correctness and usage. Group-relative advantages use ε=10^−8; the text-only clipped objective uses ε_c=0.2 and β=10^−3, with the old policy serving as both importance-ratio reference and KL anchor.
  • C. Per-Checkpoint Trajectory of Switch: The representative run reaches 70.0% overall accuracy after curriculum stage 6 with K_min=4, 78.0% at RL step 200, and 72.6% at step 800. The appendix distinguishes this trajectory from the strongest run producing the headline 79.3%/89.2% results and documents late-training policy degradation.
  • D. Algorithm Boxes: Algorithm 1, CoconutSwiModel forward, partitions text and latent execution while carrying a streaming KV cache. Algorithm 2, One Switch-GRPO step, performs hidden-state-injection rollouts, computes group-relative advantages, and applies segmented backward to text segments only.
  • E. Visible-Token CDF: Figure 10 complements the main-body histogram with an empirical length distribution. The brevity variant shifts responses toward shorter visible lengths, whereas the standard RL variant increases the median by the annotated 69 tokens relative to SFT.
  • F. Mechanistic Analysis: Additional Details: Probes use balanced positive and negative samples, C=1.0, and a single 80:20 split. Teacher-forced comparisons use the same MATH-500 problems across checkpoints; interventions use greedy decoding with K_min=4 and select the diagnostic subset from correct normal-mode responses containing a latent block.
  • G. Per-Subject and Per-Difficulty Visualisation: Reported subject accuracies are Algebra 88.7%, Prealgebra 80.5%, Number Theory 79.0%, Precalculus 67.9%, Counting & Probability 60.5%, Intermediate Algebra 56.7%, and Geometry 53.7%. Difficulty-level accuracies are 93.0%, 90.0%, 83.8%, 64.1%, and 53.7% for levels 1–5. Section G calls these headline-checkpoint results, whereas Table 9’s caption identifies the representative run; Figure 11’s latent and text branches are model-selected rather than randomized groups.
  • H. Generation Trace Analysis: Correct responses average 959 visible tokens without switching and 1,197 with switching, whereas wrong responses average 1,729 and 1,805. Correct switched responses average 1.86 blocks and 7.43 latent steps, versus 1.40 blocks and 5.60 steps for wrong switched responses. Wrong trajectories have less confident switch decisions, with entropy 0.717 versus 0.608 and p(<swi>)=0.669 versus 0.763.
  • I. Ablations: The authors sweep K_min over {0, 2, 4, 8, 16}. Removing minimum dwell yields single-pass latent blocks and a reported 53.0% curriculum-SFT accuracy; the appendix favors K_min=4 but provides no complete numerical table for the sweep. Table 8 also changes curriculum stage between its K_min=0 and K_min=4 SFT rows, so those rows do not isolate dwell alone.
  • J. Full Related Work: The extended review organizes prior work into latent representations, hybrid switching, RL optimization, and interpretation of internal states. It positions Switch as preserving Coconut-style recurrence while adding a trainable discrete control interface.
  • J.1. Latent Chain-of-Thought Reasoning: Coconut uses a gradual recurrence curriculum, CODI aligns student and teacher states through self-distillation, and CoLaR adds a latent head. Vocabulary-mixture approaches instead combine token embeddings, with recent RL variants adding sampling mechanisms; their exclusion narrows the demonstrated empirical advantage. The review also discusses earlier small-scale recurrence experiments, pause and filler tokens, implicit-CoT internalization, and multimodal latent representations.
  • J.2. Switchable / Hybrid Reasoning: SwiReasoning switches a frozen model according to an external entropy rule, whereas Switch learns boundary-token probabilities and introduces latent execution through supervised training. Adaptive visible-reasoning methods allocate extra computation without replacing decoded thinking with recurrent hidden states. Claims here that latent dynamics and dwell are optimized end to end by RL require qualification by the text-only gradient implementation and enforced minimum dwell.
  • J.3. Reinforcement Learning for Reasoning and Latents: Vocabulary-mixture RL approaches make latent actions samplable, while Switch-GRPO restricts likelihood ratios to sampled text decisions and retains deterministic recurrent execution. This supplies an RL-compatible trajectory interface, not direct policy-gradient terms at continuous latent positions.
  • J.4. Interpretability of Internal Reasoning States: The paper combines logit-lens readouts, linear probes, and causal activation interventions. These tools test complementary properties, but qualitative vocabulary readouts and linearly recoverable features do not explain the full latent computation.
  • K. Extended Discussion: The authors argue that tractable policy densities are needed at sampled decision points rather than at every deterministic computational step. They interpret reduced boundary probabilities and intervention effects as improved switch selection and meaningful latent computation. The discussion’s claims that switching halves and latent-conditional accuracy nearly doubles are not supported by Table 2’s 81%→58% and 66.7%→79.3% values, while same-norm replacement leaves the specificity of latent content unresolved.

Brief Thoughts

The boundary interface is the clearest contribution: it makes recurrent execution controllable, compatible with discrete-action RL, and easier to instrument. Matched-base-model re-implementations and disclosure of late-training degradation make the evaluation more informative, although they do not resolve differences in inference cost.

The empirical and mechanistic claims need tighter accounting. Reconciled checkpoint and decoding tables, matched total-compute comparisons, repeated runs, and step-specific interventions would help distinguish improved switch selection and visible reasoning from problem-specific latent computation.