Therefore I am. I Think | Summary
01 Apr 2026 | Paper Review Tool Use Activation Steering Chain-of-Thought FaithfulnessContents
This article explains the key points of Therefore I am. I Think.
- 2026-04-01 (arXiv)
- Esakkiraja, Esakkivel, Rajeswar, Sai, Akhiyarov, Denis, Venkatesaramani, Rajagopal.
- Khoury College of Computer Sciences, Northeastern University, Mila, ServiceNow Research, ServiceNow
- Paper
Summary
- Therefore I am. I Think. investigates whether reasoning models encode tool-use choices before producing visible chain-of-thought (CoT). Logistic regression probes predict the model’s eventual tool-versus-no-tool action from pre_gen activations with AUROC above 0.90 in all four main model–benchmark settings.
- The authors construct a tool-call steering direction from class-conditional mean activations and use it to inject or suppress tool use. On 100 held-out examples per model–benchmark setting and direction, the strongest tested interventions produce thinking-mode flip rates of 11–79%, often accompanied by longer reasoning.
- Two external LLM judges frequently classify injection-induced action changes as Confabulated Support or Constraint Override. These findings support early action predictability and susceptibility to intervention, but do not establish that all visible reasoning is post-hoc: steering continues during decoding, resistant cases remain, and the behavioral labels describe observable text rather than internal reasoning.
1 Introduction
The paper asks whether tool-use choices are predictable before visible reasoning begins, whether those choices can be steered, and whether subsequent CoT resists or justifies the intervention.
- Tool calling provides a binary, interpretable action target rather than requiring analysis of an unrestricted generated answer.
- The three stated contributions are Early decision encoding, Decision direction causality, and Rationalization behavior.
2 Related Work
The paper connects anticipatory representations of future outputs, hidden-state probes for correctness and adaptive computation, studies of CoT faithfulness, and activation steering. Its focus is the pre-generation representation of a tool-use action and the reasoning produced after perturbing that representation.
- Unlike early-exit work, the objective is not to terminate reasoning efficiently but to examine the action tendency present before visible reasoning.
- Activation steering is used to study causal effects on behavior, not as a fine-tuning or performance-optimization procedure.
3 Methods
The methodology collects baseline reasoning traces, extracts residual-stream activations, trains linear action probes, and constructs a mean-difference steering vector. Evaluation combines probe AUROC, action flip rates, reasoning-token changes, and pairwise behavioral judgments.
The pipeline separates reading an action signal from intervening on it and evaluating the consequences. Probe predictability alone cannot establish that the representation affects behavior.
3.1 Models, Data, and Benchmarks
The main experiments use Qwen3-4B and GLM-Z1-9B on When2Call and a supporting benchmark constructed from BFCL Irrelevance and BFCL Simple, v3: base + live. GPT-OSS-20B receives supplemental probing experiments but is excluded from causal analysis because, according to the authors, its mixture-of-experts architecture requires a different steering technique.
- When2Call contains 3,652 multiple-choice examples and 300 LLM-judge examples, with tool_call (∼57%), direct (∼12%), request_for_info (∼14%), and cannot_answer (∼17%) categories.
- BFCL Irrelevance supplies tool-mismatch cases, while BFCL Simple supplies straightforward solvable tool-use cases, testing generalization across benchmark construction and prompt style.
3.2 Hidden-State Extraction and Prediction Target
Baseline traces are generated with vLLM using recommended generation arguments. Forward hooks extract the post-layer residual stream from a full pass over each prompt and continuation; causal attention preserves the autoregressive representation at each position.
- pre_gen is immediately before the first thinking token; think_start marks the first thinking token; think_end marks the last reasoning token; decision_token is the first token after reasoning.
- Intermediate positions sample the reasoning span. The plotted main-model positions are pre_gen, think_start, 5%, 10%, 25%, 50%, 75%, and think_end.
- The prediction target is the model’s traced tool-versus-no-tool behavior, not whether its action matches the benchmark’s gold label.
3.3 Probe Training
A logistic regression probe maps a hidden-state vector to a tool-call probability: ŷᵢ = σ(w⊤xᵢ). The tool class is predicted when ŷᵢ ≥ 0.5, and probes are trained with Binary Cross Entropy loss independently for each sampled layer–position pair.
- Section 3.3 specifies layer sampling every four layers, including the first and final layers; Section 3.5 describes intervals of approximately four or five layers, consistent with the main-model heatmaps.
- Section 3.3 specifies nine token positions through decision_token, whereas the main-model plots show eight positions through think_end.
3.4 Activation Steering Vector
For a fixed layer and pre_gen position, the authors average activations separately for baseline tool and no-tool examples. Their steering direction is (v=\mu_+-\mu_-), and the intervention is (h’^{(L,t)}=h^{(L,t)}+\alpha v); positive and negative coefficients inject and suppress tool use, respectively.
- Most settings test coefficient magnitudes (\alpha\in{4,8,12}). GLM-Z1-9B on BFCL uses (\alpha\in{10,20,30}) because the selected layer has a substantially larger mean activation norm.
- Although the vector is derived exclusively from pre_gen activations, it is applied during both pre-fill and decoding. The experiment therefore does not isolate a perturbation at the pre-generation instant.
3.5 Evaluation Metrics
Probe performance is evaluated with 5-fold stratified cross-validation using AUROC. Steering uses 100 held-out examples per model, direction, and benchmark, excluded from both probe training and steering-vector computation; suppression subsets begin with baseline tool calls, while injection subsets begin with baseline no-tool actions.
- Suppression flip rate. The fraction of baseline tool examples that become no-tool examples after steering.
- Injection flip rate. The fraction of baseline no-tool examples that become tool examples after steering.
- Reasoning-token change. The relative change is Δreason = (rsteer − rbase)/rbase, comparing reasoning-token counts for the same example.
Behavioral analysis. GPT 5.4 and Claude Sonnet 4.6 compare baseline and steered responses at (\alpha=\pm12), or (\alpha=\pm30) for GLM-Z1-9B on BFCL. Each pair is evaluated twice with reversed presentation order at temperature 0, and the paper reports inter-judge agreement and class distributions separately for flipped and resistant outcomes.
- Seamless divergence denotes a different final action justified fluently without visible conflict; Confabulated support denotes invented facts, defaults, or user intent.
- Constraint override denotes acknowledging and weakly dismissing a constraint; Inflated deliberation denotes additional hedging or repeated reconsideration without new information.
- Decision instability denotes a visible change of reasoning direction; No meaningful difference denotes behaviorally comparable responses.
4 Results
The results examine whether early activations predict an eventual action, whether interventions alter that action, and how the resulting reasoning differs from baseline. These analyses concern action selection rather than overall task accuracy or successful tool execution.
Pre-Generation Activations Predict Action Decisions.
At the selected layer, pre_gen AUROC exceeds 0.95 in three main settings and 0.90 in all four. Predictability drops early in the reasoning span, particularly around 5–10% on When2Call, then recovers toward nearly perfect discrimination at think_end.
- Figure 2 presents layer-20 curves alongside the mean across sampled layers; the layer-average performance is generally lower.
- The early AUROC dip establishes reduced linear decodability at those positions, not by itself that the model has become uncertain or changed its internal decision.
Across the four main model–benchmark panels, the selected-layer curve starts above 0.90 AUROC before reasoning, declines early, and approaches 1.0 at think_end. The mean-across-layers curves show that this signal is not equally accessible throughout the network.
Pre_gen predictions agree with think_end predictions in 88% of When2Call examples for both main models, 92% of BFCL examples for Qwen3-4B, and 90% for GLM-Z1-9B. The think_end probe predicts the realized action with accuracies of 99.0%, 99.4%, 99.3%, and 99.7%, respectively.
- Agreement also declines during early reasoning before recovering. The early action signal persists in many examples, but early and final probe predictions differ in a nontrivial subset.
Pre_gen agreement with the final probe is 88%, 88%, 92%, and 90% across the four panels, while final-probe accuracies are 99.0–99.7%. The initial probe prediction is highly informative but not identical to the final one in every example.
Steering the Pre-Generation Signal Affects CoT and Action Decisions.
Stronger steering increases Qwen3-4B flip rates in both directions. At coefficient magnitude 12, thinking-mode suppression and injection flip rates are 49% and 62% on When2Call, and 26% and 53% on BFCL; the corresponding no-thinking rates are 18% and 29%, and 20% and 27%.
- For GLM-Z1-9B, When2Call suppression remains approximately flat at 10%, 9%, and 11%, whereas injection increases from 3% to 13% to 21%. No-thinking mode is unavailable for this model in the authors’ setup.
- On BFCL, GLM-Z1-9B suppression increases from 9% to 27% to 58%, and injection from 29% to 52% to 79%, at coefficient magnitudes 10, 20, and 30.
- Control directions derived from ProntoQA True/False examples with similar activation norms produce a reported 0% flip rate across the tested models and benchmarks.
The strength sweep shows increasing flip rates in most settings, with GLM-Z1-9B suppression on When2Call as an exception. Qwen3-4B is generally more steerable in thinking mode, although thinking and no-thinking suppression are equal at coefficient magnitude 8 on BFCL.
The representative injection example asks the model to “play baby Shark” when the only available tool is set_volume. Baseline Qwen3-4B correctly declines playback, but steering at (\alpha=12) induces set_volume(volume=100), despite the steered reasoning explicitly recognizing that volume adjustment is not music playback.
- The pre-generation probe assigns a tool-call probability of 0.16. The example illustrates Constraint Override rather than a legitimate solution to the user’s request.
The steered response recognizes that set_volume cannot play a song, then calls set_volume(volume=100) anyway. The example shows an acknowledged tool mismatch being overridden rather than successful tool use.
Steering frequently increases average CoT length, including when the final action resists the intervention. Qwen3-4B’s resistant When2Call suppression group rises from 208 to 477 tokens, a reported ratio of 2.30, while GLM-Z1-9B’s resistant BFCL suppression group rises from 259.1 to 564.8 tokens, a ratio of 2.18.
- Inflation is not universal: Qwen3-4B’s resistant injection groups have ratios of 0.98 on When2Call and 1.00 on BFCL; GLM-Z1-9B’s resistant When2Call suppression group has a ratio of 0.95.
- Table 2 reports steered-to-baseline token ratios rather than the relative-change quantity defined in Section 3.5.
- The Qwen3-4B When2Call suppression split is 45 flipped and 55 resistant examples in Table 2, whereas Table 1 reports a 49% flip rate at the same stated strength. The paper does not reconcile this discrepancy.
Both flipped and resistant groups can generate substantially more reasoning, while some resistant groups remain unchanged or slightly shorten. The table distinguishes action susceptibility from token inflation and reports a Qwen3-4B When2Call suppression count inconsistent with the flip rate in Table 1.
Behavioral Analysis Shows Rationalization.
On When2Call, the judges agree on 73 and 93 Qwen3-4B examples for suppression and injection, and on 72 and 89 GLM-Z1-9B examples. Among agreed flipped injection pairs, every case is classified as Confabulated Support or Constraint Override: 53 and 5 cases for Qwen3-4B, and 11 and 7 for GLM-Z1-9B.
- Suppression more often produces Inflated Deliberation or No Meaningful Difference. Of the agreed Inflated Deliberation cases, 18/37 flip for Qwen3-4B and 7/18 flip for GLM-Z1-9B.
- These distributions cover only examples on which both judges agree, not the full held-out sets.
Among judge-agreed When2Call injection flips, all cases fall into Confabulated Support or Constraint Override. Suppression contains many Inflated Deliberation and No Meaningful Difference cases, showing that behavioral responses differ by intervention direction.
On BFCL, agreed flipped injection pairs are again dominated by Confabulated Support and Constraint Override: 21 and 16 cases for Qwen3-4B, and 23 and 20 for GLM-Z1-9B. Suppression differs by model: Qwen3-4B has many resistant No Meaningful Difference or Inflated Deliberation cases, whereas GLM-Z1-9B has 24 flipped Decision Instability cases.
- Inter-judge agreement is lower for BFCL injection: 71/100 examples for Qwen3-4B and 62/100 for GLM-Z1-9B.
- The behavioral evidence supports rationalization-like text patterns, especially under injection, but the categories do not directly measure the internal causal role of CoT.
BFCL injection flips again concentrate in invented support and overridden constraints. GLM-Z1-9B suppression is distinctive: Decision Instability contains 24 of the 39 agreed flipped cases, whereas many resistant cases show Inflated Deliberation.
5 Discussion
The authors interpret the probing, steering, and behavioral results as evidence that action choices can be encoded before visible reasoning and that subsequent CoT can rationalize externally induced changes. Resistant examples with inflated reasoning also show that longer deliberation does not necessarily revise the selected action.
- The discussion questions the reliability of CoT as an explanation of decision causes and identifies activation manipulation as a potential security concern; it does not evaluate a concrete attack.
- A proposed future direction penalizes high pre-generation probe confidence during reinforcement-learning training. This intervention is not implemented or tested in the paper.
Appendix
- A.1 When2Call Layer-Position Heatmaps and A.2 BFCL Layer-Position Heatmaps show that strong pre_gen decodability is concentrated in middle-to-late layers rather than being uniform across the network. Both benchmarks exhibit early-reasoning dips and later recovery; the displayed pre_gen maxima are approximately 0.96 on When2Call for both models, 0.98 on BFCL for Qwen3-4B, and 0.95 on BFCL for GLM-Z1-9B.
- A.3 Supplemental GPT-OSS-20B Results includes A.3.1 Layer-Position Heatmaps, A.3.2 Position Curves, and A.3.3 Agreement Curves for medium and high reasoning. GPT-OSS-20B follows the early-predictability, early-dip, and late-recovery pattern, but When2Call pre_gen AUROC is lower than for the main models: the displayed heatmap maxima are 0.89 for medium reasoning and 0.87 for high reasoning, versus approximately 0.94 for both settings on BFCL.
- GPT-OSS-20B pre_gen agreement with the think_end probe is 80% and 78% on When2Call under medium and high reasoning, and 89% and 87% on BFCL. Corresponding think_end action-prediction accuracies are 98.9%, 97.2%, 98.5%, and 97.5%; these results extend the predictive observation but provide no steering evidence for this model.
- A.4 Behavioral Analysis Judge Prompt supplies the shared Judge Instructions. and Runtime Input Template. The rubric requires one most-prominent behavioral category, restricts judgments to visible text, separates behavior from correctness, and prefers no_meaningful_difference when evidence is weak; runtime inputs include the task context, baseline and steered responses, and their final actions.
- A.5 Additional Steering Examples and A.5.1 Illustrative Behavioral Examples document both flips and resistance at the displayed (\alpha=12). A salon-search example with probe probability 0.9992 abandons a tool call after 2.87× more reasoning; a weather example with probability 1.00 retains its call despite 3.57× more reasoning; an application-metadata example with probability (7.5\times10^{-9}) remains no-tool after briefly mentioning an unrelated function.
- A.6 Judge Disagreement Statistics reports When2Call disagreement rates of 11.0% and 28.0% for GLM-Z1-9B injection and suppression, and 7.0% and 26.3% for Qwen3-4B. BFCL rates are 38.0% and 31.0% for GLM-Z1-9B, and 29.0% and 27.0% for Qwen3-4B. The most common BFCL disagreements are Confabulated Support versus Constraint Override during injection and Inflated Deliberation versus Decision Instability during suppression; When2Call injection disagreements most often concern No Meaningful Difference versus Inflated Deliberation.
Brief Thoughts
The strongest contribution is the combination of early action decoding, strength-dependent behavioral intervention in most settings, and an unrelated-direction control with similar activation norms. The supported conclusion is that a pre-generation-derived direction predicts and influences tool-use behavior; applying steering throughout decoding prevents attributing the entire effect to an initial commitment alone.
The quantitative interpretation needs care: AUROC is not a calibrated probability or an accuracy percentage, and an early decodability dip does not establish uncertainty. The paper’s random-agreement argument uses 1/36 for agreement on a particular category; under independent uniform six-class labels, agreement on any category is 1/6, and empirical class imbalance further changes the appropriate baseline. The agreed-only behavioral distributions exclude up to 38% of examples in some settings, so they should not be treated as unbiased estimates for all steered responses. Table 5’s 26.3% Qwen3-4B When2Call suppression disagreement rate does not match Table 3’s 73 agreed examples out of 100. Together with the suppression flip-count discrepancy, this limits exact cross-table comparisons. Two causally tested models and binary tool-use benchmarks support a focused finding, not a general conclusion that reasoning models always decide before thinking.
The early action signal strengthens in middle-to-late layers, reaching displayed pre_gen AUROC values around 0.96. The early-reasoning dip appears across multiple layers, not only in the selected-layer curve.
GPT-OSS-20B’s strongest displayed pre_gen AUROC is 0.89 under medium reasoning and 0.87 under high reasoning. These values are lower than the main models’ When2Call results, although the early dip and near-final recovery remain evident.
Both GPT-OSS-20B reasoning settings reach displayed pre_gen AUROC values around 0.94 on BFCL. Early predictability is stronger than on When2Call, demonstrating benchmark dependence within the same model.
The When2Call position curves show lower early AUROC under high reasoning and a deeper early-reasoning dip than under medium reasoning. Both settings recover near the end, but the curves do not establish whether that recovery requires all intervening reasoning tokens.
BFCL position curves begin above 0.90 at the best layer in both reasoning settings and recover toward approximately 1.0. The layer-average curves remain lower, so best-layer results should not be generalized to every representation.
When2Call pre_gen agreement is 80% for medium reasoning and 78% for high reasoning, below the main models’ 88%. Final-probe accuracies of 98.9% and 97.2% show high final action decodability while leaving substantial early-versus-final disagreement.
BFCL pre_gen agreement is 89% and 87% under medium and high reasoning, with final-probe accuracies of 98.5% and 97.5%. Agreement falls during early reasoning and then increases, extending the temporal pattern to the supplemental model.
Suppression converts a salon-search call into a request for additional location details by repeatedly reconsidering whether Pleasanton needs a state. The 2.87× reasoning increase accompanies newly emphasized obstacles to making the call.
The weather call survives suppression, but reasoning expands 3.57× while debating whether the tool accepts states rather than cities. Resistance in the final action does not imply that the reasoning trace is unaffected.
Injection briefly introduces an unrelated get_current_weather reference, yet the model returns to its original conclusion that the available tools cannot retrieve application metadata. This example shows that steering does not inevitably produce an action flip.
Disagreement reaches 38.0% for GLM-Z1-9B BFCL injection, most often between Confabulated Support and Constraint Override. Frequent disagreements between related categories limit the precision of the behavioral labels; the reported 26.3% Qwen3-4B When2Call suppression rate also differs from the 27% implied by Table 3.