Strategic Choice in LLMs: Internal Incentive Pathways Differ Between Base and Instruct Models

A new activation-level study of 144 2x2 games shows base and instruct LLMs choose alike but route incentive to choice differently.

Direct answer

A new study recording activations from four open-weight models across all 144 strict ordinal 2x2 games finds that strategic incentive and eventual choice are decodable in every model, yet models differ in whether that incentive actually reaches the decision [1]. The matched Qwen2.5 base and instruct pair chose almost identically at baseline (96.4% agreement) and represented the payoff incentive to a similar degree, but post-training strengthened the link between incentive and choice [1]. This matters because earlier work on strategic choice in LLMs relied on outputs, reasoning traces, or agent-level performance, leaving the internal route from represented incentive to decision unmeasured [1][2]. The result reframes post-training as a change in how already-available information is recruited, not merely in what information is stored.

5sources cited

This article was generated with WisPaper-powered search and paper analysis.

From behavioural similarity to internal routes

Earlier evidence established that LLMs behave systematically as strategic agents but left the internal computation unmeasured. The private-agent work showed that augmenting an LLM with concealed deliberation and deception improved long-term payoffs across repeated Prisoner's Dilemma, stag hunt, chicken, and battle-of-the-sexes games, with the private agent scoring 1.43 versus 3.76 for the public agent in GPT-4 play [2]. That study inferred strategy from choices and payoffs, not from activations. Reviews of attention-head mechanisms similarly mapped internal components such as constant, single-letter, negative, and memory heads to reasoning stages, but noted that identified circuits are rarely validated across tasks and that most evidence comes from simple, specific scenarios [3]. The anchor paper's contribution is to follow a prespecified incentive from prompt, through activations, to choice within one exhaustive decision domain, the 144 strict ordinal 2x2 games [1].

The design matters for interpretation. Each game was played once with no history or feedback, and every game was rendered in four counterbalanced forms so that aggregated quantities do not depend on letter assignment or option order [1]. The canonical action was defined by a fixed priority rule: dominant action first, then the unique pure-strategy equilibrium, then the payoff-dominant equilibrium, then maximin [1]. This gives a common axis for comparing choices, decoded directions, and token-level answer scores across games, which is what allows the paper to ask whether incentive reaches choice rather than only whether both are decodable.

Availability, recruitment, and susceptibility come apart

The central empirical pattern is a three-way dissociation. Incentive, choice, and cue identities were decodable from activations in all four models, but the models differed in whether the encoded incentive was recruited into the choice, when this happened, and how susceptible the choice was to intervention along that signal [1]. In the dense models, incentive and choice representations became aligned only in deeper transformer blocks, closer to the output, and a separate prompt-position analysis showed the incentive could be decoded while the model processed its own payoffs but became strongly reflected in preference between the two answers only near the choice position [1]. GPT-OSS showed a different pattern: the same linear probes recovered the incentive more accurately from router gate scores than from selected expert routes, though these activations were recorded after reasoning and were not manipulated, so the comparison remains descriptive [1].

Causal steering sharpened the distinction between representation and use. Strengthening the incentive direction at layer 65 produced canonical-action dose-slopes of +0.060 for Qwen2.5 and +0.236 for Qwen2.5-Instruct in double-dominance games, with the instruct interval excluding zero and the base interval also excluding zero but smaller [1]. In one-dominance level-2 games the slopes were +0.019 for base and +0.152 for instruct, and pooled across 54 games they were +0.012 and +0.066 respectively, against a random-direction control of +0.001 and +0.006 [1]. Llama-3.1 showed a positive but interval-including-zero pooled slope of +0.032, and the regenerated-choice effects were weaker and overlapping with zero, so the strongest causal evidence concerns graded preference rather than regenerated decisions [1].

The matched Qwen pair isolates post-training

The clearest evidence that similar behaviour can rest on different computation comes from the matched Qwen2.5 pair, which shares pretrained weights. The base and instruction-tuned models agreed on 96.4% of baseline decisions and represented the payoff incentive to a similar degree, yet in the instruction-tuned model that information was more strongly linked to choice and more strongly reflected in its preference between the two answers [1]. Within this pair, post-training changed how already-available information reached choice, not whether the information was present [1]. This is a concrete difference that behaviour and decodability alone would have missed.

Competing evidence from repetition priming points in a compatible direction while warning against overgeneralising. That work found that base models show strong positive priming with no lag sensitivity, whereas instruct models show weaker effects and negative lag effects, with a ModelType x Lag interaction in Experiment 1 (beta = -0.26, SE = 0.04, t = -6.50, p < .001) and Experiment 2 (beta = -0.13, SE = 0.03, t = -4.52, p < .001) [4]. Targeted ablation of the top 10 prior-occurrence attention heads sharply reduced base-model priming (LLaMA3: +6.89 to +2.62; Qwen2.5: +5.52 to +1.84) but barely changed instruct models (LLaMA3: +2.05 to +1.97; Qwen2.5: +2.31 to +2.14) [4]. The authors interpret this as automatic versus controlled processing and note that the source of the gating remains open, possibly conflict resolution or broader post-training-induced repetition avoidance rather than RLHF alone [4]. Both papers thus converge on post-training reshaping a pathway rather than simply adding or removing information, but they measure different constructs and cannot be treated as one mechanism.

Fixed decision cues are internally distinct but behaviourally constrained

The anchor paper also tested five fixed decision cues plus a neutral procedural control. Cued and baseline responses were paired within game and presentation form, and game-held-out linear discriminant analysis classified which cue produced each activation difference with 0.97-1.00 accuracy, showing the fixed wordings produce distinguishable activation shifts [1]. The authors explicitly caution that this does not separate construct identity from vocabulary, syntax, or length, and is not evidence of a general risk-, loss-, or inequity-aversion representation [1]. Behaviourally, cue effects depended on the cue, the model, and the game: structure supplied the main predictive gain in nested held-out prediction, while the incremental cue-identity contribution was small [1].

This representation-versus-use gap echoes the attention-head review's warning that identified components are often task-specific and rarely validated across settings [3]. It also aligns with the private-agent finding that augmenting an agent with private deliberation improved performance in competitive scenarios but that LLMs had deficiencies in sampling from distributions and identifying opponent types [2]. Across these lines, internal or architectural interventions can change behaviour, but the mapping from a decodable signal to a decision is neither uniform nor guaranteed.

Where the conclusion stops

The study's own limitations bound the claim tightly. The game catalogue is exhaustive only within one-shot, strict ordinal 2x2 games, so generalisation to repeated, sequential, incomplete-information, or multiplayer settings remains unknown [1]. Computational demands limited the model sample to four large open-weight checkpoints, making cross-model comparisons descriptive and the post-training conclusion limited to one matched Qwen pair [1]. Each decision cue used one fixed wording, so its effects cannot be separated from that wording or generalised to a psychological disposition [1]. The causal intervention tested one prespecified incentive at one answer position and two layers in three dense models, and GPT-OSS could not receive comparable token-level or causal tests because its activations were recorded after reasoning [1]. The comparison with humans establishes behavioural similarity, not shared internal computation [1].

A further boundary comes from evidence on how AI-risk disclosure itself is framed as a strategic choice, where firm fixed effects and taxonomy construction shape what gets disclosed [5]. That work is a limitation on any claim that a single internal signal or cue maps cleanly onto a stable construct across contexts. The practical implication for mechanistic interpretability and LLM decision-making researchers is that decodability, recruitment, and susceptibility are separate evidentiary standards, and none substitutes for the others [1]. The open question is whether the base-instruct pathway difference generalises beyond this game domain and beyond one matched checkpoint pair, and whether the gating observed in repetition priming and the recruitment difference observed here share any common post-training mechanism [1][4].

About These Sources

This research page is built on 5 studies (4 peer-reviewed, 1 preprint) — published from 2024 to 2026, 5 from 2024 or later — selected as the most relevant from 5 studies that passed quality screening, drawn from 56 papers retrieved from a database of over 500 million.

Sources used in this answer

1

The Internal Anatomy of Strategic Choice in Large Language Models

The anchor paper records activations from four open-weight models across all 144 strict ordinal 2x2 games and finds that incentive and choice are decodable in every model, but models differ in whether incentive reaches choice, with the matched Qwen2.5 base and instruct pair choosing almost identically (96.4% agreement) yet differing in incentive-to-choice linkage [1].

2

Effect of Private Deliberation: Deception of Large Language Models in Game Play.

The private-agent study shows that concealed deliberation and deception improve long-term payoffs in repeated games, with GPT-4 private agents scoring 1.43 versus 3.76 for public agents in Prisoner's Dilemma, but it infers strategy from choices and payoffs rather than internal computation [2].

3

Attention heads of large language models.

The attention-head review provides a four-stage framework and categorises heads such as constant, single-letter, negative, and memory heads, while noting that identified circuits are rarely validated across tasks and that most evidence comes from simple, specific scenarios [3].

4

Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans

The repetition-priming study finds base models show strong lag-invariant facilitation while instruct models show weaker, lag-sensitive effects, with targeted ablation of top prior-occurrence attention heads sharply reducing base-model priming but barely changing instruct models, interpreted as automatic versus controlled processing [4].

5

AI-risk Disclosure is not One-size-fits-all: An Agentic Measurement System and Network Evidence

The AI-risk disclosure work frames disclosure itself as a strategic choice shaped by firm fixed effects and taxonomy construction, serving as a limitation on claims that a single internal signal or cue maps cleanly onto a stable construct across contexts [5].