State-Driven Reasoning · NeurIPS 2026

State of Thought Enables Endogenous Reasoning

SoT reframes test-time reasoning as a closed loop: the model's endogenous state selects the right historical evidence for the next step and controls when to stop, rather than following fixed external reasoning programs.

Zhiren Gong1,2, Yikun Hou1, Zihao Zeng1, Ming Xiao4, Chau Yuen3, Wei Yang Bryan Lim1

1 College of Computing and Data Science, Nanyang Technological University 2 Interdisciplinary Graduate Programme, Nanyang Technological University 3 School of Electrical and Electronic Engineering, Nanyang Technological University 4 Information Science & Engineering, KTH Royal Institute of Technology, Sweden

Across 3 LLMs and 2 VLM scales (16 text benchmarks + 3 multimodal tasks), a 582-parameter controller on frozen backbones improves accuracy while cutting generated tokens by 62.6% and latency by 44.6%.

Reasoning, from the inside

AskWhich bag for a rainy hiking trip?

Each step, the endogenous state re-selects only the past steps that matter for this step (green links) — then raises p(stop) as the answer locks in.

Chain-of-Thought520 tok
State of Thought198 tok
—
fewer tokens generated
—
lower end-to-end latency
582
param controller, frozen backbone

One tiny controller reads the endogenous state to organize evidence and stop early — consistent gains across 3 LLMs + 2 VLMs.

From external control to endogenous state-conditioned reasoning
Figure. Paradigm shift: from external control to endogenous state-conditioned reasoning.

Why This Paradigm Matters

  • External scripts are rigid: fixed reasoning formats are brittle across heterogeneous tasks.
  • Search-heavy methods are costly: quality gains often depend on large sampling and high latency.
  • SoT changes the control variable: reasoning is driven by endogenous state, not external templates.
  • Result: better quality-efficiency frontier with a reusable closed-loop controller.
SoT performance across models, tasks, and efficiency metrics
Figure. SoT performance across models, tasks, and efficiency metrics.

3 + 2

LLMs + VLM scales

16

Text benchmarks (+3 multimodal)

62.6%

Token reduction

44.6%

Latency reduction

1.34×

Quantitative gain factor

1.62×

General gain factor

1.76×

Symbolic/code gain factor

2.51×

Long-context gain factor

582

Controller parameters

84.1%

Black-box judge agreement

Abstract

Test-time compute improves LLMs, but existing reasoning paradigms rely on externally imposed control — fixed reasoning programs or costly search in constrained spaces — which limits both generalization and efficiency. We propose State of Thought (SoT), a paradigm that enables endogenous reasoning, with the model's own internal reasoning state governing how reasoning unfolds.

Concretely, SoT extracts a compact dynamics-geometric state from the model's internal information transfer and uses a 582-parameter controller on frozen backbones to selectively activate the historical reasoning support useful under the current state — framing reasoning as a state-conditioned process over evidence rather than an externally prescribed token chain.

Across quantitative (1.34×), general (1.62×), symbolic-and-code (1.76×), and long-context (2.51×) reasoning on 3 LLMs and 16 datasets, SoT consistently improves accuracy while reducing generated tokens by 62.6% and latency by 44.6%. Across 2 VLM scales and 3 tasks, it adds 3.8 points over reasoning baselines with 74.9% fewer completion tokens and 73.5% lower latency than search-based methods.

Method

1) Read the endogenous state

At each sentence-level step, SoT reads a compact dynamics-geometric state mt ∈ ℝ4 from the backbone's internal information transfer: geometry (within-step dispersion), progress (step displacement), direction (consistency between steps), and uncertainty (predictive entropy).

Output: a control signal reflecting the current reasoning regime.

2) Select state-matched evidence (𝒮)

The evidence-organization operator re-selects only the sparse subset of past sentence-level reasoning units that match the current state, and re-inserts them in order — sharpening the evidence carried into the next step, rather than broadening search.

Output: a sparse active context for the next reasoning step.

3) State-conditioned stopping (𝒯)

The same state feeds a stopping operator, pstop = σ(𝒯(mt)): reasoning ends when enough support has accumulated, so depth emerges from state rather than a fixed schedule.

Output: a closed loop balancing quality and efficiency.

State trajectories across reasoning paradigms
Figure. Different reasoning paradigms occupy distinguishable endogenous state regions; SoT spans broader adaptive regimes.
State-conditioned activation and stopping patterns
Figure. Different endogenous states induce different sparse evidence activation patterns and stop tendencies.

Mechanism Insight

  • Tiny, frozen-backbone controller: only a 582-parameter interface is trained (a 16×32 selection head + a 4→1 stopping head); the backbone stays frozen. The selector is stable across problem-held-out refits (AUC 0.779 ± 0.005).
  • Generalization source: SoT transfers as a control principle, not as a dataset-specific prompt recipe.
  • Efficiency source: compute is redirected to state-matched support, rather than uniformly longer chains.
  • Interpretability: state clusters align with distinct evidence-selection and stop behaviors.

Main Results

Main Result Table (Llama-3.1-8B · Quantitative + Symbolic/Code)

Category Method GSM8K MATH DROP QS avg FOLIO ProofWriter BBH-Temporal HumanEval MBPP S&C avg
GreedyVanilla79.058.26.253.929.626.026.025.055.233.0
ReasoningCoT80.864.21.755.54.924.570.025.051.631.2
ReasoningPS71.240.01.443.44.915.255.010.022.017.6
ReasoningSR75.844.815.650.432.531.046.025.036.432.9
ReasoningSC83.654.01.653.216.828.552.010.035.227.3
ReasoningCB48.036.519.837.140.435.834.030.049.238.6
ReasoningMCTS55.233.513.237.536.031.530.035.049.236.7
MemoryH2O79.859.22.453.66.427.848.010.042.026.3
MemorySNAP80.460.22.354.18.426.255.015.041.227.3
MemorySTREAM70.444.21.344.41.531.81.00.05.613.0
LatentCOCO81.460.03.754.845.345.251.020.047.242.5
RL-BasedGRPO0.00.50.30.22.02.27.00.00.01.8
OursSoT81.864.840.565.849.864.878.059.146.858.4

Main Result Table (Llama-3.1-8B · General + Long-Context)

Category Method CommonsenseQA StrategyQA BoolQ MMLU RACE GU avg HotpotQA NarrativeQA LongBench MultiFieldQA LCR avg
GreedyVanilla48.068.859.245.650.754.27.020.132.716.2
ReasoningCoT32.831.537.234.050.336.32.04.526.97.1
ReasoningPS30.518.221.531.242.028.11.73.020.45.3
ReasoningSR57.064.072.262.462.763.617.921.725.520.6
ReasoningSC28.534.543.231.045.735.81.94.325.86.8
ReasoningCB46.071.582.251.256.061.115.527.526.221.8
ReasoningMCTS42.065.582.246.449.757.013.527.127.821.0
MemoryH2O33.530.839.235.019.732.42.02.00.01.7
MemorySNAP33.531.038.534.616.731.82.02.00.01.7
MemorySTREAM34.826.238.822.40.025.62.50.00.01.1
LatentCOCO40.563.570.042.843.352.04.410.834.611.9
RL-BasedGRPO10.53.89.88.212.08.70.61.15.31.6
OursSoT69.568.882.365.067.070.424.939.034.731.9

Extended Experimental Modules

Click each module to show the corresponding experimental table and interpretation.

Backbone Quantitative Reasoning avg Symbolic and Code avg General Understanding avg Long-Context avg
Qwen2.5-14B (SoT)66.176.881.425.2
Mixtral-8x7B (SoT)46.351.772.623.5
  • Qwen2.5-14B shows especially strong General Understanding performance across all five GU datasets.
  • Mixtral-8x7B remains near-best in quantitative tasks and leads strongly in symbolic, GU, and long-context averages.
  • The same SoT controller transfers across markedly different backbone architectures.
Scale Task SoT acc. SoT tokens SoT latency (s) BoN acc. BoN tokens
Qwen2.5-VL-7BA-OKVQA87.51914.484.5691
Qwen2.5-VL-7BAI2D86.72155.488.7961
Qwen2.5-VL-7BM³CoT87.31984.586.7919
Qwen2.5-VL-32BA-OKVQA85.542523.285.01322
Qwen2.5-VL-32BAI2D90.036419.287.31680
Qwen2.5-VL-32BM³CoT88.040920.185.31633
  • SoT is best in 5 of 6 model–task settings, averaging 87.5% accuracy.
  • On 7B AI2D it trails BoN by 2.0 points but uses 4.5× fewer completion tokens; at 32B it exceeds the strongest alternative by 0.5–0.7 points on all three tasks with 3.1–4.6× fewer tokens.
  • Across both scales: +3.8 accuracy points over reasoning baselines, with 74.9% fewer completion tokens and 73.5% lower latency than search-based methods.
Variant Quantitative avg Symbolic and Code avg General Understanding avg Long-Context avg
Top-3 baseline average53.244.061.524.2
SoT-Training-free59.050.152.026.9
SoT-Embedding59.049.851.927.9
  • Even without full internal access, SoT variants remain competitive and often exceed Top-3 baseline averages in key domains.
  • The largest relative resilience appears in long-context tasks, where trajectory-level organization is critical.
  • These results support that the mechanism is not tied to one specific implementation interface.

Performance Interpretation

SoT consistently outperforms strong baselines across heterogeneous reasoning regimes. This pattern suggests the gain is structural: changing the control variable from external token programs to endogenous state-conditioned evidence organization.

On Llama-3.1-8B, SoT leads all four domain averages and is best or tied-best on 13 of 16 datasets, with an average domain-level gain of +10.8 points over the strongest competitor. Against the best non-SoT baseline per domain it gains +10.3 / +15.9 / +6.8 / +10.1 points (quantitative / symbolic-and-code / general / long-context) — robust gains without larger search budgets.

Efficiency & Trade-off

Tradeoff on Qwen2.5-14B
Figure. Trade-off on Qwen2.5-14B.
Tradeoff on Mixtral-8x7B
Figure. Trade-off on Mixtral-8x7B.
Accuracy-efficiency trade-off on Llama-3.1-8B by method family
Figure. Accuracy–efficiency trade-off on Llama-3.1-8B (by method family).
Tradeoff on Qwen2.5-VL 7B and 32B
Figure. Trade-off on Qwen2.5-VL (7B / 32B) — SoT stays at high accuracy with far fewer tokens.

Llama Efficiency Summary (Domain Average)

Method QS Tok / Lat S&C Tok / Lat GU Tok / Lat LCR Tok / Lat Avg Tok Avg Lat (s)
CoT271.6 / 17.5369.8 / 27.5212.9 / 12.8198.1 / 19.6263.119.4
Self-Consistency736.8 / 43.1953.9 / 63.5626.5 / 37.6533.8 / 48.9712.848.3
Constrained Beam423.9 / 22.0363.9 / 19.4341.7 / 16.8283.7 / 19.5353.319.4
MCTS422.1 / 21.4364.7 / 19.4335.2 / 16.5286.6 / 20.5352.119.4
H2O226.0 / 16.7264.4 / 34.6185.6 / 15.8164.6 / 21.5210.222.1
COCO171.5 / 11.5214.6 / 20.487.2 / 6.7131.5 / 30.3151.217.2
SoT255.7 / 5.8302.2 / 9.574.9 / 1.8263.0 / 22.0223.99.8

Efficiency Insight

  • Search-heavy baselines often move to high-token and high-latency zones.
  • Memory-only compression may cut tokens but can lose state-relevant support and hurt quality.
  • SoT frontier shift: quality gains and cost reduction are achieved jointly, not by a simple trade.

Case Studies

Case-Level Insight

Across arithmetic, long-context QA, and multimodal chart reasoning, SoT exhibits a shared pattern: it keeps only support still needed for the next decision, suppresses obsolete context, and raises stop readiness when sufficient evidence has accumulated.

Representative SoT case trajectory
Figure. Representative SoT case with step-level active evidence and stop probability.
Arithmetic case study
Figure. Arithmetic case: staged decomposition with non-monotone evidence carry.
Long-context case study
Figure. Long-context case: sparse retrieval after exploration stabilizes contrastive reasoning.
VLM case study
Figure. Multimodal case: the same state-conditioned interface transfers to chart-grounded reasoning.

Resources