Plug-and-Play Reasoning Router

LoTR: Logic-of-Thought Routing for Plug-and-Play Reasoning of LLMs

External reasoning paradigms prescribe how a model should reason, but leave its internal computation pathway fixed. LoTR strengthens any paradigm by routing the model's internal attention-head pathways to the logic state each step induces — with frozen weights, at inference time.

Zhiren Gong1,2, Ming Xiao4, Chau Yuen3, Wei Yang Bryan Lim1

1 College of Computing and Data Science, Nanyang Technological University 2 Interdisciplinary Graduate Programme, Nanyang Technological University 3 School of Electrical and Electronic Engineering, Nanyang Technological University 4 Department of Information Science and Engineering, KTH Royal Institute of Technology

LoTR overview: external paradigms, frozen LLM + probe, offline logic-state compilation, and online plug-in head routing
Figure. LoTR as plug-and-play reasoning through logic-conditioned internal pathway routing: offline it compiles a basis of logic states and per-state head templates; online a probe reads the current state and softly routes attention heads for the next step.
Llama macro score across eight paradigms: Identity vs LoTR
Figure. Llama-3.1-8B matched macro across the eight reasoning paradigms — LoTR (solid) vs Identity (dashed).
Matched-panel comparison: Identity, plug-in mean, and LoTR across three backbones
Figure. Matched panels across three backbones — Identity, four-method plug-in mean, and LoTR; lines show LoTR's margins over each.

Cross-Setting Effectiveness

Across 3 backbones, 10 benchmarks, and 8 reasoning paradigms, LoTR improves model-level matched macros by +3.60% relative on average. On the primary Llama-3.1-8B study (8 benchmarks × 8 paradigms) it gains +2.66 points (+5.37% relative), improving every benchmark-level macro and helping or tying 51 of 64 cells.

Almost-Free Efficiency

LoTR adds no extra samples and no extra search calls. On the exact 64-cell panel it changes total completion tokens by only +0.15% (final-call length −0.43%). The accuracy gain comes from logic-indexed routing, not from a larger decoding budget.

3

Backbones

10

Benchmarks

8

Reasoning paradigms

K=6

Logic states

+2.66

Llama macro gain (pts)

+5.37%

Relative gain

+0.15%

Token overhead (all calls)

+2.12

Qwen2.5-14B transfer (pts)

+0.93

Mixtral-8x7B transfer (pts)

+2.86

vs. plug-in mean (Llama)

Abstract

Reasoning paradigms prescribe external strategies — step-by-step, iterative revision, search-and-sampling — that guide how a model reasons. They rarely consider whether the model's internal computation pathway is well matched to the dynamic reasoning-logic state induced by the strategy, leaving part of each paradigm's potential unrealized.

We propose Logic-of-Thought Routing (LoTR), a plug-and-play module that strengthens existing paradigms through logic-conditioned internal pathway routing: it identifies dynamic logic states from the model's internal information transfer, couples them with attention-head pathway patterns, and uses a lightweight probe to infer the current logic-state mixture and softly route the corresponding pathway.

Across three backbones, ten benchmarks, and eight reasoning paradigms, LoTR improves model-level matched macros by 3.60% relative on average. On the primary eight-benchmark, eight-paradigm Llama study it gains 5.37% with only +0.15% total completion tokens, and it outperforms the mean of four plug-in intervention baselines on all three comparison panels — all with frozen backbone weights.

Method

LoTR · live watch one reasoning step get routed — the logic-state mixture re-blends cached head templates into the next step's gate

Ask

logic-state mixture qt  (K = 6)
internal pathway — fixed  vs  LoTR-routed  =  Σk qt(k) · Mk
same gate every step — mismatched
re-blended to the step’s logic
layer × attention headstep 1
+2.66 pts
reasoning accuracy on an 8×8 Llama study
+0.15% tokens
routing overhead, essentially free at run time
plug·and·play
frozen weights, no fine-tuning or retraining

Each sentence boundary, a lightweight probe reads the current logic-state mixture; LoTR blends the cached per-state head templates into a soft gate over attention heads for the next step. Same prompt, same weights — only the internal pathway adapts.

1) Offline Logic Basis (K=6)

Pool sentence-level internal readouts from all paradigms and cluster them into a compact basis of six recurring logic states, with distance-based soft targets.

Output: six logic-state centroids + soft state targets.

2) Layer-wise Probes

Train lightweight per-layer linear probes that read a non-exclusive state mixture from the current sentence representation, supervised by soft targets.

Output: online logic-state occupancy at each routed layer.

3) Logic–Pathway Templates

Learn one head-gating template per state — preserving heads that help successful reasoning, damping those tied to failures, with an identity regularizer — while the LM stays frozen.

Output: cached per-state templates over the layer×head grid.

Online Plug-in Routing

As reasoning unfolds, LoTR evaluates the probes only at sentence boundaries. It composes the pathway for the next step by mixing the cached state templates according to the probe activations, applying the resulting soft gates directly to attention-head outputs. The external prompt or search procedure is left completely unchanged — only the internal execution pathway adapts, step by step, to the logic the step actually requires.

Logic states across reasoning paradigms
Figure. Different reasoning paradigms occupy different regions and transition patterns over the same latent logic-state space.
Logic-state occupancy by paradigm (K=6)
Figure. Occupancy of the six logic states by paradigm — some paradigms concentrate on a few states, others spread across many.
Per-state head-gating templates over the layer-head grid
Figure. Each logic state couples to a distinct head-gating template over the layer × attention-head grid.
Stable pathway patterns under different logic states
Figure. Distinct recurring logic states couple to distinct, reusable head-level routing patterns.

Method-Level Insight

  • Logic states are recurring, probe-identifiable execution regimes inferred from sentence-level internal transfer — not static labels.
  • State readout stays non-exclusive (sigmoid, not simplex), so several execution regimes can co-activate within one step.
  • The online controller is lightweight: one probe evaluation per layer at each sentence boundary, then a weighted mix over cached templates.

Main Results

Llama-3.1-8B Paradigm Macros (Accuracy %, benchmarks weighted equally)

Paradigm Identity + LoTR Δ (pts)
Vanilla47.948.0+0.13
Chain-of-Thought49.954.5+4.61
Plan-and-Solve48.049.5+1.54
Self-Refine52.756.8+4.13
Self-Consistency47.552.1+4.64
Best-of-N45.348.3+3.02
Constrained Beam55.657.6+1.97
Mini-MCTS49.650.8+1.25
Overall macro49.552.2+2.66

Improves 44 of 64 cells and ties 7 more (51/64 helped-or-tied); every benchmark-level macro improves.

Llama-3.1-8B Full Benchmark Table (Identity vs LoTR)

Paradigm Variant BoolQ DROP HumanEval QASC ScienceQA MMLU NarrativeQA FOLIO Average
VanillaIdentity80.07.968.038.468.046.025.249.547.9
Vanilla+ LoTR80.07.670.038.468.045.624.949.548.0
Chain-of-ThoughtIdentity75.511.354.038.473.659.242.844.649.9
Chain-of-Thought+ LoTR79.513.958.048.476.458.449.152.554.5
Plan-and-SolveIdentity72.510.734.045.278.448.046.448.548.0
Plan-and-Solve+ LoTR74.014.032.044.481.653.249.347.549.5
Self-RefineIdentity72.527.148.055.277.659.639.741.652.7
Self-Refine+ LoTR73.531.758.062.083.260.443.941.656.8
Self-ConsistencyIdentity72.014.342.046.474.846.041.742.647.5
Self-Consistency+ LoTR80.017.940.047.682.454.443.051.552.1
Best-of-NIdentity72.513.444.044.877.252.825.032.745.3
Best-of-N+ LoTR78.016.250.048.478.057.629.628.748.3
Constrained BeamIdentity78.533.362.048.072.450.450.949.555.6
Constrained Beam+ LoTR82.031.562.049.677.656.849.751.557.6
Mini-MCTSIdentity79.520.552.040.066.848.047.242.649.6
Mini-MCTS+ LoTR79.523.458.040.867.251.244.841.650.8

Metrics are each benchmark's native score: accuracy (BoolQ, QASC, ScienceQA-text, MMLU, FOLIO), token-level F1 (DROP, NarrativeQA), and Pass@1 (HumanEval).

Click to Explore More Experimental Blocks

Backbone Evaluation Identity LoTR Δ W/T/L
Llama-3.1-8B8 datasets × 8 paradigms49.552.2+2.6644/7/13
Qwen2.5-14B6 datasets; Self-Refine75.077.1+2.124/1/1
Mixtral-8x7B4 datasets; CoT35.636.5+0.933/1/0
Backbone Panel ITI ActAdd STIR RISER LoTR Δ vs mean
Llama-3.1-8B5 datasets; CoT53.052.851.851.955.2+2.86
Qwen2.5-14B6 datasets; SR77.076.676.576.377.1+0.51
Mixtral-8x7B4 datasets; CoT37.036.035.736.736.5+0.16

Top: matched Identity–LoTR macros with paired win/tie/loss counts. Bottom: all methods on shared examples per backbone; LoTR beats the four-method mean on every panel.

Benchmark Type Metric Role in evaluation
BoolQBoolean QAAccuracyYes/no factual verification.
DROPDiscrete reasoning over textToken-level F1Numerical multi-span reading.
HumanEvalCode generationPass@1Program correctness under wrappers.
QASCMulti-hop science QAAccuracyTwo-fact composition.
ScienceQA-textScience QAAccuracyGrounded scientific reasoning.
MMLUBroad-domain knowledgeAccuracyCross-domain academic reasoning.
NarrativeQALong-form QAToken-level F1Narrative-level reasoning fidelity.
FOLIOFormal logicAccuracyLogical entailment consistency.
CODAHCommonsenseAccuracyAdversarial commonsense (Qwen panel).
HellaSwagCommonsense completionAccuracySentence completion (Qwen panel).

Efficiency

Token Efficiency on the Primary Llama Panel

Measure Identity LoTR Relative change
Final call (completion tokens)141.3140.7−0.43%
All calls (work proxy for multi-call scaffolds)382.4383.0+0.15%

Values are completion tokens, macro-averaged over eight datasets and eight paradigms.

Interpretation

  • LoTR changes internal routing without adding sampling or search calls: accuracy rises +2.66 points while total completion tokens move only +0.15%.
  • The gain therefore comes from logic-indexed pathway reuse, not from a larger decoding budget or brute-force expansion.
  • Wall-clock timings are excluded because measurements come from heterogeneous shared experiment servers; tokens are the controlled decoding-work proxy.

Ablation & Robustness

Aligned Mechanism Ablation (Llama-3.1-8B)

Variant BoolQ DROP ScienceQA-text Avg.
Identity routing75.011.773.453.4
Single state (true K=1)72.916.371.753.6
Full LoTR (K=6)79.413.976.256.5

Averaged over Vanilla, CoT, SC, and Best-of-N on matched items. Full LoTR improves the panel macro by +3.12 over Identity and +2.87 over a true single-state controller — the gain is not recovered by one globally routed template.

K ablation: Identity vs single-state vs LoTR K=6
Figure. State differentiation matters: collapsing the six states into one erases most of the gain.
Probe-logit scale sensitivity
Figure. Probe-scale sensitivity: the LoTR−Identity gain peaks at the deployed β = 0.50 and stays positive across a broad range.

Ablation Insight

  • State diversity is load-bearing: a distinguishable six-state basis keeps the pathway templates discriminative.
  • Conditioning on the inferred state is what carries the gains — steering without reading the state fades to a generic tweak.
  • LoTR is not knife-edge tuned: it stays positive across a broad probe-scale range and different state counts.

Case Study

Under the same external wrapper, LoTR keeps step-level constraint verification and avoids premature answer commitment. The trajectory shows that improvements come from better internal pathway alignment, not from changing the external reasoning script itself.

Qualitative case study with and without LoTR
Figure. Qualitative case study: with vs without LoTR under the same external reasoning paradigm.

Case Interpretation

LoTR does not replace the external wrapper; it reduces internal pathway mismatch. In practice this means fewer premature commitments and better late-step constraint verification under the same prompt scaffold.

Resources

Video Tutorial

A narrated, subtitled ~3-minute walkthrough of the problem, the idea, the method and the results. Watch it →

Paper

Coming soon.

Code

Coming soon.

Citation

Formal BibTeX will be released soon.