Mechanistic interpretability · self-repair · circuit completion

Conditional Co-Ablation

A circuit that explains the intact model can be incomplete under the very ablation used to test it. CoAx asks what grows once the primary set is gone — and completes the circuit.

Zhiren Gong1,2He Lu3Tiantong Wang1Yichi Zhang2,5Yixin Wang1Zihao Zeng1Ming Xiao6Chau Yuen4Wei Yang Bryan Lim1
1College of Computing and Data Science, Nanyang Technological University 2Interdisciplinary Graduate Programme, Nanyang Technological University 3School of Civil and Environmental Engineering, Nanyang Technological University 4School of Electrical and Electronic Engineering, Nanyang Technological University 5Lee Kong Chian School of Medicine, Nanyang Technological University 6Information Science & Engineering, KTH Royal Institute of Technology, Sweden
strongest intact-state baseline0.665
backup-recovery ROC-AUC →
CoAx (ours)0.941
From self-repair to conditional circuit completion
From self-repair to conditional circuit completion (GPT-2-small, IOI). (a) Ablating the primary circuit recruits a dormant backup branch while behavior is largely preserved. (b) CoAx compares each candidate's ablation effect with the primary set present and removed; that contrast distinguishes dormant backups and exposes the all-order effect-change aggregate linking each candidate to the removed set. (c) Conditioning exposes the documented backups, improves recovery over intact-state baselines, and substantially completes the missing branch.

A component's importance is not a property it carries alone. In a self-repairing circuit it becomes visible only relative to what has been removed — and that single shift is what CoAx measures, turning a blind spot into a signal.

01
The problem

The ablation we use to test a circuit changes the circuit

A behavior is localized to a circuit, and the standard test is to ablate a component and measure the resulting output change. Self-repair breaks that test: ablating a primary component can activate a dormant backup, changing which circuit carries the behavior in the first place.

Two wrong conclusions from one measurement

Remove a primary head and a dormant backup takes over. The output barely moves, so the primary reads as unimportant — the model quietly repaired the damage. And the backup was silent on the intact model, so it reads as unimportant too.

So a circuit can faithfully explain the original model and still be incomplete under the very ablation used to test it. That raises a conditional question: which components become causally important only after a known part of the circuit is removed?

0.28

of a 2.57 margin
is all that moves when the three primary name-movers are ablated — eight documented backups compensate

1.9×

non-additive
ablating the primaries together with those backups costs 1.9× the sum of their individual effects

The IOI circuit in token x layer context
The documented IOI circuit, in token×layer context. Given “Mary and John went to the store; John gave a drink to __”, the primary name-movers write “Mary” to the output (orange); the backup branch (blue dashed) is the part the intact model hides.

Scoring on the intact model cannot tell a dormant backup from an irrelevant head — both are silent. So we change the question.

02
The idea

Ask what grows once the primary set is gone

We formalize the gap as conditional circuit completion: given a primary set S supplied by the analyst, identify the components whose causal importance emerges only after S is removed. CoAx scores each candidate by the growth of its ablation effect under that conditioning, measured in the output distribution's own (Fisher) geometry:

$$\operatorname{comp}_u(S)\;=\;\underbrace{\mathcal{E}\!\left(\delta z_{u\mid S}\right)}_{\text{effect once }S\text{ is removed}}\;-\;\underbrace{\mathcal{E}\!\left(\delta z_{u}\right)}_{\text{effect alone (intact state)}}$$
CoAx · live Self-repair lets a backup head sit silent until it is needed, so single-head ablation rates it like noise. CoAx ablates the primary set first, then ranks every head by how much its effect grows.

The problem Single-head ablation on the intact model

Ablate each candidate alone. A dormant backup barely moves the output, so it scores like an irrelevant head.

Nothing clears the detection bar → backups missed, circuit incompleteness 0.75.

The fix Conditional co-ablation ablate S

Remove S, then re-score by the growth Δ of each head's ablation effect. Self-repair forces the backups to carry the behaviour.

The two backups surge past the bar → recovered at 0.941 ROC-AUC.

—
ROC-AUC recovering GPT-2 IOI self-repair backups
—→—
circuit incompleteness, before → after CoAx
forward·only
label-free ranking, no gradients or training

On the intact model a dormant backup and an irrelevant head look identical. CoAx conditions on removing the primary set S, so a backup's suppressed contribution reappears as measurable growth — recovering self-repair components that intact-state scores cannot see.

Blind-spot-proof

A dormant backup and an irrelevant head look identical on the intact model. CoAx separates them: the backup's effect grows once its primary is gone; the irrelevant head's does not.

Cheap & label-free

Two shared reference passes plus two per candidate: O(|U|) forward evaluations. No backward pass, no task gradient, no labels.

A completion, not a rewrite

The primary set can come from any intact-state method; CoAx completes that circuit by returning the branch it hides.

Why conditioning is the right instrument — two results

Proposition 1 · the blind spot is real

A perfectly dormant backup can be indistinguishable from an irrelevant component for every intact-state score — anything computed from a candidate's activations, output derivatives, or single-ablation effect on the original model. It is non-identifiability for the whole class, not weakness in one score.

Proposition 2 · what conditioning buys

The change in a candidate's ablation effect once S is removed is exactly the aggregate of every interaction order linking it to S. Enumerating those terms is combinatorial in |S|; conditioning returns the whole aggregate in O(|U|) forward evaluations.

Pairwise interaction over the IOI heads
The pairwise interaction structure over the IOI heads: the name-movers and their backups form a bright off-diagonal block — the interaction a per-head score cannot see, and the aggregate that conditioning returns without enumerating it.

The structure is there in the interaction signal. Does conditioning actually surface the documented backups?

03
Discovery

From the blind spot to the top of the ranking

On the GPT-2-small IOI circuit — the one with head-level backup ground truth — we supply the three primary name-movers, rank the other 141 heads by conditional growth, and score against the eight documented backups. Labels are used only for evaluation.

scorebackup ROC-AUC
AtP intact0.650
single ablation intact0.603
conditional gradient removed0.767
activation-space EAP-IG intact0.581
conditional energy removed0.758
AtP* GradDrop intact0.665
CoAx intact→removed0.941

The gain is the two-state contrast — not simply reading the intervened model alone. CoAx improves by 0.276 over the strongest intact-state attribution baseline and by 0.174 over the strongest removed-state control. Applying the same contrast to AtP* reaches 0.938, confirming that conditioning carries the central signal; CoAx measures the exact finite, full-distribution intervention effect without task gradients. Under the 8/141 class imbalance the separation is starker in average precision: 0.557 versus 0.098 for AtP*.

Holding the contrast fixed and replacing the Fisher geometry with plain ℓ2 still gives 0.905, so conditioning carries the main signal and the geometry sharpens it.

The component analysis also evaluates a normalized CoAx variant (0.984 AUC). It is not a separate baseline: signed growth is the default because it has the exact interaction decomposition, transfers directly across mechanisms, and avoids division by near-zero intact effects.

Intact-state versus conditional scoring
The same eight backups, two scores. Every head placed by intact-state scoring and by the CoAx score: the documented backups move from the middle of the ranking to the top.

Recovery tracks which circuit was removed

We search 5,000 random three-head sets and keep the 16 that best match the true primary set in behavioral damage, in primary-set Fisher displacement, and in depth. On held-out prompts those matched wrong sets recover the backups at only 0.40 ± 0.13 AUC, while the true primary set stays at about 0.94. The interventions are matched both in task-level damage and in output-space displacement measured in the same geometry as the score — and they still fail to expose the branch.

A second control rules out normalization: replaying the ablated forward pass with clean LayerNorm denominators raises backup AUC from 0.941 to 0.983. LayerNorm rescaling does not create the signal; removing it suppresses background growth and sharpens the separation.

A high rank is necessary, not sufficient. Are the surfaced heads mechanistically backups?

04
Mechanism

They wake up — and the wake-up is causal

Two stronger questions than ranking: does the backup response grow as more of the primary circuit is removed, and does preventing that response remove the repair?

Direct-logit hand-off
(a) Removing the primary name-movers shifts direct-logit attribution toward the documented backup branch while the IO−S margin is largely preserved.
Graded wake-up
(b) Documented backups and CoAx-selected heads both rise on the wake-up read-outs as more primaries are removed; matched random heads stay flat.
Causal load-bearingness
(c) Freezing each selector's top-8 at its intact-state outputs measures causal load-bearingness.
+0.57→+1.68

The hand-off is graded
the backup branch roughly triples its direct contribution to the answer, while the margin falls only 2.57 → 2.29

0.90

CoAx’s set is load-bearing
IOI-margin loss from freezing its top-8 — against 0.57 for the documented backups, 0.27 for a size-matched random set and 0.03 for conditional energy

One dormant backup, end to end
Head [10,6], traced end to end.

One head, two views

A single documented backup, [10,6], is silent on the intact model, yet from the answer position it already attends to the indirect-object name — the structural signature of a name-mover.

Its direct IO−S contribution grows from +0.07 intact to +0.21 once the primary name-movers are removed, while a random comparison head stays near zero. Invisible in the intact state, structurally a name-mover, causal only under conditioning.

Structural read of the name-mover branch
An independent, ablation-free structural read. Final-position attention to token roles over 96 IOI prompts: documented primary and backup name-movers attend preferentially to the IO name, controls do not. This read alone gives 0.96 backup AUC and correlates only weakly with the CoAx score (ρ = 0.09) — independent evidence for the same role.

If the branch is load-bearing, adding it back should repair the causal account that was incomplete.

05
Closing the loop

Completing the circuit

Recovery is not the goal in itself: the recovered components should repair the causal account that was incomplete under intervention. We take the same raw headline ranking used in discovery — no backup labels, no task gradient, no task-direction filter — and use its top eight heads as the completion.

Following Wang et al., completeness asks whether a circuit reproduces the full model's response to the same ablation. Writing D(C) for the gap between the full model's response to primary-set removal and the circuit's own, lower is better.

Circuit completion

circuitincompleteness D(C) ↓
primary circuit (backup branch removed)0.474 ± 0.030
documented circuit0.300 ± 0.027
+ conditional energy top-80.181 ± 0.027
+ CoAx top-80.236 ± 0.037
+ normalized CoAx variant top-80.258 ± 0.035
+ random top-80.499 ± 0.073

Adding the label-free CoAx set reduces incompleteness from 0.47 to 0.24. Conditional energy gives a smaller scalar gap but a much worse backup-role knockout, separating response magnitude from branch identity.

Backup-role knockout

top-upmargin dist. ↓output KL ↓unrel. KL ↓
documented backups0.000.000.09
+ CoAx0.990.170.14
+ normalized CoAx variant0.990.070.09
+ conditional energy1.710.460.47
+ co-activation1.790.470.15
+ CoAx next-82.031.380.86
+ AtP* GradDrop2.701.020.47
+ random1.240.470.24

Reversing the intervention: removing the selected set should reproduce the documented backup effect. A large drop alone would not identify the branch — disrupting unrelated computation also hurts IOI — so we track collateral KL on unrelated text as well.

One circuit, one model — does the conditional signal hold more broadly?

06
Generalization

Beyond one annotated circuit

Two tests without backup labels: does conditional growth agree with independently measured repair across mechanisms, and do the recovered sets complete circuits in other models?

Generalization across mechanisms and models
Conditional repair generalizes. (a) Signed growth has positive rank agreement with mechanism-matched held-out causal repair in 11 of 12 instances and in all four mechanism clusters; the bottom row is the hierarchical mean with a nested 95% CI. (b) CoAx-selected completions beat size-matched random completions on all eight non-GPT-2 models across six architecture families; hollow circles are each model's own next induction heads, a stronger role-matched control.
11 / 12

Across mechanisms
positive agreement with held-out causal repair over name-mover, S-inhibition, induction and duplicate-token clusters; hierarchical mean +0.123, nested 95% CI [+0.047, +0.201]

8 / 8

Across models
completion beats a size-matched random set on every non-GPT-2 model — Pythia 160M/410M/1.4B, GPT-Neo 1.3B, Gemma-2 2B, Qwen2.5 7B, OLMo-2 7B, Llama-3.1 8B — and beats each model's own next induction heads on 5 of 8

What a nearby control shows

Input-side co-activation gives a complementary intact-state view and nearly matches CoAx on backup-label AUC (0.93 versus 0.941). But freezing its selected set causes a much larger raw IOI-margin loss (1.97 versus 0.90) while matching the documented backup knockout substantially worse — it disrupts the wider computation rather than isolating the repair branch. High label overlap is not the same as isolating the mechanism.

Under the hood

What makes the score work — and what it costs

Conditioning carries the signal, the centered Fisher geometry sharpens it, and the price is forward passes only: no backward pass, no task gradient, no labels, and every candidate intervention is independent and batchable.

Diagnostics of the CoAx score
Diagnostics of the score. (a) Recovery is already near the full result with a small unlabeled calibration set — 16 prompts give 0.931, 32 give 0.937, and the 96-prompt main setting gives 0.941. (b) Output-geometry ablation under the same conditional contrast: plain ℓ2 reaches 0.905, centering alone 0.903 and probability weighting alone 0.909, while the full centered Fisher geometry reaches 0.941 — the two ingredients only pay off together. (c) As backups move off the task direction, full-distribution CoAx stays stable while the conditional gradient and an intact task-gradient proxy deteriorate; the latter stays near chance.
Resources

Paper, code, video, and citation

BibTeX
@misc{gong2026coax,
  title         = {Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits},
  author        = {Gong, Zhiren and Lu, He and Wang, Tiantong and Zhang, Yichi and Wang, Yixin
                   and Zeng, Zihao and Xiao, Ming and Yuen, Chau and Lim, Wei Yang Bryan},
  year          = {2026},
  eprint        = {2607.01940},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  note          = {Project page:
                   https://gongzhiren.github.io/Conditional-Co-Ablation-website}
}