The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

Explainable & Ethical AI
Published: arXiv: 2606.27510v1
Authors

Sankaran Vaidyanathan David Arbour Aaron Mueller Scott Niekum David Jensen

Abstract

Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating its natural indirect effect (NIE). Re-deriving the activation patching estimand from causal mediation analysis, we find that the NIE does not solely capture the causal effect through the specific component. It also contains interaction effects (INT) that measure how much the component's causal effect itself depends on the state of other components in the model. A natural response may be to try to eliminate INT by adjusting the estimator or unit of analysis, but each of these potential remedies has predictable failure modes. We demonstrate these failure modes in the GPT-2 IOI circuit; components whose causal importance is conditional on the state of other components are either invisible or artificially inflated, and INT variance explains the previously documented instability of faithfulness scores. We prove that INT scales with the distance between clean and patched component activations, is negligible when the model is locally affine, and decomposes combinatorially into pairwise and higher-order group interactions. Despite its inevitability, INT is not a nuisance to be eliminated, but rather a diagnostic for interpretability studies. Its individual and group-level magnitude and sign signal when causal conclusions are prompt-dependent, and when greedy NIE-based component ranking will miss mechanisms only discoverable through combinatorial search.

Paper Summary

Problem
The main problem addressed in this research paper is the limitations and failures of activation patching, a widely used method for mechanistic interpretability in neural networks. Activation patching is used to attribute causal responsibility to individual components of a model, but it has been shown to miss important components and produce inconsistent results.
Key Innovation
The key innovation of this work is the identification of "interaction effects" (INT) in activation patching. INT refers to the way in which the causal effect of a component depends on the state of other components in the model. The researchers demonstrate that INT is a natural consequence of activation patching and that it can lead to predictable failure modes.
Practical Impact
This research has important practical implications for the field of mechanistic interpretability. By recognizing the role of INT in activation patching, researchers can better understand the limitations of this method and develop new approaches that take into account the complex interactions between components. This can lead to more accurate and reliable causal attribution, which is essential for understanding how neural networks work and for developing more robust and transparent AI systems.
Analogy / Intuitive Explanation
Imagine a car with multiple engines, each contributing to the overall performance of the vehicle. Activation patching is like trying to measure the contribution of each engine individually, without considering how they interact with each other. INT represents the way in which the performance of one engine depends on the state of the other engines. For example, if one engine is malfunctioning, the performance of the other engines may be affected, even if they are not directly related to the malfunctioning engine. By recognizing the role of INT, researchers can develop more nuanced understanding of how the engines (or components) interact and how they contribute to the overall performance of the car (or model).
Paper Information
Categories:
cs.LG cs.CL
Published Date:

arXiv ID:

2606.27510v1

Quick Actions