Ablation & Variance Decomposition
4.5 Ablation & Variance Decomposition
Ablation Testing
The ablation framework tests whether each declared mechanism actually contributes to model realism. For every mechanism, the runner disables it (via the toggle_fn), re-runs Monte Carlo evaluation, and compares the resulting features to the baseline. If the effect size is below the decorative threshold (default 10%), the mechanism is flagged as decorative --- it adds complexity without improving realism.
@dataclass
class MechanismDeclaration:
name: str # Mechanism identifier
hypothesis: str # Causal hypothesis (why this should help)
toggle_fn: Callable # Callable that disables the mechanism on the model
predicted_effect: Dict[str, Tuple[str, float, float]]
# Maps feature_name -> (direction, min_change, max_change)
# direction: "increase", "decrease", or "change"
# e.g., {"volatility_clustering": ("decrease", 0.3, 0.7)}
evidence: str = "" # Supporting evidence or references
The predicted_effect field is critical: it forces the modeller to state quantitatively what they expect to happen when a mechanism is removed. If the actual effect doesn't match the prediction (wrong direction or outside the predicted magnitude range), it signals a misunderstanding of the mechanism's role.
class AblationRunner:
def __init__(self, decorative_threshold: float = 0.10) -> None
def run(
self,
model_cls: Type[ABMModel],
config: Dict[str, Any],
mechanisms: List[MechanismDeclaration],
feature_fn: Callable,
seeds: int = 10,
pairwise: bool = False, # Also test mechanism pairs for interactions
mc_spec: Optional[Dict] = None,
trace_fields: Optional[List[str]] = None,
) -> AblationReport
The AblationReport provides:
| Method / Property | Description |
|---|---|
decorative_mechanisms | List of mechanism names with effect size below threshold |
mechanism_importance() | Dict mapping mechanism names to their max absolute effect size |
summary_table() | Formatted table with baseline vs ablated features per mechanism |
pairwise_results | Interaction effects between mechanism pairs (if pairwise=True) |
When pairwise=True, the runner also disables mechanisms in pairs and compares the combined effect to the sum of individual effects. Superadditive interactions (combined effect > sum of individual effects) indicate mechanism synergy; subadditive interactions indicate redundancy.
Variance Decomposition
For LLM-ABMs, the NestedMCRunner decomposes total Monte Carlo variance into two components: seed variance (structural model randomness from the PRNG) and within-seed variance (stochasticity from LLM API responses or other non-deterministic sources). This is a nested experimental design analogous to a two-level random-effects ANOVA.
class NestedMCRunner:
def __init__(
self,
n_seeds: int = 20, # Outer level: independent seeds
n_inner_reps: int = 5, # Inner level: replications per seed
max_workers: Optional[int] = None,
) -> None
def run(
self,
model_cls: Type[ABMModel],
config: Dict[str, Any],
feature_fn: Callable,
trace_fields: Optional[List[str]] = None,
) -> NestedMCResult
The NestedMCResult contains a decompose(feature_name) method that returns a NestedANOVAResult:
@dataclass
class NestedANOVAResult:
feature_name: str
grand_mean: float
components: List[VarianceComponent] # [seed, within, total]
f_statistic: float
f_p_value: float
icc: float # Intra-class correlation coefficient
n_seeds: int
n_inner_reps: int
reliability: float
def summary_table(self) -> str
def diagnosis(self) -> str
Each VarianceComponent contains: source ("seed", "within", or "total"), sum_of_squares, degrees_of_freedom, mean_square, variance_estimate, and variance_fraction.
The diagnosis() method returns a human-readable interpretation:
| ICC Range | Diagnosis | Meaning |
|---|---|---|
| ICC > 0.75 | "EXCELLENT" | Seed variance dominates. Model structure drives results. Within-seed noise is negligible. |
| 0.50 < ICC < 0.75 | "MODERATE" | Both sources contribute. Model is usable but within-seed variance is non-trivial. |
| 0.25 < ICC < 0.50 | "POOR" | Within-seed variance is substantial. Consider simplifying prompts or increasing temperature=0 cache hits. |
| ICC < 0.25 | "UNACCEPTABLE" | Within-seed variance dominates. Model output is driven by noise, not model structure. Redesign mechanisms. |
A PowerAnalysisResult is also available via power_analysis(), which recommends the number of seeds and inner replications needed to achieve a target statistical power (default 0.80) for detecting a given effect size.