Skip to main content

Ablation & Variance Decomposition

4.5 Ablation & Variance Decomposition​

Ablation Testing​

The ablation framework tests whether each declared mechanism actually contributes to model realism. For every mechanism, the runner disables it (via the toggle_fn), re-runs Monte Carlo evaluation, and compares the resulting features to the baseline. If the effect size is below the decorative threshold (default 10%), the mechanism is flagged as decorative --- it adds complexity without improving realism.

@dataclass
class MechanismDeclaration:
name: str # Mechanism identifier
hypothesis: str # Causal hypothesis (why this should help)
toggle_fn: Callable # Callable that disables the mechanism on the model
predicted_effect: Dict[str, Tuple[str, float, float]]
# Maps feature_name -> (direction, min_change, max_change)
# direction: "increase", "decrease", or "change"
# e.g., {"volatility_clustering": ("decrease", 0.3, 0.7)}
evidence: str = "" # Supporting evidence or references

The predicted_effect field is critical: it forces the modeller to state quantitatively what they expect to happen when a mechanism is removed. If the actual effect doesn't match the prediction (wrong direction or outside the predicted magnitude range), it signals a misunderstanding of the mechanism's role.

class AblationRunner:
def __init__(self, decorative_threshold: float = 0.10) -> None

def run(
self,
model_cls: Type[ABMModel],
config: Dict[str, Any],
mechanisms: List[MechanismDeclaration],
feature_fn: Callable,
seeds: int = 10,
pairwise: bool = False, # Also test mechanism pairs for interactions
mc_spec: Optional[Dict] = None,
trace_fields: Optional[List[str]] = None,
) -> AblationReport

The AblationReport provides:

Method / PropertyDescription
decorative_mechanismsList of mechanism names with effect size below threshold
mechanism_importance()Dict mapping mechanism names to their max absolute effect size
summary_table()Formatted table with baseline vs ablated features per mechanism
pairwise_resultsInteraction effects between mechanism pairs (if pairwise=True)

When pairwise=True, the runner also disables mechanisms in pairs and compares the combined effect to the sum of individual effects. Superadditive interactions (combined effect > sum of individual effects) indicate mechanism synergy; subadditive interactions indicate redundancy.

Variance Decomposition​

For LLM-ABMs, the NestedMCRunner decomposes total Monte Carlo variance into two components: seed variance (structural model randomness from the PRNG) and within-seed variance (stochasticity from LLM API responses or other non-deterministic sources). This is a nested experimental design analogous to a two-level random-effects ANOVA.

class NestedMCRunner:
def __init__(
self,
n_seeds: int = 20, # Outer level: independent seeds
n_inner_reps: int = 5, # Inner level: replications per seed
max_workers: Optional[int] = None,
) -> None

def run(
self,
model_cls: Type[ABMModel],
config: Dict[str, Any],
feature_fn: Callable,
trace_fields: Optional[List[str]] = None,
) -> NestedMCResult

The NestedMCResult contains a decompose(feature_name) method that returns a NestedANOVAResult:

@dataclass
class NestedANOVAResult:
feature_name: str
grand_mean: float
components: List[VarianceComponent] # [seed, within, total]
f_statistic: float
f_p_value: float
icc: float # Intra-class correlation coefficient
n_seeds: int
n_inner_reps: int
reliability: float

def summary_table(self) -> str
def diagnosis(self) -> str

Each VarianceComponent contains: source ("seed", "within", or "total"), sum_of_squares, degrees_of_freedom, mean_square, variance_estimate, and variance_fraction.

The diagnosis() method returns a human-readable interpretation:

ICC RangeDiagnosisMeaning
ICC > 0.75"EXCELLENT"Seed variance dominates. Model structure drives results. Within-seed noise is negligible.
0.50 < ICC < 0.75"MODERATE"Both sources contribute. Model is usable but within-seed variance is non-trivial.
0.25 < ICC < 0.50"POOR"Within-seed variance is substantial. Consider simplifying prompts or increasing temperature=0 cache hits.
ICC < 0.25"UNACCEPTABLE"Within-seed variance dominates. Model output is driven by noise, not model structure. Redesign mechanisms.

A PowerAnalysisResult is also available via power_analysis(), which recommends the number of seeds and inner replications needed to achieve a target statistical power (default 0.80) for detecting a given effect size.