Skip to main content

Observability & Safety

6.3 Observability & Safety​

OpenTelemetry​

The SDK provides built-in OpenTelemetry instrumentation for distributed tracing and metrics collection. When enabled, every simulation step, action execution, message send, and LLM call is traced with context propagation, enabling end-to-end observability in production deployments.

Five span types are created automatically:

Span NameScopeAttributes
simulation.setupFull setup(config) callmodel_class, seed, agent_count
simulation.stepEach step() calltick, agent_count, duration_ms
simulation.actionEach Action executionaction_name, agent_type, agent_count
message.sendMessage enqueuesender_id, msg_type, target_count
llm.callLLM API invocationmodel, prompt_length, tokens, cost_usd

Seven Prometheus-compatible metrics:

MetricTypeDescription
abm.simulation.stepsCounterTotal steps completed
abm.step.durationHistogram (ms)Step execution time distribution
abm.agents.countObservable gaugeCurrent agent count by type
abm.messages.sentCounterMessages sent per step
abm.llm.callsCounterLLM API calls by model name
abm.llm.latencyHistogram (ms)LLM API response time distribution
abm.llm.costCounter (USD)Cumulative LLM spend

Telemetry is enabled via settings.json:

{
"telemetry": {
"otel_enabled": true,
"service_name": "abm-lab",
"endpoint": "http://localhost:4317",
"export_protocol": "grpc",
"trace_sample_rate": 1.0,
"trace_llm_calls": true,
"collect_agent_counts": true,
"collect_step_timing": true,
"collect_llm_metrics": true
}
}

The OTelInstrumentation class registers itself as hooks on BEFORE_STEP, AFTER_STEP, BEFORE_ACTION, AFTER_ACTION, and ON_MESSAGE_SEND events, so instrumentation requires zero changes to model code. Traces export via OTLP to any compatible backend (Jaeger, Zipkin, Grafana Tempo). Metrics export to Prometheus or any OTLP metrics receiver.

Guardrails​

The guardrail system provides runtime safety checks that execute after every simulation step (via the AFTER_STEP hook). Four built-in guardrails cover common failure modes:

Guardrail__init__What It Checks
MaxAgentsGuardrail(max_agents: int = 100000)Agent population does not exceed limit (prevents runaway spawning)
BudgetGuardrail(max_usd: float = 10.0)Cumulative LLM spend does not exceed budget
DeterminismGuardrail(check_random_module: bool = True)Seed is configured; no random module usage detected
StateInvariantGuardrail(predicate: Callable, message: str)Custom predicate returns True (e.g., all balances non-negative)

Guardrails are registered on the model's GuardrailRegistry:

from simudyne.engine.guardrails import (
GuardrailRegistry, MaxAgentsGuardrail, BudgetGuardrail, StateInvariantGuardrail,
)

model.guardrails.add(MaxAgentsGuardrail(max_agents=50000))
model.guardrails.add(BudgetGuardrail(max_usd=5.0))
model.guardrails.add(StateInvariantGuardrail(
predicate=lambda ctx: all(
b.cash_buffer >= 0 for b in ctx.model.get_agents("Bank")
),
message="Negative cash buffer detected"
))

When enforce() is called (automatically by the hook), it runs all guardrails and raises GuardrailViolation with the list of violation messages if any check fails. The check_all() method returns violations as a list of strings without raising, enabling logging-only mode.

Checkpointing​

The CheckpointManager provides save/restore with deterministic resume. Checkpoints capture the full model state: agent states, accumulator values, link topology, RNG states (both agent and environment PRNGs), and configuration. The SHA-256 checksum ensures checkpoint integrity.

from simudyne.engine.checkpoint import CheckpointManager

manager = CheckpointManager(
directory="checkpoints/",
max_kept=5, # Keep only the 5 most recent
auto_interval=50, # Auto-save every 50 steps
enabled=True,
)
MethodSignatureDescription
save(model, step=None, path=None) -> strSave checkpoint, return file path
load(path, model_class=None, validate_checksum=True) -> ABMModelRestore model from checkpoint
list_checkpoints(directory=None) -> List[Dict]List all checkpoints with metadata
resume_from_checkpoint(path, additional_steps=0) -> ABMModelLoad and continue execution
auto_checkpoint_hook(ctx) -> NoneHook callback for automatic saving
from_settings(settings, model=None) -> CheckpointManagerCreate from settings.json

Auto-checkpointing is wired via the hook system:

from simudyne.engine.hooks import HookEvent

model.hooks.on(HookEvent.AFTER_STEP, manager.auto_checkpoint_hook)

This saves a checkpoint every auto_interval steps, automatically deleting older checkpoints beyond max_kept. Checkpoint metadata (step, timestamp, agent count, accumulator values, config, checksum) is stored as a JSON sidecar file alongside the pickle.