Observability & Safety
6.3 Observability & Safety
OpenTelemetry
The SDK provides built-in OpenTelemetry instrumentation for distributed tracing and metrics collection. When enabled, every simulation step, action execution, message send, and LLM call is traced with context propagation, enabling end-to-end observability in production deployments.
Five span types are created automatically:
| Span Name | Scope | Attributes |
|---|---|---|
simulation.setup | Full setup(config) call | model_class, seed, agent_count |
simulation.step | Each step() call | tick, agent_count, duration_ms |
simulation.action | Each Action execution | action_name, agent_type, agent_count |
message.send | Message enqueue | sender_id, msg_type, target_count |
llm.call | LLM API invocation | model, prompt_length, tokens, cost_usd |
Seven Prometheus-compatible metrics:
| Metric | Type | Description |
|---|---|---|
abm.simulation.steps | Counter | Total steps completed |
abm.step.duration | Histogram (ms) | Step execution time distribution |
abm.agents.count | Observable gauge | Current agent count by type |
abm.messages.sent | Counter | Messages sent per step |
abm.llm.calls | Counter | LLM API calls by model name |
abm.llm.latency | Histogram (ms) | LLM API response time distribution |
abm.llm.cost | Counter (USD) | Cumulative LLM spend |
Telemetry is enabled via settings.json:
{
"telemetry": {
"otel_enabled": true,
"service_name": "abm-lab",
"endpoint": "http://localhost:4317",
"export_protocol": "grpc",
"trace_sample_rate": 1.0,
"trace_llm_calls": true,
"collect_agent_counts": true,
"collect_step_timing": true,
"collect_llm_metrics": true
}
}
The OTelInstrumentation class registers itself as hooks on BEFORE_STEP, AFTER_STEP, BEFORE_ACTION, AFTER_ACTION, and ON_MESSAGE_SEND events, so instrumentation requires zero changes to model code. Traces export via OTLP to any compatible backend (Jaeger, Zipkin, Grafana Tempo). Metrics export to Prometheus or any OTLP metrics receiver.
Guardrails
The guardrail system provides runtime safety checks that execute after every simulation step (via the AFTER_STEP hook). Four built-in guardrails cover common failure modes:
| Guardrail | __init__ | What It Checks |
|---|---|---|
MaxAgentsGuardrail | (max_agents: int = 100000) | Agent population does not exceed limit (prevents runaway spawning) |
BudgetGuardrail | (max_usd: float = 10.0) | Cumulative LLM spend does not exceed budget |
DeterminismGuardrail | (check_random_module: bool = True) | Seed is configured; no random module usage detected |
StateInvariantGuardrail | (predicate: Callable, message: str) | Custom predicate returns True (e.g., all balances non-negative) |
Guardrails are registered on the model's GuardrailRegistry:
from simudyne.engine.guardrails import (
GuardrailRegistry, MaxAgentsGuardrail, BudgetGuardrail, StateInvariantGuardrail,
)
model.guardrails.add(MaxAgentsGuardrail(max_agents=50000))
model.guardrails.add(BudgetGuardrail(max_usd=5.0))
model.guardrails.add(StateInvariantGuardrail(
predicate=lambda ctx: all(
b.cash_buffer >= 0 for b in ctx.model.get_agents("Bank")
),
message="Negative cash buffer detected"
))
When enforce() is called (automatically by the hook), it runs all guardrails and raises GuardrailViolation with the list of violation messages if any check fails. The check_all() method returns violations as a list of strings without raising, enabling logging-only mode.
Checkpointing
The CheckpointManager provides save/restore with deterministic resume. Checkpoints capture the full model state: agent states, accumulator values, link topology, RNG states (both agent and environment PRNGs), and configuration. The SHA-256 checksum ensures checkpoint integrity.
from simudyne.engine.checkpoint import CheckpointManager
manager = CheckpointManager(
directory="checkpoints/",
max_kept=5, # Keep only the 5 most recent
auto_interval=50, # Auto-save every 50 steps
enabled=True,
)
| Method | Signature | Description |
|---|---|---|
save | (model, step=None, path=None) -> str | Save checkpoint, return file path |
load | (path, model_class=None, validate_checksum=True) -> ABMModel | Restore model from checkpoint |
list_checkpoints | (directory=None) -> List[Dict] | List all checkpoints with metadata |
resume_from_checkpoint | (path, additional_steps=0) -> ABMModel | Load and continue execution |
auto_checkpoint_hook | (ctx) -> None | Hook callback for automatic saving |
from_settings | (settings, model=None) -> CheckpointManager | Create from settings.json |
Auto-checkpointing is wired via the hook system:
from simudyne.engine.hooks import HookEvent
model.hooks.on(HookEvent.AFTER_STEP, manager.auto_checkpoint_hook)
This saves a checkpoint every auto_interval steps, automatically deleting older checkpoints beyond max_kept. Checkpoint metadata (step, timestamp, agent count, accumulator values, config, checksum) is stored as a JSON sidecar file alongside the pickle.