Section 1: Introduction & Design Philosophy
1.1 Overview
The Agentic Simulation Lab Python SDK (abm-lab) is a production-grade, pip-installable framework for building, evaluating, calibrating, and deploying agent-based models (ABMs). It is developed and maintained by Simudyne Ltd. as the foundational Layer 1 of the Agentic Simulation Lab platform.
The SDK is written in Python 3.11+ and ships as a single package with minimal core dependencies (NumPy, SciPy, jsonschema); everything else --- FastAPI, Plotly, OpenTelemetry, Anthropic SDK, Dask, PostgreSQL drivers --- is opt-in via extras.
Most ABM toolkits available today focus on rapid prototyping: they provide a canvas for simple models with small agent populations, but offer limited support for Monte Carlo evaluation, no statistical validation framework, no calibration tooling, and no path to production deployment. They treat model building as the end goal rather than the first step in a rigorous scientific workflow. This SDK takes a fundamentally different approach. It is designed for production deployment of models that must be deterministic, auditable, scalable, and integrated into enterprise infrastructure. It provides the full lifecycle: model construction, Monte Carlo evaluation, statistical validation, Bayesian calibration, mechanism ablation, Docker/Helm deployment, and REST API serving --- all from a single pip install.
The SDK also natively supports AI-powered agents: LLM agents that reason in natural language, hybrid agents that blend LLM reasoning with rule-based fallbacks, reinforcement learning agents with pluggable policies, and external agent proxies for cross-process integration. This makes it the first ABM framework to treat AI-driven decision-making as a first-class citizen alongside traditional rule-based agents, with full deterministic reproducibility via response caching and SHA-256 seeding.
The SDK serves four primary audiences, each with different entry points into the framework. Quantitative researchers use the core primitives (Agent, Message, Link, Action, Sequence) to define agent behaviours, the Monte Carlo runner for stochastic evaluation, the Feature protocol for validation against stylized facts, and the ABC-SMC calibrator for likelihood-free parameter estimation. Simulation engineers configure execution backends via settings.json, set up multi-level output recording, instrument models with OpenTelemetry, and deploy via Docker + Helm. Risk modellers use the 27 standard metrics, the statistical testing suite (KS, permutation, bootstrap, Wasserstein/MMD/JSD distances), and the experiment runner for parameter sweeps (grid, Latin hypercube, Sobol). AI researchers use LLMAgent, HybridAgent, ResponseCache, and NestedMCRunner for variance decomposition that separates seed variance from LLM variance.
| Capability | Details |
|---|---|
| Agent types | 5: Agent (rule-based), LLMAgent (LLM-powered), HybridAgent (LLM + rules), RLAgent (reinforcement learning), ExternalAgentProxy (cross-process) |
| Execution backends | 8: LocalBackend, ThreadPoolBackend, ProcessPoolBackend, MCPBackend, EmulatorBackend, AsyncBackend, PregelBackend, DaskBackend |
| Standard metrics | 27: 21 core (mean, variance, std, autocorrelation, kurtosis, skewness, Hurst exponent, max drawdown, CV, Hill tail index, Ljung-Box, cross-correlation, realized volatility, VaR, CVaR, tail ratio, Gini, entropy, Sharpe, Sortino, sample entropy) + 6 generic aliases (peak_to_trough, rolling_std, lower_quantile, lower_tail_mean, risk_adjusted_mean, downside_adjusted_mean) |
| Calibration | ABC-SMC with 5 prior types, compare_models() for complexity-penalised selection, posterior predictive checks |
| Validation | Generic Feature protocol, FeatureSet with compute/validate/compare/distance, 3 domain libraries: financial (8), supply chain (4), epidemiology (4) |
| Statistical testing | KS, Anderson-Darling, permutation, Mann-Whitney U, bootstrap CI; Wasserstein, energy, MMD, JSD distances; Bonferroni, Holm, BH-FDR, BY-FDR corrections |
| Output | Multi-level SimulationRecorder (model/agent_group/agent/environment), CSV/CSV.GZ/Parquet/JSON, PNG plots, PostgreSQL/ClickHouse sinks, S3 |
| Deployment | Docker, Helm (8 K8s templates), FastAPI REST API (8 endpoints + WebSocket), OpenTelemetry (5 spans, 7 metrics), checkpointing |
| Reference models | 23 models across 10 domains, each with README (tags + architecture diagram), PDF report, results.json |
1.2 Getting Started
Installation
The SDK is installed from source as a local editable package:
git clone https://github.com/simudyne/agentic-sdk.git
cd agentic-sdk
python -m venv .venv
source .venv/bin/activate
# Core only --- numpy, scipy, jsonschema (no server, no LLM, no viz)
pip install -e .
# With specific extras
pip install -e ".[server]" # FastAPI + Uvicorn for REST API
pip install -e ".[llm]" # Anthropic SDK for LLM agents
pip install -e ".[notebook]" # Jupyter + Matplotlib
pip install -e ".[viz-interactive]" # Plotly interactive charts
pip install -e ".[otel]" # OpenTelemetry instrumentation
pip install -e ".[sql]" # PostgreSQL + ClickHouse drivers
pip install -e ".[cloud]" # boto3 for S3 output
pip install -e ".[parquet]" # PyArrow for Parquet output
pip install -e ".[duckdb]" # DuckDB for local analytics
pip install -e ".[streaming]" # WebSocket streaming
# Everything at once
pip install -e ".[all]"
Requirements: Python >= 3.11. The core package depends only on numpy >= 1.24, scipy >= 1.10, and jsonschema >= 4.17. All other dependencies are opt-in via the extras above.
Your First Model
The following is a complete, runnable model that demonstrates every core concept: agent state, accumulators, the init/setup/step lifecycle, action binding, sequence execution, and accumulator recording.
from simudyne.engine.sdk import ABMModel, Agent, Action, Sequence
class Greeter(Agent):
"""An agent that increments a counter each step."""
__state_schema__ = {"count": int}
def greet(self):
self.count += 1
self.get_accumulator("total").add(1)
class HelloModel(ABMModel):
def init(self):
# Register agent type and create a model-level accumulator
self.register_agent_type("Greeter", Greeter)
self.create_accumulator("total")
def setup(self, config):
# MUST call super().setup() first --- initialises SeedManager
super().setup(config)
# Create N agents (default 10)
self.create_agents("Greeter", config.get("n", 10))
def step(self):
# Reset accumulators to zero at the start of each step
self.reset_accumulators()
# Execute the greet action on all Greeter agents
self.run(Sequence(Action.create(Greeter, Greeter.greet)))
# Record accumulator values for output
self.record_accumulators()
# Run the simulation: 100 steps with seed 42
model = HelloModel.run_simulation({"seed": 42, "steps": 100, "n": 10})
print(f"Total greetings: {model.get_accumulator('total').get()}")
# Output: Total greetings: 10 (10 agents x 1 greet per step, accumulated in last step)
What happens under the hood:
HelloModel()is instantiated.init()registers theGreeteragent type and creates thetotalaccumulator.setup(config)is called.super().setup(config)initialises theSeedManagerfromconfig["seed"], sets the model tick to 0, and prepares the execution backend (default:LocalBackend). Thencreate_agents()instantiates 10Greeterobjects, each with a deterministic PRNG derived fromSHA-256(agent_id) XOR scenario_seed.- For each of the 100 steps,
step()is called.reset_accumulators()setstotalto 0.run(Sequence(...))iterates over allGreeteragents and callsgreet()on each. Each call increments the agent'scountand adds 1 to thetotalaccumulator (thread-safe viathreading.Lock).record_accumulators()writes the current accumulator values to the output recorder. - After 100 steps, the model is returned. The final
totalaccumulator value is 10 (reset each step, so only the last step's value persists).
Command-Line Interface
The SDK provides a CLI entry point abm-lab with four subcommands:
# Run a model for multiple seeds
abm-lab run --model examples/models/forest_fire/model.py \
--seeds 10 --steps 200 --settings settings.json
# Validate a model file (3-stage validation: syntax, structure, runtime)
abm-lab validate --model examples/models/gai_kapadia/model.py
# Monte Carlo evaluation with a spec file
abm-lab evaluate --model examples/models/cda_market/model.py \
--mc-spec mc_spec.json --workers 8 --output results.json
# Compare two MC evaluation results
abm-lab compare --baseline results_v1.json --challenger results_v2.json
The --model flag accepts either a file path (examples/models/forest_fire/model.py) or a dotted module path (library.models.gai_kapadia). The CLI auto-discovers the first ABMModel subclass in the module. The --settings flag points to a settings.json file that controls the execution backend, output format, LLM configuration, Monte Carlo defaults, checkpointing, and telemetry. If omitted, the SDK uses sensible defaults (local backend, CSV output, no LLM, no telemetry).
The REST API server is started with abm-lab serve --model model.py --port 8080, which launches a FastAPI server with Swagger UI at /docs, 8 REST endpoints, and a WebSocket endpoint for live step-by-step streaming.
1.3 Design Philosophy
[!bluequote] "Models should describe the system, not just fit it."
This single sentence governs every design decision in the SDK. An ABM that reproduces historical price volatility by tuning a noise parameter is a curve fit --- it tells you nothing about why volatility behaves that way. An ABM that produces realistic volatility because its fundamentalist agents provide negative feedback (mean-reversion) while its chartist agents provide positive feedback (trend-following), and the interaction between these two mechanisms generates clustered volatility as an emergent property --- that model describes the system. It gives you a causal explanation you can test, modify, and reason about.
The SDK enforces this philosophy structurally, not just as an aspiration. Every mechanism must be declared with a quantitative hypothesis (MechanismDeclaration). Every mechanism can be disabled and the effect measured (AblationRunner). Every model can be validated against domain-specific stylized facts (FeatureSet.validate()). The framework makes it structurally easier to describe systems than to fit data.
Five Design Principles
1. Mechanism hypothesis required. No model component should exist without a declared causal hypothesis about why it should improve realism. The SDK enforces this through MechanismDeclaration, which requires a name, a hypothesis (string explaining the causal mechanism), a toggle_fn (callable that disables the mechanism), and a predicted_effect (dictionary mapping feature names to expected direction and magnitude of change). A mechanism without a hypothesis is decorative complexity.
2. Ablation testing. Every mechanism must survive ablation: disable it, re-run Monte Carlo evaluation, and measure the effect on features. If removing a mechanism does not statistically change the output (effect size below the decorative_threshold, default 10%), it is decorative --- remove it. The AblationRunner automates this with single-mechanism ablation and pairwise ablation (testing mechanism interactions). This prevents the common failure mode of accumulating mechanisms that individually seem reasonable but collectively overfit.
3. Complexity must be earned. Adding complexity to a model requires evidence that simpler alternatives cannot explain the gap between model output and observed data. The five-stage complexity model (below) gates when each SDK capability is permitted. You cannot use LLM agents (Stage D) until you have demonstrated that rule-based agents (Stages A-B) and RL agents (Stage C) are insufficient. The compare_models() function applies complexity penalties to enforce this: a simpler model that reproduces 3 stylized facts beats a complex model that reproduces 5 but uses more mechanisms.
4. Glass-box explainability. Every mechanism must be traceable from cause to effect. The generic Feature protocol provides the measurement layer: a Feature is a named, described callable that extracts a quantitative signal from model output. Features are composed into FeatureSet collections that can compute, validate, compare, and measure distance. This makes model evaluation explicit and quantitative rather than relying on visual inspection of time series plots.
5. Description over fitting. A model that reproduces 3 stylized facts through mechanistically correct agent interactions beats one that reproduces 5 by tuning parameters. The calibration framework (ABCSMCCalibrator) uses the FeatureSet.distance() method as its loss function, which measures how far simulated features are from observed features in a domain-appropriate metric space. The compare_models() function adds a complexity penalty proportional to the number of free parameters, ensuring that parsimony is rewarded.
Five-Stage Complexity Model
The SDK gates capabilities behind a five-stage complexity model. Each stage unlocks specific SDK features; you may only use capabilities from your current stage or below. This prevents premature complexity and ensures that every model starts simple and adds complexity only when evidence justifies it.
| Stage | Name | Agent Types | Backends | Parameters | Evaluation |
|---|---|---|---|---|---|
| A | Simple | Agent only | LocalBackend | Constant only | MCRunner (small N), DeterminismChecker |
| B | Intermediate | Agent (heterogeneous) | ThreadPoolBackend, ProcessPoolBackend | Input + Variable, ExperimentRunner | MCRunner (larger N), sensitivity sweeps, Guardrails |
| C | Advanced | Agent + RLAgent | AsyncBackend | Full param system, sweep grids | MCRunner + ExperimentRunner, SQL sinks, Checkpoint |
| D | Hybrid | All 5 agent types | DaskBackend, PregelBackend | Full + differentiable | Full analytical layer: features, calibration, ablation, variance decomp |
| E | Production | All + ExternalAgentProxy | Production cluster | Full | REST API, Docker, Helm, OTel, audit, mandatory checkpointing |
The key insight is that complexity must be earned. You cannot use LLM agents (Stage D) until you have demonstrated that rule-based agents (Stages A-B) cannot reproduce the target behaviour. You cannot deploy to production (Stage E) until the model passes OpenTelemetry trace validation, guardrail checks, and determinism verification. Each stage gate requires evidence that simpler capabilities are insufficient.
1.4 Key Innovations & Patent Integration
Three capabilities distinguish this SDK from every other ABM framework in existence, underpinned by two Simudyne patents that solve fundamental problems in distributed deterministic simulation.
Deterministic LLM Agents
LLM-based agents are inherently stochastic --- even at temperature 0, API responses can vary across calls due to server-side batching and numerical precision. The SDK solves this with ResponseCache: a content-addressed cache (keyed by SHA-256 hash of the full prompt) that stores LLM responses in an SQLite database. On first run, LLM calls hit the API and populate the cache. On replay (same seed, same config), all prompts are identical (because agent state evolves deterministically from the SHA-256 seed tree), so the cache achieves a 100% hit rate. This makes LLM-ABM simulations fully reproducible --- a property no other framework offers.
The MockLLMBackend takes this further for testing: it returns deterministic responses without any API calls, enabling thousands of test runs at zero cost. The DebugLLMBackend logs all prompts and responses for inspection without modifying behaviour.
Bifurcated Execution
In a model with 100 rule-based agents and 10 LLM agents, the rule-based agents execute in microseconds while the LLM agents take hundreds of milliseconds per API call. The AsyncBackend handles this by partitioning agents by their BehaviorType: rule-based agents execute synchronously (zero asyncio overhead), while LLM agents execute concurrently via asyncio.gather() with a configurable semaphore limit. This means a model with 100 rule-based agents and 10 LLM agents runs the rule agents in ~100 microseconds and the LLM agents in ~200ms (1 round-trip, not 10), for a total step time of ~200ms instead of ~2 seconds. The semaphore prevents overwhelming the LLM API endpoint. For models with zero LLM agents, AsyncBackend falls back to synchronous execution with zero overhead.
Variance Decomposition for LLM-ABMs
Monte Carlo evaluation of a rule-based ABM has one source of variance: the random seed. LLM-ABMs have two: seed variance (structural model randomness) and LLM variance (AI decision stochasticity). If LLM variance dominates, the model is unstable --- small prompt changes cause large output changes, which means the model is fitting noise rather than structure.
NestedMCRunner runs a nested experimental design: n_seeds outer seeds x n_inner_reps inner replications. For each outer seed, the model runs n_inner_reps times with the same seed but different LLM cache states (cold cache). This produces a nested dataset that NestedANOVA decomposes into between-seed variance (model structure) and within-seed variance (labelled "within"). The intra-class correlation coefficient (ICC) quantifies the ratio: ICC > 0.7 means the model is healthy (seed variance dominates); ICC < 0.3 means the model is within-noise-dominated and needs mechanism redesign. The diagnosis() method returns a human-readable interpretation. (The parameter was renamed from n_llm_reps to n_inner_reps in v0.7.1; the old name is preserved as a backward-compatible alias.)
Patent 1: Topology-Aware Agent Partitioning (US20200278838)
When distributing agents across multiple processes or machines for parallel execution, the naive approach (round-robin or random assignment) can place heavily-communicating agents on different partitions, causing expensive cross-partition message passing. Patent 1 provides topology-aware partitioning: a greedy BFS algorithm that traverses the agent link graph and assigns connected clusters to the same partition, minimising the number of edges that cross partition boundaries (the "cut ratio").
The SDK implements three partitioning strategies:
TypeBasedPartition: Assigns all agents of the same type to the same partition. O(N) time. Effective when inter-type messaging dominates (e.g., all Banks talk to all Traders).TopologyAwarePartition: Greedy BFS traversal of the link graph, assigning connected components to partitions while maintaining load balance. O(N+E) time. Effective for dense graphs with cross-type links (e.g., interbank networks).ManualPartition: User-specified mapping from agent IDs to partition indices. O(N) time. Effective when the domain expert knows the optimal partitioning (e.g., geographic regions).
The PartitionMetrics dataclass quantifies partition quality: cut_edges (number of cross-partition links), cut_ratio (fraction of total edges that cross), balance (ratio of smallest to largest partition), max_load and min_load (agent counts per partition).
Patent 2: SHA-256 Hierarchical Seeding (US20200278838)
Deterministic simulation requires that the same seed produces the same output, regardless of whether the model runs on one process or ten, on one machine or a cluster. Python's built-in hash() function is not deterministic across processes (due to hash randomisation since Python 3.3), so it cannot be used for PRNG seeding.
Patent 2 provides a SHA-256-based seed hierarchy:
global_seed
+-- SHA-256(global_seed || "model") --> model_seed
+-- SHA-256(model_seed || scenario_index) --> scenario_seed
+-- SHA-256(agent_id) XOR scenario_seed --> agent_rng
+-- SHA-256("environment") XOR scenario_seed --> environment_rng
Each agent's PRNG is a numpy.random.Generator seeded with the XOR of SHA-256(agent_id) and the scenario seed. This gives three critical properties:
- Agent independence: Adding or removing agent N does not change agent M's random number stream. This is because each agent's seed depends only on its own ID, not on the total agent count or creation order.
- Environment isolation: The environment PRNG is derived from a fixed label (
"environment"), not from any agent ID, so it is independent of all agent PRNGs. - MC orthogonality: Each Monte Carlo scenario seed is derived from the model seed and the scenario index via SHA-256 mixing, producing statistically independent PRNG streams across scenarios.
The SeedManager class encapsulates this hierarchy. It is initialised in ABMModel.setup() (which is why super().setup(config) must always be called first), and provides mc_seeds(), agent_rng(), and environment_rng() methods.
1.5 Roadmap
The SDK is under active development. The following capabilities are planned for upcoming releases.
Differentiable ABMs
The current calibration framework (ABCSMCCalibrator) is likelihood-free: it runs full forward simulations and compares output features to observations, requiring hundreds or thousands of simulation runs to converge. This is effective but computationally expensive. The next major capability is a JAX-based differentiable execution backend that enables gradient-based calibration. By expressing agent update rules as differentiable operations, the model's loss function (feature distance from observed data) can be differentiated through the entire simulation, enabling calibration via gradient descent in minutes rather than hours.
The groundwork is already in place: the Input descriptor supports a differentiable=True flag that marks parameters as candidates for gradient-based optimisation. The differentiable backend will implement the same ExecutionBackend interface as all other backends, meaning existing models can switch to gradient-based calibration with a single configuration change --- no model code modifications required. This will be particularly impactful for models with high-dimensional parameter spaces (10+ free parameters) where ABC-SMC becomes sample-inefficient.
AutoCalibration
Building on both the ABC-SMC framework and the planned differentiable backend, the SDK will provide an AutoCalibration pipeline that automates the full calibration workflow: prior specification from parameter metadata, automatic feature selection from domain libraries, adaptive algorithm selection (ABC-SMC for non-differentiable models, gradient descent for differentiable ones), convergence diagnostics, posterior predictive validation, and report generation. The goal is a single function call --- auto_calibrate(model_cls, observed_data) --- that returns calibrated parameters with uncertainty estimates and a validation report, without requiring the user to manually configure priors, particle counts, or generation schedules.
AutoCalibration will also integrate with the ablation framework: after calibration, the pipeline will automatically run ablation tests on all declared mechanisms and flag any that are decorative (statistically insignificant effect when removed). This closes the loop between calibration and mechanism validation, ensuring that calibrated models are not over-parameterised.
External System Integration (Patent 3)
The SDK currently supports cross-process agent execution via ExternalAgentProxy and MCPBackend, but these mechanisms are designed for agents that live within the simulation boundary. The next step is full bidirectional integration with external systems --- trading platforms, risk engines, market data feeds, regulatory reporting systems --- where the ABM acts as a digital twin that both consumes real-time data and emits decisions back to the source system.
Patent 3 (in development) extends the MCP protocol with a standardised observation/action interface: external systems push observations (market data, order fills, position updates) into the simulation environment, and the simulation pushes agent actions (order submissions, hedging decisions, risk alerts) back to the external system. This transforms the ABM from an offline analytical tool into a live decision-support system that runs alongside production infrastructure, continuously calibrating against incoming data and surfacing emergent risks before they materialise.
The integration layer will support multiple transport protocols (HTTP/REST, WebSocket, gRPC, Kafka) and provide adapters for common financial systems (FIX protocol for order routing, ITCH/OUCH for market data). The ExternalAgentProxy will evolve into a general-purpose bridge that can represent any external decision-maker --- human traders, algorithmic systems, or other simulations --- as a first-class agent within the model.
Additional Planned Capabilities
- Out-of-sample validation protocol: Walk-forward validation that splits time series data into calibration and validation windows, preventing in-sample overfitting and providing honest estimates of predictive performance.
- Overfitting detector: Automatic warning when the ratio of free parameters to validated features exceeds a threshold, flagging models that are likely fitting noise.
- Stronger RL integration: Native support for stable-baselines3 policies, enabling agents to learn optimal strategies through reward-driven training within the simulation environment.
- GPU-accelerated execution: CUDA backend for models with large agent populations (>100K) where vectorised agent updates can be parallelised across GPU cores.
- Federated simulation: Multi-organisation simulation where each participant runs a private partition of the model and shares only aggregate statistics, enabling collaborative modelling of systemic risk without exposing proprietary positions or strategies.