ABM Lab SDK - Production Deployment Guide
This guide covers deploying the ABM Lab Python SDK in production environments using Docker, Docker Compose, and Kubernetes (Helm).
Table of Contents
- Docker Deployment
- Docker Compose Deployment
- Kubernetes Deployment
- Configuration
- Security Best Practices
- Monitoring and Observability
Docker Deployment
Building the Image
# Build with default tag
docker build -t abm-lab:0.7.0 .
# Build with custom tag
docker build -t myregistry.com/abm-lab:0.7.0 .
# Multi-platform build
docker buildx build --platform linux/amd64,linux/arm64 -t abm-lab:0.7.0 .
Running Containers
1. REST API Server
docker run -d \
--name abm-api \
-p 8080:8080 \
-v $(pwd)/models:/app/models:ro \
-v $(pwd)/settings.json:/app/config/settings.json:ro \
-v $(pwd)/output:/app/output \
-e ANTHROPIC_API_KEY=${ANTHROPIC_API_KEY} \
abm-lab:0.7.0 \
abm-lab serve \
--model examples.models.simple_economy \
--name SimpleEconomy \
--port 8080
2. Batch Simulation Runner
docker run -it \
--name abm-runner \
-v $(pwd)/models:/app/models:ro \
-v $(pwd)/settings.json:/app/config/settings.json:ro \
-v $(pwd)/output:/app/output \
abm-lab:0.7.0 \
abm-lab run \
--model examples.models.simple_economy \
--seeds 100 \
--steps 500 \
--settings /app/config/settings.json
3. MCP Partition Server
docker run -d \
--name abm-mcp-0 \
-p 8000:8000 \
-v $(pwd)/settings.json:/app/config/settings.json:ro \
-v $(pwd)/output:/app/output \
abm-lab:0.7.0 \
python -m engine.mcp_server \
--model examples.models.gai_kapadia:GaiKapadiaModel \
--partition 0 \
--num-partitions 3 \
--port 8000
Volume Mounts
| Host Path | Container Path | Purpose | Mode |
|---|---|---|---|
./models | /app/models | Custom model files | ro |
./settings.json | /app/config/settings.json | Runtime configuration | ro |
./output | /app/output | Simulation results | rw |
Environment Variables
| Variable | Required | Description | Default |
|---|---|---|---|
ANTHROPIC_API_KEY | No* | Anthropic API key for LLM agents | - |
ABM_LAB_SETTINGS | No | Path to settings.json | /app/config/settings.json |
PYTHONUNBUFFERED | No | Enable unbuffered Python output | 1 |
*Required only if using LLMAgent or HybridAgent subclasses.
Docker Compose Deployment
Quick Start
# Start API server
docker compose up abm-api
# Start API + DuckDB
docker compose up abm-api duckdb
# Start batch runner
docker compose up abm-runner
# Start 3-partition MCP deployment
docker compose up abm-mcp-partition-0 abm-mcp-partition-1 abm-mcp-partition-2
Service Overview
| Service | Purpose | Port | Mode |
|---|---|---|---|
abm-api | REST API server | 8080 | API |
abm-runner | Batch simulation runner | - | CLI |
duckdb | Persistence backend | - | Storage |
abm-mcp-partition-* | Distributed MCP partitions | 8000-8002 | MCP |
Customizing Services
Create .env file:
MODEL_PATH=library.models.my_model
MODEL_NAME=MyModel
SEEDS=50
STEPS=200
ANTHROPIC_API_KEY=your-api-key-here
Start services:
docker compose --env-file .env up
Scaling Services
# Scale API to 3 replicas
docker compose up --scale abm-api=3
# Scale with load balancer (requires additional nginx/traefik config)
docker compose -f docker-compose.yml -f docker-compose.lb.yml up
Kubernetes Deployment
Prerequisites
# Install Helm 3
curl https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash
# Verify Kubernetes cluster
kubectl cluster-info
kubectl get nodes
Installing the Helm Chart
1. Basic Installation (API Mode)
helm install abm-lab ./deploy/helm/abm-lab
2. Custom Values File
Create production-values.yaml:
deploymentMode: api
replicaCount: 3
image:
repository: myregistry.com/abm-lab
tag: 0.7.0
model:
path: examples.models.financial_contagion
name: FinancialContagion
service:
type: LoadBalancer
resources:
limits:
cpu: "4"
memory: "8Gi"
requests:
cpu: "1"
memory: "2Gi"
autoscaling:
enabled: true
minReplicas: 2
maxReplicas: 20
targetCPUUtilizationPercentage: 70
persistence:
enabled: true
size: 50Gi
storageClass: fast-ssd
settings:
orchestration:
mode: threaded
max_workers: 8
llm:
default_model: claude-sonnet-4-6
cost_budget_total: 100.0
logging:
level: INFO
log_llm_calls: true
Install with custom values:
helm install abm-lab ./deploy/helm/abm-lab -f production-values.yaml
3. Secrets Management
Create Kubernetes secret for Anthropic API key:
kubectl create secret generic abm-lab-secrets \
--from-literal=anthropic-api-key=${ANTHROPIC_API_KEY}
The Helm chart automatically mounts this secret.
4. Ingress Configuration
ingress:
enabled: true
className: nginx
annotations:
cert-manager.io/cluster-issuer: letsencrypt-prod
nginx.ingress.kubernetes.io/proxy-body-size: 100m
hosts:
- host: abm.example.com
paths:
- path: /
pathType: Prefix
tls:
- secretName: abm-lab-tls
hosts:
- abm.example.com
Deployment Modes
API Mode (Default)
Deploys REST API server for interactive model serving:
helm install abm-api ./deploy/helm/abm-lab \
--set deploymentMode=api \
--set service.type=LoadBalancer
Runner Mode
Deploys batch simulation runner:
helm install abm-runner ./deploy/helm/abm-lab \
--set deploymentMode=runner \
--set model.seeds=1000 \
--set model.steps=500
MCP Mode
Deploys distributed MCP partition (requires StatefulSet for multi-partition):
helm install abm-mcp ./deploy/helm/abm-lab \
--set deploymentMode=mcp \
--set mcp.numPartitions=5
Upgrading Deployments
# Upgrade with new image version
helm upgrade abm-lab ./deploy/helm/abm-lab \
--set image.tag=0.7.1
# Upgrade with new values
helm upgrade abm-lab ./deploy/helm/abm-lab -f updated-values.yaml
# Rollback to previous release
helm rollback abm-lab
Uninstalling
helm uninstall abm-lab
# Also delete PVCs (if persistence enabled)
kubectl delete pvc -l app.kubernetes.io/name=abm-lab
Configuration
settings.json Structure
The SDK uses a settings.json file for runtime configuration:
{
"orchestration": {
"mode": "local",
"max_workers": 4,
"mcp_config": null
},
"llm": {
"default_model": "claude-sonnet-4-6",
"default_max_tokens": 200,
"default_temperature": 0.0,
"cost_budget_total": 10.0,
"api_key_env": "ANTHROPIC_API_KEY"
},
"logging": {
"level": "INFO",
"log_llm_calls": true,
"log_messages": false
},
"mc": {
"default_seeds": 100,
"default_steps": 100,
"max_workers": 4
}
}
Orchestration Modes
| Mode | Description | Use Case |
|---|---|---|
local | Single-threaded execution | Development, debugging |
threaded | Thread pool execution | Single-node production |
mcp | Distributed MCP servers | Multi-node scaling |
Environment Variable Overrides
Settings can be overridden via environment variables:
# Override logging level
-e ABM_LAB_LOG_LEVEL=DEBUG
# Override LLM model
-e ABM_LAB_LLM_MODEL=claude-opus-4-6
# Override max workers
-e ABM_LAB_MAX_WORKERS=16
Security Best Practices
1. Container Security
- Non-root user: All containers run as UID 1000 (
abmuser) - Read-only root filesystem: Consider enabling with
--read-onlyflag - Minimal base image: Uses
python:3.14-slim(minimal attack surface) - No build tools in runtime: Multi-stage build removes gcc/g++
2. Secrets Management
DO NOT commit secrets to version control or bake into images.
Docker Secrets
# Create Docker secret
echo ${ANTHROPIC_API_KEY} | docker secret create anthropic_api_key -
# Use in Docker service
docker service create \
--name abm-api \
--secret anthropic_api_key \
abm-lab:0.7.0
Kubernetes Secrets
# Create from file
kubectl create secret generic abm-lab-secrets \
--from-file=anthropic-api-key=./api-key.txt
# Create from literal
kubectl create secret generic abm-lab-secrets \
--from-literal=anthropic-api-key=${ANTHROPIC_API_KEY}
AWS Secrets Manager / Azure Key Vault
Use CSI driver for external secrets:
apiVersion: secrets-store.csi.x-k8s.io/v1
kind: SecretProviderClass
metadata:
name: abm-lab-secrets
spec:
provider: aws
parameters:
objects: |
- objectName: "abm-lab/anthropic-api-key"
objectType: "secretsmanager"
3. Network Security
Docker Networks
# Create isolated network
docker network create --driver bridge abm-private
# Run containers on private network
docker run --network abm-private abm-lab:0.5.0
Kubernetes NetworkPolicies
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: abm-lab-netpol
spec:
podSelector:
matchLabels:
app.kubernetes.io/name: abm-lab
policyTypes:
- Ingress
- Egress
ingress:
- from:
- podSelector:
matchLabels:
app: ingress-nginx
ports:
- protocol: TCP
port: 8080
egress:
- to:
- namespaceSelector: {}
ports:
- protocol: TCP
port: 443 # HTTPS (Anthropic API)
4. Resource Limits
Always set resource limits to prevent resource exhaustion:
resources:
limits:
cpu: "2"
memory: "4Gi"
requests:
cpu: "500m"
memory: "1Gi"
5. Image Scanning
Scan images for vulnerabilities before deployment:
# Trivy
trivy image abm-lab:0.7.0
# Grype
grype abm-lab:0.7.0
# Docker Scout
docker scout cves abm-lab:0.7.0
Monitoring and Observability
Health Checks
Docker
# Check health status
docker ps --filter "name=abm-lab"
# View health check logs
docker inspect --format='{{json .State.Health}}' abm-lab | jq
Kubernetes
# Check pod health
kubectl get pods -l app.kubernetes.io/name=abm-lab
# Describe pod for probe details
kubectl describe pod <pod-name>
Logging
Docker Logs
# Tail logs
docker logs -f abm-lab
# Export logs to file
docker logs abm-lab > abm-lab.log
Kubernetes Logs
# Tail logs
kubectl logs -f deployment/abm-lab
# Previous container logs (after crash)
kubectl logs deployment/abm-lab --previous
# All pods in deployment
kubectl logs -l app.kubernetes.io/name=abm-lab --all-containers
Centralized Logging (ELK/EFK)
Configure log shipping to Elasticsearch/Loki:
# Fluent Bit sidecar
- name: fluent-bit
image: fluent/fluent-bit:latest
volumeMounts:
- name: logs
mountPath: /app/output/logs
Metrics
Prometheus Metrics
Add Prometheus annotations to Helm values:
podAnnotations:
prometheus.io/scrape: "true"
prometheus.io/port: "8080"
prometheus.io/path: "/metrics"
Custom Metrics
The SDK can expose custom simulation metrics via /metrics endpoint (requires implementation).
Tracing
For distributed tracing (MCP mode), integrate OpenTelemetry:
from opentelemetry import trace
from opentelemetry.exporter.jaeger import JaegerExporter
tracer = trace.get_tracer(__name__)
Production Checklist
- Build and scan Docker image for vulnerabilities
- Store secrets in secret manager (not in config files)
- Set resource limits on all containers/pods
- Enable persistent storage for output data
- Configure log aggregation (ELK, Loki, CloudWatch)
- Set up health check monitoring
- Enable horizontal pod autoscaling (HPA)
- Configure ingress with TLS certificates
- Set up network policies for isolation
- Test rollback procedures
- Document disaster recovery procedures
- Configure backup for persistent volumes
- Set up alerts for failures and resource exhaustion
Troubleshooting
Issue: Container exits immediately
Check logs:
docker logs abm-lab
Common causes:
- Missing required environment variables
- Invalid model path
- Settings.json syntax error
Issue: Health check failing
Check endpoint:
curl http://localhost:8080/health
Verify settings:
- Ensure port 8080 is exposed
- Check if service is listening on 0.0.0.0 (not 127.0.0.1)
Issue: Out of memory
Increase limits:
resources:
limits:
memory: "8Gi"
Optimize model:
- Reduce agent count
- Lower max_workers
- Enable agent partitioning
Issue: LLM API errors
Check API key:
kubectl get secret abm-lab-secrets -o jsonpath='{.data.anthropic-api-key}' | base64 -d
Check cost budgets:
- Increase
cost_budget_totalin settings - Monitor LLM usage with
log_llm_calls: true
Support
For issues, questions, or feature requests:
- Documentation:
/docs - GitHub Issues: https://github.com/simudyne/agentic-sdk/issues
- Email: support@simudyne.com
License: Proprietary - Simudyne Ltd. All rights reserved.