AI Chat API Documentation
Overview
The Agentic Simulation Lab SDK includes an AI chat endpoint powered by Claude via AWS Bedrock. This feature allows users to interactively query and analyze simulation results using natural language.
Architecture
Components
engine/chat.py: Bedrock client wrapper and system prompt builderPOST /api/v1/chat: REST endpoint for chat requests- Rate limiter: Simple in-memory rate limiting (10 req/min per client)
System Prompt Construction
The system prompt includes:
- Model name and basic metadata (agents, steps, runtime)
- Tracked metric names
- Summary statistics per metric (mean, std, min, max)
- Final metric values
- Monte Carlo run count
- Active scenarios
This context grounds the AI's responses in the actual simulation results.
API Reference
Endpoint
POST /api/v1/chat
Request Body
{
"model_name": "ForestFire",
"message": "What was the final burn rate?",
"conversation_id": "chat_1234567890",
"history": [
{
"role": "user",
"content": "How many agents were in this simulation?"
},
{
"role": "assistant",
"content": "This simulation had 10,000 agents..."
}
],
"context": {
"model_name": "ForestFire",
"n_agents": 10000,
"total_steps": 100,
"wall_time_s": 1.234,
"field_names": ["burned_count", "tree_count"],
"time_series_summary": {
"burned_count": {
"mean": 5123.4,
"std": 456.2,
"min": 0.0,
"max": 9500.0
}
},
"final_values": {
"burned_count": 9487,
"tree_count": 513
},
"mc_runs": 1,
"active_scenarios": ["Baseline"]
}
}
Response
{
"conversation_id": "chat_1234567890",
"response": "The final burn rate in this simulation was 94.87%...",
"usage": {
"input_tokens": 234,
"output_tokens": 89,
"total_tokens": 323
}
}
Error Responses
429 - Rate Limit Exceeded
{
"detail": "Rate limit exceeded: 10 requests per 60s"
}
503 - Service Unavailable
{
"detail": "boto3 is not installed. Install with: pip install boto3"
}
{
"detail": "AWS credentials not configured. Set up credentials via AWS CLI or environment variables."
}
500 - Internal Server Error
{
"detail": "Chat service error: <error details>"
}
Python Client Usage
Basic Usage
from engine.chat import BedrockChatClient, build_system_prompt
# Build context from simulation results
context = {
"model_name": "ForestFire",
"n_agents": 10000,
"total_steps": 100,
"wall_time_s": 1.234,
"field_names": ["burned_count", "tree_count"],
"time_series_summary": {...},
"final_values": {...},
"mc_runs": 1,
"active_scenarios": ["Baseline"],
}
# Create system prompt
system_prompt = build_system_prompt(context)
# Initialize client
client = BedrockChatClient()
# Send message
response, usage = client.send_message(
system_prompt=system_prompt,
messages=[
{"role": "user", "content": "What was the final burn rate?"}
],
max_tokens=512,
)
print(response)
print(f"Tokens used: {usage['total_tokens']}")
Conversation History
messages = []
# First message
messages.append({"role": "user", "content": "How many agents were there?"})
response1, usage1 = client.send_message(system_prompt, messages, max_tokens=512)
messages.append({"role": "assistant", "content": response1})
# Follow-up message
messages.append({"role": "user", "content": "What percentage burned?"})
response2, usage2 = client.send_message(system_prompt, messages, max_tokens=512)
messages.append({"role": "assistant", "content": response2})
Custom Model Configuration
# Use different region or model
client = BedrockChatClient(
model_id="us.anthropic.claude-sonnet-4-6",
region="us-west-2"
)
Error Handling
from engine.chat import ChatUnavailableError
try:
response, usage = client.send_message(system_prompt, messages)
except ChatUnavailableError as e:
print(f"Chat unavailable: {e}")
# Handle gracefully (e.g., show error message to user)
Setup Requirements
1. Install boto3
pip install boto3
2. Configure AWS Credentials
Option A: AWS CLI
aws configure
Option B: Environment variables
export AWS_ACCESS_KEY_ID=your_access_key
export AWS_SECRET_ACCESS_KEY=your_secret_key
export AWS_DEFAULT_REGION=us-east-1
Option C: IAM role (for EC2/ECS/Lambda)
- Attach IAM role with Bedrock access to your compute instance
3. Enable Bedrock Access
- Log into AWS Console
- Navigate to Amazon Bedrock
- Request access to Claude models (Anthropic)
- Wait for approval (usually immediate for standard accounts)
Rate Limiting
The chat endpoint implements simple in-memory rate limiting:
- Limit: 10 requests per 60 seconds per client IP
- Scope: Per-instance (not distributed)
- Response: HTTP 429 when limit exceeded
For production deployments with multiple API instances, consider:
- Redis-backed distributed rate limiting
- API Gateway rate limiting
- Per-user (not per-IP) limits with authentication
Cost Considerations
AWS Bedrock charges per token:
- Input tokens: ~$0.003 per 1K tokens
- Output tokens: ~$0.015 per 1K tokens
Typical costs per conversation turn:
- System prompt: ~200-400 input tokens
- User message: ~20-100 input tokens
- Assistant response: ~100-500 output tokens
- Total: ~$0.002-0.010 per turn
Example: 1000 users × 5 turns × $0.005/turn = $25/day
Frontend Integration
The chat endpoint is integrated into the Simudyne Nexus frontend:
- Results view: Chat panel in right sidebar
- Context auto-populated: Pulls from current simulation run
- Conversation history: Persisted in localStorage
- Error handling: Graceful fallback when chat unavailable
See frontend/python-sdk/src/components/Chat/ for implementation details.
Testing
Unit Tests
pytest tests/test_chat.py -v
Manual Testing
# Run example script
PYTHONPATH=. python examples/chat_example.py
Integration Testing
# Start API server
python -m engine.api --host 0.0.0.0 --port 8000
# In another terminal, test endpoint
curl -X POST http://localhost:8000/api/v1/chat \
-H "Content-Type: application/json" \
-d '{
"model_name": "ForestFire",
"message": "What was the final burn rate?",
"context": {...}
}'
Security Considerations
1. API Key Protection
The chat endpoint respects the ABMLAB_API_KEY environment variable. If set, all POST requests require the X-Api-Key header:
export ABMLAB_API_KEY=your_secret_key
# Client request
headers = {"X-Api-Key": "your_secret_key"}
response = requests.post(url, headers=headers, json=data)
2. Rate Limiting
The simple in-memory rate limiter is sufficient for single-instance deployments. For production:
- Use Redis for distributed rate limiting
- Implement per-user limits with authentication
- Add circuit breakers for Bedrock failures
3. Input Validation
All inputs are validated via Pydantic models:
- Message length limits (implicit via max_tokens)
- Context structure validation
- Role validation (user/assistant only)
4. AWS Credentials
Never commit AWS credentials to source control:
- Use IAM roles when possible
- Use environment variables or AWS credentials file
- Rotate keys regularly
- Use least-privilege IAM policies
5. Prompt Injection
The system prompt is constructed server-side and cannot be overridden by clients. User messages are passed to Claude with appropriate role labels, preventing prompt injection attacks.
Troubleshooting
Error: "boto3 is not installed"
pip install boto3
Error: "AWS credentials not configured"
aws configure
# OR
export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
Error: "The provided model identifier is invalid"
- Ensure you're using the correct model ID:
us.anthropic.claude-sonnet-4-6 - Verify Bedrock access is enabled in your AWS account
- Check the region (us-east-1 by default)
Error: "Rate limit exceeded"
- Wait 60 seconds and try again
- Increase rate limit by modifying
_CHAT_RATE_LIMIT_REQUESTSinengine/api.py
Error: "Bedrock API error (ThrottlingException)"
- You've hit Bedrock's service quota
- Request quota increase in AWS Console
- Implement exponential backoff and retry logic
Future Enhancements
Planned Features
- Streaming responses: Real-time token streaming via WebSocket
- Conversation persistence: Store conversations in SQLite
- Multi-model support: Allow model selection (Claude/GPT/Llama)
- Advanced context: Include mechanism summaries, agent traces
- Plot generation: Generate custom plots based on natural language requests
- Comparison queries: Compare results across multiple runs
Performance Optimizations
- Response caching: Cache common queries (e.g., "What is the model about?")
- Context compression: Summarize long time series to reduce token usage
- Async processing: Non-blocking chat requests with webhooks
- Batch processing: Combine multiple questions into single API call