30 KiB
APEX OS Engineering Standards
The definitive engineering standards governing all infrastructure, code, databases, APIs, testing, monitoring, and operations within APEX OS. All employees producing technical deliverables must adhere to these standards.
Version: 2.0 Last Updated: Phase 7 — CEO Command Center Classification: CORE — All technical work must comply Maintainer: Engineer (#1) Cross-references: APEX_CONSTITUTION.md · TOOL_REGISTRY.md · EMPLOYEE_HANDBOOK.md
Table of Contents
- 1. Architecture Principles
- 2. Docker Standards
- 3. Repository Standards
- 4. Git Workflow
- 5. Database Standards
- 6. API Standards
- 7. Testing Requirements
- 8. Logging Standards
- 9. Monitoring Standards
- 10. Error Handling
- 11. Versioning
- 12. Documentation Standards
- 13. Rollback Procedures
- 14. Infrastructure Change Process
- 15. Network Architecture
- 16. Performance Standards
- 17. Security Engineering
- 18. Change History
1. Architecture Principles
1.1 Microservices via Docker
Every service in APEX OS runs as an isolated Docker container. This provides:
- Reproducibility — Identical environments from dev to production
- Isolation — A failing service doesn't take down others
- Scalability — Services can be independently scaled
- Clean rollback — Previous image tags can be restored instantly
1.2 Single Responsibility
Each container serves one purpose. A database is a database. A proxy is a proxy. An API is an API. No monolithic containers combining multiple services.
1.3 Internal Communication
Services communicate via the Docker network (apex_apex-net). Internal service discovery uses Docker DNS (container names resolve to IPs within the network).
Service A ──(apex_apex-net)──► Service B
│
└── DNS: apex-postgres:5432
└── DNS: apex-redis:6379
└── DNS: apex-litellm:4000
1.4 External Access
All external traffic is routed through Traefik reverse proxy:
- HTTP (80) → Redirected to HTTPS (443)
- HTTPS (443) → Routed to containers via Docker labels
- Let's Encrypt certificates auto-renewed
- No service directly exposes ports to the internet (except Traefik)
1.5 State Management
- Persistent state stored in PostgreSQL (relational data, vectors, agent state)
- Cache/ephemeral state stored in Redis (LLM response cache, sessions)
- File state stored in Docker volumes (Gitea repos, Vaultwarden data, n8n workflows)
- No state in containers — containers are ephemeral and replaceable
1.6 Cost-Driven Design
Architecture decisions optimize for cost first, then speed, then quality (Constitution §5, Principle 3):
- Local models for routine tasks (Ollama → $0)
- Cloud models only when quality demands it (OpenRouter → metered)
- Caching for repeated queries (Redis)
- Batch operations over individual API calls
2. Docker Standards
2.1 docker-compose.yml Requirements
Every service in docker-compose.yml must include:
services:
apex-example:
image: example/image:version # Pinned version tag
container_name: apex-example # apex- prefix required
restart: unless-stopped # Auto-restart on failure
healthcheck: # Health check required
test: ["CMD", "curl", "-f", "http://localhost:PORT/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 40s
networks:
- apex-net # Internal network
labels:
- "traefik.enable=true" # If externally accessible
- "traefik.http.routers.example.rule=Host(`example.apex.unstuck-path.com`)"
- "traefik.http.routers.example.tls.certresolver=letsencrypt"
volumes:
- example_data:/data # Named volumes for persistence
environment:
- CONFIG_VAR=${CONFIG_VAR} # From .env file, never hardcoded
2.2 Container Naming
- Prefix: All containers start with
apex- - Format:
apex-{service-name}(lowercase, hyphenated) - Examples:
apex-postgres,apex-traefik,apex-litellm,apex-status-api
2.3 Health Checks
Required on every container. Health check types:
| Type | When to Use | Example |
|---|---|---|
| HTTP | Service has a web endpoint | curl -f http://localhost:PORT/health |
| TCP | Service listens on a port | pg_isready for PostgreSQL |
| CMD | Custom health logic | redis-cli ping |
2.4 Restart Policy
- Default:
restart: unless-stopped - This ensures containers restart after crashes but respect manual stops
- Never use
restart: always(prevents intentional stops)
2.5 Resource Awareness
- Set memory limits for resource-intensive services:
deploy: resources: limits: memory: 512M - Monitor resource usage via Prometheus/Grafana
- Document expected resource usage in TOOL_REGISTRY.md
2.6 Volumes
- Use named volumes for persistent data:
apex_postgres_data,apex_gitea_data - Use bind mounts only when necessary (e.g., configuration files)
- Mount as read-only (
:ro) when write access is not needed - Never mount the Docker socket directly unless absolutely required
2.7 Networks
- All APEX OS containers join
apex_apex-net(bridge network) - Sandbox testing uses
apex-sandbox(isolated network) - No containers on the default Docker bridge network
2.8 Prohibited Practices
- ❌
--privilegedmode (unless documented exception with justification) - ❌
--net=host(breaks isolation) - ❌ Exposing ports directly to
0.0.0.0(use Traefik) - ❌ Hardcoded credentials in docker-compose.yml
- ❌
latesttag for critical infrastructure (PostgreSQL, Redis — pin versions)
3. Repository Standards
3.1 Git Hosting
All code is hosted on Gitea at git.apex.unstuck-path.com.
3.2 Organization
- Organization:
engineer/ - Repository naming:
kebab-case(lowercase, hyphen-separated) - Examples:
apex-os-docs,apex-status-api,engineer-workspace
3.3 Repository Requirements
Every repository must contain:
| File | Purpose | Required |
|---|---|---|
README.md |
Project description, setup, usage | ✅ Always |
.gitignore |
Files to exclude from version control | ✅ Always |
LICENSE |
Software license (if applicable) | ⚠️ When publishing |
docker-compose.yml |
Container definition (if deployable service) | ⚠️ When applicable |
Dockerfile |
Custom image build (if needed) | ⚠️ When applicable |
3.4 README Template
# [Project Name]
> [One-line description]
## Overview
[2-3 paragraph description of what this project does and why it exists]
## Architecture
[How it fits into APEX OS, what it depends on, what depends on it]
## Setup
[Step-by-step setup instructions]
## Usage
[How to use the project — commands, API endpoints, etc.]
## Configuration
[Environment variables, configuration files, etc.]
## Troubleshooting
[Common issues and their solutions]
## Related Documentation
- [link to related docs]
3.5 Employee Workspaces
Each employee has a dedicated Gitea workspace:
- Format:
engineer/{role-slug}-workspace - Examples:
engineer/engineer-workspace,engineer/research-workspace - Workspaces contain working files, drafts, research, and deliverables
- Deliverables are committed to the appropriate project repo when complete
4. Git Workflow
4.1 Branch Strategy
mainbranch is production — always deployable- Feature branches for all changes:
feature/{description} - Hotfix branches for urgent fixes:
hotfix/{description} - No direct commits to
mainfor critical services
4.2 Commit Messages
Format: verb: description
| Verb | Usage |
|---|---|
feat |
New feature or capability |
fix |
Bug fix |
docs |
Documentation update |
refactor |
Code restructuring without behavior change |
test |
Adding or updating tests |
chore |
Maintenance, dependency updates |
deploy |
Deployment-related changes |
security |
Security-related changes |
Examples:
feat: add health check endpoint to status API
fix: resolve Redis connection timeout in LiteLLM
docs: update TOOL_REGISTRY with Langfuse entry
refactor: simplify auto-recovery container detection
deploy: upgrade Grafana from 10.x to 11.x
security: rotate LiteLLM master key
4.3 Git Rules
- ❌ No force pushes to
main— ever - ❌ No uncommitted deployments — commit before deploying
- ✅ Meaningful commit messages — future you will thank present you
- ✅ Atomic commits — one logical change per commit
- ✅ Commit early, commit often — small, focused commits
5. Database Standards
5.1 Primary Database
PostgreSQL 16.x with pgvector extension — hosted in apex-postgres container.
5.2 Schema Organization
| Schema | Purpose | Owner |
|---|---|---|
apex |
Company operational data (tasks, projects, decisions, reflections, etc.) | APEX OS |
mem0 |
Shared knowledge base (embeddings, knowledge entries) | mem0 |
letta |
Agent framework state (agents, memory, tools) | Letta |
5.3 Naming Conventions
| Object | Convention | Example |
|---|---|---|
| Tables | snake_case |
engineer_decisions, task_status_changes |
| Columns | snake_case |
decision_date, created_at |
| Indexes | idx_{table}_{column} |
idx_tasks_status, idx_decisions_category |
| Primary keys | id (SERIAL or UUID) |
id SERIAL PRIMARY KEY |
| Foreign keys | {referenced_table}_id |
project_id, employee_id |
| Timestamps | {action}_at |
created_at, updated_at, completed_at |
| Booleans | is_{adjective} or has_{noun} |
is_active, has_approval |
5.4 Key Tables (apex schema)
| Table | Purpose | Key Columns |
|---|---|---|
projects |
Project tracking | id, name, status, created_at |
tasks |
Task management | id, project_id, assigned_to, status, priority, result |
engineer_decisions |
Decision audit log | id, decision_date, category, decision_summary, rationale |
reflections |
Post-project reflections | id, project_ref, what_worked, what_failed, improvement_applied |
deployments |
Deployment records | id, service_name, version, deployed_at, status |
recovery_log |
Auto-recovery actions | id, container_name, action, success, timestamp |
employee_registry |
Employee directory | id, employee_number, role, agent_id, status |
constitution_violations |
Constitution breaches | id, law_number, employee_id, description, timestamp |
performance_metrics |
Employee performance | id, employee_id, metric, value, period |
task_status_changes |
Task status audit trail | id, task_id, old_status, new_status, changed_at |
token_usage |
LLM token tracking | id, employee_id, model, tokens, cost, timestamp |
lifecycle_executions |
Engineering lifecycle tracking | id, phase, status, started_at, completed_at |
5.5 pgvector Configuration
- Extension:
pgvectorenabled in PostgreSQL - Index type: HNSW (Hierarchical Navigable Small World)
- Distance metric: Cosine similarity
- Embedding model:
nomic-embed-text(Ollama, 768 dimensions) - Usage: mem0 knowledge search, document similarity, semantic matching
-- Example: Creating a vector column with HNSW index
ALTER TABLE knowledge ADD COLUMN embedding vector(768);
CREATE INDEX idx_knowledge_embedding ON knowledge
USING hnsw (embedding vector_cosine_ops);
-- Example: Similarity search
SELECT *, 1 - (embedding <=> query_vector) AS similarity
FROM knowledge
ORDER BY embedding <=> query_vector
LIMIT 10;
5.6 Migration Rules
- All schema changes require a migration script
- Migrations must be reversible (include UP and DOWN)
- Backup database before any migration (Constitution Law 4)
- No direct DDL in production without a migration script
- Test migrations on a copy of the database first
- Log migration in
engineer_decisions
5.7 Database Security
- No direct SQL executed in production without a backup
- Connection strings stored in
.envfile (permissions600) - PostgreSQL credentials in Vaultwarden
- No superuser access for application connections
- Connection pooling via application-level management
6. API Standards
6.1 Design Principles
- RESTful endpoints following REST conventions
- JSON for all request and response bodies
- Consistent error handling and response format
- Documented with examples for every endpoint
6.2 URL Structure
https://{service}.apex.unstuck-path.com/v{version}/{resource}
Examples:
GET /v1/health
GET /v1/containers
POST /v1/tasks
PUT /v1/tasks/{id}
6.3 Health Check Endpoint
Every service must expose a health check endpoint:
GET /health OR GET /v1/health
Response (healthy):
{
"status": "healthy",
"service": "apex-status-api",
"version": "1.1",
"uptime": 86400,
"timestamp": "2026-07-02T12:00:00Z"
}
Response (unhealthy):
{
"status": "unhealthy",
"service": "apex-status-api",
"error": "Database connection failed",
"timestamp": "2026-07-02T12:00:00Z"
}
6.4 Authentication
| Method | Use Case |
|---|---|
| API Keys | Service-to-service authentication (e.g., LiteLLM master key) |
| Bearer Tokens | User-facing APIs with session management |
| No Auth | Health check endpoints only |
6.5 Rate Limiting
- External-facing APIs must implement rate limiting
- Default: 100 requests/minute per client
- Rate limit headers in response:
X-RateLimit-Limit: 100 X-RateLimit-Remaining: 95 X-RateLimit-Reset: 1625000000
6.6 Response Format
Success:
{
"success": true,
"data": { ... },
"metadata": {
"timestamp": "2026-07-02T12:00:00Z",
"request_id": "req_abc123"
}
}
Error:
{
"success": false,
"error": {
"code": "VALIDATION_ERROR",
"message": "Human-readable error description",
"details": { ... }
},
"metadata": {
"timestamp": "2026-07-02T12:00:00Z",
"request_id": "req_abc123"
}
}
6.7 HTTP Status Codes
| Code | Usage |
|---|---|
| 200 | Successful GET, PUT, PATCH |
| 201 | Successful POST (resource created) |
| 204 | Successful DELETE (no content) |
| 400 | Bad request (validation error) |
| 401 | Unauthorized (missing/invalid auth) |
| 403 | Forbidden (insufficient permissions) |
| 404 | Resource not found |
| 429 | Rate limited |
| 500 | Internal server error |
| 503 | Service unavailable |
7. Testing Requirements
7.1 Test Types
| Type | When Required | Description |
|---|---|---|
| Health Check Verification | Every deployment | Verify the service responds to health check after deployment |
| Smoke Tests | Infrastructure changes | Basic end-to-end verification that critical paths work |
| Integration Tests | New integrations | Verify service-to-service communication works correctly |
| Load Tests | New services (when applicable) | Verify service handles expected load without degradation |
| Security Tests | Security changes | Verify credentials, permissions, and access controls |
7.2 Post-Deployment Verification
After every deployment:
- ✅ Container starts and reaches
healthystate - ✅ Health check endpoint returns
200 - ✅ Service responds to basic requests
- ✅ Logs show no errors (check Loki)
- ✅ Metrics appear in Prometheus/Grafana
- ✅ Monitor for 15 minutes for stability
7.3 Regression Verification
After infrastructure changes:
- ✅ All existing containers remain healthy
- ✅ Traefik routes resolve correctly
- ✅ Database connections work
- ✅ LiteLLM proxy routes to all models
- ✅ Telegram bot responds to test message
8. Logging Standards
8.1 Format
Structured logging in JSON format is preferred:
{
"timestamp": "2026-07-02T12:00:00Z",
"level": "INFO",
"service": "apex-status-api",
"message": "Health check completed",
"data": {
"containers_checked": 20,
"healthy": 19,
"unhealthy": 1
}
}
8.2 Log Levels
| Level | Usage | When |
|---|---|---|
ERROR |
Something failed and requires attention | Service errors, failed operations, unhandled exceptions |
WARN |
Something unexpected but handled | Retry attempts, degraded performance, deprecated usage |
INFO |
Normal operations worth recording | Startup, shutdown, deployment, configuration changes |
DEBUG |
Detailed diagnostic information | Request details, variable states (development only) |
8.3 Log Pipeline
Container stdout/stderr → Promtail → Loki → Grafana
- Promtail tails container logs and labels them by container name
- Loki indexes and stores logs (label-based, not full-text)
- Grafana provides log exploration, search, and correlation with metrics
8.4 Decision-Level Logging
Significant operational decisions are logged to apex.engineer_decisions (not just container logs):
- Infrastructure changes
- Tool adoption/retirement
- Architecture decisions
- Security changes
- Incident responses
8.5 Prohibited Log Content
- ❌ Credentials, API keys, or tokens
- ❌ Personally identifiable information (PII)
- ❌ Full external API request/response bodies (summaries only)
- ❌ Database query results containing sensitive data
- ❌ Base64-encoded secrets or encrypted values
9. Monitoring Standards
9.1 Dashboard Requirements
- Grafana Executive Dashboard (UID:
a5jdct) — 12-panel overview - Every critical service must have representation in the dashboard
- Dashboard accessible at
grafana.apex.unstuck-path.com
9.2 Metrics Collection
Prometheus scrapes metrics from all instrumented services:
| Metric Category | Examples |
|---|---|
| Container health | Up/down status, restart count |
| Resource usage | CPU %, memory MB, disk I/O |
| Application metrics | Request count, latency, error rate |
| LLM metrics | Token usage, model routing, cost |
| Task metrics | Completion rate, blocked tasks, average duration |
9.3 Alerting Rules
| Alert | Condition | Action |
|---|---|---|
| Container Down | Health check fails 3x | Auto-recovery script restarts |
| High Memory | >90% memory usage | Alert in Grafana |
| Error Spike | Error rate >10% in 5 min | Alert in Grafana |
| Backup Failure | Backup script exits non-zero | Log alert + Telegram notification |
| Disk Space | <10% free disk space | Alert in Grafana + Telegram |
9.4 Auto-Recovery
- Script:
/opt/apex/scripts/(cron, every 5 minutes) - Logic: Check container health → restart unhealthy → log to
recovery_log→ escalate after 3 consecutive failures - Scope: All
apex-*containers - Logging: All actions logged to
apex.recovery_log
9.5 LLM Observability
Langfuse provides detailed LLM call tracing:
- Request/response pairs for all LLM calls
- Token usage breakdown per model
- Latency analysis per model/endpoint
- Cost tracking per employee
- Quality evaluation (when configured)
10. Error Handling
10.1 Principles
- Graceful degradation — When a dependency fails, provide reduced functionality rather than total failure
- Retry with exponential backoff — For transient failures (network, rate limits)
- Circuit breaker — For persistent external service failures
- Fallback responses — For non-critical failures
- Error logging with full context — What failed, why, what was the input
10.2 Retry Strategy
Attempt 1: Immediate
Attempt 2: Wait 1 second
Attempt 3: Wait 2 seconds
Attempt 4: Wait 4 seconds
Attempt 5: Wait 8 seconds
After 5 attempts: Log failure, escalate
10.3 Circuit Breaker Pattern
For external API dependencies:
CLOSED (normal) → error threshold exceeded → OPEN (fail fast)
OPEN → cool-down period → HALF-OPEN (test request)
HALF-OPEN → success → CLOSED | failure → OPEN
10.4 Fallback Hierarchy
When LiteLLM cloud routing fails:
- Retry with same model (exponential backoff)
- Try alternative cloud model (GPT-4o → Claude or vice versa)
- Fall back to local model (Ollama phi3:mini) with quality warning
- Log failure and escalate to Engineer
11. Versioning
11.1 Semantic Versioning
Custom APEX OS services follow SemVer (MAJOR.MINOR.PATCH):
| Component | When to Increment | Example |
|---|---|---|
| MAJOR | Breaking changes (API incompatibility) | 1.0 → 2.0 |
| MINOR | New features (backward compatible) | 1.0 → 1.1 |
| PATCH | Bug fixes (backward compatible) | 1.0.0 → 1.0.1 |
11.2 Docker Image Tags
- Custom services: tagged with SemVer (e.g.,
apex-status-api:1.1) - Third-party services: use upstream version tags (e.g.,
postgres:16.3) - Never use
latestfor critical infrastructure — pin specific versions - Retain previous image tags for 30 days (rollback capability)
11.3 API Versioning
- Version via URL path:
/v1/,/v2/ - Maintain backward compatibility within a major version
- Deprecation notice at least 30 days before removal
- Document breaking changes in release notes
12. Documentation Standards
12.1 Required Documentation
| Scope | Required Documents |
|---|---|
| Every project | README.md with setup, usage, and architecture |
| Every API | Endpoint documentation with request/response examples |
| Every deployment | Entry in engineer_decisions with rationale and rollback plan |
| Every infrastructure change | Before/after state documented |
| Every phase | Formal phase report in apex-os-docs |
| Every tool adoption | Certification report (see TOOL_REGISTRY.md §5) |
| Every troubleshooting resolution | Added to relevant troubleshooting guide |
12.2 Architecture Decision Records (ADRs)
For significant architecture decisions, create an ADR:
# ADR-{number}: {Title}
**Date:** {date}
**Status:** Proposed | Accepted | Deprecated | Superseded
## Context
[What is the issue or decision to be made?]
## Decision
[What was decided and why?]
## Consequences
[What are the implications — positive and negative?]
## Alternatives Considered
[What other options were evaluated?]
12.3 Troubleshooting Guides
Every common issue should have a documented resolution:
## Issue: [Description]
**Symptoms:** [What the user/operator sees]
**Root Cause:** [Why this happens]
**Resolution:**
1. [Step 1]
2. [Step 2]
**Prevention:** [How to prevent recurrence]
13. Rollback Procedures
13.1 Docker Container Rollback
# 1. Stop the service
docker-compose stop apex-example
# 2. Update docker-compose.yml to previous version tag
# image: example/image:1.0 (was 1.1)
# 3. Start with previous version
docker-compose up -d apex-example
# 4. Verify health
docker inspect --format='{{.State.Health.Status}}' apex-example
13.2 docker-compose.yml Rollback
# Restore from backup
cp docker-compose.yml.backup.{timestamp} docker-compose.yml
# Recreate containers
docker-compose up -d
13.3 Database Rollback
# Restore from daily backup
pg_restore -d apex /opt/apex/backups/apex_YYYYMMDD.dump
# Or restore full cluster
psql -f /opt/apex/backups/full_cluster_YYYYMMDD.sql
13.4 Configuration Rollback
# Restore .env from backup
cp /opt/apex/backups/.env.backup.{timestamp} /opt/apex/.env
# Restart affected services
docker-compose up -d
13.5 Full System Recovery
In case of catastrophic failure:
- Provision new VPS (or reset existing)
- Install Docker and Docker Compose
- Restore
/opt/apex/from offsite backup - Restore docker-compose.yml and .env
- Pull images:
docker-compose pull - Start services:
docker-compose up -d - Restore PostgreSQL:
pg_restorefrom backup - Verify all services: Check health endpoints
- Verify DNS: Ensure
*.apex.unstuck-path.comresolves - Notify Human CEO via alternative channel
14. Infrastructure Change Process
Every infrastructure change follows this process:
14.1 Pre-Change
- Create backup:
cp docker-compose.yml docker-compose.yml.backup.$(date +%Y%m%d%H%M%S) - Document the change in
engineer_decisions:- What is being changed
- Why it's being changed
- What the expected outcome is
- What the rollback plan is
- Test in isolation when possible (sandbox Docker network)
- Get approval for production changes (Constitution Law 6)
14.2 During Change
- Execute the change following documented steps
- Monitor actively — watch logs and metrics in real-time
- Verify health checks pass for all affected services
14.3 Post-Change
- Monitor for 15 minutes after deployment
- Run smoke tests — verify critical paths work
- Update documentation — TOOL_REGISTRY.md, README, etc.
- Log outcome in
engineer_decisions - Commit changes to Gitea
14.4 If Change Fails
- Execute rollback plan immediately
- Log failure with root cause analysis
- Create reflection — What went wrong? How to prevent next time?
- Escalate if rollback fails
15. Network Architecture
15.1 Docker Network
Network: apex_apex-net (bridge)
┌─────────────────────────────────────────────┐
│ apex_apex-net │
│ │
│ All apex-* containers are members │
│ Internal DNS resolves container names │
│ No external access except via Traefik │
└─────────────────────────────────────────────┘
15.2 Port Mapping
| External Port | Service | Notes |
|---|---|---|
| 80 | Traefik | HTTP → HTTPS redirect |
| 443 | Traefik | HTTPS termination + routing |
| (none others) | — | All other ports internal only |
15.3 Internal Service Ports
| Service | Internal Port | Access Via |
|---|---|---|
| PostgreSQL | 5432 | apex-postgres:5432 |
| Redis | 6379 | apex-redis:6379 |
| LiteLLM | 4000 | apex-litellm:4000 |
| Ollama | 11434 | apex-ollama:11434 |
| Letta | 8283 | apex-letta:8283 |
| Gitea | 3000 | git.apex.unstuck-path.com |
| Grafana | 3000 | grafana.apex.unstuck-path.com |
| n8n | 5678 | n8n.apex.unstuck-path.com |
| Langfuse | 3000 | langfuse.apex.unstuck-path.com |
| Open WebUI | 8080 | openwebui.apex.unstuck-path.com |
| Vaultwarden | 80 | vaultwarden.apex.unstuck-path.com |
| Code-Server | 8443 | code-server.apex.unstuck-path.com |
| Dockge | 5001 | dockge.apex.unstuck-path.com |
| Socket Proxy | 2375 | apex-socket-proxy:2375 |
| Status API | 3100 | Internal only |
| Prometheus | 9090 | Internal only |
| Loki | 3100 | Internal only |
15.4 DNS Structure
All services accessible via subdomains of apex.unstuck-path.com:
- Wildcard DNS:
*.apex.unstuck-path.com→62.72.3.145 - Traefik matches
Host()rules in container labels - SSL via Let's Encrypt (ACME HTTP-01 challenge)
16. Performance Standards
16.1 Response Time Targets
| Category | Target | Measurement |
|---|---|---|
| Health check endpoints | < 500ms | Prometheus histogram |
| API endpoints | < 2s (p95) | Langfuse / Prometheus |
| LLM responses (local) | < 10s | Langfuse |
| LLM responses (cloud) | < 30s | Langfuse |
| Page load (web UIs) | < 3s | Manual verification |
16.2 Availability Targets
| Service | Target | Monitoring |
|---|---|---|
| Core infrastructure (Traefik, PostgreSQL, Redis) | 99.5% uptime | Prometheus + auto-recovery |
| AI services (LiteLLM, Ollama, Letta) | 99% uptime | Prometheus + auto-recovery |
| Supporting services (Grafana, Gitea, n8n) | 98% uptime | Prometheus |
16.3 Resource Limits
| Metric | Warning Threshold | Critical Threshold |
|---|---|---|
| CPU usage (total) | 70% | 90% |
| Memory usage (total) | 75% | 90% |
| Disk usage | 80% | 90% |
| Container restart count | 3/hour | 5/hour |
17. Security Engineering
17.1 Principle of Least Privilege
Every container, service, and agent has the minimum permissions required for its function:
- Docker Socket Proxy restricts Docker API to read-only queries
- Code-Server has access to config/data only (not full
/opt/apex/) - n8n has no Docker socket access
- Employees have role-specific database access
17.2 Credential Lifecycle
Generate → Store in Vaultwarden → Reference in .env → Pass as env var → Rotate every 60 days
17.3 Security Checklist for New Services
- No default credentials
- No privileged container mode
- Non-root user where possible
- Read-only mounts where possible
- No direct Docker socket access (use proxy)
- Health check configured
- Traefik TLS termination (no plain HTTP)
- Credentials in Vaultwarden
- CVE check on image and dependencies
18. Change History
| Date | Version | Author | Changes |
|---|---|---|---|
| Phase 2 | 0.1 | Engineer (#1) | Initial standards — Docker basics, database schema |
| Phase 3 | 0.5 | Engineer (#1) | Added Git workflow, repository standards |
| Phase 5 | 1.0 | Engineer (#1) | Added API standards, testing requirements, error handling |
| Phase 5.5 | 1.5 | Engineer (#1) | Added monitoring standards, auto-recovery, logging pipeline |
| Phase 7 | 2.0 | Engineer (#1) | Full standards formalization. Added network architecture, performance standards, security engineering, rollback procedures, infrastructure change process. Comprehensive coverage of all engineering aspects. |
These standards are not suggestions — they are requirements. Every technical deliverable in APEX OS must comply. When in doubt, err on the side of more documentation, more testing, and more caution.
Cross-references: APEX_CONSTITUTION.md · TOOL_REGISTRY.md · EMPLOYEE_HANDBOOK.md · COMPANY_STRUCTURE.md