Phase 8: Complete APEX OS documentation suite - Constitution, Structure, Handbook, Standards, Security, Evolution, NL Guide, Tool Registry, Roadmap, Changelog
This commit is contained in:
@@ -0,0 +1,903 @@
|
||||
# APEX OS Engineering Standards
|
||||
|
||||
> **The definitive engineering standards governing all infrastructure, code, databases, APIs, testing, monitoring, and operations within APEX OS.**
|
||||
> All employees producing technical deliverables must adhere to these standards.
|
||||
|
||||
**Version:** 2.0
|
||||
**Last Updated:** Phase 7 — CEO Command Center
|
||||
**Classification:** CORE — All technical work must comply
|
||||
**Maintainer:** Engineer (#1)
|
||||
**Cross-references:** [APEX_CONSTITUTION.md](APEX_CONSTITUTION.md) · [TOOL_REGISTRY.md](TOOL_REGISTRY.md) · [EMPLOYEE_HANDBOOK.md](EMPLOYEE_HANDBOOK.md)
|
||||
|
||||
---
|
||||
|
||||
## Table of Contents
|
||||
|
||||
- [1. Architecture Principles](#1-architecture-principles)
|
||||
- [2. Docker Standards](#2-docker-standards)
|
||||
- [3. Repository Standards](#3-repository-standards)
|
||||
- [4. Git Workflow](#4-git-workflow)
|
||||
- [5. Database Standards](#5-database-standards)
|
||||
- [6. API Standards](#6-api-standards)
|
||||
- [7. Testing Requirements](#7-testing-requirements)
|
||||
- [8. Logging Standards](#8-logging-standards)
|
||||
- [9. Monitoring Standards](#9-monitoring-standards)
|
||||
- [10. Error Handling](#10-error-handling)
|
||||
- [11. Versioning](#11-versioning)
|
||||
- [12. Documentation Standards](#12-documentation-standards)
|
||||
- [13. Rollback Procedures](#13-rollback-procedures)
|
||||
- [14. Infrastructure Change Process](#14-infrastructure-change-process)
|
||||
- [15. Network Architecture](#15-network-architecture)
|
||||
- [16. Performance Standards](#16-performance-standards)
|
||||
- [17. Security Engineering](#17-security-engineering)
|
||||
- [18. Change History](#18-change-history)
|
||||
|
||||
---
|
||||
|
||||
## 1. Architecture Principles
|
||||
|
||||
### 1.1 Microservices via Docker
|
||||
Every service in APEX OS runs as an isolated Docker container. This provides:
|
||||
- **Reproducibility** — Identical environments from dev to production
|
||||
- **Isolation** — A failing service doesn't take down others
|
||||
- **Scalability** — Services can be independently scaled
|
||||
- **Clean rollback** — Previous image tags can be restored instantly
|
||||
|
||||
### 1.2 Single Responsibility
|
||||
Each container serves **one purpose**. A database is a database. A proxy is a proxy. An API is an API. No monolithic containers combining multiple services.
|
||||
|
||||
### 1.3 Internal Communication
|
||||
Services communicate via the **Docker network** (`apex_apex-net`). Internal service discovery uses Docker DNS (container names resolve to IPs within the network).
|
||||
|
||||
```
|
||||
Service A ──(apex_apex-net)──► Service B
|
||||
│
|
||||
└── DNS: apex-postgres:5432
|
||||
└── DNS: apex-redis:6379
|
||||
└── DNS: apex-litellm:4000
|
||||
```
|
||||
|
||||
### 1.4 External Access
|
||||
All external traffic is routed through **Traefik** reverse proxy:
|
||||
- HTTP (80) → Redirected to HTTPS (443)
|
||||
- HTTPS (443) → Routed to containers via Docker labels
|
||||
- Let's Encrypt certificates auto-renewed
|
||||
- No service directly exposes ports to the internet (except Traefik)
|
||||
|
||||
### 1.5 State Management
|
||||
- **Persistent state** stored in PostgreSQL (relational data, vectors, agent state)
|
||||
- **Cache/ephemeral state** stored in Redis (LLM response cache, sessions)
|
||||
- **File state** stored in Docker volumes (Gitea repos, Vaultwarden data, n8n workflows)
|
||||
- **No state in containers** — containers are ephemeral and replaceable
|
||||
|
||||
### 1.6 Cost-Driven Design
|
||||
Architecture decisions optimize for cost first, then speed, then quality (Constitution §5, Principle 3):
|
||||
- Local models for routine tasks (Ollama → $0)
|
||||
- Cloud models only when quality demands it (OpenRouter → metered)
|
||||
- Caching for repeated queries (Redis)
|
||||
- Batch operations over individual API calls
|
||||
|
||||
---
|
||||
|
||||
## 2. Docker Standards
|
||||
|
||||
### 2.1 docker-compose.yml Requirements
|
||||
|
||||
Every service in `docker-compose.yml` must include:
|
||||
|
||||
```yaml
|
||||
services:
|
||||
apex-example:
|
||||
image: example/image:version # Pinned version tag
|
||||
container_name: apex-example # apex- prefix required
|
||||
restart: unless-stopped # Auto-restart on failure
|
||||
healthcheck: # Health check required
|
||||
test: ["CMD", "curl", "-f", "http://localhost:PORT/health"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
start_period: 40s
|
||||
networks:
|
||||
- apex-net # Internal network
|
||||
labels:
|
||||
- "traefik.enable=true" # If externally accessible
|
||||
- "traefik.http.routers.example.rule=Host(`example.apex.unstuck-path.com`)"
|
||||
- "traefik.http.routers.example.tls.certresolver=letsencrypt"
|
||||
volumes:
|
||||
- example_data:/data # Named volumes for persistence
|
||||
environment:
|
||||
- CONFIG_VAR=${CONFIG_VAR} # From .env file, never hardcoded
|
||||
```
|
||||
|
||||
### 2.2 Container Naming
|
||||
- **Prefix:** All containers start with `apex-`
|
||||
- **Format:** `apex-{service-name}` (lowercase, hyphenated)
|
||||
- **Examples:** `apex-postgres`, `apex-traefik`, `apex-litellm`, `apex-status-api`
|
||||
|
||||
### 2.3 Health Checks
|
||||
**Required** on every container. Health check types:
|
||||
|
||||
| Type | When to Use | Example |
|
||||
|------|-------------|---------|
|
||||
| HTTP | Service has a web endpoint | `curl -f http://localhost:PORT/health` |
|
||||
| TCP | Service listens on a port | `pg_isready` for PostgreSQL |
|
||||
| CMD | Custom health logic | `redis-cli ping` |
|
||||
|
||||
### 2.4 Restart Policy
|
||||
- **Default:** `restart: unless-stopped`
|
||||
- This ensures containers restart after crashes but respect manual stops
|
||||
- Never use `restart: always` (prevents intentional stops)
|
||||
|
||||
### 2.5 Resource Awareness
|
||||
- Set memory limits for resource-intensive services:
|
||||
```yaml
|
||||
deploy:
|
||||
resources:
|
||||
limits:
|
||||
memory: 512M
|
||||
```
|
||||
- Monitor resource usage via Prometheus/Grafana
|
||||
- Document expected resource usage in TOOL_REGISTRY.md
|
||||
|
||||
### 2.6 Volumes
|
||||
- Use **named volumes** for persistent data: `apex_postgres_data`, `apex_gitea_data`
|
||||
- Use **bind mounts** only when necessary (e.g., configuration files)
|
||||
- Mount as **read-only** (`:ro`) when write access is not needed
|
||||
- Never mount the Docker socket directly unless absolutely required
|
||||
|
||||
### 2.7 Networks
|
||||
- All APEX OS containers join `apex_apex-net` (bridge network)
|
||||
- Sandbox testing uses `apex-sandbox` (isolated network)
|
||||
- No containers on the default Docker bridge network
|
||||
|
||||
### 2.8 Prohibited Practices
|
||||
- ❌ `--privileged` mode (unless documented exception with justification)
|
||||
- ❌ `--net=host` (breaks isolation)
|
||||
- ❌ Exposing ports directly to `0.0.0.0` (use Traefik)
|
||||
- ❌ Hardcoded credentials in docker-compose.yml
|
||||
- ❌ `latest` tag for critical infrastructure (PostgreSQL, Redis — pin versions)
|
||||
|
||||
---
|
||||
|
||||
## 3. Repository Standards
|
||||
|
||||
### 3.1 Git Hosting
|
||||
All code is hosted on **Gitea** at `git.apex.unstuck-path.com`.
|
||||
|
||||
### 3.2 Organization
|
||||
- **Organization:** `engineer/`
|
||||
- **Repository naming:** `kebab-case` (lowercase, hyphen-separated)
|
||||
- **Examples:** `apex-os-docs`, `apex-status-api`, `engineer-workspace`
|
||||
|
||||
### 3.3 Repository Requirements
|
||||
Every repository must contain:
|
||||
|
||||
| File | Purpose | Required |
|
||||
|------|---------|----------|
|
||||
| `README.md` | Project description, setup, usage | ✅ Always |
|
||||
| `.gitignore` | Files to exclude from version control | ✅ Always |
|
||||
| `LICENSE` | Software license (if applicable) | ⚠️ When publishing |
|
||||
| `docker-compose.yml` | Container definition (if deployable service) | ⚠️ When applicable |
|
||||
| `Dockerfile` | Custom image build (if needed) | ⚠️ When applicable |
|
||||
|
||||
### 3.4 README Template
|
||||
```markdown
|
||||
# [Project Name]
|
||||
|
||||
> [One-line description]
|
||||
|
||||
## Overview
|
||||
[2-3 paragraph description of what this project does and why it exists]
|
||||
|
||||
## Architecture
|
||||
[How it fits into APEX OS, what it depends on, what depends on it]
|
||||
|
||||
## Setup
|
||||
[Step-by-step setup instructions]
|
||||
|
||||
## Usage
|
||||
[How to use the project — commands, API endpoints, etc.]
|
||||
|
||||
## Configuration
|
||||
[Environment variables, configuration files, etc.]
|
||||
|
||||
## Troubleshooting
|
||||
[Common issues and their solutions]
|
||||
|
||||
## Related Documentation
|
||||
- [link to related docs]
|
||||
```
|
||||
|
||||
### 3.5 Employee Workspaces
|
||||
Each employee has a dedicated Gitea workspace:
|
||||
- Format: `engineer/{role-slug}-workspace`
|
||||
- Examples: `engineer/engineer-workspace`, `engineer/research-workspace`
|
||||
- Workspaces contain working files, drafts, research, and deliverables
|
||||
- Deliverables are committed to the appropriate project repo when complete
|
||||
|
||||
---
|
||||
|
||||
## 4. Git Workflow
|
||||
|
||||
### 4.1 Branch Strategy
|
||||
- **`main`** branch is production — always deployable
|
||||
- **Feature branches** for all changes: `feature/{description}`
|
||||
- **Hotfix branches** for urgent fixes: `hotfix/{description}`
|
||||
- **No direct commits to `main`** for critical services
|
||||
|
||||
### 4.2 Commit Messages
|
||||
Format: `verb: description`
|
||||
|
||||
| Verb | Usage |
|
||||
|------|-------|
|
||||
| `feat` | New feature or capability |
|
||||
| `fix` | Bug fix |
|
||||
| `docs` | Documentation update |
|
||||
| `refactor` | Code restructuring without behavior change |
|
||||
| `test` | Adding or updating tests |
|
||||
| `chore` | Maintenance, dependency updates |
|
||||
| `deploy` | Deployment-related changes |
|
||||
| `security` | Security-related changes |
|
||||
|
||||
**Examples:**
|
||||
```
|
||||
feat: add health check endpoint to status API
|
||||
fix: resolve Redis connection timeout in LiteLLM
|
||||
docs: update TOOL_REGISTRY with Langfuse entry
|
||||
refactor: simplify auto-recovery container detection
|
||||
deploy: upgrade Grafana from 10.x to 11.x
|
||||
security: rotate LiteLLM master key
|
||||
```
|
||||
|
||||
### 4.3 Git Rules
|
||||
- ❌ **No force pushes to `main`** — ever
|
||||
- ❌ **No uncommitted deployments** — commit before deploying
|
||||
- ✅ **Meaningful commit messages** — future you will thank present you
|
||||
- ✅ **Atomic commits** — one logical change per commit
|
||||
- ✅ **Commit early, commit often** — small, focused commits
|
||||
|
||||
---
|
||||
|
||||
## 5. Database Standards
|
||||
|
||||
### 5.1 Primary Database
|
||||
**PostgreSQL 16.x** with `pgvector` extension — hosted in `apex-postgres` container.
|
||||
|
||||
### 5.2 Schema Organization
|
||||
|
||||
| Schema | Purpose | Owner |
|
||||
|--------|---------|-------|
|
||||
| `apex` | Company operational data (tasks, projects, decisions, reflections, etc.) | APEX OS |
|
||||
| `mem0` | Shared knowledge base (embeddings, knowledge entries) | mem0 |
|
||||
| `letta` | Agent framework state (agents, memory, tools) | Letta |
|
||||
|
||||
### 5.3 Naming Conventions
|
||||
|
||||
| Object | Convention | Example |
|
||||
|--------|-----------|---------|
|
||||
| Tables | `snake_case` | `engineer_decisions`, `task_status_changes` |
|
||||
| Columns | `snake_case` | `decision_date`, `created_at` |
|
||||
| Indexes | `idx_{table}_{column}` | `idx_tasks_status`, `idx_decisions_category` |
|
||||
| Primary keys | `id` (SERIAL or UUID) | `id SERIAL PRIMARY KEY` |
|
||||
| Foreign keys | `{referenced_table}_id` | `project_id`, `employee_id` |
|
||||
| Timestamps | `{action}_at` | `created_at`, `updated_at`, `completed_at` |
|
||||
| Booleans | `is_{adjective}` or `has_{noun}` | `is_active`, `has_approval` |
|
||||
|
||||
### 5.4 Key Tables (apex schema)
|
||||
|
||||
| Table | Purpose | Key Columns |
|
||||
|-------|---------|-------------|
|
||||
| `projects` | Project tracking | `id`, `name`, `status`, `created_at` |
|
||||
| `tasks` | Task management | `id`, `project_id`, `assigned_to`, `status`, `priority`, `result` |
|
||||
| `engineer_decisions` | Decision audit log | `id`, `decision_date`, `category`, `decision_summary`, `rationale` |
|
||||
| `reflections` | Post-project reflections | `id`, `project_ref`, `what_worked`, `what_failed`, `improvement_applied` |
|
||||
| `deployments` | Deployment records | `id`, `service_name`, `version`, `deployed_at`, `status` |
|
||||
| `recovery_log` | Auto-recovery actions | `id`, `container_name`, `action`, `success`, `timestamp` |
|
||||
| `employee_registry` | Employee directory | `id`, `employee_number`, `role`, `agent_id`, `status` |
|
||||
| `constitution_violations` | Constitution breaches | `id`, `law_number`, `employee_id`, `description`, `timestamp` |
|
||||
| `performance_metrics` | Employee performance | `id`, `employee_id`, `metric`, `value`, `period` |
|
||||
| `task_status_changes` | Task status audit trail | `id`, `task_id`, `old_status`, `new_status`, `changed_at` |
|
||||
| `token_usage` | LLM token tracking | `id`, `employee_id`, `model`, `tokens`, `cost`, `timestamp` |
|
||||
| `lifecycle_executions` | Engineering lifecycle tracking | `id`, `phase`, `status`, `started_at`, `completed_at` |
|
||||
|
||||
### 5.5 pgvector Configuration
|
||||
- **Extension:** `pgvector` enabled in PostgreSQL
|
||||
- **Index type:** HNSW (Hierarchical Navigable Small World)
|
||||
- **Distance metric:** Cosine similarity
|
||||
- **Embedding model:** `nomic-embed-text` (Ollama, 768 dimensions)
|
||||
- **Usage:** mem0 knowledge search, document similarity, semantic matching
|
||||
|
||||
```sql
|
||||
-- Example: Creating a vector column with HNSW index
|
||||
ALTER TABLE knowledge ADD COLUMN embedding vector(768);
|
||||
CREATE INDEX idx_knowledge_embedding ON knowledge
|
||||
USING hnsw (embedding vector_cosine_ops);
|
||||
|
||||
-- Example: Similarity search
|
||||
SELECT *, 1 - (embedding <=> query_vector) AS similarity
|
||||
FROM knowledge
|
||||
ORDER BY embedding <=> query_vector
|
||||
LIMIT 10;
|
||||
```
|
||||
|
||||
### 5.6 Migration Rules
|
||||
- All schema changes require a **migration script**
|
||||
- Migrations must be **reversible** (include UP and DOWN)
|
||||
- **Backup database before any migration** (Constitution Law 4)
|
||||
- No direct DDL in production without a migration script
|
||||
- Test migrations on a copy of the database first
|
||||
- Log migration in `engineer_decisions`
|
||||
|
||||
### 5.7 Database Security
|
||||
- No direct SQL executed in production without a backup
|
||||
- Connection strings stored in `.env` file (permissions `600`)
|
||||
- PostgreSQL credentials in Vaultwarden
|
||||
- No superuser access for application connections
|
||||
- Connection pooling via application-level management
|
||||
|
||||
---
|
||||
|
||||
## 6. API Standards
|
||||
|
||||
### 6.1 Design Principles
|
||||
- **RESTful** endpoints following REST conventions
|
||||
- **JSON** for all request and response bodies
|
||||
- **Consistent** error handling and response format
|
||||
- **Documented** with examples for every endpoint
|
||||
|
||||
### 6.2 URL Structure
|
||||
```
|
||||
https://{service}.apex.unstuck-path.com/v{version}/{resource}
|
||||
|
||||
Examples:
|
||||
GET /v1/health
|
||||
GET /v1/containers
|
||||
POST /v1/tasks
|
||||
PUT /v1/tasks/{id}
|
||||
```
|
||||
|
||||
### 6.3 Health Check Endpoint
|
||||
**Every service must expose a health check endpoint:**
|
||||
|
||||
```
|
||||
GET /health OR GET /v1/health
|
||||
|
||||
Response (healthy):
|
||||
{
|
||||
"status": "healthy",
|
||||
"service": "apex-status-api",
|
||||
"version": "1.1",
|
||||
"uptime": 86400,
|
||||
"timestamp": "2026-07-02T12:00:00Z"
|
||||
}
|
||||
|
||||
Response (unhealthy):
|
||||
{
|
||||
"status": "unhealthy",
|
||||
"service": "apex-status-api",
|
||||
"error": "Database connection failed",
|
||||
"timestamp": "2026-07-02T12:00:00Z"
|
||||
}
|
||||
```
|
||||
|
||||
### 6.4 Authentication
|
||||
| Method | Use Case |
|
||||
|--------|----------|
|
||||
| API Keys | Service-to-service authentication (e.g., LiteLLM master key) |
|
||||
| Bearer Tokens | User-facing APIs with session management |
|
||||
| No Auth | Health check endpoints only |
|
||||
|
||||
### 6.5 Rate Limiting
|
||||
- External-facing APIs must implement rate limiting
|
||||
- Default: 100 requests/minute per client
|
||||
- Rate limit headers in response:
|
||||
```
|
||||
X-RateLimit-Limit: 100
|
||||
X-RateLimit-Remaining: 95
|
||||
X-RateLimit-Reset: 1625000000
|
||||
```
|
||||
|
||||
### 6.6 Response Format
|
||||
|
||||
**Success:**
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"data": { ... },
|
||||
"metadata": {
|
||||
"timestamp": "2026-07-02T12:00:00Z",
|
||||
"request_id": "req_abc123"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
**Error:**
|
||||
```json
|
||||
{
|
||||
"success": false,
|
||||
"error": {
|
||||
"code": "VALIDATION_ERROR",
|
||||
"message": "Human-readable error description",
|
||||
"details": { ... }
|
||||
},
|
||||
"metadata": {
|
||||
"timestamp": "2026-07-02T12:00:00Z",
|
||||
"request_id": "req_abc123"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 6.7 HTTP Status Codes
|
||||
|
||||
| Code | Usage |
|
||||
|------|-------|
|
||||
| 200 | Successful GET, PUT, PATCH |
|
||||
| 201 | Successful POST (resource created) |
|
||||
| 204 | Successful DELETE (no content) |
|
||||
| 400 | Bad request (validation error) |
|
||||
| 401 | Unauthorized (missing/invalid auth) |
|
||||
| 403 | Forbidden (insufficient permissions) |
|
||||
| 404 | Resource not found |
|
||||
| 429 | Rate limited |
|
||||
| 500 | Internal server error |
|
||||
| 503 | Service unavailable |
|
||||
|
||||
---
|
||||
|
||||
## 7. Testing Requirements
|
||||
|
||||
### 7.1 Test Types
|
||||
|
||||
| Type | When Required | Description |
|
||||
|------|---------------|-------------|
|
||||
| **Health Check Verification** | Every deployment | Verify the service responds to health check after deployment |
|
||||
| **Smoke Tests** | Infrastructure changes | Basic end-to-end verification that critical paths work |
|
||||
| **Integration Tests** | New integrations | Verify service-to-service communication works correctly |
|
||||
| **Load Tests** | New services (when applicable) | Verify service handles expected load without degradation |
|
||||
| **Security Tests** | Security changes | Verify credentials, permissions, and access controls |
|
||||
|
||||
### 7.2 Post-Deployment Verification
|
||||
After every deployment:
|
||||
1. ✅ Container starts and reaches `healthy` state
|
||||
2. ✅ Health check endpoint returns `200`
|
||||
3. ✅ Service responds to basic requests
|
||||
4. ✅ Logs show no errors (check Loki)
|
||||
5. ✅ Metrics appear in Prometheus/Grafana
|
||||
6. ✅ Monitor for 15 minutes for stability
|
||||
|
||||
### 7.3 Regression Verification
|
||||
After infrastructure changes:
|
||||
1. ✅ All existing containers remain healthy
|
||||
2. ✅ Traefik routes resolve correctly
|
||||
3. ✅ Database connections work
|
||||
4. ✅ LiteLLM proxy routes to all models
|
||||
5. ✅ Telegram bot responds to test message
|
||||
|
||||
---
|
||||
|
||||
## 8. Logging Standards
|
||||
|
||||
### 8.1 Format
|
||||
**Structured logging in JSON format** is preferred:
|
||||
```json
|
||||
{
|
||||
"timestamp": "2026-07-02T12:00:00Z",
|
||||
"level": "INFO",
|
||||
"service": "apex-status-api",
|
||||
"message": "Health check completed",
|
||||
"data": {
|
||||
"containers_checked": 20,
|
||||
"healthy": 19,
|
||||
"unhealthy": 1
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 8.2 Log Levels
|
||||
|
||||
| Level | Usage | When |
|
||||
|-------|-------|------|
|
||||
| `ERROR` | Something failed and requires attention | Service errors, failed operations, unhandled exceptions |
|
||||
| `WARN` | Something unexpected but handled | Retry attempts, degraded performance, deprecated usage |
|
||||
| `INFO` | Normal operations worth recording | Startup, shutdown, deployment, configuration changes |
|
||||
| `DEBUG` | Detailed diagnostic information | Request details, variable states (development only) |
|
||||
|
||||
### 8.3 Log Pipeline
|
||||
```
|
||||
Container stdout/stderr → Promtail → Loki → Grafana
|
||||
```
|
||||
- **Promtail** tails container logs and labels them by container name
|
||||
- **Loki** indexes and stores logs (label-based, not full-text)
|
||||
- **Grafana** provides log exploration, search, and correlation with metrics
|
||||
|
||||
### 8.4 Decision-Level Logging
|
||||
Significant operational decisions are logged to `apex.engineer_decisions` (not just container logs):
|
||||
- Infrastructure changes
|
||||
- Tool adoption/retirement
|
||||
- Architecture decisions
|
||||
- Security changes
|
||||
- Incident responses
|
||||
|
||||
### 8.5 Prohibited Log Content
|
||||
- ❌ Credentials, API keys, or tokens
|
||||
- ❌ Personally identifiable information (PII)
|
||||
- ❌ Full external API request/response bodies (summaries only)
|
||||
- ❌ Database query results containing sensitive data
|
||||
- ❌ Base64-encoded secrets or encrypted values
|
||||
|
||||
---
|
||||
|
||||
## 9. Monitoring Standards
|
||||
|
||||
### 9.1 Dashboard Requirements
|
||||
- **Grafana Executive Dashboard** (UID: `a5jdct`) — 12-panel overview
|
||||
- Every critical service must have representation in the dashboard
|
||||
- Dashboard accessible at `grafana.apex.unstuck-path.com`
|
||||
|
||||
### 9.2 Metrics Collection
|
||||
**Prometheus** scrapes metrics from all instrumented services:
|
||||
|
||||
| Metric Category | Examples |
|
||||
|-----------------|---------|
|
||||
| Container health | Up/down status, restart count |
|
||||
| Resource usage | CPU %, memory MB, disk I/O |
|
||||
| Application metrics | Request count, latency, error rate |
|
||||
| LLM metrics | Token usage, model routing, cost |
|
||||
| Task metrics | Completion rate, blocked tasks, average duration |
|
||||
|
||||
### 9.3 Alerting Rules
|
||||
| Alert | Condition | Action |
|
||||
|-------|-----------|--------|
|
||||
| Container Down | Health check fails 3x | Auto-recovery script restarts |
|
||||
| High Memory | >90% memory usage | Alert in Grafana |
|
||||
| Error Spike | Error rate >10% in 5 min | Alert in Grafana |
|
||||
| Backup Failure | Backup script exits non-zero | Log alert + Telegram notification |
|
||||
| Disk Space | <10% free disk space | Alert in Grafana + Telegram |
|
||||
|
||||
### 9.4 Auto-Recovery
|
||||
- **Script:** `/opt/apex/scripts/` (cron, every 5 minutes)
|
||||
- **Logic:** Check container health → restart unhealthy → log to `recovery_log` → escalate after 3 consecutive failures
|
||||
- **Scope:** All `apex-*` containers
|
||||
- **Logging:** All actions logged to `apex.recovery_log`
|
||||
|
||||
### 9.5 LLM Observability
|
||||
**Langfuse** provides detailed LLM call tracing:
|
||||
- Request/response pairs for all LLM calls
|
||||
- Token usage breakdown per model
|
||||
- Latency analysis per model/endpoint
|
||||
- Cost tracking per employee
|
||||
- Quality evaluation (when configured)
|
||||
|
||||
---
|
||||
|
||||
## 10. Error Handling
|
||||
|
||||
### 10.1 Principles
|
||||
1. **Graceful degradation** — When a dependency fails, provide reduced functionality rather than total failure
|
||||
2. **Retry with exponential backoff** — For transient failures (network, rate limits)
|
||||
3. **Circuit breaker** — For persistent external service failures
|
||||
4. **Fallback responses** — For non-critical failures
|
||||
5. **Error logging with full context** — What failed, why, what was the input
|
||||
|
||||
### 10.2 Retry Strategy
|
||||
```
|
||||
Attempt 1: Immediate
|
||||
Attempt 2: Wait 1 second
|
||||
Attempt 3: Wait 2 seconds
|
||||
Attempt 4: Wait 4 seconds
|
||||
Attempt 5: Wait 8 seconds
|
||||
After 5 attempts: Log failure, escalate
|
||||
```
|
||||
|
||||
### 10.3 Circuit Breaker Pattern
|
||||
For external API dependencies:
|
||||
```
|
||||
CLOSED (normal) → error threshold exceeded → OPEN (fail fast)
|
||||
OPEN → cool-down period → HALF-OPEN (test request)
|
||||
HALF-OPEN → success → CLOSED | failure → OPEN
|
||||
```
|
||||
|
||||
### 10.4 Fallback Hierarchy
|
||||
When LiteLLM cloud routing fails:
|
||||
1. Retry with same model (exponential backoff)
|
||||
2. Try alternative cloud model (GPT-4o → Claude or vice versa)
|
||||
3. Fall back to local model (Ollama phi3:mini) with quality warning
|
||||
4. Log failure and escalate to Engineer
|
||||
|
||||
---
|
||||
|
||||
## 11. Versioning
|
||||
|
||||
### 11.1 Semantic Versioning
|
||||
Custom APEX OS services follow **SemVer** (MAJOR.MINOR.PATCH):
|
||||
|
||||
| Component | When to Increment | Example |
|
||||
|-----------|-------------------|---------|
|
||||
| MAJOR | Breaking changes (API incompatibility) | 1.0 → 2.0 |
|
||||
| MINOR | New features (backward compatible) | 1.0 → 1.1 |
|
||||
| PATCH | Bug fixes (backward compatible) | 1.0.0 → 1.0.1 |
|
||||
|
||||
### 11.2 Docker Image Tags
|
||||
- Custom services: tagged with SemVer (e.g., `apex-status-api:1.1`)
|
||||
- Third-party services: use upstream version tags (e.g., `postgres:16.3`)
|
||||
- **Never use `latest` for critical infrastructure** — pin specific versions
|
||||
- Retain previous image tags for 30 days (rollback capability)
|
||||
|
||||
### 11.3 API Versioning
|
||||
- Version via URL path: `/v1/`, `/v2/`
|
||||
- Maintain backward compatibility within a major version
|
||||
- Deprecation notice at least 30 days before removal
|
||||
- Document breaking changes in release notes
|
||||
|
||||
---
|
||||
|
||||
## 12. Documentation Standards
|
||||
|
||||
### 12.1 Required Documentation
|
||||
|
||||
| Scope | Required Documents |
|
||||
|-------|--------------------|
|
||||
| Every project | `README.md` with setup, usage, and architecture |
|
||||
| Every API | Endpoint documentation with request/response examples |
|
||||
| Every deployment | Entry in `engineer_decisions` with rationale and rollback plan |
|
||||
| Every infrastructure change | Before/after state documented |
|
||||
| Every phase | Formal phase report in `apex-os-docs` |
|
||||
| Every tool adoption | Certification report (see TOOL_REGISTRY.md §5) |
|
||||
| Every troubleshooting resolution | Added to relevant troubleshooting guide |
|
||||
|
||||
### 12.2 Architecture Decision Records (ADRs)
|
||||
For significant architecture decisions, create an ADR:
|
||||
```markdown
|
||||
# ADR-{number}: {Title}
|
||||
|
||||
**Date:** {date}
|
||||
**Status:** Proposed | Accepted | Deprecated | Superseded
|
||||
|
||||
## Context
|
||||
[What is the issue or decision to be made?]
|
||||
|
||||
## Decision
|
||||
[What was decided and why?]
|
||||
|
||||
## Consequences
|
||||
[What are the implications — positive and negative?]
|
||||
|
||||
## Alternatives Considered
|
||||
[What other options were evaluated?]
|
||||
```
|
||||
|
||||
### 12.3 Troubleshooting Guides
|
||||
Every common issue should have a documented resolution:
|
||||
```markdown
|
||||
## Issue: [Description]
|
||||
|
||||
**Symptoms:** [What the user/operator sees]
|
||||
**Root Cause:** [Why this happens]
|
||||
**Resolution:**
|
||||
1. [Step 1]
|
||||
2. [Step 2]
|
||||
**Prevention:** [How to prevent recurrence]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 13. Rollback Procedures
|
||||
|
||||
### 13.1 Docker Container Rollback
|
||||
```bash
|
||||
# 1. Stop the service
|
||||
docker-compose stop apex-example
|
||||
|
||||
# 2. Update docker-compose.yml to previous version tag
|
||||
# image: example/image:1.0 (was 1.1)
|
||||
|
||||
# 3. Start with previous version
|
||||
docker-compose up -d apex-example
|
||||
|
||||
# 4. Verify health
|
||||
docker inspect --format='{{.State.Health.Status}}' apex-example
|
||||
```
|
||||
|
||||
### 13.2 docker-compose.yml Rollback
|
||||
```bash
|
||||
# Restore from backup
|
||||
cp docker-compose.yml.backup.{timestamp} docker-compose.yml
|
||||
|
||||
# Recreate containers
|
||||
docker-compose up -d
|
||||
```
|
||||
|
||||
### 13.3 Database Rollback
|
||||
```bash
|
||||
# Restore from daily backup
|
||||
pg_restore -d apex /opt/apex/backups/apex_YYYYMMDD.dump
|
||||
|
||||
# Or restore full cluster
|
||||
psql -f /opt/apex/backups/full_cluster_YYYYMMDD.sql
|
||||
```
|
||||
|
||||
### 13.4 Configuration Rollback
|
||||
```bash
|
||||
# Restore .env from backup
|
||||
cp /opt/apex/backups/.env.backup.{timestamp} /opt/apex/.env
|
||||
|
||||
# Restart affected services
|
||||
docker-compose up -d
|
||||
```
|
||||
|
||||
### 13.5 Full System Recovery
|
||||
In case of catastrophic failure:
|
||||
1. **Provision new VPS** (or reset existing)
|
||||
2. **Install Docker and Docker Compose**
|
||||
3. **Restore `/opt/apex/` from offsite backup**
|
||||
4. **Restore docker-compose.yml and .env**
|
||||
5. **Pull images:** `docker-compose pull`
|
||||
6. **Start services:** `docker-compose up -d`
|
||||
7. **Restore PostgreSQL:** `pg_restore` from backup
|
||||
8. **Verify all services:** Check health endpoints
|
||||
9. **Verify DNS:** Ensure `*.apex.unstuck-path.com` resolves
|
||||
10. **Notify Human CEO** via alternative channel
|
||||
|
||||
---
|
||||
|
||||
## 14. Infrastructure Change Process
|
||||
|
||||
Every infrastructure change follows this process:
|
||||
|
||||
### 14.1 Pre-Change
|
||||
1. **Create backup:**
|
||||
```bash
|
||||
cp docker-compose.yml docker-compose.yml.backup.$(date +%Y%m%d%H%M%S)
|
||||
```
|
||||
2. **Document the change** in `engineer_decisions`:
|
||||
- What is being changed
|
||||
- Why it's being changed
|
||||
- What the expected outcome is
|
||||
- What the rollback plan is
|
||||
3. **Test in isolation** when possible (sandbox Docker network)
|
||||
4. **Get approval** for production changes (Constitution Law 6)
|
||||
|
||||
### 14.2 During Change
|
||||
5. **Execute the change** following documented steps
|
||||
6. **Monitor actively** — watch logs and metrics in real-time
|
||||
7. **Verify health checks** pass for all affected services
|
||||
|
||||
### 14.3 Post-Change
|
||||
8. **Monitor for 15 minutes** after deployment
|
||||
9. **Run smoke tests** — verify critical paths work
|
||||
10. **Update documentation** — TOOL_REGISTRY.md, README, etc.
|
||||
11. **Log outcome** in `engineer_decisions`
|
||||
12. **Commit changes** to Gitea
|
||||
|
||||
### 14.4 If Change Fails
|
||||
13. **Execute rollback plan** immediately
|
||||
14. **Log failure** with root cause analysis
|
||||
15. **Create reflection** — What went wrong? How to prevent next time?
|
||||
16. **Escalate** if rollback fails
|
||||
|
||||
---
|
||||
|
||||
## 15. Network Architecture
|
||||
|
||||
### 15.1 Docker Network
|
||||
```
|
||||
Network: apex_apex-net (bridge)
|
||||
|
||||
┌─────────────────────────────────────────────┐
|
||||
│ apex_apex-net │
|
||||
│ │
|
||||
│ All apex-* containers are members │
|
||||
│ Internal DNS resolves container names │
|
||||
│ No external access except via Traefik │
|
||||
└─────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
### 15.2 Port Mapping
|
||||
|
||||
| External Port | Service | Notes |
|
||||
|--------------|---------|-------|
|
||||
| 80 | Traefik | HTTP → HTTPS redirect |
|
||||
| 443 | Traefik | HTTPS termination + routing |
|
||||
| (none others) | — | All other ports internal only |
|
||||
|
||||
### 15.3 Internal Service Ports
|
||||
|
||||
| Service | Internal Port | Access Via |
|
||||
|---------|--------------|------------|
|
||||
| PostgreSQL | 5432 | `apex-postgres:5432` |
|
||||
| Redis | 6379 | `apex-redis:6379` |
|
||||
| LiteLLM | 4000 | `apex-litellm:4000` |
|
||||
| Ollama | 11434 | `apex-ollama:11434` |
|
||||
| Letta | 8283 | `apex-letta:8283` |
|
||||
| Gitea | 3000 | `git.apex.unstuck-path.com` |
|
||||
| Grafana | 3000 | `grafana.apex.unstuck-path.com` |
|
||||
| n8n | 5678 | `n8n.apex.unstuck-path.com` |
|
||||
| Langfuse | 3000 | `langfuse.apex.unstuck-path.com` |
|
||||
| Open WebUI | 8080 | `openwebui.apex.unstuck-path.com` |
|
||||
| Vaultwarden | 80 | `vaultwarden.apex.unstuck-path.com` |
|
||||
| Code-Server | 8443 | `code-server.apex.unstuck-path.com` |
|
||||
| Dockge | 5001 | `dockge.apex.unstuck-path.com` |
|
||||
| Socket Proxy | 2375 | `apex-socket-proxy:2375` |
|
||||
| Status API | 3100 | Internal only |
|
||||
| Prometheus | 9090 | Internal only |
|
||||
| Loki | 3100 | Internal only |
|
||||
|
||||
### 15.4 DNS Structure
|
||||
All services accessible via subdomains of `apex.unstuck-path.com`:
|
||||
- Wildcard DNS: `*.apex.unstuck-path.com` → `62.72.3.145`
|
||||
- Traefik matches `Host()` rules in container labels
|
||||
- SSL via Let's Encrypt (ACME HTTP-01 challenge)
|
||||
|
||||
---
|
||||
|
||||
## 16. Performance Standards
|
||||
|
||||
### 16.1 Response Time Targets
|
||||
|
||||
| Category | Target | Measurement |
|
||||
|----------|--------|-------------|
|
||||
| Health check endpoints | < 500ms | Prometheus histogram |
|
||||
| API endpoints | < 2s (p95) | Langfuse / Prometheus |
|
||||
| LLM responses (local) | < 10s | Langfuse |
|
||||
| LLM responses (cloud) | < 30s | Langfuse |
|
||||
| Page load (web UIs) | < 3s | Manual verification |
|
||||
|
||||
### 16.2 Availability Targets
|
||||
| Service | Target | Monitoring |
|
||||
|---------|--------|------------|
|
||||
| Core infrastructure (Traefik, PostgreSQL, Redis) | 99.5% uptime | Prometheus + auto-recovery |
|
||||
| AI services (LiteLLM, Ollama, Letta) | 99% uptime | Prometheus + auto-recovery |
|
||||
| Supporting services (Grafana, Gitea, n8n) | 98% uptime | Prometheus |
|
||||
|
||||
### 16.3 Resource Limits
|
||||
| Metric | Warning Threshold | Critical Threshold |
|
||||
|--------|-------------------|-------------------|
|
||||
| CPU usage (total) | 70% | 90% |
|
||||
| Memory usage (total) | 75% | 90% |
|
||||
| Disk usage | 80% | 90% |
|
||||
| Container restart count | 3/hour | 5/hour |
|
||||
|
||||
---
|
||||
|
||||
## 17. Security Engineering
|
||||
|
||||
### 17.1 Principle of Least Privilege
|
||||
Every container, service, and agent has the **minimum permissions** required for its function:
|
||||
- Docker Socket Proxy restricts Docker API to read-only queries
|
||||
- Code-Server has access to config/data only (not full `/opt/apex/`)
|
||||
- n8n has no Docker socket access
|
||||
- Employees have role-specific database access
|
||||
|
||||
### 17.2 Credential Lifecycle
|
||||
```
|
||||
Generate → Store in Vaultwarden → Reference in .env → Pass as env var → Rotate every 60 days
|
||||
```
|
||||
|
||||
### 17.3 Security Checklist for New Services
|
||||
- [ ] No default credentials
|
||||
- [ ] No privileged container mode
|
||||
- [ ] Non-root user where possible
|
||||
- [ ] Read-only mounts where possible
|
||||
- [ ] No direct Docker socket access (use proxy)
|
||||
- [ ] Health check configured
|
||||
- [ ] Traefik TLS termination (no plain HTTP)
|
||||
- [ ] Credentials in Vaultwarden
|
||||
- [ ] CVE check on image and dependencies
|
||||
|
||||
---
|
||||
|
||||
## 18. Change History
|
||||
|
||||
| Date | Version | Author | Changes |
|
||||
|------|---------|--------|---------|
|
||||
| Phase 2 | 0.1 | Engineer (#1) | Initial standards — Docker basics, database schema |
|
||||
| Phase 3 | 0.5 | Engineer (#1) | Added Git workflow, repository standards |
|
||||
| Phase 5 | 1.0 | Engineer (#1) | Added API standards, testing requirements, error handling |
|
||||
| Phase 5.5 | 1.5 | Engineer (#1) | Added monitoring standards, auto-recovery, logging pipeline |
|
||||
| Phase 7 | 2.0 | Engineer (#1) | Full standards formalization. Added network architecture, performance standards, security engineering, rollback procedures, infrastructure change process. Comprehensive coverage of all engineering aspects. |
|
||||
|
||||
---
|
||||
|
||||
> **These standards are not suggestions — they are requirements. Every technical deliverable in APEX OS must comply. When in doubt, err on the side of more documentation, more testing, and more caution.**
|
||||
|
||||
*Cross-references: [APEX_CONSTITUTION.md](APEX_CONSTITUTION.md) · [TOOL_REGISTRY.md](TOOL_REGISTRY.md) · [EMPLOYEE_HANDBOOK.md](EMPLOYEE_HANDBOOK.md) · [COMPANY_STRUCTURE.md](COMPANY_STRUCTURE.md)*
|
||||
Reference in New Issue
Block a user