Files
apex-os-docs/SECURITY_POLICY.md
T

11 KiB

APEX OS — Security Policy

Security is not a feature. It is a fundamental property of the system.

Table of Contents

  1. Principles
  2. Secrets Management
  3. Authentication & Access
  4. Container Security
  5. Network Security
  6. Backup & Recovery
  7. Audit & Logging
  8. Incident Response
  9. Credential Rotation
  10. Human Approval Gates
  11. Recovery Procedures
  12. Compliance Checklist
  13. Change History

Principles

  1. Defense in Depth — Multiple layers of security, never rely on a single control
  2. Least Privilege — Every service gets minimum required access
  3. Secrets Never in Code — All credentials in Vaultwarden or environment variables
  4. Assume Breach — Design for containment and rapid recovery
  5. Audit Everything — Every access, change, and decision is logged
  6. Human in the Loop — Security-critical changes require human approval

Secrets Management

Vaultwarden (Primary Secrets Store)

Secret Categories

Category Storage Rotation
SSH Keys /root/.ssh/ (VPS) Annually
Database Passwords .env file (chmod 600) 90 days
API Keys (LiteLLM) .env + Vaultwarden 60 days
Gitea Tokens Vaultwarden 90 days
Grafana Admin Vaultwarden 90 days
n8n Credentials Vaultwarden 90 days
Telegram Bot Token Vaultwarden As needed
OpenRouter API Key .env + Vaultwarden 90 days

Rules

  • Never commit secrets to Git repositories
  • Never log secrets (mask in all output)
  • Never transmit secrets in plain text over the network
  • Never store secrets in container images
  • .env file must be chmod 600 (owner read/write only)

Authentication & Access

SSH Access

  • Method: Key-only authentication (password disabled)
  • Key Location: Engineer SSH key managed separately
  • Root Login: Allowed via key only
  • Port: Standard (22)
  • Failed Login: Monitored via auth.log

Service Authentication

Service Auth Method Notes
LiteLLM API Key (LITELLM_MASTER_KEY) Required for all API calls
Gitea Username/Password + API Token Token for automation
Grafana Username/Password Admin access restricted
n8n Username/Password Admin account only
Letta No auth (internal only) Not exposed externally
Ollama No auth (internal only) Not exposed externally
Traefik Dashboard BasicAuth Rate-limited
Dockge Username/Password Docker management
Vaultwarden Master Password Encrypted vault

External Access Points

Only these services are accessible from the internet (via Traefik):

  • openwebui.apex.unstuck-path.com
  • n8n.apex.unstuck-path.com
  • gitea.apex.unstuck-path.com (or git.apex.unstuck-path.com)
  • grafana.apex.unstuck-path.com
  • code-server.apex.unstuck-path.com
  • vaultwarden.apex.unstuck-path.com
  • status.apex.unstuck-path.com
  • dockge.apex.unstuck-path.com
  • langfuse.apex.unstuck-path.com

All external access is HTTPS via Let's Encrypt certificates (auto-renewed by Traefik).


Container Security

Docker Standards

  • No privileged containers — None run with --privileged
  • Read-only filesystems where possible
  • No host network mode — All containers on Docker bridge networks
  • Resource limits — Memory and CPU limits recommended
  • Health checks — Required on all containers
  • Restart policyunless-stopped for production services
  • Docker Socket — Mediated by Docker Socket Proxy (read-only for most services)

Docker Socket Proxy

  • Container: apex-socket-proxy
  • Purpose: Mediates Docker API access
  • Read-Only Access: Most services (monitoring, status API)
  • Read-Write Access: Dockge only (for container management)
  • Forbidden: No container can directly mount /var/run/docker.sock except Socket Proxy

Image Security

  • Use official images from Docker Hub
  • Pin image versions (avoid latest in production when possible)
  • Scan images for vulnerabilities periodically
  • Update images during controlled maintenance windows

Network Security

Docker Networks

  • apex_apex-net — Primary internal network for all services
  • All inter-service communication happens on this network
  • External access only through Traefik reverse proxy

Firewall Rules

  • Port 22 (SSH): Open
  • Port 80 (HTTP): Open (redirects to HTTPS)
  • Port 443 (HTTPS): Open
  • All other ports: Closed to external access
  • Internal Docker ports accessible only within apex_apex-net

TLS/SSL

  • All external traffic encrypted via HTTPS
  • Certificates managed by Traefik + Let's Encrypt
  • Auto-renewal configured
  • HSTS headers recommended

Backup & Recovery

Automated Backups

  • Schedule: Daily at 2:00 AM UTC
  • Script: /opt/apex/scripts/backup.sh
  • Retention: 7 days
  • Contents:
    • docker-compose.yml
    • .env file
    • All Docker volumes (PostgreSQL data, Gitea repos, Grafana dashboards)
    • Configuration files
    • SSL certificates

Backup Storage

  • Location: /opt/apex/backups/
  • Format: Compressed tar archives
  • Naming: apex-backup-YYYYMMDD-HHMMSS.tar.gz

Recovery Procedures

Single Service Recovery

# Stop the service
docker-compose stop {service}

# Restore volume from backup
tar -xzf /opt/apex/backups/{backup-file} -C /

# Restart the service
docker-compose up -d {service}

Full System Recovery

# Stop all services
docker-compose down

# Restore all volumes
tar -xzf /opt/apex/backups/{backup-file} -C /

# Restore configuration
cp backup/.env /opt/apex/.env
cp backup/docker-compose.yml /opt/apex/docker-compose.yml

# Start all services
docker-compose up -d

# Verify health
curl http://localhost:3100/health

Auto-Recovery

  • Script: /opt/apex/scripts/auto-recovery.sh
  • Schedule: Every 5 minutes via cron
  • Actions: Checks container health, restarts unhealthy/stopped containers
  • Logging: All actions logged to apex.recovery_log
  • Escalation: If container fails 3+ restarts, alert via Telegram

Audit & Logging

Decision Audit Trail

Every significant action is logged:

  • apex.engineer_decisions — All planning and deployment decisions
  • apex.reflections — Post-action analysis
  • apex.constitution_violations — Any rule violations (should be 0)
  • apex.lifecycle_executions — Project lifecycle tracking
  • apex.task_status_changes — Task state transitions

System Logs

  • Loki — Centralized log aggregation
  • Promtail — Log shipping from all containers
  • Grafana — Log visualization and search
  • Docker logs — Container-level logging

Monitoring

  • Prometheus — Metrics collection
  • Grafana — Dashboard visualization (12 panels)
  • APEX Status API — Real-time health checks for 6 critical services
  • Auto-Recovery — Automated container health monitoring

Incident Response

Severity Levels

Level Description Response Time Action
P0 - Critical Data loss, security breach, total outage Immediate Halt + Human CEO alert
P1 - High Service outage, failed recovery 5 minutes Auto-recovery + Telegram alert
P2 - Medium Performance degradation, single service issue 15 minutes Auto-recovery, CEO Agent notified
P3 - Low Minor issue, cosmetic, non-impacting Next business cycle Logged for review

Incident Response Steps

  1. Detect — Auto-recovery or monitoring catches the issue
  2. Contain — Isolate affected service if necessary
  3. Diagnose — Check logs, identify root cause
  4. Fix — Apply correction (backup first)
  5. Verify — Confirm fix resolves the issue
  6. Document — Log to engineer_decisions and reflections
  7. Learn — Update knowledge base to prevent recurrence

Emergency Contacts

  • Human CEO: Telegram Chat ID 1775182448
  • Emergency Bot: @JimmysalesBot sends critical alerts

Credential Rotation

Schedule

Credential Rotation Period Last Rotated Next Due
LiteLLM Master Key 60 days Phase 1 As scheduled
Gitea API Token 90 days Phase 1 As scheduled
Grafana Admin 90 days Phase 1 As scheduled
n8n Admin 90 days Phase 3 As scheduled
Dockge Admin 90 days Phase 3 As scheduled
Vaultwarden Admin 90 days Phase 1 As scheduled
SSH Keys Annual Phase 1 As scheduled

Rotation Procedure

  1. Generate new credential
  2. Update Vaultwarden entry
  3. Update .env file (if applicable)
  4. Restart affected services
  5. Verify services are operational
  6. Log rotation in engineer_decisions
  7. Update this table

Human Approval Gates

These actions ALWAYS require human CEO approval via Telegram:

Action Reason
Production deployment Irreversible system change
New tool installation Security and stability risk
External API registration Cost and security implications
Marketing spend Financial commitment
Employee hiring Workforce change
Infrastructure changes Stability risk
Credential rotation Access change
Data deletion Irreversible
Security policy changes Governance
Constitution amendments Foundational rules

Recovery Procedures

Docker Compose Recovery

# If docker-compose.yml is corrupted
cp /opt/apex/docker-compose.yml.backup.{latest} /opt/apex/docker-compose.yml
docker-compose up -d

Database Recovery

# Stop PostgreSQL
docker-compose stop postgres

# Restore from backup
docker exec apex-postgres psql -U apex -d apex < /opt/apex/backups/db-backup.sql

# Restart
docker-compose start postgres

Complete System Recovery

  1. Re-provision VPS (if needed)
  2. Install Docker and Docker Compose
  3. Restore /opt/apex/ from backup
  4. docker-compose up -d
  5. Verify all services via Status API
  6. Test Telegram bot connectivity

Compliance Checklist

Run this checklist monthly:

  • .env file permissions are 600
  • SSH password authentication is disabled
  • No containers running in privileged mode
  • Traefik dashboard has BasicAuth
  • Daily backups running successfully
  • Auto-recovery script active
  • No plaintext credentials in Git repositories
  • All external services behind HTTPS
  • Vaultwarden accessible and synced
  • Constitution violations count = 0
  • Docker images up to date (no critical CVEs)
  • Credential rotation on schedule

Change History

Date Version Author Changes
Phase 1 1.0 Engineer Initial hardening (SSH, .env, Traefik)
Phase 5.5 1.1 Engineer Added auto-recovery, security audit tool
Phase 7.1 2.0 Engineer Comprehensive security policy formalization