8.4 KiB
8.4 KiB
APEX OS — Platform Evolution Guide
The platform grows through controlled, documented, human-approved evolution cycles.
Table of Contents
- Philosophy
- The Evolution Loop
- Capability Detection
- Research Phase
- Evaluation Framework
- Sandbox Protocol
- Testing Standards
- Benchmarking
- Documentation Requirements
- Approval Process
- Production Deployment
- Monitoring & Reflection
- Retirement Process
- Self-Improvement Rules
- Change History
Philosophy
APEX OS is a living platform. It must evolve to meet new requirements, adopt better tools, and improve its own capabilities. However, evolution must be:
- Controlled — No unreviewed changes
- Documented — Every change has a paper trail
- Reversible — Every deployment has a rollback plan
- Approved — Human CEO has final say on production changes
- Measured — Before and after benchmarks required
The Evolution Loop
1. DETECT → Identify missing capability or improvement opportunity
↓
2. RESEARCH → Find 3+ alternatives, compare features/costs/risks
↓
3. EVALUATE → Score against evaluation criteria
↓
4. SANDBOX → Deploy in isolated Docker network for testing
↓
5. TEST → Run functional, integration, and security tests
↓
6. BENCHMARK → Measure performance, resource usage, cost
↓
7. DOCUMENT → Write findings, recommendations, migration plan
↓
8. RECOMMEND → CEO Agent presents recommendation to Human CEO
↓
9. APPROVE → Human CEO approves/rejects via Telegram
↓
10. DEPLOY → Production deployment following Engineering Standards
↓
11. MONITOR → Watch for 24 hours post-deployment
↓
12. REFLECT → Log lessons learned, update knowledge base
↓
13. LEARN → Update shared knowledge for future decisions
Capability Detection
Capabilities can be detected through:
Automatic Detection
- Task failures logged in reflections with
improvement_applied = false - Repeated similar errors in recovery_log
- Performance metrics degradation
- Employee requests for tools they don't have
Manual Detection
- Human CEO request ("Find a better X")
- Engineer identifies technical debt
- Security audit findings
- Cost optimization opportunities
Detection Triggers
| Trigger | Action |
|---|---|
| 3+ similar task failures | Auto-generate research task |
| Employee reports missing tool | Log to engineer_decisions, research |
| CEO request | Immediate research assignment |
| Security vulnerability | Emergency research + sandbox |
| Cost threshold exceeded | Cost optimization research |
Research Phase
Requirements
- Minimum 3 alternatives evaluated
- Open source options prioritized
- Docker compatibility verified
- Resource requirements documented
- Security track record checked
- Community activity verified (commits in last 6 months)
- License compatibility confirmed
Output
Research report saved to:
- Researcher's Gitea workspace
- mem0.shared_knowledge (for future reference)
- apex.engineer_decisions (decision context)
Evaluation Framework
Scoring Criteria (1-10 each)
| Criterion | Weight | Description |
|---|---|---|
| Functionality | 25% | Does it solve the problem completely? |
| Open Source | 15% | Is it fully open source? |
| Docker Ready | 15% | Can it run as a Docker container? |
| Resource Efficiency | 10% | CPU, memory, storage requirements |
| Security | 15% | Track record, isolation capability |
| Maintenance | 10% | Active development, documentation |
| Integration | 10% | Works with existing APEX stack |
Minimum Score: 7.0/10 for production consideration
Sandbox Protocol
Setup
# Create isolated network
docker network create apex-sandbox
# Deploy candidate tool
docker run -d \
--name sandbox-{tool-name} \
--network apex-sandbox \
--memory 512m \
--cpus 1.0 \
{tool-image}:{version}
Rules
- No connection to production network (apex_apex-net)
- No access to production databases
- No access to Docker socket
- Maximum 512MB memory, 1 CPU core
- 48-hour maximum sandbox duration
- Automatic cleanup after evaluation
Testing Standards
Functional Tests
- Core features work as documented
- API endpoints respond correctly
- Error handling is graceful
- Configuration options work
Integration Tests
- Connects to PostgreSQL (sandbox instance)
- Works with Traefik reverse proxy
- Compatible with existing monitoring (Prometheus/Grafana)
- Logs are captured by Loki/Promtail
Security Tests
- No privileged mode required
- No host network access needed
- Credentials properly isolated
- No known CVEs in current version
- Network requests are expected and documented
Performance Tests
- Response time under load
- Memory usage over time (check for leaks)
- CPU usage patterns
- Disk I/O patterns
- Startup time
Benchmarking
Before/After Comparison
| Metric | Current Tool | Candidate | Improvement |
|---|---|---|---|
| Response Time | X ms | Y ms | Z% |
| Memory Usage | X MB | Y MB | Z% |
| CPU Usage | X% | Y% | Z% |
| Features | List | List | Delta |
| Cost | $X | $Y | Savings |
Minimum Improvement Threshold
- Performance: 20% improvement to justify migration
- Cost: 10% savings to justify migration
- Features: Must not lose existing capabilities
Documentation Requirements
Pre-Deployment Documentation
- Research report with alternatives comparison
- Evaluation scorecard
- Sandbox test results
- Benchmark comparison
- Migration plan
- Rollback plan
- Updated TOOL_REGISTRY.md entry
Post-Deployment Documentation
- Updated architecture diagram
- Configuration documentation
- Monitoring setup
- Troubleshooting guide
- Lessons learned
Approval Process
Recommendation Format (sent via Telegram)
🔧 Platform Evolution Recommendation
Tool: {name}
Purpose: {what it replaces/adds}
Score: {X.X}/10
Highlights:
- {key benefit 1}
- {key benefit 2}
- {key benefit 3}
Risks:
- {risk 1}
- {risk 2}
Cost Impact: {increase/decrease/neutral}
Resource Impact: {CPU/memory delta}
Recommendation: {ADOPT / TRIAL / REJECT}
[Approve] [Reject] [More Info]
Production Deployment
Follow Engineering Standards (ENGINEERING_STANDARDS.md):
- Create backup of current docker-compose.yml
- Update docker-compose.yml with new service
- Pull new image
- Deploy with
docker-compose up -d - Verify health checks pass
- Monitor for 15 minutes
- Update TOOL_REGISTRY.md
- Commit changes to Gitea
Monitoring & Reflection
24-Hour Watch Period
- Check logs every 4 hours
- Monitor resource usage
- Verify integration with existing services
- Watch for error spikes
Post-Deployment Reflection
Log to reflections table:
- What went well?
- What was unexpected?
- Would we make the same decision again?
- What can we improve in the evaluation process?
Retirement Process
When replacing a tool:
- Ensure data migration is complete
- Update all references in documentation
- Remove from docker-compose.yml
- Clean up Docker volumes
- Update TOOL_REGISTRY.md (status: Retired)
- Archive configuration for reference
- Commit all changes to Gitea
Self-Improvement Rules
What APEX Can Do Autonomously
- Research new tools and report findings
- Run sandbox evaluations
- Generate benchmark reports
- Suggest optimizations
What Requires Human Approval
- Installing new production tools
- Removing existing tools
- Changing infrastructure configuration
- Modifying the Constitution
- Upgrading core services (PostgreSQL, LiteLLM, Letta)
Absolute Prohibitions
- Never modify the Constitution without human approval
- Never disable security measures
- Never remove backup systems
- Never bypass the approval process
- Never self-modify core agent logic without approval
Change History
| Date | Version | Author | Changes |
|---|---|---|---|
| Phase 7.1 | 1.0 | Engineer | Initial platform evolution guide |