Operations & Runbooks
Disaster Recovery & Emergency Procedures
High-severity incident protocols, host outage migration, and data corruption recovery.
Disaster Recovery & Emergency Procedures
IMPLEMENTED
Protocols for catastrophic VPS hardware failure, database corruption, or security incidents.
1. Severity 1 Incident Protocol (Total Outage)
- Assess Impact: Check
/healthz/and/api/v1/health-check/. - Database Verification: Confirm PostgreSQL service health and disk volume integrity.
- Failover Execution: If the primary VPS is unrecoverable, provision a new Linux server in Dokploy and deploy from the latest verified Git tag and offsite SQL backup.
2. MikroTik Edge Router Outage
If a primary edge router fails:
- Promote secondary standby router in
/routers. - Run bulk subscriber re-sync task to provision PPPoE secrets on the backup router.
- Update edge switch VLAN gateways to route traffic through the standby BNG.