Incident Center
Evidence source: REAL MOCK STATIC UNKNOWN· REAL = live read-only evidence · MOCK = demo, never verified · STATIC = documented/committed (not a live probe) · UNKNOWN = no evidence. Only PASS counts as verified.
READ-ONLY / SIMULATION. All incidents below are clearly-labelled MOCK scenarios. There is no incident mutation API and no action button executes a command.
A. Current incidents
INC-DEMO-001 · MOCK
PostgreSQL unavailable
CRITICAL INVESTIGATING
Demo scenario: DCP /health reports the database unavailable; private-network checks in progress. MOCK only.
service PostgreSQL · opened 2026-10-05T02:14:00+07:00 · last update 2026-10-05T02:31:00+07:00
Recovery: PostgreSQL unhealthy
Affected services
| Service | Impact | Detail | Source |
|---|
| DCP ownership | SERVICE_DOWN | Primary datastore unavailable. | STATIC |
| n8n ownership | SERVICE_DOWN | Automation datastore unavailable. | STATIC |
| AI Company Portal | DEGRADED | DCP read-only API unreachable. | STATIC |
Timeline
monitor · DETECTED · DCP /health reports database unavailable. MOCK
2026-10-05T02:14:00+07:00
devops · INVESTIGATING · Checking PostgreSQL service health on the private network. MOCK
2026-10-05T02:18:00+07:00
devops · NOTE · Distinguishing connectivity fault from data fault. MOCK
2026-10-05T02:31:00+07:00
INC-DEMO-002 · MOCK
Backup older than threshold
MAJOR MONITORING
Demo scenario: backup age exceeds the expected window. Recovery capability degraded, production unaffected. MOCK only.
service Backup · opened 2026-10-05T01:00:00+07:00 · last update 2026-10-05T03:05:00+07:00
Recovery: Backup failed
Affected services
| Service | Impact | Detail | Source |
|---|
| Home Server | RECOVERY_DEGRADED | Off-disk copies are stale. | STATIC |
Timeline
monitor · DETECTED · backup_age_seconds exceeded threshold. MOCK
2026-10-05T01:00:00+07:00
devops · MITIGATING · Verifying destination mount and free space. MOCK
2026-10-05T01:20:00+07:00
devops · MONITORING · Awaiting admin-run retry; old backups retained. MOCK
2026-10-05T03:05:00+07:00
INC-DEMO-003 · MOCK
DCP health degraded
WARNING MITIGATING
Demo scenario: DCP /health degraded while the database is healthy; application rollback being prepared. MOCK only.
service DCP · opened 2026-10-05T04:02:00+07:00 · last update 2026-10-05T04:12:00+07:00
Recovery: DCP unhealthy, PostgreSQL healthy
Affected services
| Service | Impact | Detail | Source |
|---|
| DCP ownership | DEGRADED | App degraded; DB component healthy. | STATIC |
| AI Company Portal | DEGRADED | Partial DCP data. | STATIC |
Timeline
monitor · DETECTED · DCP /health reports degraded. MOCK
2026-10-05T04:02:00+07:00
devops · INVESTIGATING · Checking app logs; DB component healthy. MOCK
2026-10-05T04:07:00+07:00
devops · MITIGATING · Preparing previous-image rollback (admin-run). MOCK
2026-10-05T04:12:00+07:00
INC-DEMO-004 · MOCK
Tailscale connectivity unavailable
MAJOR OPEN
Demo scenario: remote private access to DCP is unavailable; local services believed healthy. MOCK only.
service Tailscale · opened 2026-10-05T05:40:00+07:00 · last update 2026-10-05T05:44:00+07:00
Recovery: Tailscale unavailable
Affected services
| Service | Impact | Detail | Source |
|---|
| DCP ownership | CONNECTIVITY | Remote reachability only; local health separate. | STATIC |
| Home Server | CONNECTIVITY | Private mesh path affected. | STATIC |
Timeline
monitor · DETECTED · Remote access to DCP dashboard failing. MOCK
2026-10-05T05:40:00+07:00
devops · INVESTIGATING · Distinguishing local service health from remote connectivity. MOCK
2026-10-05T05:44:00+07:00
B. Recent incidents
INC-DEMO-005 · MOCK
Transient n8n notification delay
INFO RESOLVED
Demo scenario: a transient notification delay self-cleared; no data impact. MOCK only.
service n8n · opened 2026-10-04T21:10:00+07:00 · last update 2026-10-04T21:26:00+07:00
Recovery: Backup failed
Affected services
| Service | Impact | Detail | Source |
|---|
| n8n ownership | DEGRADED | Notification delivery delayed. | STATIC |
Timeline
monitor · DETECTED · Notification delivery latency elevated. MOCK
2026-10-04T21:10:00+07:00
devops · MONITORING · Latency trending down; no action taken. MOCK
2026-10-04T21:20:00+07:00
devops · RESOLVED · Self-cleared; no data impact. MOCK
2026-10-04T21:26:00+07:00
D. Incident timeline
MAJOR OPEN Tailscale connectivity unavailable MOCK
INC-DEMO-004 · Tailscale · 2026-10-05T05:44:00+07:00
WARNING MITIGATING DCP health degraded MOCK
INC-DEMO-003 · DCP · 2026-10-05T04:12:00+07:00
MAJOR MONITORING Backup older than threshold MOCK
INC-DEMO-002 · Backup · 2026-10-05T03:05:00+07:00
CRITICAL INVESTIGATING PostgreSQL unavailable MOCK
INC-DEMO-001 · PostgreSQL · 2026-10-05T02:31:00+07:00
INFO RESOLVED Transient n8n notification delay MOCK
INC-DEMO-005 · n8n · 2026-10-04T21:26:00+07:00
E. Recovery options
Read-only recovery guidance per scenario. Use the wizard below.
DCP unhealthy, PostgreSQL healthy — The control-plane app is failing its healthcheck while the datastore looks fine. Prefer an application rollback over touching the database.
PostgreSQL unhealthy — The datastore is failing. Establish whether it is connectivity, resource pressure, or actual data corruption before considering a restore.
Backup failed — The backup job did not complete. Recovery safety degrades, but active production continues. Investigate before any cleanup.
Seagate backup disk unavailable — The external backup destination is offline. Active production keeps running; only backup/recovery capability is degraded.
Tailscale unavailable — Remote private connectivity is affected. Local service health is separate: distinguish a local fault from a remote-connectivity fault before acting.
Bad application deployment — A new image/version regressed the app. Roll back the image first; only touch the DB if the data state is proven damaged.
Recovery Wizard (read-only)
DCP unhealthy, PostgreSQL healthy
DCPSTATIC
The control-plane app is failing its healthcheck while the datastore looks fine. Prefer an application rollback over touching the database.
Order: SYMPTOM (GEJALA) → CHECK (PEMERIKSAAN) → DIAGNOSIS → SAFEST ACTION (SOLUSI) → VERIFY (VERIFIKASI) → ROLLBACK.
SYMPTOMGEJALA
DCP reports unhealthy
DCP /health is not OK (or the dashboard is unreachable) while the database component still reports healthy.
CHECKPEMERIKSAAN
Check health + logs
Read the compose healthcheck output and the DCP container logs. Confirm the database component inside /health is healthy.
Runbook: DCP rollback (path only — not browser-accessible)docs/PRODUCTION-DEPLOYMENT.md (DCP repo)
DIAGNOSISDIAGNOSIS
Separate app vs data
If only the app is unhealthy and the DB is healthy, this is an application/configuration fault — not a data-loss fault.
SAFEST ACTIONSOLUSI
Prefer application rollback first
Revert DCP_IMAGE_TAG to the previous immutable image and recreate the DCP service (admin-run). Do NOT restore the database by default.
⚠ STOP: do not restore the DB unless data corruption is proven. A bad app deploy is fixed by rolling back the app.
Runbook: DCP rollback (path only — not browser-accessible)docs/PRODUCTION-DEPLOYMENT.md (DCP repo)
VERIFYVERIFIKASI
Verify /health + database
Re-check DCP /health (status + database) and confirm the dashboard loads before declaring recovery.
ROLLBACKROLLBACK
If rollback fails
Escalate to a manual, admin-approved recovery using the restore order. Production restore only after a disposable drill passes.
Runbook: Restore order (path only — not browser-accessible)homelab-infra/docs/RESTORE-ORDER.md
Recovery options
- Check health + logs
Read the healthcheck output and DCP logs (admin-run).
- Roll back DCP image
Revert DCP_IMAGE_TAG to the previous immutable image and recreate the service (admin-run).
- Restore database REQUIRES ADMIN APPROVALDESTRUCTIVE — NOT DEFAULT
Only if data corruption is proven, after explicit admin approval and a disposable drill.
NEVER DO- Do not restore the DB for an app-only fault.
- Do not use `docker system prune`.
Runbook: DCP rollback (path only — not browser-accessible)docs/PRODUCTION-DEPLOYMENT.md (DCP repo)
Runbook: Restore order (path only — not browser-accessible)homelab-infra/docs/RESTORE-ORDER.md
Read-only guidance. No action button executes a command.
PostgreSQL unhealthy
PostgreSQLSTATIC
The datastore is failing. Establish whether it is connectivity, resource pressure, or actual data corruption before considering a restore.
Order: SYMPTOM (GEJALA) → CHECK (PEMERIKSAAN) → DIAGNOSIS → SAFEST ACTION (SOLUSI) → VERIFY (VERIFIKASI) → ROLLBACK.
SYMPTOMGEJALA
DCP/DATABASE reports unavailable
DCP /health reports the database unavailable, or dependent services (n8n) lose their datastore.
CHECKPEMERIKSAAN
Verify container/service health
Check the PostgreSQL service/container health and logs (admin-run). Confirm the private `homelab-backend` network and that the data volume is mounted.
Runbook: Host baseline (path only — not browser-accessible)homelab-infra/host/BASELINE.md
DIAGNOSISDIAGNOSIS
Inspect DB connectivity
Attempt a read-only connection from a disposable client on the private network. Rule out a network/DSN problem before suspecting the data.
SAFEST ACTIONSOLUSI
Restore only if corruption is proven
Use backup restore ONLY when corruption/data loss is demonstrated. Restore to a disposable target first.
⚠ STOP: never restore over the production database without explicit admin approval and a recorded decision.
Runbook: Automated restore drill (path only — not browser-accessible)homelab-infra/docs/AUTOMATED-RESTORE-DRILL.md
VERIFYVERIFIKASI
Validate restored data
Compare row counts, ids, relations and provenance against the source backup (see the drill validation queries).
Runbook: Automated restore drill (path only — not browser-accessible)homelab-infra/docs/AUTOMATED-RESTORE-DRILL.md
ROLLBACKROLLBACK
Fall back to previous backup
If the restore is unverified or fails, revert to the previous verified backup; do not stack unverified attempts.
Runbook: Restore order (path only — not browser-accessible)homelab-infra/docs/RESTORE-ORDER.md
Recovery options
- Verify container/service health
Check the PostgreSQL service/container health and logs (admin-run).
- Inspect DB connectivity
Attempt a read-only connection from a disposable client on the private network.
- Restore from backup REQUIRES ADMIN APPROVALDESTRUCTIVE — NOT DEFAULT
Only if corruption/data loss is proven; restore to a disposable target first, with explicit admin approval.
NEVER DO- Do not restore over production without approval.
- Do not delete the previous backup while recovering.
Runbook: Restore order (path only — not browser-accessible)homelab-infra/docs/RESTORE-ORDER.md
Runbook: Automated restore drill (path only — not browser-accessible)homelab-infra/docs/AUTOMATED-RESTORE-DRILL.md
Read-only guidance. No action button executes a command.
Backup failed
BackupSTATIC
The backup job did not complete. Recovery safety degrades, but active production continues. Investigate before any cleanup.
Order: SYMPTOM (GEJALA) → CHECK (PEMERIKSAAN) → DIAGNOSIS → SAFEST ACTION (SOLUSI) → VERIFY (VERIFIKASI) → ROLLBACK.
SYMPTOMGEJALA
Backup status is FAILED / stale
The backup-status contract reports status=FAILED or the backup age exceeds the expected window.
Runbook: Backup-status schema (path only — not browser-accessible)docs/BACKUP-STATUS-SCHEMA.md
CHECKPEMERIKSAAN
Verify mount + disk identity + free space
Confirm the destination is mounted and is the expected device, and that free space is not exhausted.
Runbook: Backup architecture (path only — not browser-accessible)homelab-infra/backup/ARCHITECTURE.md
DIAGNOSISDIAGNOSIS
Inspect the latest backup + checksum
Read the latest run result and checksum status; distinguish a space/permission fault from an integrity fault.
Runbook: Backup-status schema (path only — not browser-accessible)docs/BACKUP-STATUS-SCHEMA.md
SAFEST ACTIONSOLUSI
Fix the cause, re-run (admin)
Resolve the cause and let the admin-run job retry. Do NOT automatically delete old backups to free space.
⚠ STOP: never delete old backups automatically — a failed new backup makes old backups the only recovery point.
Runbook: systemd units (path only — not browser-accessible)homelab-infra/systemd/README.md
VERIFYVERIFIKASI
Confirm a fresh verified backup
Confirm a new backup completes with checksum PASS and the backup age drops back inside the window.
Runbook: Backup-status schema (path only — not browser-accessible)docs/BACKUP-STATUS-SCHEMA.md
ROLLBACKROLLBACK
If the destination is unhealthy
Fall back to the alternative destination if available; keep the existing backups intact until a fresh one is verified.
Runbook: Backup architecture (path only — not browser-accessible)homelab-infra/backup/ARCHITECTURE.md
Recovery options
- Verify mount + disk identity + free space
Confirm the destination is mounted and is the expected device, with free space available.
- Inspect latest backup + checksum
Read the latest run result and checksum status.
- Fix cause and let the job retry (admin)
Resolve the cause and allow an admin-run retry.
NEVER DO- Do not delete old backups automatically.
- Do not mark backup PASS from a design or a mock fixture.
Runbook: Backup architecture (path only — not browser-accessible)homelab-infra/backup/ARCHITECTURE.md
Runbook: Backup-status schema (path only — not browser-accessible)docs/BACKUP-STATUS-SCHEMA.md
Read-only guidance. No action button executes a command.
Seagate backup disk unavailable
Seagate Backup DiskSTATIC
The external backup destination is offline. Active production keeps running; only backup/recovery capability is degraded.
Order: SYMPTOM (GEJALA) → CHECK (PEMERIKSAAN) → DIAGNOSIS → SAFEST ACTION (SOLUSI) → VERIFY (VERIFIKASI) → ROLLBACK.
SYMPTOMGEJALA
Backup destination not mounted
The Seagate drive is not mounted; backups cannot write to the external destination.
Runbook: Host baseline (path only — not browser-accessible)homelab-infra/host/BASELINE.md
CHECKPEMERIKSAAN
Confirm production is unaffected
Confirm all active services still run from `/srv/data` and are serving traffic; the external disk is a backup destination.
Runbook: Host baseline (path only — not browser-accessible)homelab-infra/host/BASELINE.md
DIAGNOSISDIAGNOSIS
Establish degradation scope
Recovery SAFETY is degraded (no fresh off-disk copies). Active production is NOT down.
Runbook: Backup architecture (path only — not browser-accessible)homelab-infra/backup/ARCHITECTURE.md
SAFEST ACTIONSOLUSI
Restore the destination (admin)
Reconnect/mount the external drive (admin-run). Do NOT move active application data onto the backup disk.
⚠ STOP: never move active app data onto the backup disk; it is a backup destination, not primary storage.
Runbook: Host baseline (path only — not browser-accessible)homelab-infra/host/BASELINE.md
VERIFYVERIFIKASI
Verify mount + fresh backup
Confirm the device identity and that a new backup succeeds to the restored destination.
Runbook: Backup-status schema (path only — not browser-accessible)docs/BACKUP-STATUS-SCHEMA.md
ROLLBACKROLLBACK
No rollback needed for production
No production rollback is required; production never stopped. Escalate only if a restore is needed while the backup is cold.
Runbook: Restore order (path only — not browser-accessible)homelab-infra/docs/RESTORE-ORDER.md
Recovery options
NEVER DO- Do not move active app data onto the backup disk.
- Do not treat degraded recovery as a production outage.
Runbook: Host baseline (path only — not browser-accessible)homelab-infra/host/BASELINE.md
Runbook: Backup architecture (path only — not browser-accessible)homelab-infra/backup/ARCHITECTURE.md
Read-only guidance. No action button executes a command.
Tailscale unavailable
TailscaleSTATIC
Remote private connectivity is affected. Local service health is separate: distinguish a local fault from a remote-connectivity fault before acting.
Order: SYMPTOM (GEJALA) → CHECK (PEMERIKSAAN) → DIAGNOSIS → SAFEST ACTION (SOLUSI) → VERIFY (VERIFIKASI) → ROLLBACK.
SYMPTOMGEJALA
Remote access to DCP fails
Remote users cannot reach the DCP dashboard over the private mesh (or the VPS->home path is broken).
Runbook: Network architecture (path only — not browser-accessible)homelab-infra/network/ARCHITECTURE.md
CHECKPEMERIKSAAN
Distinguish local vs remote
Check local service health directly on the home host. If local services are healthy, this is a connectivity fault, not a service outage.
Runbook: Network architecture (path only — not browser-accessible)homelab-infra/network/ARCHITECTURE.md
DIAGNOSISDIAGNOSIS
Scope the connectivity fault
Confirm whether it is the mesh, the VPS relay, or a single client. Do not change ACLs/firewall outside an approved window.
Runbook: Network architecture (path only — not browser-accessible)homelab-infra/network/ARCHITECTURE.md
SAFEST ACTIONSOLUSI
Restore connectivity (admin)
Reconnect/re-authenticate the mesh node (admin-run). Treat ACL/firewall changes as admin-only and reversible.
⚠ STOP: do not make unreviewed Tailscale ACL or firewall changes; that can widen access.
Runbook: Network architecture (path only — not browser-accessible)homelab-infra/network/ARCHITECTURE.md
VERIFYVERIFIKASI
Verify private reachability
Confirm the private path and DCP dashboard are reachable again; confirm no public expose was introduced.
Runbook: Network architecture (path only — not browser-accessible)homelab-infra/network/ARCHITECTURE.md
ROLLBACKROLLBACK
Revert connectivity changes
Revert any change and fall back to the documented fallback path. Keep changes reversible.
Runbook: Network architecture (path only — not browser-accessible)homelab-infra/network/ARCHITECTURE.md
Recovery options
NEVER DO- Do not make unreviewed ACL/firewall changes.
- Do not expose services publicly to work around the mesh.
Runbook: Network architecture (path only — not browser-accessible)homelab-infra/network/ARCHITECTURE.md
Read-only guidance. No action button executes a command.
Bad application deployment
DCPSTATIC
A new image/version regressed the app. Roll back the image first; only touch the DB if the data state is proven damaged.
Order: SYMPTOM (GEJALA) → CHECK (PEMERIKSAAN) → DIAGNOSIS → SAFEST ACTION (SOLUSI) → VERIFY (VERIFIKASI) → ROLLBACK.
SYMPTOMGEJALA
Regression after a deploy
Health degrades or errors appear immediately after a deployment.
Runbook: DCP rollback (path only — not browser-accessible)docs/PRODUCTION-DEPLOYMENT.md (DCP repo)
CHECKPEMERIKSAAN
Confirm timing + correlation
Correlate the regression with the deploy timestamp and image tag; check /health and logs.
Runbook: DCP rollback (path only — not browser-accessible)docs/PRODUCTION-DEPLOYMENT.md (DCP repo)
DIAGNOSISDIAGNOSIS
App fault vs data fault
If the DB is healthy and the fault is app-level, this is a bad deployment — not a data-loss event.
SAFEST ACTIONSOLUSI
Roll back the previous image first
Redeploy the previous immutable image tag and recreate the service (admin-run).
⚠ STOP: restore the DB only if the DB state is proven damaged. A bad deploy is fixed by an image rollback.
Runbook: DCP rollback (path only — not browser-accessible)docs/PRODUCTION-DEPLOYMENT.md (DCP repo)
VERIFYVERIFIKASI
Verify health after rollback
Confirm /health (status + database) is OK and the dashboard works before re-attempting any change.
ROLLBACKROLLBACK
If the image rollback fails
Escalate to a manual, admin-approved recovery; use the restore order and a disposable drill for data.
Runbook: Restore order (path only — not browser-accessible)homelab-infra/docs/RESTORE-ORDER.md
Recovery options
- Confirm timing + correlation
Correlate the regression with the deploy timestamp and image tag.
- Roll back the previous image
Redeploy the previous immutable image tag and recreate the service (admin-run).
- Restore database REQUIRES ADMIN APPROVALDESTRUCTIVE — NOT DEFAULT
Only if the DB state is proven damaged, after approval and a disposable drill.
NEVER DO- Do not restore the DB for an app-only regression.
- Do not redeploy an unpinned `:latest` image.
Runbook: DCP rollback (path only — not browser-accessible)docs/PRODUCTION-DEPLOYMENT.md (DCP repo)
Runbook: Restore order (path only — not browser-accessible)homelab-infra/docs/RESTORE-ORDER.md
Read-only guidance. No action button executes a command.