- url-resolver.js: add healthCheckUrl priority (bypasses SSO for health checks) - DC-PRODUCTION-GRADE-BACKLOG.md: v2 backlog with 19 tasks (DC-062–DC-080) based on full codebase audit: 1539 tests, 86.55% coverage, 0 ESLint errors v2 backlog replaces completed v1 (P0-1 through P2-7 all done). New priorities: P0: OpenAPI spec update, branch coverage gap, Dockerfile resource limits P1: Console sweep remainder, billing E2E test, graceful shutdown, lint sweep, health notification spam P2: CI/CD pipeline, Sentry, source maps, request logging, multi-stage Docker, health endpoint P3: WebSocket, i18n, config backup/restore, mobile, plugin system
14 KiB
14 KiB
DashCaddy Production-Grade Backlog (v2)
Generated 2026-08-12 from a full codebase audit. v1 items (P0-1 through P2-7) are ALL DONE. Current state: 1539 tests, 86.55% statement coverage, 0 ESLint errors, 173 warnings.
Current Health Snapshot
- Tests: 1539 passing across 63 suites
- Coverage: Statements 86.55% | Branches 72.14% (below 80% gate) | Functions 80.8% | Lines 90.67%
- ESLint: 0 errors, 173 warnings (all pre-existing)
- Remaining console. calls in src/:* 21 across 10 files
- Dockerfile: Runs as root (documented — needs Docker socket), no resource limits
- OpenAPI spec: Present but stale (says v1.0.0, actual is v1.15.0)
- Unhandled rejection/exception handlers: Present in server.js ✓
- Rate limiting: Present on auth + general routes ✓
- npm audit: 4 remaining vulns (semver-major transitive deps, deferred)
P0 — Must Fix (blocks public release)
DC-062: OpenAPI spec is stale — update to match actual v1.15.0 API surface
- status: pending
- details:
openapi.yamlsaysversion: 1.0.0and describes only a fraction of the API. Since DC-046/047 (auth providers), DC-053 (share), DC-055 (billing), DC-058 (share UI), and the tailscale-admin routes were added, the spec is significantly out of date. A stale spec is worse than no spec — it misleads API consumers and breaks any code generation from it. Fix: audit all route files (grep -rn 'router\.\(get\|post\|put\|delete\|patch\)' routes/), update openapi.yaml with every endpoint, bump version to 1.15.0, add it to the test suite (DC-017-style source-of-truth test that fails if a route exists but has no spec entry). Effort: ~3 hr. - impact: Public API trust. No paying customer can integrate against an undocumented API.
DC-063: Branch coverage at 72% — below the 80% gate
- status: pending
- details: Jest coverage report shows branches at 72.14% (303/420), failing the 80% threshold. The uncovered branches are concentrated in error-handling paths (catch blocks, fallback returns, edge-case conditionals). Fix: run
npx jest --coverage --coverageReporters=textto identify the files with the lowest branch coverage, then add targeted tests for the uncovered conditional paths. Priority files: backup-manager.js (multiple catch blocks), health-checker.js (timeout/retry branches), tailscale-coord.js (API error branches). Effort: ~2 hr. - impact: Error paths are where production incidents hide. Every untested catch block is a potential crash.
DC-064: Dockerfile runs as root with no resource limits
- status: pending
- details: The Dockerfile has no
USERdirective andstart.shhas no--memoryor--cpusflags. While root is needed for Docker socket access, the container can still OOM the host. Fix: (1) Add--memory=512m --memory-swap=1g --cpus=1.5to thedocker runin start.sh. (2) Create a non-root userdashcaddyfor the application process, and use a Docker socket proxy (liketecnativa/docker-socket-proxy) that exposes a limited subset of Docker API endpoints — the app only needs read access for monitoring + controlled container lifecycle. (3) Add--restart=unless-stoppedif not already present. Effort: ~2 hr. Risk: medium — socket proxy may break some Docker API calls, needs testing. - impact: Without limits, a memory leak in the API can take down the entire host. This is a production safety issue.
P1 — Code Quality & Reliability
DC-065: Remaining 21 console.* calls — sweep to structured logger
- status: pending
- details: After DC-060 (update-manager) and P1-3 through P1-8, 21 console calls remain across 10 files:
error-handler.js(2),email.js(1),dns-providers/registry.js(2),audit-logger.js(3),csrf-protection.js(3),config-drift-detector.js(1),auto-restart-manager.js(1),http.js(1),logging.js(6 intentional — the logger itself),routes/backups.js(1). The logging.js calls are fine (the logger IS console internally). The rest should route throughlog.info/warn/error. Some are fallbacks:ctx.logError || ((_c, err) => console.error(err))— these fire when ctx isn't available, which is exactly when structured logging matters most. Effort: ~45 min. - impact: Consistency. The logger write to error.log and supports structured JSON — console does not.
DC-066: No API integration test for the billing flow end-to-end
- status: pending
- details: DC-057 shipped contract tests and unit tests for the Stripe bridge, but there is no test that exercises the full flow: pricing page → Stripe Checkout → webhook → license-key delivery → license activation → Pro unlock. Build a single integration test that mocks Stripe's API, walks the complete flow, and asserts the license works at the end. This is the revenue path — it must be tested as a chain, not just individual pieces. Effort: ~2 hr.
- impact: Confidence in the revenue pipeline. A broken webhook or catalog mismatch silently loses sales.
DC-067: No graceful shutdown — SIGTERM kills in-flight requests
- status: pending
- details: server.js handles
uncaughtExceptionandunhandledRejection, but there is noSIGTERMhandler that callsserver.close()to drain connections. Docker stop sends SIGTERM (the Dockerfile hasSTOPSIGNAL SIGTERM), but without a handler the process exits immediately, dropping any in-flight API calls. Fix: add aSIGTERMhandler in server.js that (1) stops accepting new connections viaserver.close(), (2) waits up to 10s for in-flight requests, (3) closes DB/file handles, (4) exits cleanly. Also emit ashutdownevent so managers (health checker, SSL monitor, workflow engine) can stop their timers. Effort: ~1 hr. - impact: Zero-downtime deployments. Currently, every
docker stopdrops active requests.
DC-068: ESLint warnings sweep — 173 pre-existing warnings
- status: pending
- details: While there are 0 ESLint errors, 173 warnings remain. Top files:
dns-providers/base.js(27),update-manager.js(14),backup-manager.js(10),keychain-manager.js(10),bundled-workflows.js(10),auth/providers/base.js(9),log-digest.js(8). Most areno-unused-vars,require-await,no-nested-ternary. Fix: sweep through the top 10 files, fix what's actionable (unused vars → remove, nested ternaries → extract to named variables, false-positive require-await → mark_or restructure). Set a ceiling: warnings should never increase. Effort: ~2 hr. - impact: Clean codebase. 173 warnings is noise that hides real issues when new ones are added.
DC-069: Health check notification spam — add failure threshold + cooldown
- status: pending
- details: The workflow engine sends a notification on EVERY health check failure (every 15 min). If a service is down for a day, that's 96 identical notifications. There is no backoff, no deduplication, no "service recovered" message. Fix: (1) Only notify on state TRANSITIONS (up→down, down→up), not every failure. (2) Add a
consecutiveFailuresthreshold (e.g., 2 failures before first alert) to avoid flapping noise. (3) Send a recovery notification when a service comes back up. (4) Optional: daily digest of uptime stats instead of per-failure alerts. Effort: ~1.5 hr. - impact: Operator sanity. The current notification volume is exactly why people mute alerting channels — and then miss real incidents.
P2 — Polish & Developer Experience
DC-070: No CI/CD pipeline — tests run manually
- status: pending
- details: There is no GitHub Actions / CI configuration. Tests are run manually before push. This means a bad commit can reach main if someone forgets to test. Fix: add
.github/workflows/test.yml(or Gitea Actions equivalent) that runsnpm ci && npx jest --coverageon every PR and push to main. Cache node_modules. Upload coverage report as artifact. Block merge on test failure or coverage decrease. Effort: ~1 hr. - impact: Automated quality gate. No bad commit reaches production.
DC-071: No error tracking / Sentry integration
- status: pending
- details: Errors go to
error.loginside the container. If the container is recreated (DC-050 migration), the error log is lost. There is no external error tracking. Fix: add an optional Sentry (or GlitchTip for self-hosted) integration. IfSENTRY_DSNenv var is set, initialize Sentry before Express. Wrap async handlers to capture exceptions. The error-handler.js middleware should forward to Sentry before returning the generic error response. Make it opt-in (no DSN = no Sentry, zero behavior change). Effort: ~1 hr. - impact: Production visibility. Right now, errors are invisible unless someone SSHs in and reads the log.
DC-072: Frontend bundle has no source maps in production
- status: pending
- details:
status/build.jsuses esbuild but the production build doesn't emit source maps. When a frontend error occurs in production, the stack trace points to minified bundle lines — useless for debugging. Fix: addsourcemap: trueto the esbuild production config. Serve.mapfiles from Caddy (they're already indist/). Optionally upload source maps to Sentry (DC-071). Effort: ~30 min. - impact: Frontend bug reports become actionable instead of "line 1 of core.js".
DC-073: No API request/response logging middleware for debugging
- status: pending
- details: While there is an audit logger for POST/PUT/DELETE, there's no request/response logging middleware for debugging purposes (like morgan or a custom equivalent). When an operator reports "the dashboard is slow" or "this endpoint returns 500 sometimes", there's no way to trace the request through the system. Fix: add an optional debug-level request logger that logs method, path, status, duration, and request ID. Gated behind
LOG_LEVEL=debugso it's off in production by default. Effort: ~45 min. - impact: Drastically reduces time-to-resolution for production issues.
DC-074: Docker image is not multi-stage — build artifacts bloat the image
- status: pending
- details: The Dockerfile copies source files into a single stage based on
node:20-alpine. The image includesdevDependenciesbecausenpm install --productionstill installs some optional deps, and there's no.dockerignore(so__tests__/,.git/,node_modules/from the host can leak in). Fix: (1) Add a.dockerignorefile excluding__tests__/,.git/,node_modules/,*.md,coverage/. (2) Convert to multi-stage: build stage installs all deps, production stage copies onlynode_modules/(production) + source. (3) Pin Node.js version:FROM node:20.10-alpineinstead ofnode:20-alpine(floating). Effort: ~1 hr. - impact: Smaller image = faster pulls = faster deploys. Current image size carries unnecessary weight.
DC-075: No health check dashboard endpoint for operators
- status: pending
- details: The
/api/v1/monitoring/statsendpoint returns container stats, but there's no single "is everything OK" endpoint that returns a human-readable system health summary. Fix: addGET /api/v1/system/healththat returns{ status: "healthy"|"degraded"|"unhealthy", checks: { database: "ok", diskSpace: "ok", memory: "ok", uptime: ..., activeServices: N/M, lastError: "..." } }. This is useful for uptime monitoring services (UptimeRobot, BetterStack) and for a quick operator glance. Effort: ~1 hr. - impact: Operators can plug DashCaddy into external monitoring without parsing container stats.
P3 — Future & Nice-to-Have
DC-076: WebSocket support for real-time dashboard updates
- status: pending
- details: The dashboard polls the API every N seconds for service status updates. For a "live" dashboard experience, WebSocket (or SSE) push would be better — status changes appear instantly without polling overhead. Fix: add a WebSocket server (using
wslibrary) that pushes service status changes, health check results, and container events to connected dashboard clients. Keep polling as fallback for clients without WS support. Effort: ~3 hr. - impact: Dashboard feels "live". Reduces API load from polling.
DC-077: Multi-language (i18n) support
- status: pending
- details: All UI text is hardcoded English. For a public product, internationalization is a step toward wider reach. Fix: extract all user-facing strings into a locale file, add an i18n library (like i18next), provide at minimum an English + Arabic locale (Sami's audience). Effort: ~4 hr.
- impact: Market expansion. Arabic-speaking homelab community is underserved.
DC-078: Backup and restore of DashCaddy's own configuration
- status: pending
- details: While DashCaddy can backup app data, there's no one-click "backup my entire DashCaddy setup" (services.json, config.json, health-config.json, credentials, Caddyfile, license) that could be restored on a fresh install. Fix: add
GET /api/v1/system/export(returns a signed JSON bundle) andPOST /api/v1/system/import(restores from bundle). The credentials file should be encrypted with a user-provided passphrase. Effort: ~2 hr. - impact: Migration story. "Moving DashCaddy to a new host" is currently a multi-hour manual process.
DC-079: Mobile-responsive dashboard improvements
- status: pending
- details: While the dashboard is somewhat responsive, it's not optimized for mobile use. For operators checking services on their phone, the experience should be touch-first. Fix: audit all dashboard pages on mobile viewport, fix any horizontal scroll, ensure buttons are touch-target sized (min 44px), add a mobile-specific layout for the service grid. Effort: ~3 hr.
- impact: Operators check services on their phone. Current mobile experience is usable but not polished.
DC-080: Plugin/extension system for custom services
- status: pending
- details: DashCaddy supports a fixed set of service templates. A plugin system would allow community-contributed service definitions (e.g., "Home Assistant", "Vaultwarden", "Nextcloud") without modifying core code. Fix: define a plugin manifest schema (name, logo, health check URL pattern, config fields), load plugins from
/data/plugins/, add a community plugin registry page. Effort: ~4 hr. - impact: Community growth. Extensibility is what makes a tool ecosystem vs. a product.
Summary by Priority
| Priority | Count | Effort | Theme |
|---|---|---|---|
| P0 | 3 (DC-062–064) | ~7 hr | Public release blockers |
| P1 | 5 (DC-065–069) | ~7 hr | Reliability & code quality |
| P2 | 6 (DC-070–075) | ~5.5 hr | Polish & DX |
| P3 | 5 (DC-076–080) | ~16 hr | Future growth |
| Total | 19 | ~35.5 hr |