Files
dashcaddy/DC-PRODUCTION-GRADE-BACKLOG.md
T
Hermes dc788e5dd3
CI / Test & Lint (push) Canceled after 0s
CI / Security audit (push) Canceled after 0s
DC-061: healthCheckUrl override + v2 production-grade backlog
- url-resolver.js: add healthCheckUrl priority (bypasses SSO for health checks)
- DC-PRODUCTION-GRADE-BACKLOG.md: v2 backlog with 19 tasks (DC-062–DC-080)
  based on full codebase audit: 1539 tests, 86.55% coverage, 0 ESLint errors

v2 backlog replaces completed v1 (P0-1 through P2-7 all done).
New priorities:
  P0: OpenAPI spec update, branch coverage gap, Dockerfile resource limits
  P1: Console sweep remainder, billing E2E test, graceful shutdown, lint sweep, health notification spam
  P2: CI/CD pipeline, Sentry, source maps, request logging, multi-stage Docker, health endpoint
  P3: WebSocket, i18n, config backup/restore, mobile, plugin system
2026-08-12 02:05:08 -07:00

14 KiB
Raw Blame History

DashCaddy Production-Grade Backlog (v2)

Generated 2026-08-12 from a full codebase audit. v1 items (P0-1 through P2-7) are ALL DONE. Current state: 1539 tests, 86.55% statement coverage, 0 ESLint errors, 173 warnings.

Current Health Snapshot

  • Tests: 1539 passing across 63 suites
  • Coverage: Statements 86.55% | Branches 72.14% (below 80% gate) | Functions 80.8% | Lines 90.67%
  • ESLint: 0 errors, 173 warnings (all pre-existing)
  • Remaining console. calls in src/:* 21 across 10 files
  • Dockerfile: Runs as root (documented — needs Docker socket), no resource limits
  • OpenAPI spec: Present but stale (says v1.0.0, actual is v1.15.0)
  • Unhandled rejection/exception handlers: Present in server.js ✓
  • Rate limiting: Present on auth + general routes ✓
  • npm audit: 4 remaining vulns (semver-major transitive deps, deferred)

P0 — Must Fix (blocks public release)

DC-062: OpenAPI spec is stale — update to match actual v1.15.0 API surface

  • status: pending
  • details: openapi.yaml says version: 1.0.0 and describes only a fraction of the API. Since DC-046/047 (auth providers), DC-053 (share), DC-055 (billing), DC-058 (share UI), and the tailscale-admin routes were added, the spec is significantly out of date. A stale spec is worse than no spec — it misleads API consumers and breaks any code generation from it. Fix: audit all route files (grep -rn 'router\.\(get\|post\|put\|delete\|patch\)' routes/), update openapi.yaml with every endpoint, bump version to 1.15.0, add it to the test suite (DC-017-style source-of-truth test that fails if a route exists but has no spec entry). Effort: ~3 hr.
  • impact: Public API trust. No paying customer can integrate against an undocumented API.

DC-063: Branch coverage at 72% — below the 80% gate

  • status: pending
  • details: Jest coverage report shows branches at 72.14% (303/420), failing the 80% threshold. The uncovered branches are concentrated in error-handling paths (catch blocks, fallback returns, edge-case conditionals). Fix: run npx jest --coverage --coverageReporters=text to identify the files with the lowest branch coverage, then add targeted tests for the uncovered conditional paths. Priority files: backup-manager.js (multiple catch blocks), health-checker.js (timeout/retry branches), tailscale-coord.js (API error branches). Effort: ~2 hr.
  • impact: Error paths are where production incidents hide. Every untested catch block is a potential crash.

DC-064: Dockerfile runs as root with no resource limits

  • status: pending
  • details: The Dockerfile has no USER directive and start.sh has no --memory or --cpus flags. While root is needed for Docker socket access, the container can still OOM the host. Fix: (1) Add --memory=512m --memory-swap=1g --cpus=1.5 to the docker run in start.sh. (2) Create a non-root user dashcaddy for the application process, and use a Docker socket proxy (like tecnativa/docker-socket-proxy) that exposes a limited subset of Docker API endpoints — the app only needs read access for monitoring + controlled container lifecycle. (3) Add --restart=unless-stopped if not already present. Effort: ~2 hr. Risk: medium — socket proxy may break some Docker API calls, needs testing.
  • impact: Without limits, a memory leak in the API can take down the entire host. This is a production safety issue.

P1 — Code Quality & Reliability

DC-065: Remaining 21 console.* calls — sweep to structured logger

  • status: pending
  • details: After DC-060 (update-manager) and P1-3 through P1-8, 21 console calls remain across 10 files: error-handler.js (2), email.js (1), dns-providers/registry.js (2), audit-logger.js (3), csrf-protection.js (3), config-drift-detector.js (1), auto-restart-manager.js (1), http.js (1), logging.js (6 intentional — the logger itself), routes/backups.js (1). The logging.js calls are fine (the logger IS console internally). The rest should route through log.info/warn/error. Some are fallbacks: ctx.logError || ((_c, err) => console.error(err)) — these fire when ctx isn't available, which is exactly when structured logging matters most. Effort: ~45 min.
  • impact: Consistency. The logger write to error.log and supports structured JSON — console does not.

DC-066: No API integration test for the billing flow end-to-end

  • status: pending
  • details: DC-057 shipped contract tests and unit tests for the Stripe bridge, but there is no test that exercises the full flow: pricing page → Stripe Checkout → webhook → license-key delivery → license activation → Pro unlock. Build a single integration test that mocks Stripe's API, walks the complete flow, and asserts the license works at the end. This is the revenue path — it must be tested as a chain, not just individual pieces. Effort: ~2 hr.
  • impact: Confidence in the revenue pipeline. A broken webhook or catalog mismatch silently loses sales.

DC-067: No graceful shutdown — SIGTERM kills in-flight requests

  • status: pending
  • details: server.js handles uncaughtException and unhandledRejection, but there is no SIGTERM handler that calls server.close() to drain connections. Docker stop sends SIGTERM (the Dockerfile has STOPSIGNAL SIGTERM), but without a handler the process exits immediately, dropping any in-flight API calls. Fix: add a SIGTERM handler in server.js that (1) stops accepting new connections via server.close(), (2) waits up to 10s for in-flight requests, (3) closes DB/file handles, (4) exits cleanly. Also emit a shutdown event so managers (health checker, SSL monitor, workflow engine) can stop their timers. Effort: ~1 hr.
  • impact: Zero-downtime deployments. Currently, every docker stop drops active requests.

DC-068: ESLint warnings sweep — 173 pre-existing warnings

  • status: pending
  • details: While there are 0 ESLint errors, 173 warnings remain. Top files: dns-providers/base.js (27), update-manager.js (14), backup-manager.js (10), keychain-manager.js (10), bundled-workflows.js (10), auth/providers/base.js (9), log-digest.js (8). Most are no-unused-vars, require-await, no-nested-ternary. Fix: sweep through the top 10 files, fix what's actionable (unused vars → remove, nested ternaries → extract to named variables, false-positive require-await → mark _ or restructure). Set a ceiling: warnings should never increase. Effort: ~2 hr.
  • impact: Clean codebase. 173 warnings is noise that hides real issues when new ones are added.

DC-069: Health check notification spam — add failure threshold + cooldown

  • status: pending
  • details: The workflow engine sends a notification on EVERY health check failure (every 15 min). If a service is down for a day, that's 96 identical notifications. There is no backoff, no deduplication, no "service recovered" message. Fix: (1) Only notify on state TRANSITIONS (up→down, down→up), not every failure. (2) Add a consecutiveFailures threshold (e.g., 2 failures before first alert) to avoid flapping noise. (3) Send a recovery notification when a service comes back up. (4) Optional: daily digest of uptime stats instead of per-failure alerts. Effort: ~1.5 hr.
  • impact: Operator sanity. The current notification volume is exactly why people mute alerting channels — and then miss real incidents.

P2 — Polish & Developer Experience

DC-070: No CI/CD pipeline — tests run manually

  • status: pending
  • details: There is no GitHub Actions / CI configuration. Tests are run manually before push. This means a bad commit can reach main if someone forgets to test. Fix: add .github/workflows/test.yml (or Gitea Actions equivalent) that runs npm ci && npx jest --coverage on every PR and push to main. Cache node_modules. Upload coverage report as artifact. Block merge on test failure or coverage decrease. Effort: ~1 hr.
  • impact: Automated quality gate. No bad commit reaches production.

DC-071: No error tracking / Sentry integration

  • status: pending
  • details: Errors go to error.log inside the container. If the container is recreated (DC-050 migration), the error log is lost. There is no external error tracking. Fix: add an optional Sentry (or GlitchTip for self-hosted) integration. If SENTRY_DSN env var is set, initialize Sentry before Express. Wrap async handlers to capture exceptions. The error-handler.js middleware should forward to Sentry before returning the generic error response. Make it opt-in (no DSN = no Sentry, zero behavior change). Effort: ~1 hr.
  • impact: Production visibility. Right now, errors are invisible unless someone SSHs in and reads the log.

DC-072: Frontend bundle has no source maps in production

  • status: pending
  • details: status/build.js uses esbuild but the production build doesn't emit source maps. When a frontend error occurs in production, the stack trace points to minified bundle lines — useless for debugging. Fix: add sourcemap: true to the esbuild production config. Serve .map files from Caddy (they're already in dist/). Optionally upload source maps to Sentry (DC-071). Effort: ~30 min.
  • impact: Frontend bug reports become actionable instead of "line 1 of core.js".

DC-073: No API request/response logging middleware for debugging

  • status: pending
  • details: While there is an audit logger for POST/PUT/DELETE, there's no request/response logging middleware for debugging purposes (like morgan or a custom equivalent). When an operator reports "the dashboard is slow" or "this endpoint returns 500 sometimes", there's no way to trace the request through the system. Fix: add an optional debug-level request logger that logs method, path, status, duration, and request ID. Gated behind LOG_LEVEL=debug so it's off in production by default. Effort: ~45 min.
  • impact: Drastically reduces time-to-resolution for production issues.

DC-074: Docker image is not multi-stage — build artifacts bloat the image

  • status: pending
  • details: The Dockerfile copies source files into a single stage based on node:20-alpine. The image includes devDependencies because npm install --production still installs some optional deps, and there's no .dockerignore (so __tests__/, .git/, node_modules/ from the host can leak in). Fix: (1) Add a .dockerignore file excluding __tests__/, .git/, node_modules/, *.md, coverage/. (2) Convert to multi-stage: build stage installs all deps, production stage copies only node_modules/ (production) + source. (3) Pin Node.js version: FROM node:20.10-alpine instead of node:20-alpine (floating). Effort: ~1 hr.
  • impact: Smaller image = faster pulls = faster deploys. Current image size carries unnecessary weight.

DC-075: No health check dashboard endpoint for operators

  • status: pending
  • details: The /api/v1/monitoring/stats endpoint returns container stats, but there's no single "is everything OK" endpoint that returns a human-readable system health summary. Fix: add GET /api/v1/system/health that returns { status: "healthy"|"degraded"|"unhealthy", checks: { database: "ok", diskSpace: "ok", memory: "ok", uptime: ..., activeServices: N/M, lastError: "..." } }. This is useful for uptime monitoring services (UptimeRobot, BetterStack) and for a quick operator glance. Effort: ~1 hr.
  • impact: Operators can plug DashCaddy into external monitoring without parsing container stats.

P3 — Future & Nice-to-Have

DC-076: WebSocket support for real-time dashboard updates

  • status: pending
  • details: The dashboard polls the API every N seconds for service status updates. For a "live" dashboard experience, WebSocket (or SSE) push would be better — status changes appear instantly without polling overhead. Fix: add a WebSocket server (using ws library) that pushes service status changes, health check results, and container events to connected dashboard clients. Keep polling as fallback for clients without WS support. Effort: ~3 hr.
  • impact: Dashboard feels "live". Reduces API load from polling.

DC-077: Multi-language (i18n) support

  • status: pending
  • details: All UI text is hardcoded English. For a public product, internationalization is a step toward wider reach. Fix: extract all user-facing strings into a locale file, add an i18n library (like i18next), provide at minimum an English + Arabic locale (Sami's audience). Effort: ~4 hr.
  • impact: Market expansion. Arabic-speaking homelab community is underserved.

DC-078: Backup and restore of DashCaddy's own configuration

  • status: pending
  • details: While DashCaddy can backup app data, there's no one-click "backup my entire DashCaddy setup" (services.json, config.json, health-config.json, credentials, Caddyfile, license) that could be restored on a fresh install. Fix: add GET /api/v1/system/export (returns a signed JSON bundle) and POST /api/v1/system/import (restores from bundle). The credentials file should be encrypted with a user-provided passphrase. Effort: ~2 hr.
  • impact: Migration story. "Moving DashCaddy to a new host" is currently a multi-hour manual process.

DC-079: Mobile-responsive dashboard improvements

  • status: pending
  • details: While the dashboard is somewhat responsive, it's not optimized for mobile use. For operators checking services on their phone, the experience should be touch-first. Fix: audit all dashboard pages on mobile viewport, fix any horizontal scroll, ensure buttons are touch-target sized (min 44px), add a mobile-specific layout for the service grid. Effort: ~3 hr.
  • impact: Operators check services on their phone. Current mobile experience is usable but not polished.

DC-080: Plugin/extension system for custom services

  • status: pending
  • details: DashCaddy supports a fixed set of service templates. A plugin system would allow community-contributed service definitions (e.g., "Home Assistant", "Vaultwarden", "Nextcloud") without modifying core code. Fix: define a plugin manifest schema (name, logo, health check URL pattern, config fields), load plugins from /data/plugins/, add a community plugin registry page. Effort: ~4 hr.
  • impact: Community growth. Extensibility is what makes a tool ecosystem vs. a product.

Summary by Priority

Priority Count Effort Theme
P0 3 (DC-062064) ~7 hr Public release blockers
P1 5 (DC-065069) ~7 hr Reliability & code quality
P2 6 (DC-070075) ~5.5 hr Polish & DX
P3 5 (DC-076080) ~16 hr Future growth
Total 19 ~35.5 hr