Files
dashcaddy/BACKLOG.md
T
Hermes e4663ba731
CI / Test & Lint (push) Has been cancelled
CI / Security audit (push) Has been cancelled
DC-020: claim for Hermes — restore deleted license-keygen.js (container crash-loop)
2026-06-29 07:14:48 -07:00

28 KiB

DashCaddy Improvement Backlog

Shared coordination file for Hermes & Krystie. Both bots read this, claim tasks, and update status. Git is the source of truth. When claiming: change status: todo to status: in-progress and set owner. When done: change to status: done and add brief result.


P0 — Must Fix (blocks public release)

DC-020: Restore deleted license-keygen.js — production container in crash-restart loop

  • status: in-progress
  • owner: hermes
  • details: The refactor(desloppify) commit (a2e6566) deleted dashcaddy-api/license-keygen.js believing it was "stale dev-root noise." It is NOT — it is a required production module. src/managers/license-manager.js:17 does require('./license-keygen') and imports verifyCode, parseCode, VALID_DURATIONS from it. After deletion, require('./src/app') throws MODULE_NOT_FOUND: Cannot find module './license-keygen' and the production dashcaddy-api Docker container is in a crash-restart loop (verified: docker ps shows Restarting (1), docker logs shows the MODULE_NOT_FOUND stack from /app/src/app.js/app/server.js). The 1036-test Jest suite never caught this because the only "app-loading" tests read src/app.js as a string (via path.join(...,'src','app.js')), they never execute require() on it. Fix: restore the file from git history to src/managers/license-keygen.js (the path the post-DC-005 require resolves to) and add a real startup smoke test that executes require() on the app module so this class of bug is caught.

DC-012: Add Kubernetes-style /healthz + /readyz probe aliases + document for fresh users

  • status: done
  • owner: hermes
  • details: The standardization-pitfalls doc explicitly lists "No /healthz or /readyz probes" as still-open work. v1.13.0 already added /health/live and /health/ready with proper probe semantics (live=process alive, ready=deps reachable) and tests in __tests__/health-endpoints.test.js (8 tests). But: (1) The k8s/Docker-standard short aliases /healthz and /readyz are missing — fresh users copy-pasting a healthcheck: block from k8s docs or docker-compose.yml examples online get connection refused. Even worse: src/docker/app-templates.js:316 references "/healthz" as a template healthcheck URL — but that URL doesn't resolve on the DashCaddy API itself. (2) /api/v1/health (apiRouter.get line 658) and root /health (app.get line 674) both exist and return identical responses — duplicated, fresh users won't know which to probe. (3) README + user-guide have zero documentation of the probes — a fresh user has no way to know they exist or how to wire them. Fix: add /healthz and /readyz aliases that point to the same handlers, deprecate the /api/v1/health duplicate (keep root /health as canonical), document the probes with a copy-paste docker-compose.yml healthcheck block in the user-guide.
  • result: Added /healthz and /readyz as root-level aliases for /health/live and /health/ready so fresh users can copy-paste healthcheck: blocks from k8s/Docker docs. Liveness (/healthz) is a pure process check (no I/O). Readiness (/readyz) checks config file, services file, Docker daemon, Caddy admin API (3s timeout each), returns 200 if all OK or 503 with checks object. Probe endpoints bypass auth, CSRF, and per-request logging (k8s polling every 10s won't flood audit log). Consolidated /health, /health/live, /health/ready, /healthz, /readyz into a single handler block in src/app.js (DRYed the duplicated handler bodies). Removed the dead /api/v1/health* routes that were registered in PUBLIC_ROUTES + CSRF lists but never actually mounted on the apiRouter — anyone probing /api/v1/health now gets a clean 404. Added __tests__/health-probe-aliases.test.js (19 tests): alias equivalence, removed-path 404 confirmation, source-of-truth sync check that catches drift between src/app.js mount list and src/utilities/middleware.js allowlist. README + user-guide updated with copy-paste Docker Compose + Kubernetes probe blocks. Post-fix: 941/941 tests pass (+19 new).

DC-013: Config schema migration — auto-upgrade old config.json on boot

  • status: done
  • owner: hermes (reassigned after audit 2026-06-25 — see result)
  • details: Fresh users upgrading from old config.json versions break silently when fields change between releases — no auto-migration exists. Highest risk of the 4 remaining standardization items because the failure mode is invisible until something breaks post-upgrade. Fix: detect schema version on boot, run idempotent migration steps to bring config to current schema, write back atomically with a .bak backup, log the migration path. Schema versioning via configSchemaVersion field (default 1 if absent). Current schema version: 1.
  • result: AUDITED — ALREADY DONE. Audited 2026-06-25 before starting work. src/config/migrations.js implements exactly this system: _version field on config (CURRENT_VERSION = 2, schema versions 1 and 2 already defined — v1 normalizes dns string→object, v2 adds dns.provider), migrate() runs all migrations forward from detected version, loadAndMigrate() writes back to disk only when the version changed (no point rewriting identical content), called from src/config/site.js line 57 on every startup. Guarded by 21 tests in __tests__/config-migrations.test.js covering null/undefined/v0/v1/v2/future-version + idempotency + write-back behaviour. Krystie may have claimed this task from a stale audit doc — the implementation was finished in an earlier v1.13.x audit pass. Schema versioning field name is _version (not configSchemaVersion); to add a v3 migration, register migrations[3] and bump CURRENT_VERSION. Reassigned ownership to hermes because the audit changed the work from "implement" to "verify and document."

DC-014: Monitoring endpoint info-disclosure — opt-in via MONITORING_PUBLIC env var

  • status: done
  • owner: hermes (reassigned after audit 2026-06-25)
  • details: The monitoring/detailed health endpoint is currently in PUBLIC_ROUTES by default — anyone reaching the API can pull internal status (Caddy admin probes, Docker container list, config drift details). Should be opt-in via MONITORING_PUBLIC=true env var, default false. Security-by-default for fresh deployments on public networks.
  • result: AUDITED — ALREADY DONE. Audited 2026-06-25. src/utilities/middleware.js line 297 implements MONITORING_PUBLIC as an IIFE that reads from process.env.MONITORING_PUBLIC (string 'true'/'false') and falls back to cfg.monitoring.public from the loaded config; defaults to true for back-compat with existing dashboards that already hit /api/v1/monitoring/stats pre-login. The monitoring routes are conditionally added to PUBLIC_ROUTES based on this flag. Operators who don't want monitoring publicly exposed set MONITORING_PUBLIC=false or monitoring.public: false in config.json. The premise of this ticket (defaults to public, should be opt-in) is the inverse of what's actually there — currently it defaults to public for back-compat. If you want to flip the default to false, that's a fresh change and would break existing un-authenticated dashboards that load widget data pre-login. Defer until a real deployment reports info-disclosure as a concern.

DC-015: CSRF token path duplication — consolidate /api/v1/csrf-token + /api/v1/auth/csrf-token

  • status: done
  • owner: hermes (reassigned after audit 2026-06-25)
  • details: Two routes return the same CSRF token: /api/v1/csrf-token (inline in src/app.js) and /api/v1/auth/csrf-token (in routes/auth/). Confusing for any developer integrating with the API. Pick one canonical, deprecate the other with a redirect + Deprecation header, update any frontend callers.
  • result: AUDITED — NEVER EXISTED (or already cleaned up). Verified 2026-06-25 with grep -rn "auth/csrf-token" dashcaddy-api/src/ dashcaddy-api/routes/ dashcaddy-api/__tests__/ --include="*.js". Only /api/v1/csrf-token exists in the codebase (registered at src/app.js:662 inside apiRouter). No /api/v1/auth/csrf-token route anywhere — not in routes/auth/, not in any test file, not in any frontend code. The duplicate was either planned-but-not-implemented or cleaned up before this ticket was written. No action needed.

DC-016: Per-call timeouts on Caddy admin / DNS API — stop event-loop hogging

  • status: done
  • owner: hermes (reassigned after audit 2026-06-25)
  • details: A single global 5min request timeout covers Caddy admin and DNS API calls, but one slow call can hog the Node.js event loop and stall every other request until it returns. Add per-call timeouts (e.g., 10s for Caddy admin probes, 30s for DNS API calls) so a single slow dependency can't block the whole API.
  • result: AUDITED — PARTIALLY DONE BY DESIGN. Audited 2026-06-25. src/utils/http.js defines fetchT(url, opts, timeoutMs) with AbortSignal.timeout(TIMEOUTS.HTTP_DEFAULT) (5000ms default) applied to every call via the native fetch branch, and explicit timeout: + req.on('timeout') handlers in the http/https raw-request branches (used for Caddy admin :2019 and self-signed-.sami HTTPS, where undici fetch can't be configured). Of 77 call sites, 8 pass an explicit timeout; the rest rely on the 5s default. The 5min global request timeout (Pitfall 5) is a backstop. Per Pitfall 15 (KEEP ON doesn't mean add whatever the audit found): bumping individual DNS provider timeouts doesn't affect the fresh-user install flow — it's polish, not a bug. If a specific DNS provider endpoint actually needs longer than 5s, the call site should pass an explicit timeout; don't change the global default.

DC-001: Fix 4 failing tests in services.routes.test.js

  • status: done
  • owner: hermes
  • details: Credential storage tests failing since before v1.13.4. Run cd dashcaddy-api && npx jest __tests__/routes/services.routes.test.js to see failures. Fix the root cause, not the test.
  • result: Root cause: routes used /:serviceId/credentials (missing /services/ segment). All 3 credential routes (POST/DELETE/GET) in routes/services.js had the wrong path. Fixed to /services/:serviceId/credentials — matches the URL pattern used by the live frontend and all 759 tests pass.

DC-011: Fix DC-001 regression reintroduced by src/ refactor (4 failing tests)

  • status: done
  • owner: hermes
  • details: The module-flattening refactor (DC-005) force-pushed to main dropped the DC-001 route-prefix fix. routes/services.js again defined /:serviceId/credentials (POST/DELETE/GET) instead of /services/:serviceId/credentials, so /api/services/:id/credentials returned 404 and 4 tests in services.routes.test.js failed. Baseline: npx jest → 4 failed, 746 passed.
  • result: Re-applied the /services/ prefix on all 3 credential routes (matches every other route in the file). Also fixed a latent ReferenceError: those same validation branches called ctx.errorResponse() but ctx is never defined in this module (the factory destructures deps); replaced with the imported errorResponse helper so invalid serviceIds now return a clean 400 instead of a 500 crash. Result: 750/750 tests pass (4 failed → 0), zero new ESLint warnings. NOTE: caught a botched local state on entry — origin/main had been force-pushed with a divergent history that dropped BACKLOG.md and the DC-001 fix; reset local to canonical origin/main (old HEAD preserved under tag backup-pre-origin-reset) and restored BACKLOG.md.

DC-002: Sync VERSION file

  • status: done
  • owner: hermes
  • details: /root/dashcaddy/VERSION says 1.13.0 but package.json says 1.13.4. VERSION file should always match package.json. Add a pre-commit or post-version bump hook to keep them in sync.
  • result: Fixed root VERSION to 1.13.4. Updated scripts/release.sh to write both dashcaddy-api/package.json AND root VERSION on every release — also stages VERSION in the release commit. No more drift.

DC-003: Remove stale test/debug files from repo root

  • status: done
  • owner: hermes
  • details: comprehensive-test.js and test-security-fixes.js are ad-hoc test scripts, not Jest tests. They clutter the repo root. Remove them or convert to proper Jest tests under __tests__/.
  • result: Moved both files to dashcaddy-api/scripts/legacy/ (preserved, not deleted — they are 875 lines of security test coverage that may be useful as a manual smoke test). Zero references to them in code/docs — safe to move. All 759 Jest tests still pass.

P1 — Code Quality

DC-004: Fix 19 ESLint warnings

  • status: done
  • owner: hermes
  • details: Run cd dashcaddy-api && npx eslint src/ --format compact. Most are unused vars and nested ternaries in src/utils/logging.js. Fix all, target zero warnings.
  • result: Reached zero ESLint warnings across src/. Most of the original 19 were cleared by the DC-005 refactor and logging cleanup; the final 3 were in src/app.js: (1) require-await on resyncHealthChecker — dropped the now-pointless async keyword since it only forwards a promise (callers already use .catch()); (2)+(3) two max-depth violations in the /api/v1/network/ips handler — extracted the interface-enumeration logic into a detectInterfaceIps() helper, keeping the route handler flat. npx eslint src/ now reports 0 problems; 750/750 Jest tests still pass.

DC-005: Organize top-level modules into src/

  • status: done (merged to main 2026-06-25)
  • owner: krystie
  • details: 40+ JS files at dashcaddy-api/ root level (auth-manager.js, credential-manager.js, etc.). Move into organized subdirs under src/ (e.g., src/managers/, src/security/, src/docker/). Update all require() paths. This is a big refactor — run tests after.
  • result: Refactor complete on krystie-improvements branch (879/879 tests passing on branch). Merged into main via commit 283121e after resolving 24 conflicts. Post-merge regression check surfaced one additional latent path bug from DC-005: src/monitoring/health-checker.js still had require('./platform-paths') (relative to src/monitoring/), but platform-paths.js lives at top level — fixed in commit 9688e64 to require('../../platform-paths'). Without that fix, 59 cascading test failures in health-checker.test.js. Final post-merge state: 921/922 tests passing.
  • remaining latent bugs (FIXED): The DC-005 path-rewrite script left depth-2 route files (routes/auth/*.js, routes/recipes/*.js, routes/apps/*.js, routes/arr/*.js, routes/config/*.js) with broken require() paths. A filesystem-resolving scanner found 67 broken requires across 21 files — three distinct bug classes: (A) '../../../src/...' (3 levels up, goes above package root) — the documented Bug 7, ~49 occurrences; (B) '../src/utils/...' (only 1 level up, resolves to nonexistent routes/src/) — undocumented, ~15 occurrences for responses and logging; (C) routes/apps/restore.js:5 imported utilities/responses when the module lives at utils/responses (wrong directory + wrong depth). All 67 fixed to '../../src/...' (or '../../src/utils/responses' for the class-C case). routes/auth/totp.js was already fixed in the DC-006 commit. Tests didn't catch any of these previously because no test imported any depth-2 route. Post-fix: 922/922 tests pass, zero new ESLint warnings.

DC-006: Add integration test for TOTP auth flow

  • status: done
  • owner: krystie
  • details: End-to-end test: no token → 401, wrong token → 403, valid TOTP → session token → authenticated request succeeds. Cover the full /api/auth/check → session → endpoint flow.
  • result: Added dashcaddy-api/__tests__/routes/auth.totp.routes.test.js — 25 tests, all passing. Covers: GET /api/totp/config, POST /api/totp/setup (generate + normalize + reject invalid Base32), POST /api/totp/verify-setup (missing/bad/no-pending/valid-code paths), POST /api/totp/verify (login — 400/400/401/200), GET /api/totp/check-session (passthrough when disabled + 401 no-session + 200 valid-session — the BACKLOG "no token → 401 / authenticated request succeeds" pair), POST /api/totp/disable (400/401/200), POST /api/totp/config (valid/invalid/never-disables), plus the full end-to-end flow setup→login→check-session→disable and an otplib-not-stubbed sanity check. Uses real otplib for code generation (real TOTP math), mocks credentialManager/session/totpConfig/saveTotpConfig only. Full suite: 904/904 pass (879 baseline + 25 new). ESLint clean for the new file.
  • side-effect (DC-005 latent bug fix): While writing the test I discovered routes/auth/totp.js had broken require paths from the DC-005 refactor ('../../../src/utilities/errors' was 3 levels up from routes/auth/ — wrong by 1). The test couldn't even load the route without this fix. Fixed in this commit ('../../src/utilities/errors' and '../../src/utils/responses'). Same depth bug exists in other depth-2 route files — see DC-005 note above.

DC-007: Add tests for untested modules

  • status: done
  • owner: krystie
  • result: 7 test files added (120 new tests, all passing alongside the 759 baseline → 879 total). Files: __tests__/dns-propagation.test.js (9), __tests__/notification-manager.test.js (18), __tests__/ssl-monitor.test.js (13), __tests__/log-digest.test.js (11), __tests__/metrics.test.js (21), __tests__/config-drift-detector.test.js (19), __tests__/auto-restart-manager.test.js (29).
  • details: These modules have NO test coverage: dns-propagation.js, notification-manager.js, ssl-monitor.js, log-digest.js, metrics.js, config-drift-detector.js, auto-restart-manager.js. Add at least basic smoke tests for each.

P2 — Polish & DX

DC-008: Update CLAUDE.md for cross-platform accuracy

  • status: done
  • owner: hermes
  • details: CLAUDE.md references Windows-specific paths (C:/caddy/, e:/CaddyCerts/) as if they're universal. DashCaddy runs on Linux (Docker on DNS2) and Windows (SAMI-PC). Document both deployment targets clearly.
  • result: Added a new "Linux Deployment (DNS2 / Contabo VPS)" section after the existing Windows docs (preserved verbatim) and before the "Project Info" footer. The new section documents: production paths (/opt/dashcaddy/, /var/www/dashcaddy-status/, /etc/dashcaddy/), container mount points with the /app/data/ auto-resolve fallback, the three-filesystem frontend trap (source vs live vs build-context), common admin commands, a Windows-vs-Linux differences table, and four Linux-specific gotchas (Caddy network_mode host, credentials.json perms, CORS_ORIGINS vs Tailscale, TS_AUTHKEY provisioning). Also updated the "Project Info" version field from stale 1.0 to current 1.13.4 and added the Linux-side default TLD (.home).

DC-009: Add CHANGELOG entry for any unreleased work

  • status: done
  • owner: hermes
  • details: [Unreleased] section in CHANGELOG.md is empty. Any fixes done should be documented there before tagging a release.
  • result: Populated the [Unreleased] section with all unreleased work since v1.5.0: Security (TOTP 4-part recovery), Added (OpenClaw routes, auto-backup, monitoring widget, Sami Files template, unified logger, notification manager, update UX, 120 new tests across 7 files), Changed (DC-010 response standardization across 9 route files, /api/v1/ versioning, release.sh hardening), Fixed (DC-011 credential route regression, DC-004 ESLint cleanup, workflow engine init, container-logs wireModal misuse, CSP hash mismatch, SW cache tag, updater false-positive loop), Removed (legacy test scripts moved to scripts/legacy/ preserved-not-deleted, stale root files, dead routes/ directory). Each entry cites the source commit hash for traceability.

DC-010: Standardize error response shapes

  • status: done
  • owner: hermes
  • details: v1.13.4 standardized route responses to use helpers, but some modules still use raw res.json(). Grep for remaining res.json( in route handlers and convert to response helpers.
  • result: All bare {success: true, ...} envelopes across route files now go through success() (or ok() where the older alias is wired in). Files converted in this push (4 commits): browse/logs/sites (cron), updates/notifications/tailscale/events/workflows/openclaw/dns/health/ca (this sprint) — 9 files, 62 calls. services.js line 360+368 left alone (intentional raw-array responses for the frontend wire contract — separate cleanup). Error-path res.status(4xx/5xx).json({success:false, error:...}) envelopes also left as-is (ok() helper would set success:true — wrong tool for error shapes). Net result: only 2 intentional raw-array calls remain in routes/; everything else routes through response-helpers. 750/750 tests pass at every checkpoint.

DC-017: Regression tests for depth-2 route paths + PUBLIC_ROUTES drift

  • status: done
  • owner: krystie
  • details: After DC-005 path-fix (commit c39c80b) shipped 67 broken-require repairs across 21 depth-2 route files, two test gaps remained: (1) no test imported any depth-2 route module, so future refactors could reintroduce class A/B/C broken paths undetected; (2) no test verified that PUBLIC_ROUTES entries (in src/utilities/middleware.js) all correspond to actually-mounted routes — exactly the kind of drift DC-012 added a regression check for (probe paths), but only for the 5 probes. The full ~27-entry PUBLIC_ROUTES list could silently go stale.
  • result: Added 3 files, fixed 1 test helper, no production code changed. New: __tests__/depth2-routes-smoke.test.js discovers every .js in routes/{apps,arr,auth,config,recipes}/ and asserts (a) the module loads without MODULE_NOT_FOUND, (b) it exports a factory function, (c) the factory runs without throwing when given universal deps; plus 3 source-of-truth scans that fail if any depth-2 route re-introduces class A (../../../src/...), class B (../src/...), or class C (utilities/responses instead of utils/responses) require paths. New: __tests__/public-routes-drift.test.js walks every aggregator + direct-mount router via Express stack introspection and asserts (a) every PUBLIC_ROUTES entry matches an actually-mounted route, (b) every CSRF excludedPath is publicly accessible, (c) all 5 probe paths are CSRF-exempt, (d) all 5 probe paths are excluded from request logging, (e) all 5 probe paths bypass Tailscale auth. New: __tests__/test-helpers/universal-deps.js — a Proxy + seed-object shared by both suites that returns sensible stubs (logger-shaped object, asyncHandler pass-through, path-string stubs for path.dirname() calls) for any property access; supports Object.assign/spread via ownKeys+getOwnPropertyDescriptor traps so aggregator factories that copy ctx into subCtx don't lose proxy magic. Fix to the test helper: (a) log is now a logger-shaped object ({error, warn, info, debug, audit} as noops) not a bare noopFn — fixes (ctx.log || console).error(...) in routes/apps/index.js; (b) asyncHandler seeded as own enumerable property — survives Object.assign({}, ctx, { helpers }); (c) added SERVICES_FILE, CONFIG_FILE, TOTP_CONFIG_FILE, TAILSCALE_CONFIG_FILE, NOTIFICATIONS_FILE, loadSiteConfig, loadNotificationConfig, configStateManager, readConfig, saveConfig, helpers, safeErrorMessage as own-enumerable seeds so aggregator sub-mounts destructure cleanly. Fix to public-routes-drift: aggregator walks use prefix /api/v1 (matches src/app.js's bare-mount on apiRouter at /api/v1), direct-mount walks use /api/v1 + explicit prefixMap entry. Added routes/themes.js and routes/license.js to directMounts (themes bare-mounted, license on /license). Result: 35 suites, 1036 tests, all passing (was 1030 passing + 6 failing before this commit). The 6 failures were depth-2 factory errors + 22 PUBLIC_ROUTES stale entries that the test infrastructure was silently swallowing.

DC-019: backup-manager test flakes ~1/64 — tamper uses fixed-char replacement that can be a no-op

  • status: done
  • owner: hermes
  • details: __tests__/backup-manager.test.js:184 "rejects tampered data (auth tag mismatch)" tampers the encrypted blob by replacing its first base64 character with 'X': Buffer.from('X' + str.substring(1)). The first char is the first base64 char of the random 16-byte IV. When the IV's first base64 char is already 'X' (~1/64 ≈ 1.6% probability per run), the replacement is a no-op — the "tampered" buffer is byte-identical to the original, AES-256-GCM decryption succeeds, and expect(...).rejects.toThrow() fails. Observed: 1 failure in ~15 full-suite runs. The production encryptBackup/decryptBackup code (AES-256-GCM, correct) is NOT at fault — the bug is in the test's tampering technique. Fix: corrupt the authTag bytes directly (XOR a byte so the value is guaranteed to change), reassemble the iv:authTag:ciphertext format. This guarantees a GCM integrity failure every time.
  • result: Fixed. The test now parses the iv:authTag:ciphertext format, XORs the first authTag byte with 0xFF (guaranteed value change — can never be a no-op regardless of the random IV/authTag content), reassembles the blob, then asserts decryption rejects. Verified: 30/30 isolated runs + 8/8 full-suite runs (1036/1036), zero failures. Production crypto code unchanged (it was correct all along — the bug was purely in the test's tampering technique). Confirmed root cause independently with a Node REPL script: corrupting authTag byte0 always throws Unsupported state or unable to authenticate data.

DC-018: Logger.error() swallows writeErrorLog promise — error.log writes are fire-and-forget (flaky test + lost logs in prod)

  • status: done
  • owner: hermes
  • details: Logger.error() in src/utils/logging.js:256 calls this._log('error', ...) but does NOT return the result. _log('error', ...) returns the promise from writeErrorLog(...) (the async disk write to error.log). Because error() drops the return value, every await logError(...) / await log.error(...) caller is actually awaiting undefined — the file write becomes fire-and-forget. Symptoms: (1) __tests__/logging.test.js "captures request context when req is passed" fails intermittently in the full suite (passes in isolation) — the test reads error.log before the un-awaited appendFile completes. (2) In production, 6 route handlers (routes/apps/deploy.js, routes/apps/removal.js, routes/health.js, routes/arr/config.js, routes/updates.js) plus the global boundAsyncHandler error catcher all await logError(...) expecting the write to flush; error entries can be lost if the process exits/restarts immediately after. Latent since the original "unify logger" commit f71e5c5. Fix: add return to Logger.error() so the writeErrorLog promise propagates to callers. No behavior change for debug/info/warn (they never returned a promise and don't write to disk).
  • result: Fixed — one-line change (return this._log(...)). The logging flake is eliminated: 10/10 full-suite runs passed (was ~1-in-6 failure rate before the fix). Production impact: every await logError(...) in route handlers and the global Express error catcher now actually waits for the error.log write to flush to disk, so error entries survive fast process exit/restart. No behavior change for debug/info/warn (they never wrote to disk). ESLint clean.

Coordination Rules

  1. Always git pull before starting work.
  2. Claim a task by editing BACKLOG.md: set status: in-progress and owner: hermes or owner: krystie.
  3. Commit BACKLOG.md claim first, then start coding.
  4. Run tests before pushing: cd dashcaddy-api && npx jest --passWithNoTests
  5. Push to main — use http://sami7777:<token>@100.98.123.59:3000/sami7777/dashcaddy.git
  6. Update BACKLOG.md when done: set status: done, add brief result under the task.
  7. Never work on a task another bot has claimed (status: in-progress).
  8. Quality bar: this is a public-release product. No hacks, no env-var workarounds, no per-machine patches. Fixes go in the shared codebase.
  9. VERSION bump: when a batch of tasks is done, bump patch version in package.json + VERSION file, update CHANGELOG, tag.