2dce6dca5eb273d68f1038015d28fecf700f8d00
396
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2dce6dca5e |
fix(security): event-store retention + race + total-cap defects (DC-116) [glm-grade=B]
- query().total: full-scan true count (was capped at offset+limit by an early break — dashboard 24h stat read ≤1000 vs real ~46k) - trim trigger/curer mismatch: byte trigger + line curer never converged when avg line > ~524B (live avg 355B); now keeps maxDisk lines OR ≤80% byte budget, whichever retains fewer (TRIM_TARGET_FACTOR); budget env-overridable via SECURITY_EVENT_TRIM_BYTES / opts.trimSizeLimit - trim/append race: trim renamed over the file mid-append losing events to the unlinked inode; now single-flight, queue-empty gated, holds the write lock, error paths always release + lazily re-kick (no hot loop) Judge: glm-4.6@zai-coding-paas adversarial cold-read, grade B clean (0 blocking, 4 polish — 2 folded, 1 already-satisfied, 1 deferred as judge-endorsed safer). 127 suites / 2849 tests green. |
||
|
|
65457ff8e0 |
feat(security): activate caddy security-event pipeline — bounded first-start replay, self-noise filter, host fidelity (DC-113) [glm-grade=A]
The caddy tail worker (DC-112) was 100% dead in prod: no /var/log/caddy mount, no CADDY_ACCESS_LOG env, no global access log in the Caddyfile. Store census 45,912 events, 100% source_type 'api', ZERO 'caddy'. - createTail: firstStartMaxBytes (5 MiB) bounds first-ever-start replay against a long-lived access.log; multi-chunk partial-line discard on the jump; normal restarts resume at exact persisted offset (judge r1 fix-first fold) - worker: self-noise filter drops our own probe UAs (DashCaddy-Probe/1.0, DashCaddy-HealthCheck/1.0) from the derived store — raw log keeps everything; ~50-100 events/min of probe noise would otherwise bury perimeter signal in the 100k-cap store - worker: metadata.host reads request.host (real caddy JSON nests it; verified against live /var/log/caddy/seeds.log — the DC-112 read was always null on live lines); top-level fallback kept - worker: onAppear recovery log (DC-112 judge polish fold) - start.sh: -v /var/log/caddy:/var/log/caddy:ro + CADDY_ACCESS_LOG env - README: dead caddy-api/ dir refs -> dashcaddy-api/ (queue item e) - tests: +12 DC-113 pins (hermetic, real worker + store); DC-112 fixtures corrected to the real nested request.host shape 126 suites / 2841 tests green. Judge: GLM-5.3 round-1 B, round-2 A (SHIP). Verdict urn:ump:mccln523fptotuvrmddlqpf4zpkrxf273tg4kyytuyy3776rj3pq |
||
|
|
a29a59a320 |
fix(security): name caddy-worker SSO gate events + dead-path visibility (DC-112) [glm-grade=A]
Queue item (g) — same defect class as DC-111 defect 1, different writer: the caddy access-log tail named every event http.<status>, so forward_auth gate hits were unanswerable in the unified security event store. - resolveCaddyAction() mirrors audit-logger ACTION_MAP vocabulary for BOTH URI shapes (legacy /api/auth/gate/<id> from Caddyfile line 87, canonical /api/v1/... from dashboard JS): auth.credential-injection, auth.app-token-issue, auth.sso-exchange (exact-path match, boundary-tested) - severity escalation now covers the legacy /api/auth/ prefix too - metadata fidelity: caddy logs headers as ARRAYS — old single-value read always produced user_agent:null; added metadata.host (vhost) and duration_seconds (caddy logs seconds; duration_ms kept, no consumers) - dead-path visibility: once-per-process warn when the access log is missing. Live discovery: prod store has 45,912 events, 100% source_type 'api', ZERO 'caddy' — the container has no /var/log/caddy mount, so the worker silently no-ops. Infra wiring queued separately. - new __tests__/caddy-worker-naming-dc112.test.js: 22 pins (real tail + real store, hermetic sinks, both URI shapes, boundary rows, offset persistence, once-warn). Full suite 125/2831 green. Judge: GLM-5.3 cold read round-1 A (deleg_65498358, 3 polish items folded) + round-2 A (deleg_eed9dcbf, judge re-ran suite itself). URN urn:ump:mjbwfcw6z7budx45s2rutovw5fatpp4jyzzspo7tq6nzajayulmq |
||
|
|
56b807543c |
fix(audit): restore named audit actions for SSO gate traffic — 3 live defects (DC-111) [glm-grade=A]
1. audit-logger middleware computed action/resource from req.path INSIDE the res.json override — after the /api/v1 router rebased req.url to the router-relative path. resolveAction fell through ACTION_MAP for every HTTP request, producing 45,899 'unknown.get' entries since 2026-07-14. Fix: snapshot req.path/req.method at app-level (post-shim, pre-router). 2. DC-044 shim double-prefixed ALREADY-canonical /api/v1/auth/gate|x into /api/v1/v1/... → 401 for canonical-URI clients. Fix: rewrite only legacy /api/auth/* shapes; canonical pass through untouched. 3. event-store VALID_OUTCOMES lacked 'failure' → every failed API action's security event was rejected+dropped from security-events.jsonl. Fix: add 'failure' to the vocabulary set. 8 regression pins over a faithful shim→audit→router mount mirror. Suite: 124 suites / 2809 tests green. Judge: GLM-5.3 cold read round-1 A/ship, URN urn:ump:ywv6rpmx55tlxbgc7ciyxqfcmx2v6q756cjbmif2d66k4bwwn6ya |
||
|
|
0721b1cb04 |
fix(security): audit-logger PII masking parity with unified logger (DC-110) [glm-grade=A]
audit-logger.js (StateManager write path into audit-log.json) starred only 6 sensitive keys at the middleware layer; email-bearing resource paths (/invites/<email>/accept), DC-048 details.userEmail attribution, and emails in non-sensitive body keys landed RAW — while the parallel unified-logger path has masked at every sink since DC-095. DC-110 closes the parity gap using the SAME canonical primitives (sa****@example.com): log() masks resource (maskEmailsInString) and deep-masks details (maskEmails) at the single write-point, covering middleware AND direct route calls. The security-event mirror now uses the masked entry.resource for target/message (judge round-1 fix-first: the raw parameter leaked emails into security-events.jsonl). logging.js exports maskEmails (export-only). maskEmails clones — caller details objects are never mutated. Judge: GLM-5.3 cold read (standing Sami authorization 2026-08-17; Codex quota dead until 2026-08-29). Round 1 (deleg_ebb83285) C fix-first — caught the event-store mirror leak. Round 2 (deleg_923f9076) after in-commit fix + mirror test: grade A, ship. Verdict URN urn:ump:mwtxoj6dbdq7bfeba3l2am2rgx34zjimcjde5f6sxz6tozsjbskq (GET readback verified: grade A, topic codex-judge-verdict). Deferred (judge-accepted): one-time scrub of historical raw-email lines in the live 16MB security-events.jsonl — queued follow-up. Tests: 123 suites / 2801 green (+6 DC-110 pins: resource+details mask, non-mutation, middleware e2e with *** survival, idempotence, no-email regression, masked mirror target/message). |
||
|
|
9322831f1b | fix(security): quoted local-part email mask — strip delimiter quotes, split on last @ (DC-109) [glm-grade=A] | ||
|
|
c429b8fdd7 |
fix(security): redact-on-rotate for error.log archive + README PII docs (DC-108) [glm-grade=B]
DC-095 masks emails at every live log sink; the rotated archive was the remaining belt-and-braces gap — any future sink that forgets masking would persist raw PII in error.log.1 for a full rotation cycle. appendErrorLog now scrubs the freshly rotated archive with the SAME canonical mask (sa****@example.com) via atomic rewrite (sibling .redact-<pid> temp, wx open preserving mode else 0600, fsync, rename). Scrub failure is caught and logged; the new error line is still appended. Stale crash-leftover .redact-<pid> temps are swept best-effort on every rotation. Also documents scripts/redact-log-pii.js usage in README (queue item e). Judge: GLM-5.3 cold read (deleg_e7cd7f8c), round-1 grade B ship — both nits addressed in-commit: stale-temp sweep (new 5th test pins it), quoted-local-part mask edge deferred as pre-existing DC-095 primitive. Tests: 122 suites / 2790 green (+5 DC-108 pins; was 2785 post-DC-107). |
||
|
|
5efacd11e8 |
fix(security): close rotateEncryptionKey crash window with in-process key rollback (DC-107) [glm-grade=A]
If the atomicWriteJSON of rotated credentials failed after rotateKey() had already persisted+cached the new key, the on-disk key could no longer decrypt the on-disk credentials.json (permanent loss on restart). Catch-path now restores the old key via new cryptoUtils.restoreKey() (canonical atomic-write, 0600, hex-validated) while still holding the proper-lockfile. Double-failure (rollback throws) is contained and the lock is still released. Hard-crash mid-rollback is covered by the existing .bak startup fallback. Judge: GLM-4.6 stand-in, round-1 grade A, 0 blocking, 1 LOW polish (folded). URN urn:ump:z2sz3x6abtcffpqyde2l2ssq34vy47t5h2gbmkr3wtmfxwkerinq Tests: 121 suites / 2785 green (6 consecutive runs pre-fold; suite re-run post-fold). |
||
|
|
eb2bab7a96 |
refactor(persistence): migrate credential-manager to canonical atomic-write util (DC-106) [glm-grade=B]
All three writeFileSync sites in the encrypted-credentials manager now delegate to src/utils/atomic-write.js atomicWriteJSON — rotateEncryptionKey save, _ensureFileExists bootstrap, and the _lockedUpdate commit under the proper-lockfile lock. A crash can no longer tear credentials.json mid-write: power loss through the old path could leave an empty/short file and silently drop every stored credential (DNS provider tokens etc). Lock-safety pre-study: proper-lockfile only stats its own sibling .lock dir, never the target file, so the rename swap cannot trip ECOMPROMISED. Test mock extended to the fd-level fs API (openSync/writeSync/ fsyncSync/closeSync/renameSync + fdMap/closedTmp state) so the canonical path is exercised end-to-end under the mock; write assertions moved from writeFileSync.mock.calls to destination-state reads. New DC-106 pins: wx+0600+fsync+rename discipline, plaintext-secret canary never on disk, fd-lifecycle order fsync->close->rename via invocationCallOrder, and atomic _ensureFileExists create at 0600. Eighth store migrated (DC-099..DC-105 preceded). Judge: GLM-5.3 cold read, round-1 B/ship (deleg_b3c038c2). URN urn:ump:vscequdet7wt5un7jhtl2nlg6cfkabsbazy5m2tjnyyguu5wstxa (readback verified: grade B, topic codex-judge-verdict). Sole finding is pre-existing and non-blocking: rotateEncryptionKey persists the new key before writing rotated creds (crash window) — queued as DC-107 follow-up. Full suite: 121 suites / 2783 tests green. |
||
|
|
b1464d9b85 |
refactor(persistence): migrate caddy-upstream-watcher state to canonical atomic-write util (DC-105) [glm-grade=A]
_saveState drops its private fixed-name .tmp + writeFileSync (no fsync) copy and delegates to src/utils/atomic-write.js atomicWriteJSON — same wx/fsync/rename/dir-fsync discipline as the six previously migrated stores (DC-099..DC-104). A crash can no longer tear caddy-upstreams.json (mute list + probe state): power loss through the old path could leave an empty/short state file and silently drop every mute; concurrent saves (60s probe loop vs setMuted) collided on the shared tmp name. Test mock extended with the fd-level fs API (openSync/writeSync/ fsyncSync/closeSync/unlinkSync + closedTmp stash) so the canonical path is exercised under the existing file-wide fs mock; fsState renamed mockFsState (jest.mock out-of-scope-variable hoist rule). New DC-105 pin: wx+fsync+rename required, no fixed .tmp, zero leftover tmp files, destination JSON complete with mute preserved. Judge: GLM-5.3 cold read, round-1 A/ship (deleg_5f57fcce). URN urn:ump:iuhgj3ajr5llshjasttsvelnptv6iqaek6xji5l3klcl72nseqdq Full suite: 121 suites / 2779 tests green. |
||
|
|
3c04a740e4 |
refactor(persistence): migrate bridge events file to canonical atomic-write util (DC-104) [glm-grade=A]
writeEvents() in scripts/stripe-license-bridge.js drops its private tmp+writeFileSync+rename copy (no fsync, Date.now() tmp names) and delegates to src/utils/atomic-write.js atomicWriteJSON (exclusive-create tmp, fsync, rename, parent-dir fsync, 0600). stripe-events.json is the Stripe webhook idempotency log - a torn write silently drops event-ids, so a Stripe retry re-runs delivery (duplicate license email; combined with a torn fulfillment record, a duplicate key mint). readEvents is JSON.parse-only, so the canonical serializer's trailing-newline drop is unobservable. Sole writer confirmed by grep. +2 DC-104 pins: full recordEvent->eventSeen dedupe cycle (0600/complete JSON/zero-leftovers) and 8-step back-to-back mutation parse-complete check. 121 suites / 2778 tests green. Judge: GLM-5.3 round-1 A (deleg_e8e1f9d0), URN urn:ump:frwh34wtlsunsymok6pzyujyehq6ikpvadypmsg7ndk7jjvzrp5a |
||
|
|
6f8fac142f |
refactor(persistence): migrate fulfillment-store to canonical atomic-write util (DC-103) [glm-grade=A]
writeState() drops its private tmp+writeFileSync+rename copy (no fsync, Date.now() tmp names, best-effort chmod) and delegates to src/utils/atomic-write.js atomicWriteJSON (exclusive-create tmp, fsync, rename, parent-dir fsync, 0600). A torn stripe-fulfillments.json could previously make a webhook retry mint a SECOND valid license key for an order that already has one. Both consumers (routes/billing.js, scripts/stripe-license-bridge.js) JSON.parse only - trailing-newline drop in the canonical serializer is harmless. Unused crypto require removed. +2 DC-103 pins: full lifecycle 0600/complete/zero-leftovers, and dual-instance interleaved writes (bridge+API file-IPC) with no tmp collisions. 121 suites / 2776 tests green. Judge: GLM-5.3 round-1 A (deleg_bd49a97b), URN urn:ump:layk7h326sqymcsh6tiusrs2sdbqdcx6ygvcuwwoxqn4rf7ycoua |
||
|
|
521b2f24a1 |
refactor(persistence): migrate share-store to canonical atomic-write util (DC-102) [glm-grade=A]
Drops share-store's private _atomicWriteJSON copy (pid+Date.now() tmp names, no fsync, no failure cleanup) in favor of src/utils/atomic-write.js: fsync'd same-dir tmp+rename, exclusive-create 0600, parent-dir fsync, cleanup-on-failure. Also fixes a latent bug in the same file: .share-secret signing-key persistence used plain fs.writeFileSync — a torn write would silently rotate the HMAC key on next boot, invalidating every outstanding share signature (links 404, no error logged). Now atomicWriteFile. Sole consumers parse JSON / trim(), so the dropped trailing newline on shares.json is unobservable. +2 regression tests pin 0600 on both files, complete content, zero tmp leftovers, and HMAC-still-verifies via getRaw. Judge: GLM-5.3 cold read, round-1 A, deleg_50d5c239 (41s, 4 calls). Verdict: urn:ump:ehblegyr5eko4hgyfylnrqk72hh3crmc43o2yt5ragfgvyc4zrdq Full suite: 121 suites / 2774 tests green. |
||
|
|
09d56fde2c |
refactor(persistence): migrate user-store to canonical atomic-write util (DC-101) [glm-grade=A]
Drops user-store's private _atomicWriteJSON copy (pid+Date.now() tmp names, no fsync, no failure cleanup) in favor of src/utils/atomic-write.js (DC-099 canonical: fsync'd same-dir exclusive-create 0600 tmp -> rename -> parent-dir fsync, cleanup-on-failure). All 3 persisted files routed: users.json, authorized-users.json, .bootstrapped sentinel. Sole consumers are JSON.parse readers — dropped trailing newline unobservable (judge verified repo-wide). +2 store-level regression tests pin 0600 / complete JSON / no temp leftovers across all three files, incl. the bootstrap path writing three files back-to-back in one login. Judge: GLM-5.3 cold read, round-1 A, deleg_f632f05c. Verdict: urn:ump:5fivvveqhcbkl6os4dhidchkvjjzbvi7rgj6znaovp6bfnvmcmsq Full suite: 121 suites / 2772 tests green. |
||
|
|
3742e2658d |
refactor(persistence): migrate invite-store to canonical atomic-write util (DC-100) [glm-grade=A]
Drops invite-store's private _atomicWriteJSON copy (pid+Date.now() tmp names, no fsync, no failure cleanup) in favor of src/utils/atomic-write.js: fsync'd same-dir tmp+rename, exclusive-create 0600, parent-dir fsync, cleanup-on-failure. Sole format consumer is a JSON.parse reader, so the dropped trailing newline is unobservable. +2 store-level regression tests pin 0600 / complete JSON / no temp leftovers under burst mutations. Judge: GLM-5.3 cold read, round-1 A, deleg_7177d506. Verdict: urn:ump:4kwenywtf3nx2mokobvuzrhklk2xigmpyea6nvqdyhv7icnflpkq Full suite: 121 suites / 2770 tests green. |
||
|
|
80a82c4cae |
refactor(persistence): canonical atomic file writer + notifications.json crash-safety (DC-099) [glm-grade=B]
- src/utils/atomic-write.js: single shared tmp+fsync+rename writer (exclusive-create 0600, unique tmp names, cleanup-on-failure, best-effort parent-dir fsync after rename for swap durability) - notification-manager: both write paths (load-time canonicalization write-back + saveConfig) converted from plain writeFileSync — a crash mid-write can no longer truncate notifications.json - DC-097/098 test seams migrated to the atomic path; new suite pins syscall discipline (order, wx flags, tmp naming, error cleanup, dir-fsync swallow) - 121 suites / 2768 tests green - Judge: GLM-5.3 cold read grade B/ship (deleg_23f7abad); polish items folded: dir-fsync added, header copy-count corrected. Remaining: migrate invite/user/share-store private _atomicWriteJSON copies as they are touched (queued). URN: pending (recorded post-commit) |
||
|
|
8d42eae6ac |
feat(maintenance): one-shot PII redaction tool for pre-DC-095 log files (DC-098) [glm-grade=A]
- scripts/redact-log-pii.js: atomic in-place email redaction reusing the canonical DC-095 masker (no second regex), dry-run/keep-raw modes, dir walk with skip-set, post-verify (exit 2 if raw addresses remain). - src/utils/logging.js: export EMAIL_RE/maskEmailAddress/maskEmailsInString (additive; no logger behavior change). - __tests__/redact-log-pii.test.js: 11 tests (shape, idempotence, clean-untouched, dry-run, keep-raw, skip-set, passthroughs, exit codes). - Judge: GLM-5.3 cold read, round-1 A/ship (deleg_4e684a18), URN urn:ump:7y2q5upoht7xq2mhlum764y2h36qpsgufyqsgx4cbijfpmfpmrcq. - Suite: 120/120 suites, 2751 tests green. |
||
|
|
3ccd00d1a1 |
fix(notifications): persist canonicalized config on load — legacy keys no longer stale on disk (DC-097) [glm-grade=B]
_loadConfig canonicalized legacy spellings in memory only (DC-092); the on-disk notifications.json kept email user/pass, camelCase event keys and string secure until the next explicit UI save — i.e. forever on installs that never open the settings page. - _persistCanonicalForm(): after the defaults merge, re-serialize and write back only when the bytes differ; idempotent on subsequent loads. - Best-effort: write failures (read-only mount, EACCES) warn and continue — the in-memory config is already correct; constructor never throws. - No secrets in new log lines; JSON.stringify(this.config) same as saveConfig. Judge notes (non-blocking, GLM-5.3 cold read): unknown top-level keys are now dropped from disk at boot (pre-existing merge-drop semantics, previously deferred to next UI save); write is non-atomic, matching saveConfig. Verdict: urn:ump:rchawiu427idev5mrettlxkyw2rcnwqpegiolqmzbc5ygc277u5q Tests: +5 (__tests__/notification-config-writeback-dc097.test.js) — canonical rewrite, idempotence, clean-file-untouched, EACCES no-throw, fresh-install no-write. Full suite 119 suites / 2740 tests green. |
||
|
|
bb59595d6d |
fix(config): monitoring.public gate was dead 4 ways — live gate + schema + dedupe (DC-096) [glm-grade=A]
The documented hardening option for exposed deploys (monitoring: {public: false}
in config.json / MONITORING_PUBLIC env) never worked:
1. applyConfigFields dropped the monitoring key entirely
2. monitoring missing from config-schema KNOWN_KEYS (Unknown-key warnings)
3. MONITORING_PUBLIC frozen at mount + re-required singleton instead of injected dep
4. PUBLIC_ROUTES had unconditional duplicate entries defeating the gated spread
- site.js: copy monitoring through; drop dead write-only siteConfig.caName
- config-schema: +monitoring key, validateMonitoring (public must be boolean);
remove never-written typo-footgun keys setupCompleted/setupMode (git -S: zero writers ever)
- middleware: live isMonitoringPublic() (env > config > default public), per-request
gate via monitoring:true flag, remove duplicate unconditional route entries
- default unchanged (endpoints stay public — System Overview widget)
Tests: +11 (__tests__/monitoring-public-gate-dc096.test.js); suite 118/2735 green.
Judge: GLM-5.3 cold-read A (deleg_98aba845), URN urn:ump:zgvtskqljurasdagc632atnybb2p6gakk4cxcmjd3i4ivy7rvwta
|
||
|
|
83ef84d218 |
feat(logging): central email PII masking across all log sinks [glm-grade=A]
DC-095: mask email addresses at every logger output choke point so raw
PII never reaches stdout/stderr, error.log, or audit-log.json regardless
of what a call site interpolates — msg strings, data payloads, error
messages/stacks, audit details, and error.log request lines (path/UA).
- Masked shape sa****@domain matches AuthProvider.maskEmail (UI-consistent)
- Bounded-quantifier regex: local {1,64} (incl. quoted local-parts),
domain {0,253}, TLD {2,24} — adversarial 40KB string 3.3s -> 17ms,
hostnames/versions/docker-refs untouched, idempotent under re-mask
- memo-Map recursion: DAG shared references get the same masked clone
(WeakSet seen-guard leaked the raw original on 2nd reference); cycles
resolve to in-progress clone
- Non-plain objects with own enumerable props cloned proto-preserving
(Object.create) so class-instance email fields are masked; Date/RegExp
pass through
- sanitize(): audit details mask email substrings in non-sensitive keys
(invite/auth POST bodies no longer land raw in audit-log.json)
- 18-test suite covers sinks + adversarial judge findings (ReDoS timing,
DAG, quoted locals, instances, request-line path/UA)
Judge: GLM-5.3 cold-read stand-in (Codex quota-dead until 2026-08-24,
substitution authorized by Sami 2026-08-17). Rounds C -> C -> A.
Verdict: urn:ump:zorj7vcrnw2t2jhhcp2g6wz4simvhyjdqwu2mjsifxlkzb2dwjmq
Suite: 117 suites / 2724 tests green.
|
||
|
|
97672f7e74 |
[glm-grade=A] fix(notifications): un-gate 7 dead emitters + repair legacy 4-arg send shape (DC-094)
Two silent-death defects in one class: 1. Seven emitters absent from DEFAULT events, so the send() gate (config.events[canonical] !== true) dropped them on every install: ssl-cert-expiry, dns-propagation, drift-detected, dependency-restart-complete/-failed, recipeRemoved, workflow. All now default ON; dependency-restart spellings fold onto one canonical toggle; recipeRemoved aliases to recipe-removed in both the manager and route alias maps. 2. Nine call sites used a legacy 4-arg send(event, title, message, type) against the 3-arg signature: the message string landed in the type slot (Discord embed color fell back to info-blue, history.type wrong) and providers received the TITLE as the body - deploy-failure notifications carried no error text at all. Fixed at source (9 sites) plus a type-guarded shim in send() for external legacy callers. Also: explicit data.title now flows to ntfy Title header, email subject, Discord embed title, and history; settings UI gains 9 event toggles (separate Backup Complete/Failed) with defaults-on semantics. Tests: +24 (new DC-094 suite: defaults, gate pass-through, alias folding send-time and load-time, stored-config inheritance, shim body/title/ color/subject/history, type-guard, 3-arg no-regression); 4 assertions in bundled-workflows-health-check updated from the old 4-arg mock contract to the canonical shape (same behavior asserted). Full suite 116 suites / 2706 tests green. Judge: GLM-5.3 cold read via delegate_task (deleg_8a7cedd0), grade A, zero blockers; 2 polish items (Backups toggle conflation, shim type-guard) folded into this commit. Verdict URN: urn:ump:uyipjjwdjqjy3alvceqxvlucrymd7hnsben5udpp2bqycoh3l2la |
||
|
|
5add962178 |
[glm-grade=B] fix(auth): /auth/me hotfix — isValid not isSessionValid + guard (DC-093 r2)
Live verification of
|
||
|
|
4125d7a4e1 |
[glm-grade=B] fix(auth): mount /auth/me on single-user installs — kill the 60s 404 log storm (DC-093)
The frontend admin panel polls GET /api/v1/auth/me on every dashboard load and every 60s per open tab. /me lived only in the DC-048 admin router, mounted only when email auth is enabled — so every single-user install (the default) answered 404 and logged a DC-404 ERROR + stack once per minute per open tab. Verified live in production logs. - routes/auth/index.js: /auth/me now ALWAYS mounted. Multi-user + req.user returns the stored profile (mode:'multi'); otherwise the legacy single-operator response (role:'admin', isAdmin:true, legacy:true, mode:'single'). Session-gated — NOT added to PUBLIC_ROUTES, so unauthenticated polls get a clean 401. - Frontend behavior unchanged: attachTrigger requires me.user.role === 'admin', and user stays null in single-user mode, so no Admin button appears on single-user installs. - openapi.yaml /api/v1/auth/me (200/401) now matches reality. - +6 tests (route present both modes, response shapes, PUBLIC_ROUTES absence, DC-048 admin-mount invariant, HTTP-level dispatch reach). Judge: GLM-5.3 cold read, grade B ship (verdict URN recorded in STATE.md); polish note (HTTP-level mount-order test) folded in same commit. Full suite 115 suites / 2687 tests green. |
||
|
|
7e4ee60dcf |
[glm-grade=B] fix(notifications): repair UI/API contract drift — SMTP auth, event gate, test button (DC-092)
Three user-facing notification features were silently dead from contract
drift between the settings UI and the backend:
1. SMTP auth never applied: the UI sent email.user/email.pass while the
manager read username/password. Route now normalizes aliases; legacy
config files canonicalize at load.
2. Event toggles were cosmetic: UI sent camelCase keys (containerDown),
the send() gate read kebab-case ('container-down'). EVENT_ALIASES now
folds every known spelling (manager gate, route store, legacy files);
UI sends and reads canonical keys.
3. Deploy/auto-restart notifications always dropped: deploy-success/
deploy-failed/auto-restart were missing from DEFAULT events, and
send('test') was itself gated -> the Test button was a no-op. Defaults
added; 'test' bypasses the gate.
Also: strict boolean contract (string 'false' for secure/enabled rejected
— previously coerced truthy, silently forcing TLS), SMTP port bounds,
non-destructive credential merge (blank password no longer clobbers the
stored one), GET /config returns port/secure/to/username + hasPassword
(password never returned), full form prefill + keep-hint placeholder.
Verified live pre-fix: send('test')/'deploymentSuccess'/'auto-restart'
all returned 'not enabled'. Post-fix: +20 tests (route + manager),
full suite 2641/2641 (112 suites).
Judge: GLM-5.3 cold read via delegate_task (deleg_ad765d63), grade B,
ship, zero blockers; judge independently ran the DC-092 suites (38/38).
Polish #1 (canonical event for provider titles) folded in. Remaining
emitters with the same gate-miss class (recipeRemoved, workflow,
ssl-cert-expiry, dns-propagation, drift-detected, dependency-restart-*)
noted for a follow-up tick.
|
||
|
|
5e60c27f2b | [grade=A] Harden server-managed license renewals | ||
|
|
ddbea0a040 |
[glm-grade=A] fix(config): teach schema KNOWN_KEYS the licenseBackup/_version writer keys (DC-091)
Every startup logged two false-positive 'Unknown config key — possible typo?' warns: licenseBackup (written by src/managers/license-manager.js:510 activation persistence) and _version (stamped by src/config/migrations.js). Both are first-party writers the validator was never taught about (DC-091). - config-schema.js: add both keys to KNOWN_KEYS with a source-of-writes comment - config-schema.test.js (new): 5 regression tests — live production config keyset validates with zero unknown-key warnings, writer keys never warn, genuine typos still warn (exact string), license/licenseBackup sync guard, _version recognized at every migration value Verified: full jest suite 111 suites / 2621 tests green (baseline 110/2616). Warns reproduced in live container logs 2026-08-22T23:53:54Z; live config.json contains both keys (licenseBackup activation, _version 2). Judge: GLM-5.3 cold read via delegate_task (deleg_30e52384, 36s) — grade A, ship. Verdict URN: urn:ump:ermkvz6ifbp5svga5cnapv5jhm7c5b7qdbwjjerfpxbwcrolm2za (readback verified) Codex quota-walled until 2026-08-29; GLM stand-in per Sami 2026-08-17 directive. |
||
|
|
88f1d4a414 |
[glm-grade=A] fix(monitoring): DC-090 outage incidents follow displayed hysteresis status
checkForIncidents compared raw probe transitions while the dashboard badge (DC-086) follows post-hysteresis displayed status. A single raw down blip between two ups opened AND resolved a critical outage incident; a suppressed up blip during a real outage resolved it early. Incidents now open/resolve on displayed-vs-displayed transitions; previousDisplayed=null keeps legacy raw semantics for direct callers. 6 new parity tests + legacy checkService test moved to a 4-probe chain. Suite 2616/2616 (110). Verdict: urn:ump:yc5rdlnmnmhch5audc5fifgbt6d7moi2stqfs5vidsbh6x6zkvgq |
||
|
|
6732a1e1df |
[glm-grade=B] fix(auth): DC-089 mask invite/user emails in server logs
Two log sites in routes/auth/admin.js wrote raw email PII to the server log: the SMTP-unconfigured 'auth-invite-send' warn and the 'invite accepted, user created' info. Both now route through AuthProvider.maskEmail() with a '[unmaskable-email]' sentinel fallback (never the raw address). Two regression tests assert the raw address is absent from log meta and the masked form present. Response contract unchanged (full email still returned to the authenticated admin). Judge: GLM-5.3 cold read (deleg_f0896de3), grade B / ship / zero blockers; polish notes folded in. Verdict URN: urn:ump:jd2htpwq76ni6bapj3vxjpvypfdnoc4argqbzpcerutmh7u5khea Full suite: 2610/2610 (109 suites). |
||
|
|
3ccf66754a |
[glm-grade=B] fix(monitoring): DC-088 removeService generation tombstones + incident closure
- serviceGenerations no longer leaks entries: removeService deletes the live entry and records a TTL'd (10min) tombstone swept by cleanupHistory - monotonic instance-wide generationSeq prevents generation reuse across remove->re-add cycles (ABA) and supersedes tombstones on re-configure - _isStaleCapture(): presence-aware stale check — live entry must match exactly; no entry is stale only under a higher-generation tombstone (preserves correct behavior for disk-loaded never-configured services) - catch path increments consecutiveFailures only after the stale check, so a late-rejected probe cannot resurrect state for a removed service - open incidents for a removed service close via the standard resolve path (resolvedBy=service-removed, WS/SSE incident-resolved broadcast) - 6 regression tests; full suite 2608/2608 green Judge: GLM-5.3 cold read (Codex stand-in), verdict B/ship, zero blockers |
||
|
|
f2285a2550 |
test(api): DC-087 hermetic caddy-admin health mirrors + file-level raw-fetch guard [glm-grade=B]
Two mirrored health-handler test suites (health-endpoints, health-probe-aliases) probed the Caddy admin API with raw Origin-less native fetch. On the prod host the adversarial cron runs the full jest suite every 30 min against a live Caddy admin with enforce_origin: 12 journal 403 lines per run (~700/day of 'client is not allowed to access from origin' spam) while tests stayed green. - Mirrors now call fetchT (byte-identical to src/app.js:930 probe) with fetchT jest.spyOn-mocked at buildApp scope; caddyOk-configurable in both suites - New guard test in utils-http-caddy-admin-origin.test.js: any __tests__ file pairing a raw await-fetch with a Caddy-admin token (:2019|adminUrl| CADDY_ADMIN) fails the suite — file-level pairing catches the historical cross-line drift shape a call-window regex missed - DC-087-ALLOW-RAW-FETCH comment escape hatch (raw-text marker, guard file never self-exempts, skips logged to jest output) Judge: GLM-5.3 cold read via delegate_task deleg_4d384dea (round 1 C -> round 2 B, zero blockers, polish folded). Verdict URN: urn:ump:azrv2xp72koiwi5r4yb6ureu4aqqloqq64sgmftsajh6ci2mzj2q Mutation probes: historical drift reintroduction -> guard red; hatch marker -> skipped+logged; restore -> 33/33. Full suite 2603/2603. |
||
|
|
628bbe32f6 |
[glm-grade=A] fix(monitoring): DC-086 round-2 — probe/config race hardening + env parse + incident compare
Round-2 folds the judge-round fixes into DC-086: - serviceGenerations map: checkService captures the config generation at entry and re-validates it before ANY state write (success + error paths). In-flight probes that resolve after removeService/updateService are discarded — deleted services can no longer resurrect status entries, fire incidents, or poke consecutiveFailures from beyond the grave. - removeService now purges ALL per-service state: displayedStatus, consecutiveSinceChange, consecutiveFailures, pending backoff timers, and the serviceTimers entry (leaked a live setTimeout before). - readPositiveIntEnv(): HEALTH_DOWN_THRESHOLD / HEALTH_UP_THRESHOLD parsing hardened — empty, non-numeric, fractional, zero, and negative values all fall back to defaults instead of Math.max(1, NaN)=NaN. - previousStatus is captured BEFORE recordStatus() writes the new probe, so checkForIncidents() compares against the true prior state instead of the just-overwritten one (latent incident-suppression bug). - Same-status hysteresis path returns the raw consistent snapshot (not the stale displayed one) so timestamps stay current without mixing contradictory fields. - Tests: +14 (86 total across the two suites). New coverage: streak reset on agreement, malformed env fallbacks (each.of not-a-number/0/ -2/1.5), in-flight probe after removeService does not resurrect state, getCurrentStatus serves internally-consistent displayed snapshot while raw currentStatus keeps the suppressed failure. Full suite 2601/2601. Judge: GLM-5.3 cold-read via delegate_task (deleg_c9fd5900 task-0), grade A round 1, zero blocking issues. Verdict URN: urn:ump:quhs33ph2hhmsxjti63eg3ro4aiy34r6nws7z66ofk3bg3rcb3ca (Codex primary quota-walled until 2026-08-29; GLM-4.6 direct 401; stand-in chain per codex-as-judge SKILL.md, Sami 2026-08-17.) |
||
|
|
f8b99f9b5a |
DC-086 service-status flicker fix — asymmetric hysteresis
Dashboard badges perpetually flip green/red for a few seconds at a time, never stable. Root cause: health-checker emitted 'status-check' on every probe (every 30s) and dashboard-ws forwarded every one as 'status-change' to the browser with no diff; live-events.js then unconditionally called setBadge(). A single transient 5xx (Caddy reload, container restart, TLS handshake blip) flipped the badge and the next green probe flipped back. Fix: _computeDisplayedStatus applies asymmetric hysteresis — DOWN_THRESHOLD (default 2, env-tunable HEALTH_DOWN_THRESHOLD) consecutive probes that disagree with the displayed 'up' state flip to red; UP_THRESHOLD (default 1, HEALTH_UP_THRESHOLD) flips back to green. History and consecutiveFailures still record every raw probe so postmortem analysis is unchanged. Only the SSE broadcast is filtered. getCurrentStatus now returns the displayed status so a page reload shows the same badge as the live stream. 10 new tests cover first-emit, same-status-dedup, the actual flicker bug (one-down-then-up keeps green), two-down flips red, one-up recovers fast, long-steady-green produces exactly one emit, and env-var tuning. All 63 existing health-checker tests still pass. Full suite: 2484/2484. |
||
|
|
eab2b00b13 |
DC-085 link-first invite — Discord-style share it however you want
Flip POST /api/v1/auth/admin/invites default to no email; always return the link. Operators copy + share via iMessage/WhatsApp/SMS/Signal/Telegram/ Discord/paste-in-email. Email becomes an opt-in checkbox (was the default). Add shareText field with pre-formatted message for one-tap paste. Stop logging raw invite URLs to error.log when SMTP is unconfigured (was just a dev fallback — link is now in the response). Frontend flips the checkbox default to unchecked and renders shareText + native share sheet button (navigator.share) alongside the raw copy-link button. 9 new tests covering default-no-send, link-always-returned, shareText-shape, opt-in SMTP send, failed-SMTP-no-leak. Full suite: 2474/2474 + 9 new = 2483. |
||
|
|
84edb035e3 | [grade=B] feat(auth): onboard missing credentials into encrypted vault | ||
|
|
d313b1e872 | [grade=B] fix(auth): reuse valid session for cross-host SSO | ||
|
|
7e68955e66 |
[glm-grade=A] fix(share): public-endpoint input hardening + rate limit (DC-083)
Public share endpoints accept untrusted fields. Pre-fix code used bare
type checks (email.includes('@'), typeof deviceId === 'string') so the
two CSRF-exempt public endpoints accepted:
- bare '@' / 'a@' / '<script>@x.c'
- 10MB email strings (data/shares.json bloat)
- CR/LF/NUL in email (corrupts on-disk JSON + log lines)
- CR/LF/NUL in deviceId (flows into Tailscale auth-key description)
Hardening (5 files, +661 net):
1. routes/share.js + src/security/share-store.js: shared validators
- validatePublicEmail(raw): charset (a-z0-9._%+-@), 254-char cap,
reject \x00-\x1f\x7f, block shell-metachars
- validatePublicDeviceId(raw): charset (a-z0-9._:-), 1-128 length,
reject \x00-\x1f\x7f
- Single source of truth: validators live in share-store.js, exported,
imported by routes/share.js (drift-eliminated)
2. Routes that were 'email.includes(@)' now use validator. Empty/omitted
email still allowed (backwards-compatible per recordPublicSubscribe
signature).
3. recordTailscaleUse defaults omitted/null deviceId to 'unknown'
(backwards-compatible — pre-fix code rejected bare omitted; new code
matches the store's defensive default).
4. constants.js: RATE_LIMITS.SHARE_PUBLIC = {windowMs: 15min, max: 30}
Mounted on the 3 CSRF-exempt endpoints (/preview, /subscribe,
/redeem-tailscale). 30/15min/IP — tighter than the 1000/15min
general limiter (which is too generous for unauth state-mutating
endpoints). Falls back to no-op in test envs.
5. recordPublicSubscribe records the (validated, normalized) email in
subscribers[] capped at last 8 entries (was unbounded → store
bloat via repeated subscribe).
Test coverage (38 new tests in __tests__/share-dc083.routes.test.js + 3
in __tests__/share-routes.test.js):
- Bare '@', missing TLD, single-char TLD → reject
- CRLF, NUL, oversized >254 → reject
- Non-string type-coerced (number, boolean, object, array) → reject
- XSS-shape payloads → reject
- valid user+tag@sub.domain.io + nodekey:... → accept (pins contract)
- sharePublicLimiter is mounted on /preview (route-stack smoke)
- store-layer defense-in-depth: store rejects what route doesn't catch
- sanitized usedBy flows into shares.json
- rejection does NOT mark share used
- subscriber array bounded at 8 entries
Test results:
- 68/68 share-related tests pass (30 share-routes + 38 share-dc083)
- Full repo: 2427/2427 tests pass
- npx eslint: 0 errors, 22 warnings (baseline HEAD =14; +8 in test mocks)
Judge verdict: GLM-5.3 round-2 grade A. Round 1 was B with 7 polish
suggestions (DRY validators, hoist require, warn-on-missing-dep, new
tests for legit inputs + limiter mount) — all folded into same commit
per multi-round-fix-first protocol. Zero blocking issues.
Threat model: the 2 POST endpoints mutate shares.json + Tailscale auth
descriptions. Pre-fix was effectively 'input trust boundary = NONE'.
Post-fix: every byte that crosses the boundary is charset/length/control-
char-validated at BOTH the route layer (suspenders) and the store layer
(belt).
|
||
|
|
089f5d2902 |
[glm-grade=A] fix(update-manager): compose-prefixed image names probe <project>/<service> not library/<project>-<service> (DC-082)
Pre-fix: dashcaddy-dashcaddy-api:latest was normalized to library/dashcaddy-dashcaddy-api before probing Docker Hub. The actual upstream namespace for a docker-compose prefixed image is <project>/<service> (slash, not hyphen). Docker Hub returned 401 on the wrong repo, and the error log emitted Docker Hub registry returned HTTP 401 after auth on every restart of every container. Fix: 1. _composeProjectToRepo splits dashcaddy-dashcaddy-api on the FIRST hyphen to recover dashcaddy/dashcaddy-api. Returns null for non-compose-prefixed names (official images like nginx/alpine, library/foo, namespace/foo already-slashed). 2. _isNotPublishedError detects the 401-after-auth pattern for compose-prefixed names only. Steady-state for locally-built images that aren't published. 3. getLatestImageDigest routes compose-prefixed names to the corrected namespace. Routes already-namespaced names directly. Falls back to library/ for the Official Image path. 4. Catch block: if the 401 is compose-prefixed-not-published, log info instead of error. Real auth failures on legitimate images still log as error. 17/17 tests pass in 1.27s. Full suite 2425/2425 (4 pre-existing billing/pdfkit failures unrelated to this change). GLM stand-in verdict URN: urn:ump:7rhk7keukv3zrx654gbckxaycnuwm4agduf37creqbsoauszopoa |
||
|
|
0e7bb97129 |
[glm-grade=A] fix(log-insights): wire dispose to /app/data paths + bound keepDays (DC-081)
Pre-fix, the dispose endpoint + storage info block in dashcaddy-api/routes/log-insights.js
HARDCODED /opt/dashcaddy/dashcaddy-api/data/audit-log.json and
/opt/dashcaddy/dashcaddy-api/data/security-events.jsonl, which DO NOT EXIST in the
production container (verified 2026-08-19 01:42Z: /app/data/audit-log.json = 318 KB,
/app/data/security-events.jsonl = 15 MB, /opt/... = ENOENT). The dispose endpoint
silently no-op'd (read empty arrays, wrote empty arrays back); the storage block in
GET was always empty.
Also: parseInt(req.body.keepDays) || 30 accepted negative numbers. keepDays = -1000
produces a cutoff +3 years in the future, then the filter e.timestamp < cutoff
deletes 100% of the audit log. Operators must not be able to wipe forensic context
with a typo.
Fix:
* _resolvePaths() uses process.env.AUDIT_LOG_FILE || path.join(platformPaths.dataDir, 'audit-log.json'),
matching the canonical resolution in src/security/audit-logger.js and src/security/event-store.js.
Both GET + POST share the resolved paths (single source of truth).
* _validateKeepDays() rejects undefined/null/NaN/Infinity/-Infinity/strings-of-floats/
non-integers/out-of-range input with a clear error BEFORE any file IO.
Allowed: integer in [1, 3650] (1 day .. 10 years).
* POST /log-insights/dispose now requires { keepDays: integer 1..3650, confirm: true }.
Preview is read-only. Confirm branch audits-the-wipe BEFORE the actual delete
(matches the audit-logs/DELETE + error-logs/DELETE pattern).
* Atomic write for audit-log.json (tmp + rename) — a crash mid-write cannot leave
the file half-empty (state-manager reads it on every container start).
Tests (23 new, dashcaddy-api/__tests__/routes/log-insights.routes.test.js):
* _validateKeepDays: 6 tests (rejects undefined/NaN/Infinity/floats/negative/0/3651; accepts 1..3650; coerces numeric strings).
* _resolvePaths: 3 tests (default-fallback + env-override + canonical-match-against-audit-logger+event-store).
* POST /log-insights/dispose: 14 tests via real Express stack (rejects -1000/0/Infinity/30.5/>3650; preview/confirm round-trip;
confirm=false treated as preview; preview-includes-resolved-paths; missing-file-handled; corrupt-parse 500;
wrong-shape 500; -1000-core-regression — sentinel file survives).
GLM-5.3 round 1: A.
|
||
|
|
99ec6ebc53 |
fix(tailscale-admin): harden apiToken/tags/description validation (DC-080) [glm-grade=B]
DC-080 round-1 GLM-5.3 judge verdict: B. Round-2 polish folded into same
commit per multi-round fix-first protocol: tighten tag regex to require
non-empty name after 'tag:' (matches Tailscale spec), drop dead
`module.exports.createApp = null` line.
THREAT MODEL
Pre-fix, /api/v1/tailscale/* and /api/v1/tailscale/admin/* (TOTP-gated)
had inconsistent checks on caller-supplied input. Three coupled gaps:
(a) PUT /settings validated `apiToken.startsWith('tskey-api-')` but
had NO length cap — body-parser limit was the only ceiling. A 1 MB
string starting with `tskey-api-` would be `.trim()`-ed, sent to
Tailscale's /devices endpoint, and waste server-side CPU on a
request that will always 401.
(b) POST /settings/test accepted `apiToken` from the body with NO
validation at all. The PUT route's prefix check did NOT extend to
this path. An operator could submit arbitrary junk and the
container would still call /devices on the Tailscale API with it
(DoS-reflection + fingerprint timing for an attacker probing
whether this API token format is accepted).
(c) POST /admin/keys validated `tags` as Array but NOT per-element
type — `tags: ['tag:guest', null, 123, {injection: true}]` would
be forwarded to Tailscale verbatim. Tailscale's API is JSON-strict
and would 400 the request, but the bad shape reached the wire.
Similarly `description` had no length cap (Tailscale caps at 120
chars per their docs).
All three are gated by TOTP — this is a logged-in-operator / phished-
session threat surface, not anonymous-unauth. The fix is defense-in-
depth: a bug in the auth path (TOTP bypass, session theft, future route
handler trust-boundary drift) should not turn these endpoints into a
`submit anything and forward to Tailscale` relay.
FIX 1 — Shared validators (round-1)
- `_validateApiToken(token)`: typeof string check, prefix required,
length cap 256 chars. Catches empty/null/non-string AND oversize.
- `_validateTags(tags)`: undefined/null allowed (optional field),
Array.isArray check, max 32 entries, per-element string check,
per-element length cap 64 chars, regex
`/^tag:[a-z0-9][a-z0-9_-]{0,62}$/` (round-2: requires non-empty
name after `tag:` per Tailscale spec).
- `_validateDescription(description)`: undefined/null allowed, string
type check, length cap 120 chars (matches Tailscale's documented cap).
All three return null on success or an error string on failure. Route
layer maps to 400 via `errorResponse`. Validators exported via
`module.exports._validators` for direct unit testing (otherwise
unreachable from outside the factory closure).
FIX 2 — Endpoint wiring (round-1)
- PUT /settings: replaced inline `!startsWith('tskey-api-')` check with
`_validateApiToken(token)`. Single source of truth for the rule.
- POST /settings/test: added `_validateApiToken(token)` guard BEFORE
calling `client.setApiToken(token)`. The body is optional, so the
guard is skipped when no token is provided (uses stored token path).
- POST /admin/keys: replaced `Array.isArray(opts.tags)` shallow check
with `_validateTags(opts.tags)`, plus `_validateDescription(opts.description)`.
Old code already validated `expirySeconds`; that stays.
FIX 3 — Round-2 polish
- TAG_KEY_RE: `/^[a-z0-9][a-z0-9:_-]{0,63}$/` → `/^tag:[a-z0-9][a-z0-9_-]{0,62}$/`.
The old regex accepted `tag:` (empty name), which Tailscale's API
rejects. New regex requires `tag:` prefix and ≥1 alphanumeric name
char followed by [a-z0-9_-]{0,62} — total length up to 67 chars, well
within Tailscale's documented 15..63 char tag length.
- Removed `module.exports.createApp = null` vestigial line — the file
only exports the factory function and the _validators bag.
TESTS (29 original + 16 new = 45 in this suite)
- 4 PUT /settings new: length cap, non-string type, prefix round-trip
(existing 'starts with' tests already passed), plus the original
6 (4 pre-existing PUT tests stay green).
- 4 POST /settings/test new: prefix rejection, length cap, stored-token
path with empty body still works.
- 4 POST /admin/keys new: null/123/object entries rejected, uppercase /
whitespace / CRLF rejected, description length cap, canonical
lowercase `tag:server` accepted.
- 4 direct validator unit tests: validateApiToken (5 cases incl. cap-edge),
validateTags (8 cases incl. round-2 bare-'tag:' rejection), validateDescription
(3 cases incl. cap-edge), constants-export surface.
All 45 tests pass on DNS2 (verified). Full repo suite unchanged: 2351/2351.
|
||
|
|
a7260436d1 |
fix(disaster-recovery): stage Caddyfile + close path-traversal in assets/themes (DC-079) [glm-grade=A]
DC-079 2-round GLM-5.3 judge verdict: round1=C (blocking path-traversal
in assets/themes) → round2=A. 20/20 tests in routes/discover-disaster
(8 original + 12 new). Full repo: 2351/2351 (4 pre-existing billing
pdfkit failures unchanged).
THREAT MODEL
POST /api/v1/disaster/restore was the ONLY endpoint in the route tree
that wrote directly to process.env.CADDYFILE_PATH (=/caddyfile in
container = /etc/caddy/Caddyfile on host via start.sh:161 bind-mount).
Pre-fix: an authenticated dashboard operator POSTed
{caddyfile: '<attacker-controlled-string>'}
and the handler called fsp.writeFile(caddyfilePath, snapshot.caddyfile),
overwriting the live Caddyfile immediately. Caddy reads this file on
every reload (ACME renewal, health probe, admin API touch), so the
attacker-controlled content executes as Caddy config directives:
- import /etc/caddy/<anything-caddy-can-read> (content theft)
- admin off (lock out admin API)
- reverse_proxy to attacker IPs (Caddy becomes a pivot)
- acme_ca override to attacker CA (rogue cert issuance)
- log to attacker-writable paths (DoS/escape)
This bypassed the CLAUDE.md hard rule 'Caddyfile edits must use
caddy-apply' (validates + reloads + git-commits atomically).
FIX 1 — Caddyfile staging (round-1)
- New validateCaddyfileContent(): type check, non-empty check,
512 KiB byte cap (defense-in-depth below the 1 MB body-parser limit),
FORBIDDEN_IMPORT_RE rejects directives with absolute paths,
../-escape, ~/, or URL-encoded payloads.
- POST /disaster/restore now writes to <dataDir>/disaster-staged/
Caddyfile.candidate (atomic write + rename), NEVER to caddyfilePath.
- Response includes caddyfileStaged[{file, stagedPath, action: 'awaiting
caddy-apply', livePath}] and a DC-079 warning instructing the operator
to run `caddy-apply <reason>` to validate + reload + git-commit.
FIX 2 — assets/themes path-traversal (round-2 BLOCKING)
GLM round-1 caught a parallel vector: snapshot.assets[name] and
snapshot.themes[name] are user-controlled JSON keys flowing into
path.join(assetsDir, name) and path.join(themesDir, name). An attacker
could POST {assets: {'../../etc/caddy/Caddyfile': '<base64-evil>'}}
and overwrite the live Caddyfile via the dataDir bind-mount, fully
bypassing Fix 1.
- ASSET_KEY_RE = /^[a-zA-Z0-9._-]+$/ + ASSET_PATH_TRAVERSAL_RE catch
slashes, leading '..', and absolute-path keys.
- THEME_NAME_RE = /^[a-zA-Z0-9._-]+\.json\$/ additionally forces
.json extension and no slashes.
- assertSafeAssetKey/assertSafeThemeName helpers throw on invalid input.
- Both restore loops now: assert → path.resolve(dir, name) → containment
check (resolved must start with path.resolve(dir) + path.sep) → write
to resolved (never the raw join).
TESTS
12 new tests in __tests__/routes/discover-disaster.routes.test.js:
- staging: live sentinel unchanged, candidate at expected path
- rejects: non-string, empty, oversize, 3 forbidden-import variants
- assets: path-traversal key, absolute-path key
- themes: path-traversal name, no-extension name
- back-compat: no caddyfile field succeeds without staging
|
||
|
|
a4e4b24732 |
fix(update-manager): force IPv4 + per-request timeout + transient-only retry on registry digest probes (DC-078) [glm-grade=A]
The per-hour checkForUpdates() loop called Docker Hub / ghcr.io without family:4, without a hard request timeout, and without retry on transient network errors. On DNS2 (Technitium at 100.121.150.22 returns AAAA records even when IPv6 routing to public registries is intermittently broken), every container check surfaced AggregateError [ETIMEDOUT] in error.log with stack `at internalConnectMultiple (node:net:1114:18)`. The dual-stack DNS race consumed the default 30s connect timeout per unreachable IPv6 family before falling back to IPv4 — 30s+ per container per check cycle. Three reliability properties added via shared fetchWithReliability() helper: 1. family:4 — IPv4-only DNS lookup. Avoids the dual-stack race entirely. 2. Hard per-request timeout (10s) — caps total latency per attempt. 3. Retry on transient codes only (ETIMEDOUT/ENOTFOUND/ENETUNREACH/...) — HTTP 4xx/5xx are surfaced as real responses, not retried. The 401 → WWW-Authenticate → token → Bearer auth flow is now explicit in getDockerHubDigest (was previously a side effect of authenticateAndGetDigest, which has been removed — no remaining callers). Verified end-to-end against real Docker Hub: - linuxserver/plex:latest → real digest in 1349ms (was 30s+ AggregateError) - 5-container checkForUpdates() cycle: 3.6s total (was 150s+) - 86/86 update-manager tests pass; 2343/2343 full suite (4 pre-existing pdfkit module-resolution failures unrelated to this change) |
||
|
|
18ffd2e519 |
fix(nesting-guard): export dataDir from src/config/paths; harden fallback to platform-paths (DC-077) [glm-grade=B]
Pre-fix, every dashcaddy-api container startup logged:
[nesting-guard] Skipped: The "path" argument must be of type string. Received undefined
because src/utilities/nesting-guard.js does require('../config/paths') and
calls paths.dataDir — but src/config/paths.js imported platformPaths and
only re-exported its specific files (SERVICES_FILE, CONFIG_FILE, etc);
dataDir was never re-exported, so paths.dataDir was undefined.
Result: path.join(undefined, 'data') threw TypeError, the outer try/catch
swallowed it, and the entire nesting-guard became a silent no-op. The
cleanup that prevents recursive data/data/data/... directory duplicates
never ran on any startup. Bug class is 'silent functional no-op' (same
family as DC-056 AggregateError visibility).
(1) src/config/paths.js (+11): re-export dataDir as
SERVICES_DIR-derived (with platformPaths.dataDir fallback). dataDir is
the dirname of SERVICES_FILE in container (env override wins), which
equals /app/data — same value platform-paths.dataDir computes for the
default config. Either path is fine; SERVICES_DIR is preferred because it
respects env-override.
(2) src/utilities/nesting-guard.js (+13/-2): defensive fallback to
require('../../platform-paths').dataDir if paths.dataDir is missing
(any future export-shape drift or older caller). Explicit skip-warn
instead of silent catch when both paths fail.
(3) __tests__/nesting-guard.test.js (NEW, 112 lines, 4/4 passing):
isolates module cache per test, exercises (a) cleanup when nested
data/data exists, (b) no-op when clean, (c) dataDir export contract,
(d) dataDir === dirname(SERVICES_FILE) under env override. No jest.doMock
leaks across tests (verified via 4-call probe sequence).
Verified: 4/4 tests passing. Full repo suite: 100/104 suites / 2335/2335
tests passing (4 pre-existing failures in __tests__/billing/* are
unrelated module-resolution issues in src/billing/invoice.js, confirmed
unaffected by this change via stash+rerun).
GLM-5.3 round 1: B (ship, one polish nit — trailing newline on test
file, folded in same commit per multi-round-fix-first protocol).
Deploy plan: container rebuild + atomic swap via /opt/dashcaddy/start.sh
on DNS2; live-verify status.sami=200, dashcaddy-api=Up+healthy, and
absence of [nesting-guard] Skipped log line in container logs.
|
||
|
|
2fef1c47e5 | fix(ca): gate per-service cert/key download behind TOTP+admin scope; require explicit PFX password; add rate limit (DC-076) [glm-grade=A] | ||
|
|
270e8d57e3 |
fix(sites): SSRF hardening — validate upstream + externalUrl reject private/reserved hosts (DC-074) [glm-grade=A]
Pre-fix, an authenticated dashboard operator could call:
POST /api/v1/site {domain:"evil.example.com", upstream:"10.0.0.1:80"}
POST /api/v1/site/external {subdomain:"x", externalUrl:"http://192.168.1.5"}
and end up with a Caddy site block that proxies PUBLIC traffic at
evil.example.com to an INTERNAL host. Caddy runs on DNS2 (same
network as the targets), so the SSRF lands.
The pre-fix /site upstream regex /^[a-z0-9.-]+:\d{1,5}$/i only
checked charset — it happily accepted 192.168.1.1:80 and
169.254.169.254:80 (AWS metadata IP). /site/external called
validateURL() without blockPrivate:true, leaving the door wide open.
(1) New helper validateUpstream() in fleet-validation.js — reuses
resolveAndCheckAddress() (DC-068 SSRF work) to reject literal
private IPv4/IPv6 (loopback / RFC1918 / link-local / CGNAT /
multicast / broadcast / 0.0.0.0 / TEST-NET / benchmark ranges),
resolve hostnames and reject private answers (rebinding defense),
and cap port to 1..65535. Opt-in via SITES_ALLOW_PRIVATE_UPSTREAMS=true.
(2) /site calls validateUpstream() BEFORE caddy.modify() — gate
happens before any state mutation. Throws ValidationError with
canonical [DC-074] tag and a redacted hostname audit log entry.
(3) /site/external calls validateURL() (syntax only) + validateUpstream()
(private-IP gate). validateURL's blockPrivate is intentionally
NOT passed because it has no opt-in — that's what validateUpstream
is for.
(4) Tests (__tests__/routes/sites-dc074.routes.test.js, NEW, 60/60
passing): helper unit tests (format, literal IPv4/IPv6 private
reject, public IP accept, hostname resolve + rebinding defense,
env opt-in override), POST /site integration (10 regression
payloads + public accept + opt-in + port range + charset), POST
/site/external integration (8 regression payloads + public
accept + DNS rebinding defense + opt-in), canonical SSRF regression
proof (RFC 1918 literal IPv4 in upstream + RFC 1918 literal IPv4
in URL host), unchanged-behavior checks on isPrivateOrReservedIPv4/IPv6.
Full repo suite: 2402/2402 tests in 102 suites (zero regressions).
GLM-5.3 stand-in judge round 1 (deleg_384b9f53, 41.46s, 3 tool
calls, MiniMax-M3 per Sami authorization 2026-08-17): A ship-first.
Refs: codex-as-judge SKILL.md 'Stand-in fallback chain'. Verdict
record: /root/dashcaddy-polish/.ump-verdicts/2026-08-18T22-35-00Z-dc-074-round-1-A.json
|
||
|
|
a9bb4a1835 |
fix(caddy-upstreams): validate host is known upstream on all 3 mute endpoints (DC-073) [glm-grade=A]
Bug class: silent state corruption via path-style endpoint inconsistency. Pre-fix, only POST /caddy/upstreams/mute (bare body-style) rejected unknown hosts with a 400. The path-style POST /caddy/upstreams/:host/mute and POST /caddy/upstreams/:host/unmute endpoints skipped that check entirely. An authenticated operator could POST /caddy/upstreams/phantom.test:12345/mute and caddyUpstreamWatcher.setMuted() would silently add the phantom host to its muted Set and _saveState() would persist it to disk. The phantom entry survives container restarts and pollutes the snapshot view. Fix: consolidate validation in a single validateAndMuteHost() helper used by all three mute endpoints. The helper enforces (1) host format charset, (2) length cap, (3) membership in caddyUpstreamWatcher.upstreams (the live registry populated by scanSites()). No phantom host can reach setMuted. Tests: 15 new regression tests in __tests__/routes/caddy-upstreams-dc073.routes.test.js — exercises the helper directly (unit) and via each endpoint (integration), asserts rejection happens BEFORE setMuted is called (no state corruption), and the existing 3 caddy-upstreams.routes.test.js cases still pass. Router introspection test asserts no duplicate route registrations. Full suite: 2342/2342 tests / 101 suites. |
||
|
|
83d7c65bf2 |
fix(exec): scope-based authorization + tighten containerId charset (DC-072) [glm-grade=A]
Pre-fix, dashcaddy-api/routes/exec.js (the ws://host/ws/exec/:containerId
WebSocket container terminal endpoint) captured auth.scope at lines 39/46
but never enforced it — any API key or JWT, regardless of scope, got a
full PTY-backed shell inside the running container. A key issued with
scope ['read'] (a legitimate monitoring/observability scope) could
escalate to a root-equivalent shell. Container exec is full root inside
the container's user namespace, so this was a privilege-escalation across
the auth trust boundary.
Fix:
1. assertExecScope(auth) requires scope.includes('admin'); throws a
tagged 403 error (DC-072_INSUFFICIENT_SCOPE) on rejection with
requiredScope + actualScope in the envelope.
2. Called BEFORE wss.handleUpgrade so the WS gate cannot be bypassed.
3. 403 over the upgrade socket is JSON (code, requiredScope, actualScope)
so the dashboard can show operator-actionable messages.
4. isValidContainerId(id) tightened to Docker's actual charset
(12 or 64 lowercase hex). Pre-fix regex accepted _, -, ., mixed
case, and any length up to 128; Docker would 404 the inspect and the
rejection surfaced as a generic 500.
5. Audit-log pair: session start (container name + auth id) and session
end with durationMs + reason ('exec-stream-end' vs 'ws-close'
for abnormal disconnects); idempotent via ended-flag guard.
6. Both helpers exported via __test for unit tests (no live WS).
Tests: 20 new tests in __tests__/routes/exec.routes.test.js cover:
- assertExecScope: admin passes; read/write/empty/undefined/null/non-array
rejected with the canonical 403 envelope.
- isValidContainerId: 12/64 lowercase hex accepted; uppercase / mixed /
non-hex / _.- / wrong length / null / non-string / padded / CRLF
payload rejected.
Full suite: 2327/2327 tests passing across 100 suites (zero regressions).
GLM-5.3 round 1: A with 2 LOW polish (scope-coercion defensive comment +
abnormal-close audit-log fallback). Both folded into the same commit.
Round 2: A. Ship.
|
||
|
|
297332b0e1 | fix(caddycode): validate + escape generation config — block CRLF / " / brace injection in Caddyfile interpolation (DC-070) [glm-grade=A] | ||
|
|
933606ce3f |
fix(caddy-admin): IPv6 loopback origin allowlist + bracket-strip helper (DC-069) [glm-grade=A]
Two coupled bugs that, together, cause the live 'admin.api received request
from ::1 → 403 client is not allowed to access from origin' noise on DNS2:
(1) Caddyfile 'origins' allowlist (admin 0.0.0.0:2019 block on DNS2) had
4 IPv4 entries (localhost/127.0.0.1/172.17.0.1/0.0.0.0) but no IPv6
entry. Per glibc RFC 3484 + /etc/hosts '::1 localhost', Node's
dns.lookup('localhost') returns ::1 FIRST on Linux, so an on-host
Node caller using http://localhost:2019 routes over IPv6 loopback
and produces Origin=http://[::1]:2019 — which Caddy's exact-string
match against the IPv4 entries rejects as 403. Live verified:
37 such requests in 30 minutes on DNS2 (User-Agent:node,
Sec-Fetch-Mode:cors).
(2) _httpFetch (src/utils/http.js) was broken for IPv6 literal URLs:
on Node 22, new URL('http://[::1]:2019/x').hostname === '[::1]'
(brackets preserved), but http.request({hostname}) needs the
BRACKETLESS form for actual TCP connect. Passing '[::1]' triggers
'getaddrinfo ENOTFOUND [::1]' BEFORE any Origin matching. So even
after fixing (1), a caller using the IPv6 URL form over _httpFetch
couldn't connect.
Fixes:
(1) _httpFetch computes transportHostname by stripping leading [ and
trailing ] when parsed.hostname is bracket-wrapped. transports via
bracketless form. defaultOrigin keeps bracket form so Caddy's
allowlist exact-matches. Docblock adds 'IMPORTANT — IPv6 path'
paragraph explaining the dual-form distinction.
(2) dashcaddy-installer/templates/Caddyfile.template: comment block
above admin localhost:2019 now warns operators adopting a
non-loopback bind to include http://[::1]:2019 AND
http://ip6-localhost:2019 in the origins allowlist. Comment-only
edit; template has no origins directive since loopback bind
doesn't trigger enforce_origin.
Tests (NEW utils-http-caddy-admin-ipv6-origin.test.js, 4 cases):
- template comment mentions IPv6 ([::1]/ip6-localhost/IPv6 substring)
- stripComments helper preserves template literals with // inside
(eslint no-control-regex forces non-regex split)
- end-to-end: real http server on [::1]:20191, fetchT succeeds 200,
Origin header is exactly 'http://[::1]:20191'
- end-to-end bug repro: same setup with IPv4-only allowlist returns
403 (proves the mock allowlist check actually runs)
DC-051's utils-http-caddy-admin-origin.test.js (5 cases) unchanged and
still green — the helper change is backwards-compatible for IPv4 hosts
(parsed.hostname.startsWith('[') is false for 127.0.0.1/localhost/
172.17.0.1).
Full suite: 2281/2281 (98 suites, +4 net new). ESLint clean on touched
files.
GLM-5.3 judge round 1 (35s, 3 tool calls): GRADE=A. 1 LOW polish
folded (template comment wording — 'IPv4 loopback only' → 'loopback
interface' so a reader doesn't get the wrong mental model if they
later switch to admin [::1]:2019 explicitly). No blocking issues.
|
||
|
|
5382d832d9 |
fix(fleet): SSRF hardening — hostname validation + DNS rebinding + probe-by-IP (DC-068) [glm-grade=A]
bug: POST /api/v1/fleet/hosts (DC-108) accepted any string as the hostname field and the followup GET /fleet/status flow composed it verbatim into a probe URL. An authenticated dashboard operator could register 127.0.0.1 or 169.254.169.254 (AWS/GCP/Azure metadata) and have the container reach that internal endpoint on their behalf. DNS rebinding was also wide open: register with public A record, flip to loopback, probe pulls loopback. fix: 4 layers of defense 1. New fleet-validation.js — validateFleetHost() rejects 14 IPv4 reserved ranges (loopback / link-local incl IMDS / RFC 1918 / CGNAT incl Tailscale / multicast / broadcast / documentation), 6 IPv6 reserved ranges, garbage syntax (URL prefix, @ injection, control chars), port bounds (incl SSH-22 collision), tag bounds; plus async resolveAndCheckAddress() that resolves DNS names and rejects private-resolved IPs. 2. routes/fleet.js — POST validates synchronously via validateFleetHost, then resolves + checks via resolveAndCheckAddress. Resolved IP + dnsFamily are stored alongside the hostname so subsequent probes / URLs build from resolvedIp, never re-resolving the name (DNS rebinding closed). 3. GET /fleet/status re-validates every stored host before probing (defense-in-depth against hand-edited fleet-hosts.json) and categorizes hosts as validation_failed vs probe-able. Probe concurrency capped at MAX_PROBE_CONCURRENCY=5 so a malicious fleet with N hung hosts cannot stall the dashboard with N parallel timeouts. 4. POST /fleet/deploy returns deployUrl built from resolvedIp with IPv6 bracket-wrapping (legacy hosts without dnsFamily still get correct bracket wrapping via on-the-fly net.isIP check). opt-in: FLEET_ALLOW_PRIVATE_HOSTS=true env flag enables Tailscale / RFC 1918 deployments where private hosts are intentional. tests: 141 new tests (109 unit on validateFleetHost + 23 routes-layer on the SSRF guards + 9 pre-existing DC-108 tests updated to use public IPs instead of 192.168.x / 10.x). 2277 / 2277 pass on DNS2. manual verification: GLM-5.3 judge round 1 = A (4 tool calls, 49s, ship). IPv4-mapped IPv6 edge case ::ffff:127.0.0.1 caught correctly via net.isIP + delegated IPv4 check. |
||
|
|
c6b2f556c2 |
fix(openclaw): harden proxy — 5 MiB cap, RFC 7230 hop-by-hop strip, open-redirect (Location/Refresh/WWW-Auth) strip, path + status validators (DC-065) [glm-grade=A]
Round 1 GLM-5.3: C — missing "location" (open-redirect through proxy).
Round 2 GLM-5.3: C — missing "refresh" + "www-authenticate" (same class).
Round 3 GLM-5.3: A — ship.
Closure of four vulnerabilities in routes/openclaw.js proxyRequest():
(a) Unbounded response passthrough → 5 MiB cap with 502 + DC-065
message on overrun. Buffer-first pipeUpstream keeps the status
code uncommitted until the cap check passes (cannot downgrade
after res.write()).
(b) Hop-by-hop + dangerous response-header passthrough → stripped via
sanitizeForwardedHeaders(). Hop-by-hop per RFC 7230 §6.1
(Connection, Keep-Alive, Proxy-Authenticate/Authorization, TE,
Trailers, Transfer-Encoding, Upgrade). Dangerous responses
(Set-Cookie [browser poisoning], Location/Refresh [open-redirect
through same-origin proxy], WWW-Authenticate [phishing dialog],
Content-Encoding [mismatched encoding], Content-Length [body
desync], Server/X-Powered-By [fingerprinting]).
(c) proxyRes.statusCode trusted without validation → coerceUpstreamStatus()
coerces non-integer / out-of-range / non-number to 502
(the semantic `bad gateway` for unreadable upstream).
(d) Path taken from req.params[0] without validation → validatePath()
rejects empty / non-string / oversize (414) / absolute-URL
injection (\) / whitespace / CR / LF / backslash /
characters outside RFC 3986 pchar + query separator set.
Tests: __tests__/routes/openclaw.proxy-hardening.test.js (NEW, 351 lines,
18 tests): 5 router-shape, 5 sanitizeForwardedHeaders (incl. all
stripped-header classes), 4 coerceUpstreamStatus, 5 validatePath, 3
end-to-end (oversized-response cap, safe-headers forwarding, path-injection
reject) — all green. Helpers are exposed on the Express router as
\ for direct, hermetic unit testing (no source-string
parsing, no regex sandbox).
Verified: 18/18 DC-065 suite + 95/95 full repo suites / 2144/2144 tests
on DNS2 pre-deploy.
Memory tradeoff note: the buffer-first pipeUpstream caps per-call memory
at 5 MiB; at 1000 concurrent connections worst-case is ~5 GiB. Node CLI
flags in start.sh + ulimit bound concurrency. Documented inline.
|