Two things:
1. FIX a live regression the async conversion missed: chat.js calls repos via the
lazy repos() helper (not the R. prefix), so my sweep skipped it — effectiveStatus
/ broadcastPresence read `repos().users.byId(userId)` synchronously, but that's a
Promise now, so presence broadcasts always reported status 'active' and dropped
last_seen. Now awaited (effectiveStatus/broadcastPresence async); touchSeen is a
fire-and-forget UPDATE with .catch. Audited all non-R. repo calls — only chat.js
was affected (media.js backfill was already awaited).
2. Swappable pub/sub for multi-instance real-time fan-out (the actual blocker to
running >1 instance — not the DB). server/pubsub.js picks a backend by
PUBSUB_BACKEND (default 'memory'). Local socket delivery is UNCHANGED; publish is
additive — memory = no-op (zero hot-path cost, identical single-instance
behaviour), redis = fan-out to other instances with a self-echo guard. chat.js
pushToUser/broadcastPresence now also publish; each instance subscribes to deliver
remote events to its local sockets. Interface is tiny so Redis is one swappable
file (Postgres LISTEN/NOTIFY or NATS could drop in the same way — never hardwired,
as requested). Dormant redis service added to compose behind the 'scale' profile;
redis dep added; PUBSUB_BACKEND/REDIS_URL documented.
Validated: smoke 22/22 (memory), e2e chat delivery green. NOTE: full multi-instance
also needs distributed presence (isOnline is per-process) + meeting-signaling
sharing — chat/presence fan out via this layer; those are follow-ups.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- db/pg.js: the pg backend (prepare/exec/tx/init) — ?→$N translation, BIGINT parsed
as Number (matches sqlite; else expires_at<Date.now() compares string<number),
transactions on one pooled client, init() applies schema.pg.sql. Same interface as
db/sqlite.js, so repos are unchanged.
- repos.js: the ~7 SQLite-only queries rewritten to run on BOTH engines —
audit.add @named→positional; email lookups COLLATE NOCASE→LOWER()=LOWER();
INSERT OR IGNORE→ON CONFLICT DO NOTHING (addMember/poll vote/favorite);
mergeInto's UPDATE OR IGNORE→UPDATE…WHERE NOT EXISTS/NOT IN and INSERT OR
REPLACE→ON CONFLICT DO UPDATE. Re-validated on sqlite: db-smoke still 22/22.
- server.js: boot now `await db.init()` before listening (pg creates tables; sqlite
no-op), so the first request can't hit a missing table.
- db/migrate-sqlite-to-pg.js: one-shot row copy in FK order (bulk insert, TRUNCATE
first so re-runnable). audit_log id left to PG's identity.
- package.json: add pg ^8.13.1.
Next: validate DB_BACKEND=pg smoke against a real Postgres on the server, then merge.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The full sync→async conversion is complete and green on the SQLite backend. Every
DB call across the app now awaits the async adapter, so the identical code runs on
Postgres at cutover.
Converted (this commit finishes Phase 3):
- session.js: currentUser/apiKeyFromReq async → 63 route awaits + WS + static.
- routes.js: all ~250 R.* awaited; DTO helpers (namesFor, avatarsFor, buildMsgDTO,
buildPollDTO, reactionsForMessage, postSystemMessage, pushGroupUpdate,
issueRefreshToken, provisionFromBizgaze) made async; every `.map(x=>buildDTO(x))`
restructured to `await Promise.all(...map(async...))` preserving order; `.filter`
predicates that hit the DB moved to an `asyncFilter` helper; chained
`R.x.y(...).length/.map/.filter` wrapped as `(await R.x.y(...)).method`; stream
upload handlers (recording/transcript/attachment) made async.
- calls.js / signaling.js: all call/meeting fns async; leaveMeeting AWAITS
persistCallHistory + finalizeTranscript BEFORE endCallByRoom (ordering matters —
fire-and-forget would race the map teardown); WS handle()/cleanup() async with
.catch guards.
- static.js: authAttachment(Raw) async (the .some carrier check became a loop),
handleGet async; server.js dispatch catches handler rejections → 500 not a hang.
- media.js backfill, push.js, reminders.js, webhooks.js await their repo calls.
Validation on DB_BACKEND=sqlite: db-smoke 22/22; legacy e2e 80 checks pass with zero
FAILs (throws only at a PRE-EXISTING WS lobby-drift assertion, unrelated). Every
server file `node --check` clean.
Still on the branch — master untouched. Next: Phase 5 (pg backend + ~7 dialect
queries + data migration + Docker Postgres + cutover), then merge.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
THE ANSWER to "why does an already-downloaded video still buffer?" — it was never
the download, and it was not the server. Probing the real uploads on the box:
d0e49e58… 1920x1080 19.4 Mbps 75 MB / 31 s
ad929d0b… 1920x1080 19.0 Mbps 27 MB / 11 s
9f4e0865… 720x1584 3.6 Mbps 14 MB / 31 s
To play a 19 Mbps file the client has to SUSTAIN a 19 Mbps download for the whole
clip. No mobile link does, so the <video> buffer drains every few seconds: buffers,
plays, buffers, plays. Server-side disk read was instant and load was 1.7 on 20
cores throughout — the bottleneck is the media itself, not the delivery path.
Second, independent defect: phone MP4s store `moov` AFTER `mdat` (verified on two
uploads), so the player must fetch the file's tail before it can start at all.
Fix — keep the original bytes untouched (that is what the download button serves,
full quality) and build <id>.web.mp4 beside it: longest side capped at 1280,
~2.5 Mbps ceiling, +faststart. Measured on the 19 Mbps file:
27.3 MB @ 19.0 Mbps -> 2.55 MB @ 1.78 Mbps (10.7x less bandwidth)
transcode took 2.4 s for an 11.5 s clip
- server/media.js (new): probe, decide, 2-at-a-time background queue. Already
light + correctly sized + faststart => no rendition at all. Light but wrong atom
order => remux -c copy (seconds, no re-encode). Otherwise re-encode. A rendition
that lands bigger than the original is discarded. MP4 box-walker for the
faststart test is unit-checked against known fast/slow files, both directions.
- /stream/<id> serves the rendition, falling back to the original while it is still
transcoding, so a video is never unplayable. /files/<id> is unchanged and still
serves the pristine original for download.
- Renditions are queued at upload, and backfilled 15 s after boot for the videos
that predate this. Range serving is now one shared helper for both routes.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
User-facing
- New post-login home (/home): chat rail + Share/Connect (embedded) + Meeting; login lives here when logged out
- Landing: "Log in with BizGaze" + no-login screen share
- Console replaced by a role-scoped Dashboard (/dashboard): admins see all team sessions, others see only their own; stats + CSV/PDF export
- Recordings saved as MP4 (H.264/AAC) with WebM fallback; old .webm still downloadable
- Fix: duplicate "Sign in" on the login card
Auth / integration
- BizGaze as identity provider: /api/login validates against BIZGAZE_LOGIN_URL (env-gated) and provisions a local user
- Phase 2 start: /api/v1 alias for all /api routes; Authorization: Bearer accepted across HTTP + WS; login returns a token (for native desktop/mobile clients)
Backend refactor (Phase 1, behavior-preserving)
- Split server.js into config/lib/session/presence/routes/static/signaling + repos (data-access) + bizgaze (service)
- All SQL behind repos.js, tenant-scoped (tenantId == team_id for now)
- e2e updated to current flow (21/21 pass before and after)
Docs: ARCHITECTURE.md (target architecture + phased plan), CLAUDE.md repo layout, .env.example BIZGAZE_LOGIN_URL
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- server: /api/ice endpoint reads TURN creds from env (TURN_URLS/USERNAME/CREDENTIAL)
- share/connect: load ICE config at page open
- fixes: stop icon, bright chat notification, beep audio-unlock,
customer screen cleanup on session end, Home link, Remember-me on agent login, Time spent fixed from 90 seconds to actual time spent