2026-07-05/06 — Post-audit launch hardening (PR #26 cont.)
2026-07-05/06 — Post-audit launch hardening (PR #26 cont.)
Observability — Langfuse LLM tracing
- Every LLM call is now traced (OpenRouter lanes + the Anthropic Batches lane): input/output, token usage, cost, and grouping metadata. Isolated
NodeTracerProvider+LangfuseSpanProcessorcoexisting with Sentry's global OTel. Full guide in `OBSERVABILITY.md`. - Two silent failures found + fixed: spans were dropped by Sentry's request-span sampling (fixed with
AlwaysOnSampler+ aROOT_CONTEXTdetach), and generation input/output rendered null until the tracer was named"ai"— the only scope nameLangfuseSpanProcessormaps I/O for. - Conventions: Langfuse
sessionId=scan_session.id(a scan's four platform calls group into one session),userId=organizationId; scan traces carryqueryId+topicIdfor cross-run filtering.LANGFUSE_*keys are optional — unset = zero-overhead no-op.
Billing / quota
- Manual/comp tier grants take effect immediately — enforcement read the raw
organization_quota.tierand ignoredmanualTieruntil a Stripe sync ran, so a granted org stayed capped at its stale tier ("out of scans" on a comped account). NeweffectiveTier(row) = higherTier(tier, manualTier)drives all three gates + the quota display; unknown tiers still fail closed tofree.
Server-side pagination on every list page
- Queries, Topics, and Scans now paginate + filter server-side. Scans already did; Topics and Queries were converted. Queries was the worst offender — it fetched _every_ query (1000+ for a large org) and built a sparkline for each on every visit, then filtered/paged in the browser (so client-side filters silently only saw the loaded page). Now search / topic / importance / status filters, sort, and pagination are all SQL; state lives in the URL (deep-linkable, back-button friendly); sparklines build for the visible page only. Bulk multi-select preserved across navigation.
Insights LLM cost — ~45% cheaper per narrative, thin-topic calls eliminated
- Thin data no longer reaches the LLM — topic narratives (sync + batch) now enforce the
INSIGHTS_MIN_SCANS(5) floor the workspace already used, plus aresponses === 0skip that closes a raw-walk empty-snapshot edge (≥5 scans but no content). A 1–4 scan topic used to burn a full Sonnet call to say "not much yet." - Output clamped 3000 → 1000 tokens on all four lanes (output is ~68% of cost; the schema caps at 6+6 items). Brevity via
.describe()hints + prompt guidance, _not_.max()(a hard cap would fail Zod and burn a second billed call through the text fallback — see the documented rationale ininsights-prompt.ts). - Prompt payload ~67% smaller — new
compactSnapshotForPromptdrops zero-response platform slices, null context fields, and per-sourcesampleUrl/sampleTitle(the stored snapshot keeps them for the UI). Serialized compact, not pretty-printed. Measured ~1760 → ~589 tokens on a rich real snapshot. - Better cache hit-rate —
hashSnapshotno longer keys on the volatile per-row sample citation, so citation churn stops invalidating an otherwise-identical narrative. - Honest batch metrics — the Anthropic Batches lane records its cost with the 50% discount (Opus added to
pricing.ts); it was previously un-priced in Langfuse. - Verified live: a rich-snapshot narrative comes in at ~730 output tokens (under the 1000 cap), complete, calibration intact.