Skip to content

Add alerting on usage table write rate (zero-rows SLO) #70

Description

@vaprdev

Background

On 2026-06-27, a PlanetScale schema deploy (DR #18) dropped the go_quota_allocation_id column from the usage table. The production console server was still running old code that included that column in every INSERT. Every usage insert failed silently with ERROR 1054: Unknown column 'go_quota_allocation_id' in 'field list'.

Because the proxy is designed to fail open on usage writes (apps/server/src/services/proxy/record-usage.ts catches DatabaseError and returns void), inference continued normally for customers — but zero usage rows were written for ~1 hour 40 minutes (17:11 UTC to ~18:50 UTC).

The outage was not detected by any automated alert. It was noticed manually after ~1.5 hours.

What happened

Metric Value
Last usage row before break 2026-06-27 17:11:38
Deploy completed 2026-06-27 17:11:49
Rows written during outage 0
Rows in the 90 min before outage 1,627
Estimated unbilled revenue ~$173 (anomaly internal) + ~$0.23 (CarMax)
Time to detection ~1.5 hours (manual)

What we need

An alert that pages immediately when the usage table stops receiving rows. This is the revenue signal — if it goes to zero, billing is broken.

Proposed SLO

  • Metric: COUNT(*) FROM usage WHERE created_at > NOW() - INTERVAL 5 MINUTE
  • Threshold: 0 rows in a 5-minute window (during normal hours, anomaly generates hundreds of rows per 5 minutes)
  • Page: immediately (P0)
  • Secondary metric: UsageRepo.record failure counter — currently swallowed by the fail-open catch in record-usage.ts. Export a counter/metric so failures are visible even when the response succeeds.

Implementation notes

  • The usage table write rate is a direct proxy for "is billing recording revenue." It should be treated with the same urgency as inference uptime.
  • Consider a Datadog/Dash0/Grafana query that scrapes this count periodically, or emit a metric from the server itself.
  • The fail-open behavior in record-usage.ts is correct for availability (do not drop customer responses), but the swallowed errors must be observable. Today they are only Effect.logError calls — not alerted on.

Related

  • Critical path: docs/critical-paths.md §4 "Usage / ledger writing"
  • Fail-open design: apps/server/src/services/proxy/record-usage.ts:56-58
  • Root cause writeup: tmp/db-1.md

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions