Why rewrite a backend that works
The original OpsPing backend was TypeScript on Node, and we don't regret it — a small team with a deadline should pick the productive, boring stack. It got the product built. But a paging backend has an unforgiving job description: accept alerts from every webhook and agent pointed at it, write every one of them down, and stay responsive at 3 AM while everything around it is on fire. That job is concurrency, I/O, and JSON — and under sustained concurrent load, the runtime's cut of every request stopped being negligible.
So we rewrote it in Go. Same routes, same JSON payloads, same auth, same OpsGenie-compatible API surface — nothing that speaks to OpsPing's API can tell the two implementations apart. We swapped the engine and kept the car. What follows is the honest accounting of what that bought, measured rather than felt.
Twelve-year-old silicon, on purpose
The benchmark machine for both versions is a 2013-era Intel Xeon E5-2660 v2. That's a deliberate choice, not a budget constraint. Modern hardware is very forgiving of inefficient code — enough cores and cache will make almost anything look acceptable. Old silicon has no such manners. Running the benchmarks on a 2013 Xeon stresses the code instead of the CPU's marketing budget.
It's also the honest ceiling check: if the backend is fast on 2013 hardware, it'll be fast on anything you're likely to run — including the modest box a private single-tenant deployment might live on. Both versions got identical conditions: same machine, same Postgres, identical load, side by side.
The numbers
Every number below follows the same protocol: identical hardware, identical database, identical load. Ingest throughput is counted from database rows — what actually landed in Postgres — not from what the load tool claimed it sent. Client-side counters drift in exactly the direction you'd hope.
| Benchmark | TypeScript | Go | Difference |
|---|---|---|---|
| Sustained alert ingest (DB ground truth) | 138 creates/s | 197.9 creates/s | +43% throughput |
| p99 latency, same ingest load | 941 ms | 390 ms | 2.4× better tail |
| CPU-bound burst — login/bcrypt path, concurrent load | 1× | ~10× | ~10× throughput |
| Read / list / create API endpoints | 1× | 2–3× | 2–3× faster |
| Idle memory, production | ~200 MB Node process | ~17 MB RSS | ~10× smaller |
| Deploy artifact | Node.js process + runtime | Single ~15 MB static binary | One file, no runtime |
The memory figures are production numbers, not benchmark numbers: idle RSS of the live process. The login path is bcrypt-bound — deliberately slow by design — which makes it a clean CPU-bound concurrency benchmark. The endpoint gains come from the framework and JSON layers alone.
Thirty minutes, flat out
Sustained load tells the truth that burst benchmarks skip. We ran the Go backend at a flat 400/s ingest for 30 minutes straight: 728,659 alerts written, zero server errors. The limiting factor was our load generator, which ran out of headroom before the server did. That result set the floor, not the ceiling.
Then we qualified 1,000 alerts per second
On October 1 we restored a 757,009-alert database and ran the real mixed profile against one OpsPing process and PostgreSQL 17: alert creation plus 15 API pollers, five live-event streams, five responders acknowledging and snoozing alerts, and the normal reping workers. The load ramped to exactly 1,000 creates/s and held there for two minutes.
The database confirmed 119,892 alerts across the aligned 120-second peak window — 999.1/s, with 10-second buckets peaking at 1,000.4/s. Read p95 was 67.5 ms, write p95 107.7 ms, and all 95 live-event holds completed cleanly.
The failures on the way there mattered more than the victory lap. Our first attempt exhausted ephemeral ports in a disposable bridge-network nginx hop, even though the app and database were healthy. Replacing that test-only hop with the same host Caddy topology we use in production removed every 502. That exposed the next honest limit: a 16-connection database pool. On the dedicated 24-core qualification host, a measured 64-connection pool cleared the gate.
What this proves — and what it doesn't. One OpsPing process on the documented 24-core, 31 GiB test host can absorb a two-minute 1,000 alerts/s burst while normal read, response, SSE, and reping work stays active. It does not claim that a small 2-vCPU silo can do the same, that 1,000/s is sustained forever, or that provider latency is free. Capacity claims without the hardware, duration, and concurrent workload are just adjectives. See the dated evidence in our Trust Center.
How we made the swap boring
A rewrite is only as good as its cutover. Ours was engineered to be an event nobody noticed:
- Identical API. Same routes, same payloads, same GenieKey-compatible keys — existing integrations kept working without a single change.
- Zero-downtime data migration. Same Postgres, same schema. 126 tenants, 183 users, and 4,851 alerts moved while the service stayed live. The storage layer never changed, so "migration" mostly meant verification.
- Full end-to-end suite before cutover. The complete Playwright suite ran green — 232 of 233. The one failure was a pre-existing flake we could reproduce against the old backend, so it didn't get a vote.
Methodology, in short. Identical hardware (2013-era Intel Xeon E5-2660 v2), identical Postgres, identical load — TypeScript and Go A/B. Load generated with k6 plus a custom load driver for the ingest path. Ingest counts are ground truth from database rows, not client-side counters.
What's next
We didn't do this to win a benchmark. We did it because the tail latency and the memory bill were the two lines we didn't like on our own graphs — the numbers above are the receipt.
Practically, a ~15 MB static binary that idles at ~17 MB RSS runs comfortably on small, cheap machines. That's good for our costs and better for anyone running OpsPing as a private single-tenant deployment on hardware they already own. Next up is more of what made this rewrite safe: identical APIs, boring migrations, and benchmarks we'd be glad to see someone else reproduce.
— The OpsPing team