All posts
Node.jsPerformanceBenchmarkAPM

How I Built a Zero-Config Node.js APM Without Tanking Event Loop Performance

I was stating "under 0.3ms per request" and "under 50 nanoseconds per DB call" on my own site without ever actually measuring either number. Here's the real benchmark, the real result, and why I published it even though it wasn't as flattering as the number I made up.

Rahul Patel·6 min read·October 6, 2026

We made up a number, and it was wrong

For a while, two different pages on this site said two different things about how much latency APILens adds to a request. One said "under 0.3ms per request." Another said "less than 1ms per request... under 50 nanoseconds per DB call." Neither number came from an actual benchmark. They were estimates that sounded reasonable, got typed into a FAQ answer, and sat there unquestioned.

That's a bad habit for an observability tool specifically — the entire pitch is "stop guessing, see the real number." Shipping guessed numbers about my own product undercuts that pitch directly.

So I ran the benchmark I should have run before writing either claim.

How the middleware is supposed to be cheap

Before the numbers, the architecture, because it explains why the result should be small in the first place.

auto-api-observe does two things per request:

1. Synchronous capture — timing, route, status, and (via auto-patched DB/HTTP clients) query and outbound-call details, all collected inline as the request already flows through your handler. No extra I/O, no extra await.

2. Asynchronous shipping — captured events go into an in-memory queue and get batched off to the ingest endpoint on a timer (every 5s or 100 events, whichever comes first), via plain http/https, completely decoupled from the request/response cycle.

shipper.push(entry) in the hot path is a synchronous array push — it never awaits a network call before your response goes out. That's the whole reason to expect the overhead to be small: there's no network round-trip sitting between your handler and your response.

"Should be small" isn't a benchmark, though. So:

The benchmark

Two identical single-route Express apps — GET /ping → res.json({ ok, ts }) — differing in exactly one thing: whether observe() is installed. The instrumented app's endpoint pointed at a local mock sink instead of the real apilens.rest, so the test measures the middleware's own CPU cost, not network latency to a remote server (shipping is async either way, per the architecture above — this isolation doesn't change what's being measured, it just removes a variable).

Load generated with autocannon: 50 concurrent connections, 10,000 requests per run, 3 runs each, against localhost.

The full setup — both server scripts, the mock sink, and exact rerun instructions — is in the repo: benchmark/. This isn't a number to trust because I said so; it's a number you can reproduce in about two minutes.

The result

BaselineWith auto-api-observe
Avg latency (3 runs)1.78ms / 1.15ms / 1.23ms2.80ms / 2.70ms / 2.12ms
p50~1–2ms~2ms
p97.5~2–4ms~5–8ms
p99~3–4ms~6–11ms
Throughput~10,000 req/sec~10,000 req/sec

Average overhead: roughly 1.1–1.5ms per request. Not 0.3ms. Not under a millisecond, consistently. Throughput held steady at this concurrency — the cost showed up as added latency per request, not a lower request ceiling.

That's a real number, and it's worse than what I was claiming before. I'm publishing it anyway, because an honest 1.5ms beats a flattering 0.3ms that nobody ever checked.

Why real routes should see less than this

This test hit the cheapest possible route on purpose — a trivial JSON response with no database query, no outbound call, nothing for the instrumentation to actually describe beyond "a request happened." That's close to a worst case for *proportional* overhead: the middleware's fixed per-request cost is being measured against an almost-zero baseline cost.

A real route doing a database query or two has a baseline latency of several milliseconds to begin with. The same ~1–1.5ms fixed instrumentation cost becomes a smaller fraction of a larger number. I haven't benchmarked that specific case yet — it's the obvious next one to run, and I'll publish that too when I have it.

What I still haven't measured

In the interest of not repeating the same mistake with a new number: this benchmark does not cover DB query instrumentation overhead specifically (the test route touches no database), outbound HTTP call instrumentation, or behavior under sustained load over a long-running process rather than a short burst. The "under 50 nanoseconds per DB call" claim I removed is unverified by this test and won't be reused until it has its own real benchmark behind it.

Try it

npm install auto-api-observe
app.use(observe({ apiKey: process.env.APILENS_KEY }));

- Benchmark source: github.com/rahhuul/auto-api-observe/tree/master/benchmark

- Dashboard: apilens.rest/dashboard

- GitHub: github.com/rahhuul/auto-api-observe

- npm: auto-api-observe

Free to start. Zero dependencies. And now, a latency claim you can actually reproduce.