Skip to content

Neon cold starts, where the half second goes and what to do about it

Scale-to-zero means an idle branch takes roughly 500 ms to wake, and a serverless function adds its own cold start on top. How to measure the parts and fix the ones that matter.

Neon5 min readships at docs/solutions/neon/cold-start-latency.md

Tags: neon · performance · cold-start · serverless · latency · vercel

Someone reports that the app is slow. You open it and the first page takes two seconds; you reload and it takes 200 ms. Nothing in the code is conditional on "first request". Two independent cold starts are stacking up, and they have different fixes.

Where the time actually goes

A first request to an idle serverless app on an idle Neon branch pays, in order:

StageTypicalWhat it is
Function cold start100-400 msRuntime boot, your module graph evaluated
Neon compute wake~500 msThe branch's compute was scaled to zero
TLS + auth to Postgres20-80 msOnly if you open a socket
First queryyour queryThe actual work

The second and third are the ones people misattribute. Neon scales a branch's compute to zero after a period of inactivity (five minutes by default) and resumes it on the next connection. That resume is roughly half a second. It is not a bug, it is the reason the branch costs nothing while idle, and it is why branch-per-preview is affordable at all.

Measure before you fix

const started = Date.now();
const rows = await sql`select 1 as ok`;
console.log("db round trip", Date.now() - started, "ms");

checkDatabase() in src/db/health.ts already reports exactly this, and bun run verify prints it. Run it twice in a row:

  • First call slow, second call fast → compute wake. Read on.
  • Both slow → not a cold start. Look at the query, the region pair, or a missing index.
  • First call slow only after a deploy → function cold start, not the database.

Fix 1: use the HTTP driver

getSql() sends one HTTPS request per statement. There is no socket to open, no TLS handshake to a Postgres backend, no pool to warm. On a cold instance this is measurably the fastest way to get your first row, and it is why it is the default in src/db/client.ts.

getPool() opens a real connection: TCP, TLS, Postgres auth, and a WebSocket upgrade through Neon's proxy. Worth it when a request needs an interactive transaction; pure overhead when it does not.

If you find yourself using the pool because two writes must be atomic, use batchTransaction() instead: one HTTP round trip, one transaction, no connection.

Fix 2: put the compute next to the function

A query is a round trip. A function in iad1 talking to a Neon branch in eu-central-1 pays 90-120 ms per statement, before the database does any work. Five sequential queries in a page is half a second of pure geography, and no amount of indexing helps.

Pick the Neon region to match your primary Vercel region, and pin the function's region if your host lets you. This is the single largest latency fix available to most projects, and it is a dropdown.

While you are there: reduce the number of round trips. Two sequential awaits that do not depend on each other should be Promise.all. A page that issues seven queries against a database 100 ms away has a 700 ms floor.

Fix 3: cache the reads that can be stale

Most "the database is slow" traffic is the same handful of queries running on every request. Next.js gives you two levers before you reach for infrastructure:

export const revalidate = 60; // route segment: rebuild at most once a minute
import { unstable_cache } from "next/cache";

const getPlans = unstable_cache(
  async () => getSql()`select slug, name from plans where active = true`,
  ["plans"],
  { revalidate: 300, tags: ["plans"] },
);

A cached read never wakes the compute at all, which means it also never pays the half second. For a marketing page or a pricing table this turns the cold-start question into a non-question.

Fix 4: decide deliberately about keeping compute warm

Neon lets you raise the scale-to-zero delay (and on paid plans, disable it) per branch. That trades money for latency: an always-on compute bills continuously.

The honest framing:

  • Production with steady traffic never scales to zero anyway. Do nothing.
  • Production with bursty, latency-sensitive traffic: a raised suspend delay is a reasonable purchase.
  • Preview branches should absolutely scale to zero. A reviewer waiting half a second on their first click is a fair trade for previews being nearly free.

A cron job that pings the database to keep it warm is a worse version of raising the suspend delay: it costs the same compute, adds a moving part, and lies to your metrics.

What not to blame

Do not blame the pooler. The -pooler endpoint does not add meaningful latency; it removes connection-count problems. If someone proposes moving the app to the direct URL "for speed", that is a connection-exhaustion incident being scheduled.

Do not blame the branch. A branch is not slower than its parent. It has its own compute, so it has its own wake: that is all.

Do not add a read replica for cold-start latency. A replica has its own compute with its own scale-to-zero behaviour, so the first read after idle is just as slow, and you now have two of everything.

Reporting a cold start honestly

When you are asked whether the app is slow, separate the numbers: function cold start, compute wake, query time. "First request after five minutes idle: 1.8 s, of which ~500 ms is compute wake; warm requests: 180 ms" is an answer someone can act on. "The database is slow" is not.