Someone reports that the app is slow. You open it and the first page takes two seconds; you reload and it takes 200 ms. Nothing in the code is conditional on "first request". Two independent cold starts are stacking up, and they have different fixes.
Where the time actually goes
A first request to an idle serverless app on an idle Neon branch pays, in order:
| Stage | Typical | What it is |
|---|---|---|
| Function cold start | 100-400 ms | Runtime boot, your module graph evaluated |
| Neon compute wake | ~500 ms | The branch's compute was scaled to zero |
| TLS + auth to Postgres | 20-80 ms | Only if you open a socket |
| First query | your query | The actual work |
The second and third are the ones people misattribute. Neon scales a branch's compute to zero after a period of inactivity (five minutes by default) and resumes it on the next connection. That resume is roughly half a second. It is not a bug, it is the reason the branch costs nothing while idle, and it is why branch-per-preview is affordable at all.
Measure before you fix
const started = Date.now();
const rows = await sql`select 1 as ok`;
console.log("db round trip", Date.now() - started, "ms");
checkDatabase() in src/db/health.ts already reports exactly this, and
bun run verify prints it. Run it twice in a row:
- First call slow, second call fast → compute wake. Read on.
- Both slow → not a cold start. Look at the query, the region pair, or a missing index.
- First call slow only after a deploy → function cold start, not the database.
Fix 1: use the HTTP driver
getSql() sends one HTTPS request per statement. There is no socket to open, no
TLS handshake to a Postgres backend, no pool to warm. On a cold instance this is
measurably the fastest way to get your first row, and it is why it is the default
in src/db/client.ts.
getPool() opens a real connection: TCP, TLS, Postgres auth, and a WebSocket
upgrade through Neon's proxy. Worth it when a request needs an interactive
transaction; pure overhead when it does not.
If you find yourself using the pool because two writes must be atomic, use
batchTransaction() instead: one HTTP round trip, one transaction, no
connection.
Fix 2: put the compute next to the function
A query is a round trip. A function in iad1 talking to a Neon branch in
eu-central-1 pays 90-120 ms per statement, before the database does any
work. Five sequential queries in a page is half a second of pure geography, and
no amount of indexing helps.
Pick the Neon region to match your primary Vercel region, and pin the function's region if your host lets you. This is the single largest latency fix available to most projects, and it is a dropdown.
While you are there: reduce the number of round trips. Two sequential awaits
that do not depend on each other should be Promise.all. A page that issues
seven queries against a database 100 ms away has a 700 ms floor.
Fix 3: cache the reads that can be stale
Most "the database is slow" traffic is the same handful of queries running on every request. Next.js gives you two levers before you reach for infrastructure:
export const revalidate = 60; // route segment: rebuild at most once a minute
import { unstable_cache } from "next/cache";
const getPlans = unstable_cache(
async () => getSql()`select slug, name from plans where active = true`,
["plans"],
{ revalidate: 300, tags: ["plans"] },
);
A cached read never wakes the compute at all, which means it also never pays the half second. For a marketing page or a pricing table this turns the cold-start question into a non-question.
Fix 4: decide deliberately about keeping compute warm
Neon lets you raise the scale-to-zero delay (and on paid plans, disable it) per branch. That trades money for latency: an always-on compute bills continuously.
The honest framing:
- Production with steady traffic never scales to zero anyway. Do nothing.
- Production with bursty, latency-sensitive traffic: a raised suspend delay is a reasonable purchase.
- Preview branches should absolutely scale to zero. A reviewer waiting half a second on their first click is a fair trade for previews being nearly free.
A cron job that pings the database to keep it warm is a worse version of raising the suspend delay: it costs the same compute, adds a moving part, and lies to your metrics.
What not to blame
Do not blame the pooler. The -pooler endpoint does not add meaningful
latency; it removes connection-count problems. If someone proposes moving the app
to the direct URL "for speed", that is a connection-exhaustion incident being
scheduled.
Do not blame the branch. A branch is not slower than its parent. It has its own compute, so it has its own wake: that is all.
Do not add a read replica for cold-start latency. A replica has its own compute with its own scale-to-zero behaviour, so the first read after idle is just as slow, and you now have two of everything.
Reporting a cold start honestly
When you are asked whether the app is slow, separate the numbers: function cold start, compute wake, query time. "First request after five minutes idle: 1.8 s, of which ~500 ms is compute wake; warm requests: 180 ms" is an answer someone can act on. "The database is slow" is not.