This guide gives an instance the two probes an orchestrator asks, and a
shutdown that lets the load balancer move traffic away before anything is
cut. The shutdown is @tetsujs/lifecycle; its
package page lists every option.
Liveness asks whether the process works at all. A failure restarts
it, so it checks nothing but the process: a liveness check on the
database would restart every instance when the database goes down.
Readiness asks whether this instance can serve traffic now. A failure
takes it out of rotation until it passes again. It checks the
dependencies the instance cannot answer without, and it fails while the
instance is shutting down.
: the method is
half of a route's identity, and an application that remembers its
routes — for a generated client, for tooling — needs to know which one
this is.
method: "GET",
path: "/livez"
Route path with :param segments, e.g. "/orders/:id/cancel".
Must start with /, contain no empty segments and no trailing slash;
a malformed literal is a compile error.
path: "/livez",
docs?: RouteDocs |undefined
Documentation metadata for OpenAPI generation.
docs: {
hidden?: boolean |undefined
Keeps the route out of the generated document.
For endpoints that exist but are nobody's business to call: an
internal probe, an admin escape hatch, a route kept alive for one
legacy client. The route is served exactly as before — this is a
statement about the document, not about access, and a hidden route is
as reachable as any other.
false says the opposite out loud: the route is in the document even
when its handler was annotated to keep the routes it answers out of
it — a package's handler serving files, say. Left out, the handler
decides; written, the route does.
deprecated is the other half of the pair: an endpoint on its way out
stays in the document and says so, an endpoint that was never public
is simply absent.
hidden: true },
handler: (ctx: {
readonlyreq:Request& {
readonlycookies?:Bun.CookieMap;
};
readonlyserver:Bun.Server<unknown>;
readonlyout:Outgoing;
readonlyroute:RouteInfo;
readonlystartedAt:number;
readonlyparams: {};
}) => string
The endpoint logic; ctx is fully inferred, never annotate it.
The return type is inferred rather than demanded, and checked twice
over. Against the route's own contract, by the
HandlerResult
bound on R: answering with something response never declared is a
compile error that says so, instead of a structural diff against
Response. And against what the framework can serialize at all, by
the intersected
ValidateResult
: a stream handed over bare is
refused whether or not the route declared anything, because that is
the case no contract covers — without a response schema
HandlerResult is unknown and accepts every value there is.
handler: () =>"ok" }),
ready: route({
method: "GET"
The method this route answers.
Kept as a literal rather than widened to
Method
: the method is
half of a route's identity, and an application that remembers its
routes — for a generated client, for tooling — needs to know which one
this is.
method: "GET",
path: "/readyz"
Route path with :param segments, e.g. "/orders/:id/cancel".
Must start with /, contain no empty segments and no trailing slash;
a malformed literal is a compile error.
path: "/readyz",
docs?: RouteDocs |undefined
Documentation metadata for OpenAPI generation.
docs: {
hidden?: boolean |undefined
Keeps the route out of the generated document.
For endpoints that exist but are nobody's business to call: an
internal probe, an admin escape hatch, a route kept alive for one
legacy client. The route is served exactly as before — this is a
statement about the document, not about access, and a hidden route is
as reachable as any other.
false says the opposite out loud: the route is in the document even
when its handler was annotated to keep the routes it answers out of
it — a package's handler serving files, say. Left out, the handler
decides; written, the route does.
deprecated is the other half of the pair: an endpoint on its way out
stays in the document and says so, an endpoint that was never public
is simply absent.
hidden: true },
handler: (ctx: {
readonlyreq:Request& {
readonlycookies?:Bun.CookieMap;
};
readonlyserver:Bun.Server<unknown>;
readonlyout:Outgoing;
readonlyroute:RouteInfo;
readonlystartedAt:number;
readonlyparams: {};
}) =>Promise<string>
The endpoint logic; ctx is fully inferred, never annotate it.
The return type is inferred rather than demanded, and checked twice
over. Against the route's own contract, by the
HandlerResult
bound on R: answering with something response never declared is a
compile error that says so, instead of a structural diff against
Response. And against what the framework can serialize at all, by
the intersected
ValidateResult
: a stream handed over bare is
refused whether or not the route declared anything, because that is
the case no contract covers — without a response schema
HandlerResult is unknown and accepts every value there is.
The checks run side by side, each with its own deadline, so one that hangs
fails alone and the probe still answers in time. A refusal names the
checks that failed:
stopping is a function because the shutdown signal exists only once the
server does, and the server is built from this controller. The next
section wires it.
Keep each check cheap — select 1 on the pool, not a query over a table —
since every balancer asks every few seconds. Leave out dependencies the
instance can serve without: a shared dependency fails every instance’s
check at once, and the balancer is left with nowhere to send traffic.
docs: { hidden: true } keeps the probes out of the
OpenAPI document. To keep them out of
request logs and metrics, skip
them by record.route in accessLog()’s write. Mount a rate limit on
the groups it protects rather than on the application, so the probes are
not counted against it.
On SIGTERM or SIGINT, onShutdownSignals() aborts stopping, so
readiness starts failing, keeps serving for preStopDelayMs, then aborts
draining and stops the server, giving the requests in flight graceMs
(10 seconds) to finish. It then cuts what is left within forceMs
(1 second), runs the close functions and exits. The package page
lists every step and option.
Aborts the moment a stop is asked for, before anything else happens.
This is what makes preStopDelayMs worth having: a readiness endpoint
that reads it starts failing while the delay runs, the balancer stops
routing here, and only then does the server stop. Without it the delay
postpones the same cut instead of avoiding it.
A standard AbortSignal, so it composes with everything that already
takes one — a poll loop, a queue consumer, a fetch to an upstream
that is no longer worth waiting for.
It exists only once the server does, and the server is built from the
routes: a handler reads it when a request comes in, by which time it
is there, and a controller built from its dependencies takes it as a
function.
stopping.
aborted: boolean
The aborted read-only property returns a value that indicates whether the asynchronous operations the signal is communicating with are aborted (true) or not (false).
How long to keep serving after the stop was asked for, before the
server is told to stop at all. Zero by default.
This is the half of a graceful shutdown that stopping gracefully does
not cover. In Kubernetes the SIGTERM and the pod's removal from the
Service travel in parallel: the signal arrives at once, while the
endpoint change has to reach kube-proxy, the ingress and whatever
balancer sits in front, which takes hundreds of milliseconds and
sometimes seconds. Stopping the moment the signal lands therefore cuts
exactly the requests that are still being routed here — the ones this
package exists to protect.
So the sequence is: say we are not ready, keep serving for this long
while the news travels, and only then drain. A second signal cuts the
wait short, the same way it cuts the grace period short.
Zero by default because the delay is pure cost anywhere the caller is
not behind a balancer that has to be told — a test, a CLI, a process
nobody is routing to. Set it where something has to hear.
A delay on its own changes nothing. Something has to start
answering readiness with a failure while it runs, or the balancer goes
on sending traffic and the wait only postpones the same cut. See
stopping on the return of
onShutdownSignals
.
preStopDelayMs: 5_000,
close?: readonly Closer[] |undefined
Resources to release, in order, after the server has stopped.
After, never before: a request still in flight may reach for the pool
that closing it early would have taken away.
close: [() =>
constpool: {
query(sql:string):Promise<unknown>;
end():Promise<void>;
}
pool.end()],
});
The closers run after the server has stopped, so a request in flight does
not lose the pool it is using.
In Kubernetes, SIGTERM and the pod’s removal from the Service happen in
parallel, and the removal takes time to reach every proxy. A server that
stops the moment the signal lands cuts the requests still being routed to
it. So readiness fails first, and the server keeps serving for
preStopDelayMs while the balancers catch up.
Use the delay together with a failing readiness check. A balancer that
polls each instance learns about the shutdown only from the failing
answer; without it, the delay only postpones the same cut. Where nothing
has to be told — a test, a CLI — leave the delay at 0.
A shutdown takes at most preStopDelayMs + graceMs + forceMs — 16 seconds
with the settings above — plus the closers, which have no deadline. The
platform must wait at least that long before it kills the process; see
Kubernetes below. A second signal skips the delay and the
grace period but still cuts what is left and runs the closers; a third
ends the process at once.
An event stream or a long poll never finishes on its own. Left open, it
holds the stop for the whole of graceMs, and the process then exits with
1. Close it on draining, which aborts after the delay, when the server
starts to stop. An event stream takes it as
until.
The signal exists only once the server does, and the server is built from
the controllers, so a controller takes it as a function, as the health
controller takes stopping, and calls it when a request comes in:
: the method is
half of a route's identity, and an application that remembers its
routes — for a generated client, for tooling — needs to know which one
this is.
method: "GET",
path: "/feed"
Route path with :param segments, e.g. "/orders/:id/cancel".
Must start with /, contain no empty segments and no trailing slash;
a malformed literal is a compile error.
path: "/feed",
handler: (ctx: {
readonlyreq:Request& {
readonlycookies?:Bun.CookieMap;
};
readonlyserver:Bun.Server<unknown>;
readonlyout:Outgoing;
readonlyroute:RouteInfo;
readonlystartedAt:number;
readonlyparams: {};
}) => Response
The endpoint logic; ctx is fully inferred, never annotate it.
The return type is inferred rather than demanded, and checked twice
over. Against the route's own contract, by the
HandlerResult
bound on R: answering with something response never declared is a
compile error that says so, instead of a structural diff against
Response. And against what the framework can serialize at all, by
the intersected
ValidateResult
: a stream handed over bare is
refused whether or not the route declared anything, because that is
the case no contract covers — without a response schema
HandlerResult is unknown and accepts every value there is.
handler: (
ctx: {
readonly req: Request & {
readonly cookies?: Bun.CookieMap;
};
readonly server: Bun.Server<unknown>;
readonly out: Outgoing;
readonly route: RouteInfo;
readonly startedAt: number;
readonly params: {};
}
ctx) =>sse(
ctx: {
readonly req: Request & {
readonly cookies?: Bun.CookieMap;
};
readonly server: Bun.Server<unknown>;
readonly out: Outgoing;
readonly route: RouteInfo;
readonly startedAt: number;
readonly params: {};
}
ctx, feed, {
until?: AbortSignal |undefined
Ends the stream when it fires, as a client leaving would. What a
server that is stopping closes its streams on — see
StreamOptions.until
.
draining exists only once the server does, and the server is built
from the routes: the handler reads it when a request comes in, by
which time it is there.
until:
draining: () => AbortSignal
draining() }),
}),
}),
);
const
constserver:Bun.Server<unknown>
server= Bun.serve({
...createApp({
routes: feedController({
draining: () => AbortSignal
draining: () =>
constshutdown:ShutdownHandle
shutdown.
draining: AbortSignal
Aborts when the server starts to stop: after the pre-stop delay, the
moment server.stop() is called — or with stopping, when there is
no delay.
What a response that never ends on its own closes on. server.stop()
waits for every request in flight, and an event stream or a long poll
is one that is always in flight: left open, it holds the stop for the
whole of graceMs, and the process then exits with 1, its
connections cut. Closed here, it ends at once, and its client
reconnects to a server the balancer is still sending traffic to.
Not stopping: that one fires while this server is still being sent
traffic, so a client that reconnects at once lands here again, to be
closed again, for as long as the delay runs.
Read, like stopping, when a request comes in: the routes are built
before the server, and the signal only with it.
draining }),
}),
port?: string | number |undefined
The port the server listens on
@default ― process.env.PORT || "3000"
port: 3000,
});
const
constshutdown:ShutdownHandle
shutdown=onShutdownSignals(
constserver:Bun.Server<unknown>
server, {
preStopDelayMs?: number |undefined
How long to keep serving after the stop was asked for, before the
server is told to stop at all. Zero by default.
This is the half of a graceful shutdown that stopping gracefully does
not cover. In Kubernetes the SIGTERM and the pod's removal from the
Service travel in parallel: the signal arrives at once, while the
endpoint change has to reach kube-proxy, the ingress and whatever
balancer sits in front, which takes hundreds of milliseconds and
sometimes seconds. Stopping the moment the signal lands therefore cuts
exactly the requests that are still being routed here — the ones this
package exists to protect.
So the sequence is: say we are not ready, keep serving for this long
while the news travels, and only then drain. A second signal cuts the
wait short, the same way it cuts the grace period short.
Zero by default because the delay is pure cost anywhere the caller is
not behind a balancer that has to be told — a test, a CLI, a process
nobody is routing to. Set it where something has to hear.
A delay on its own changes nothing. Something has to start
answering readiness with a failure while it runs, or the balancer goes
on sending traffic and the wait only postpones the same cut. See
stopping on the return of
onShutdownSignals
.
preStopDelayMs: 5_000 });
By the time a request comes in, shutdown exists. A route written in the
same file as the server can read shutdown.draining in its handler
directly.
Use draining, not stopping: during the delay the balancer still sends
traffic here, and a client that reconnects at once would land on this
server again.
A WebSocket endpoint takes until too, and
closes its sockets with 1001, going away. The endpoint is declared before
the server exists, so until is a function, called as each socket opens:
Closes the endpoint's open sockets when it fires, with 1001 — going
away — and every socket opened after it at once.
What a server that is stopping closes its sockets on: draining from
@tetsujs/lifecycle. server.stop() waits for every open socket, and
a socket stays open for as long as its client wants: left open, one
held every deploy for the whole grace period, and was then cut with
1006 and no reason, and the process exited as a forced stop. Closed
here, its client sees a server going away and reconnects to one that
stays; close runs for each as for any other.
A signal, or a function that returns one, called as a socket opens.
The function is the usual form: an endpoint is declared before the
server exists, and draining only after — onShutdownSignals takes
the server the endpoint is part of. A socket for which it returns
nothing, or throws — reported, as websocket — is not closed by it.
until: () =>
constshutdown:ShutdownHandle
shutdown.
draining: AbortSignal
Aborts when the server starts to stop: after the pre-stop delay, the
moment server.stop() is called — or with stopping, when there is
no delay.
What a response that never ends on its own closes on. server.stop()
waits for every request in flight, and an event stream or a long poll
is one that is always in flight: left open, it holds the stop for the
whole of graceMs, and the process then exits with 1, its
connections cut. Closed here, it ends at once, and its client
reconnects to a server the balancer is still sending traffic to.
Not stopping: that one fires while this server is still being sent
traffic, so a client that reconnects at once lands here again, to be
closed again, for as long as the delay runs.
Read, like stopping, when a request comes in: the routes are built
before the server, and the signal only with it.
How long to keep serving after the stop was asked for, before the
server is told to stop at all. Zero by default.
This is the half of a graceful shutdown that stopping gracefully does
not cover. In Kubernetes the SIGTERM and the pod's removal from the
Service travel in parallel: the signal arrives at once, while the
endpoint change has to reach kube-proxy, the ingress and whatever
balancer sits in front, which takes hundreds of milliseconds and
sometimes seconds. Stopping the moment the signal lands therefore cuts
exactly the requests that are still being routed here — the ones this
package exists to protect.
So the sequence is: say we are not ready, keep serving for this long
while the news travels, and only then drain. A second signal cuts the
wait short, the same way it cuts the grace period short.
Zero by default because the delay is pure cost anywhere the caller is
not behind a balancer that has to be told — a test, a CLI, a process
nobody is routing to. Set it where something has to hear.
A delay on its own changes nothing. Something has to start
answering readiness with a failure while it runs, or the balancer goes
on sending traffic and the wait only postpones the same cut. See
stopping on the return of
onShutdownSignals
.
preStopDelayMs: 5_000 });
Without until, an open socket holds every stop for graceMs and is then
cut with 1006.
terminationGracePeriodSeconds is how long Kubernetes waits before
it sends SIGKILL. Keep it above preStopDelayMs + graceMs + forceMs
plus the time the closers take.
The readiness probe has to fail within preStopDelayMs: a short
periodSeconds and a failureThreshold of one or two.
The liveness probe is slower and more forgiving: a restart is
expensive, and a busy instance that misses one probe is not dead.
A preStop hook that sleeps does the same job as preStopDelayMs,
and the two add up, so use one or the other. preStopDelayMs needs
nothing in the image; a sleep command needs a binary that a minimal
image or a compiled executable may not have.
Startup needs nothing from the package: open the pool and run migrations
with await before Bun.serve, and readiness fails until the server
listens.
A process with more than one server — the API and a
metrics port, or a public API
and an admin one — passes them all to one onShutdownSignals call; see
Several servers. Scheduled
work stops with the server the same way; see
Background jobs.