Skip to main content

Rolling restart

Restarting every node of a tier, one at a time, without triggering a cluster-wide replica rebuild. This is how a fleet-wide plugin install actually gets done: the accelerator and content plugins take effect on restart, and doing that by hand across fifty nodes is how clusters get hurt.

Asking for one​

A roll is requested by writing a document to the cluster. The controller picks it up on its next interval, plans the roll over every node in that tier, and works through it.

curl -XPOST "$CLUSTER/.octane-restart-requests/_doc?refresh=true" \
-H 'Content-Type: application/json' -d '{
"tier": "hot",
"roles": ["data"],
"reason": "install octane-accelerator 1.1.0.0",
"requested_at_millis": '"$(date +%s000)"'
}'

A document rather than an endpoint, because the controller is a polling loop with no HTTP surface of its own. Giving it one would mean a port to expose, a certificate to manage and an authentication story to invent — for a process whose whole job is talking to a cluster you already have credentials for. It also makes the request auditable: who asked, when, and why, in an index that outlives the controller process.

roles selects the nodes and is not optional in practice. A tier named with no roles matches nothing, and the roll is refused rather than silently doing nothing.

What you will see in the controller's log:

[prod] rolling restart requested for tier [hot]: install octane-accelerator 1.1.0.0 (6 nodes)
[prod] rolling restart: RESTART_REQUESTED - restarting prod-hot-2
[prod] rolling restart: WAITING_FOR_NODE - prod-hot-2 has not returned

What happens to the request​

It is consumed as the plan is created — not before. Consuming it first would lose the request if planning failed, and an operator whose request vanished with nothing happening cannot tell that from one that was never written. Consuming it at all is what stops a controller restart from rolling the fleet a second time.

A request the controller refuses — a tier with no matching nodes, a license that does not cover it, a scale target that cannot restart nodes — is also consumed, with the reason logged. The alternative is retrying the same impossible request every interval forever.

A request that names no tier is left in place rather than consumed. Deleting it would hide the mistake while you waited for a roll that was never going to happen.

Watching and stopping​

# progress and history
curl -s "$CLUSTER/.octane-rolling-restarts/_search?pretty"

Deleting the active plan stops the roll; the controller restores the allocation setting it found and stands down. A roll already running refuses a second request rather than interleaving two plans over one fleet.

A rolling restart is not a drain​

The obvious implementation is to reuse the scale-in path — it already knows how to empty a node safely. It is the wrong answer, and the reason is worth understanding because it is the difference between an operation that takes minutes and one nobody will ever run.

   DRAIN each node                    CLAMP, then restart each node

node 1: move all shards off allocation.enable: primaries
node 2: move all shards off node 1: restart, comes back with its data
node 3: move all shards off node 2: restart, comes back with its data
... × 50 ... × 50
restore allocation.enable

the whole dataset moves almost nothing moves
once per node

A drained node is one the cluster has been taught to live without. A restarted node keeps its data on disk and brings it back. Draining first moves the entire dataset once per node to reach a state the restart reaches anyway.

What replaces the drain​

cluster.routing.allocation.enable: primaries

This stops the cluster rebuilding replicas for a node that is coming back shortly. Without it, a node away for ninety seconds triggers a full replica rebuild, and the roll generates more data movement than a drain would have.

The clamp is written once for the whole plan and restored when it ends — including when it ends badly. That is the failure mode worth guarding: a clamp left in place is invisible until the day the cluster needed to rebalance and quietly did not.

The prior value is restored verbatim, not reset to a default. A cluster that was deliberately running with allocation restricted before the roll is left as it was found.

Ordering​

Cluster-manager-eligible nodes go last. An election is cheap but not free, and doing the data nodes first means most of the roll happens against a stable manager.

Knowing a node came back​

Return is detected by the node's JVM start time changing, not by it being reachable. A node that never went down would otherwise look like a node that restarted instantly, and the roll would march on having restarted nothing.

Each node gets 15 minutes to go away and come back. That is generous on purpose — a large node replaying a long translog can take a while, and a roll that declares failure early leaves a mixed-version fleet for no reason. The deadline exists for the case that genuinely does not resolve: a pod that cannot be scheduled, an image that does not pull. Waiting forever there would mean the clamp stays on forever.

Safety still applies​

Each node is checked before it is restarted — a snapshot in progress, active shard operations, high recovery pressure, or restarting the active cluster-manager all produce a Wait or Unsafe verdict. See the safety verdicts.

Resumability​

Restart plans live in .octane-rolling-restarts with compare-and-set writes. A controller killed mid-roll resumes the plan — including restoring the allocation clamp — rather than leaving a half-rolled tier and a clamped cluster.

Licensing​

Planning a rolling restart is a gated operation. Finishing one already under way is not, and must never be: a lapsed license part-way through a roll would otherwise leave allocation clamped with no way to clear it. See Licensing.