Autoscale
A controller that grows, shrinks and rolls an OpenSearch cluster without losing a shard copy. It runs beside the cluster and speaks to the REST API — nothing is installed on a node.
Why that shape
┌──────────────┐ REST ┌────────────────────┐
│ controller │ ──────────────────▶ │ OpenSearch │
│ (a process) │ │ (stock, unmodified)│
└──────┬───────┘ └────────────────────┘
│
│ scale / restart
▼
StatefulSet, or an operator's custom resource
Being outside the cluster is the point:
- No per-node install, so adopting it needs no restart.
- No version pin. It works against any 3.x, and there is no plugin to rebuild on upgrade.
- It works on a managed service, where Octane's plugins cannot be installed.
- It survives the cluster. A controller inside the cluster cannot safely restart the cluster it lives in; one outside can.
What it does each cycle
requested roll? → an operator asked for a rolling restart; do that instead
│
evaluate deciders → what does each tier need?
│
safety check → is it safe to remove this node right now?
│
drain → write an allocation exclusion, wait for shards to leave
│
poll → has the node actually emptied?
│
scale → resize the StatefulSet or custom resource
│
release → give back the lease, clear the exclusion
A rolling restart outranks autoscaling for its duration. The two would otherwise disagree about the same nodes — a roll deliberately takes a node away while the deciders see the capacity drop and ask for a replacement, and the cluster oscillates in the middle of an upgrade. Asking for one is a single document; see rolling restart.
Every step is resumable. State lives in indices on the cluster with sequence-number compare-and-set, so a controller that is killed mid-drain picks the drain back up rather than abandoning a half-emptied node. That is also why there is no single-writer election to run: the compare-and-set is the mutual exclusion.
Pages
- Configuration — the YAML file, and what each mode requires
- Deciders — how capacity is decided, and every threshold
- Scale targets — StatefulSet, custom resource, recommendation-only
- Rolling restart — how a fleet-wide plugin install gets done safely
- Troubleshooting
It defaults to doing nothing
The default mode is recommend: evaluate everything, log what it would have done, change nothing.
A controller that is running but misconfigured therefore produces a log entry rather than a resize.
Turn on actuation deliberately, once the recommendations look right.
Safety is a veto, not a vote
Before any node is drained or restarted, the safety service returns one of three verdicts:
| Verdict | Meaning |
|---|---|
Safe | Proceed. |
Wait(reason, retryAfter) | Not now — a transient condition. Retry after the stated interval. |
Unsafe(reason) | Never, under this configuration. Do not retry blindly. |
Wait and Unsafe are deliberately distinct. A snapshot in progress is a Wait — it will finish.
Draining the active cluster-manager is Unsafe — waiting will not make it acceptable. Collapsing
them into one "no" would either spin forever on a permanent condition or give up on a temporary one.
Conditions the service checks include: a snapshot in progress (cluster-wide or on a specific index), active shard operations on an index, high recovery pressure, draining the active cluster-manager, closing or reducing replicas on a system index, and whether draining would leave a primary shard without a copy elsewhere.
Licensing
The controller verifies a license file at startup, because it cannot read cluster state. Without one it runs in recommendation mode. Committing operations — scale out, start scale-in, plan a rolling restart — are gated; unwinds never are. See Licensing.