Skip to main content

Scale targets

A scale target is the thing the controller actually resizes. Which one you want depends on who owns the cluster's desired state, and getting that wrong produces a specific, confusing failure.

ModeTargetActuatesHandles drain
recommendnothingnon/a
kubernetesa StatefulSetyesno — the controller drains
custom_resourcean operator's CRyesyes — the operator drains

recommend​

The default. Evaluates everything, logs what it would have done, changes nothing.

Use it to build confidence: run it for a week against production and read the recommendations before you let it act. It is also the right permanent mode when a human approves every capacity change.

kubernetes — StatefulSet​

Scales a StatefulSet directly.

octane:
mode: kubernetes
kubernetes:
namespace: search
statefulsets:
hot: opensearch-hot
warm: opensearch-warm

The controller drains the node itself before removing it, because a StatefulSet does not know what a shard is. Scaling in removes the highest-ordinal pod, which is what nodeRemovedOnScaleIn predicts so the right node is drained rather than an arbitrary one.

custom_resource — an operator's CR​

Patches a field in a custom resource and lets the operator do the rest.

octane:
mode: custom_resource
kubernetes:
namespace: search
custom_resource: clusters.lucenia.io/v1/luceniaclusters/prod
replica_paths:
hot: spec.nodePools.hot.replicas
warm: spec.nodePools.warm.replicas

Use this whenever an operator owns the cluster​

This is not a preference. An operator reconciles a StatefulSet it owns straight back to its own spec. In kubernetes mode against an operator-managed cluster, the sequence is:

   controller  ──▶  StatefulSet: replicas 6 → 5
│
operator ──▶ reconciles: replicas 5 → 6 (its spec still says 6)
│
controller ──▶ still over capacity: 6 → 5
│
... forever

The tier returns to its old size seconds later, repeatedly, and the drain work is wasted each cycle. Patch the resource the operator reads from, and the operator scales the StatefulSet itself.

The patch is a compare-and-set​

The controller writes an RFC 6902 JSON Patch with a test operation on metadata.resourceVersion. If the resource changed since it was read, the patch is rejected rather than applied to a version the controller never saw. That is what keeps a controller and a human editing the same resource from overwriting each other.

Drains belong to the operator here​

An operator that manages node pools generally drains for itself. The target declares handlesDrain() == true and the controller does not write its own allocation exclusion — two components draining the same node would fight, and the loser leaves an exclusion behind.

The ScaleTarget contract​

A target implements a small interface, so a new one is a class rather than a fork:

MethodMeaning
describe()What this target is, for logs.
actuates()Whether it changes anything. false for recommendation-only.
scaleTo(tier, desiredNodes)Resize.
handlesDrain()Whether the target drains, so the controller should not.
nodeRemovedOnScaleIn(tier, currentNodes)Which node will go, so the right one is drained.
supportsRestart()Whether the target can restart a node.
restartNode(tier, nodeName)Restart one node.

nodeRemovedOnScaleIn is the method that makes a drain correct rather than hopeful. Draining a node the platform was not about to remove empties the wrong node and then removes a full one.

What is not implemented​

There is no AWS/managed-service scale target. A target that called a provider's domain-config API would fit the interface, and nothing here implements one — the three above are what ships. For a managed cluster today, recommend mode works (it is REST-only) and the resize is yours to apply.