Skip to main content

Troubleshooting

The controller will not start​

It refuses a configuration it cannot act on, and names the problem. See Configuration for the full list. The most common are a missing namespace in a Kubernetes mode, and custom_resource mode with no replica_paths.

Nothing is being scaled​

Check the mode. recommend is the default, and it changes nothing by design:

octane:
mode: recommend # ← logs what it would do; actuates nothing

If the mode is right, check the license. Without octane.license_path, or with an expired license, committing operations are refused and the controller falls back to recommending. The log names the operation and the reason.

A tier scales and immediately scales back​

This is the operator-reconcile loop. You are in kubernetes mode against a cluster an operator owns: the controller resizes the StatefulSet, and the operator puts it back.

Switch to custom_resource mode and patch what the operator reads from. See Scale targets.

A drain never completes​

The safety service returns Wait for transient conditions, and it will keep waiting as long as the condition holds. The verdict's reason says which:

ReasonWhat to do
Snapshot in progressWait, or reschedule the snapshot.
Active shard operations on an indexWait.
Recovery pressure is highWait; the cluster is already moving data.
Draining would leave a primary shard without a copyAdd capacity, or reduce the scale-in step.
Cannot drain the active cluster-manager nodeUnsafe, not Wait — it will never proceed.

Unsafe verdicts will not resolve by waiting. Something about the request has to change.

A node is excluded and nothing is reclaiming it​

An allocation exclusion left behind means a drain was interrupted. The controller resumes drains from .octane-drains on startup, so the first thing to check is that a controller is actually running.

Cancelling a drain is never license-gated, precisely so this state is always recoverable.

The cluster will not rebalance after a rolling restart​

Check whether the allocation clamp was left on:

curl -s localhost:9200/_cluster/settings | grep allocation.enable

The coordinator restores the prior value when a plan ends, including when it ends badly, and resumes plans from .octane-rolling-restarts after a crash. If a clamp is still set and no plan is running, clear it:

curl -XPUT localhost:9200/_cluster/settings -H 'Content-Type: application/json' \
-d '{"transient":{"cluster.routing.allocation.enable":null}}'

A rolling restart gave up on a node​

Each node has 15 minutes to go away and come back, measured by its JVM start time changing. A node that exceeds it is usually a pod that cannot be scheduled or an image that will not pull — check the platform, not the cluster.

Where the controller keeps state​

IndexHolds
.octane-autoscale-policiesPolicies, one document per tier
.octane-drainsDrains in progress
.octane-drain-leasesNode reservations held by external readers
.octane-rolling-restartsRestart plans

All are hidden, and all use sequence-number compare-and-set. Deleting them loses in-progress operations — including the record of an allocation exclusion that then has to be cleared by hand.

Drain leases​

An external reader — a Spark job reading shards directly — can take a lease that pins nodes against drain. While it is held, the safety service returns Wait for those nodes.

Leases expire by absolute wall-clock time. There is no background reaper: expiry is evaluated lazily on read and pruned opportunistically on write, so a controller that is down does not extend anybody's lease.