Troubleshooting
The controller will not start
It refuses a configuration it cannot act on, and names the problem. See
Configuration for the
full list. The most common are a missing namespace in a Kubernetes mode, and custom_resource mode
with no replica_paths.
Nothing is being scaled
Check the mode. recommend is the default, and it changes nothing by design:
octane:
mode: recommend # ← logs what it would do; actuates nothing
If the mode is right, check the license. Without octane.license_path, or with an expired license,
committing operations are refused and the controller falls back to recommending. The log names the
operation and the reason.
A tier scales and immediately scales back
This is the operator-reconcile loop. You are in kubernetes mode against a cluster an operator owns:
the controller resizes the StatefulSet, and the operator puts it back.
Switch to custom_resource mode and patch what the operator reads from. See
Scale targets.
A drain never completes
The safety service returns Wait for transient conditions, and it will keep waiting as long as the
condition holds. The verdict's reason says which:
| Reason | What to do |
|---|---|
| Snapshot in progress | Wait, or reschedule the snapshot. |
| Active shard operations on an index | Wait. |
| Recovery pressure is high | Wait; the cluster is already moving data. |
| Draining would leave a primary shard without a copy | Add capacity, or reduce the scale-in step. |
| Cannot drain the active cluster-manager node | Unsafe, not Wait — it will never proceed. |
Unsafe verdicts will not resolve by waiting. Something about the request has to change.
A node is excluded and nothing is reclaiming it
An allocation exclusion left behind means a drain was interrupted. The controller resumes drains from
.octane-drains on startup, so the first thing to check is that a controller is actually running.
Cancelling a drain is never license-gated, precisely so this state is always recoverable.
The cluster will not rebalance after a rolling restart
Check whether the allocation clamp was left on:
curl -s localhost:9200/_cluster/settings | grep allocation.enable
The coordinator restores the prior value when a plan ends, including when it ends badly, and resumes
plans from .octane-rolling-restarts after a crash. If a clamp is still set and no plan is running,
clear it:
curl -XPUT localhost:9200/_cluster/settings -H 'Content-Type: application/json' \
-d '{"transient":{"cluster.routing.allocation.enable":null}}'
A rolling restart gave up on a node
Each node has 15 minutes to go away and come back, measured by its JVM start time changing. A node that exceeds it is usually a pod that cannot be scheduled or an image that will not pull — check the platform, not the cluster.
Where the controller keeps state
| Index | Holds |
|---|---|
.octane-autoscale-policies | Policies, one document per tier |
.octane-drains | Drains in progress |
.octane-drain-leases | Node reservations held by external readers |
.octane-rolling-restarts | Restart plans |
All are hidden, and all use sequence-number compare-and-set. Deleting them loses in-progress operations — including the record of an allocation exclusion that then has to be cleared by hand.
Drain leases
An external reader — a Spark job reading shards directly — can take a lease that pins nodes against
drain. While it is held, the safety service returns Wait for those nodes.
Leases expire by absolute wall-clock time. There is no background reaper: expiry is evaluated lazily on read and pruned opportunistically on write, so a controller that is down does not extend anybody's lease.