Skip to main content

Monitoring

Octane exposes no metrics endpoint of its own. What it gives you is a thread pool that appears in node stats, a license endpoint that answers in one call, and a controller log written to be read. This page says what is worth watching and — as importantly — what looks alarming and is not.

The query embedding pool​

query_embedding is a dedicated pool, so embedding a query cannot consume the search pool. It shows up in node stats under its own name:

curl -s 'localhost:9200/_nodes/stats/thread_pool/query_embedding?pretty'
SettingDefaultMeaning
thread_pool.query_embedding.size8Embeddings in flight at once
thread_pool.query_embedding.queue_size64How many may wait

The pool is wider than a compute-bound one on purpose: its threads block on a remote provider rather than burning CPU.

rejected is the signal. The queue is bounded deliberately, so provider slowness surfaces as a visible rejection rather than as creeping query latency. A non-zero and rising rejected means the provider cannot keep up — widen the pool only if the provider can actually take the concurrency, otherwise you have moved the queue rather than drained it.

A rising queue with a flat rejected is the earlier warning for the same thing.

License status​

One call, and it answers without consulting anything outside the cluster:

curl -s localhost:9200/_lucenia/license

Worth alerting on two things: status leaving active, and expires_in_days falling below whatever notice period you need to get a renewal installed. Expiry is not an outage — existing indices and pipelines keep working — but it stops you creating anything new, and that tends to be discovered at the worst moment. See Licensing.

The controller's log​

The controller's output is its interface in recommend mode, and its levels are chosen so the log can be alerted on directly — with one exception.

LevelWhat it meansAlert?
errorThe controller could not do its jobYes
warnA pass or a policy did not completeMostly — see below
infoWhat each pass did, per policyNo

The exception: a licensing refusal is logged at warn, and it is not a fault. When the license does not cover an operation the controller reports it and carries on managing the other policies. The message names the policy and the reason. Alerting on every warn will page somebody for a cluster that is behaving exactly as designed on an expired license.

Warnings that are worth paging on: a pass failing repeatedly for the same cluster, a policy that cannot be evaluated, and the shutdown warning that a pass did not finish within 30 seconds.

A controller managing several clusters contains each cluster's failures — one unreachable cluster does not stop the other nine — so a warning names the cluster id and it is worth alerting per id rather than in aggregate.

What a healthy controller looks like​

Each pass logs one line per policy. A steady tier logs at debug, so a quiet log is the normal state once everything is settled — but the controller deliberately logs once per cluster on startup, because a controller that connects, evaluates a healthy cluster and then says nothing is indistinguishable from one that is broken.

If you see nothing at all after startup, check mode before assuming a fault: Troubleshooting.

Graph traversal memory​

Graph traversal keeps per-segment caches — adjacency, and the contraction hierarchy used for routing — outside the Lucene commit. They are rebuilt on demand and never enter a snapshot.

There is no metric for them. What to watch instead is node heap on the nodes serving traversal queries, and the traversal scratch directory's disk usage. See Graph.

What is not instrumented​

Being explicit so nobody builds a dashboard expecting it:

  • No Prometheus endpoint, no stats API, and no per-processor counters.
  • Ingest processors report through OpenSearch's own ingest stats, not through anything Octane adds.
  • The controller keeps no history — every decision is made from the cluster as it stands at that moment, so there is no trend series to scrape from it. The record of what it did is the log and the documents in .octane-drains and .octane-rolling-restarts.