Skip to main content

Points

The accelerator can write a point field as a multi-way tree (MKD) instead of Lucene's binary BKD tree. It does so per field, and only where a measurement supports it — not for every point field.

Which fields actually change​

A multi-way tree carries a fixed cost per leaf visit and only pays for it where there are enough indexed dimensions for exactly-tightened child bounds to prune substantially harder than Lucene's reconstructed cells. Measured:

Indexed dimensionsTypical fieldsStructureMeasured
1long, date, double, ipbinary (BKD)8.9% slower as a multi-way tree, 1.75x dearer to merge
2geo_pointbinary (BKD)7.4% slower on materialised box queries; distance queries at parity
4+geo_shape, range fieldsmulti-way (MKD)24% faster, reproducing on every clean run

So on a typical index the thing that gets faster is geo_shape. Numerics, dates, IP addresses and geo_point keep Lucene's binary tree, because that is what the measurements say is faster for them.

Routing one-dimensional fields to the multi-way tree is available as an opt-in — see One-dimensional fields — and is off by default.

Changing this policy never strands data. The structure chosen for a field is recorded on the field itself when it is written, and the read path uses that record rather than re-deciding. A segment is always read by whatever actually wrote it, so no migration or reindex is needed. Merges do re-apply the current policy, which means an index converges on it as it merges.

One-dimensional fields​

With one indexed dimension every fanout — and Lucene's binary tree — produce the identical partition, the same leaves holding the same points, so there is no cell shape to improve. What is left is the traversal, and with the bulk leaf read that measures 14.1% faster than the binary tree, at roughly 30% more merge cost.

Which of those dominates depends on the index, so it is a setting rather than a default:

PUT /reference-data
{
"settings": {
"index.lucenia.accelerator.enabled": true,
"index.lucenia.accelerator.points.one_dimension": true
}
}

Worth it for an index queried far more often than it is written — reference data, logs past their rollover, an analytics index rebuilt nightly. Not worth it for one under constant churn.

It opts in one indexed dimension and nothing else: deliberately not a lower threshold, because a threshold would sweep up two and three indexed dimensions on the way past — the shapes measured slower — and move an index's geometry along with its numerics.

Like index.lucenia.accelerator.enabled, it is final: it decides how segments are laid out, and a mid-life change would leave one index answering the same query two different ways depending on which segment it lands in.

It is a drop-in replacement​

Field encodings are entirely unaffected. LatLonPoint's quantised coordinates, LatLonShape's tessellated triangles, the sortable byte encodings behind IntPoint, LongPoint and DoublePoint, InetAddressPoint, and range fields all produce exactly the same packed bytes as before. Only the structure those bytes are organised into changes.

That means no mapping changes and no query changes. A geo_distance query, a date range, a numeric filter — all unchanged. You enable the setting and reindex.

PUT /events
{
"settings": { "index.lucenia.accelerator.enabled": true },
"mappings": {
"properties": {
"location": { "type": "geo_point" },
"timestamp": { "type": "date" }
}
}
}

Why a multi-way tree​

A binary tree splits one dimension in two at each level. A multi-way tree splits into more children per level, so the tree is shallower for the same number of points, and each descent step reads a block rather than following a pointer.

   BKD (binary)                     MKD (multi-way)

● ●
/ \ / / | \ \
● ● ● ● ● ● ●
/ \ / \ ...
● ● ● ● fewer levels,
deeper walk wider blocks

For a range or distance query, which descends the tree and then scans leaves, fewer levels means fewer dependent reads before the scan starts.

On-disk files​

ExtensionContents
.mkdmmeta — per-field geometry, global bounds, counts, root pointer
.mkdiindex — inner-node blocks, written as a forward-only depth-first stream
.mkdddata — leaf blocks

The index file being a forward-only stream is the point worth noticing: descending the tree moves forward through the file, so the access pattern is sequential rather than a scatter of seeks. That is what turns a shallower tree into less I/O rather than merely fewer levels.

Compatibility​

These files are written by the accelerator's codec. A node without the plugin cannot read them — see Compatibility and the warning about removing the plugin in the accelerator overview.