Skip to main content

Vectors

The accelerator adds a vector field type backed by binary quantization: one bit per dimension, stored alongside the raw vectors, with an HNSW graph over it.

The field type​

PUT /docs
{
"settings": { "index.lucenia.accelerator.enabled": true },
"mappings": {
"properties": {
"embedding": {
"type": "lucenia_vector",
"dimension": 768,
"similarity": "cosine"
}
}
}
}
ParameterRequiredDefaultValues
dimensionyes—Number of dimensions.
similaritynoeuclideaneuclidean, cosine (alias cosinesimil), dot_product (alias innerproduct), max_inner_product

The aliases exist so mappings written for other vector plugins port without editing.

Querying​

GET /docs/_search
{
"query": {
"lucenia_knn": {
"embedding": {
"vector": [0.12, 0.44, 0.91, 0.03],
"k": 10
}
}
}
}

The object key is the field name, and the vector must have exactly the dimension that field was mapped with — the example above is shortened for readability. boost and _name are supported as on any query.

Why not k-NN's knn_vector?​

lucenia_vector does not layer on top of the k-NN plugin's field type, and it cannot, in either of k-NN's configurations:

  • With index.knn: true, k-NN claims the one codec service an index is allowed. An index cannot be both a k-NN index and an accelerated one.
  • With it unset, k-NN's flat mapper writes into binary doc values rather than a Lucene vector field — so there is no graph to search, and a query returns nothing.

The two coexist in a cluster on different indices. They cannot coexist on the same index.

What binary quantization actually costs​

The compression figure is easy to misread, so it is worth stating plainly.

   per vector, 768 dims, float32
┌─────────────────────────────────────────────┐
│ raw vectors 3072 bytes │ ← still written
├─────────────────────────────────────────────┤
│ quantized 96 bytes (1 bit/dim) │ ← added
└─────────────────────────────────────────────┘
total on disk: slightly MORE than before
hot working set: 96 bytes instead of 3072

This format increases total on-disk size slightly. The raw vectors are still stored, and the quantized form is written in addition. The win is in memory and traversal I/O — the part that is read on every query — not in disk footprint.

The raw vectors cannot be dropped, for two independent reasons:

  • Merging. Quantization is lossy, and a quantized vector is only meaningful against the centroid and rotation it was encoded with. A merged segment generally wants a new centroid, so it cannot be produced by transcoding quantized payloads — re-quantizing needs full precision.
  • Rescoring. One-bit scores are estimates. Recovering exact ordering means rescoring an oversampled window of candidates against the originals.

Lucene's own quantized formats take the same approach, so this follows that precedent rather than inventing a variation.

Measured on 2,000 documents of 768 dimensions, with and without the format:

  .vec   raw float32   6,144,080 -> 6,144,080    unchanged, byte for byte
.bbqd quantized 0 -> 216,073 added
TOTAL 6,161,369 -> 6,380,643 1.036x

So the quantized form really is 32× smaller than the raw vectors — 192,000 bytes against 6,144,000 — and the index is nonetheless 3.6% larger, because the one is added to the other rather than replacing it. BinaryQuantizedFootprintTests re-derives those numbers on every build.

Making the index smaller: derived source​

The disk saving is real, but it is not in the vector data. It is in _source.

A vector is written three times by default: as the Lucene vector data the graph walks, as the quantized form the first pass scores against, and as JSON text inside _source. The third copy is the largest of the three and the only one nothing reads. A float32 value is 4 bytes of vector data; rendered as a decimal number with a separator it is typically 10 to 12 bytes, and those digits are close to random so they compress badly.

index.lucenia.accelerator.vectors.derived_source.enabled: true

With this on, the stored-fields writer drops those values on the way in and the reader puts them back from the vector data on the way out. Measured on 500 documents of 768 dimensions:

  .fdt   stored fields  3,042,924 ->     7,680    396x smaller
.vec raw float32 1,536,144 -> 1,536,144 unchanged - it is what reconstruction reads
TOTAL 4,575,712 -> 1,548,192 0.338x

The whole index is about three times smaller, and the saving grows with dimension count because the JSON copy does while the vector data does not.

Reconstruction is exact — the same float32 values, read back from the copy that was kept — so _get, search hits, reindex and update-by-query all still see a complete document. The cost is that a source fetch now reads the vector data, which is the right trade for almost any vector index and is still a decision to make rather than one to have happen on upgrade.

Two properties worth knowing before turning it on:

  • It is Final. An index that stripped its source cannot un-strip the segments it already wrote. A toggle would leave one index answering _get two ways depending on which segment a document landed in.
  • It is recorded per segment, not read from settings at query time. Lucene resolves a segment's codec by name with no index settings in hand, so each segment says for itself whether its source was stripped — which is also what lets segments written before you enabled it keep working.

It is done in the codec rather than in an indexing listener, so the translog still holds the document you sent: recovery, replication and reindex-from-remote need no special handling.

How a query runs​

   query vector
│
├─▶ HNSW walk over 1-bit codes cheap, approximate
│ │
│ └─▶ oversampled candidate window
│
└─▶ rescore the window against raw vectors exact ordering
│
└─▶ top k

The approximate pass decides which vectors are worth looking at; the rescore decides the order they come back in. Recall comes from the width and patience of the walk, and precision from the rescore.

Widening the first pass​

The walk is asked for more candidates than you want, because a one-bit code finds the right neighbourhood and cannot rank it. How much more is oversample:

{
"query": {
"lucenia_knn": {
"embedding": {
"vector": [0.12, -0.04, 0.88],
"k": 10,
"oversample": 3.0
}
}
}
}
OptionDefaultMeaning
oversample3.0How much wider than k the first pass searches. Must be at least 1.

The default is 3.0 because one bit per dimension is the lossiest encoding this format writes — lossier than the scalar-quantized cases smaller factors were chosen for.

The floor matters more than the factor. The first pass never collects fewer than 100 candidates and never more than 10,000. At a typical k of 10 the factor alone would rescore thirty candidates, which is close enough to k that the reordering would barely change the answer; the floor is what makes the rescore meaningful for the queries people actually run. The cap is on the candidate count rather than on the ratio, because the count is what the walk costs — bounding the ratio instead would remove usable operating points at small k while protecting nothing.

Raise oversample when recall matters more than latency: it widens the candidate set, and the candidate set is the recall — the rescore only orders what the walk already found. Lowering it toward 1 moves the answer back toward the raw one-bit ordering.

When it helps​

Binary quantization pays off when the vector working set no longer fits comfortably in memory — that is the regime it was built for, and where reading 96 bytes instead of 3072 per candidate changes the shape of the query. Below that threshold, a plain float32 index may well be faster, and is certainly simpler.

Measure on your own data before switching. Recall is data-dependent: quantization error interacts with how your embeddings are distributed, and a model whose vectors cluster tightly loses more information per bit than one that spreads them out.

If recall is short of what you expect, raise oversample before concluding the format is wrong for your data — a wider first pass is the cheapest lever available and needs no reindex.