Skip to main content
Version: 0.12.0

Tools and processors reference

This page is the single entry point for everything Lucenia can do to your data: the tools an agent can call, the ingest processors that run while documents are written, and the search processors that run while queries are served.

note

The word tool means two different things in these docs. The Ingestion tools and utilities section covers third-party data shippers (Beats, Logstash, Fluentd, OpenTelemetry), the Lucenia CLI, and migration tooling — programs that run outside the cluster. This page covers the tools and processors that ship inside the cluster.

Ask the cluster, not the docs

A running cluster is always the authoritative inventory for the version you are on.

SurfaceRequestNotes
Agent toolsGET /_plugins/_ml/toolsReturns name, type, description, and version for every registered tool type.
Ingest processorsGET /_nodes/ingest?filter_path=nodes.*.ingest.processorsComplete. Includes processors contributed by every installed module.
Search processorsGET /_nodes/search_pipelinesIncomplete — see the warning below.
warning

GET /_nodes/search_pipelines reports only request_processors and response_processors. It omits search phase results processors, which is where normalization-processor lives — the processor that makes hybrid search work. Do not treat its output as the full search processor list. The complete list is in Search processors.

Agent tools

A tool is a reusable building block an agent calls to perform one specific task — search an index, run a deployed model, call an external service. You reference a tool by its type when registering an agent. For parameters, examples, and per-tool detail, see Tools.

Tool typeWhat it does
AgentToolRuns another agent by its agent ID.
CatIndexToolRetrieves detailed index information for the Lucenia cluster (health, status, document counts, and store sizes).
ConnectorToolInvokes an external service through a configured connector.
IndexMappingToolRetrieves mapping and setting information for one or more indexes.
ListIndexToolLists the indexes in the cluster, along with their health, status, and document counts.
McpSseToolInvokes a tool hosted on a remote Model Context Protocol (MCP) server. See Using MCP tools.
MLModelToolRuns any deployed machine learning model by its model ID.
QueryPlanningToolTurns a natural-language question into a Lucenia query (query DSL) using an LLM.
RAGToolRetrieves with a k-NN search, then asks a generation model to answer using the retrieved context.
ReadFromScratchPadToolReads back the notes an agent saved to its per-conversation scratchpad.
SearchIndexToolSearches an index using a query written in query domain-specific language (DSL).
VectorDBToolEmbeds your query text and runs a k-NN search, returning the matching documents as context.
VisualizationToolFinds saved visualizations by matching a search term against their titles.
WriteToScratchPadToolSaves a short note to an agent's per-conversation scratchpad for later recall.
note

Two names in the shipped ML Commons module do not match the type you write in a request:

  • The class VisualizationsTool registers the type VisualizationTool (no s). Use the type.
  • McpSseTool is registered as a type but is constructed by the cluster when an MCP connector discovers a remote tool. It is bound to an MCP client at that point, so it is not listed by the MCP built-in tool registry when unbound.

Ingest processors

Ingest processors transform documents on the write path, inside an ingest pipeline. The complete table, with a link to each processor's own page, is in Ingest processors.

The processors below have dedicated pages in this version:

ProcessorWhat it does
chunkSplits text content into overlapping chunks using recursive, fixed, semantic, or topic-shift algorithms. Designed for RAG and vector search pipelines.
content_extractExtracts structured content blocks from documents in various formats (PDF, DOCX, HTML, images) via inline text, S3/HTTPS references, or base64 attachments.
embedGenerates vector embeddings from text, image, or multimodal content using Bedrock, OpenAI, or HTTP providers.
ellipseConverts an ellipse geometry (WKT or GeoJSON) into a polygon approximation and indexes it as a geo_shape or shape. Useful for directional or asymmetric coverage areas such as cellular sectors.
image_tilingSplits large images into fixed-size tiles for multimodal vector search. Supports GeoTIFF/COG with HTTP Range reads and geographic bounding box computation.
ocrPerforms optical character recognition on image blocks using LLM vision models (Claude on Bedrock). Handles charts, diagrams, tables, and handwriting.
rerank_prepareAnnotates chunks with document-level metadata and position scores for search-time reranking.

Search processors

Search processors transform requests, responses, and intermediate phase results on the read path, inside a search pipeline. The complete tables are in Search processors.

They run in three places: