Compliance processor
The compliance processor finds sensitive values in a document and redacts them before indexing,
so the raw value is never written to a segment. It also records what it found, which is what makes it
auditable rather than merely destructive.
{
"compliance": {
"fields": ["content", "notes"],
"profile": "hipaa"
}
}
Options
| Option | Required | Default | Meaning |
|---|---|---|---|
fields | yes | — | Fields to scan. At least one; an empty list is a configuration error. |
profile | no | none | A named regulatory profile (below). |
policy | no | — | Use a cluster-stored named policy instead of an inline one. |
metadata_field | no | _compliance | Where the findings report is written. |
ignore_missing | no | true | Skip a listed field that is absent rather than failing the document. |
overrides | no | — | Per-entity redaction mode, overriding the profile. |
exclude | no | — | Entities the profile would catch but you want left alone. |
detectors | no | — | Custom detectors defined at runtime. |
import | no | — | Import detectors from an external definition. |
hash_salt | no | "" | Salt for hash mode, so the same value hashes the same way within a deployment. |
audit | no | false | Detect and report but change nothing. See Audit first. |
policy cannot be combined with any inline option — profile, hash_salt, exclude, overrides,
detectors, import or audit. A named policy is the whole policy; half of it coming from the
pipeline would make the stored policy a lie about what is enforced.
Profiles
A profile is a regulation-shaped starting point: which categories of data it covers, and the default redaction mode for them.
| Profile | Default mode | Covers |
|---|---|---|
none | mask | nothing by default — build it yourself with overrides |
pii_basic | mask | PII |
gdpr | mask | PII |
ccpa | mask | PII |
hipaa | mask | PHI, PII |
pci_dss | partial | PCI, financial |
soc2 | mask | PII, secrets |
fedramp | mask | PII, PHI, secrets |
iso_27001 | mask | PII, secrets |
nist_800_53 | mask | PII, PHI, PCI, financial, secrets |
pci_dss defaults to partial rather than mask for a practical reason: PCI requires that a card
number not be stored in full, but the last four digits remain legitimately useful for customer
service, so masking everything would break the workflow the rule was written to allow.
Redaction modes
| Mode | Effect on 4111 1111 1111 1111 |
|---|---|
mask | **************** |
partial | ************1111 |
hash | a stable salted digest |
tokenize | a reversible token, if you hold the mapping |
remove | the value is deleted |
tag | unchanged; the finding is only reported |
Set them per entity:
{
"compliance": {
"fields": ["content"],
"profile": "pci_dss",
"overrides": { "email_address": "hash", "credit_card": "partial" },
"exclude": ["date_of_birth"],
"hash_salt": "${COMPLIANCE_SALT}"
}
}
Audit first
audit: true forces every entity to tag mode: detection runs, findings are recorded, and the
document is otherwise untouched.
{ "compliance": { "fields": ["content"], "profile": "gdpr", "audit": true } }
Run this against real traffic before you enforce. It tells you what a policy would redact, which is the only safe way to discover that a detector is over-matching on your data — a discovery that is expensive to make after the original values have been masked on the way in.
The findings report
Each document gets a _compliance field (or whatever metadata_field names) describing what was
found and what was done. This is what makes the pipeline auditable: you can answer "was this document
redacted, and against which policy" without re-scanning it.
Named policies
A policy can live in cluster state and be referenced by name, so many pipelines share one definition and changing it does not mean editing every pipeline.
PUT /_plugins/_compliance/policies/{name}
GET /_plugins/_compliance/policies/{name}
GET /_plugins/_compliance/policies
DELETE /_plugins/_compliance/policies/{name}
{ "compliance": { "fields": ["content"], "policy": "corporate-baseline" } }
Custom detectors
Detectors can be defined at runtime — no plugin rebuild, no node restart — for the identifiers that are specific to your business and that no shipped profile could know about.
{
"compliance": {
"fields": ["content"],
"profile": "pii_basic",
"detectors": [
{ "name": "employee_id", "pattern": "EMP-[0-9]{6}", "category": "pii", "mode": "hash" }
]
}
}
A pattern is validated at definition time, not first use, so a bad rule is rejected when you configure the pipeline rather than when a document arrives.
The ReDoS guard
A user-supplied regular expression is an availability risk: a catastrophically backtracking pattern against a large field can hang the ingest thread. Every detector runs against a bounded character budget, and a pattern that exhausts it is abandoned.
The guard fails closed — exhausting the budget means the document is rejected, not quietly passed through unredacted. A redaction that silently did not happen is worse than a document that did not index, because only one of the two is visible.
Reference catalogues
The shipped detectors, frameworks and governance templates are all introspectable:
GET /_plugins/_compliance/presets # and /{id}
GET /_plugins/_compliance/frameworks # and /{id}
GET /_plugins/_compliance/governance # and /{id}
Frameworks map to control identifiers — gdpr, hipaa, pci_dss, soc2, fedramp, iso_27001,
nist_800_53, nist_800_171, cmmc_level2, ccpa — and carry the specific controls a redaction
satisfies (SC-28, SI-19, MP-6, AC-4, PT-3, pseudonymisation, de-identification), which is
what an auditor actually asks for.
Migrating from Presidio
POST /_plugins/_compliance/_convert/presidio
Converts a Presidio recogniser definition into detectors, so an existing deployment's rules come across rather than being rewritten by hand.
Redacting at query time
The compliance processor redacts on the way in, so nobody can ever read the raw value. Sometimes
the value has to be stored as written — a clinician needs the note — and hidden only from everyone
else. compliance_redact is a search response processor for that: the value stays in the index,
and the same detectors run over search results before they leave the node.
{
"response_processors": [
{
"compliance_redact": {
"fields": ["notes"],
"profile": "hipaa",
"exempt_roles": ["clinician"],
"exempt_backend_roles": ["phi-cleared"]
}
}
]
}
Create it with PUT /_search/pipeline/{id}, then search with ?search_pipeline={id}, or make it the
index default with index.search.default_pipeline.
| Option | Required | Default | Meaning |
|---|---|---|---|
fields | yes | — | Fields to redact in results. A field covers everything beneath it and its sub-fields. |
exempt_roles | no | — | Security roles whose members see the clear text. |
exempt_backend_roles | no | — | External groups (LDAP, SAML, OIDC) whose members see the clear text. |
exempt_when_no_identity | no | false | Let a caller with no verified identity see the clear text. |
profile, policy, detectors, ... | no | — | The same policy options as compliance. |
Who is exempt
Exemption comes from the identity the OpenSearch security plugin verified. exempt_roles matches
the roles a user is mapped to, and exempt_backend_roles matches the groups their identity provider
reports. A group is never treated as a role, even if the names match.
It fails closed. With no verified identity — the security plugin not installed, or a request made
with a node certificate — results are redacted, unless the pipeline sets exempt_when_no_identity.
Every way a value can leave
Redaction covers each path a value comes back by, not only the obvious one:
- nested objects (
patient.notes), flattened dotted keys, and arrays - a whole object under a configured field
- highlight fragments, scanned with highlight tags removed, since
123-<em>45</em>-6789would otherwise hide the value from the detector - stored and doc-value
fields, including sub-fields such asnotes.keyword, which hold the same raw text - inner hits
A failure fails the search
ignore_failure is rejected when the pipeline is created. A search pipeline that ignores a failed
processor returns the results as they were, which here means unredacted.
Auditing who read it
compliance_redact does not write its own audit records. A record per search, written from inside
the search path, would add a write to every query — and would need a system index, which a plugin on a
managed service cannot register. The record an auditor asks for, who accessed which sensitive field,
and when, is already kept by the OpenSearch security plugin's compliance audit log.
Turn it on for the fields you redact, in audit.yml:
compliance:
enabled: true
read_metadata_only: true
read_watched_fields:
"clinical-notes": ["notes"]
or at runtime:
PATCH _plugins/_security/api/audit
[
{ "op": "add", "path": "/config/compliance/read_watched_fields", "value": { "clinical-notes": ["notes"] } },
{ "op": "add", "path": "/config/compliance/read_metadata_only", "value": true }
]
Keep read_metadata_only: true. The audit log records the read when the shard fetches the document,
before the search response is redacted, so with it false the log copies the very values the pipeline
exists to hide — a second place the data lives, with its own retention and readers.
Read the two together: the audit log says who read a watched field; the pipeline's exempt_roles and
exempt_backend_roles say which of those readers saw it in the clear. To audit changes as well as
reads, add the index to write_watched_indices.
What it does not do
It redacts search results. A GET of the document, and aggregations over a keyword field, still
return the stored value. Restrict those with the security plugin's document- and field-level security,
or redact at ingest instead.
Licensing
Creating a pipeline that uses compliance requires a license. Running one does not — see
Licensing for why execution is never gated.
compliance_redact needs no license at all. Refusing to create a redaction pipeline when a license
lapses would push an operator to serve results without it.