Skip to main content

Compliance processor

The compliance processor finds sensitive values in a document and redacts them before indexing, so the raw value is never written to a segment. It also records what it found, which is what makes it auditable rather than merely destructive.

{
"compliance": {
"fields": ["content", "notes"],
"profile": "hipaa"
}
}

Options​

OptionRequiredDefaultMeaning
fieldsyes—Fields to scan. At least one; an empty list is a configuration error.
profilenononeA named regulatory profile (below).
policyno—Use a cluster-stored named policy instead of an inline one.
metadata_fieldno_complianceWhere the findings report is written.
ignore_missingnotrueSkip a listed field that is absent rather than failing the document.
overridesno—Per-entity redaction mode, overriding the profile.
excludeno—Entities the profile would catch but you want left alone.
detectorsno—Custom detectors defined at runtime.
importno—Import detectors from an external definition.
hash_saltno""Salt for hash mode, so the same value hashes the same way within a deployment.
auditnofalseDetect and report but change nothing. See Audit first.

policy cannot be combined with any inline option — profile, hash_salt, exclude, overrides, detectors, import or audit. A named policy is the whole policy; half of it coming from the pipeline would make the stored policy a lie about what is enforced.

Profiles​

A profile is a regulation-shaped starting point: which categories of data it covers, and the default redaction mode for them.

ProfileDefault modeCovers
nonemasknothing by default — build it yourself with overrides
pii_basicmaskPII
gdprmaskPII
ccpamaskPII
hipaamaskPHI, PII
pci_dsspartialPCI, financial
soc2maskPII, secrets
fedrampmaskPII, PHI, secrets
iso_27001maskPII, secrets
nist_800_53maskPII, PHI, PCI, financial, secrets

pci_dss defaults to partial rather than mask for a practical reason: PCI requires that a card number not be stored in full, but the last four digits remain legitimately useful for customer service, so masking everything would break the workflow the rule was written to allow.

Redaction modes​

ModeEffect on 4111 1111 1111 1111
mask****************
partial************1111
hasha stable salted digest
tokenizea reversible token, if you hold the mapping
removethe value is deleted
tagunchanged; the finding is only reported

Set them per entity:

{
"compliance": {
"fields": ["content"],
"profile": "pci_dss",
"overrides": { "email_address": "hash", "credit_card": "partial" },
"exclude": ["date_of_birth"],
"hash_salt": "${COMPLIANCE_SALT}"
}
}

Audit first​

audit: true forces every entity to tag mode: detection runs, findings are recorded, and the document is otherwise untouched.

{ "compliance": { "fields": ["content"], "profile": "gdpr", "audit": true } }

Run this against real traffic before you enforce. It tells you what a policy would redact, which is the only safe way to discover that a detector is over-matching on your data — a discovery that is expensive to make after the original values have been masked on the way in.

The findings report​

Each document gets a _compliance field (or whatever metadata_field names) describing what was found and what was done. This is what makes the pipeline auditable: you can answer "was this document redacted, and against which policy" without re-scanning it.

Named policies​

A policy can live in cluster state and be referenced by name, so many pipelines share one definition and changing it does not mean editing every pipeline.

PUT  /_plugins/_compliance/policies/{name}
GET /_plugins/_compliance/policies/{name}
GET /_plugins/_compliance/policies
DELETE /_plugins/_compliance/policies/{name}
{ "compliance": { "fields": ["content"], "policy": "corporate-baseline" } }

Custom detectors​

Detectors can be defined at runtime — no plugin rebuild, no node restart — for the identifiers that are specific to your business and that no shipped profile could know about.

{
"compliance": {
"fields": ["content"],
"profile": "pii_basic",
"detectors": [
{ "name": "employee_id", "pattern": "EMP-[0-9]{6}", "category": "pii", "mode": "hash" }
]
}
}

A pattern is validated at definition time, not first use, so a bad rule is rejected when you configure the pipeline rather than when a document arrives.

The ReDoS guard​

A user-supplied regular expression is an availability risk: a catastrophically backtracking pattern against a large field can hang the ingest thread. Every detector runs against a bounded character budget, and a pattern that exhausts it is abandoned.

The guard fails closed — exhausting the budget means the document is rejected, not quietly passed through unredacted. A redaction that silently did not happen is worse than a document that did not index, because only one of the two is visible.

Reference catalogues​

The shipped detectors, frameworks and governance templates are all introspectable:

GET /_plugins/_compliance/presets              # and /{id}
GET /_plugins/_compliance/frameworks # and /{id}
GET /_plugins/_compliance/governance # and /{id}

Frameworks map to control identifiers — gdpr, hipaa, pci_dss, soc2, fedramp, iso_27001, nist_800_53, nist_800_171, cmmc_level2, ccpa — and carry the specific controls a redaction satisfies (SC-28, SI-19, MP-6, AC-4, PT-3, pseudonymisation, de-identification), which is what an auditor actually asks for.

Migrating from Presidio​

POST /_plugins/_compliance/_convert/presidio

Converts a Presidio recogniser definition into detectors, so an existing deployment's rules come across rather than being rewritten by hand.

Redacting at query time​

The compliance processor redacts on the way in, so nobody can ever read the raw value. Sometimes the value has to be stored as written — a clinician needs the note — and hidden only from everyone else. compliance_redact is a search response processor for that: the value stays in the index, and the same detectors run over search results before they leave the node.

{
"response_processors": [
{
"compliance_redact": {
"fields": ["notes"],
"profile": "hipaa",
"exempt_roles": ["clinician"],
"exempt_backend_roles": ["phi-cleared"]
}
}
]
}

Create it with PUT /_search/pipeline/{id}, then search with ?search_pipeline={id}, or make it the index default with index.search.default_pipeline.

OptionRequiredDefaultMeaning
fieldsyes—Fields to redact in results. A field covers everything beneath it and its sub-fields.
exempt_rolesno—Security roles whose members see the clear text.
exempt_backend_rolesno—External groups (LDAP, SAML, OIDC) whose members see the clear text.
exempt_when_no_identitynofalseLet a caller with no verified identity see the clear text.
profile, policy, detectors, ...no—The same policy options as compliance.

Who is exempt​

Exemption comes from the identity the OpenSearch security plugin verified. exempt_roles matches the roles a user is mapped to, and exempt_backend_roles matches the groups their identity provider reports. A group is never treated as a role, even if the names match.

It fails closed. With no verified identity — the security plugin not installed, or a request made with a node certificate — results are redacted, unless the pipeline sets exempt_when_no_identity.

Every way a value can leave​

Redaction covers each path a value comes back by, not only the obvious one:

  • nested objects (patient.notes), flattened dotted keys, and arrays
  • a whole object under a configured field
  • highlight fragments, scanned with highlight tags removed, since 123-<em>45</em>-6789 would otherwise hide the value from the detector
  • stored and doc-value fields, including sub-fields such as notes.keyword, which hold the same raw text
  • inner hits

ignore_failure is rejected when the pipeline is created. A search pipeline that ignores a failed processor returns the results as they were, which here means unredacted.

Auditing who read it​

compliance_redact does not write its own audit records. A record per search, written from inside the search path, would add a write to every query — and would need a system index, which a plugin on a managed service cannot register. The record an auditor asks for, who accessed which sensitive field, and when, is already kept by the OpenSearch security plugin's compliance audit log.

Turn it on for the fields you redact, in audit.yml:

compliance:
enabled: true
read_metadata_only: true
read_watched_fields:
"clinical-notes": ["notes"]

or at runtime:

PATCH _plugins/_security/api/audit
[
{ "op": "add", "path": "/config/compliance/read_watched_fields", "value": { "clinical-notes": ["notes"] } },
{ "op": "add", "path": "/config/compliance/read_metadata_only", "value": true }
]

Keep read_metadata_only: true. The audit log records the read when the shard fetches the document, before the search response is redacted, so with it false the log copies the very values the pipeline exists to hide — a second place the data lives, with its own retention and readers.

Read the two together: the audit log says who read a watched field; the pipeline's exempt_roles and exempt_backend_roles say which of those readers saw it in the clear. To audit changes as well as reads, add the index to write_watched_indices.

What it does not do​

It redacts search results. A GET of the document, and aggregations over a keyword field, still return the stored value. Restrict those with the security plugin's document- and field-level security, or redact at ingest instead.

Licensing​

Creating a pipeline that uses compliance requires a license. Running one does not — see Licensing for why execution is never gated.

compliance_redact needs no license at all. Refusing to create a redaction pipeline when a license lapses would push an operator to serve results without it.