Encryption at rest
There are two different questions behind that phrase and this page keeps them apart, because the answers live in different places and only one of them is this service’s job.
- What does this service encrypt before it hands a value to a store?
A fixed set of credentials and private keys, under one key it never
generates. That is below, and the live answer is at
/admin/encryption. - What encrypts everything else — the rest of the database, the WAL, the backups? Nothing here does. That is the storage layer, it is the operator’s, and the options are further down.
What this service seals itself
In product mode — and only where the key-encryption key outlives the process — this service encrypts a specific list of values with AES-256-GCM before they reach the persistence store: its own signing keys, the certificate authority’s key pairs, the private key of every assertion key pair it issues to an application or to a person, authenticator-app shared secrets, the Kerberos long-term keys it stores for directory people and service principals (password-equivalent, and withheld from every page and search even as ciphertext), and every row it mints (sessions, tokens, codes, artifacts, the audit log).
/admin/encryption is the list and this page is deliberately not a second
copy of it. That page names each kind, where it lives, and whether it is
sealed; it is generated from the same labels the code passes at the call site,
so it cannot drift. Read it there.
Two properties of the mechanism matter for what follows:
- The key is never generated by this service and never stored by it.
keys.kekProviderselects where it is read from: a mounted file (the default), AWS Secrets Manager, GCP Secret Manager, Azure Key Vault or HashiCorp Vault — or a key in a key management service, which wraps the data keys itself and never gives the service its bytes (below). - In development mode almost none of this applies. The key-encryption key is ephemeral there, so directory attributes are written in the clear deliberately — sealing an attribute under a key that will not survive the restart the attribute does would turn a certificate into permanent garbage.
How values are encrypted: a key-encryption key and data encryption keys
This service uses envelope encryption (#391). Two kinds of key work together, and they never swap roles:
| Key-encryption key (KEK) | Data encryption key (DEK) | |
|---|---|---|
| What it encrypts | Data encryption keys, and nothing else | Stored values: private keys, secrets, sessions, tokens and the rest of the list above |
| How many | One for the whole deployment (plus one per cell’s own data, in a multi-cell deployment) | One per trust realm per kind of data, and more over time as they are rotated |
| Who makes it | You, or your key management service. This service never generates or writes a KEK | This service: 256 random bits (512 for AES-256-SIV) each |
| Where it is kept | Outside the database: a mounted file, a secret store, or inside a key management service it never leaves | In the database, wrapped (encrypted) by the KEK, in the dek:<scope>:<realm> rows of the key table |
| Where it is used | Once per data key: to wrap a new one, or unwrap a stored one | Once per value, in memory, every time a value is sealed or opened |
| How it is rotated | Supply a successor and restart, or rotate it in its key management service (Rotating the key-encryption key, below). The data keys are re-wrapped; no value is re-encrypted | Yearly by schedule, or by hand. The values are re-encrypted in the background (Rotating the data encryption keys, below) |
The point of the split is that the expensive, sensitive key is used rarely. A busy service seals thousands of values a minute (every session, code and token it stores). If each of those were a call to a key management service, it would be slow, costly and fragile. Instead, each process asks for each data key once, holds it in memory, and seals values locally. And because the KEK only wraps data keys, rotating it means re-wrapping a handful of keys rather than re-encrypting the database.
Sealing and opening a value
To seal a value (for example, an application’s client secret in the
acme realm):
- The service picks the data key for that realm and kind of data
(
acme/client-secret). If none exists yet, it generates one, has the KEK wrap it, and stores the wrapped copy before anything is sealed under it. - It encrypts the value with that data key (AES-256-GCM, or AES-256-SIV for
directory data when
keys.directoryCiphersays so) under a fresh random nonce. -
It stores the result, which names the data key that sealed it:
$aesgcm$2$<data key id>$<iv>$<tag>$<ciphertext> $aessiv$2$<data key id>$<nonce>$<siv>$<ciphertext>
To open a value, the service reads the data key id from the value, finds that data key in memory, and decrypts. Opening never involves the KEK, and never needs to know the realm or kind of data: the value says which key it was sealed under.
A wrapped data key is stored in one of two forms:
$dekwrap$1$... wrapped in this process by a KEK it read
$dekkms$1$<provider>$<key reference>$<ciphertext> wrapped by a key management service
What each key is bound to
Both layers are authenticated encryption, and each binds what it protects to where it belongs:
- A value is bound to its data key. The envelope’s version and data key id are authenticated with the value. A value copied into a row that claims a different key does not decrypt.
- A data key is bound to its id, scope, realm and kind of data. The KEK
wraps it with those four as associated data (Transit’s
associated_data, AWS’s encryption context, Cloud KMS’s additional authenticated data). A wrapped data key copied into another realm’s row does not unwrap, and the key management service itself refuses it. For an Azure RSA key, which takes no associated data, this service checks the binding instead. - Tampering fails closed. A changed byte in a value or a wrapped key is a failed decryption, never a different plaintext. A signing key that decrypted to the wrong bytes would produce signatures nothing can verify, far from the cause, so this matters more than it looks.
One data key per realm per kind of data
No data key is ever shared between realms. The acme realm’s sessions, its
signing keys and its people’s authenticator secrets are three keys of
acme’s own, and the globex realm’s are three others. The kinds of data
are the labels /admin/encryption lists (client-secret, totp-secret,
minted-rows, signing-keys and so on).
Data keys also have a scope. service keys are readable by every node
of the deployment. In a multi-cell deployment, a cell’s resident data
(people homed in that cell) is sealed under cell.<id> keys, wrapped by the
cell’s own KEK (keys.cellKek*). Another cell holds those rows but cannot
open them.
Where the KEK comes from
keys.kekProvider decides, and the difference between the two groups is
whether the KEK’s bytes ever enter this process:
| Providers | The KEK is | What happens at start |
|---|---|---|
file, aws, gcp, azure, vault |
Read into the process: 32 bytes from a mounted file or a secret store (AWS Secrets Manager, GCP Secret Manager, an Azure Key Vault secret, Vault KV) | The KEK is read once, the stored data keys are unwrapped locally, and the KEK stays in memory to wrap new ones |
vault-transit, aws-kms, gcp-kms, azure-keys |
A key in a key management service that never leaves it | The service holds only a handle. Each stored data key is sent to the KMS once to be unwrapped, and each new one to be wrapped |
The key management service options are described in detail below.
At start, and across processes and nodes
- At start, before it listens, the service reads the KEK (or reaches its
key management service) and unwraps every stored data key it may need. A
data key that cannot be unwrapped stops the start
(
STS-KEYS-0091). Usually it means the configured KEK is not the one the data keys were wrapped under, and starting anyway would mean every value under them was lost. - Every process and every node shares the data keys through the database, not through a channel of its own. The key rows are merged under a lock, so two nodes making a data key for the same slot at once both keep theirs, and all agree on one as current.
- A new data key is published before it is used. A rotated key is
written first and used
keys.dataKeyActivationLeadSeconds(300) later, so every node has read it before any value names it. - A value under a data key this process does not hold yet (one another
node made a moment ago) is refused (
STS-KEYS-0092), and the key rows are read again in the background.
Development mode
When nothing outlives the process (development mode, or any configuration
whose KEK is not durable), data keys are derived from an ephemeral
per-run KEK and never stored. The same envelopes are produced, so the code
path is the same, but there is nothing to rotate, count or re-encrypt, and
/admin/encryption says so.
What this protects against, and what it does not
- A copy of the database alone is useless. Every sealed value needs a data key, and every data key needs the KEK, which is not in the database. This covers a stolen backup, a replica, or a database administrator reading rows.
- The KEK opens everything. Anyone who can read it, or use it in its key management service, can unwrap every realm’s data keys. A realm has keys of its own but is not an independent boundary at rest. A KEK per realm is not built; see A key-encryption key per realm, at the end of this page.
- A running process holds its data keys in memory. Someone who can read the service’s memory can read them. A key management service keeps the KEK out of the process, but not the data keys it unwraps.
- Not everything is sealed.
/admin/encryptionlists what is and what is not. On PostgreSQL every directory entry’s attributes are sealed as one blob. The entry’s DN and its attribute names stay readable, and so do timestamps, counters and the risk history. The values the directory looks up (login names, mail addresses, entry UUIDs, object classes) are stored as keyed digests: equal values can be matched, but not read. For everything else, use storage-level encryption (The rest of the store, below).
A store written before #391 (each value under a subkey of the KEK, with no
data keys) is not read by this build; recreate it. The service refuses to
start on one (STS-KEYS-0095) rather than generate new keys over it.
The key-encryption key in a key management service
Set keys.kekProvider to a KMS key and the key-encryption key never leaves
the KMS: the service sends each data key to it to be wrapped, and asks it to
unwrap each data key once when it starts. Values are never sent.
keys.kekProvider |
keys.kekRef |
The key must be |
|---|---|---|
vault-transit |
the Transit key’s name; the engine’s mount is keys.kekTransitMount (transit) |
of type aes256-gcm96, aes128-gcm96 or chacha20-poly1305 |
aws-kms |
the key’s ID, ARN or alias, in keys.kekRegion |
a symmetric encryption key (SYMMETRIC_DEFAULT, ENCRYPT_DECRYPT), enabled |
gcp-kms |
the key’s resource name, projects/…/locations/…/keyRings/…/cryptoKeys/… (the key, not a version) |
ENCRYPT_DECRYPT, GOOGLE_SYMMETRIC_ENCRYPTION, with an enabled primary version |
azure-keys |
the key’s name; the vault (or Managed HSM) URL is keys.kekVault |
an oct-HSM key (Managed HSM) allowing encrypt and decrypt, or an RSA / RSA-HSM key of at least 3072 bits allowing wrapKey and unwrapKey |
Vault Transit is reached with the same login as the vault provider. The
policy needs read on <mount>/keys/<name> and update on
<mount>/encrypt/<name> and <mount>/decrypt/<name>. The other three use
their cloud’s own identity (an instance or task role, a service account, a
managed identity):
- AWS KMS:
kms:DescribeKey,kms:Encryptandkms:Decrypton the key (andkms:GetKeyRotationStatusfor/admin/secrets). Every node must name the key the same way: for a multi-region key, use itsmrk-…key ID and setkeys.kekRegionper node. - Cloud KMS:
roles/cloudkms.cryptoKeyEncrypterDecrypterandroles/cloudkms.vieweron the key. The service reads the key at start, and the encrypter role alone cannot. - Azure: the Key Vault Crypto User role on the key.
Azure RSA keys. RSA-OAEP takes no associated data, so for an RSA key the service wraps each data key with a SHA-256 digest of its binding to its realm and kind of data. The service, not the vault, refuses a key moved to another realm’s row. RSA is also not post-quantum. Where a Managed HSM is available, an
oct-HSM(AES-256) key avoids both.
Each wrapped data key is bound to its realm and kind of data (Transit’s associated data, AWS’s encryption context). The KMS refuses to unwrap one that was moved to another realm’s row.
- Rotating the key in the KMS. Rotating a Transit, Cloud KMS or Azure key creates a new version, and every data key wrapped under an older version is re-wrapped at the next start. Keep the old version enabled until every node has restarted. AWS rotates the key material behind one key ID and needs nothing from the service.
- Moving to a KMS from a file, or back. Use the steps in Rotating the
key-encryption key, below: put the old key in
keys.previousKek*and the new one inkeys.kek*.
Warning. With a KMS key-encryption key the service cannot start while the KMS is unreachable: it has no other way to unwrap its data keys.
The cipher for data stored in the directory
keys.directoryCipher sets the cipher for the values stored on directory
entries: private keys, client secrets, authenticator-app secrets, Kerberos
keys, recovery codes and the other kinds /admin/encryption lists as
directory data.
aes-256-gcm(the default).aes-256-siv: AES-SIV (RFC 5297) with a 512-bit key. That key is two AES-256 keys, one to authenticate and one to encrypt. Unlike GCM, it stays safe when a nonce is repeated.
There is no AES-512. AES has 128-, 192- and 256-bit keys only. AES-256-SIV’s 512-bit key gives 256-bit AES strength, the same as AES-256-GCM. Choose it for its tolerance of a repeated nonce, not for a longer key.
The setting applies to data keys made after it changes. Changing it makes
those kinds of data due a rotation at the next keys.data-key-rotate run (or
Rotate now). The re-encryption pass then re-seals the existing values
under the new cipher. Every other kind of data stays AES-256-GCM.
Rotating the data encryption keys
Every data encryption key is replaced after keys.dataKeyRotationDays (365;
0 turns the schedule off) by the scheduler job keys.data-key-rotate. You can
also rotate them now — every key, one realm’s, or one kind of data’s — from
Monitoring → Encryption (/admin/encryption, Admin Write) or with
POST /admin-api/encryption/rotate-data-keys.
- The new key is published first and used
keys.dataKeyActivationLeadSecondslater (300), so every node has it before anything is sealed under it. - The key it replaces is superseded: it still opens what it sealed, and nothing new is sealed under it.
- Every hour
keys.data-key-reencryptre-seals what is still under a superseded key. Run it now with the console’s Re-encrypt now orPOST /admin-api/encryption/reencrypt-data-keys. - A superseded key is destroyed once nothing in the store is sealed
under it and it has been superseded for
keys.dataKeyRetireAfterDays(7). Destruction cannot be undone.
Warning. Only the PostgreSQL store can count what is sealed under a key. On the
ldifstore superseded keys are never destroyed: they stay, wrapped, in the key table. That is safe, and it means a rotation there does not remove an old key from the store.
The page lists every data key with its realm, kind of data, cipher, state (current, waiting to be used, superseded, destroyed), age, and how many values are sealed under it. It never shows a key.
Value counts come from the scheduler job keys.data-key-count, which
runs daily. Run it now with Count now or
POST /admin-api/encryption/count-data-keys. Counting reads the whole store,
so the page shows the last count and when it was taken; a key not yet counted
shows —. Only the PostgreSQL store can count.
Rotating the key-encryption key
Rotating the key-encryption key re-wraps the data keys; it does not re-encrypt the store.
- Put the new key where the service reads it (
keys.kek*), and pointkeys.previousKekProvider,keys.previousKekRef(and…Field,…Region,…Vaultas the provider needs) at the old key. - Restart. Every data key still wrapped under the old key is re-wrapped
under the new one and written back (logged as
STS-KEYS-0096). - Once every node has started with the new key, set
keys.previousKekProviderback tononeand restart again. The old key is no longer needed.
A node started with only the new key, before step 2 has run anywhere, refuses
to start (STS-KEYS-0091): it cannot unwrap the data keys.
A key in a key management service can also be rotated from Monitoring →
Encryption (Rotate the key-encryption key, Admin Write) or with
POST /admin-api/encryption/rotate-kek. The key service makes a new version,
and every data key is re-wrapped under it at once, with no restart. AWS KMS
rotates on demand behind the same key ID, so nothing is re-wrapped. The
identity the service runs as needs permission to rotate the key (for example
kms:RotateKeyOnDemand, or Key Vault Crypto Officer). The AWS, GCP and Azure
deployments in this repository grant only use of the key, and rotate it on
the key service’s own schedule. A key read into the process cannot be
rotated this way: its successor has to be supplied, as above.
The keyed digests behind the cell routing index and the cell locator tags are made under a stored digest key, wrapped like the data keys, so a rotated key-encryption key changes none of them. That key is never rotated.
Not yet covered. A cell’s own key-encryption key (
keys.cellKek*, #98) has no “previous key” setting yet, so it cannot be rotated this way.
The stack ships with a secret store
docker compose up brings up OpenBao beside the service
and the database, and the service reads BOTH secrets out of it. Nothing is
configured for that: it is what the stack does.
What the bootstrap builds, in three one-shot steps you can watch in the log:
- a TLS listener certificate, minted before the store starts by this repository’s own X.509 encoder;
- the store, with
seal "static"— a real auto-unseal, so it comes up unsealed with no operator and no key ceremony, and keeps its storage across restarts (dev mode would lose the key-encryption key on every start, and a service whose KEK changes cannot read back anything it sealed); - the secrets, a certificate authority inside the store, a client certificate issued from it for this service, and a policy binding that certificate to read on two paths and write nothing.
The service then authenticates with that certificate rather than a token — the store issued the identity, can revoke it, and no bearer credential sits in a file waiting to be copied.
The policy is verified rather than asserted. The seeder finishes by logging in as the service, reading what it reads, and attempting the write it must never be able to make; if that write is accepted the stack does not start. A widened capability list therefore fails at boot rather than in an audit.
sts-bao-seed: proved it with the certificate itself: policies [default, sts-read],
the two secrets readable, a write to them refused 403, and no other
path reachable.
What it does not protect you from, said plainly: the unseal key travels in
the stack’s environment, so the store protects its contents from a reader of the
database and not from somebody who can read the stack. Set STS_BAO_SEAL_KEY
from your orchestrator’s own secret, or replace the seal stanza in
openbao/bao.hcl with transit, awskms, gcpckms or azurekeyvault — one
stanza, same behaviour, key somewhere the stack cannot read.
Seeing the state of it: /admin/secrets
Under Monitoring in the admin console. It answers the question none of the settings pages can: did this service actually get its secrets, and is the thing holding them healthy?
- For a mounted file — the path, whether it is there, its mode, owner and mtime, its symlink target if it has one (a Kubernetes Secret mount is a symlink that is replaced on rotation, so the target is how you see that a rotation landed), and which members it holds if it is JSON. The member names, never the values.
- For OpenBao or HashiCorp Vault — initialised or not, sealed or unsealed
and the seal type, the version and build, the cluster and which node leads
it, the store’s clock against this service’s, the client certificate this
service presents and how many days before it expires, the policies the
token it obtained carries, and every version of the secret the store has
kept. And the row worth opening first: what this identity may actually do,
asked of the store rather than read off
openbao/read-only.hcl— the seeder’s final proof, repeated by the running service against the running store. - For AWS, GCP and Azure — whatever each publishes about the secret: rotation and its schedule, the KMS key, version states, staging labels, replication. Through a describe rather than a get.
Beside each secret it says whether this process has read it, when, and —
if the last attempt failed — why. That row matters more than it looks: in
development mode nothing is persisted, so the key-encryption key is never asked
for, and a completely broken provider configuration looks exactly like a working
one until the day you set global.mode=product.
Three things it will not do, and all three are deliberate: it will not show
you a secret, it has no rotate button (the key-encryption key is rotated by
pointing keys.kek* at the new key and keys.previousKek* at the old one —
Rotating the key-encryption key, above — and the store is the operator’s to
change, not this service’s), and it has no test-read, because that would be a
console page causing the key to be in memory. Every probe behind it reads metadata only.
A probe refused with 403 is usually good news and the page says so: the identity is bound to two read paths, so a refusal anywhere else is the policy working.
GET /admin-api/secrets is the same report as JSON.
The database password comes from the same place, if you want it to
persistence.databaseUrl carries its password in plain text
— right for a throwaway database of mock identities, wrong for anything you
would call a deployment — and persistence.databasePasswordProvider reads it
from a secret store instead, using the same five providers and the same code
as the key-encryption key above.
# One mounted file holding both secrets:
# { "kek": "…32 bytes base64…", "databasePassword": "…" }
STS_KEYS_KEK_PROVIDER: file
STS_KEYS_KEK_FILE: /run/secrets/sts
STS_KEYS_KEK_FIELD: kek
STS_DATABASE_PASSWORD_PROVIDER: file
# no location: empty means "wherever the key-encryption key is"
STS_DATABASE_URL: postgres://sts_app@postgres:5432/sts?sslmode=require
The location defaults to the key’s, which is the whole point: a deployment
already mounts one file, or already keeps one secret in AWS, and being made to
provision a second one is work this service would have invented. What tells the
two apart inside it is persistence.databasePasswordField, databasePassword
by default.
Except when the key-encryption key is in a key management service. Its
location is the name of a key, not a place a secret is stored, so nothing is
borrowed from it: set persistence.databasePasswordRef (and its provider’s
other settings) explicitly. This repository’s compose stack does:
STS_DATABASE_PASSWORD_REF=secret/data/sts, beside the Transit key
sts-kek.
A secret of its own needs no field — point persistence.databasePasswordRef
at a Vault path or a file holding nothing but the password and it is taken
whole. For the {"username": …, "password": …} shape AWS Secrets Manager writes
for a database credential, set the field to password.
What it refuses, and why each one
- A shared location holding something that is not a JSON object. What is in a plain key file is the key, and returning it as a password would send this service’s master key to a database server as a credential, in the clear, on the wire — with a failed connection as the only symptom. This is the refusal the feature is built around.
- An empty secret. An empty password is not a password; dialling with one fails at the database as authentication failed, which sends you to look at the wrong end.
- A connection string in libpq’s
host=… user=…keyword/value form.pgaccepts it and this cannot edit one safely, so it says so rather than dialling without the password you configured. Write it as a URL and leave the password out. - A read that fails at startup. The service does not start, exactly as it does not start for a store it cannot open: a process that was told where the password lives and carried on with the one in the URL would be ignoring the configuration that exists to keep it out of the URL.
Two details you would otherwise find the hard way
A password already in the URL is replaced, and the log says so without saying what with. Two passwords for one connection is a question with no good answer, and the configured provider is the one somebody chose deliberately.
The password is injected into the string rather than passed beside it, and
that is pg’s doing rather than a preference: its ConnectionParameters lets
everything parsed out of a connectionString override an explicit field, and
its parser returns password: '' even for a string carrying none — so a
password option beside a connectionString is silently thrown away. The value
is percent-encoded on the way in, because pg decodes what it finds; a password
containing % is mangled or throws without it. tests/database_password.js
asserts the round-trip through pg’s own parser for every character that has
ever caused this.
/admin/persistence and GET /admin-api/persistence report where the
password came from and never what it is — a Password row naming the provider,
the location and whether it is shared with the key.
The rest of the store: four layers, and what each one is worth
Everything this service does not seal — directory entries, group memberships, applications, realms, settings, and in development mode the lot — is plaintext in whatever store you configured. (The section above is about the credential this service USES to reach that store; this one is about what is inside it.) PostgreSQL has no encryption of its own in the community build, so the options are underneath it or beside it.
| Layer | Protects against | Useless against |
|---|---|---|
| Block / volume (LUKS, cloud disk) | a stolen disk, decommissioned hardware, a cloned snapshot | anybody with access to the running host |
| Filesystem (ZFS, fscrypt) | the same, plus per-dataset key separation | the same |
| In-database TDE (forks, extensions) | reading the data files without the server | a database superuser, SQL injection |
| Application envelope (what this service does) | a compromised database, a leaked dump | a compromised application process |
They compose. The bottom two are cheap and cover everything uniformly; the top one is expensive per field and is why this service only applies it to credentials and keys.
Block level — the usual answer
- LUKS / dm-crypt on the partition holding
PGDATA. Transparent to PostgreSQL and effectively free with AES-NI. The real decision is how it unlocks unattended: a keyfile on separate media, TPM 2.0 (systemd-cryptenroll --tpm2-device=auto), or Clevis + Tang so that the key is network-bound and a disk that leaves the building does not unlock. - Cloud disk encryption with a customer-managed key — EBS, GCP Persistent Disk, Azure Disk. This is also what RDS / Aurora / Cloud SQL encryption at rest actually is. Gotcha worth knowing before you need it: on RDS you cannot turn it on in place — you snapshot, copy the snapshot encrypted, and restore.
- Self-encrypting drives or array-level encryption protect against the drive walking out and nothing else; the key usually lives with the array.
Filesystem level
- ZFS native encryption (
encryption=aes-256-gcm, keys per dataset,zfs load-key) is the best fit of the filesystem options: you can encrypt just the PostgreSQL dataset, andzfs sendof an encrypted dataset stays encrypted, which covers the backup path as well. Pair it with the ordinary PostgreSQL-on-ZFS tuning (recordsize=8K,compression=lz4,atime=off,logbias=throughput). - ext4 / F2FS
fscrypt— native, per-directory keys, no FUSE overhead. The catch is timing: the policy is set on an empty directory, so it is aninitdb-time decision rather than something to retrofit. - gocryptfs, eCryptfs and friends — fine over a documents directory, poor under a database’s random I/O and fsync behaviour. Not recommended here.
Inside PostgreSQL
- There is no TDE in community PostgreSQL. Cluster-file encryption has been proposed repeatedly and has not landed; check the release notes of whichever version you are on rather than trusting this sentence indefinitely.
- Percona
pg_tde(Percona Server for PostgreSQL 17 onward) is the closest open equivalent, with per-table keys and an external key provider. EDB Postgres Advanced Server, Fujitsu Enterprise Postgres and CYBERTEC’s patched fork are the commercial ones. pgcryptodoes column-level encryption in SQL. The cryptography is not the problem; the key is. It arrives inside a statement, which means it can reachpg_stat_statements,log_statementand error logs. If the key never comes from the client, you have reinvented the application layer — which this service already implements, with a proper key provider behind it.
What column-level encryption misses, and block-level does not
Encrypting individual columns leaves the plaintext in the write-ahead log,
in temporary files when a sort spills to disk, in pg_dump output, on
replicas, and in query logs. That asymmetry is the main argument for
putting the layer underneath the database rather than inside it.
This repository’s own stack
docker-compose.yml runs postgres:18 with a named volume, so the encryption
is the host’s job: put the docker volume — or /var/lib/docker — on a LUKS
partition or an encrypted ZFS dataset. Two specifics are easy to get wrong:
- The key-encryption key is a Transit key in the stack’s OpenBao
(
sts-kek,keys.kekProvider=vault-transit). It never enters the service, which asks OpenBao to wrap and unwrap each data key. OpenBao’s data is in a docker volume on the same host, though, so whoever has the host’s disk has both the store and the key. That is right for a development stack and wrong for anything else. In a real deployment, put the key in a key management service or secret store that does not share a disk with the database. - TLS to the database is already on and is a different property.
The stack’s PostgreSQL refuses a plaintext connection (
hostsslon every rule) and this service asks forsslmode=require, so both ends insist. It does not authenticate that server — the certificate is generated in the container and signed by nobody — and/admin/persistencereports those as two facts rather than one tick. Encryption in transit is not encryption at rest, and neither implies the other.
A recommendation
For a deployment that matters: LUKS or ZFS under the data directory with the
key held by a TPM, a Tang server or a cloud KMS; the application-level sealing
left exactly as it is; and pg_tde only if you specifically need per-tenant
keys visible inside the database. The first covers the stolen-media and
leaked-snapshot cases for everything in the store; the second is what protects
the credentials if the database itself is read by somebody who should not; the
third is a large operational commitment for a narrower gain.
A key-encryption key per realm
If one tenant’s data must be unreadable with another tenant’s key, the data keys are already per realm; what is shared is the key-encryption key that wraps them. A separate key-encryption key per realm is not implemented: the secret provider would need a keyed read, the keystore could not read one key at startup before the realm registry exists, and creating a realm would become a key-provisioning act where today it is one API call.