A three-node cluster in AWS
deploy/aws/ builds a test environment with Terraform: three iya-sts nodes
running active-active on ECS Fargate, one in each of three availability zones,
behind an internet-facing Network Load Balancer, against RDS PostgreSQL 18 with
a read replica. The encryption key and the database password come from AWS
Secrets Manager. A GitHub Actions workflow creates an environment, runs the
protocol suite against it, and destroys it.
you (allowed CIDRs) ──443──▶ NLB ──8081, PROXY v2──▶ node-a | node-b | node-c
(us-west-2a | 2b | 2c)
│ TLS, verified
RDS primary (2a) ───▶ read replica (2b)
What you get
| Piece | Details |
|---|---|
| Nodes | ECS Fargate, 1 vCPU and 3 GB each, one ECS service per AZ; cluster mode active-active; development mode by default (sts_mode) |
| Front door | NLB, TCP 443 → 8081 plus 389 (the directory) and 8082 (CRLs and OCSP) on the same numbers, TLS passed through to the nodes, PROXY protocol v2, a security group admitting only allowed_cidrs and the suite runner’s address |
| Suite runner | an ECS task started per run in a subnet of its own behind a NAT gateway (suite_runner, on by default), which runs every job from inside the VPC and uploads the report to S3 |
| Database | RDS PostgreSQL 18, db.t4g.small, 20 GB gp3; not publicly accessible; rds.force_ssl = 1; encrypted with the project KMS key; 14-day automated backups; one read replica in a second AZ |
| Key-encryption key | a multi-region AWS KMS key made once by foundation/ (alias/iya-sts-kek), never leaving KMS: each node wraps and unwraps its data encryption keys with it, naming the key by its ID and its own region (kek_provider = "kms", the default). kek_provider = "secret" keeps it in Secrets Manager instead |
| Secrets | Secrets Manager, generated by Terraform: the application and master database passwords, the /admin-api client secret, and the secret key-encryption key (read only with kek_provider = "secret", or while migrating from it) |
| Risk dataset uploads | a gp3 EBS volume per node, 10 GiB, encrypted with the project KMS key, mounted at risk.uploadDirectory — created when the node’s task starts and deleted when it stops (see Where a risk dataset upload goes, below) |
| Logs | CloudWatch log group /iya-sts/containers, stream prefix <environment>-<node>, kept 14 days — after the environment is destroyed |
| Identity | the containers run as a task role that can read their secrets and use the key-encryption key and nothing else; the environment is created by a dedicated deployer role |
What is reachable: https://<nlb-dns-name> on 443, and the same name on
389 (LDAP) and 8082 (the plain-HTTP CRL and OCSP listener), because the service
writes those addresses into what it publishes. Kerberos and LDAPS are bound
inside the nodes but not published. There is no separate mutual-TLS
port for the certificate sign-in: a certificate is presented to 443 like everything else — the main port asks
every connection for one and requires none (GET /tls/sign-in).
Where a risk dataset upload goes
A file uploaded on Monitoring → Risk, or to POST /admin-api/risk/upload
(Risk scoring), is written to
risk.uploadDirectory on the node the load balancer sent the request to,
imported by that node into the shared database, and deleted. On AWS that
directory is /usr/src/sts/data/risk-uploads, and each node has an Amazon EBS
volume of its own mounted there:
- gp3, 10 GiB by default, five times the largest upload the service
accepts (
risk.uploadMaxBytes, 2 GiB). An upload the volume has no room for is refused before it is stored. Change it with the Terraform variablesrisk_upload_volume_gib,risk_upload_volume_throughputandrisk_upload_volume_iops; if you raiserisk.uploadMaxBytes(throughextra_environment), raise the volume with it. - Encrypted with the project’s KMS key, the one that already encrypts the secrets, the database and the logs.
- Temporary. ECS creates the volume when the node’s task starts and deletes it when the task stops. Nothing on it survives a restart, and nothing needs to: an import interrupted by a restart is marked refused, and you upload the file again.
- Not the task’s ephemeral storage, which stays at Fargate’s default and is shared with the container image.
It costs about $0.80 a month per node at the default size. An administrator
has to re-apply foundation/ once before the first environment with the
volume is applied: the role ECS uses to manage the volumes needs a
permissions boundary that only foundation/ creates.
Cost: about $0.35 for an idle hour and $0.53 for an hour with a suite run
(us-west-2 on-demand, 2026-09; most of it the three nodes, the two database
instances and the suite runner’s NAT gateway — suite_runner = false saves
about $0.05 an hour). The long-lived pieces — the KMS key, the image
repository and the two buckets — cost about $1.10 a month.
Running it
Once, with administrator credentials:
deploy/aws/bootstrap-state.sh
terraform -chdir=deploy/aws/foundation init -backend-config=bucket=iya-sts-terraform-state-<account>
terraform -chdir=deploy/aws/foundation apply
aws iam create-access-key --user-name iya-sts-deployer
That creates the state bucket and the deployer identity. Everything below runs
as the iya-sts-deployer role.
An environment by hand (from the repository root):
repo=<account>.dkr.ecr.us-west-2.amazonaws.com/iya-sts
tag=$(git rev-parse --short=12 HEAD)
docker build -t $repo:$tag \
--build-arg STS_CLOUD_SDKS="@aws-sdk/client-secrets-manager @aws-sdk/client-kms" \
--build-arg STS_DATABASE_CA_URL=https://truststore.pki.rds.amazonaws.com/global/global-bundle.pem .
docker build -t $repo:schema-$tag -f deploy/aws/schema-init/Dockerfile .
docker build -t iya-sts-tests -f tests/Dockerfile .
docker build -t $repo:runner-$tag --build-arg TESTS_IMAGE=iya-sts-tests \
-f deploy/aws/runner/Dockerfile deploy/aws/runner
docker build -t $repo:pep-$tag -f xacml-pep/Dockerfile .
for i in $tag schema-$tag runner-$tag pep-$tag; do docker push $repo:$i; done
cd deploy/aws/environment
terraform init -backend-config=bucket=iya-sts-terraform-state-<account> \
-backend-config=key=environment/dev.tfstate
terraform apply -var environment=dev -var image_tag=$tag \
-var 'allowed_cidrs=["<your-ip>/32"]'
cd ../../..
deploy/aws/run-suite-in-aws.sh dev # every job, from inside the VPC
deploy/aws/run-suite.sh dev # or from your machine, less two jobs
terraform -chdir=deploy/aws/environment destroy -var environment=dev \
-var image_tag=$tag -var 'allowed_cidrs=["<your-ip>/32"]'
deploy/aws/terraform-local.sh dev apply|suite|destroy does the Terraform and
suite steps inside a container, with nothing but Docker installed. An
environment can be reused: each suite run first removes what the previous
runs left — every realm but the default one, and in the default realm the
bulk-load people and groups, the applications the suite registered and the
runtime overrides (deploy/aws/reset-environment.js; --dry-run lists them;
STS_SUITE_KEEP_REALMS=1 keeps everything).
From GitHub Actions: the AWS cluster test workflow
(.github/workflows/aws-cluster.yml), started by hand, with the actions
apply-and-test, apply, test, plan and destroy. It needs two
repository secrets — AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY, the key
of the workflow’s own IAM user, which may only assume the deployer role — and
admits the runner’s own address to the load balancer. The keep input leaves
an apply-and-test environment running; the report is uploaded as an
artifact either way.
Several regions: cells
An environment can also be built in several AWS regions at once, as cells (issue #98). Each cell is a complete copy of the environment above in one region: its own nodes, load balancer, database and logs. The cells answer to one public name. A global stack joins them.
test-idp.iyasec.io (Route 53)
Canada ─────────┘ └───────── everyone else: the nearest
│ healthy cell (latency)
▼ ▼
cell cac1 (ca-central-1) ◀── peering, 8446 ──▶ cell usw2 (us-west-2)
nodes, cell database nodes, cell database
global read replica ◀──── RDS replication ──── global database (writer)
What stays in a cell: its own database and its own key-encryption key. That key is sealed by a KMS key that exists only in the cell’s region and is never copied.
What every cell shares: the global database (one writer, and a read replica in each other cell), the global key-encryption key — by default the multi-region KMS key, with a replica in each cell’s region — and a few secrets that must be the same everywhere. Secrets Manager copies these into each cell’s region.
How clients reach a cell:
- A country can be pinned to a jurisdiction (a legal area, such as the
EU). Its clients always go to one of that jurisdiction’s cells: the nearest
healthy one, or, if they are all down, one of them anyway, never a cell
elsewhere. Canada is pinned to
ca(cellcac1). - Everyone else goes to the nearest cell that passes its health check.
- The cells reach one another on port 8446, over VPC peering, by private names. That port is on no public load balancer.
Before the first apply, an administrator re-applies foundation/ with
every cell’s region in permitted_regions. Its default lists every region
testidpna and globalidp use. A region launched since 2019 is opt-in:
it must be enabled on the account first (Account → AWS Regions).
Building and removing it: a multi-cell environment is described by
deploy/aws/environment/envs/<env>.cells.tfvars.json. The first one,
testidpna, has two cells, usw2 and cac1. You apply and destroy it as
one environment, and the entrypoint takes the steps in order:
IMAGE_TAG=<tag> deploy/aws/terraform-local.sh testidpna apply
deploy/aws/terraform-local.sh testidpna destroy
TF_CELL=cac1 deploy/aws/terraform-local.sh testidpna output # one cell
You can also use the testidp deploy and testidp destroy workflows with
environment: testidpna.
testidp and testidpna cannot both exist, because they use the same
public name. Destroy one before you apply the other.
After the first cell, the cells of each step are applied at the same time
(six at once by default; set TF_CELL_PARALLEL to change it). Each cell’s
output lines start with its id, for example [euw1].
Cost: each cell costs about as much as testidp, roughly $0.50 an hour.
Add a global replica per extra cell and data sent between regions.
globalidp: six regions
globalidp is the six-region test case, at global-idp.iyasec.io. It has
its own name, so it can exist beside testidp or testidpna.
| Cell | Region | Jurisdiction | Pinned countries |
|---|---|---|---|
usw2 |
us-west-2 (Oregon) | us |
none (holds the global database) |
use2 |
us-east-2 (Ohio) | us |
none |
euc1 |
eu-central-1 (Frankfurt) | eu |
the EU countries, Iceland, Liechtenstein, Norway |
euw1 |
eu-west-1 (Ireland) | eu |
(the same) |
apse1 |
ap-southeast-1 (Singapore) | sg |
Singapore |
apse5 |
ap-southeast-5 (Malaysia) | my |
Malaysia |
Each region’s own console is at <cell>.global-idp.iyasec.io, for example
euw1.global-idp.iyasec.io.
IMAGE_TAG=<tag> deploy/aws/terraform-local.sh globalidp apply
deploy/aws/terraform-local.sh globalidp destroy
Before the first apply, an administrator enables ap-southeast-5 on the
account (it is an opt-in region) and re-applies foundation/. It costs
about $3 an hour while it exists.
Adding a region
- If the region is opt-in, enable it on the account.
- Add it to
permitted_regionsindeploy/aws/foundation/variables.tf, and have an administrator re-applyfoundation/. - Add a cell to the environment’s
.cells.tfvars.json. Its id is the region’s name shortened: the area, the first letter of each direction word, and the number.eu-west-1iseuw1andap-southeast-5isapse5. Give it a jurisdiction and a VPC CIDR that no other cell uses. To pin countries to a new jurisdiction, add them underjurisdictions. - Apply the environment.
Before anything is created, the apply checks each cell’s region: that it is enabled, that the deployer may use it, and that it offers the database version.
Converting testidp into testidpna without losing its data
testidp’s database becomes the database of cell usw2, so its people,
applications, keys and risk datasets and history carry over. Destroying
testidp deletes its database with no final snapshot and deletes its secrets
straight away, so two things have to exist before you destroy it:
- a snapshot of its database, copied so that it is encrypted with the
usw2cell’s key (a restored database keeps its snapshot’s key); - a carry-over secret,
iya-sts/carryover/testidp, that holds the values the database was written with: the key-encryption key, the management API client’s secret, the bootstrap administrator’s password and the Kerberos service password.
deploy/aws/convert-to-cells.sh checks both and makes them. Run it with no
flags first. It only reads, and it prints what is missing and the whole
sequence:
deploy/aws/convert-to-cells.sh testidp testidpna
Then:
- An administrator re-applies
foundation/with both regions inpermitted_regions. This also lets the deployer copy and restore snapshots. deploy/aws/convert-to-cells.sh testidp testidpna --carry-secretsdeploy/aws/convert-to-cells.sh testidp testidpna --copy-snapshot, then the check again, until nothing is missing.- Destroy
testidp. The service is down from here until step 5 finishes. Anything written after the snapshot was taken is lost; the check prints how to stop writes and take a new snapshot first.deploy/aws/terraform-local.sh testidp destroy - Apply
testidpnawith the conversion:TF_CONVERT=1 IMAGE_TAG=<tag> deploy/aws/terraform-local.sh testidpna applyThe
usw2database is restored from the snapshot. Its nodes are held at zero while a one-off task converts the data. If the conversion fails, the apply stops with the nodes still at zero and nothing lost; fix the cause and run the same command again. - Sign in with the same bootstrap password as before
(
iya-sts/testidpna/bootstrap-admin-password) and check Monitoring → Risk.
After that, apply testidpna without TF_CONVERT. When you are sure
the cell is good, delete the carry-over secret and the two snapshots.
The snapshot and the secret’s name are in
deploy/aws/environment/envs/testidpna.conversion.tfvars.json, which is
used only when TF_CONVERT=1 is set. A new testidpna built without it
starts empty.
Trying two cells on one machine
The test suite can run two cells locally, with no AWS account:
./run-tests.sh --modes=cells --only=sts_cells --protocol=only --no-browser
This starts cell cella (jurisdiction us) and cell cellb (jurisdiction
ca), each with its own database, and a third database for the global tier.
Both cells answer as https://sts:8081. The cells mode runs only when you
name it; a plain ./run-tests.sh does not start it.
The nine sts_cells_* jobs check, over HTTP:
- each cell sees the other and never reports where it is;
- a login name is unique across cells, and a person can be created in the other cell;
- a person homed in
cellbwho starts signing in atcellais sent back to the start of the flow, finishes it incellb, and gets tokens that verify against the one key set of the realm; - where a realm permits it (
cells.permittedTransfersset toca>us), the session moves tocella, and disabling the person incellbends it there; cellb’s people can be listed fromcellaonly where the realm permits it.- a person moved from
cellbtocellaloses what they held, keeps theirsub, and signs in atcella; a move the realm’s jurisdictions forbid is refused; - an administrator can sign in to the console through
cellb, and a directory bind atcellbis checked in the person’s home cell; - when
cellacannot reachcellb, a sign-in, a refresh and a directory bind are refused, andfail-openlets a session held atcellabe refreshed.
Reading the logs
aws logs tail /iya-sts/containers --log-stream-name-prefix dev-node-a --since 1h
Each node’s stream carries two containers: schema-init, which applies
postgres/schema.sql as the RDS master user before the node starts, and
iya-sts. A node that never became healthy usually says why in the first of
them.
GET /admin-api/cluster (with an /admin-api token) lists the live nodes, and
GET /admin-api/secrets shows where the key-encryption key is (the AWS KMS
key and its rotation status, or the Secrets Manager secret) and that the
database password was read from AWS Secrets Manager, and when.