Runbook: isolated signer recovery
The signer is the separate process that holds private-key operations. In an
isolated Helm topology it runs as its own Deployment from
deploy/helm/trstctl/templates/signer-deployment.yaml, speaks only gRPC over
mTLS, and uses the signer key store plus either a local KEK mount or Helm
externalKMS DEK wrapping. Recovery means getting that same signer identity and
key store healthy again, not silently making a new CA.
Prerequisites
- Decide whether this is a restart failure, storage failure, or key compromise. Use incident response for suspected compromise.
- Confirm the signer PVC or configured key store, signer auth Secret, signer mTLS Secret, and the configured custody input still exist. Local-KEK installs need the KEK Secret; externalKMS installs need the same provider, keyRef, wrapper adapter, and provider credentials.
- Confirm the control plane can still answer
/healthz;/readyzmay be503because it checks the signer. - Before changing the signer, capture a redacted offline diagnostic:
trstctl support-bundle --output signer-recovery-support.tar.gz --log-file <control-plane-log>. It runs even when the control-plane HTTP listener never started and distinguishes signer, datastore, event-spine, migration, and queue posture without recording socket paths, DSNs, tenant IDs, raw environment values, or credentials. - Capture
trstctl_signer_up, signer pod logs, control-plane logs, agent heartbeat age, and inventory counts before changing anything. - Have a recent full backup if the signer key store, signer auth Secret, local KEK, or externalKMS adapter configuration must be restored.
Commands: inspect and restart
kubectl -n trstctl get deploy,pod,pvc,secret | grep -E 'trstctl|signer'
kubectl -n trstctl logs deploy/trstctl-signer --tail=200
kubectl -n trstctl describe deploy/trstctl-signer
kubectl -n trstctl rollout restart deployment/trstctl-signer
kubectl -n trstctl rollout status deployment/trstctl-signer --timeout=10m
Then check the control plane:
curl -fksS https://cp.example.com/readyz
curl -fksS https://cp.example.com/metrics | grep trstctl_signer_up
trstctl-cli agents list
trstctl-cli certificates list
Commands: restore signer storage
If the signer pod cannot open /data/signer/keys, /etc/trstctl/kek/kek.bin, or
the configured externalKMS wrapper, restore the key store and the matching
custody input from the last known-good operational backup. Do not delete and
recreate the key store as a "quick fix"; that changes the CA identity.
kubectl -n trstctl scale deployment/trstctl --replicas=0
kubectl -n trstctl scale deployment/trstctl-signer --replicas=0
# Restore the signer PVC and custody input using your storage backup tooling.
# Local KEK mode must restore /data/signer/keys and /etc/trstctl/kek/kek.bin.
# externalKMS mode must restore /data/signer/keys and the same provider/keyRef/wrapper adapter.
kubectl -n trstctl scale deployment/trstctl-signer --replicas=1
kubectl -n trstctl rollout status deployment/trstctl-signer --timeout=10m
kubectl -n trstctl scale deployment/trstctl --replicas=1
kubectl -n trstctl rollout status deployment/trstctl --timeout=10m
If the key store is lost and no backup exists, stop. Open a CA key ceremony and re-issue under a new CA. That is a planned trust migration, not signer recovery.
Expected metrics and logs
/readyzchanges from signer-degraded to200.trstctl_signer_upchanges from0to1.- Signer logs show a listener on
:9443and no KEK, externalKMS, or key-store open errors. - Control-plane logs stop reporting signer health failures.
- Existing agents heartbeat again without re-enrollment.
- Inventory counts remain stable; signer restart should not erase certificates, agents, or SSH keys.
Abort criteria
Stop and escalate when:
trstctl_signer_upremains0after one clean signer rollout./readyzstill reports signer failure after the pod is Ready.- The signer starts with an empty key store when a non-empty one was expected.
- Agent heartbeats resume under a different tenant or CA bundle.
- Inventory counts drop after signer recovery.
- Any log suggests private-key compromise, KEK mismatch, externalKMS unwrap failure, or unauthorized handle use.
Rollback commands
Rollback means return to the last known-good signer Deployment and storage snapshot:
helm history trstctl -n trstctl
helm rollback trstctl <last-good-revision> -n trstctl --wait --timeout=10m
kubectl -n trstctl rollout status deployment/trstctl-signer --timeout=10m
If the new signer pod is unsafe, stop issuance while you investigate:
kubectl -n trstctl scale deployment/trstctl-signer --replicas=0
This makes issuance and renewal fail closed; it is better than serving with an unknown key store.
Post-checks
/readyzis200.trstctl_signer_up == 1.- Agent heartbeat age returns to the normal interval.
trstctl-cli agents liststill has the expected fleet count.trstctl-cli certificates listand other inventory counts match the pre-incident baseline.- Record whether recovery used restart only, storage restore, or CA key ceremony.