Appearance
Secret protection (Auth module)
Protection at rest of the secrets GrydAuth must be able to read back. Configuration section:
GrydAuth:Security:SecretProtectionDesign decisions: ADR 0005
1. What this protects, and what it does not
| Secret | Column | Why it must be reversible |
|---|---|---|
| User TOTP seed | MfaFactors.EncryptedSecret | The seed is needed to compute and verify the code |
| Enterprise IdP client secret | IdentityProviderConnections.ClientSecretEncrypted | It is sent to the IdP on every token exchange |
Passwords are not here and never will be. They are hashed one-way with Argon2id (GeraltArgon2idPasswordHasher) and are never recovered — login compares hash against hash. These two are a different category: the plaintext genuinely has to come back, so they are encrypted, not hashed. Do not "harden" this by hashing them; it would end TOTP and enterprise SSO outright.
2. Configuration
jsonc
"GrydAuth": {
"Security": {
"SecretProtection": {
"Scheme": "KeyRing", // KeyRing (default) | KmsEnvelope
"ActiveKeyId": "gryd-secret-2026-07",
"Keys": {
"gryd-secret-2026-07": "${GRYD_SECRET_PROTECTION_KEY}"
}
}
}
}Generate a key:
bash
openssl rand 32 | basenc --base64url | tr -d '='| Property | Required | Notes |
|---|---|---|
Scheme | no | KeyRing (keys from configuration) or KmsEnvelope (per-record data key wrapped by a KMS-held KEK). |
ActiveKeyId | yes (KeyRing) | Id of the key new secrets are protected with. Must be present in Keys. |
Keys | yes (KeyRing) | keyId → 32-byte Base64URL key. Retired ids stay here until the re-wrap finishes. |
Kms:Provider | yes (KmsEnvelope) | AzureKeyVault. |
Kms:KeyId | yes (KmsEnvelope) | Full key identifier, e.g. https://vault.vault.azure.net/keys/gryd/<version>. |
Kms:DataKeyCacheLifetime | no | Default 00:10:00, hard ceiling 00:15:00. |
Kms:DataKeyCacheSize | no | Default 1024 unwrapped data keys held in process. |
Environment-variable form: GrydAuth__Security__SecretProtection__ActiveKeyId, GrydAuth__Security__SecretProtection__Keys__<keyId>.
This is validated at startup, in every environment
There is no Development exemption. A host without a usable key does not boot, and the failure message names the section and shows the command that generates a key.
That is deliberate. The previous validator let Development start with an empty section, so the first write failed instead — as an HTTP 409 Conflict whose message mentioned MFA, returned from an endpoint that was creating an identity provider. Configuration that is mandatory has to fail where it is cheap to diagnose.
3. How a protected value is built
g2.<scheme>.<keyRef>.<nonce>.<tag>.<ciphertext>g2— magic + format version.scheme—kr(configured key ring) orkms(KMS envelope).keyRef— opaque to everything except the owning scheme: forkrit is the key id; forkmsit isbase64url(kekId)~base64url(wrappedDataKey).- the last three segments are unpadded Base64URL.
Everything before the ciphertext is public metadata. It is authenticated but not encrypted, so altering any of it fails the AEAD tag.
The associated data is the whole point
g2 | scheme | keyRef | purpose | contextpurpose and context come from the SecretProtectionScope the caller passes:
| Purpose token | Context | Bound to |
|---|---|---|
mfa.totp-secret | user:<userId>/factor:<factorId> | the user and the factor |
federation.idp-client-secret | tenant:<tenantId>/connection:<connectionId> | the tenant and the connection row |
The associated data is never stored — it is rebuilt on read from the row's own identity. So a ciphertext only opens in the exact place it was written:
- a TOTP seed moved into the client-secret column fails;
- one tenant's client secret placed in another tenant's connection fails;
- a seed lifted onto another user's factor fails.
Before this design, all three silently succeeded. That is the defect this subsystem was rebuilt to close.
Both tokens and both context formats are wire values. They are mixed into key derivation and into every ciphertext at rest. Renaming a C# member is safe; changing a token or a context format is a breaking change that makes existing ciphertext unreadable.
Keys are derived, not used directly
The configured key is a master key. The key that actually encrypts is:
HKDF-SHA256(masterKey, salt = keyId, info = "gryd.secret.v2|" + purpose)One key per keyId for the operator, one independent subkey per purpose in practice: leaking the MFA subkey exposes no client secret, and vice versa.
4. Errors an operator will see
| Code | HTTP | Meaning | Fix |
|---|---|---|---|
SECRET_PROTECTION_UNAVAILABLE | 503 | No active key, or the envelope names a key this instance does not hold. | Check the section is deployed; if a key was retired, put it back and finish the re-wrap first. |
SECRET_PROTECTION_INVALID | 500 | The stored value is malformed, or failed its integrity check. | Investigate the row — a tag failure means the ciphertext does not belong where it is stored. |
Responses never carry the key id, the scope context, plaintext or ciphertext; the structured log carries the diagnosis.
A federated login whose client secret will not decrypt is refused as FederationException(ProviderUnavailable) rather than attempted with an empty secret — otherwise the operator would be debugging an authentication error at the customer's IdP instead of a key problem here.
5. Rotating a key
The rotation surface is admin:system only and deliberately not tenant-scoped: one key protects every tenant's secrets, so a rotation that could only see one tenant could never retire it.
GET /api/v1/admin/security/secret-protection/status
POST /api/v1/admin/security/secret-protection/rewrap- Generate a new key and add it to
Keys, keeping the old one. - Point
ActiveKeyIdat the new key and reload configuration — the key ring reads throughIOptionsMonitor, so no restart is needed. POST .../rewrapwith{"dryRun": true}and check the counts. (An empty body is a dry run.)POST .../rewrapwith{"dryRun": false}. Repeat untiltotalRewrappedis0— the operation is idempotent.GET .../statusuntilisFullyRotatedistrueand nothing is counted under the old key.- Only then remove the old key from
Keys.
Skipping step 5 is the mistake that produces the classic failure: a customer's first federated login of the day returns 503 because its secret still names a key that was deleted last night.
A re-wrap is also a tamper sweep
Each row is decrypted and re-encrypted under the scope rebuilt from the row itself. So the re-wrap cannot repair a mis-bound ciphertext — it can only refuse it. A non-zero failures count is a signal, not noise: that row's ciphertext does not belong where it is stored. Failures are reported by row id and never abort the run.
6. Envelope encryption with a KMS (optional)
With Scheme: "KmsEnvelope" each record gets its own AES-256 data key, wrapped by a KEK that never leaves the vault. The record itself is still sealed with the same associated data, so nothing about the domain separation changes — only the custody of the key does.
Reading always follows the scheme named in the ciphertext, so the migration is additive:
- Deploy with the KMS scheme available but
Scheme: "KeyRing"— nothing changes. - Provision the KEK and grant the deployment identity
wrapKey/unwrapKey. - Flip to
Scheme: "KmsEnvelope". New secrets use the KMS; existing ones keep opening underkr. - Run the re-wrap (§5) to convert the backlog.
- Remove the
krkeys and their environment variables.
The data-key cache is in-process and must stay that way. Unwrapped data keys are cached in a private MemoryCache (bounded lifetime and entry count) so a federated login does not pay a network round trip. They must never reach the distributed cache: the federation descriptor deliberately carries the client secret to Redis still encrypted so a cache read is not a credential read — caching the key that opens it in the same Redis would undo that in one step.
Watch gryd_auth.secret_protection.kms.* and gryd_auth.secret_protection.data_key_cache.lookups (meter GrydAuth.Security.SecretProtection). An unwrap rate that tracks the login rate means the cache has regressed.
7. Migrating to this design (one-way)
The ResetSecretProtectionForV2 migration clears every secret protected under the old format. There is no reader for it by design, and the plaintext was never stored, so those values cannot be recovered.
Effect:
- every MFA factor is set to
Disabledand its owner must re-enrol; - every connection returns to
Draft; an admin must re-enter the client secret and re-activate.
Deploy order
- Publish
GrydAuth__Security__SecretProtection__*in every environment first. - Deploy — the migration runs and clears the secrets.
- Admins re-enter each connection's client secret and re-activate it.
- Notify users that MFA re-enrolment is required.
- Verify: one federated login end to end, and one TOTP challenge after re-enrolment.
Before the deploy, list the tenants with enforceSso = true and check their break-glass accounts: while their connection sits in Draft, those users have no login path except break-glass.
Rollback restores the binary, not the secrets. Treat the deploy as one-way and rehearse it in staging with representative data.