Secrets leak through git history, CI logs, crash dumps, and screenshots more often than through cinematic nation-state attacks. If a production database password ever lived in a Slack thread, treat it as burned—even if “only three people” saw it.
What counts as a secret
Anything that grants access or signing power:
- API keys and access tokens
- Database passwords and connection strings
- Private keys, TLS materials, signing secrets for JWTs/webhooks
- OAuth client secrets
- Encryption keys (separate from passwords when possible)
Non-secrets that teams still over-restrict: public client IDs, non-sensitive feature flags, and docs URLs. Over-classifying clutter makes real secrets harder to govern.
BAD PATH BETTER PATH
──────── ───────────
Code → .env committed Code → fetch at runtime
CI prints env on failure CI uses masked vars + OIDC
Laptop .env shared via zip Secret manager + short TTLs
Same key in all environments Per-env secrets, least privilegeDo this instead: the control loop
1. Inventory
List every secret by service and environment. Include “shadow” secrets in old serverless env vars, container task defs, and mobile apps (mobile cannot truly hide secrets—treat client-shipped keys as public).
2. Move out of code
Store in a manager: AWS Secrets Manager / SSM Parameter Store (SecureString), GCP Secret Manager, Azure Key Vault, HashiCorp Vault, etc. Reference by name/ARN in platform config—not by pasted value in Terraform state if you can avoid it (and lock down state storage either way).
3. Inject at runtime
Prefer workload identity (IRSA, Workload Identity Federation, managed identities) over long-lived static cloud keys on disks.
// Pseudocode: fetch at startup or on rotation tick — never bake into the image
import { SecretsManager } from "@aws-sdk/client-secrets-manager";
const sm = new SecretsManager({ region: process.env.AWS_REGION });
const out = await sm.getSecretValue({ SecretId: "prod/payments/stripe" });
const stripeKey = JSON.parse(out.SecretString!).apiKey;4. Least privilege IAM
Grant GetSecretValue on specific resource ARNs to specific roles. No wildcard secretsmanager:* on * for app roles. Separate human break-glass roles from runtime roles.
5. Rotate on a schedule and after incidents
Rotation without consumer reload is theater. Pair rotation with:
- dual-key windows (accept old + new briefly), or
- rolling restarts that refresh caches, or
- sidecars/agents that hot-reload
Warning: Revoke any key that ever appeared in a chat, ticket, screenshot, or CI log. Assume scrapers and indexes already copied it.
Never do this
- Commit
.envfiles with production keys (or “temporary” prod copies) - Bake secrets into Docker layers (
ENV DB_PASSWORD=...) - Log headers that contain
Authorizationor cookies - Share prod credentials into staging “so we can reproduce”
- Email private keys as zip attachments
- Store the only copy of a root key in one engineer’s password manager without escrow
Git history is forever (almost)
If a secret hit git:
- Rotate/revoke first (history scrub does not un-leak clones).
- Purge with documented history rewriting only if policy requires—and coordinate with every fork/clone.
- Add scanning:
gitleaks, vendor secret scanning, pre-commit hooks.
# Example: fail the build if high-entropy secrets appear
gitleaks detect --source . --verbose --exit-code 1Real incidents (condensed)
| Mistake | What happened | Fix |
|---|---|---|
| Debug logging of config map | API keys in centralized logs | Redact + structured allowlists |
Terraform sensitive forgotten in output | Key in PR comment | Mark sensitive; break glass review |
| Copied kube secret to Slack for “quick debug” | Token replayed | Short-lived tokens + Session Manager |
| Same JWT secret across staging/prod | Staging leak → prod forge | Per-env signing keys |
Checklist you can paste into a ticket
- Inventory secrets per service/env
- Remove secrets from repo and container images
- Store in a manager; wire runtime fetch or platform injection
- Restrict IAM/Kubernetes RBAC to named secrets
- Enable automatic rotation where supported
- Add secret scanning in CI and pre-commit
- Document break-glass and revocation steps
- Verify logs/metrics/traces scrub sensitive fields
- Split prod vs non-prod credentials completely
- Test restore: can on-call rotate under 15 minutes?
App-level patterns that help
// Cache with TTL; refresh before expiry; fail closed if fetch fails in prod
type SecretCache = { value: string; expiresAt: number };
async function getDbPassword(cache: SecretCache | null) {
if (cache && cache.expiresAt > Date.now() + 60_000) return cache.value;
const value = await fetchSecret("prod/db/password");
return value;
}For webhooks, prefer HMAC verification with rotating secrets and reject replayed timestamps. For service-to-service auth, prefer mTLS or short-lived JWTs over static shared passwords.
CI/CD identity beats long-lived keys
Static cloud access keys in CI are a classic breach pattern. Prefer OIDC federation from GitHub Actions, GitLab, or your runner identity into temporary cloud credentials scoped to one repo and one environment. Mask secrets in logs, block workflow printenv debugging in prod pipelines, and separate staging/prod OIDC trust policies so a compromised PR workflow cannot mint prod roles.
GitHub OIDC token → Cloud STS AssumeRoleWithWebIdentity
→ short-lived creds
→ deploy job (no AWS_SECRET_ACCESS_KEY in repo secrets)Developer ergonomics so people do not bypass you
If local setup requires a scavenger hunt, engineers will paste prod values into .env. Provide a documented dev secret path, sealed local templates with fake values, and a one-command login to the secret manager for authorized humans. Security that fights developer experience loses quietly.
Closing
Secret management is operational hygiene: inventory, manager, identity-based access, rotation, and scanning. The elite attack path is rare; the pasted key in a ticket is common. Treat every exposure as a revoke event, then make the next exposure harder with automation—not with another wiki page nobody reads.
