Troubleshooting internals
The Troubleshooting pages give the fix for each error. This page holds the mechanism behind some of those entries, for someone changing cdkd or reading a run closely. Nothing here is needed to apply a fix.
Lock acquisition
cdkd deploy, cdkd rollback and cdkd scrub retry a held lock 3 times at
2-second intervals: 4 attempts, about 6 seconds. Two refinements:
- When an attempt fails but cdkd then finds no lock, the holder released it in between. cdkd tries again at once without counting that as a retry, up to 3 extra attempts per command.
- Past those, an attempt that finds no readable lock waits the same 2 seconds.
If the last one still fails, the message says
No lock could be read after the last failed attemptinstead of naming a holder. Alock.jsonthat is not a readable lock also fails this way.
The commands that write state without deploying fail fast on contention:
cdkd destroy, cdkd state destroy, cdkd import, cdkd export,
cdkd orphan, cdkd drift --accept, cdkd drift --revert and
cdkd state refresh-observed. The exception is cdkd export's nested-stack
children, which retry 6 times at 5-second intervals, a ceiling of about 30
seconds.
Nested-stack children are locked separately. A child teardown reached from
cdkd deploy (because the template no longer declares the child) or from
cdkd rollback fails fast on the child's lock, even though the parent's lock
was acquired with retry.
cdkd orphan and cdkd state orphan differ in how they treat the lock:
| Command | Lock behaviour |
|---|---|
cdkd orphan '<path>'... |
Takes the lock; fails fast when it is held |
cdkd state orphan '<stack>'... |
Takes no lock; refuses while one is held; --force deletes that lock, including a live one |
cdkd state orphan '<stack>' --resource <logicalId> |
Takes the lock like cdkd orphan; --force does not bypass a held one |
The lock record
lock.json holds owner, timestamp, expiresAt and an optional
operation. expiresAt is what the expiry check, the expires in clause of
the error and the renewal loop all read. It is timestamp plus the 30-minute
TTL, rewritten every 2 minutes while the holder runs. The TTL therefore
measures silence, not duration: it tolerates fourteen consecutive missed
renewals before it lapses. A truncated lock.json with no expiresAt is
treated as expired.
Signals
deploy, destroy, state destroy and rollback handle both SIGINT and
SIGTERM. The first signal finishes in-flight operations, saves state and
releases the lock. On a deploy it also records a rollback journal.
GitHub Actions escalates SIGINT to SIGTERM about 7.5 s later and to
SIGKILL about 2.5 s after that. A job whose in-flight AWS operation outlasts
that window is killed before the lock release runs.
A second signal force-quits with exit 130. destroy and state destroy
fire an un-awaited lock release. deploy does not attempt one, because a
deploy may hold locks on several stacks. rollback has no second-signal path:
further signals are absorbed by its graceful handler, so stopping it takes
SIGKILL, and the lock then has to be cleared.
The state save and its conflict check
The write of state.json is guarded by an S3 precondition on the ETag cdkd
read. A conflicting write is refused instead of overwritten, and cdkd does not
retry the save: a re-run re-reads state and plans against whatever the other
writer left. The lock excludes other cdkd processes; the ETag precondition is
a separate guard on the object.
Cross-region state bucket
When the state bucket is in a different region from the SDK client, S3
answers a HEAD request with an empty-body 301. The AWS SDK v3 region-redirect
middleware does not handle that response, and the protocol parser produces a
synthetic Unknown exception with the literal message UnknownError.
cdkd avoids the condition by looking up the bucket region with
GetBucketLocation (a GET, which does not hit the same path) and rebuilding
its S3 clients for that region. Three components do this before any
operation: the state backend, the lock manager, and the custom-resource
response path that reads the pre-signed ResponseURL.
On the state-bucket path cdkd rewrites the synthetic error into a sentence keyed on the HTTP status, which names the region. The lock path does not rewrite and surfaces the raw 301.
How cdkd prevents orphans
- In-memory state update per resource. Each resource updates the in-memory state as soon as it is provisioned.
- Partial state save per resource. After each successful resource the state is persisted to S3, serialized through a save chain to avoid ETag conflicts. A crash mid-deploy therefore loses at most the in-flight operations.
- Save before rollback. If a resource fails, cdkd saves the in-memory state to S3 before attempting rollback, so resources completed concurrently with the failed one are tracked.
- Save after rollback. After rollback completes, or is skipped with
--no-rollback, state is saved again. - Rollback journal. On a
--no-rollbackfailure, an interruption, or before an automatic rollback, cdkd writesrollback-journal.jsonbesidestate.json, recording which operations completed.cdkd rollbackreplays it.
The journal's lifetime
- The next successful deploy and
cdkd destroydelete the journal. Each first deletes, per itsDeletionPolicy, any resource a failed create made that only the journal records. - A successful deploy that cannot act on such a resource keeps the journal,
reduced to that entry where it can. One that a state record may own is
warned about and left in AWS. Either way the deploy exits
2. - After a clean automatic rollback the journal is reduced to a failed-only
segment instead of deleted. The completed operations are already reverted,
but the failed resource's pre-operation record is kept so
cdkd rollback --revert-failedcan still revert a possibly half-applied resource. The next successful deploy clears it. - An automatic rollback that skipped an operation it could not revert
(a
ROLLBACK_RESOURCE_SKIPPEDevent) is not clean and keeps the full journal. - A nested stack's successful deploy keeps its journal until the top-level stack's deploy succeeds, so a failure of the parent can revert the child.
cdkd rollback exit 2
A partial rollback exits 2. Besides failed operations (journal kept) and
operations skipped with a warning (segment cleared, ROLLBACK_RESOURCE_SKIPPED
recorded), two more cases exit 2:
- A reverted operation left a resource cdkd no longer tracks: a new copy
retained by
UpdateReplacePolicy: Retain, or one whose delete failed. ItsROLLBACK_RESOURCE_SUCCEEDEDevent carries the survivor's id and areason. - A reverse-replacement whose re-create returned the live new resource, so the
replacement was not fully reversed. Its event carries a
reasononly.
Retry schedules
cdkd retries create, update and delete operations during cdkd deploy and
the replays performed by cdkd rollback and cdkd drift --revert. The
schedule is chosen per error class.
| Class | Schedule | Retries | Sleep |
|---|---|---|---|
| Throttling and transient | 1s, 2s, 4s, 8s, 8s, 8s, 8s, 8s |
8 | 47s |
| IAM propagation | 0.25s, 0.5s, 1s, 2s, 2s, ... |
26 | 47.75s |
| Name cooldown | 2s, 4s, 8s, 10s, 10s, 10s, 10s, 10s |
8 | 64s |
Throttling and transient covers rate limits, a resource still leaving
Pending, an async delete releasing a dependency, and HTTP 500 / 502 / 503 /
504, the same four the AWS SDK's own retry strategy treats as transient.
IAM propagation covers Invalid IAM Instance Profile, cannot be assumed, not authorized to perform, Policy Error: PrincipalNotFound,
not authorized to access the Log Destination, does not have a trust relationship allowing and similar. cdkd creates an IAM entity and consumes it
one to three seconds later, faster than IAM propagates, and this class usually
resolves in single-digit seconds. The dense schedule re-probes about every 2
seconds instead of idling through a 4s or 8s step. Its window is slightly
longer than the generic one, so nothing that recovers under the generic
schedule stops recovering.
A few of these patterns are also emitted for permanent problems. On those the budget is spent and the permanent failure surfaces afterwards. The Step Functions log-destination rejection is the main example.
Name cooldown covers QueueDeletedRecently, Step Functions'
StateMachineDeleting and S3's conflicting conditional operation. It is a
separate class because SQS's own message names a 60-second window, and a 47s
budget against a 60s window does not converge.
Counters on the give-up line
- The seconds on the line count propagation backoff only. An interleaved throttle's own wait is excluded, so the figure compares against the 47.75s budget and is not total elapsed time.
- A throttle mid-sequence consumes an attempt without counting as a propagation retry, so the budget can run out at 25 or fewer retries.
- A sequence that spent its budget on transient server errors says
transient server-error retries (HTTP 5xx), and a mixed one reports both. - A rollback's reverse-replacement re-create, the arm that revives the old resource after a replacement failed, retries on the same dense schedule and prints the same line.
Providers outside the outer retry
Custom::* / AWS::CloudFormation::CustomResource and
AWS::CloudFormation::Stack are never wrapped by the outer schedule, so no
give-up line is produced for them. A custom resource's Lambda handler is an
ordinary AWS::Lambda::Function and does retry on the outer schedule.
The custom-resource provider retries internally, on two separate budgets:
| Failure | Budget |
|---|---|
The handler returns FAILED with an authorization-shaped reason |
CDKD_CR_AUTHZ_MAX_RETRIES re-invocations (default 2, clamped to 10) |
| An authorization-shaped error thrown by an SDK call the provider makes before the request reaches the handler | 26 retries over 47.75s, as the outer schedule |
A re-invocation re-runs the user's handler, so that budget stays small. A
pre-delivery throw reached no handler. CDKD_CR_AUTHZ_MAX_RETRIES=0 disables
re-invocations only; it does not disable the pre-delivery retry or the retry
of the response-placeholder PutObject.
Two cases are never replayed:
- A throw after the invoke was accepted. The handler is running and will write to that attempt's response URL.
- The readiness waiters, which have already polled
lambda:GetFunctionfor their own 600 seconds. A permanent denial there surfaces after one waiter timeout.
AWS::CloudFormation::Stack has no internal retry.
Cloud Control polling
Polling is a different mechanism from the retry above. It waits for an
accepted operation to finish, on a 1s, 1.5s, 2.25s, 3.4s, 5.1s, 7.6s, 10s
schedule (a 1.5x multiplier capped at 10s), against a 15-minute deadline. The
deadline is longer for known-slow types such as OpenSearch domains and RDS,
Redshift and ElastiCache clusters.
When cdkd cannot reach Cloud Control, the AWS SDK retries a few times and cdkd
re-polls the same request token for up to two more minutes. For the
deletion-protection flip that cdkd destroy --remove-protection does first,
the allowance is ten seconds, since that step is best-effort and the delete
reports on it. A local-address failure (EADDRNOTAVAIL, seen under heavy
parallel load) is treated like an outage: checking the status is a read, so
cdkd re-polls.
Routing between the SDK provider and Cloud Control
A resource recorded provisionedBy: cc-api stays on Cloud Control on later
deploys. A cdkd release that adds SDK coverage for the property that sent it
there does not move it back by itself, because doing so for every type would
mean destroy-and-recreate churn on each release.
Types carrying an 'sdk-coverage' exemption are the exception: Cloud Control
manages the type but is slower, and cdkd covers every property the resource
uses, so the next mutating deploy moves it back in place. For
AWS::ElasticLoadBalancingV2::Listener, Cloud Control also leaves a removed
ListenerAttributes key at its old value.
cdkd diff annotates routing with [via CC API: <property>],
[via CC API: sticky] and [returning to SDK provider]. The last one does
not share the via CC API: prefix because the resource is leaving Cloud
Control.
A type exempt for the opposite reason, because Cloud Control cannot manage it
correctly, moves for correctness, logs a different line and ignores
--pin-cc-api.
Ambiguous creates
The AWS SDK retries HTTP 5xx responses by default. For create calls that carry no idempotency token, cdkd turns that retry off so the failure reaches its own retry, which first looks for what the failed attempt may have made.
CreateDistribution is handled through CallerReference, CloudFront's
idempotency key. cdkd sends a value that is stable across every attempt of
one logical create, so CloudFront refuses the replay. The first attempt's
distribution cannot be adopted: ListDistributions returns Comment but not
CallerReference, and GetDistribution needs the Id the lost response
carried.
Proxy handling
The AWS SDK for JavaScript v3 does not read the proxy environment variables
the way botocore (the AWS CLI) and Go's net/http do. Its
guide
states that a proxy is supplied "through a third-party HTTP agent" by whoever
constructs the client. cdkd therefore installs a routing agent on every SDK
client it builds.
Node's NODE_USE_ENV_PROXY=1 is not a substitute. It rewires the global
agent, while every SDK client builds its own.
On spelling precedence, the AWS CLI prefers the lower-case variable as cdkd
does, while Go's net/http prefers the upper-case one.
A CIDR entry in NO_PROXY is ignored because a hostname never contains /,
so the entry falls into the exact-match branch and can never match.