State schema internals
This page is the contributor reference behind State Management. It holds the version-by-version history of the state schema, the field-by-field reference, the lifecycle of every object in the state bucket, and the implementation of locking and of the noncurrent-version purges.
A user of the cdkd binary does not need this page: the public page covers
where state lives, how to read a record, how the lock behaves, and how to back
up and secure the bucket. The authority for everything here is
src/types/state.ts and src/state/**.
| Section | What it covers |
|---|---|
| Which region a record belongs to | The key decides; when a destroy or rollback refuses |
| Sidecar objects in the state bucket | Response placeholders, the create-token ledger, the rollback journal, the bootstrap marker |
| Bucket name and region resolution | The legacy-name fallback and the GetBucketLocation probe |
| Schema version history | version: 1 through version: 11 |
| Fields added without a version bump | skippedOutputs, conditionVerdicts, the masked-property fingerprints |
| Schema reference | StackState, ResourceState, LockInfo, malformed-record handling, physical ids |
| Lock implementation | Acquire, release, retry, renewal |
| State writes during deploy and destroy | Incremental saves and what a partial run leaves |
| Noncurrent-version purges | Which objects are purged, the warning, the replication check |
Which region a record belongs to
The key is what decides a record's region, not the region field inside
it. cdkd writes the two to agree — every save stamps the key's region into
the body — so on a record cdkd wrote there is nothing to choose between. When
they do disagree, cdkd uses the key's region and prints a warning that the
record was not written by cdkd; the operation then acts on the region you
pointed it at. The warning names the kind of value the body held (a string,
a number) rather than the value itself, since a region a record supplies can
be chosen by anyone able to write that key; run with --verbose to see it. A
body with no region at all is read as belonging to its key's region, the
same way a legacy (version: 1) record with no region is readable from any
region.
A destroy and a cdkd rollback are the operations that refuse on such a
record, and only when it still lists resources. Reading a record is safe under either answer,
but DELETING is not: if the key is the dishonest half, every delete would be
issued where the resources are not, come back not-found — which a destroy
reads as "already deleted" — and the run would report success, remove the
record, and leave your resources standing with nothing naming them. cdkd
cannot tell which half is honest, so cdkd destroy / cdkd state destroy
stop and say so — as does a cdkd deploy that removes a nested stack, which
destroys the child's resources through the same path. cdkd rollback refuses
for the same reason before replaying anything: its replay deletes and reverts
in the key's region and reads a not-found delete as already rolled back. It
stops saving, keeps the journal and exits partially when the record is
rewritten with a disagreeing region while the rollback runs. A record with no
resources is not refused: there is nothing to strand, so cleaning one up still
works. To act on the refusal, destroy
against the region the resources are really in, or correct the record's
region field to match the key it is stored under; cdkd state orphan still
drops the record and leaves the live resources standing.
Sidecar objects in the state bucket
The layout is drawn in State Management. This section covers the lifecycle of each object that sits beside state.json.
Custom-resource response placeholders
The custom-resource-responses/{requestId}.json placeholders are written by
CustomResourceProvider before each invocation, so the handler has a
pre-signed URL to PUT its cfn-response to. They are transient and the happy
paths delete them again, but three shapes strand one: an interrupted deploy
between the PUT and any cleanup, a throw on a path that reaches no cleanup
call, and a LATE handler PUT landing after cdkd stopped polling (the only one
that leaves real Data content rather than an empty body). cdkd gc collects
the stranded ones — see
cdkd gc for the
staleness rule and why an in-flight run's key is never taken.
Deleting one is not the same as removing it, because this bucket is
VERSIONED. cdkd bootstrap turns versioning on, so a plain DeleteObject
writes a DELETE MARKER and every earlier version of the key stays readable
through GetObject with a VersionId. That matters here more than anywhere
else under the state bucket: the object at this key is the handler's FULL
cfn-response, Data included, so a handler that mints a secret (a generated
password, an issued API key) put that value in the state bucket. Previously
both delete paths — the
provider's own cleanup and cdkd gc's collection — left it retrievable after
reporting the object gone. Both now purge the key's noncurrent versions as
well, scoped to that exact key so a concurrent deploy's live placeholder under
the same shared prefix is never touched.
This is not the whole account of where that value lives. A handler-minted
secret returned in Data also flows into state.json — into the custom
resource's own attributes, and into the resolved properties of every
resource that consumed it through Fn::GetAtt. That is a separate object with
a separate lifetime, and purging the response sidecar does nothing for it.
A handler can
opt those values out of the state file, by declaring the response sensitive —
see NoEcho custom-resource responses
below. Without that declaration the state file still carries the value.
That purge is conditional on s3:ListBucketVersions and
s3:DeleteObjectVersion on the state bucket — as is every other
noncurrent-version purge cdkd runs (the rollback journal and the bootstrap
marker below, the transient CFn template upload, and the deployments/ event
store) — see
IAM actions cdkd needs on the bucket
— older recommended policies did not grant either. It fails soft by design, because it
runs on a cleanup path that must never abort the operation it follows: without
those actions the deploy, destroy or cdkd gc run still succeeds, a warning
names the two grants, and the body stays retrievable by VersionId. A purge
that SUCCEEDS still leaves the bodies behind if the state bucket is replicated
— see
S3 replication keeps copies cdkd cannot delete.
Create-token ledger (create-tokens.json)
The create-tokens.json sibling is the stack's create-token ledger. EFS file systems, FSx
file systems and CloudFront origin access identities are created with a token
the service binds to the resource for its whole life, so the token decides which
resource a create gets back. The ledger holds a random nonce folded into those
tokens, and a sent entry per logical id: the token its latest create sent and
when it was first sent, written BEFORE the create is sent. It holds no secret.
- The nonce is replaced whenever cdkd lets go of a resource it made while the
resource still exists. The ledger is deleted, BEFORE
state.json, with the state record (cdkd destroy,cdkd state destroy,cdkd state orphan, an export, a rolled-back first deploy); if it cannot be deleted, the record is left in place and the command fails. Its nonce is rotated, and the let-go logical id'ssententry dropped, bycdkd orphan(before it saves the record, failing if it cannot) and by a deploy orcdkd rollbackthat keeps a resource (DeletionPolicy: Retain,UpdateReplacePolicy: Retain,cdkd rollback --orphan). So a file system or identity left behind that way does not hold the token the stack's next create sends, and that create makes a new one, as the AWS CDK CLI does. - A
sententry is kept until the deploy that wrote it SUCCEEDS and its state record names the resource. So a deploy interrupted anywhere -- during the create, or after it but before the state record named the resource -- is re-run with the same token, and takes the resource that interrupted run made. - A deploy that finds no state record but a ledger that one was saved under (an earlier cdkd version deleted the record and left the ledger) starts a fresh ledger instead of resuming, so it does not take over resources that deployment kept. A first deploy interrupted before its first state save still resumes.
- A ledger cdkd cannot read or write (including one written by a newer cdkd) refuses each EFS, FSx and origin access identity create in that deploy, naming the stack and the logical id, and is never overwritten: a create sent without the ledger's token could not be found again by a re-run. Re-run once the ledger is reachable.
- The first time a stack's ledger is started, an EFS create looks up a file system holding the token cdkd versions without the ledger sent, and warns naming it: an earlier destroy or replacement kept it, or an interrupted create left it, and this create does not take it over. It may hold data, so inspect it before deleting it. FSx has no lookup by token, so an FSx file system left that way is not named.
s3:DeleteObjecton the ledger key is required wherever deleting the stack record is: a denied delete stops the command before the record is touched, even when no ledger exists, because S3 checks the permission first.
Rollback journal (rollback-journal.json)
The rollback-journal.json sibling is written whenever a
deploy ends without a completed rollback — a --no-rollback failure, a
Ctrl+C interruption, or before an automatic rollback (so a rollback that
dies partway is resumable). It records the exact operations the failed
deploy completed (one segment per failed attempt) so cdkd rollback can
revert them with no synth. Each segment also
carries the op(s) that FAILED mid-deploy (failedOperations[] — pre-op
state + attempted properties; an additive field, no journalVersion bump)
so cdkd rollback --revert-failed can optionally revert them too. The one
place a deploy acts on the journal's contents while provisioning is an AWS::EC2::SecurityGroupIngress create that
AWS rejects as a duplicate: from the journal, it adopts the existing rule only
when a segment holds a FAILED attempt of the same logical id at the same rule that
no later completed op of the id superseded. (It also adopts a rule another record
of the same stack already holds, which is state rather than the journal.)
Otherwise it refuses. A completed op is never that
evidence: the failed deploy recorded its resource in state, so a later create of
the id means the resource was reverted, destroyed or replaced. Neither is a failed
CREATE that carries a physical id, nor a failed replacement UPDATE whose physical
id differs from its previous record's: both were recorded in state too. Removing a segment
(a clean rollback's settle, cdkd rollback, a nested child's settled or
orphaned segments, a successful deploy whose journal delete failed, or one that
keeps only the resources a failed CREATE made that it could not delete) leaves the
ids of its completed ops on the nearest older
segment (supersededLogicalIds, additive, no journalVersion bump), so the
attempt a reverted adoption consumed does not count again. cdkd import writes the
journal too: before its state write, it records each resource it adopts, as its
logical id, physical id and type, on every segment holding an op of that logical id
(importedResources, additive, no journalVersion bump), and supersedes each
marked id on the newest segment. cdkd rollback then runs none of a marked
segment's ops of that id. It warns about any op that recorded another resource (cdkd rollback). A failed op refused because the
resource it met belongs to someone else, or whose write AWS provably did not
apply (one send answered with a 4xx or a throttle), is recorded WITHOUT its
attempted properties, so it is never read back as an attempt; any other
failed write keeps them, since it may have landed. Every completed
UPDATE op additionally records whether the deploy left the OLD physical
resource alive (oldResourceRetained — also additive, also no bump; it is read
only for a replacement, but recording it unconditionally is what keeps an
ABSENT value meaning "written by an older binary"), so the
rollback acts on the verdict the deploy reached rather than re-deriving it from
the previous state record, which disagrees on the deploy that adds or removes
UpdateReplacePolicy: Retain. It also records the type the resource had BEFORE
the update (previousResourceType — additive, no bump): an op's own
resourceType is the template's, so after a Type change it names only the new
resource, and the rollback needs the old one to pick the provider that
re-creates it. A journal written before that field falls back to the previous
resource record the op already carries. It also records whether the deploy
deleted the old resource BEFORE creating the new one (oldDeletedBeforeCreate
— additive, no bump; --recreate-via-cc-api / --recreate-via-sdk-provider,
the update-unsupported fallback, --replace's delete-first fallback, a child
lost with its re-created parent): the
rollback then deletes the new resource before re-creating the old one, so a
port or name only one of them can hold does not collide. An absent value (an
older binary's journal) keeps the create-first order. A failed UPDATE carries
it too, so the rollback knows the old resource it names is gone. It is deliberately not part of the state
schema (its own journalVersion field, no StackState.version bump) and
not under the deployments/ prefix (that layer survives destroy by
design; the journal must not). Lifecycle: created on a failed / interrupted
deploy and before an auto-rollback; each replayed segment is popped; the
object is deleted on the next successful deploy, after a clean
cdkd rollback, and by cdkd destroy / cdkd state destroy; the last
three first delete (per its DeletionPolicy) any resource a failed CREATE
made that only the journal records. A clean
automatic rollback deletes such a resource too, then settles the journal to a failed-only
segment instead of
deleting it (operations: [] plus the remaining failed op records, reason: auto-rollback-clean) so cdkd rollback --revert-failed works in the
default deploy flow too. An automatic rollback is clean only with no failed
AND no skipped op: one that left an op unreverted (a ROLLBACK_RESOURCE_SKIPPED
event) keeps the full segment, and an attempt that wrote no segment of its
own (nothing completed, and its only failures creates refused before
anything was applied, or a segment write that failed) settles nothing, leaving older segments as they are. A nested stack ({Parent}~{Child}) differs:
its successful deploy appends a nested-pending-parent segment instead of
deleting the journal, and the journal is deleted when its top-level stack's
deploy succeeds; the parent's rollback replays it to revert the child (see
cdkd rollback). It carries
resolved properties, the same sensitivity class as state.json (no new
secret-exposure class). Every writer holds the lock of the stack whose journal
it writes — for a nested child's journal, the child's lock — so no optimistic
locking is needed.
Deleting the journal purges its noncurrent versions too, on every one of
those paths. The bucket is
versioned, so a plain DeleteObject would leave each earlier body readable
through GetObject with a VersionId — and failedOperations[] holds the
attempted properties of the FAILED write verbatim, which is where a literal
password lands when the resource that failed had one. Unlike state.json,
whose noncurrent versions ARE the recovery capability versioning exists for
and are left alone everywhere except an explicit
cdkd scrub, the
journal is transient by design, so nothing weighs against removing them. Like the sidecar purge above it fails
soft: without the two version grants the deploy / rollback / destroy still
succeeds and a warning names them.
Asset-storage bootstrap marker (cdkd-bootstrap/{region}.json)
The cdkd-bootstrap/{region}.json marker is written by cdkd bootstrap
(unless --no-assets) and records that the region opted into cdkd-owned
asset storage — its body names the region's asset bucket
(default cdkd-assets-{accountId}-{region}) and container-asset ECR repo
(default cdkd-container-assets-{accountId}-{region}; custom names via
cdkd bootstrap --asset-bucket '<name>' / --container-repo '<name>' —
every consumer reads
the names from the marker, never from the naming convention). Deploys read
the marker per
(account, region) to pick the asset mode: absent → legacy (publish to the
CDK bootstrap destinations verbatim, byte-identical to the behavior before
cdkd-owned asset storage existed);
present → cdkd-assets mode (asset publishing redirects to the cdkd storage
and template references are rewritten to match — see the asset-destinations
section in cdkd bootstrap; no state schema
change, the deployed properties simply carry the cdkd names); present but
bucket/repo deleted → hard error
(never a silent fallback). cdkd bootstrap --destroy removes the marker and
purges its noncurrent
versions as well — the marker carries no secret (it names the region's asset
bucket and container repo), so that is class completeness rather than a
disclosure fix. The marker deliberately lives OUTSIDE the
{STATE_PREFIX}/ prefix so stack listing never mistakes it for a stack, and
per-region keys mean concurrent bootstraps of two regions cannot race on a
shared object. cdkd state info lists the opted-in regions. Full design in
the asset-storage design note.
To opt a region back out, cdkd bootstrap --destroy --region '<r>' tears
down the region's asset bucket + ECR repo and deletes the marker last
(the reverse of the create-side marker-written-last ordering); add
--include-state-bucket to also delete the state bucket once every stack
is destroyed. See the teardown section in
cdkd bootstrap.
Because assets are content-addressed and never deleted on cdkd destroy,
the asset bucket / ECR repo grow over time; cdkd gc reclaims
unreferenced objects / images by scanning every state file in the state
bucket for asset references (with a 30d default age guard). See the gc
section in
cdkd gc.
Bucket name and region resolution
Legacy bucket-name fallback
Pre-v0.8 cdkd used cdkd-state-{accountId}-{region} as the default name.
For users who already bootstrapped under that scheme, the lookup chain in
resolveStateBucketWithDefault is:
Probe
cdkd-state-{accountId}(current default). If it exists, use it.If not found (
HeadBucketreturns 404 /NoSuchBucket), probecdkd-state-{accountId}-{profileRegion}(legacy default). If it exists, use it and emit a deprecation warning:Using legacy state bucket name 'cdkd-state-123456789012-us-east-1'. The default has changed to 'cdkd-state-123456789012'. To migrate, run: cdkd state migrate --region us-east-1 (add --remove-legacy to delete the legacy bucket after a successful copy; legacy support will be dropped in a future release.)If neither exists, fail with a "run cdkd bootstrap" error pointing at the new name.
The legacy fallback is temporary. It will be dropped in a future
release together with the cdkd-state-{accountId}-{region} legacy
bucket name. Users who already bootstrapped under that name should
migrate via
cdkd state migrate.
State bucket region probe
The state bucket can live in any AWS region — it does not have to match
your CLI's profile region or the regions you deploy stacks into. cdkd
auto-detects the bucket's region via GetBucketLocation (a GET, not a
HEAD — has a body and avoids the AWS SDK v3 region-redirect parsing
glitch on empty-body 301 HEAD responses) and rebuilds its state-bucket
S3 client to that region before any state operation.
All four S3 consumers of the state bucket do this: the state backend
(state.json reads/writes), the lock manager
(lock.json acquire/release — previously state
operations succeeded against a cross-region bucket but every lock
acquisition failed with S3's 301 PermanentRedirect), the exports
index store (_index/{region}/exports.json writes/removes for
Fn::ImportValue tracking — previously the index
write/remove also hit the 301; non-fatal, so the cross-region index was
silently never maintained), and the custom-resource response path
(custom-resource-responses/*.json placeholder writes + the pre-signed
ResponseURL the Lambda handler PUTs its cfn-response to —
previously a cross-region deploy of any stack carrying a
Lambda-backed Custom Resource failed hard with the 301, because the
pre-signed URL was signed against the deploy region's endpoint). A
SUCCESSFUL bucket-region lookup is cached per bucket name for the process
lifetime, so all four consumers share a single GetBucketLocation
call. A FAILED probe is deliberately not cached: the resolver never
throws, so a failure degrades to a best guess, and caching that guess
pinned every later consumer in the process to one transient error's
answer with no way to heal.
The probe itself is aimed at the caller's own region — falling back to
the AWS SDK's region chain (AWS_REGION, the shared config profile) and
only then to us-east-1. GetBucketLocation is answered by any regional
S3 endpoint for a bucket in the same partition, so the probe never needs
to know the answer to ask the question; it does have to REACH the right
partition, and the hardcoded us-east-1 endpoint it previously used
is unreachable from aws-cn / us-iso* — so outside the
commercial partition the probe could not run at all and every consumer
above silently proceeded against the commercial default.
This is intentionally scoped to the state-bucket S3 clients only.
Provisioning clients (Cloud Control API, Lambda, IAM, etc.) continue to
use the stack's env.region so resources are still created in the
region the CDK app declares.
Schema version history
Writers always emit the current version, version: 11. Readers accept every
earlier version and upgrade it in memory; the next write persists the current
shape. A binary that meets a version newer than it knows refuses the record
with an "Upgrade cdkd" error.
| Version | What it added |
|---|---|
| 1 | Legacy flat key layout, {prefix}/{stackName}/state.json (read path only) |
| 2 | Region-prefixed key layout, {prefix}/{stackName}/{region}/state.json |
| 3 | observedProperties on each resource |
| 4 | imports[] on the stack (Fn::ImportValue references) |
| 5 | deletionPolicy / updateReplacePolicy on each resource |
| 6 | parentStack / parentLogicalId / parentRegion on a nested-stack child |
| 7 | provisionedBy on each resource |
| 8 | outputReads[] on the stack (Fn::GetStackOutput references) |
| 9 | exportNames[] on the stack |
| 10 | observedBaselineRefused on each resource |
| 11 | noEchoLeaves / noEchoAttributeNames on each resource; NoEcho values stored as *** |
version: 2 and version: 4 have no section of their own below: the first is
the key layout every later section assumes, and the second is described with
imports in the schema reference.
Legacy layout (version: 1) — read path only
State files written by early cdkd versions used a flat per-stack layout:
s3://{STATE_BUCKET}/{STATE_PREFIX}/
└── {StackName}/
├── lock.json # not region-scoped
└── state.json # version: 1, region recorded inside the body
cdkd still reads this layout (looking up the legacy key only when its
embedded region field matches the requested region), and the next write
auto-migrates: it writes the new region-scoped key, then deletes the legacy
key. The legacy read path is temporary and will be removed in a future
release.
An older cdkd binary that only knows an earlier version will fail with
a clear error if it sees a higher-versioned blob (e.g. Unsupported state schema version 3. Upgrade cdkd.) instead of silently mishandling
unknown fields.
version: 3 adds observedProperties (v3+ writers)
Schema version: 3 adds an optional observedProperties field to each
ResourceState. Writers emit version: 3 or later. The on-disk key layout
(cdkd/{stackName}/{region}/state.json) is unchanged from version: 2 —
only the per-resource shape grew. v2 readers see a version: 3 blob and
fail clearly with the same "upgrade cdkd" error as above.
observedProperties is the AWS-current snapshot of a resource's
properties as captured by provider.readCurrentState immediately after
each successful create / update. The cdkd drift comparator prefers it
as the baseline so changes the user did not template (a manual tag added
in the AWS console, an inline policy attached out-of-band, etc.) surface
as drift instead of being silently ignored. Resources with
observedProperties: undefined (older state, or providers without
readCurrentState) fall back to comparing against properties.
One carve-out: a top-level key the template never declared whose
captured value was EMPTY ([] / {} / null) is skipped by the
comparator — such keys are typically populated AFTER the capture by a
sibling resource in the same stack (capacity-provider associations,
standalone lifecycle hooks / security-group rules) or by AWS itself,
and comparing them produced permanent phantom drift that
drift --revert then destructively "fixed". An
undeclared key captured with a real value is still compared.
Because it records what AWS returned, observedProperties can hold a value
the template never references — a password an operator set in the console
over a placeholder literal, for example. Secret redaction rewrites a value
only where the template spells a {{resolve:...}} reference, so such a value
is stored as AWS returned it, by design; see
cdkd import's note on it
and Values that are stored in the clear.
v2 → v3 upgrade is automatic on the next cdkd deploy. When the
deploy engine loads state and finds resources without
observedProperties (typical the first time you deploy after upgrading
from cdkd <0.47), it kicks off provider.readCurrentState for each in
parallel with the rest of the deploy and drains the result into state at
the final save. The deploy critical path does NOT wait on these reads —
the cost is bounded by the longest single readCurrentState (~200-300ms
in practice) once at the end of the deploy. NO_CHANGE-only deploys (no
diff to apply) still drain and persist the refreshed baseline so the
next cdkd drift run sees a real AWS-current snapshot. Pass
--no-capture-observed-state to disable both regular capture and this
upgrade refresh; cdkd state refresh-observed '<stack>' remains the
manual / non-deploy path for refreshing the baseline.
version: 5 adds deletionPolicy / updateReplacePolicy (pre-v6 writers)
Schema version: 5 adds two optional template-attribute fields to each
ResourceState: deletionPolicy and updateReplacePolicy. They mirror the
CloudFormation DeletionPolicy / UpdateReplacePolicy attributes that the
synth template carried at the resource's last successful create / update.
Writers emit version: 5 or later. The on-disk key layout is unchanged from
version: 2; only the per-resource shape grew. v4 readers see a version: 5
blob and fail clearly with the same "upgrade cdkd" error.
DiffCalculator (v5+) compares both attributes against the template on
every deploy / diff. A change there — typically a user removing
removalPolicy: RemovalPolicy.DESTROY from a CDK construct (CDK then emits
DeletionPolicy: Retain instead of Delete) — is now classified as
UPDATE rather than silently swallowed as No changes detected. The
attribute flip has no per-resource AWS API, so cdkd's deploy engine
refreshes the cdkd state record only — no provider call. v4 → v5
upgrade is automatic on the next cdkd deploy: state-update sites write
the current template attributes (or undefined when the template does not
carry the attribute) into the resource record, and the next deploy's
comparator has a real baseline to diff against. cdkd destroy and
cdkd state destroy honor state.deletionPolicy for the
Retain / RetainExceptOnCreate skip (the AWS resource is kept; the
cdkd state record is dropped). Both read state.deletionPolicy only,
never the template's attribute, so pre-v5 state has no signal to skip on
(redeploy under v5 to populate the field). A deploy's DELETE of a resource
removed from the template falls back to the template's DeletionPolicy
when state has no recorded value. With no recorded policy, an AWS::RDS::DBCluster or a
standalone AWS::RDS::DBInstance takes CloudFormation's default, Snapshot.
DeletionPolicy: Snapshot is honored on the same
paths: cdkd creates the final snapshot CloudFormation
promises before deleting (see the "DeletionPolicy: Snapshot" section in
cdkd destroy for the per-type mechanics and the
--skip-final-snapshot opt-out).
Upgrade note (v4 → v5) — the first
cdkd deployafter upgrading from a v0.99.x binary will classify every resource whose template carries aDeletionPolicyorUpdateReplacePolicyasUPDATEand print one↻ <logicalId> attribute update: ...line + aUpdated: N (metadata)summary entry. No AWS API call fires for any of these resources — cdkd is just recording the attribute value into its own state file so the next diff has a baseline. The deploy finishes in seconds regardless of resource count. Subsequent deploys only surfaceUPDATEfor resources whose template attribute actually changed.
version: 6 adds parentStack / parentLogicalId / parentRegion (v6+ writers)
Schema version: 6 adds three optional stack-level fields to StackState:
parentStack, parentLogicalId, parentRegion. They are populated only on
nested-stack child state records — the
AWS::CloudFormation::Stack adoption. Top-level stack state
files leave all three undefined; a v6 reader treats absence as "I am a
top-level stack" (= the default semantics for every state file v1..v5
binaries wrote).
Child state files live at cdkd/{parentStack}~{parentLogicalId}/{region}/state.json
— the ~ separator avoids ambiguity with CDK Stage's /-separated
display paths. The on-disk shape is otherwise identical to v5.
Writers emit version: 6 or later. v5 readers see a version: 6 blob
and fail with the same "upgrade cdkd" error. v5 → v6 upgrade is
fully transparent — read a v5 state file with a v6 binary and the
parser tolerates the missing fields (degrades to "top-level stack");
the next write persists version: 6 silently. No cdkd state migrate-schema command, no env flag, no manual JSON edit. The
tests/integration/schema-v5-to-v6-migration/
integ test proves the round-trip against real AWS.
The fields are consumed by
NestedStackProvider:
when a parent stack contains an AWS::CloudFormation::Stack
resource, the provider runs a recursive child deploy / destroy and the
child's state file lives at
cdkd/{parentStackName}~{NestedStackLogicalId}/{region}/state.json
with the three fields populated. Top-level deploys (the common case)
leave the three fields undefined on every write — the v6 reader treats
absence as "I am a top-level stack" and degrades cleanly.
cdkd import --migrate-from-cloudformation recursively adopts existing
CFn-managed nested-stack hierarchies — each nested child gets its own v6-keyed state file with all three
parent-link fields populated, and the source CFn stacks are retired via a
single parent-side DeleteStack cascade after recursive DeletionPolicy: Retain
injection. cdkd export of a cdkd-managed nested stack back into
CloudFormation is supported as well — the orchestrator submits one IMPORT changeset per cdkd-managed
stack in leaf-first order, non-leaf parents adopt their just-imported
children via the AWS-docs "Nest an existing stack" pattern, and cdkd
state for every stack in the tree is deleted leaf-first after the
CFn-side IMPORT loop completes. Fresh cdkd deploy of new nested
stacks is supported too.
version: 7 adds provisionedBy (v7+ writers)
Schema version: 7 adds an optional provisionedBy field to each
ResourceState — Provisioning Layers covers the same
routing from the user's side. The value is 'sdk'
(cdkd's preferred fast path — direct synchronous AWS SDK calls) or 'cc-api'
(the Cloud Control API fallback), i.e. which provisioning layer owns the
resource. A Custom Resource is recorded 'sdk' too, so the field is always
populated on a v7+ write; it has no SDK-vs-Cloud-Control dichotomy of its own,
so read the value there as "not Cloud Control" rather than as a literal claim
about synchronous SDK calls.
Pre-v7 every resource was implicitly SDK-managed, so a v7 reader treats the
absent field on a
v6-and-earlier record as the legacy SDK default. Precisely, an absent field
means the record is not PINNED: routing re-decides from scratch, so such a
resource can still be auto-routed to Cloud Control by the silent-drop
check — the same decision it got before v7 existed. Only a recorded
'cc-api' pins. v7+ writers (cdkd deploy and cdkd import alike) emit the
field explicitly so the decision is durable across deploys.
The field is sticky by default: once a resource is 'cc-api', a later
SDK-provider backfill does not by itself migrate it back, because doing that
unconditionally would mean physical-ID churn (destroy + recreate) on every
backfill release. The stickiness has narrow exemptions, listed in
STICKY_CC_MIGRATION_EXEMPT in
src/provisioning/provider-registry.ts —
consult the constant rather than a list here, since its membership changes.
Every entry must satisfy the same hard requirement: the SDK provider
addresses the resource by the SAME physicalId the Cloud Control path stored,
so the re-route costs no churn and the record flips to 'sdk' transparently on
its next write. That parity is a per-type empirical fact, not an argument, so
each entry names an integration fixture that OBSERVED it on a live resource;
a unit test refuses an entry whose fixture does not exist or has never run.
Entries then differ in WHY they were admitted, and the difference decides how conditional the escape is:
'cc-broken'— Cloud Control cannot correctly manage the type at all, so staying pinned keeps a live bug alive. The escape is unconditional.AWS::Scheduler::Schedule(a schedule in a customScheduleGroupis unaddressable via Cloud Control),AWS::RDS::DBProxyTargetGroup(the read and delete handlers cannot derive the proxy name from the TargetGroupArn),AWS::Lambda::EventInvokeConfig(every Cloud Control update fails validation) andAWS::Pipes::Pipe(a Cloud Control UPDATE cannot change a stream or broker source's write-onlySourceParameters) are the members today.'sdk-coverage'— Cloud Control manages the type, apart from at most a defect that only an update reaches, and is slower; cdkd has since gained full property coverage. The escape is conditional on this resource: it happens only on a mutating deploy where neither the template's property bag nor the recorded one carries a property cdkd would silently drop. Reading the recorded bag too is what keeps a removal deploy correct — a property applied under Cloud Control and since deleted from the template still needs Cloud Control to unset it, so the flip waits one deploy.AWS::SNS::TopicandAWS::ElasticLoadBalancingV2::Listenerare the members today. The listener has one more reason to leave: Cloud Control leaves aListenerAttributeskey removed from the template at its old value, and the SDK provider resets it, on the same deploy that flips the record.
When a 'sdk-coverage' flip is about to happen, cdkd diff annotates the
resource [returning to SDK provider], and --pin-cc-api <LogicalId> declines it for that deploy. A 'cc-broken' entry ignores the
pin — honoring it would re-pin the resource to a handler that cannot manage it.
So a provisionedBy: 'cc-api' record is NOT proof the resource will keep being
managed through Cloud Control.
cdkd destroy reads the field to pick the delete path, cdkd drift to pick
readCurrentState, and cdkd state show displays it
(ProvisionedBy: sdk | cc-api | (sdk, legacy default)).
v6 → v7 upgrade is fully transparent — a v6 state file read by a v7 binary
parses with the field undefined, and the next write persists version: 7
silently. No command, no flag, no manual JSON edit. The
tests/integration/schema-v6-to-v7-migration/
integ test proves the round-trip against real AWS.
version: 8 adds outputReads
Schema version: 8 adds an optional stack-level outputReads array — one
StateOutputReadEntry per Fn::GetStackOutput resolution that was served from
a cdkd state record during the consumer stack's deploy. Two resolutions are
deliberately NOT recorded, so the array is a subset of the references a
template carries rather than an inventory of them: a cross-account
(RoleArn) read (deferred to a future bump alongside a sourceAccountId
field), and one served by the CloudFormation fallback
(the producer is not
cdkd-managed, so cdkd never recreates it and there is no warning to attach the
consumer to). Same-account cross-REGION reads ARE recorded (sourceRegion
carries the producer's region). It is the sibling of v4's
imports, with one deliberate difference: outputReads is informational
only. There is no destroy-time refusal for Fn::GetStackOutput, because that
intrinsic is a weak reference by design — the producer stays deletable
independently of its consumers. The entries are consumed by
findDownstreamConsumers to name affected downstream stacks in the
--recreate-via-cc-api / --recreate-via-sdk-provider warn block.
The field is omitted from the JSON when the recorded set is empty, so the
on-the-wire shape is byte-identical to v7 for a stack whose references were all
of the two unrecorded kinds above — and for one that uses no
Fn::GetStackOutput at all.
v7 → v8 upgrade is fully transparent — outputReads === undefined on a
pre-v8 record reads as "no Fn::GetStackOutput consumers known" and the
enumeration degrades to imports-only (the v4-shipped behavior); the next deploy
under a v8 binary repopulates the field and persists version: 8 silently. The
tests/integration/schema-v7-to-v8-migration/
integ test proves the round-trip against real AWS.
version: 9 adds exportNames
Schema version: 9 adds a stack-level exportNames array: the keys of
outputs that are Export.Name aliases, i.e. the ONLY names an
Fn::ImportValue may bind to. The outputs bag has
always held plain Output names and export aliases side by side, and nothing in
the record said which was which — so the exports index (on update and on
rebuild) and the resolver's state.json scan treated EVERY key as an export. A
plain CfnOutput('VpcId') in an unrelated stack was indexed as the producer of
export VpcId, last writer wins, and a consumer's Fn::ImportValue: VpcId
resolved to whichever stack deployed most recently — silently, and to a value
CloudFormation would never hand out (its export namespace is separate from its
output names, and it refuses a second producer of one name). All four readers
— those three plus the cdkd local commands' --from-state Fn::ImportValue
fallback scan — now go through one predicate (importableOutputKeys in
src/types/state.ts).
Two shapes of the field mean two different things, so unlike imports /
outputReads an EMPTY array is written, not omitted: [] means the stack is
known to export nothing (its plain outputs are not importable), while an
ABSENT field means the set is not known and the record keeps the legacy
"every key is importable" rule. Absent is what a pre-v9 record carries, and
also what a v9 failure-path save writes when it carries a pre-v9 bag forward
unchanged — the set travels with the bag it describes, and cdkd never invents
[] for a bag it did not re-resolve.
Two stacks that both EXPORT one name keep the index's latest-writer policy but now produce a warning on the producer's deploy (and on an index rebuild); CloudFormation refuses the second producer outright, so rename one of them. During the upgrade window that warning can also name a stack that merely holds a same-named PLAIN output under a pre-v9 record — that is its stale entry from before the field existed, and the stack's next deploy clears it.
v8 → v9 upgrade is fully transparent — exportNames === undefined on a
pre-v9 record reads as "every output key is importable" (the v8-shipped
behavior), so no existing cross-stack reference breaks; the next deploy of the
producer under a v9 binary writes the set and persists version: 9 silently.
That includes a deploy with NO template change: the no-change path persists the
set and re-feeds the exports index with the exports only whenever the effective
export set changed while the outputs values did not, so a producer whose
template never changes still stops publishing its plain output names after one
deploy. The same path also handles a self-named export toggled on a v9
record — adding or removing Export.Name equal to an output's own key rewrites
the same key with the same value (byte-equal bag), and the effective-set
comparison is what persists and re-indexes the flip so a newly-exported name
becomes importable (and a newly-unexported one stops being served). The
tests/integration/schema-v8-to-v9-migration/
integ test proves the round-trip against real AWS — and first reproduces the
shadowing under the v8 binary (a consumer bound to a decoy stack's plain
output) before the v9 binary rebinds it to the real export.
A hand-edited exportNames that is not an array reads as an empty set, not
as an absent one. The distinction matters because absent means "not known" and
falls back to the legacy every-key rule — so reading a corrupt field that way
would republish every plain output name as an export, which is exactly the
shadowing v9 exists to close. A string, a number, an object or null there is
therefore read as "this stack exports nothing", and any element that is not a
string is dropped — key lookup coerces rather than throwing, so
exportNames: [0] against a bag holding a "0" key would otherwise publish
it. A healthy string[] is unaffected.
version: 10 adds observedBaselineRefused
Schema version: 10 adds a per-resource observedBaselineRefused flag, set
by cdkd import when it DECLINES to capture an observedProperties baseline
for a resource.
cdkd import refuses that capture when it cannot vouch that the resource's
recorded properties still spell the dynamic reference the template had — a
resolve that threw, one that lost a {{resolve: opener, or one that discarded
a subtree it could not prove inert. The redaction protecting a captured
readback is POSITION-based, so an unvouched bag gives it no evidence and the
decrypted value would be written to state.json in the clear.
Before v10 the refusal left only observedProperties: undefined behind — the
same thing a pre-v3 record and a provider without readCurrentState leave —
so the commands whose job is to FILL a missing baseline could not tell a
refusal from a resource that simply never had one, and filled it anyway:
| command | what it did |
|---|---|
cdkd deploy (the deploy-start baseline auto-refresh) |
captured a readback positioned against the untrusted properties |
cdkd state refresh-observed |
the same, for every resource in the stack |
cdkd drift --accept |
wrote the readback INTO properties |
cdkd drift --revert |
pushed properties to AWS, which can overwrite a live secret with a placeholder the stack never deployed |
cdkd drift (detection) |
compared AWS against those properties and printed the live value — a decrypted secret among them — with nothing able to mask it |
cdkd import (a later run) |
re-captured a baseline for a resource a SELECTIVE import left in place — its recorded properties are still the ones an earlier run refused |
All six now decline a refused resource. cdkd drift reports it under
notCompared with the cause baselineRefused — it does not read the resource
back from AWS at all, so the live value never enters the report. The three you
invoke to act on that resource — cdkd state refresh-observed, cdkd drift --accept and --revert — say so at normal verbosity; the deploy-start refresh
and the import skip report only under --verbose, since neither is a command
you ran to refresh that resource in the first place.
cdkd state show renders an ObservedBaseline: REFUSED ... line for one, cdkd state refresh-observed
reports them in their own tally rather than as unsupported, and cdkd export
lists them apart from the resources a refresh really can help.
How to clear it (every refusal but the one class below): deploy a change to the resource. A CREATE, UPDATE or
replacement rebuilds its state record from the template — the evidence the
import lacked — and captures a trustworthy baseline. A NO_CHANGE deploy does
NOT clear it, and neither does re-running cdkd state refresh-observed. Until
then cdkd drift compares that resource against its recorded properties, as
it did for any resource without a baseline before schema v3.
One refusal class is NOT cleared by an in-place update, and the record says
which: observedBaselineRefusalReason: "unverifiable-parameter" (an optional
field beside the marker, no version bump — for a record written before it
existed, see the next paragraph). cdkd import
writes it when the resource depends on a template parameter whose deployed value
it could not prove equal to the Default it bound. cdkd deploy binds that same
Default, so an update that does not rewrite the parameter-bound property
leaves the deployed value in AWS, and a baseline captured against the
placeholder would record it. For that reason an in-place UPDATE keeps both
fields and takes no readback; a replacement or a fresh CREATE clears them, and
so does a cdkd import that re-imports the resource while a CloudFormation
stack proves the parameter (a re-import with no such stack carries them forward
on an unchanged physical id). The reason is never present without the marker.
A cdkd binary that writes version: 10 but predates the reason field ignores
it and clears the marker on any UPDATE, and 0.290.36 — which has the field —
still clears a marker recorded WITHOUT one (next paragraph). So do not deploy
such a stack with any cdkd older than the one that reads a reason-less marker
fail closed.
Every refusal cdkd writes now carries a reason: "incomplete-resolution" is the
other class (the import-time resolve threw, lost a {{resolve: reference, or
discarded part of the properties), which an UPDATE clears as described above.
A marker with NO reason was written by an older cdkd and does not say
which class it is — 0.290.35 wrote unverifiable-parameter refusals that way — so
it is read fail closed: when the resource's definition in the template at hand
reads a declared template parameter (or cdkd cannot tell: a template it cannot
read, or an intrinsic it does not know in a template where something reads a
declared parameter), cdkd deploy stamps it "unverifiable-parameter" at the start of the deploy and
cdkd import carries it as one; otherwise an UPDATE clears it. See
import.md for the cost and the one known gap.
Migration is transparent in both directions a user can observe: a version: 9 record has the field absent, absence means "not refused", and that is
exactly how those commands behaved before the field existed. The next write
persists version: 10 silently.
As with every bump since v2, an OLDER cdkd binary refuses a version: 10 blob
with an explicit "upgrade cdkd" error. That refusal is the point here rather
than a side effect — a binary that did not know the flag would ignore it and
refill the refused baseline — but it is a trade, not a free win, and it is
worth knowing before you upgrade one machine in a fleet: cdkd stamps the
current schema version on every state file it writes, so once a v10 binary has
deployed a stack, every older binary fails on that stack, whether or not it
holds a refused record. Upgrade the whole fleet together.
version: 11 stores NoEcho values as *** (current writers)
Schema version: 11 keeps the value of a NoEcho: true template parameter
out of everything cdkd writes down. Before it, cdkd masked such a value in its
log output only, and state.json held it in the clear wherever a resource or
an output used it.
Where a NoEcho parameter supplied a value, cdkd now stores ***:
| surface | what is stored |
|---|---|
properties, observedProperties |
*** at every position the parameter fills, whatever the value's type or length (a value embedded in a longer string from 4 characters) |
attributes |
*** for an attribute DECLARED NoEcho, whatever its type or length, including one that echoes the value back (an AWS::SSM::Parameter's Value); another attribute holding the value is masked from 4 characters |
outputs, the exports index |
*** for an output whose value reads the parameter |
rollback-journal.json |
the same masks, for the records and outputs it saves |
A string that only CONTAINS the value is stored as *** whole, as for a
NoEcho custom-resource response (see
NoEcho custom-resource responses).
Two optional per-resource fields record what was masked, and a third how AWS reports it; none holds any part of a value:
| field | meaning |
|---|---|
noEchoLeaves |
the positions in properties (and so in observedProperties) stored as *** because a NoEcho parameter, or an attribute declared NoEcho, supplied them; each position is a list of keys and array indexes |
noEchoAttributeNames |
the record's own attributes its provider declared NoEcho: every attribute of a custom resource that answered NoEcho: true, a nested stack's masked outputs, and an attribute that echoes a NoEcho value. Each is stored as *** whatever its type or length |
noEchoExactEchoLeaves |
the noEchoLeaves positions, under a create-only property, at which AWS was seen to report exactly what cdkd sent (no version bump; see the table below) |
An ABSENT field means "not known", which is every record an older cdkd wrote.
How a deploy compares a masked value
AWS still receives the real value: cdkd resolves the parameter on every deploy
and sends it. What changes is how a deploy decides whether to send it, since
the record only says ***. Every resource that reads a NoEcho parameter is
re-resolved on every deploy, and:
| the property | what the deploy does |
|---|---|
| can be updated, and AWS reports it back | reads the resource back; an unchanged value is skipped, a changed one updated |
can be updated, but AWS does not report it (write-only, such as an RDS MasterUserPassword, or a type cdkd cannot read back) |
sends it on every deploy, with one info line per resource saying why |
| cannot change without a replacement (create-only), and AWS was seen to report it exactly | reads the resource back; an unchanged value is skipped, a changed one REPLACES the resource, as before version: 11 |
| create-only otherwise, including a write-only one or one whose type schema cdkd could not look up | never replaced: when the readback cannot confirm the value, every deploy warns, and --recreate-via-cc-api / --recreate-via-sdk-provider is how to apply a new value |
| create-only, and the readback FAILED | the resource fails with a message to re-run; it is never replaced on a failed read |
"Seen to report it exactly" is recorded per position in noEchoExactEchoLeaves:
a readback handed the
record, which holds *** there, reported exactly the string cdkd sent. It
describes how the provider reports the property, never the value. The deploy
that creates the resource, or replaces it on a create-only change, reads it
back once to set it, and so does the migration deploy below; any later
readback that holds the value sets it too. A replacement a provider's update
falls back to takes no such readback: the new resource starts without it until
a later readback holds the value. A readback that differs never clears it, and one that fails or cannot
report the property leaves it unset. Only a whole string value under a
create-only property is eligible: a list, a number, or a value inside a list is
never trusted, since a provider may reorder or retype what it reports. Without
it, a provider that normalizes what it reports (lowercases a name, reorders a
list) would read differs on an unchanged value and replace the resource on
every deploy, so the warning says the provider is not known to report the
property exactly.
The replacement goes through the same create-first path, stateful-resource
guard (--force-stateful-recreation) and name-collision checks as any other.
The approval prompt before the deploy sees no replacement in a diff that
cannot read AWS, so under --require-approval=destructive or any-change the
deploy asks again when it reaches such a replacement (--yes approves it). A
"no", a terminal that cannot be asked, or a resource deadline that already
expired keeps the resource and warns, and
--recreate-via-* applies the value.
One case the flag cannot catch: a provider whose readback right after a create reports exactly what was sent, while AWS normalizes the value later. Each rotation-free deploy then reads a different value and replaces the resource. cdkd stores nothing derived from the value, so it cannot tell this from a change; the stateful-resource guard still stops a stateful type.
A stack whose resources read a NoEcho parameter, with nothing else changed,
is reported as No changes: those resources are compared as above and do not
count as changes, or as destructive changes for the approval prompt.
A provider can receive *** as the PREVIOUS value of such a property on an
update whose readback could not confirm it, since the record holds nothing
else. A readback that confirms the value sends nothing.
Custom resources that read a NoEcho parameter
cdkd cannot read a custom resource back, so its handler is the "AWS does not report it" row above:
| request | what the handler receives |
|---|---|
Update |
sent on every deploy: ResourceProperties holds the real value, OldResourceProperties holds *** at that position. A ServiceToken fed this way is not compared with the record, so a changed value is not refused (a changed custom-resource ServiceToken) |
Delete |
cdkd destroy with the app: ResourceProperties holds the real value, re-resolved from today's template and parameters. cdkd deploy --recreate-via-cc-api <logicalId> re-resolves it the same way: it is the one deploy route that replaces a custom resource still in the template (deleting through its recorded handler and creating through the template's, on the custom-resource provider, so the record stays provisionedBy: sdk). No other deploy does (a changed ServiceToken is refused, and a Type change keeps the skip). Otherwise not sent: the delete is skipped, as for any resource whose DELETE needs a NoEcho-filled property (below) |
On a Delete, cdkd re-resolves a position only while today's template still
reads a NoEcho parameter there, for a resource of the same type, and, where
the record holds the hashes the last deploy took of that property's template
text and of its resolved non-secret inputs (maskedPropertyFingerprints,
maskedPropertyInputFingerprints), only while today's are the same: a changed
expression, condition, list element, or Default of a non-NoEcho input is
refused (a NoEcho parameter's own Default is not hashed, so a change to it
is accepted and the handler gets today's value). A hash the record holds but
cannot compare (a refused one, an input unknown today) refuses too; a property
the record never hashed (an older cdkd) is accepted. On cdkd destroy an input
that reads a resource is unknown, so a masked property that also reads one
elsewhere (Config: {Token: {Ref: P}, Bucket: {Ref: MyBucket}}) is refused
when the record hashed its inputs.
On cdkd destroy the expression must be built from parameters, the pseudo
parameters AWS::Region, AWS::Partition, AWS::URLSuffix, AWS::AccountId
and AWS::StackName, and literals; a nested stack's child gets the value its
parent's row hands it. The handler receives the value bound TODAY, as a deploy
of that template would send it: if the parameter's value changed since the
last deploy, it is the new one. The delete stays skipped when:
- the command holds no template:
cdkd state destroy, a deploy that removed the resource, and a rollback; - the position read an attribute a custom resource or nested stack declared
NoEcho, and the template still does: there is no template value to re-resolve; - the template no longer reads a
NoEchoparameter there, that property's text or resolved inputs changed since the last deploy, the resource changed type, the template carries aTransform(destroy does not expand macros), or it was synthesized for another region; - on
cdkd destroy, the expression reads anything else (a resource, a condition, another stack, a dynamic reference), a parameter cannot be bound, or a nested child reads a row parameter its parent could not re-resolve; - the record holds
***at a position itsnoEchoLeavesdoes not name (one embedded throughFn::Join, or a record an earlier cdkd wrote): nothing names what it stood for. Acdkd deployof the app (or, for a record with nonoEchoLeaves, acdkd scrubof the stack) first records theNoEchopositions, after whichcdkd destroysends the delete.
The value goes into the handler's request only; the record keeps ***. The
warnings, errors and handler log lines cdkd prints are masked like a create's:
a value shorter than 4 characters embedded in a longer line, or one the
handler re-encodes (JSON-escaped quotes or backslashes), can still show.
A handler whose Update is not idempotent, or that compares the two bags to
decide what to do, should take a name or an ARN (for example of a Secrets
Manager secret) and read the value itself, rather than the value, which is the
same pattern cdkd already requires for a secure dynamic reference.
cdkd diff cannot read AWS, so it compares the masks and says once per stack
how many unchanged resources read a NoEcho parameter; that note does not
count as a change for --fail.
What else changes
- A nested stack's child treats each parameter its parent fills from a
NoEchosource (aNoEchoparameter, or an attribute declaredNoEcho) as aNoEchoparameter, whatever the child template declares: its records, outputs andcdkd diffmask and compare that value as above. - A stack that reads another stack's output served by a
NoEchoparameter gets the value only within ONEcdkd deployrun that also deploys the producer; a separate run reads***and is refused, as for a custom-resourceNoEchooutput. Deploy producer and consumer together (cdkd deploy --all). - An
Export.Nameholding aNoEchovalue is not published (the deploy warns), so no consumer can bind to it. - A resource that reads an attribute a custom resource or a nested stack
declared
NoEcho, out of a record an EARLIER run wrote, is refused with the attribute named; cdkd does not re-run the producer to recover the value. Change the producer in the same deploy (for a custom resource, change one of its properties so its handler runs again). - An attribute that echoes a
NoEchovalue (such as an SSM parameter'sValue) is not refused: when the deploy gave the producer that value, its readers are served from an AWS readback of the producer, and the value is never written to state. cdkd driftreports a masked position in its own group, without printing either side and without affecting the exit code.cdkd drift --accept/--revertleave such a position alone: the baseline either one rebuilds holds***there, and--revertkeeps AWS's value rather than sending the mask (see Redacted (NoEcho) baselines). A custom-resource mask is still refused by--accept;--revertrefuses it only when it cannot tell which live value belongs there.cdkd exportexports a record whose only masks sit atNoEchopositions: the exported template reads the parameter.- A rollback revert reads a masked position back from AWS and sends the value
AWS holds there, and refuses when it cannot read it (see
cdkd rollback). - The observed baseline (captured at the start of a
cdkd deploy, or bycdkd state refresh-observed) is masked at every position where the record holds***, named innoEchoLeavesor not. A leaf that is***for another reason (a custom resource'sNoEchovalue) is masked there too, socdkd driftcan report it as drift until a deploy updates or replaces that resource. This errs toward hiding a value, never toward storing one. cdkd importandcdkd scrubstore***at every position today's template fills from aNoEchoparameter, whatever the value's type or length, and writenoEchoLeavesfor it, as a deploy does.cdkd scrubalso masks a plaintext the record still holds there (an olderDefault) wherever else the same record holds it, and a declared output the parameter serves. A value a nested child received from its parent's row is masked by a deploy and bycdkd scrub(positioned by the row; seecdkd scrub);cdkd importleaves it to a deploy.- A resource's physical id is never masked, whether it embeds a
NoEchovalue or IS one (a name-identified resource, such as an RDS parameter group named by the parameter). ANoEchovalue used as a NAME is published by AWS, and the deploy warns once per such resource. Other resources' resolved copies of that name, and the outputs, exports index and rollback journal that carry it, hold it in the clear too; see A resource named after a secret keeps that name in state. - A resource whose DELETE needs a property a
NoEchoparameter fills (a name, a policy target, or any property of a custom resource, whose handler would receive***) cannot be addressed from its record, which holds***(except a custom resource's, re-resolved as described in Custom resources that read aNoEchoparameter).cdkd destroy, and a deploy that removes the resource, skip that delete, keep the record and exit non-zero (a deploy exits zero with--allow-unaddressed); delete the resource by hand (for a custom resource, whatever its handler manages), then drop the record withcdkd state orphan. When a REPLACEMENT creates the new resource first, the delete of the old one is skipped with a warning that it is no longer tracked, and it is left in AWS; the resource's row is reported as a partial update, and the deploy exits non-zero for it unless--allow-unaddressed. - A
Numberparameter whose resource reports the value back as its STRING spelling ("5432"for5432) never reads as held: an updatable property is re-sent on every deploy, and a create-only one is warned about on every deploy, with no replacement.
What stays in plain text
- A value shorter than 4 characters, or a number, that reaches state other
than at a position the template names: embedded in a longer string read
through a declared attribute, or reaching a nested stack's child other than
through a row parameter the parent fills from a
NoEchosource (such a parameter is positioned like aNoEchoone, a list's elements included). A value of 4 or more characters is stored as***wherever it lands. - A record the template no longer names (as the same logical id and type), such as a resource being deleted, an orphan record, or the previous copy of a resource whose type changed, which the rollback journal saves: the positions come from today's template.
- Outputs saved after a failed outputs pass, and an output the template changed since an older cdkd wrote it.
- An output of fewer than 4 characters, or a number, that a
NoEchoparameter serves in another stack and that a consumer imports (Fn::ImportValue/Fn::GetStackOutput) in the samecdkd deploy: the consumer's record holds it in the clear.
Migration
Nothing to do. A version: 10 record is read unchanged, and the first
cdkd deploy after the upgrade migrates the stack: each unchanged value is
compared with the plaintext the record still holds, which is exactly what cdkd
last sent, so the migration deploy neither updates nor replaces a resource for
it, and its save stores *** and the new fields. A record that deploy did not
reach is masked by the template's positions too, while the template still
names it as the same logical id and type.
A value that DID change since the last deploy is applied by that first deploy
as before: a replacement where a change to the property replaces the resource,
an in-place update where cdkd updates that property in place. A replacement's
warning names the cause, never the value: "a NoEcho parameter's value changed
since the last deploy" where the property is the parameter itself, otherwise
"the value at its NoEcho position changed since the last deploy". The recorded
plaintext is exact evidence. The same deploy reads each create-only resource
it reaches back once, handed a copy of the record with the plaintext already
masked, to record whether AWS reports the value exactly (above). Later deploys
compare against ***, and replace a create-only property on a readback only
where that was recorded.
Rotate any NoEcho value a stack ever held in the clear. Earlier object
versions of state.json written before the upgrade still contain it, and cdkd
does not purge them: they are the state's recovery path. Treat the value as
exposed to anyone who can read the state bucket's object versions, as
cdkd scrub advises for a secret.
As with every bump, an OLDER cdkd binary refuses a version: 11 blob with the
"Upgrade cdkd" error, so upgrade every machine that deploys the stack together.
A v11 binary stamps version: 11 on every state file it writes, including by
commands that hold no template (cdkd state refresh-observed, cdkd drift --accept, cdkd orphan). They do not position a record that has no
noEchoLeaves. cdkd state refresh-observed masks the baseline it captures
wherever the record names a position, and wherever the record's own property
is already ***.
version: 11 alone therefore does not mean a stack's values are masked, and
only a cdkd deploy migrates it.
Fields added without a version bump
skippedOutputs (informational, no version bump)
An Output the deploy could NOT resolve is SKIPPED — warned about when the
resolver threw, silently when the resolver returned nothing (both under the
default arm; --strict-getatt aborts the deploy instead) — and cdkd deploy
stores nothing for it, so a bag the deploy re-resolved lacks the key (a
no-change deploy keeps a failed output's stored value, so a key that resolved
on an earlier deploy can keep that value beside a record — the diff then
ignores the record for it). When that
failure belongs to a secret reference — a {{resolve:secretsmanager:...}}
naming a JSON key the secret does not hold, or a reference assembled from
another secret's value, refused before its lookup — cdkd diff cannot
reproduce it: the diff resolves
outputs with secret references left as their tokens, so the value assembles
cleanly, and the diff used to preview an ADD the deploy would never perform
on every run of the unchanged stack, keeping cdkd diff --fail red.
skippedOutputs is the deploy telling the diff what it learned: each skipped
Outputs key mapped to a sha256 over the template inputs its resolution reads
— the output's own entry (Value, Export, Condition) and every top-level
section except Resources and the sibling Outputs (Parameters,
Conditions, Mappings, ...), digested from the template as handed in,
before any parameter binding or condition evaluation, on both sides. One value
is deliberately excluded: a NoEcho: true parameter's Default is hashed as a
constant, so this field cannot become a confirm oracle for a low-entropy one.
It is not a claim that the value is otherwise absent from state. Everything
else about such a parameter is still hashed, so only a change to that default
alone fails to un-bind. The diff
previews a recorded key as absent — no row, exactly what the next deploy
will leave in state, and the sibling outputs are previewed normally, so a
genuine change beside the broken output still renders and --fail still
exits 1 for it — only while the key is still absent from outputs AND today's
digest equals the recorded one. Any change to those inputs (the Value
repaired, an Export.Name added or removed, a parameter default, a condition,
a mapping) puts it back under the ordinary preview rules — usually an ADD
row; an intrinsic Export.Name the diff still cannot resolve keeps omitting
the section, as it did before this field — and the next deploy re-decides it:
it publishes the output if the repair took, or records it again under the new
digest.
Repairing the RESOURCE an output reads is handled separately, because
Resources is deliberately not digested (hashing it would discard the record
on every unrelated resource edit). cdkd diff declines to use the record for
an output whose Value or Export.Name references a logical id this run's
resource diff reports as changing: the deploy that follows re-resolves every
output, so the record cannot speak for it. The test is the reference, not
whether the edit could actually repair the output — undecidable from a
template — so an unrelated edit to a referenced resource also stops the record
binding. Bounded on purpose: that diff already reports the resource's own
change.
Lifecycle: written by every deploy that re-resolves outputs (the changed and
the no-change path alike; the no-change path saves on a record change alone,
and writes THIS pass's record beside the outputs that did resolve, while a
failed output keeps its earlier value — or beside the whole previous outputs
when they cannot be merged safely: a failed output with an earlier value whose
Export.Name is an intrinsic, or a save that would put the first secret
reference beside a kept value, checked on the outputs as they will be saved; a
kept value is not repositioned onto a reference from today's template, though
the ordinary secret scan still redacts it), omitted when nothing was skipped, cleared for a key that
resolves or leaves the template, and carried forward unchanged by the saves
that carry the outputs bag forward without re-resolving it (a failed
deploy's partial saves, and the snapshot a partial cdkd destroy leaves —
every resource it removed returns as a CREATE on the next diff, which un-binds
any record that references it). Every command that rebuilds state OUTSIDE a
deploy DROPS it instead — cdkd import, cdkd drift --accept,
cdkd drift --revert, cdkd rollback, cdkd scrub, cdkd orphan and
cdkd state refresh-observed. All but the last can change the values an
output's resolution reads while every resource still reports NO_CHANGE, so
the diff has nothing to un-bind on; refresh-observed writes only
observedProperties, which the resolver does not read, and drops anyway on
any run that refreshed at least one resource (a run that refreshed nothing
keeps the record). The rule is deliberately flat rather than per-command:
of three attempts to argue a particular command safe, two were shown wrong and
the third could not be settled either way. Those keys return to pre-record
behaviour until the next deploy.
Upgrade is transparent — the field is absent on a record written before
it existed, the diff then behaves as before, and the next no-change deploy of
the affected stack writes it. Limitation: five repairs are invisible to the
digest, so the record stays (and the diff stays silent about that output)
until the next deploy re-resolves it and clears the entry. Two are outside the
template and outside anything the diff looks up: the secret gained the key, or
the SSM parameter was created. One is a nested stack's input VALUE changing on
the parent's side, which the diff does resolve but the digest does not hash,
since hashing supplied values would tie the record to a caller's arguments
rather than to the template. One is cdkd itself being upgraded so that a
provider now builds an attribute an output reads. The fifth is deliberate: a
NoEcho parameter's default is masked out of the digest (above), so a change
to that default alone does not un-bind. Where the deploy warns, that
warn names the broken output and cdkd diff is not a second signal for it —
but the quiet arm, a resolver returning nothing for an attribute it cannot
construct, emits none, so for that one there is no signal at all until the next
deploy re-resolves the output.
conditionVerdicts (no version bump)
A condition that reads a parameter fed a secret {{resolve:...}} reference
has no verdict at plan time: cdkd diff never resolves a secret. The deploy
evaluates it against the real value and records, for each such condition the
diff reads (by an Fn::If, or a resource's or output's Condition), its
verdict and a fingerprint: sha256: over the condition's definition,
every condition it names, and the inputs of the parameters they read. A
secret-fed parameter contributes its {{resolve:...}} reference, never its
value, and a value carrying a secret is never an input, except a value equal
to the parameter's Default, which is template text. cdkd diff reuses a
verdict only when the fingerprint it recomputes is equal; otherwise it takes
the condition's FALSE branch, as before
(cdkd diff).
- Written by a deploy's final save and by its no-change save (the latter also when only the record changed). Only a nested child records anything, since only a child receives a secret-fed parameter. A parent skips an unchanged nested-stack row, so a child last deployed by an older cdkd gains its record on the next deploy that changes it.
- Dropped by every other save: the per-resource, rollback and output-failure saves rebuild state without it. So a failed or interrupted deploy leaves NO record, and the next diff falls back to the FALSE branch, never to a stale verdict.
- No version bump. An older binary ignores the field and drops it on its
next save; this binary reads its absence, or any malformed shape, as no
record. A writer that carries the loaded record forward (rollback, drift,
scrub, orphan rewrite,
state refresh-observed) may carry it, since none of them changes a definition or an input the fingerprint covers.
maskedPropertyFingerprints / maskedPropertyInputFingerprints (no version bump)
A property whose resolved value carries a secret in a form state cannot
record as a reference is persisted as ***. The common case is the EC2
UserData shape: Fn::Base64 over a script that embeds a
{{resolve:...}} reference, whose encoding decodes straight back to the
secret. *** identifies nothing, so for each such top-level property the
record also keeps two fingerprints of what the property was built from.
cdkd diff and cdkd deploy treat a property whose recorded fingerprint no
longer matches as changed, so the change is shown and sent. A secret rotated
behind an unchanged template leaves both equal and sends nothing, as
CloudFormation does.
maskedPropertyFingerprints(sha256:) hashes the property's template TEXT, so an edit to the script around the reference, or a retarget of the reference, is sent.maskedPropertyInputFingerprints(inputs-sha256:, then the text hash it belongs to after a+) hashes the template value with each NON-SECRET input replaced by what it resolved to: a parameter's value, aRef/Fn::GetAttresult (so aRefto a resource the deploy replaced), a cross-stack read, and the branch an evaluated condition selects. So a new parameter value, a replaced resource's new name and a flipped condition are sent too.
Secrets stay in their template form and never reach the hash, decided by
where an input comes from, never by comparing its value with a secret it did
not read. A {{resolve:...}} reference is hashed as the reference, a NoEcho
parameter as its Ref, and these are kept as written:
- a condition that reads a
NoEchoparameter, a reference, a cross-stack value or an attribute (as its wholeFn::If); - a
Ref/Fn::GetAttto a resource whose own definition reads a secret, aNoEchoparameter, a cross-stack value, a name the template does not declare or anFn::FindInMapover a mapping holding a reference anywhere (or one it cannot name), directly or through another resource; a condition reading such a mapping is kept whole the same way; - an
Fn::GetAtton a custom resource, whose attributes may beNoEcho(its physical id is hashed), and one whose attribute NAME is built from any of these; - a nested stack's output (
Fn::GetAtt [Child, Outputs.X], or${Child.Outputs.X}in anFn::Sub) unless it is classified clean, below. A resource whose own definition reads a nested stack's output counts as reading a secret, as before; - an input whose resolution read a secret (a
NoEchocustom resource'sData, a redacted***read), and an attribute that the save redacts because the resource it belongs to read that secret in the same deploy; - in a nested stack, a parameter value its parent passed that the PARENT
built from any of the above, and every resource and condition that reads
one. The parent classifies each expression in its stack row's
Parametersby these same rules and hands the result to the child; a value built only from non-secret inputs (aRefto a parent resource, say) enters the child's hash like any input. A passed value the parent did not classify (a rollback, which replays the child without the parent's template) is kept as written too, even when it equals theDefault; the next deploy that classifies it sees the property's hash move and sends it once. One the parent could not read this time is neither compared nor hashed, nor is a property that reads it directly or through a resource or condition. A parameter the parent does not pass binds the child'sDefault, which is template text, and is hashed.
Nested-stack outputs. A masked property that reads a nested stack's
output directly is hashed with the output's value when the output is CLEAN.
The class is read from the templates in the cloud assembly, never from the
child's state, so the deploy (before and after the child runs) and
cdkd diff decide it alike. The parent follows the output's Value through
the child's template by the rules above, without resolving anything there:
- a parameter the parent passes takes the class the parent gives that value
(the same class the child is handed); one it does not pass reads its
Default; aNoEchoparameter is a secret whatever was passed; - a
{{resolve:...}}reference, a cross-stack read, a name the child does not declare, a resource whose definition reads any of these, anFn::FindInMapover a mapping that holds a reference anywhere, and an attribute of a custom resource keep the output as written; - an
Fn::Ifis clean only when its condition reads only clean inputs and both branches are clean, since the child's verdict is not evaluated here; an output with aConditionover a secret is kept as written; - an output read from a grandchild is classified the same way one level down, with the values the child passes classified over the child's template.
An output that cannot be classified (no assembly, a missing, unreadable or
cyclic template, an undeclared output) is kept as written, as before. One
that reads a value the parent could not read this time is neither compared
nor hashed. A clean output's value still takes the checks every input takes:
a ***, a reference, or a value whose read recorded a secret is kept as
written. A child template edit that changes an output's class moves the hash
once and sends the property once; an unchanged tree sends nothing.
The same rule classifies a value a parent passes to one nested stack from another's output: a clean sibling output is now a clean passed value, so a masked property in the receiving child that reads it is sent when the output's value changes (it used to be kept as written).
So these are NOT sent through the mask: a flip of a condition over a NoEcho
parameter, and a new value of anything above. (A new value of the NoEcho
parameter itself is found by reading the resource back from AWS; see
version: 11.) A hash
that moved with such a value would let anyone holding the state file test
guesses of it. CloudFormation would update the resource; change the
property's template text, or replace the resource, to push one.
A non-NoEcho parameter value and a cross-stack output value are treated as
public: they enter the hash, so a low-entropy one may be recoverable from the
state file by guessing. Declare a sensitive parameter NoEcho. A secret the
redaction itself does not mask (one shorter than 4 characters, embedded in a
longer value) is persisted in the clear already, and is not kept out of the
hash either.
A property whose template text holds, as a literal, the value
of one of the stack's NoEcho parameters, or a value the same resource
resolved as a secret, is the exception: it gets no hash in either field and
is compared as before the fields existed. A NoEcho parameter value is known
when the deploy starts, so this holds from the first deploy. A resolved secret
is known only to a deploy that resolves the resource, so a hash the first
deploy under this version filled in for a resource it did not change is
checked by the next deploy that resolves it.
The check is a plain text match, so a short secret that also occurs as
ordinary text in the property (a word in a script) costs that property its
hash too, and edits to it are not seen through the mask.
- Written by the save of a deploy that created, updated or replaced the resource, from the template it deployed. A failed update keeps the previous record and its previous fingerprints, so the retry still sends the edit.
- No version bump. A record without the fields (an older cdkd's) is
compared exactly as before. The first deploy under a cdkd that knows the
text field fills it in, per masked property, from the template it deploys,
and saves even when nothing else changed; that deploy cannot tell an edit
made since the last deploy, so such an edit is not sent until the property
changes again. To push one anyway, change the property once more, or replace
the resource with
--recreate-via-cc-api/--recreate-via-sdk-provider. - An input fingerprint is bound to the text fingerprint it was computed with. A property with a text fingerprint but no bound input fingerprint (a record written before the input fingerprint existed, or one an older cdkd rewrote since: it updates the text field and carries the other one as it was) is compared as text, so an edit is sent as before. While the text is unchanged, the deploy fills in the input fingerprint from today's inputs and sends nothing: it cannot tell whether an input moved since (a parameter changed under the older cdkd), so such a change is sent only once the input moves again. An older cdkd reads the text field only, exactly as it always did, so after a downgrade it still sends a template edit.
cdkd diffcompares the input fingerprint for the stack it was given, classifying its nested stacks' outputs from the same assembly the deploy reads. For a nested child, and wherever it cannot bind an input (a parameter it cannot resolve), it compares the text field only: a template edit shows, an input change shows only when the deploy sends it.- Nested stacks. A nested child is deployed only when its parent row changed, which is usually because the child's own template changed. So the first deploy that reaches a child last deployed by an older cdkd is, most likely, one that carries an edit, and an edit to a masked property in that deploy is not sent. Check such a child after that deploy, and change the property once more if it did not take.
- A malformed field reads as absent. A record a rollback orphans gets them from the same save, and a writer that spreads an existing record (rollback, drift, scrub, orphan adoption) carries them.
Schema reference
StackState (state.json)
interface StackState {
version: 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 // 1 = legacy, 2 = region-prefixed, 3 = +observedProperties, 4 = +imports[], 5 = +deletionPolicy/updateReplacePolicy, 6 = +parentStack/parentLogicalId/parentRegion (nested-stack adoption), 7 = +provisionedBy on ResourceState, 8 = +outputReads[], 9 = +exportNames[], 10 = +observedBaselineRefused on ResourceState, 11 = +noEchoLeaves/noEchoAttributeNames on ResourceState
stackName: string // Stack name
region?: string // Required on version >= 2
resources: Record<string, ResourceState> // Logical ID → Resource state
outputs: Record<string, unknown> // Output name → Resolved value (NOT coerced to string)
imports?: StateImportEntry[] // v4+: Fn::ImportValue refs (strong reference — blocks the producer's destroy)
outputReads?: StateOutputReadEntry[] // v8+: Fn::GetStackOutput refs (informational — weak reference, never destroy-blocking)
exportNames?: string[] // v9+: which `outputs` keys are Export.Name aliases — the ONLY names Fn::ImportValue may bind to (undefined = pre-v9 record, every key importable until its next deploy; [] = exports nothing)
skippedOutputs?: Record<string, string> // informational, no version bump: Outputs keys the last deploy could not resolve and skipped → digest of their template inputs (issue #2740); absent = nothing skipped, or a record older than the field
orphans?: StackOrphanRecord[] // no version bump: resources a rollback retained under `Retain` and moved out of `resources`, kept so a later deploy can re-adopt one instead of colliding with the name it holds; absent = none
conditionVerdicts?: Record<string, { verdict: boolean; fingerprint: string }> // no version bump: the deployed verdict of each condition `cdkd diff` reads that depends on a secret-fed parameter, with a fingerprint of its definitions and inputs (issue #4479); absent = no record
parentStack?: string // v6+: populated on nested-stack child state records (undefined on top-level)
parentLogicalId?: string // v6+: child's AWS::CloudFormation::Stack logical id in the parent's template
parentRegion?: string // v6+: parent's region (always equals `region` until cross-region nested stacks ship)
lastModified: number // Unix timestamp (milliseconds)
}
outputs values are unknown, not string — cdkd persists whatever the
intrinsic resolver produced for the Output's Value, with no stringification
step. Most Outputs do resolve to a string, but an Fn::GetAtt that
CloudFormation defines as a LIST persists a JSON array when it is used as
the Output value directly (rather than wrapped in Fn::Join) — e.g.
AWS::Route53::HostedZone.NameServers, whose list shape the Route 53 provider
preserves end to end. Do not write code (or docs) that assumes a
state.outputs value is a string; a consumer reading one back must handle the
non-string shapes too.
The Outputs: block cdkd deploy prints is not evidence of the stored
shape. That summary renders each value with JavaScript's String(value), and
for an array that is a comma join with no brackets or spaces — a persisted
["ns-1.awsdns-00.com", "ns-2.awsdns-01.net"] prints as
ns-1.awsdns-00.com,ns-2.awsdns-01.net, indistinguishable from a genuine
comma-separated string. (An object prints as [object Object], and an
unresolved output is dropped from the block entirely rather than printed as
undefined.) To see what was actually stored, read the state file —
aws s3 cp s3://<bucket>/cdkd/{stackName}/{region}/state.json - — or run
cdkd state show '<stack>', which renders any non-scalar through
JSON.stringify and so preserves the distinction.
When resources is not an object
resources is the map of logical id to resource record, and it is unchecked in
the same way: a hand-edited or truncated record can hold a string, a list, a
number, a boolean or null there. Object.keys answers three different ways
over those, and the middle answer is the dangerous one — a [], a number or a
boolean enumerates no keys, which is indistinguishable from a stack that
genuinely has none. A string enumerates one fabricated logical id per character.
| Command | Answer |
|---|---|
cdkd deploy |
Refuses at the load (STATE_RESOURCES_MALFORMED, exit 1) — a map read as empty makes every resource the template declares plan as a CREATE, so the deploy re-provisions a stack that already exists, then saves a well-formed record over the evidence |
cdkd deploy --dry-run |
Refuses, identically — the plan a dry run prints comes from the same comparison |
cdkd destroy / cdkd state destroy |
Refuses before the prompt and before the lock (STATE_RESOURCES_MALFORMED, exit 1) — the map is the list of what to delete, so an unreadable one counted as zero resources and the run removed state.json down the empty-stack fast path |
cdkd orphan, cdkd import, cdkd rollback |
Refuse (STATE_RESOURCES_MALFORMED, exit 1) — each carries the bag into a save |
cdkd export |
Refuses the named stack at the load, before the lock, and every nested child record before any changeset (STATE_RESOURCES_MALFORMED, exit 1), under --dry-run too — a child whose map cannot be read contributes no nested stacks of its own, so migrating past it would leave that subtree out |
cdkd scrub |
Refuses on a real run (exit 2); audits and reports under --dry-run |
cdkd diff |
Repairs in memory and warns — it never writes state; on the stack you named it also reports the deploy's refusal under Blocking and exits 3; see cdkd diff |
cdkd state show |
Repairs in memory and warns; --json still emits the stored value — see cdkd state |
cdkd state resources |
Repairs in memory and warns; --json emits [], because that mode is the resource array cdkd derived rather than a view of the stored value |
cdkd local * (--from-state) |
Repairs in memory and warns — it writes no state record; every Ref / Fn::GetAtt in the run's environment that names a resource of this record resolves to nothing and is dropped, and a bare --assume-role falls back to the developer's credentials |
The destroy row is the one where repairing would be unsafe rather than
merely lossy. Read as empty, the count comes back zero, the empty-stack fast
path removes state.json with no confirmation, and the destroy reports
success having deleted nothing — every resource the record named left live
in AWS with nothing to say what they were. So reading the map as empty is not
the safe alternative here; it is that outcome. An empty {} and an unreadable []
both count zero, so the two are separated by the container's shape, never by
its size, and a legitimately empty stack still takes the fast path exactly as
before.
Refusing a cleanup command does not leave you stuck, because proceeding would
not have torn anything down either — the list of what to delete is precisely
what is unreadable. If what you want is the record gone with the live resources
left standing, that is cdkd state orphan '<stack>' --stack-region '<region>',
which the refusal names. To act on the resources instead, repair the record and
re-run.
An absent resources field is a defect, unlike an absent outputs — a
stack always has a resource map, even an empty one — and is refused the same
way. An empty {} is healthy.
When outputs is not an object
A state record is parsed as JSON and used as typed data without a
field-by-field shape check, so a hand-edited or truncated one can hold a
string, a list, a number, a boolean or null where the outputs map belongs.
Object.entries walks a string as readily as a map, so anything that rebuilds
the bag from one produces a well-formed map of fabricated keys — "abcdef"
becomes {"0":"a", …, "5":"f"}, and null becomes {}.
Which answer a command gives depends on what it would DO with the bag — write it
back, decide from it, or re-apply its value — and the answers are deliberately
opposite. Writing back is the plain case; the other two are why "can it write"
alone no longer predicts the table: cdkd destroy never rewrites the bag but
deletes the record on the strength of it, and the resolver hands a value to a
deploy that applies it to a live system:
| Command | Answer |
|---|---|
cdkd deploy |
Refuses at the load (STATE_RESOURCES_MALFORMED, exit 1) — it rebuilds the bag, saves it, and republishes the result into the shared exports index |
cdkd destroy / cdkd state destroy |
Refuses before the prompt (STATE_RESOURCES_MALFORMED, exit 1) — it reads the bag to decide whether the stack might export anything, and that decision gates the cross-stack check below |
cdkd orphan |
Refuses (STATE_RESOURCES_MALFORMED, exit 1) — it rebuilds the bag and saves the result |
cdkd import |
Refuses (STATE_RESOURCES_MALFORMED, exit 1) — it carries the bag into a save |
| A nested stack's child record | Refuses the parent's deploy AND its destroy (STATE_RESOURCES_MALFORMED, exit 1) — the parent's Outputs.<Key> attributes are rebuilt from the child's bag and persisted into the parent's record, and the parent's destroy reaches the child through runDestroyForStack, which carries the same refusal |
cdkd scrub |
Refuses on a real run (exit 2); audits and reports under --dry-run — see cdkd scrub |
cdkd diff |
Repairs in memory and warns — it never writes state; on the stack you named it also reports the deploy's refusal under Blocking and exits 3; see cdkd diff |
cdkd state show / state resources |
Repairs in memory and warns; --json still emits the stored value — see cdkd state |
cdkd local * (--from-state) |
Repairs in memory and warns — it writes no state record, so the run continues with no outputs from that record |
An Fn::GetStackOutput read of that record |
Refuses the reference — the deploy fails rather than resolving a fabricated value into the consumer's template |
| The exports index rebuild | Publishes nothing from that record, warns, and indexes every other producer — see cross-stack internals |
The destroy row is the one whose repair answer would be unsafe rather than
merely lossy. cdkd destroy refuses to delete a stack another stack still
imports from, and it decides whether to run that check by asking whether the
outputs bag holds anything. A string or a list invents one export name per
character or element; a null, a number or a boolean reads as "exports
nothing" and skips the check entirely, deleting the record while consumers
still resolve against it. Reading the bag as empty is that second answer, so
there is no repair available — only a refusal.
The Fn::GetStackOutput row is the only one where the fabricated value would
be applied rather than displayed. Object.hasOwn('abcdef', '0') is true,
so an OutputName: "0" against a six-character producer bag used to resolve
the single character a and the deploy sent it to AWS as a live resource's
property. Fn::ImportValue was never affected: it binds through the export-set
predicate, which fails closed.
Refusing is what keeps the damaged record readable. Saving over it replaces the
only signal that anything is wrong with a legitimate-looking one, permanently —
and the next deploy would republish the fabricated keys into the shared exports
index that every other stack's Fn::ImportValue resolves against.
An absent outputs field is not a defect and is never refused: a record
with no outputs is one cdkd writes on purpose (a deploy's failure-path save
emits outputs: currentState.outputs, which JSON.stringify drops when it is
undefined), and cdkd scrub round-trips such a record rather than
materializing {} over it. An empty {} is healthy too — a stack can
legitimately publish no outputs.
Each container is judged on its own: a record whose resources map is fine and
whose outputs is damaged is refused with a message naming outputs, and vice
versa.
When one resources RECORD cannot be read
The map being an object says nothing about the records in it. A record is
readable only if it is an object carrying a string resourceType — the one
field every reader touches before any other. A null, a string, a list, or an
object with no type is a row nothing can route, and every command that reaches
it answers the way it answers a damaged map:
| Command | Answer |
|---|---|
cdkd deploy |
Refuses at the load (STATE_RESOURCES_MALFORMED, exit 1), naming the rows — under --dry-run too. The same refusal runs on the pre-lock --recreate-via-cc-api / --recreate-via-sdk-provider check, which reads the record itself: without it a null named row was reported as missing from state, with advice to drop the flag for it |
cdkd destroy / cdkd state destroy |
Refuses before the prompt and before the lock, and again on the record the empty-stack path re-reads under the lock, naming the rows — a falsy row was skipped as "not found in state" and the record removed with its resource live; a row with no type was routed to a provider on no type with a physical id nothing checked; see cdkd destroy |
cdkd import |
Refuses a SELECTIVE merge over an unreadable row it does NOT re-import (STATE_RESOURCES_MALFORMED, exit 1) — it copies every such row into the record it saves. A row named by --resource / --resource-mapping is replaced from the provider's answer, so cdkd import <stack> --resource <id>=<physicalId> --force is the repair of that row and is not refused pre-flight — if that row's import then does not succeed, the import refuses before saving rather than writing the row back; a whole-stack --force import or --migrate-from-cloudformation REPLACES the map from the template and is not refused either |
cdkd orphan |
Refuses for a row on a record it would keep, under --dry-run too; a row you are orphaning is dropped as usual |
cdkd scrub |
Refuses on a real run (exit 2) — the rewrite reads each row and saves the rebuilt map; under --dry-run it DROPS the rows, warns, and reports them in the audited-record refusal — see cdkd scrub |
cdkd state refresh-observed, cdkd drift --accept / --revert |
Refuse before the lock, naming the rows — see cdkd drift |
cdkd diff, plain cdkd drift |
Drop the rows, warn, and report the rest; cdkd diff also reports the deploy's refusal under Blocking on the stack you named and exits 3 |
cdkd local * (--from-state) |
Drops the rows in memory and warns — it writes no state record. A Ref or Fn::GetAtt naming a dropped row resolves to nothing and is dropped like any other unresolvable reference, and a bare --assume-role read through one falls back to the developer's credentials. An unreadable map is read as empty the same way, with its own warning |
The rollback-orphan claim scan (cdkd deploy / cdkd diff) |
Skips the row and keeps reading the rest of that sibling's record, at debug level — a row with no readable physical id claims nothing. The scan exists to stop this stack adopting a resource another stack still owns, so it reads past a damaged row rather than stopping at it |
A row that names its type but no physicalId is readable here: it can be
routed, and what its missing id costs is reported by the command that reaches
it, in that command's own terms. Inspect the record with
cdkd state show '<stack>' --stack-region '<region>' --json and repair the row.
When a resource properties map is not an object
One level down from the two maps above, each resource record carries its own
properties map. It is unchecked for the same reason, and the consequence is
different again — this is the map the change calculation compares the template
against.
A non-object compares unequal to any declared property, so every property the
template declares reads as missing from the deployed resource. For an
ordinary property that is a spurious in-place update; for a create-only one
— an S3 BucketName, a DynamoDB TableName, anything CloudFormation lists
under createOnlyProperties — it is a replacement: cdkd deploy would
delete the live resource and create a new one, from a record nobody asked it to
act on. A string map adds one fabricated change per character on top.
| Command | Answer |
|---|---|
cdkd deploy |
Refuses at the state load, before any provider call — the observed-state refresh's reads included — (STATE_RESOURCES_MALFORMED, exit 1), naming the resource records it could not read |
cdkd deploy --dry-run |
Refuses, identically — the plan a dry run prints comes from the same comparison, so it would show the replacement as though the template asked for it |
cdkd orphan |
Refuses (STATE_RESOURCES_MALFORMED, exit 1) for a map on a record it would keep, under --dry-run and --force too — but never for one on a record you are orphaning, which it removes as usual |
cdkd export |
Refuses the named stack at the load, before any lock, and every nested child record before any child stack is planned or locked — after the root's plan (under --dry-run too), and on a real run under the root lock, which it releases (STATE_RESOURCES_MALFORMED, exit 1), under --dry-run too — the export reads the map to build import identifiers and the phase-2 pre-deletes, then deletes the record |
cdkd destroy / cdkd state destroy |
Refuses before the prompt and before the lock (STATE_RESOURCES_MALFORMED, exit 1), naming the records it could not read — a retained row included, since its Ref / Fn::GetAtt edges still order the deletes. The map is handed to each resource's delete, which reads keys off it (an RDS final snapshot, emptying an ECR repository first), and reading it as empty answers each of those keys as absent. A nested child record reached through its parent's destroy inherits it. See cdkd destroy |
cdkd diff |
Repairs those maps to empty in memory and warns, naming the same records — it writes nothing, and a preview of the rest of the stack is worth more than an abort; on the stack you named it also reports the deploy's refusal under Blocking and exits 3 |
Reading the map as empty is not the safe answer here, which is why deploy
refuses rather than repairing: an empty map declares nothing either, so it
reaches the identical replacement verdict. The refusal is the only answer that
does not act on the damage. cdkd diff can take the lossy one precisely
because it never provisions, and its warning says the preview is wrong and that
cdkd deploy will refuse on the same record.
The cdkd orphan row is the one that is scoped rather than record-wide,
and the reason is that this command is itself a way out. It rewrites every
surviving resource's reference to the resources you are orphaning and saves the
result, so a map it could not read is carried into that save untouched — the
stored value is kept verbatim rather than fabricated into a well-formed one,
but the command reports success over a record the next cdkd deploy refuses.
What it cannot do is its own job: an unreadable map hides whichever references
to the orphan it holds, so the run's audit table is incomplete and the
--force-less failure on unresolvable references cannot fire for it. Because
the save cannot persist a record it is deleting, the refusal names only the
records that would survive. Three ways out, and the order matters:
Repair the record by hand.
cdkd state show '<stack>' --stack-region '<region>' --jsonshows the stored value; fix the map and put the record back. This is the only option that keeps the resource under cdkd's management.Drop the whole record with
cdkd state orphan '<stack>' --stack-region '<region>'. It needs no CDK app and leaves every live AWS resource standing.Orphan just the damaged resource — but only while your CDK app still declares it:
cdkd orphan MyStack/TheDamagedResourceThat removes it and repairs the rest of the record in one step, leaving the live AWS resource standing like any other orphan.
cdkd orphanaddresses resources by construct path, and construct paths come from the synthesized template — so a resource your app no longer declares (a record left behind after the construct was deleted) has no path there. For that one, and without a CDK app at all, address it by logical id instead:cdkd state orphan MyStack --stack-region us-east-1 --resource TheDamagedResourceIt makes the same scoped refusals and the same rewrite as
cdkd orphan(--resource).
A legacy record (<prefix>/<stack>/state.json) that cdkd state list
shows with no region (its body names none, or could not be read) is the
exception to both commands above: cdkd state orphan '<stack>' without
--stack-region is the form that selects it, and cdkd state show cannot read
it at all, so read the object from the state bucket directly. The refusal
prints those forms for that record, and names the object's path only when the
stack name renders exactly. cdkd orphan's refusals of an unreadable
resources or outputs map name the object the same way for such a record,
instead of a cdkd state show command that cannot read it.
The refusal prints its commands at the end, each on a line of its own after a
label; copy the command after the label. For the legacy record, the object's
key and bucket are printed the same way, on Object key: and State bucket:
lines; a stack name that would not be inert with its quotes stripped shows as
'<stack>' in the key, to be filled from cdkd state list --json. The cdkd state orphan and cdkd state show
commands carry the
--profile, --state-bucket and non-default --state-prefix the run was
given, so pasting them reaches the same bucket. If the stack name would not
survive display unchanged, or would not be inert with its quotes stripped
(whitespace, or a character a shell acts on, such as ', ; or $), both
become templates with the name left as a hole; if the region would not, the
name and the region are both left as holes. The
account flags stay either way, and the message says where to take the exact
name from. A --profile, --state-bucket or --state-prefix value that would
not survive display unchanged, or would not be inert with its quotes stripped,
is itself printed as a hole ('<profile>',
'<bucket>', '<prefix>' — quoted, so a pasted hole is one literal argument
rather than a shell redirection), the message says so and tells you to fill it
from the value you passed, and the object path then names neither that bucket
nor that prefix. An empty --state-bucket counts as no bucket, so no
State bucket: line is printed; an empty --state-prefix is a real key space
and IS printed, both on the object key and as --state-prefix ''.
cdkd deploy refuses the same record until it is repaired or removed, so being
blocked in both commands is the intended state rather than an extra restriction
this row adds.
An absent properties map is a defect and is refused: every writer in cdkd
records an object there, and JSON.stringify never drops an empty one. An
empty {} is healthy — a resource can legitimately declare no properties.
A resource record that is not an object at all (a null entry, a string), or
carries no resourceType, is a different defect. cdkd orphan refuses it on a
record it would keep, scoped and with the same three ways out as above: its
save rebuilds each kept record by copying fields, so a string entry would be
saved as one key per character and a number as a record with no physical id. A
reference from another resource to such a record you are orphaning, or to one
with no physical id, is reported as unresolvable rather than substituted.
cdkd deploy refuses such a record before creating, updating or deleting
any resource (STATE_RESOURCES_MALFORMED, exit 1), under --dry-run too,
naming the records it could not read. Otherwise the entry reads as absent and
deploy plans a CREATE for a resource it already manages. cdkd diff drops
those records, warns, and previews the rest; on the stack you named it also
reports the deploy's refusal under Blocking and exits 3.
The same scoped refusal covers a kept record's attributes map — the cache
Fn::GetAtt of it is read from — when it is null or not an object; an absent
one is healthy. Under --force, cdkd orphan never reads a value out of an
orphaned record's unreadable attributes cache: the reference keeps its
original intrinsic, as it does when the cache lacks the attribute.
When orphans is not a list
orphans records resources an earlier failed deploy left live in AWS under a
Retain policy, so the next deploy can adopt them instead of re-creating them.
It is a list, it is unchecked in the same way as the maps above, and a
hand-edited or truncated record can hold a string, a number, an object or
null there. Every reader reached it through a ?? [] or a ?.length, which
admits all four.
| Command | Answer |
|---|---|
cdkd deploy |
Refuses at the load, before any resource operation (STATE_RESOURCES_MALFORMED, exit 1) — the adoption pass writes the container back, so an unreadable one would be rewritten |
cdkd destroy / cdkd state destroy |
Refuses, at its first read and again at the re-read it takes under the lock — otherwise the run deletes every resource and then the record, having never reported the orphans it could not read |
cdkd rollback |
Refuses before any replay — its own bookkeeping walks the container and saves the result |
cdkd import |
Refuses — it carries the container into the record it writes, so importing over a damaged one would leave a record every other command then refuses |
cdkd orphan |
Refuses, under --dry-run too — it carries the container into its save without reading it, so it would report success over a record the next deploy refuses |
cdkd scrub |
Refuses on a real run (exit 2); audits and reports under --dry-run |
cdkd diff |
Repairs in memory and warns — it writes nothing, names (orphans container) in the preview and lists orphans in --json's unreadableContainers, which --fail counts; on the stack you named it also reports the deploy's refusal under Blocking and exits 3 |
A string is the shape that makes this worse than a lost preview: walking it
character by character yields one garbage orphan record per character, and
cdkd rollback saved exactly that — a damaged container rewritten into a
differently damaged one, with nothing said. The other shapes read as no
orphans at all, so cdkd diff previewed no adoption and cdkd destroy
removed the record with its evidence unread.
An absent orphans container is the ordinary record, not a defect: a stack
that never had a failed deploy has no orphan list, and no command writes an
empty one over it. An empty [] is healthy too. Damage INSIDE a readable list is
a separate question, answered by the section below.
When one orphans RECORD cannot be read
The field being a list says nothing about the records in it. A record is usable
only if it is an object with a string logicalId whose state is a readable
resource entry carrying a NON-EMPTY string physicalId — including that entry's
properties and attributes maps — and no OTHER record in the list carries
that same logicalId. Each command answers a damaged one the same way it
answers a damaged container:
| Command | Answer |
|---|---|
cdkd deploy |
Refuses at the load. The adoption pass dereferences every record, so one whose state is absent or null aborts the run; one already in resources is dropped silently before state is read; and a primitive or type-less state is kept with a notice, since the provider lookup fails inside the pass's own try |
cdkd destroy / cdkd state destroy |
Refuses, at both reads. The listing that tells you which resources stop being tracked prints each record's own fields |
cdkd rollback |
Refuses before any replay. This is where the loss is worst: records MISSING a logicalId all key ONE entry of the merge map, so those collapse into one and the record saved keeps only that one (two distinct NUMERIC ids stay distinct keys) — and records SHARING one collapse the same way, the other rows' resources left live in AWS with nothing tracking them |
cdkd import |
Refuses — it carries the records into the record it writes, verbatim |
cdkd orphan |
Refuses, under --dry-run too |
cdkd scrub |
Refuses on a real run (exit 2); under --dry-run it DROPS the record, warns, and reports it in the audited-record refusal |
cdkd diff |
Drops the record, names it in the preview, in --json's unreadableOrphans and in the --fail count, and previews the rest; on the stack you named the deploy's refusal is also reported under Blocking and exits 3. It drops only what the preview cannot read — an object, a string logicalId, and a readable state with a non-empty string physicalId (the preview resolves that id against AWS and against other stacks' records) — plus EVERY record whose logicalId another record also carries, since the preview keys its adoptions by that id and would show one adoption for two resources; a record whose properties or attributes map is torn is KEPT, and the preview then WARNS naming the row — at every node the run reaches with an adoption preview; a plain run visits only the top-level stack, and a state-only child being DELETED runs no preview at all — saying that cdkd deploy refuses the record over it; the TOP-LEVEL stack also exits 3, so a clean run never precedes a deploy that will not start. A kept row that is ADOPTED additionally has its properties map repaired and named by the properties repair |
Inspect the record with cdkd state show '<stack>' --stack-region '<region>' --json
and repair the row rather than deleting the record: no command removes a single
orphans row — the per-resource commands act on resources — and the record is
the only evidence that an earlier failed deploy left its resource live in AWS.
Two records sharing a logicalId are never written by cdkd — the rollback save
merges by that id and every other save carries the list unchanged — so they come
from a hand edit or a damaged file. Each of them is named, since nothing in the
record says which is the resource the stack should re-adopt; the repair is to
keep ONE record for that id, and the other resource is then no longer tracked by
cdkd.
Example
{
"version": 11,
"stackName": "MyAppStack",
"region": "us-east-1",
"resources": {
"MyBucket": {
"physicalId": "myappstack-mybucket-abc123xyz",
"resourceType": "AWS::S3::Bucket",
"properties": {
"BucketName": "myappstack-mybucket-abc123xyz",
"VersioningConfiguration": {
"Status": "Enabled"
}
},
"attributes": {
"Arn": "arn:aws:s3:::myappstack-mybucket-abc123xyz",
"DomainName": "myappstack-mybucket-abc123xyz.s3.amazonaws.com",
"RegionalDomainName": "myappstack-mybucket-abc123xyz.s3.us-east-1.amazonaws.com"
},
"dependencies": [],
"provisionedBy": "sdk"
},
"MyFunction": {
"physicalId": "arn:aws:lambda:us-east-1:123456789012:function:MyAppStack-MyFunction",
"resourceType": "AWS::Lambda::Function",
"properties": {
"FunctionName": "MyAppStack-MyFunction",
"Runtime": "nodejs20.x",
"Handler": "index.handler",
"Code": {
"S3Bucket": "cdk-hnb659fds-assets-123456789012-us-east-1",
"S3Key": "abc123.zip"
},
"Role": "arn:aws:iam::123456789012:role/MyAppStack-MyFunctionRole"
},
"attributes": {
"Arn": "arn:aws:lambda:us-east-1:123456789012:function:MyAppStack-MyFunction"
},
"dependencies": ["MyFunctionRole", "MyBucket"],
"provisionedBy": "sdk"
}
},
"outputs": {
"BucketName": "myappstack-mybucket-abc123xyz",
"BucketArn": "arn:aws:s3:::myappstack-mybucket-abc123xyz",
"FunctionArn": "arn:aws:lambda:us-east-1:123456789012:function:MyAppStack-MyFunction"
},
"exportNames": ["BucketArn"],
"lastModified": 1710835200000
}
ResourceState
interface ResourceState {
physicalId: string // AWS physical ID (ARN, name, etc.)
resourceType: string // CloudFormation resource type
properties: Record<string, unknown> // Resolved template intent (what cdkd was asked to deploy)
observedProperties?: Record<string, unknown> // AWS-current snapshot at deploy time (drift baseline)
attributes?: Record<string, unknown> // Attributes for Fn::GetAtt
dependencies?: string[] // List of dependent logical IDs
metadata?: Record<string, unknown> // Additional metadata
constructPath?: string // optional, no bump: the resource's `aws:cdk:path` at its last deploy; informational, no behavior reads it
deletionPolicy?: 'Delete' | 'Retain' | 'Snapshot' | 'RetainExceptOnCreate' // v5+: template attribute recorded at deploy time
updateReplacePolicy?: 'Delete' | 'Retain' | 'Snapshot' | 'RetainExceptOnCreate' // v5+: template attribute recorded at deploy time
provisionedBy?: 'sdk' | 'cc-api' // v7+: provisioning layer (absent = SDK legacy default)
observedBaselineRefused?: true // v10+: `cdkd import` declined to capture a baseline
observedBaselineRefusalReason?: 'unverifiable-parameter' | 'incomplete-resolution' // optional, no bump: only the first survives an in-place UPDATE
noEchoLeaves?: (string | number)[][] // v11+: positions in `properties` stored as `***` for a NoEcho value
noEchoAttributeNames?: string[] // v11+: `attributes` the provider declared NoEcho, each stored as `***`
noEchoExactEchoLeaves?: string[][] // optional, no bump: `noEchoLeaves` positions AWS reports exactly
acceptedCreateOnlyDrops?: string[] // optional, no bump: create-only properties the SDK route was told to drop, so never sent
maskedPropertyFingerprints?: Record<string, string> // optional, no bump: per property `properties` holds as `***`, a hash of its template text (issue #4451)
maskedPropertyInputFingerprints?: Record<string, string> // optional, no bump: per such property, a hash of its template value with its non-secret inputs resolved, bound to the text hash (issue #4543)
}
properties records the resolved CloudFormation template values cdkd
asked AWS to apply — the values it actually sent, which is not always
everything the template declared. A provider that deliberately narrows what
it sends records the narrowed bag, and a top-level property the SDK provider
has no wiring for is likewise absent whenever
--prefer-sdk-route kept the
resource on the SDK route (that flag is the opt-in to the property not being
written at all) — unless the property is create-only, which cdkd keeps in the
record because removing it would classify the next deploy as a replacement,
and names in acceptedCreateOnlyDrops. That name is what lets a later deploy
without the flag refuse the replacement instead of reporting no change
(CREATE_ONLY_DROP_NEEDS_REPLACEMENT).
One key is never sent: a nested stack's AWS::CloudFormation::Stack row
carries cdkd:PendingChildDeletes while its child stack (or one below it)
still holds a DELETE cdkd skipped, so the next cdkd deploy re-runs that
child and re-attempts the delete; it disappears once the delete lands. observedProperties records what AWS actually has — captured
by provider.readCurrentState immediately after each create/update so it
includes AWS-side defaults the user did not template. The cdkd drift
comparator prefers observedProperties as its baseline for richer detection;
resources without it fall back to properties (the pre-version: 3 behavior).
One consequence worth knowing before you opt into a drop: cdkd export
reconstructs a CloudFormation template from properties, so a property
accepted via --prefer-sdk-route is absent from the exported
template too. That matches what AWS actually holds — which is what makes the
exported stack importable — but it means the export is not a round-trip of your
CDK source for that field.
NoEcho custom-resource responses
A Lambda-backed custom resource's handler can declare its response Data
sensitive by setting the documented NoEcho: true field on the cfn-response
envelope. cdkd honours it: every
string value in that Data is stored as *** instead of the value itself —
in the custom resource's own attributes, in the resolved properties and
observedProperties of every resource that consumed it through Fn::GetAtt,
and in state.outputs.
A consumer value that builds a LONGER string around the value
(postgres://user:${cr.getAttString('Password')}@db/app) is stored as ***
whole, not with the value masked in place: an inline *** could not be told
from text you wrote, so nothing below could recognise and refuse it. The rest
of that string is therefore not in state either. A longer string is kept as it
is in two cases. The first is a value that occurs only inside the text of a
{{resolve:...}} reference to a service cdkd resolves: the reference is kept
so it can be resolved again, which means a reference NAME built from a
NoEcho value stores that value in the clear when the reference resolves to a
public value. One that resolves to a secret is refused instead: see
the troubleshooting entry. The second is a value that IS
public text state already holds: the region, the stack name, one of the custom
resource's literal template properties, its ServiceToken or one of that
ARN's :-separated parts (its account id, for example). For a value read
from another stack's output, only the regions and the stack names count. A
string EQUAL to such a value is still stored as ***. Any other string of the same consumer that happens to contain the value
is stored as *** too.
// in the handler
return {
PhysicalResourceId: id,
Data: { Token: mintedToken },
NoEcho: true, // <- this is the whole opt-in
};
AWS still gets the real value. CloudFormation delivers a NoEcho custom
resource's Data to a dependent resource in the clear, and cdkd matches that:
Fn::GetAtt resolves to the real value and the dependent is created with it.
Only what cdkd WRITES DOWN changes. (This is worth stating because the AWS
documentation's "masked with asterisks" sentence describes the display channel;
masking at resolution time would make a template feeding such a value into
AWS::SecretsManager::Secret.SecretString store the literal *** as the
secret.)
A new value reaches every consumer. When the handler runs and returns a
value, each resource reading it is updated with it — including a resource in a
nested stack that receives the value as a stack parameter — although the stored value
and the new one both read *** in state: two masks say nothing about the
values behind them. When the handler returns the same value again, cdkd cannot
tell, so:
- each consumer takes one redundant update;
- a consumer that is itself a custom resource has its OWN handler invoked again, with whatever side effects that handler has;
- a consumer holding the value in a property that cannot change in place is
read back from AWS first, because cdkd keeps only the mask and cannot compare
the new value with the old one. If AWS already holds exactly the new value
there, the consumer is not replaced: it is updated in place when something
else about it changed, and left alone when nothing did. The readback is
compared in memory and never stored, so nothing derived from the value lands
in state. The consumer is still REPLACED (or, for a stateful type, the deploy
stops and asks for
--force-stateful-recreation) whenever the readback cannot confirm the value: AWS holds a different one, the resource type has no readback, or the read fails. The deploy log says which, as a warning. A write-only property, which AWS never returns, is replaced only on the few resource types cdkd's own replacement rules name. On any other type, it is updated in place and never read back. So is any property of a type whose write-only list cdkd cannot look up, for example whencloudformation:DescribeTypeis denied.
There is a cost, and it is not hidden from you. cdkd has nothing to
re-derive the value from — a handler-generated value has no
{{resolve:...}} reference behind it — so once the mask is in state, cdkd will
not invent a value for it:
- A LATER deploy in which the custom resource is UNCHANGED does not re-invoke
the handler (CloudFormation semantics), so the only thing
Fn::GetAttcan read is the mask. If some resource actually has to be written in that deploy using that attribute, cdkd REFUSES it rather than sending***to AWS, and names the remedy: change one of the custom resource's properties (a nonce or version property is the usual way) so its handler runs again in the same deploy. A deploy that does not have to write the value is unaffected. cdkd driftreports the position but masks the live value, and--acceptrefuses to write that value into the baseline (accepting would undo the redaction).cdkd drift --revertleaves the position exactly as AWS has it rather than pushing the mask; when AWS reports nothing there, it refuses the resource.cdkd rollbackrefuses to replay a recorded baseline holding the mask, for the same reason. Force the custom resource to update and re-deploy to restore the property — a plain re-deploy leaves it unchanged, so the handler does not run and the mask stays.cdkd exportblocks a resource whose recorded properties hold the mask: the exported CloudFormation template would declare the literal***. Forcing the custom resource to update does NOT clear this one, because the export reads STATE and state is exactly where the mask lives. Stop settingNoEchoon that response and re-deploy, or export the stack without that resource.
ACROSS STACKS the value is available only within ONE run. Every cross-stack
route reads the PRODUCER's persisted outputs — a nested stack's
Fn::GetAtt [<Child>, 'Outputs.<Key>'], Fn::ImportValue, and
Fn::GetStackOutput — so a masked output has no plaintext for a consumer to
read. cdkd bridges the case it can: while the producer was deployed by the SAME
cdkd deploy process (a nested-stack child, or another stack in the same
cdkd deploy --all), the plaintext is still in memory and is handed to the
consumer, which then masks it in its OWN state record. Outside that:
- a
cdkd deploy Consumerrun whose producer was deployed EARLIER reads the mask, and is refused rather than writing***to AWS; - re-deploying the producer by itself does not help — it re-masks the value
on the way into its own state, so the consumer's next run reads the mask
again. Deploy the producer and the consumer in one run, with the producer's
custom resource actually running (force it to update), or stop marking that
response
NoEcho.
The Fn::Base64 encoding of a secret is masked the same way, with
different remedies. cdkd stores *** where a value is the Fn::Base64
encoding of a {{resolve:...}} reference (an EC2 UserData script, for
example), because the encoding decodes straight back to the secret. No custom
resource is involved, so forcing one to update does nothing, and cdkd cannot
tell this mask from a NoEcho one: each refusal names both causes.
cdkd rollbackrefuses to replay such a baseline. Acdkd deploythat changes the resource restores the value, because the deploy resolves the reference again and sends the encoding; a re-deploy that leaves the resource unchanged sends nothing. State still holds***afterwards, so a later rollback to it refuses again until the secret is no longer encoded into the property — have the resource read the secret at run time instead, not by writing its plaintext into the template.- A deploy refuses a resource that reads such a value back out of state, for
example through an
AWS::SSM::Parameter'sValueattribute when the parameter is unchanged. Have the reading resource build the value itself from the secret's own reference (the{{resolve:secretsmanager:...}}or{{resolve:ssm-secure:...}}reference, or a{{resolve:ssm:...}}of aSecureStringparameter, that the encoding was made from) under its ownFn::Base64. Do not read the parameter holding the encoding with a{{resolve:ssm:...}}reference instead: cdkd treats aStringparameter as public and would record the encoding in state in the clear. A re-import does not clear this mask. cdkd exportblocks the resource; see cdkd export.
Since version: 11 the record names the attributes a custom resource declared
NoEcho (noEchoAttributeNames), so a refused read names that cause alone; a
mask with no such record still names every cause.
Known bound: a value used as a NAME. A resource's physical id is what
cdkd uses to find it again, so it is never masked. A NoEcho value passed as a
create-only name (QueueName, TableName, a parameter Name) is therefore
stored in the clear as that resource's physical id. The attributes AWS builds
around the name on that resource (a queue URL, an ARN) are stored as ***,
because they contain the value (so a later deploy that has to write one of
them is refused, like any masked read), but any other resource's property or output
that reads one of those holds it in the clear, and so does any command output
that shows a physical id. CloudFormation behaves the same way: DescribeStackResources
returns the physical id in the clear, whatever NoEcho said. Use NoEcho
values as values, never as names.
physicalId Format
Varies by resource type. Examples:
| Resource Type | physicalId Example |
|---|---|
AWS::S3::Bucket |
my-bucket-name |
AWS::Lambda::Function |
arn:aws:lambda:us-east-1:123456789012:function:MyFunc |
AWS::IAM::Role |
MyRole (role name) |
AWS::DynamoDB::Table |
MyTable (table name) |
AWS::SQS::Queue |
https://sqs.us-east-1.amazonaws.com/123456789012/MyQueue |
Custom::MyResource |
Any string returned by custom resource |
Note: cdkd supports all resource types supported by Cloud Control API. The table above shows only a few examples. For resources not supported by Cloud Control API, custom SDK Providers can be implemented (see Provider Development).
The physicalId is provider-defined, and it may differ from the value
CloudFormation records for the same resource. cdkd stores whatever the
provider that created the resource returned — the value that provider needs
to address the resource again on update / delete / drift. For most types
that is the same scalar CloudFormation's Ref returns (a bucket name, a
function ARN), but it is not guaranteed to be: see the composite forms
below. Always read the id you must reuse from cdkd itself
(cdkd state show '<stack>' / cdkd state resources '<stack>') rather than
from the AWS console or CloudFormation's DescribeStackResources.
Composite (pipe-delimited) physicalIds
Some resources have no single AWS-side identifier — a Glue table is only
addressable as (database, table); an API Gateway method as (restApi,
resource, httpMethod). For those types cdkd stores a composite physical
id: the identifying segments joined with a | pipe. That is deliberately
the same convention Cloud Control API uses for a multi-part
primaryIdentifier, so a type that moves between an SDK Provider and the
Cloud Control fallback keeps a compatible id (AWS::EC2::EIP is the
explicit case — its SDK Provider reproduces the id shape the Cloud Control
path had produced).
The composite value is what state records, what cdkd state show /
cdkd state resources print, and what
cdkd import --resource '<logicalId>=<physicalId>' expects. A few types also
accept a looser form on import — see
Importing Existing Resources for the
per-type notes.
| Resource Type | physicalId format |
|---|---|
AWS::ApiGateway::Method |
<restApiId>|<resourceId>|<httpMethod> |
AWS::AppSync::ApiKey |
<apiId>|<apiKeyId> |
AWS::AppSync::DataSource |
<apiId>|<name> |
AWS::AppSync::Resolver |
<apiId>|<typeName>|<fieldName> |
AWS::EC2::EIP |
<publicIp>|<allocationId> |
AWS::EC2::NetworkAclEntry |
<networkAclId>|<ruleNumber>|<egress> (egress is true / false) |
AWS::EC2::Route |
<routeTableId>|<destination> (destination is the DestinationCidrBlock, DestinationIpv6CidrBlock, or DestinationPrefixListId the route declares) |
AWS::EC2::SecurityGroupIngress |
<groupId>|<ipProtocol>|<fromPort>|<toPort> (an omitted port is recorded as -1) |
AWS::EC2::VPCGatewayAttachment |
<internetGatewayId>|<vpcId> (note the order — CloudFormation's own identifier is VpcId first) |
AWS::Glue::Table |
<databaseName>|<tableName> (either name may itself contain |: cdkd reads the table name as everything after the recorded DatabaseName) |
AWS::Lambda::EventInvokeConfig |
<functionName>|<qualifier> (a bare function name is read as qualifier $LATEST) |
AWS::Route53::RecordSet |
<hostedZoneId>|<name>|<type> |
AWS::S3Tables::Namespace |
<tableBucketARN>|<namespaceName> |
AWS::S3Tables::Table |
<tableBucketARN>|<namespace>|<name> |
Examples as they appear in a real state file (resources map, abridged):
{
"MyGlueTable": { "physicalId": "my_database|my_table" },
"MyGetMethod": { "physicalId": "a1b2c3d4e5|xy9z8w|GET" },
"MyARecord": { "physicalId": "Z1D633PJN98FT9|www.example.com.|A" },
"MyEip": { "physicalId": "52.1.2.3|eipalloc-0abc123def456789a" }
}
The composite id is NOT what Ref returns
CloudFormation's Ref for these types returns a value of its own, which is
usually only a PART of cdkd's composite — and sometimes not a part of it at
all. cdkd translates the stored id back to CloudFormation's value before
handing it to any consumer (Fn::Join / Fn::Sub / a CfnOutput), so a
template gets the same value it would from cdk deploy. You do not need to do
anything; the table is here because the difference is visible when you compare
cdkd state show against a stack output.
| Resource Type | CloudFormation Ref returns |
|---|---|
AWS::ApiGateway::Method |
an AWS-generated id (no segment reconstructs it — cdkd passes the composite through) |
AWS::AppSync::ApiKey |
the API key ARN |
AWS::AppSync::DataSource |
the data source ARN |
AWS::AppSync::Resolver |
the resolver ARN |
AWS::EC2::EIP |
the public IP (the segment before the first |) |
AWS::Glue::Table |
the table name — everything after the recorded DatabaseName and its |, so a table named a|b resolves to a|b |
AWS::Route53::RecordSet |
the record name. A three-part id returns its middle segment without reading state. When the name itself contains | (a longer id), it is everything between the first and the last |, only when that matches the recorded Name and the last segment the recorded Type; an id that does not match passes through raw |
AWS::S3Tables::Namespace / ::Table |
the namespace / table name (the segment after the last |) |
The three AWS::AppSync::* children are the case where the Ref value is not
a segment at all: cdkd recovers the ARN from the attribute the provider records.
cdkd import records the same attribute a fresh deploy does — it reconstructs
the ARN from the composite id you supply — so an adopted child's Ref and
Fn::GetAtt resolve immediately.
Some records can still lack the attribute. Ref on such a record falls back to
the raw composite id. For Fn::GetAtt on the ARN attribute, cdkd deploy
re-reads the resource from AWS once and records the real ARN; when that read
cannot supply one it FAILS rather than serving a value CloudFormation would not
return. The records are:
- one written by a cdkd older than the fix that started recording the real ARN;
- one whose import could not reach STS, so cdkd could not determine the account. It deliberately records NOTHING rather than an ARN built from a placeholder account id, which would look valid and be wrong;
- one whose import could not build the ARN for some other reason.
Each of the import cases names itself in a warning at import time.
A deploy that resolves a Fn::GetAtt on the ARN heals the record, as does the
resource's next in-place update — see item 4 under
Purpose of attributes.
…and it is not what cdkd export sends CloudFormation either
cdkd export hands a stack to CloudFormation via an IMPORT changeset,
which addresses each resource by its CFn primaryIdentifier. For most
composite types that identifier is multi-field and cdkd splits the id
into it. Five types are different — their CFn identifier is a SINGLE
field holding a value that is not cdkd's physical id (for four of them
not any segment of cdkd's composite; for the GraphQL API, whose physical
id is the bare apiId, the ARN CloudFormation has identified it by since
September 2026):
| Resource Type | CloudFormation IMPORT identifies it by | cdkd resolves it from |
|---|---|---|
AWS::AppSync::DataSource |
DataSourceArn |
the recorded DataSourceArn attribute |
AWS::AppSync::Resolver |
ResolverArn |
the recorded ResolverArn attribute |
AWS::AppSync::GraphQLApi |
Arn |
the recorded Arn attribute |
AWS::S3Tables::Table |
TableARN |
the recorded TableARN attribute |
AWS::EC2::SecurityGroupIngress |
Id (the sgr-… rule id) |
the recorded Id attribute |
You do not need to do anything for the ARN-identified four on a stack
deployed by a current cdkd: a fresh deploy and cdkd import both record
the attribute. A record that lacks it — the degraded cases listed above —
makes cdkd export block that resource with a message naming the
attribute; re-deploy the stack once to heal the record, then re-run the
export.
AWS::EC2::SecurityGroupIngress has two ways to lack its Id, and
only one of them is healed by re-deploying:
- The rule declares more than one source. A single ingress resource
setting both
CidrIpandCidrIpv6makes AWS mint one rule per source, and cdkd deliberately records NEITHER id — neither one is "the" identifier for that resource, and picking one would name the wrong rule in the import changeset. Re-deploying never heals this; split the resource into oneAWS::EC2::SecurityGroupIngressper source, which is also the shape CloudFormation manages after the export. - The rule predates id recording — it was created by a cdkd older
than the one that started recording the id at all. This is the one
exception to "re-deploy once": AWS returns the
sgr-…id only fromAuthorizeSecurityGroupIngressitself, so a no-op deploy issues no call and records nothing. You do not have to do anything about this one —cdkd exportrecovers the id itself, by looking the rule up in AWS (see below). Only if that lookup cannot answer do you need the manual remedy: cdkd updates this type by revoking and re-authorizing, so changing ANY property of the rule mints a fresh id — as does destroying and re-deploying it. Either way the rule's traffic is interrupted for the moment between the revoke and the re-authorize, so pick the window.
The live-read backfill. For a row with no usable recorded Id,
cdkd export issues a paginated DescribeSecurityGroupRules on the
security group its physical id names and adopts the rule only when
EXACTLY ONE ingress rule on that group carries the composite's
(protocol, port range) tuple. Zero matches is refused with a message
naming the row and the tuple cdkd searched for, since nothing matched
and there is nothing to name; more than one is refused with a message
naming the row and EVERY candidate sgr-… id — cdkd's physical id
identifies a rule only by group, protocol and port range, so two rules
sharing that tuple are two rules cdkd cannot tell apart either, and
adopting one would import the wrong rule. Matching rules are counted
BEFORE any is set aside, so a rule AWS reports without a usable sgr-…
id refuses too rather than letting its sibling pass as "exactly one" —
and when more than one rule matched, that refusal carries the two-cause
remedy below as well, since such a row is ambiguous no matter how
readable the ids are. "More than one" has two causes with different
remedies: the multi-source rule above (split the resource), and two
DISTINCT ingress resources differing only by SOURCE — port 443 from a
CIDR and port 443 from a peer security group — which cdkd's composite
cannot tell apart because it carries no source. Those are already one
resource per source, so their remedy is to set the row's attributes.Id
to the sgr-… id that belongs to it, or to remove the row before
exporting. The lookup needs ec2:DescribeSecurityGroupRules; without
that permission the row is blocked with a message saying so, while a
THROTTLED lookup is retried with backoff and, if it still fails,
reported as a throttle rather than as a missing permission. A row that
already records the Id — everything a current cdkd deploys — issues no
live read at all.
In both cases you can instead remove the rule from the stack before exporting: it stays in AWS and can be re-declared in CloudFormation afterwards.
Some composite types cannot be exported at all, for an unrelated reason:
CloudFormation itself refuses AWS::Glue::Table,
AWS::Route53::RecordSet, AWS::AppSync::ApiKey and
AWS::EC2::NetworkAclEntry in IMPORT changesets. cdkd export detects
that up front and names every affected resource — see
cdkd export.
Two more types accept a composite id without producing one:
AWS::ECS::Service— cdkd stores the service ARN, but<clusterArn>|<serviceName>is also accepted on--resource.AWS::Lambda::Permission— cdkd stores the bare statement id; state written by the older Cloud Control path may instead hold<functionArn>|<statementId>, and both are read correctly.
Important
|is the shell pipe character. Always quote a composite id when you pass it on a command line:cdkd import MyStack --resource 'MyGlueTable=my_database|my_table'Unquoted, the shell splits the command at the
|and the import runs against a truncated id. JSON mapping files (--resource-mapping/--resource-mapping-inline) need no escaping —|is an ordinary character in JSON.
Important
The separator is not escaped, so a segment that contains a
|would make the id ambiguous.cdkd deployrefuses at pre-flight, naming the offending segment, rather than record such an id. There are two exceptions. One isAWS::Glue::Table, where both the table name and the database name may contain|: a table nameda|bin databasex|yis recorded asx|y|a|b, and cdkd reads the table name back as everything after the recordedDatabaseName— for update, destroy, drift andRefalike. A record whoseDatabaseNameis not a plain string (cdkd importcan leave it unresolved) cannot be placed that way: destroy then skips it and says to setproperties.DatabaseNamein the state file. The other isAWS::Route53::RecordSet, whose record name may contain|: a record nameda|b.example.comof typeAin zoneZ1is recorded asZ1|a|b.example.com|A, and cdkd reads the name back as everything between the first and the last|only when it matches the recordedNameand the last segment matches the recordedType. An id that does not match is read as CloudFormation's own physical id (the record name), and the zone comes from the recorded properties. For every other type, AWS's own naming rules and generated ids keep|out of the value.
Purpose of attributes
Stored to resolve attribute references via Fn::GetAtt.
Example:
# CloudFormation template
!GetAtt MyBucket.Arn
↓ cdkd resolves
const bucketState = state.resources['MyBucket'];
const arn = bucketState.attributes['Arn'];
// => "arn:aws:s3:::myappstack-mybucket-abc123xyz"
How Attributes are Collected:
Cloud Control API: Automatically collected from
GetResourceresponseSDK Provider: Provider explicitly returns in
create()/update()cdkd import: Provider returns them fromimport(), so an adopted resource carries the same attribute snapshot a deployed one does. When a provider'simport()returns no attributes — whether it omits the field or returns an empty{}, which is what most providers do — cdkd falls back to the map already in state, but only if the resource is being re-imported at the same physical id. A re-import that repoints a logical id at a different physical resource never inherits the old one's attributes. With neither source the map is empty ({}).Providers deliberately omit an attribute key rather than storing an empty string when a read-back cannot supply the value: the intrinsic resolver treats any non-
undefinedstored attribute as a hit, so a persisted''would shadow its computed fallback and makeFn::GetAttresolve to the empty string.cdkd deploy, on a miss: when aFn::GetAttis about to fall back to the physical ID for a resource this deploy does not update, cdkd re-reads the resource's attributes once through the provider's read-onlyimport()and adds them to the record at the next state save. Only keys the record does not already hold are added (a wildcard placeholder ARN from an old release is the one value that is overwritten), an empty value is never added, and no other field of the record is touched.--dry-runreads but records nothing.cdkd diffissues the same read for its preview and records nothing; other read-only commands (cdkd drift) never re-read. See "Cannot resolve" a GetAtt on a resource an older cdkd deployed.
// IAM Role Provider example
return {
physicalId: roleName,
attributes: {
Arn: response.Role?.Arn,
RoleId: response.Role?.RoleId,
},
};
Purpose of dependencies
Used to determine proper deletion order in destroy command.
Dependency Recording Timing: Extracted from DAG during deployment
// deploy-engine.ts
const resourceState: ResourceState = {
// ...
dependencies: dagNode.dependencies.map(dep => dep.logicalId),
};
Determining Deletion Order: Topological sort in reverse of dependencies
Creation order: Bucket → Role → Function
Deletion order: Function → Role → Bucket (reverse)
LockInfo (lock.json)
interface LockInfo {
owner: string // Process identifier (e.g., "user@hostname:12345")
timestamp: number // Lock acquisition time (Unix timestamp, milliseconds)
expiresAt: number // Lock expiry (Unix timestamp, milliseconds); RENEWED while the holder lives
operation?: string // Operation in progress (e.g., "deploy", "destroy")
}
expiresAt moves forward roughly every two minutes for as long as the holding
process is alive (see "Lock renewal" below), so it is not the time the
operation started plus the TTL -- it is the deadline by which the holder must
next check in.
Example
{
"owner": "goto@macbook:12345",
"timestamp": 1710835200000,
"expiresAt": 1710837000000,
"operation": "deploy"
}
Lock implementation
The lock is an S3 object written and deleted with conditional requests. Its behaviour as a user sees it is in State locking.
Like the state backend, the lock manager resolves the state bucket's
actual region via GetBucketLocation before its first S3 operation and
rebuilds its S3 client when the bucket lives in a different region from
the CLI's base region, so locking works against a
cross-region state bucket too. The per-bucket region lookup is cached, so
this adds no extra API call when the state backend already resolved the
same bucket.
Lock Acquisition (Acquire)
// Using If-None-Match: "*"
// → Succeeds only if object doesn't exist
await s3Client.send(
new PutObjectCommand({
Bucket: stateBucket,
Key: `cdkd/${stackName}/${region}/lock.json`,
Body: JSON.stringify(lockInfo),
IfNoneMatch: '*', // ← Important: only if object doesn't exist
})
);
Success: Lock acquired → Continue processing
Failure (PreconditionFailed): Lock already exists → Another process is running
Lock Release (Release)
The DELETE is conditional on the ETag this process last wrote, so a process can only ever delete the lock object it still owns:
// IfMatch: the ETag returned when this process wrote (or last renewed) the lock
await s3Client.send(
new DeleteObjectCommand({
Bucket: stateBucket,
Key: `cdkd/${stackName}/${region}/lock.json`,
IfMatch: heldEtag,
})
);
A PreconditionFailed here means the lock present is somebody else's; cdkd
leaves it in place and warns rather than raising, because the operation itself
has already finished and the caller has nothing to do about it.
The condition is dropped for exactly one class of failure: the endpoint or
the policy will not evaluate it at all. A conditional delete with a specific
ETag additionally requires s3:GetObject, so a policy granting only
s3:DeleteObject answers 403, and an S3-compatible endpoint that has not
implemented the header answers 501. Those fall back to an unconditional
delete so such a setup cannot end up with a stranded lock.
Even then the ownership is re-checked by hand before the condition is dropped,
because S3 authorizes a request before it evaluates a precondition: a policy
that scopes s3:GetObject away from lock.json turns a genuine 412 into a
403, and an unconditional retry there would delete the lock of whoever took
over. The re-read happens unconditionally -- not skipped when cdkd's own
deadline is still in the future, because cdkd force-unlock deletes regardless
of expiry, so a user running it mid-operation is a legitimate takeover no
deadline can rule out (cross-machine clock skew reaches the same state with
nobody running anything). A read that FAILS refuses: on the very policy this
fallback exists for, the read fails too, so answering "proceed" there would
leave the check inert in exactly the situation it was added to catch.
The expired-lock takeover has no such fallback, deliberately. Its IfMatch
is what makes concurrent reaping safe -- two processes that both judge a lock
expired race to delete it, the first wins and the second gets a 412 and
reports contention. Without the condition both would win, each would then
acquire against the key it just emptied, and the stack would have two holders.
An expired lock under a policy that cannot evaluate the condition is cleared
with cdkd force-unlock.
Every other failure raises, which is what release has always done. In
particular a 409 (S3's answer to a concurrent operation on the key) and a
503 are not fallback-worthy: the first is the contended case by
definition, and the second may mean the conditional delete already succeeded
with the response lost, so an unconditional retry would delete whichever lock
exists by then. The heartbeat is already stopped at that point, so the worst
outcome is a lock that lapses at its TTL -- recoverable, unlike a lock deleted
out from under a live writer.
A failed release never fails the command. Every caller wraps it and logs a
warning, so a throttled or conflicted release is reported without replacing the
error the command was actually about -- and without aborting a cdkd destroy --all run at the first stack over a lock that clears itself.
A second releaseLock for the same key is a no-op rather than an owner-blind
delete: the entry is tombstoned, not dropped. This matters because the
force-quit paths fire an un-awaited release while the main finally may still
be in one.
cdkd force-unlock is deliberately not conditional: it exists precisely to
remove a lock this process does not own.
In older cdkd versions this was an owner-blind unconditional delete, which is what turned a single lapsed lock into a cascade -- a process whose lock had been taken over deleted the new owner's lock on its way out, freeing the stack for a third writer.
Retry Logic
async acquireLockWithRetry(
stackName: string,
region: string,
owner?: string,
operation?: string,
maxRetries = 3,
retryDelay = 2000 // 2 seconds
): Promise<void> {
let lockInfo = null;
let releasedReacquires = 0;
for (let attempt = 0; attempt <= maxRetries; attempt++) {
// acquireLock reaps an EXPIRED foreign lock itself and retries once, so a
// `false` here means someone else held the lock at the PUT.
if (await this.acquireLock(stackName, region, owner, operation)) return;
lockInfo = await this.getLockInfo(stackName, region);
if (!lockInfo && releasedReacquires < RELEASED_LOCK_REACQUIRE_LIMIT) {
// Released since the PUT: try again now, without spending a retry.
releasedReacquires++;
attempt--;
continue;
}
if (attempt < maxRetries) {
// Reports the holder (or that none could be read), then waits.
await sleep(retryDelay);
}
}
// Renders the read taken after the LAST failed acquire -- the holder, or
// "No lock could be read" -- never a second read.
throw new LockError('Failed to acquire lock after retries');
}
RELEASED_LOCK_REACQUIRE_LIMIT is 3. The bound matters because getLockInfo
also reads a lock.json whose body is not an object as "no lock", and that
object never lets the PUT through.
Expiry is decided by the lock's own expiresAt field, not by its age: a live
holder keeps pushing that field forward, so "old" and "abandoned" are different
questions and only the second one frees the lock.
Lock renewal
The holding process renews its lock in the background, re-writing
expiresAt at most every 2 minutes (or every quarter of the TTL, whichever
is shorter) for as long as the operation runs. Each renewal is a conditional
PutObject carrying IfMatch with the ETag of the object this process last
wrote, so a process that has already lost the lock cannot resurrect its own
expiry on top of the new owner's.
This is what makes the TTL mean "the owner has been silent for 30 minutes"
rather than "the operation has been running for 30 minutes". The default TTL
tolerates fourteen consecutive missed renewals (a throttle, a network blip)
before it lapses; a renewal that fails for any reason other than "this lock is
no longer mine" is simply retried on the next tick. If they fail long enough
that the deadline actually passes, cdkd says so once at warn -- otherwise
half an hour of failing renewals would read exactly like a healthy run while
another process becomes free to take the lock.
A 412 on a renewal is not taken at face value. A conditional PutObject that
S3 applied but whose response was lost -- or an SDK-internal retry of it --
leaves the cached ETag one version behind, so the next attempt legitimately
conflicts with cdkd's own write. cdkd reads the object once to tell the two
apart and adopts the renewal when the stored body is byte-for-byte what it just
wrote (same owner, same acquisition timestamp, same millisecond deadline).
Without that check the process would declare a lock it still owns lost, warn
about a concurrent writer that does not exist, and then refuse to release its
own lock.
In older cdkd versions there was no
renewal at all, so any operation slower than the TTL silently stopped being
mutually exclusive while it was still running. That is reachable without
anything exotic: AWS::FSx::FileSystem, AWS::EMR::Cluster and Custom
Resources each wait up to an hour on their own, and a large enough stack
exceeds 30 minutes in aggregate regardless of resource type.
Two consequences worth knowing:
A lock whose
expiresAtis not a finite number counts as EXPIRED. That field arrives from the state bucket unvalidated, andInfinity/NaN/ a string would otherwise pin the stack forever: no acquisition would ever succeed again and onlycdkd force-unlockcould clear it. Treating it as expired grants no new power -- anyone who can write that value could equally have deleted the object -- and it is the recoverable direction.A lock that reaches its
expiresAtnow genuinely means an absent owner -- a crashed process, aSIGKILL, or a machine that slept. cdkd logs the takeover atwarnlevel naming the previous owner, because on the remaining chance that the process IS alive, two writers are now operating on the stack.The state bucket is versioned, so each renewal adds one
lock.jsonobject version. A 30-minute deploy writes about fifteen, and they go noncurrent the moment the next renewal lands. ADeleteObjecton a versioned bucket writes a DELETE MARKER and leaves every earlier version readable throughGetObjectwith aVersionId, so previously the release left the whole chain behind and the count grew for the life of the bucket -- 452 versions on a single measured key, invisible toaws s3 ls, still billed, and still paged through by every version listing the other purge sites issue. cdkd now purges the lock key's noncurrent versions wherever it deletes the lock: on release, on the expired-lock takeover, and oncdkd force-unlock. Nothing sensitive is in them (lock.jsoncarries onlyowner,timestamp,expiresAtand an optionaloperation), so this is bucket cost rather than disclosure -- which is exactly why the reporting differs from the four secret-bearing sites: a purge failure on the RELEASE path is logged atdebug, notwarn, so a least-privilege principal withouts3:ListBucketVersions/s3:DeleteObjectVersiondoes not get a warning at the tail of every mutating command about a heartbeat record. The two rare reap paths (takeover,force-unlock) do warn, matching the cost profile the bucket-policy section describes -- that section is also where the two grants and the never-throw contract are spelled out. Nothing fails either way.The trade that split makes, stated rather than left implicit: a principal who has
s3:ListBucketVersionsbut lackss3:DeleteObjectVersionnow gets no routine signal that the chain is still growing, because the release path -- the one that runs on every mutating command -- reports atdebug. (A principal missings3:ListBucketVersionsis already warned on every successful deploy by the rollback journal's own purge, so nothing changes for them.) That is deliberate (a warning per command about a heartbeat record is worse than the growth it reports), but it means the growth is silent for exactly the population that cannot stop it. The fallback is not a lifecycle rule -- see below for why one is not expressible against this key layout -- it is to grant the two actions, at which point the purge simply works. To check whether it is happening, run any mutating command with--verboseand look for the purge line, or count the versions directly:# Both halves matter: the purge removes noncurrent BODIES and noncurrent # DELETE MARKERS, and a released lock key accumulates one marker per cycle, # so counting bodies alone undercounts the growth. The `|| ` + "[]" default is # not decoration either -- on a fully purged key the response carries no # `Versions` array at all, and `length(null)` is a JMESPath ERROR, so the # naive query fails exactly when the answer should be 0. aws s3api list-object-versions --bucket <state-bucket> \ --prefix "cdkd/<stack>/<region>/lock.json" \ --query 'length(Versions[?IsLatest==`false`] || `[]`)' aws s3api list-object-versions --bucket <state-bucket> \ --prefix "cdkd/<stack>/<region>/lock.json" \ --query 'length(DeleteMarkers[?IsLatest==`false`] || `[]`)'Two things this does NOT do. The CURRENT version is never touched -- the purge filters on
IsLatest, so it can neither delete the live lock nor remove the delete marker whose removal would resurrect a stale one. And it is scoped to thelock.jsonKEY, never to the stack prefix, sostate.json's history is untouched.A bucket lifecycle rule is not an alternative here, contrary to what this document said before the site shipped. The lock key is
cdkd/{stackName}/{region}/lock.json--lock.jsonis a key SUFFIX, interleaved under the same per-stack prefix asstate.json,rollback-journal.jsonanddeployments/. S3 lifecycle filters supportPrefix,TagandObjectSizeonly, so no rule can select the lock keys and sparestate.json; the one expressible prefix rule,cdkd/, would expirestate.json's noncurrent versions too -- the recovery capability the exempted keys are deliberately held open to protect, done in bucket configuration instead of in code. A tag-scoped rule would needs3:PutObjectTaggingon the lock write, i.e. on the hot path.Still deliberately NOT purged:
state.jsonitself and the v1 -> v2 migration delete, because those noncurrent versions ARE the state-recovery capability versioning is enabled for. The one exception is an explicitcdkd scrub, which purges the history of astate.jsonit rewrites (and, under--purge-history, of every record it examines and does not refuse), short of deleting the bucket itself (cdkd bootstrap --destroy,cdkd state migrate --remove-legacy) — seecdkd scrub.
If the holding process dies without releasing, the lock stops being renewed and
is reclaimed by the next cdkd invocation once expiresAt passes -- or
immediately with cdkd force-unlock '<stack>'.
State writes during deploy and destroy
A deploy acquires the stack lock, loads state.json together with its ETag,
and then saves more than once:
- after each resource operation completes, so an interrupted run leaves a record of what it finished;
- before an automatic rollback, so resources the failed run created are tracked;
- once at the end, with the resolved outputs.
The per-resource saves and the pre-rollback save pass the ETag of the previous
write as If-Match, so a write is refused when the record moved underneath it
(State has been modified by another process). The final save of a successful
deploy is unconditional; it runs under the stack lock. The save that migrates a
legacy version: 1 key writes the region-prefixed key unconditionally and
then deletes the legacy key. S3 returns an ETag wrapped in double quotes, and
cdkd stores and sends it as returned, quotes included.
A failed incremental save is logged as a warning and does not fail the deploy.
On the next run, the diff sees the resources that were never recorded as
CREATE and retries them.
Destroy
Incremental state persistence during destroy: the destroy path
mirrors deploy's per-resource state saves. Each successfully deleted
resource (including resources found already deleted on a re-run) is removed
from the state object and the trimmed state is written back to S3
immediately, serialized under the stack lock the destroy already holds. An
interrupted (Ctrl-C) or partially-failed destroy therefore preserves a state
file that only lists resources that still exist — a re-run does not replay
deletes against already-deleted resources (which previously caused, for
example, a 10-minute stall per Custom Resource whose backing Lambda had
already been deleted). Resources retained via DeletionPolicy: Retain stay
in every intermediate snapshot; their record is only dropped by the
wholesale state-file delete at the end of a fully successful destroy. A
failed incremental write is logged and never fails the destroy — the final
write (state-file delete on success, preserve-write on failure) remains
authoritative.
Every persisted destroy snapshot (both the incremental writes and the final
partial-failure preserve-write) clears outputs and drops imports /
outputReads. outputs is keyed by output name, not logical id, so it
cannot be pruned precisely as the backing resources are deleted; a
partially- or fully-destroyed stack has no meaningful outputs, and leaving
them in the preserved state would advertise an export whose backing resource
is gone — a phantom export the
exports index or another producer's
strong-reference consumer scan (scanActiveConsumers) could pick up.
Clearing them removes that hazard. This does not affect the destroy's
own strong-reference check: that reads the in-memory state.outputs
before the delete loop, and the in-memory state object is never mutated
— only the persisted snapshot copies are cleared. On a clean destroy the
stack's entry is removed from the exports index outright
(exportIndexStore.removeStack); on a partial destroy the index may briefly
still list stale entries, but that index is a perf-only derived view that
self-heals on the next deploy / fallback scan, while the canonical
state.json no longer carries the phantom outputs.
Noncurrent-version purges
The state bucket is versioned, so a plain DeleteObject writes a delete
marker and leaves every earlier version of the key readable through
GetObject with a VersionId. Wherever cdkd deletes an object that is
transient by design, it also purges that key's noncurrent versions. The purge
needs s3:ListBucketVersions and s3:DeleteObjectVersion; the policy that
grants them is in
State backup and bucket security.
Six kinds of object need those two version actions, not one — the five in
this table and the deployments/** event store below it, plus state.json and
the exports index for a principal that runs cdkd scrub. The set has grown
over time, and the ordinary
commands are now in it:
| object | purged by | what its previous versions hold |
|---|---|---|
rollback-journal.json |
every successful cdkd deploy, every clean cdkd rollback, cdkd destroy / cdkd state destroy |
failedOperations[].attemptedProperties — the properties of the FAILED write, verbatim. Measured on a repo fixture as four versions each carrying a literal "MasterUserPassword" |
| custom-resource response object | cdkd deploy (the provider's own cleanup) and cdkd gc |
the handler's FULL cfn-response, Data included — where a handler-minted password or API key lands |
| transient CFn template | cdkd import --migrate-from-cloudformation, cdkd export, and MACRO EXPANSION during cdkd deploy / cdkd diff (any template over the 51,200-byte inline ceiling) |
the template body, which carries a secret only if the template does (an inline Code.ZipFile, a hand-written literal) |
cdkd-bootstrap/{region}.json |
cdkd bootstrap --destroy |
the asset bucket and container-repo names. No secret; listed for completeness |
lock.json |
every command that RELEASES a stack lock, plus the expired-lock takeover and cdkd force-unlock |
the lock heartbeat: owner, timestamp, expiresAt, operation. No secret -- one row per two-minute renewal, which is why it was the fastest-growing key in the bucket. Reported differently: see below |
The journal is the one to note if you are deciding whether this matters to you:
it is written by an ORDINARY failed or interrupted deploy, not by an opt-in
feature, and it is swept by an ordinary cdkd destroy. state.json is
deliberately NOT in this table — its previous versions are the state-recovery
capability versioning is enabled for, and short of deleting the bucket only an
explicit cdkd scrub purges them (see
cdkd scrub).
The deployments/** event store needs the two version actions as well.
Every path that deletes from it — the writer's self-bounding prune,
cdkd events prune, and cdkd destroy --purge-events — also purges the
noncurrent versions of the keys it deletes, and cdkd events prune --all /
cdkd destroy --purge-events purge every noncurrent version under the stack's
deployments/ prefix. Each run's stream is re-written in
full per flush, so one run leaves one version per flush, and the repo classes
this content as sensitive: see
Deleting a run stream also purges its earlier versions
for what that purge does not reach.
One other key family is not in the table either. Short of deleting the
bucket, only cdkd scrub purges it, so the previous versions every other write leaves accumulate and
stay readable:
_index/{region}/exports.json— the exports index, which holds resolved Output values.cdkd deployrewrites it, and so doescdkd scrub, one entry at a time, for the stacks that run scrubbed. A deploy's write leaves the previous body as a noncurrent version of a key SHARED by every cdkd-managed stack in the region;cdkd scrubpurges the key's noncurrent versions once per region whose entries it wrote. Seecdkd scrub.
Without the two grants, nothing fails — and that is the point to
understand. The purge runs on a cleanup path and must never abort the
operation it follows, so it logs a warning and the deploy, diff, rollback,
destroy, cdkd import, cdkd export, cdkd gc, cdkd events prune or cdkd scrub run still succeeds. What does not
happen is the removal: the value stays retrievable by anyone who can read the
state bucket with a VersionId. The warning counts KEYS, names them (a
prefix-wide sweep whose listing failed is counted as every key under 1 prefix(es) and named <prefix>* (every key under this prefix)), names
WHICH object it failed on, and spells the two actions exactly:
Could not purge noncurrent versions of 1 key(s) in s3://cdkd-state-bucket. Their
previous versions survive and remain readable via GetObject with a VersionId
(the rollback journal, whose `failedOperations[].attemptedProperties` records the
properties of the failed write verbatim). Grant s3:ListBucketVersions and
s3:DeleteObjectVersion on the state bucket, or purge the key(s) by hand.
Failures: cdkd/MyStack/us-east-1/rollback-journal.json (AccessDenied:
s3:ListBucketVersions)
The parenthetical is per-object — a custom-resource response object, the transient template and the bootstrap marker each name themselves — so the warning always says what to go and look at.
(Line-wrapped here; cdkd emits it as one line. It names up to five keys and
appends (and N more) beyond that, so the tail first appears at six.)
The two grants fail in different ways, and only one of them fails loudly on
its own. Missing s3:ListBucketVersions denies the listing, so the whole
purge stops. Missing s3:DeleteObjectVersion does NOT throw: DeleteObjects
reports per-key refusals in a response.Errors array and returns success
overall, so cdkd has to read that array to notice. It does — a partial failure
across a batch is counted key by key and named the same way — but it is why
granting one of the two and not the other is worth avoiding: everything looks
normal except the warning.
lock.json is the one exception to the warning, deliberately. A release
runs at the tail of EVERY mutating cdkd command, so inheriting the warning
would mean a principal on the older four-action policy -- who sees a silent
clean deploy today -- getting one after every single command, about a
heartbeat record with no secret in it. Release-path purge failures therefore go
to debug (visible under --verbose); the two rare reap paths, the
expired-lock takeover and cdkd force-unlock, still warn. Every other object
in the table warns as described above. Nothing fails on any of them.
A per-key NoSuchVersion counts as SUCCESS rather than as a failure, on every
object. The version named is already gone, which is the state the purge exists
to produce -- and on the lock key it is reachable in normal operation, because
a process reaping an abandoned lock and its original owner waking up to release
it can legitimately purge the same key at once. Reporting it would tell a
blameless user to grant permissions they already hold.
If you are on the older four-action policy, adding the two version actions is
the whole fix; the objects already stranded before the change have to be purged
by hand (aws s3api list-object-versions + delete-object --version-id). The
third addition, s3:GetReplicationConfiguration, is diagnostic: nothing fails
without it.
How cdkd detects replication after a purge
cdkd tells you when it applies — if it can see your replication
configuration. A purge that removed a BODY — or that could not establish
whether there was one — ends with one GetBucketReplication on the state
bucket, cached for the rest of the run. When
a rule covers the keys just purged, the warning names the destination:
S3 replication is enabled on s3://cdkd-state-bucket and covers the key(s) cdkd
just purged. S3 NEVER replicates a version-id delete, so the purge removed those
versions from THIS bucket only — the copies in the destination bucket survive and
remain readable there via GetObject with a VersionId (the rollback journal, whose
`failedOperations[].attemptedProperties` records the properties of the failed
write verbatim). cdkd cannot delete them. Remove them in the destination bucket
yourself (aws s3api list-object-versions, then delete-object --version-id), or
narrow the replication rule so it excludes the prefixes cdkd purges under.
Destination(s): cdkd-state-replica
Four things about that check, all deliberate:
- It needs
s3:GetReplicationConfiguration. The action is in the least-privilege policy, marked as the one DIAGNOSTIC entry: removing it costs you this check and nothing else. Without it the probe is denied, logs atdebug, and never warns — most state buckets are not replicated, so warning there would demand a permission from everyone in order to inform almost nobody. The purge itself is unaffected either way. - It runs when a BODY was purged, or when cdkd could not tell. A key with
no noncurrent version — and a key whose only noncurrent entry was a delete
MARKER, which carries no body — leaves nothing for the replica to be holding,
so neither triggers it; a listing that failed or stopped early does, because
what remains under that key is unknown and silence would be the wrong
mistake. Both exclusions matter:
cdkd deploydeletes the rollback journal on every success, marker and all, so without either exclusion an ordinary green deploy would have announced a surviving journal for a stack that has never had one. - It errs toward warning. A rule that filters on object TAGS cannot be
evaluated without reading each object, so cdkd treats it as covering
everything; and a rule whose
StatusisDisabledis still reported, flagged as disabled, because disabling a rule stops FUTURE replication without removing what it already copied. A warning you look into and dismiss is the cheaper mistake. - It repeats once per (object kind, destination set), not once per object.
A stack with thirty custom resources gets one warning, not thirty; the
repeats go to
debug. The destinations are part of that identity on purpose — if a later purge under the same object kind matches a rule pointing somewhere NEW, you are told about the new replica rather than silenced by the first warning.