cdkd deploy safety internals
The rules behind the guards that cdkd deploy: safety & compatibility flags and its sub-pages summarise. Those pages say what a refusal means and which flag gets through it. This page records the edge cases and the reasoning, for someone changing the code or chasing an unusual result.
When the auto-route needs a recreate
Adding a silent-drop property to an already-deployed provisionedBy: 'sdk'
resource does not by itself require --recreate-via-cc-api: the routing
decision is re-made every deploy, so the next one auto-routes that resource
through Cloud Control. Where the type's SDK-stored physical id is also a valid
Cloud Control identifier, the property reaches AWS as an update in place, with
the physical id preserved. Where it is not, that update fails rather than
silently doing nothing.
This is measured on a live resource by
tests/integration/sdk-to-cc-autoroute/,
which adds EvaluationWindow to a deployed AWS::CloudWatch::Alarm with no
flag and then asserts: the deploy renders the per-resource verb updated and
not replaced, the record moved to 'cc-api', the physical id is unchanged,
the property reads back from DescribeAlarms, and a tag attached out of band
before the redeploy survived. A later phase deliberately recreates the alarm
and checks that tag dies, so its survival above means something.
An earlier opt-out deploy does not defeat this, with one exception. cdkd
records only what the SDK provider actually sent, so after a
--prefer-sdk-route deploy a removable drop is
simply absent from the record: the later flag-less deploy sees a genuine
addition, re-routes the resource, and Cloud Control sends the field. The same
sdk-to-cc-autoroute fixture measures that pair — its --prefer-sdk-route
phase asserts the record does NOT carry EvaluationWindow, and the flag-less
phase after it reads the property back off the live alarm.
The exception is a create-only drop. cdkd keeps such a property IN the
record, so a redeploy that still passes the flag sees no difference and does not
replace the resource. A deploy WITHOUT the flag does see one: AWS never held
the value, and an added create-only property is a replacement. cdkd refuses
that replacement rather than doing it on its own — see
CREATE_ONLY_DROP_NEEDS_REPLACEMENT.
The same sdk-to-cc-autoroute fixture measures it on an AWS::EC2::Subnet's
AvailabilityZoneId.
The flag is for the cases the auto-routed update cannot deliver:
- The property is create-only. No update on either layer can set it, so the
resource has to be created again with the property present. The CFn resource
schema's
createOnlyPropertiesis what to check. Whether you need the flag depends on what the record holds. If it does NOT already claim the value — the ordinary case, where you just added the property — the change reads as an addition, and unless one of cdkd's own rules says that property can change in place it becomes a property-driven replacement, so cdkd recreates the resource for you with no flag. (Where cdkd has no rule of its own, that verdict is read from the type's CFn schema throughcloudformation:DescribeType. Without that permission cdkd warns and reads the verdict from its bundled schema snapshot, for the types it ships one for. For any other type it classifies the change as in-place instead, so that property-driven replacement does not happen — and where nothing rejects the in-place update, the deploy can report success with the property unapplied.) A stateful type is refused until--force-stateful-recreation, unless it declaresUpdateReplacePolicy: Retain— that is exempt from the consent flag, because the old resource is orphaned rather than deleted. Where the record DOES claim the value — the--prefer-sdk-routesequence above — the deploy refuses instead, and--recreate-via-cc-api(or--replace) is how to opt in. - The SDK-created resource's physical id is not a valid Cloud Control
identifier. Cloud Control addresses a resource by the
Identifierits schema'sprimaryIdentifierdefines, while an SDK provider stores whatever its create returned; the two agree per type or they do not, and cdkd treats that as an empirical per-type fact rather than a guarantee (see theSTICKY_CC_MIGRATION_EXEMPTadmission bar insrc/provisioning/provider-registry.ts). Where they disagree, the auto-routed update fails and the recreate is the way through.
Both bullets are reasoned from the routing model, not measured — unlike the
sdk-to-cc-autoroute paragraph that opens this section. They do not fail the
same way, so they are not reached for the same way. The second one fails
loudly: the auto-routed update errors, and the flag is a remedy you reach for
after a failure rather than a precaution you take before one. The first one
fails loudly too: when the record does not already claim the property the
replacement is planned (and refused for a stateful type without
UpdateReplacePolicy: Retain), and when it does the deploy refuses with
CREATE_ONLY_DROP_NEEDS_REPLACEMENT, naming the flag.
Resources that follow a recreated resource
Readers of the target
A recreated resource whose physical id AWS assigns (an
AWS::EC2::SecurityGroup, an AWS::ApiGateway::RestApi) comes back under a
new id, and the resources in the same stack that read it through Ref or
Fn::GetAtt follow it in the same deploy, as they do for a replacement a
property change causes:
- A reader whose referencing property can be updated in place is updated to the new id.
- A reader whose referencing property is create-only (an
AWS::EC2::Volume'sKmsKeyId, say) is replaced, and so are the create-only readers of that reader in turn. The confirmation prompt lists them under "Replaced if the id changes", with DATA LOSS on a stateful one. - A stateful reader that would be replaced needs
--force-stateful-recreationexactly as a stateful target does, and the same flag covers both. Without it, the deploy (and--dry-run) is refused after the diff and before any resource is touched (STATEFUL_REPLACE_BLOCKED), naming the reader. A reader withUpdateReplacePolicy: Retainkeeps its old resource and is not refused, and neither is one a falseConditionremoves from the deploy or one that reads the target only on an untakenFn::Ifarm. The prompt's list is read before conditions are evaluated, so it can name such a reader. - A target whose physical id is its name, fixed as a literal in the template
and equal to the recorded one (a function with an explicit
FunctionName), keeps that id across the recreate. A reader that references it only byRef(or${Target}in anFn::Sub) is neither listed nor refused, and the deploy leaves it as it is. A reader of one of its attributes (Fn::GetAtt,${Target.Attr}) still is: a fixed-name recreate keeps the id but can change an attribute, such as a DynamoDB table'sStreamArnor a database's endpoint. Any other target (an AWS-assigned id, a generated or computed name, a type whose name cdkd rewrites) is treated as one whose id moves. - The list and that early refusal cover a direct chain of create-only
references only. A stateful resource reached through an in-place hop (it
reads a custom resource's
Data, and the custom resource reads the target), through a nested stack'sParametersinto a resource of the child stack, or replaced by the Cloud ControlUnsupportedActionExceptionupdate fallback, is neither listed nor refused early: the replacement guard refuses it mid-deploy, after the target was recreated, and--force-stateful-recreationcovers it there too.
The plan, --dry-run included, counts every reader as an update, since whether
the id moves is known for certain only once the target is recreated. Where the
recreate keeps the id after all, a create-only reader that the plan (and, for a
target not known to keep it, the prompt) shows as a possible replacement is
left in place, and a reader with nothing to change is skipped. Readers in other
stacks are covered by
Other stacks that read the target.
Children deleted with the target
Some resources live inside another one and are deleted with it: a Lambda function's permissions, versions, aliases and event invoke config, an SNS topic's subscriptions and policies, an SQS queue's or S3 bucket's policy, a log group's streams and metric and subscription filters, an IAM role's, user's or group's inline policies.
When a deploy destroys such a parent and re-creates it under the same physical
id (a fixed-name function recreated by --recreate-via-cc-api, or the
delete-first --replace of a replacement), each of those children in the same
stack is re-created too, without a delete: the old one went with the old
parent.
- A child the same deploy moves from another parent onto the recreated one is replaced as usual, which removes its copy from the parent it left.
- A policy that names several parents (a topic or queue policy, an IAM policy on several roles, users or groups) is written again in place instead.
- What is merely attached to a re-created IAM role, user or group (a managed
policy, an instance profile's role, a group's members through a user's
Groupsor aUserToGroupAddition) survives it, but IAM refuses to delete a principal that still has one. cdkd detaches it first, then updates the attached resource in place, which attaches it to the re-created principal again (also when the same deploy edits the list). - A Lambda function URL is not handled this way.
The same holds when the parent is re-created by the update-failure fallback (an
in-place update the provider refuses, re-created under --replace, or
automatically when Cloud Control reports UnsupportedAction): its children are
re-created as soon as it is, before any pending resource that reads them.
If the deploy fails after re-creating the parent and before restoring such a child:
- cdkd drops the child's state record, since AWS no longer has it, and the next deploy creates it again. Until then the parent runs without it.
- A policy that also names parents the deploy did not re-create keeps its record, minus the re-created ones, so the next deploy writes it to them again.
- A child whose own restore was attempted and then failed may be in AWS after
all. It keeps its record when a
--recreate-via-*flag can write it again, and the warning names that flag. Otherwise its record is dropped as above, and the warning says the write may still be on the parent, to be removed by hand if the child leaves the template before the next deploy. - An attached resource's record is never dropped. It is kept minus the
re-created principals (a
UserToGroupAdditionon a re-created group keeps no members), so the next deploy attaches it to them again.
The recreate prompt's three display states
The plan the --recreate-via-* prompt prints shows a target's data in one of
three ways.
Stateful targets — those that reached pre-flight only because
--force-stateful-recreationwas passed — get a**DATA LOSS**prefix on their plan row plus an explicitDATA: all data in <logical id> will be lost (no automatic data migration)line. That is the third stop-and-think moment on top of the two-flag opt-in.The two conditionally stateful types get that prefix too, even though
--force-stateful-recreationskips the emptiness probes entirely: with no probe result to go on, every S3 bucket and every log group in the plan is shown as data-bearing. The plan errs toward warning, because an emptiness nothing measured is not an emptiness.A bucket whose emptiness probe ran and failed is shown as a third case, neither of the two above — a role without
s3:ListBucketVersions, a rate limit that outlived the retries, a region mismatch, or anything else that stops the call from answering. It still proceeds, because the S3 probe fails open by design, but its row says so:- MyBucket (AWS::S3::Bucket) [SDK → CC] — emptiness NOT established: the live probe failed, so cdkd does not know whether this resource holds data UNKNOWN: if MyBucket holds data, the destroy + recreate loses it (no automatic data migration)No
**DATA LOSS**prefix, because cdkd observed no contents and will not assert any; no silence either, which would have made it indistinguishable from a bucket the probe measured and found empty. A bucket AWS reports as not existing is not this case — that is an answer, so it passes through silently like a measured-empty one. A log group cannot reach this case at all: its probe promotes on both failure paths.
Name collisions
Creates that adopt a taken name
Some create APIs do not collide at all: SQS CreateQueue, SNS CreateTopic,
Step Functions CreateStateMachine and ECS CreateCluster return the resource
already holding the name, as do ELBv2 CreateLoadBalancer and
CreateTargetGroup when the settings match, EventBridge PutRule and
CloudWatch PutMetricAlarm overwrite it, and
cdkd's CloudWatch Logs provider reads ResourceAlreadyExistsException as
success, as its S3 provider reads BucketAlreadyOwnedByYou for a generated
bucket name (an explicit BucketName a bucket already holds is refused, except the bucket an earlier attempt of the same create made and could not delete, which its retry takes back while that bucket's name, region and creation date are unchanged). For those types
on cdkd's SDK providers, a replacement that changes the name — or moves an
EventBridge rule to another bus, or changes Type onto one of these types, or,
for an ELBv2 load balancer or target group, sends another name only because
--prefix-user-supplied-names differs from the deploy that created it —
first looks the new name up. When another resource holds it, the deploy fails
with NAMED_REPLACEMENT_COLLISION and nothing is created or deleted — the
create would otherwise take that resource over and record it as the stack's,
for a later cdkd destroy to delete. A lookup that cannot run fails the same
way.
A plain create of one of these types with an explicit name looks the name up
too. When a resource already holds it, or the lookup cannot run, the deploy
fails with NAMED_CREATE_COLLISION and nothing is created, as
CloudFormation's create fails with "already exists". This holds even when the resource is this stack's
own, left by an earlier interrupted deploy or kept by a cdkd destroy under
DeletionPolicy: Retain: nothing in AWS tells the two apart. Delete it, or
adopt it with cdkd import, then re-run; the error ends on the
cdkd import <stack> --resource <logicalId>=<physicalId> command for the
resource it found — confirm the resource is yours before running it. An S3
bucket gets no command, since the lookup also finds a bucket another account
owns and lets you list. When the
holder is this stack's own resource under another logical id (a construct moved
or renamed, keeping its name), the error names that id: give the new resource
another name, or deploy that id's removal first. A log group declared
explicitly that something else already created — for example a Lambda
function's /aws/lambda/<name> group, created on its first invocation — is
refused the same way. A create under a name cdkd generates is not looked up.
An ELBv2 load balancer or target group is looked up under the name the create
sends, which carries the stack-name prefix under --prefix-user-supplied-names.
How --replace proves the old resource holds the name
A collision says a name is taken, not who holds it. An orphan left by an earlier
failed attempt, a retried create, or a resource made outside the stack collides
exactly like the old resource does. So before --replace deletes anything, cdkd
checks that the old resource holds the name the create actually sent:
| The create sent | Shown to be the old resource's when |
|---|---|
| An explicit name from the template | State records that name for the old resource, or its physical id names it |
| No name (cdkd generates one) | The old resource's physical id names the name cdkd generates for this deploy — for types whose generated name cdkd can predict |
| A name the provider rewrites (IAM, ELBv2 prefix the stack name) | The old resource's physical id names the rewritten name this deploy sends |
| A type cdkd has no name property for (Cloud Control only) | Both were created through Cloud Control, the create sent the old physical id itself as a ...Name / ...Identifier property, and every such property it sent matches the old resource's |
A name placed inside a parent (an API's stage, a cluster's service) must also
be in the same parent, and the old resource's state record and its last
read-back must not name it differently (a resource renamed outside cdkd). When the check fails or cannot decide, the deploy fails
with NAMED_REPLACEMENT_COLLISION and nothing is deleted — with or without
--replace, and without advising --replace, which would refuse the same way.
Remove or rename whatever holds the name if it is yours — if that is the
resource being replaced, delete it by hand — then re-run the deploy.
UpdateReplacePolicy: Retain hard-fails in both shapes regardless of
--replace: with Retain the old resource keeps the name, so a same-name
replacement can never proceed. The same two shapes, and the same two error
codes, are reachable from the update-failure fallback — under Retain that
path also creates first, so it inherits the same constraint.
Equal physical ids across a type change
Two physical ids that happen to be equal across the two types — an SSM parameter
and a log group can share a bare name — are treated as two resources. The
exception is a change between two custom resource types (Custom::*,
AWS::CloudFormation::CustomResource), whose handler picks the id: there an
equal id is the existing resource, and the replacement is refused with
NAMED_REPLACEMENT_IDEMPOTENT_CREATE rather than deleting what it just created
(--replace deletes the old resource first instead, as for any name-idempotent
create).
Within one type an equal id is the same resource, with one exception: two
AWS::Glue::Table records in DIFFERENT databases can share an id when either
name contains | (table db|orders in database my, table orders in
database my|db). When both records' DatabaseName prefix the shared id and
differ, cdkd treats them as two tables, so such a replacement keeps the new
table and deletes the old one through its own record, and a rollback reverses
it.
When the new resource's create collides on a name instead, the error says that
the holder may be an unrelated resource of the new type, which --replace
cannot free. Under --replace, cdkd deletes the old resource first only
between two types that share one name space (RDS, DocumentDB and Neptune
clusters, instances and subnet groups; DynamoDB tables and global tables);
any other pair fails with nothing deleted.
cdkd rollback and the automatic rollback reverse such a replacement the same
way: the old resource is re-created through its own type's provider and the new
one deleted through its own. A rollback journal written by an older cdkd names
only the new type on the operation, so the old type is read from the previous
resource record the journal also carries. When a journal names no old type at
all, or names two different ones, that one operation is refused with
ROLLBACK_REPLACEMENT_UNROUTABLE and the journal is kept; fix forward with
cdkd deploy, or pass --orphan <LogicalId> to leave the resource as it is and
let the rest of the rollback proceed.
Deletion protection: what each refusal reads
Deletion protection blocks a replacement and cdkd deploy cannot clear it.
Two kinds of refusal say so explicitly:
- The stateful-resource guard's refusals on
--recreate-via-*,--replace, the Cloud Control auto-fallback and a property-driven replacement add a note for any stateful type in--remove-protection's table whose protection flag is on. The note names the flag, both outcomes (a failed delete-first deploy, an untracked old resource after a create-first one) and the public section, and leaves turning the flag off to the console or the service's own API. A target underUpdateReplacePolicy: Retaingets no note. - Some types' own refusals read their own protection property and name the
command that turns it off:
AWS::Logs::LogGroup,AWS::ElasticLoadBalancingV2::LoadBalancer,AWS::EMR::Cluster,AWS::Cognito::UserPool,AWS::DynamoDB::GlobalTableandAWS::AutoScaling::AutoScalingGroup. They point at the console instead when the resource's id cannot be printed safely on a command line: it would be changed by sanitizing, it holds whitespace or a character a shell acts on (a quote, a backtick), or something the AWS CLI itself acts on (a leadingfile://or-).
The guard's note does not ask AWS: it reads the properties cdkd recorded and
the AWS read-back it stored after its last write. Protection you enabled out of
band after that read is in neither, so such a resource gets the guard's shorter
message. The six types' own refusals do not read the stored read-back at all:
they read the properties cdkd recorded (AWS::Cognito::UserPool reads the
template's value first), so protection that only the read-back shows gets their
shorter message. AWS refuses the delete either way, for every type in
--remove-protection's table whenever a deploy has to replace one, whether or
not its refusal says so.
AWS::AutoScaling::AutoScalingGroup is the one whose refusal is narrower than
the type's protection setting, and deliberately: the group's three levels are
none, prevent-force-deletion and prevent-all-deletion, and only the last
blocks a replacement, because the deploy path's delete does not pass
ForceDelete. At prevent-force-deletion there is nothing to disable.
AWS::Cognito::UserPool is the one exception to the two-deploys rule.
Its refusal fires AFTER UpdateUserPool has already applied the template's
DeletionProtection, so clearing it in the template DOES take effect in that
same (failed) deploy — the next run with the replace flags then succeeds. Its
refusal reads the desired value first for exactly that reason. Do not
generalize it: every other type above refuses before applying anything.
How the stateful list is kept complete
The list has mechanical lower bounds cdkd enforces in unit tests, so it is checked rather than only hand-curated.
- Every type cdkd takes a final snapshot of before a destroy is on it. cdkd
snapshots the types CloudFormation lets you tag
DeletionPolicy: Snapshot, and CloudFormation permits that attribute exactly where deleting the resource destroys data worth capturing first — so a type cdkd snapshots on destroy must not be replaceable mid-deploy without consent.AWS::Redshift::Cluster,AWS::ElastiCache::CacheClusterandAWS::ElastiCache::ReplicationGroupjoined the guard for that reason. - Every type whose delete consumes the
--force-stateful-recreationconsent is on it. A resource whose deletion needs that flag to clear its own data guard is by definition data-bearing.AWS::S3Express::DirectoryBucketjoined for that reason. - Every type a tier-2 sweep proposes is either on it or written off with a
reason. The two bounds above are derived from cdkd's own SDK providers, so
neither can see the 1371 CloudFormation types that have no provider and route
a replacement through Cloud Control API — the larger population by two orders
of magnitude, and the one the guard was reaching with no flag at all. cdkd now
reads every one of their registry schemas and proposes the types that declare
an immutable property (so a rename replaces the resource on a plain
cdkd deploy) and look like they store something. Each proposal must end up on this list or be written off in the sweep's own file with the reason; a proposal in neither fails the build. Most of the always-stateful table joined this way.
Unlike the first two, that third bound is a heuristic: no AWS-published
artifact says "deleting this destroys user data", so what it buys is that the
next widening is checkable, not that the current list is complete. Where it
proposes a type whose answer is not knowable from the outside, the type is
guarded — an unprovable emptiness must not read as empty. The write-offs are
the cases where the schema settles it: AWS::RDS::GlobalCluster,
AWS::Neptune::GlobalCluster and AWS::DocDB::GlobalCluster group regional
clusters that outlive them; AWS::RedshiftServerless::Workgroup is compute
against a namespace that is guarded; AWS::Amplify::Domain,
AWS::Cognito::UserPoolDomain and AWS::Lightsail::Domain are DNS rather than
stores; AWS::EC2::TransitGatewayRouteTable and its siblings hold nothing but
tags, their routes being separate template resources. The full list of
write-offs, each with its reason, is in
scripts/audit-stateful-candidates.ts, and the proposals themselves — with the
immutable properties that make each one reachable — in
docs/_generated/stateful-candidates.md.
The rest are hand-curated, because no lower bound can see a delete that
destroys data with no opt-in at all: AWS::S3Tables::TableBucket and
AWS::S3Vectors::VectorBucket empty themselves first, AWS::S3Tables::Table
holds the rows themselves rather than a catalog entry, AWS::KMS::Key
schedules the key material for deletion (AWS::KMS::ReplicaKey is guarded on
the same footing, but routes through Cloud Control and is unmeasured here), and
AWS::CodeCommit::Repository drops the git history. AWS::KMS::Alias is
deliberately not guarded — deleting an alias removes a pointer, not key
material.
AWS::EC2::Instance and AWS::SQS::Queue are deliberately not guarded either.
An instance's root volume is ephemeral by design — keep persistent data on an
AWS::EC2::Volume, which is guarded — and a queue's backlog is transient.
Asking for --force-stateful-recreation on every AMI refresh or queue rename
would cost a dev/test workflow more than it protects, and CloudFormation
replaces both without asking as well.
AWS::S3Tables::Namespace is deliberately not guarded: AWS refuses to delete a namespace that still holds a
table, answering BadRequestException: The namespace that you tried to delete is not empty., and the Cloud Control delete fails the same way. cdkd's own
delete for a namespace issues a bare DeleteNamespace and enumerates no
tables, so a namespace rename (a replacement a plain cdkd deploy reaches with
no flag) cannot take a table with it. The replacement creates the new
namespace first; if the old one still holds tables, its delete is refused and
cdkd warns Failed to delete old resource and carries on. The old namespace
and its tables stay in AWS, no longer tracked in state, for you to move or
delete by hand. A delete-first replacement (--recreate-via-*, or --replace
when the create collides) fails at that delete instead — unless the template
also renames the resource, which makes --recreate-via-* create first and warn
the same way.
Why each category is on the list
The public table lists the types by category without comment. These are the reasons, for the categories where the data is not obvious:
- Data warehouse. An
AWS::RedshiftServerless::Namespaceowns the databases, and a snapshot is a copy of them. - Table / vector storage. Deleting a table bucket or a vector bucket empties it first, with no opt-in. A namespace is not guarded, as described above.
- Managed compute with local storage. Terminating an
AWS::EMR::Clusterdestroys the HDFS volumes on its core nodes, and the replacement comes back empty.AWS::EKS::Clusteris guarded for the same reason one level up: the etcd store behind it holds every Kubernetes object the user created, and nothing in the template describes it.AWS::SageMaker::Clustercarries local and tiered storage holding training checkpoints. - Streaming / messaging. Each type retains records on its own storage rather than passing them straight through.
- Search / index / collection. The indexed documents, face vectors and geofences are written through the service API, never from the template.
- Identity / config. An
AWS::AppConfig::ConfigurationProfilewhose location ishostedowns its configuration versions. - Runtime-written stores.
AWS::CloudFront::KeyValueStoreandAWS::Connect::DataTableare seeded from the template at most once, then written through the service API. - Backup vaults. A vault holds the recovery points, the data whose whole purpose is to outlive the resource it was taken from.
- Service domains holding records. These hold cases, profiles, a catalog and every user's home directory.
- Encryption keys. Deleting an
AWS::CloudHSM::Clusterdestroys the key material inside it, and every ciphertext produced under those keys with it. AnAWS::KMS::Keydelete schedules the key for deletion, and once the window elapses every ciphertext encrypted under it is unrecoverable, including data in other stacks that merely reference the key.AWS::KMS::ReplicaKeyis guarded on the same terms, though whether a destroyed replica's ciphertexts survive through another key in its multi-region set is unmeasured, so the guard assumes they do not. - Source control / artifacts. An
AWS::CodeCommit::Repositorydelete destroys the repository's entire git history.AWS::CodeArtifact::Repositoryholds the packages, andAWS::CodeArtifact::Domainis not a mere grouping: it owns the deduplicated asset storage every repository in it references. - Retained records.
AWS::IoTSiteWise::Workspaceis guarded on an open question: AWS makes encryption at rest required on it, but whether deleting one cascades to the datasets inside is unmeasured.AWS::AIOps::InvestigationGroupandAWS::SES::MailManagerArchiveboth retain content for a configured period.AWS::Rbin::Rulejoins them on the fail-safe side of an open question: the rule itself is fully template-declared, but what happens to the snapshots and AMIs already sitting in the Recycle Bin under it when it is deleted is unmeasured. - Nested stacks. Replacing an
AWS::CloudFormation::Stackdestroys the whole child stack, every resource it owns, with no per-resource guard. ItsStackNameis immutable, so aStackNameedit deployed with--prefer-sdk-route AWS::CloudFormation::Stack:StackNameis a replacement. - Edge / identifier immutability. For an
AWS::CloudFront::Distributionthe URL changes, which breaks consumers, and propagation takes roughly 20 minutes.AWS::SMSVOICE::PhoneNumberandAWS::SMSVOICE::SenderIdare the same class: a release returns the identifier to the pool, the replacement gets a different one, and the original may be unobtainable.
Log group retention
Retention is not an emptiness signal. An unset or zero RetentionInDays
is CloudWatch Logs' never expire setting — the most data-bearing
configuration the type has, and the one cdkd records as 0. An unset retention
therefore defers: at pre-flight the emptiness probe below decides, and
mid-deploy, where no probe can run, the log group counts as stateful.
Which retention cdkd reads. A state record carries two property bags —
properties, what the last deploy applied, and observedProperties, what it
read back from AWS — and a positive RetentionInDays in EITHER settles the
guard. Neither bag takes precedence; either one proving a retention is
enough. That is what lets a retention set OUT OF BAND (the console,
aws logs put-retention-policy) count, and a record imported by
cdkd import --migrate-from-cloudformation whose template never declared the
property. The value is COERCED rather than type-tested, so the stringly-typed
RetentionInDays: '30' a hand-written L1 or an Fn::Sub result produces
counts as 30 rather than as no retention. A zero recorded in one bag never
cancels a positive recorded in the other: zero is never-expire, which is not a
statement that the group is empty.
The guard's coercion is JavaScript's Number(), filtered to finite values.
That is wider than CloudFormation's own Integer parsing — Number() also
accepts '0x1e', '0o36', '1e3' and '30.5' — and in the GUARD the
difference is safe in the only direction that matters: a wider accepted set
can only produce MORE has-retention verdicts, i.e. more refusals.
The provider is the half that forwards the number to
logs:PutRetentionPolicy, and it reads the property
the way CloudFormation does, measured rather than assumed (live A/B on
AWS::Logs::LogGroup, us-east-1, 2026-09-14): an optional sign and decimal
digits, with surrounding whitespace trimmed — 30, '30', '+30' and
' 30 ' all deploy as 30 — while '0x1e', '1e3', '30.5' and '30.0'
are REFUSED before any AWS call, as CloudFormation refuses them. The falsy
family was measured in the same pass: an absent property, an
empty string and a whitespace-only string are CloudFormation's spellings of
"no retention" and remove the live policy on an update, while 0, '0',
false and null are rejected by CloudFormation and are refused by cdkd
rather than silently removing a retention you set on purpose. The one
cdkd-side exception is cdkd drift --revert, where a numeric 0 is cdkd's
own readback spelling of a never-expiring log group and reverts a
console-added retention as expected.
The emptiness probes
At pre-flight (--recreate-via-cc-api / --recreate-via-sdk-provider)
cdkd issues a single-page s3:ListObjectVersions(MaxKeys=1) against each
targeted bucket's recorded physical id, in the stack's deploy region. Empty
buckets pass through; non-empty ones are refused. cdkd uses
ListObjectVersions rather than ListObjectsV2 so the probe's view of "empty"
matches what the destroy-and-recreate cycle would actually wipe — a versioned
bucket whose current keys are all soft-deleted still holds prior versions and
delete-markers.
A page carrying a continuation marker with no entry in either list does not settle the question — the listing is unfinished, so that page's emptiness is not the bucket's — and such a bucket is refused rather than passed. A page that simply OMITS the version and delete-marker lists is different, and does count as empty: S3 omits an empty collection rather than sending an empty list, so omission is how an empty bucket answers.
Both emptiness probes retry a throttling response — a throttling error code, or HTTP 429 / 503 — up to three times with exponential backoff, at most 3.5 seconds per target. Every other failure goes straight to the per-type behaviour described below, because it is either an answer or something an identical retry will not change. When the retries are exhausted the probe lands in that same per-type behaviour.
If the probe itself fails — permission denied, bucket not found mid-flight, a
transient network error — cdkd logs a warning and leaves the target
un-promoted, which means the guard does not fire and the recreate proceeds
without --force-stateful-recreation. The probe fails open, so treat that
warning as a prompt to decide for yourself: pass
--force-stateful-recreation if the bucket might hold data.
For a log group, the same pre-flight issues a single-page
logs:DescribeLogStreams(limit=1) against the recorded log group name. A log
group with no log stream can hold no log event — every event belongs to a
stream — so zero streams is the one signal that proves the group empty, and
cdkd uses it rather than a byte count: LogStream.storedBytes has been
reported as zero by the API since June 2019, and stream presence needs no size
semantics at all. A group holding only empty streams therefore counts as
non-empty, which is the safe direction.
Only one answer clears the guard: a log group whose response carries a
present, empty stream list and no continuation token. A response with
no stream list at all, or an empty page that still carries a nextToken, has
not settled the question, so cdkd warns and treats the group as stateful — the
same direction a failed probe takes.
Unlike the bucket probe, the log-group probe fails CLOSED: if
DescribeLogStreams errors, cdkd warns and treats the log group as stateful,
so you get a refusal naming --force-stateful-recreation rather than a silent
recreate. The asymmetry is deliberate — the whole point of the log-group
condition is that an emptiness cdkd cannot prove must not read as empty.
One error is the exception, because it is an answer rather than a failure to
get one: a ResourceNotFoundException means AWS says the log group does not
exist, so it provably holds no events and the guard is cleared. cdkd trusts
that only after confirming the CloudWatch Logs client is pointing at the region
cdkd's state records for the resource — a not-found from the wrong region says
nothing about the log group. When that check cannot be satisfied, the group is
treated as not provably empty like any other unsettled answer.
--pin-cc-api: why an unknown id is an error
A logical id present in no stack of the run is an error, not a no-op. The
flag produces no output when it works, so a typo would otherwise give you
exactly the routing change you passed it to decline, indistinguishable from
success. The check is run-level rather than per-stack on purpose: under
--all, an id that belongs to one stack is legitimately absent from the
others, and failing per-stack would abort the run over a correct invocation.
It is raised before any stack deploys.
Fn::GetAtt attributes read live
An attribute cdkd reads live never falls back to the physical ID. A
few attributes are read from AWS at resolution time when the state record
does not hold them: an EC2 instance's PrivateIp / PublicIp /
PrivateDnsName / PublicDnsName / AvailabilityZone, a VPC's
DefaultSecurityGroup, a CloudFront distribution's DomainName, a
security group's VpcId (recorded at create from DescribeSecurityGroups,
so a group declared without VpcId resolves to the default VPC's id as
CloudFormation answers, and re-read when the record lacks it — or holds
'', which an older cdkd wrote for such a group; that record keeps ''
until the group's next update, only the resolution changes — so on such a
record cdkd diff issues one DescribeSecurityGroups and shows a one-time
'' → vpc-… Output delta, which --fail exits 1 on once and the next
deploy's Outputs persist heals; if the read is refused, the Output lands in
the failed keys and the Outputs section is suppressed with a warning). When
the read finds the value not yet assigned (an instance still pending
under --no-wait) or the read fails, cdkd refuses to resolve the reference
rather than substituting the instance / VPC / distribution / group ID,
which can never be the right value there; the message names the resource, the
attribute, what was observed (the instance state, or the error class —
--verbose shows the AWS text) and the remedy. Nothing is cached, so the
next deploy re-reads. An RDS DBProxy / DBProxyEndpoint VpcId the
record lacks is refused the same way without a live read. The
--no-wait section of cdkd deploy
says what the refusal does in a resource property versus an Output.
Cloud assembly checks
The nested-template walk
The check runs for every stack in the deploy set, including stacks pulled in as
dependencies, and before anything else happens to any of them: no macro is
expanded, no asset is published, no lock is taken, no state record is written
and no resource of the parent stack is created. One malformed stack stops the
whole run, --dry-run included, with exit code 1. A stack outside the deploy
set is not checked.
The whole tree is checked up front, however deep the cycle sits. Each nested level is a real deployment, which is why the check does not wait until the walk reaches the repeat.
The nested-stack provider repeats the same walk on the subtree under its own
row before it deploys the first nested level. With the check above in place that
second walk finds nothing; it is what still refuses the tree if a nested deploy
is ever started some other way, and its message begins with
NestedStackProvider: and ends with Refusing to deploy any level of it.
Asset paths
The two asset rows of the path table (a file asset's source.path, a Docker
asset's source.directory) are measured against the app's output directory,
not the manifest's own. A Stage's assets are staged into the app's cdk.out while
the Stage's asset manifest sits in cdk.out/assembly-<Stage>/, so CDK writes
source.path: "../asset.<hash>" there by design — and a nested Stage reaches
up further still. Those paths load normally. What is refused is a relative
path leaving the output directory, from a Stage manifest and a top-level one
alike.
An ABSOLUTE source.path or source.directory is accepted, and nothing
refuses it. cdk synth --no-staging emits exactly that shape — under
aws:cdk:disable-asset-staging CDK writes each asset's absolute SOURCE
directory instead of a staged copy — and the CDK CLI publishes such a path
without any containment check of its own, so refusing would reject the output
of a documented flag and be stricter than the tool cdkd complements.
For an absolute path the warning is the entire protection. cdkd prints one when the path falls outside the output directory, naming the directory and naming where the bytes go — the destination bucket and key, or the image build. There is no second gate behind it. A Cloud Assembly you did not synthesize can name any directory your user account can read, and cdkd will package it and upload it to a bucket that same manifest names, using your credentials.
A value naming the output directory itself is warned about "by any spelling",
and that is meant literally: a symbolic link to the output directory,
or the path a realpath would print for it, is that directory, and cdkd says
so for all of them.
The relative refusal catches an accidental or legacy .. and costs nothing,
which is why it stays. It is not a boundary against a value someone chose: the
absolute spelling of the same path is accepted with a warning.
The cdkd local * commands apply the same rule to SOME of the assembly they
read — including a path the deploy side has no equivalent of, a Lambda's
Handler for an inline Code.ZipFile, which cdkd materializes as a file
before running it. cdkd local invoke, cdkd local start-api,
cdkd local start-alb and cdkd local start-cloudfront — every command that
bind-mounts Lambda code — refuse a relative escaping aws:asset:path
before mounting it into the container, and accept an absolute one with the
same warning, for the same reason. Other assembly-supplied paths are not
covered yet; Local Execution states the trade and lists
which are which.
Why a pre-synthesized assembly is not refused
Each distinct command is announced once per run, not once per build: a
long-running cdkd local start-service rebuilds per replica and again after a
crash-loop restart, and repeating a paragraph that size would bury it. Repeats
go to --verbose.
The destination check is a name-shape check, not a proof of ownership: a bucket named like a CDK bootstrap bucket for your account can still live in someone else's. It narrows what a careless manifest gets away with, nothing more.
Pointing -a at an assembly you did not produce is the same decision as running
someone else's build output. cdkd cannot make that decision for you: anyone who
can rewrite a manifest can equally rewrite the Dockerfile, the Lambda asset and
the template, so a refusal here would stop nothing while breaking the
split-synth/deploy pipelines that are the normal shape. Synthesize it yourself,
or read it first.
The nested-stack walk separately refuses an absolute aws:asset:path, which
is a "not CDK-generated" tripwire rather than an escape. It is a different
question from the asset rows above and keeps a different answer: there the value
is a nested-stack TEMPLATE that CDK always writes into the output directory, so
--no-staging does not relocate it and an absolute one means the assembly was
not CDK-generated. Its message says which is absolute where the containment
one says which resolves to ..., so the two are told apart at a glance. Also refused is a tree nested more than 512 levels
deep, which could not deploy anyway: each level lengthens the child's state key,
and S3 caps a key at 1024 bytes. A tree with more than 10,000 nested-stack rows to
follow is refused as well. That is far beyond any CDK-generated assembly; in
practice it takes symlinked directories, which give one template file many
paths.