Provider implementation rules
The rules an SDK provider follows, and why each exists. Most were written after a defect got past review, so each carries the shape of the failure it prevents rather than only the rule.
This page is for someone writing or reviewing a provider. Start from Provider Development, which walks the interface and the steps end to end; come here for the decisions each step implies.
Error handling
- Wrap all AWS SDK calls in try-catch
- Use
ProvisioningErrorto provide detailed context
try {
await this.client.send(new CreateXxxCommand({ ... }));
} catch (error) {
throw new ProvisioningError(
`Failed to create ${logicalId}: ${String(error)}`,
resourceType,
logicalId,
physicalId,
error instanceof Error ? error : undefined
);
}
Never infer a default from a possibly-malformed value
Reading a string out of a nested config block with || — or with ?? — looks
harmless and is not:
// WRONG — a string / array / unresolved intrinsic container indexes to
// `undefined`, and the `||` silently substitutes the OPPOSITE of the
// declared intent.
const status = (versioningConfig['Status'] as string) || 'Suspended';
// EQUALLY WRONG — `??` defaults on exactly the same `undefined` (issue #1493).
const type = (source['Type'] as string) ?? 'NO_SOURCE';
VersioningConfiguration: 'Enabled' (a hand-written L1 template, or an
intrinsic the resolver could not resolve) therefore turned versioning off
on a live bucket, with no error anywhere. Use the shared guards instead:
import { readConfigString, requireConfigString } from '../config-shape.js';
// nested container, which may itself be malformed
const status = readConfigString(
versioningConfig,
'Status',
'Suspended',
'AWS::S3::Bucket VersioningConfiguration'
);
// top-level field — keep the `properties['X']` read at the call site so the
// handled-property-wiring critic can still see the property is consumed
const scope = requireConfigString(properties['Scope'], 'REGIONAL', 'AWS::WAFv2::WebACL Scope');
An absent container and an absent key still take the default ({} legitimately
means "defaulted"); a container that is present but not an object, and a key
that is present but not a non-blank string, are refused by name.
A container you never read a string out of needs its own guard. The two
above can only fire while reading a FIELD, so a block whose members you merely
probe for PRESENCE — or hand to .map — slips past both, and a malformed value
there reads as an EMPTY block rather than as an error:
import { requireConfigArray, requireConfigObject } from '../config-shape.js';
// LIST block: a truthy non-array reaches `.map` and dies with a raw TypeError,
// a falsy one is silently dropped by the truthiness gate in front of it.
const tagFilters = requireConfigArray(raw, 'AWS::S3::Bucket …TagFilters');
// OBJECT block: every probe of a malformed container indexes to `undefined`,
// so the block reads as empty and the caller proceeds WITHOUT it — an S3
// lifecycle rule losing its whole location scope and applying bucket-wide.
const filter = requireConfigObject(raw, 'AWS::S3::Bucket …Rules[].Filter');
Both leave the ABSENT case to you (raw != null && …), because an omitted
block legitimately means "no entries" / "defaulted". Both also take the
onUnusable downgrade below, and when you pass it they return undefined
instead of throwing — you then decide the skip UNIT, and that decision is
the whole point: skipping must never be the misbehavior you were refusing.
S3's per-item Puts skip the single configuration item, while its lifecycle Put
skips the WHOLE configuration, because that call replaces every rule and
applying the valid siblings alone would DELETE the malformed one from AWS.
A per-ITEM STRING read may need that same skip, not readConfigString's
default. readConfigString's onUnusable downgrade is warn-and-DEFAULT,
which is right for a field whose default is inert — and wrong wherever the
default is applied to a LIVE resource and turns something ON. All four of S3's
per-item reads are that shape (issue
#1595): defaulting a replayed
lifecycle / intelligent-tiering / replication Status to Enabled starts
expiring, archiving or replicating objects for a rule the template had
DISABLED. Ask the question without taking the answer:
import { configStringRefusal } from '../config-shape.js';
// `undefined` when the read would have succeeded; otherwise the refusal
// SENTENCE, with no action clause — you supply the one that is true here.
const refusal = configStringRefusal(rule, 'Status', 'Enabled', '…Rules[]');
if (onUnusable && refusal !== undefined) {
// Keep the clause PATH-NEUTRAL. The replay-CREATE arm reaches this too
// (a reverse-replacement revives the resource), so a message asserting a
// "live" configuration would be false there.
onUnusable(`${refusal}. Leaving the whole configuration unapplied here; …`);
return; // ...or `continue`, per the skip UNIT this API implies
}
On the create path you do not probe at all, so the original read still refuses
exactly as before. Use this helper rather than a hand-written typeof check:
it shares requireConfigString's predicate, and a twin disagrees with the read
it fronts on precisely the values that matter — a blank string, an explicit
null, a coerced number.
Pass the desired side only. previousProperties comes from cdkd state
rather than the user's template, so refusing a malformed value recorded there
by an older binary would make the stack undeployable with no way out short of
hand-editing the state file.
A list whose removals are the gap between the two sides is the exception
(issue #3948): Auto Scaling
group attachment and entry lists (tags, metrics, lifecycle hooks, traffic
sources, notifications), Firehose tags, an ELBv2 target group's Targets,
Budgets NotificationsWithSubscribers / ResourceTags, and CodeCommit
Triggers / Tags (issue #3989),
and most providers' CloudFormation Tags diffs, read through
src/provisioning/tag-list.ts (issue #3994;
Glue and log groups: #4073).
There, reading a malformed
DESIRED side as empty detaches or deletes everything the record holds, so it
is refused before any call on every path, a state replay included. A malformed
RECORDED side only misses removals: apply it ADD-only where you can — read
the live resource (keep the live entries the desired side names, warn about
the rest), or, where every add is idempotent (an upsert, or a create
whose duplicate is success), send only the adds and warn — otherwise refuse it with a repair that never asks for a secret in
state.json. undefined / null stays the empty list.
A top-level read takes two further decisions, both per site (issue
#1513), expressed as options on
requireConfigString:
// CFn coerces scalars and cdkd does not, so an unquoted YAML `IpProtocol: -1`
// arrives as a NUMBER and deploys fine today — refusing it would break a
// working template. Only for genuinely numeric-looking fields; a number where
// an enum belongs (`InstanceType`) stays a refusal.
//
// The real site wraps this call in `narrowIngressIpProtocol` so the DIFF can
// resolve the protocol the same way — see the `effectiveProperties` section
// above for why a coerced value has to be recorded as well as sent (#1633).
const ipProtocol = requireConfigString(
properties['IpProtocol'],
'-1',
'AWS::EC2::SecurityGroupIngress IpProtocol',
{ coerceNumber: true }
);
// UPDATE-path sites WARN instead of throwing when the desired bag is a state
// record (a rollback revert or `cdkd drift --revert`) — see [Pre-flight refusal](#pre-flight-refusal-when-a-provider-may-reject-what-cloudformation-forwards):
// a refusal there can leave the resource un-rollbackable with no template-side
// remedy. A template-path update refuses the same value, before any call.
// (`context` is `update()`'s `UpdateContext`; `previousStatus` is the recorded
// status, the warn fallback, so a warning never ENABLES a key.)
const stateBorneDesired =
context?.replayingState === true || context?.desiredFromAwsReadback === true;
const status = requireConfigString(
properties['Status'],
previousStatus,
'AWS::IAM::AccessKey Status',
stateBorneDesired ? { onUnusable: (message) => this.logger.warn(message) } : {}
);
And check WHERE the read lives before guarding it: a helper the delete() or
diff paths also reach is fed state-borne values, so guard the create CALL SITE
instead (EC2Provider.buildIpPermission is the tree's example — it is shared
with deleteSecurityGroupIngress and with the revoke half of the inline-rule
update diff).
Pre-flight refusal: when a provider may reject what CloudFormation forwards
cdkd's compatibility target is CloudFormation, so the default for a property cdkd cannot handle is to forward it and let AWS answer, never to invent a validation CloudFormation does not have. A provider that refuses a template CloudFormation would accept is a parity break, and parity breaks are how a tool that claims template compatibility stops being trustworthy.
There is one narrow exception, and it has a high bar. A provider MAY refuse a property before the AWS call when all of the following hold:
- The property is undeployable on cdkd's OWN path, proven by a live probe against the real API the provider calls — not inferred from CloudFormation also rejecting it. This is the load-bearing condition: "CFn rejects it too" is not sufficient, because cdkd calls the service API directly and could in principle succeed where a CFn resource handler fails.
- No shape of it works, so no user loses a working deployment. If some combination deploys, forward it.
- The refusal names the working alternative, concretely enough to copy. A refusal that only says "no" is worse than AWS's own error.
- The rationale is recorded at the check site, as a comment naming the probe (date, region, what was tried, what AWS said) and flagging it as a deliberate parity divergence — so the next reader can re-evaluate it when AWS changes.
GlueProvider's enforceIcebergTableInputAbsent (issue
#1454) is the reference
implementation. Three details are worth copying:
The check runs before the
tryblock, so the typedProvisioningErroris not caught and re-labelled by the provider's own error wrapper. A test that only asserts message CONTENT cannot catch this being moved inside thetry, because the wrapper EMBEDS the original message — assert the raw message PREFIX and an absentcauseinstead.Refuse a TEMPLATE-driven call; WARN on a STATE REPLAY. This asymmetry is the rule, not a Glue quirk, and the axis is the ORIGIN of the properties, not the operation name.
cdkd rollbackreplays from cdkd STATE, not from the template. A resource whose state record already carries the offending value (written by an older cdkd build, or by an import) would become not just un-updatable but UN-RESTORABLE, and unlike the template case the user has no remedy at all — only hand-editingstate.json. Concretely:update()— refuse on the template path, warn on a replay. The rollback executor's two revert arms callprovider.update(..., op.previousState.properties, ...)withUpdateContext.replayingState, andcdkd drift --revertsetsdesiredFromAwsReadback; the deploy engine sets neither. Refuse only when both are unset AND the refused value is template-borne on that path — a value the recorded resource already carries, such as an identity segment that only a replacement could change, keeps the warning, with the reason stated at the site (issue #3728). Refuse BEFORE the first write: a throw from a mid-update arm strands whatever the earlier arms applied (a read such as a hosted-zone lookup may precede it). Where the arm itself runs mid-update, ask its OWN predicate from a pre-flight rather than restating it — S3's per-config appliers run on a probe whose client writes nothing (issue #3740). Gate the refusal on the value having changed wherever an unchanged one sends nothing, and do NOT gate it where the value goes out on every update: Glue'sUpdateDatabasereplacesDatabaseInputwholesale, so its malformed blocks are refused whether or not they changed (#3740). A list whose removals are the gap between the two sides refuses on a replay too (see "A list whose removals are the gap between the two sides is the exception" above, #3948): warning and reading it as empty would remove everything the record holds. Separately, a create-only value such asAWS::RDS::DBProxyTargetGroupTargetGroupNameorAWS::Lambda::EventInvokeConfigQualifierkeeps the warning on purpose, for the reason above. On the warn path, pick the FALLBACK per site (issue #1551): warning and then applying the CREATE DEFAULT is frequently worse than the refusal was, because the default lands on a LIVE resource — it flipped an IAM-guarded Lambda function URL to PUBLIC, re-pointed a live DynamoDB stream, and read a malformed GSI block as "delete every index". Keep the PREVIOUS value where one exists (omitting the field entirely when the API has merge semantics), otherwise SKIP the block or SUPPRESS that part of the diff. Whichever you choose, remember the warn path now records the unusable value as the new state, so seed the next comparison from AWS's live value where the provider already holds it (issue #1552).create()— refuse, unlessCreateContext.replayingStateis set.create()is a replay path only through the rollback executor's reverse-replacement arm, which revives the OLD resource frompreviousState.properties(issue #1199). That arm — and nothing else in cdkd — passes the optional 4th parametercontext?: CreateContextwithreplayingState: true(issue #1463). The deploy engine's six create sites (CREATE, the property-driven replacement, the--recreate-via-*destroy-then-create, the--replacedelete-first fallback, the update-failure replacement, andcreateFirstThenDeleteOld, which--recreate-via-*and the update-failure replacement take for a renamed resource) are driven by freshly resolved TEMPLATE properties and never setreplayingState, so the refusal stands where the user can actually act on it. (They DO pass a context — every provider call now carriesmaskSecrets, see themaskSecretsbullet below — so the test fences for this read the context's key set rather than the call's arity.)Do not re-create inside
update()if you have a create-side refusal. Several providers callthis.create(...)from their ownupdate()(ACM certificate, IAM managed policy, IAM role, Lambda permission, SNS subscription). Those internal re-creates never receivereplayingState—update()'s own context is anUpdateContext, and the most any of the five builds from it is aCreateContextcarrying onlymaskSecrets(IAM role, IAM managed policy and Lambda permission, issue #2177) — and thepropertiesthey forward ARE a state record during a rollback replay. So the refusal would fire on a replay with no way to detect it. None of those providers has a pre-flight refusal today (they validate required fields only, which correctly stays a hard error), so there is no live gap; this is a constraint on the NEXT provider, not a description of the current tree. Since issue #3141UpdateContextcarries its ownreplayingState— set by the rollback executor's two revert arms — so the next provider CAN build aCreateContextfrom it instead of relying on the constraint.Retire what your create materialized before the error leaves. A create shaped
<one API call that materializes the resource>then<a wait for it to become usable>can fail with the resource already alive in AWS. Before issue #2169 anAWS::CertificateManager::Certificatein that state was simply LOST: the create never returned, so nothing was written to state, and it was invisible tocdkd state show, unreachable bycdkd destroy, and re-requested by the nextcdkd deploy— an orphan per attempt. Delete it in your owncatch, best-effort, and re-throw the ORIGINAL error.Three points that decide whether you get this right:
- Track the id from the API response, not from the name you computed.
Set the local only once AWS has confirmed the resource exists, so a
failure BEFORE the call deletes nothing. Do NOT reuse
ProvisioningError.physicalIdas this signal — 462 call sites across 75 provider files pass one, and on a create path it is usually the intended name computed beforehand (IAMRoleProvider.createhands its catch theroleNameit derived, whether or notCreateRoleever ran). - Never let the cleanup replace the diagnosis. The create's own error is what the user needs; a cleanup that FAILED appends a line naming the survivor and the manual command to retire it, and a cleanup that threw must never surface instead of the original.
- Do not record the remnant in state as an alternative. It looks like
the tidier answer and it is a trap: the recorded properties ARE the
template's, so the next deploy diffs the resource as
NO_CHANGEand never touches it again —cdkd deployprints "No changes detected" and exits 0 over a resource that is still unusable. Turning a loud failure into a silent one is worse than the orphan. Making the diff re-provision it instead needs a marker on the state record, i.e. a schema bump.
Whether deleting is SAFE is a per-service question you have to answer, not assume. For ACM it is, and the reason is documented rather than inferred: the DNS validation CNAME is derived from the domain and the account, not from the certificate, and AWS states you can "replace a deleted certificate" without repeating validation — so the records a user adds after the failure validate the retry's certificate too. If your service has no such property, say so and choose differently.
A failing AUXILIARY call in
create()— a tag, policy, attribute or other write after the main create — must be marked withmarkAuxiliaryFailure(error, logicalId)(src/provisioning/auxiliary-failure.ts, issue #3826), or that object's "already exists" can be classified as the resource's own name collision, and--replaceor a replacement rollback then deletes the live resource.A
create()that fails AFTER its main create call returned must also mark the error it throws withmarkCreatedBeforeFailure(error, ownerLogicalId, resourceType, physicalId)(same module, issue #1710). The resource then exists with no state record; the mark is what lets the failed-CREATE journal name it, so the automatic rollback,cdkd rollbackandcdkd destroycan delete it. When the create call's response carries an immutable AWS-generated id for the resource (RDS'sDbClusterResourceId/DbiResourceId), pass it as the optional fifth argument,createdResourceIdentity. It is journaled without a later read, and it must be the value the provider'sresourceIdentityreturns for that resource, since a later successful deploy deletes a name-keyed orphan only when the two match. Set it only behind a flag the create call's success sets (Kinesis'sstreamCreated):ProvisioningError.physicalIdis NOT that proof, since providers put the intended name on refusals and on the create call's own "already exists", which names another owner's resource. Mark the outermost errorcreate()throws, with the iddelete()takes (the success path'sphysicalId). A catch that deletes its own resource marks only when that cleanup FAILED; a resource the create adopted or found held (an idempotent create returning an existing one,heldBefore !== 'free') is never marked; a provider whose delete would reach beyond what the create wrote (AWS::IAM::Policywalks every listed principal) removes its own writes in the catch instead (issue #4583).- Track the id from the API response, not from the name you computed.
Set the local only once AWS has confirmed the resource exists, so a
failure BEFORE the call deletes nothing. Do NOT reuse
maskSecrets— mask a resolved property value before you log it. BothCreateContextandUpdateContextextend a sharedSecretMaskingContext, whose one optional field ismaskSecrets?: (text: string) => string(issue #1932). Any provider log line that interpolates a value from thepropertiesbag MUST run through it: those values arrive RESOLVED, so a{{resolve:secretsmanager:...}}property is already plaintext by the time a provider sees it, and cdkd's other masking boundaries (the deploy engine's error / reason text, the resolver's own debug line) do not cover a provider's ownlogger.warn. Read it defensively —const mask = context?.maskSecrets ?? ((t: string) => t)— sincecreate()/update()are also called bycdkd drift --revert, by the import path, and by tests. It is per-CALL, so never cache it onthis: providers are registered as singletons and serve concurrent resources. A provider whosecreate()/update()reach many private helpers may instead re-enter itself on a fresh per-call object whose prototype is the singleton and whose logger is masked, so every helper'sthis.loggerline is masked by construction (S3BucketProvider.maskedView, issue #2177). Mask the VALUE before it is stringified or interpolated; the finished message is a FALLBACK, not an equivalent. Two independent reasons: (1) escaping — a masker matches by literal occurrence, andJSON.stringifyescapes",\and newlines, so a secret containing any of them (i.e. any Secrets Manager JSON document, the commonest real secret shape) no longer occurs in the finished line and passes through verbatim; (2) length —maskSecretsInTextmasks an exact whole-value match at ANY length but only scans for SUBSTRING needles of at leastMIN_NEEDLE_LENGTH(4) characters, and a message is always longer than the value inside it, so a 1-3 character secret survives. Do BOTH: mask each string leaf before interpolating (catches escaped and short secrets), and route the assembled message through the masker too (catches interpolations added later, and text the leaf pass never sees). The mask is idempotent, so the layers compose.createMaskedLogSinks(inmasked-retry-logger.ts, below) builds both halves for one operation: a rawvaluemasker plus maskeddebug/warnsinks. A physical name derived from a secret bygenerateResourceNameWithFallback(stack prefix, folded characters, truncation) no longer OCCURS literally, sowithDerivedNameMasksadds the derived spelling as a needle when its raw value is a secret. The same helper covers a rotated secret onupdate(): pair the PREVIOUS value with the recorded name, and a previous value still spelling{{resolve:, or persisted as exactly the redaction mask***, makes that name a needle although its old plaintext is in no bag of this deploy. UsemaskDeepfrom src/provisioning/masked-retry-logger.ts for the leaf pass — do NOT hand-roll one. Issue #2176 found SIX private copies of that walk, FOUR of which had lost the depth cap. A walk that encodes a security contract drifts the moment it is copied: a hardening applied to one copy silently leaves the rest behind. Grep for the walk's SHAPE rather than a name — two of the six were nearly missed twice because they spell itmaskLeaf/maskLeafValueand declare no named depth constant to grep for. The leaf pass is now MECHANICALLY ENFORCED for theJSON.stringifycase (issue #2178).vp run audit:provider-secret-mask:check(scripts/check-provider-secret-mask.ts, a CI step) fails on any${JSON.stringify(X)}interpolated into a message undersrc/provisioning/providers/**(pluscomposite-id.ts) whereXreaches no masker. It is DATAFLOW-aware, so the mask may sit anywhere upstream — through aconstbinding, a?:/??arm, an array or object literal, a.map()callback, or a helper — and both real upstream-mask sites (asg-provider.ts,cloudfront-distribution-provider.ts) pass unchanged. It became a check rather than another sentence because the rule above was already written and violated anyway: the #2176 sweep found raw sites in files already hardened for this exact contract.An IDENTITY masker does NOT satisfy it. A binding such as
const M: MaskerFn = maskerOrIdentity(undefined)type-checks, reads as protection, and masks nothing — so the critic classifies a site masked only by one asraw. Issue #2007 is why that is worse than having no masker at all: its presence stops the next author looking. This holds in the ARGUMENT position too, which is the one you are most likely to write:maskDeep(value, maskerOrIdentity(undefined))walks every leaf and changes nothing, so the critic classifies itrawexactly as it does a bare value. That position is an ALLOW-LIST — it accepts only what resolves to the shared capability and refuses anything else, so an inline(t) => t, a cast, a hand-rolledfunction noMask(v) { return v; }and an arbitrary object property are all refused without the check having to enumerate them.What the check does NOT prove. It is syntactic, so a no-op DECLARED to be the capability is believed:
const m: MaskerFn = someNoOpand{ maskSecrets: (t) => t }both pass. A green run means the value reached something declared to be the project's masker — not that it was masked. Thread the real capability; do not satisfy the checker. Where a path genuinely has no masker to thread —DeleteContextcarries none by theSecretMaskingContextcontract — thread the capability first, or record the site in the critic'sEXEMPTlist, where it is COUNTED and re-audited on every run and fails the moment the capability arrives. Do not reach for an identity default to make the check pass. A parameter's identity DEFAULT is a different thing and stays legitimate: it is the contract's back-compatible absent-means-unmasked answer, and what the critic checks there is that every CALL SITE threads a real masker — but only for an UNTYPED parameter of a non-exported function. A parameter ANNOTATEDMaskerFn/SecretMaskeris treated as a masker by TYPE, so it joins the file's masker set unconditionally and its call sites are NOT checked; the same is true of any exported function. Prefer the REQUIRED spelling (mask: MaskerFn, no identity default) for new sites: the compiler then enforces threading, which is the only mechanism that does not depend on the critic noticing. Every in-repo call site does thread today — verified, not assumed — so this is a shape to prefer, not a live gap.Back to the masker itself, whose scope the paragraphs above interrupted.
maskDeepmasks what the caller's resolution recorded: the dynamic-reference secrets, and since issue #1998 the value of aNoEcho: truetemplate PARAMETER that aReforFn::Subvariable served, recorded as a LOG-ONLY needle that the masker reads and, since state schema v11, also as a mask-only needle, so state persists***where the value stood. No provider change was needed for either, which is the point of handing providers a function. A DIFFERENTNoEcho— the custom-resource RESPONSE field of the same name — IS covered since issue #2274, throughnoEchoAttributesabove; the two share only a spelling, so do not read one as closing the other.delete()has no masker — seeSecretMaskingContextfor why that is a statement about providers rather than about the bag, and why you must thread the capability BEFORE adding any delete-side log line that names a property value. Mask where the message is CONSTRUCTED, not where the arm catches it, when the service reports rejections through an operation poller. For an operation-based API (AWS Cloud Map'sCreate*Namespace/Update*Namespacereturn anOperationId, and the rejection arrives later asOperation.ErrorMessagewith the offending value quoted back), a shared poller builds the error and every arm's catch opens withif (error instanceof ProvisioningError) throw error;— so the poller's error is re-thrown VERBATIM and masking the arm does nothing for it. Threading the masker into the arm therefore looks like a fix and is inert on the path that actually fires, since for these types a FAILED operation is the NORMAL rejection route rather than an edge case.pollOperationin src/provisioning/providers/servicediscovery-provider.ts is the worked example (issue #2063): it masks the rawErrorMessageat construction, and it takes the masker from every CREATE / UPDATE caller rather than only the newly-threaded ones — an arm fixed by an earlier pass is not evidence that its poller was. The DELETE callers stay unthreaded, which is the contract above rather than an oversight:DeleteContextcarries no masker, and their payload is a physical id rather than a resolved property bag. The reference implementation isbuildMfaConfigRequestin src/provisioning/providers/cognito-provider.ts, which routes every warning through one masked sink rather than masking at each call;create()in src/provisioning/providers/ssm-parameter-provider.ts is the same shape for a whole operation (onemask, onewarn, onedebug, built at the top and used everywhere below). Per-site masking DOES drift, and that is measured rather than predicted. Issue #2176's sweep found raw${JSON.stringify(...)}sites in files ALREADY hardened for this contract —dynamodb-table-provider.tsmasked one argument of awarncall and left the argument beside it raw. A sink is what makes the next line added below it correct by default.
The parameter is optional, so a provider with no pre-flight needs no change. A provider that HAS one threads
contextfromcreate()to its check and emits a warning instead of throwing when the flag is set. What the flag means (and, just as importantly, what it does NOT license — nothing about the properties' content, no relaxing of data-safety guards, no skipping the validation that protects the AWS call itself) is spelled out onCreateContextin src/types/resource.ts, next to theResourceProviderinterface that consumes it. Its siblingDeleteContextlives in src/provisioning/region-check.ts instead, becauseexpectedRegionfeeds that module'sassertRegionMatchhelper;CreateContexthas no region-checking role, so it is not filed there for symmetry alone. A pointer next toDeleteContextlinks the two.One honest consequence: warning on a replay means the value IS forwarded on the create path, unlike the update path where the SDK command has no member for it. So the re-created resource is degraded in whatever way the original was, and the AWS call may still fail. That is strictly better than refusing — a refusal guarantees the resource is not restored — but the warning must SAY so and name the fix-forward (
cdkd deploywith the working shape).Where update genuinely does not forward the property anyway (cdkd does not wire Glue's update-only
UpdateOpenTableFormatInputshape), warning costs nothing: no bad value can reach AWS from that path, and the user still gets the full message. Share ONE message builder between the refusal and the warning so they cannot drift.
A pre-flight in a provider only covers the SDK route. A resource whose
state records provisionedBy: 'cc-api' is sticky-routed to
CloudControlProvider and bypasses it entirely. That is usually acceptable
(the deploy still fails, just later and less helpfully) — but say so in the
docs rather than letting a reader assume the refusal is total.
If a property is merely unimplemented rather than undeployable, this is the
wrong mechanism — move it to unhandledByDesign, which converts the silent
drop into the Cloud Control auto-route (see handledProperties against the CFn schema).
Idempotency
- Handle when
createis called on existing resource - Handle when
deleteis called on non-existent resource
Region verification on *NotFound: A *NotFound error during DELETE
must NOT be treated as idempotent success without confirming that the AWS
client's region matches the region the resource was deployed to. A destroy
run pointing at the wrong region would otherwise receive NotFound for
every resource and silently strip them all from state, leaving the actual
AWS resources orphaned in the real region (this is the silent-failure
incident that motivated PR 2 of the region/state refactor).
Providers MUST call assertRegionMatch() from
src/provisioning/region-check.ts before returning early on a *NotFound
error:
import { assertRegionMatch, type DeleteContext } from '../region-check.js';
async delete(
logicalId: string,
physicalId: string,
resourceType: string,
_properties?: Record<string, unknown>,
context?: DeleteContext,
): Promise<void> {
try {
await this.client.send(new DeleteXxxCommand({ Id: physicalId }));
} catch (error) {
if (error instanceof ResourceNotFoundException) {
const clientRegion = await this.client.config.region();
assertRegionMatch(
clientRegion,
context?.expectedRegion,
resourceType,
logicalId,
physicalId,
);
this.logger.info('Resource not found, skipping deletion');
return;
}
throw error;
}
}
assertRegionMatch is a no-op when context.expectedRegion is undefined,
preserving the existing idempotent semantics for callers that have not
been threaded with state region. When set, a region mismatch throws a
ProvisioningError that surfaces both regions and a hint to rerun with
the correct --region.
The NotFound branch is the REACTIVE use, and it is not the only one
(issue #2301). A wrong-region
call usually never REACHES that branch: most physical ids are NAMES, and the
same name commonly exists in the client's region as well — the same stack
deployed twice, or resource-name.ts deriving an identical name from an
identical stack plus logical id — so the call succeeds against the WRONG
resource instead of erroring. assertRegionMatch therefore takes an optional
sixth phase argument: the default 'not-found' keeps the wording above,
while 'pre-delete' / 'pre-update' word the same refusal for a call that has
not been issued yet. The comparison and the three outcomes are one
implementation across all three phases; only the message differs.
CloudControlProvider uses those phases UNCONDITIONALLY, for every
Cloud-Control-routed type, at the top of both delete() and update() —
before the --remove-protection flips, before the SDK delegations, before the
AWS::S3::Bucket identity probe, and before the DescribeType the update
path's write-only lookup performs. The update side is why UpdateContext
carries an expectedRegion of its own (src/types/resource.ts), threaded by
deploy-engine.ts, rollback-executor.ts and drift --revert. An SDK
provider MAY read the same field, but nothing requires it to: absent, the
guard is a no-op, exactly as on the delete side.
The create side has the same question, and it is NOT the same answer.
"Handle when create is called on an existing resource" above is the usual
idempotent short-circuit: an *AlreadyExists / *AlreadyOwned error is
swallowed and the create reported as success. Before writing one, ask what
namespace that error is scoped to and whether it is narrower, equal to, or
WIDER than the region the client points at. It is sound only when the error
can mean nothing but "the resource I want, where I want it, already exists".
For nearly every type that is automatic — either the namespace is REGIONAL
(SNS, CloudWatch Logs, SSM, EC2), so the conflict is by construction in the
region you called, or the namespace is global AND so is the resource (IAM,
Route 53, CloudFront), so there is no other region for it to be in.
AWS::S3::Bucket is the one type where the two come apart, a globally unique
name over a regionally located resource, and it broke exactly there: measured
on real AWS, a bucket in us-west-2 answers BucketAlreadyOwnedByYou to a
CreateBucket in eu-west-1 and in us-east-1 alike, so the short-circuit
adopted another region's bucket and applied the whole stack's configuration to
it while reporting success (issue
#2227).
A conflict that carries no name at all is never sound to swallow: an
AWS::EC2::SecurityGroupIngress duplicate means only that an identical rule is
on the group, whoever added it. Such a create adopts only on this stack's own
evidence — its rollback journal (getPriorAttempts), a rollback replay, or
another record of the same stack holding the rule (getStackRecords; two
template rules CDK could not dedupe, issue
#4492, waiting for such a twin
still being created in the same deploy, as CloudFormation accepts both) — and
refuses otherwise.
Two records then share one rule, so a delete leaves it in place while a record
that outlives the operation still holds it, and the last holder revokes it.
S3BucketProvider.assertExistingBucketRegion is the create-side twin of
assertRegionMatch. It reads the bucket's region from the
x-amz-bucket-region header on the 409 itself — no extra call, no extra IAM —
and falls back to GetBucketLocation. Note assertRegionMatch is a no-op
without expectedRegion while this one always has a region to compare against:
the create knows where it is deploying.
There are FOUR of these guards, not two, and assertRegionMatch is only the
generic one. The two above are the SDK route; the Cloud Control route has its
own pair, because a bucket declaring a silent-drop property (AccessControl) is
auto-routed there and S3BucketProvider.delete never runs at all (issue
#2283, fixed in
#2309):
| Guard | Route / phase | On an INDETERMINATE answer |
|---|---|---|
S3BucketProvider.assertExistingBucketRegion |
SDK, create | refuses |
assertRegionMatch (src/provisioning/region-check.ts) |
generic, delete; update only where a provider opts in (S3BucketProvider alone today) |
no-op without expectedRegion |
CloudControlProvider.confirmDeleteTargetIdentity |
CC, delete, per-type (S3 today) | warns and PROCEEDS |
CloudControlProvider.assertRecordedRegionAgainstClient |
CC, delete + update | refuses |
The last two sit in the same file and answer indeterminacy in OPPOSITE
directions, which is deliberate rather than an inconsistency. The probe asks a
REMOTE service where a globally unique name lives, so a least-privilege role
never granted s3:GetBucketLocation would be stranded by a refusal; the
comparison is LOCAL and free against a region the caller positively recorded, so
a client that cannot say where it points cannot be shown to point there. Two
further details of the probe are decisions, not oversights: it is keyed on a
per-type set rather than being generic, and it deliberately omits
ExpectedBucketOwner — the opposite of the convention state.ts and
aws-region-resolver.ts follow — because the hazard is a name resolving to a
bucket in ANOTHER ACCOUNT, and the guard has to hear that foreign answer to
refuse. Passing the parameter would turn every cross-account collision into a
403, i.e. into the indeterminate arm, i.e. into warn-and-proceed.
Deliberately NOT HeadBucket, which 301s cross-region and which SDK v3 turns
into a synthetic UnknownError (src/utils/aws-region-resolver.ts records the
mechanism), and deliberately not that module's resolveBucketRegion, which
never throws and returns a fallback region — wiring it here would turn a
fail-closed guard into a fail-open one. See
.claude/rules/provider-resource-identity.md for the full rule, including the
two legacy GetBucketLocation spellings the fold must absorb and why the
refusal is markNonRetryable. (.claude/rules/providers.md is a routing index
now — the rules corpus was split under the corpus-wide byte ceiling that issue
#2310 has since retired in
favour of a per-paths:-glob budget — so it points on rather than carrying the
rule itself.)
That guard is scoped to the BucketAlreadyOwnedByYou catch on the CREATE path,
which leaves two neighbouring routes to the wrong bucket that it cannot see —
us-east-1's legacy 200 (issue
#2241) and a state record written
before the guard existed (issue
#2245). S3BucketProvider now
also pre-flights the bucket NAME before CreateBucket when the target region is
us-east-1, and shares one assertStateBucketRegion between update() and
delete(). .claude/rules/provider-resource-identity.md carries the mechanism
and the three rules worth reusing: the pre-flight informs the partial-create CLEANUP GATE (and,
after an ambiguous attempt of the same create, a refusal that adopts nothing, issue
#4639) and never licenses an adoption in place of
CreateBucket (the ownership oracle), an unanswered probe is
kept distinct from a confirmed absence, and the state-record guard PROCEEDS on
both rather than stranding every update and destroy for a least-privilege role.
Region is not the only question. A bucket under an EXPLICIT BucketName that
this create did not make is refused even in the right region, on both the
BucketAlreadyOwnedByYou arm and the us-east-1 pre-flight (before the send,
so the legacy 200 never resets its ACLs; when the location lookup cannot
answer, the account's own bucket list decides, since the legacy 200 adopts
only a bucket the caller owns), as CloudFormation fails the same
create (issue #4684): a name is
not attribution. Only a generated name's holder is still adopted, presumed this
stack's own (#4345). The refusal
is marked a name collision, so a create-first replacement and a rollback's
reverse replacement take their record-proven collision paths rather than
recording the bucket. It is deliberately retryable by text ("already exists"):
a delete-first re-create retries that text while its own deleted name is
released, and no other caller's classifier retries a collision.
A bucket has no immutable AWS id, so its resourceIdentity (the token a
fix-forward settle compares before deleting a failed CREATE's orphan, issue
#4606) is its name, its region
and the CreationDate this account's own ListBuckets reports. A bucket the
list does not show, or shows in another region, has no token. A bucket
re-created under the name reports the new create's second (measured in
us-east-1 and us-west-2, and asserted by the s3-fix-forward-orphan fixture
before it trusts the token). Outside us-east-1 the date also moves to the
second of a versioning, tagging, encryption or policy write, so the token is
read only after the failed CREATE's last write; a moved date reads as another
bucket and keeps the orphan, the safe direction.
Note what that last choice COSTS, because it is a new IAM dependency rather than
a free win: the guard works only where the caller can call GetBucketLocation
on the target. A principal without that grant is NOT an unrelated population —
it is the same population, unprobeable, and for it the guard is inert on every
call; a bucket POLICY can Deny the call as well, which is indistinguishable on
the wire from a missing grant. That is why the delete-side degrade logs at
warn rather than debug: proceeding unverified into an irreversible delete
must never look identical to a healthy one.
That rule is also where the probe-first lesson lives, and it cuts both ways here: the mid-delete hazard the issue was FILED for is not reachable, so a guard against it could never have fired — while the guard that WAS written could not fire either, because its unit tests encoded the AWS CLI's wire shape rather than the SDK's. A mutation probe cannot catch that: it perturbs the code and reads the test, and both read the same fixture, so a premise shared by code and mock is invariant under it. Only a recorded real response or a live arm falsifies a fixture.
The S3 pre-flight's partial-create cleanup gate has a sibling for every create
API that answers success with a resource already holding the name: ELBv2
CreateLoadBalancer / CreateTargetGroup on identical settings, SNS
CreateTopic, and EventBridge PutRule, which overwrites (issue
#4403). Before the create, and
only when a wiring step that can fail is declared, the provider looks up the
name it is about to send (src/provisioning/providers/create-ownership.ts).
Only the service's own not-found answer licenses the partial-create cleanup. A
name that was held, or a lookup that could not answer, leaves the resource in
place with a warning saying why and the manual delete command. A new provider
of such a type gates its cleanup the same way. A create that hands back by
idempotency TOKEN rather than by name (FSx CreateFileSystem) cannot be looked
up beforehand, so FSxFileSystemProvider takes -- records, and so may later
clean up -- only a file system whose CreationTime is no earlier than that
token's first send in this process, on AWS's clock, and refuses any other
(#4428).
Reporting a skipped delete
Written for issue #1752.
delete() returning normally means "the resource is gone" — a delete
call succeeded, or the resource was already absent (the *NotFound
idempotent arm above, CustomResourceProvider's backing-Lambda-is-gone
pre-check). Both are honest, so both count toward N deleted.
A warn-and-continue arm that issues no AWS call at all is different, and
returning void from one is a bug. cdkd destroy's only signal used to be
"did not throw", so such an arm printed ✓ <id> (<type>) deleted, counted
toward N deleted, dropped the state record, and exited 0 — over a
resource that may still be alive and billing. Report the skip instead:
const decoded = decodeTableId(physicalId, properties);
if (!decoded) {
this.logger.warn(
compositeIdFormatMessage(GLUE_TABLE_ID_FORMAT, logicalId, physicalId, { skipping: true })
);
return compositeIdSkipResult(); // { outcome: 'skipped', reason }
}
What the destroy runner then does, and why each half matters:
- prints
⚠ <id> (<type>) skipped (<reason>)— a yellow warning glyph, so it can be read as neither the✓ … deletedsuccess line nor the✗ Failed to deletefailure line; - counts it in a separate
skippedcolumn (2 deleted, 1 skipped, 0 errors) rather than as an error — nothing was attempted, so nothing failed; - KEEPS the state record. This is the half that is easy to miss: dropping it leaves the user with neither the AWS resource deleted nor an id to go and delete it with;
- preserves
state.jsonand exits 2 (PartialFailureError), the same contract a partial / interrupted destroy already carries. Exiting 0 while leaving a resource behind is the mis-report this mechanism exists to prevent; - emits a
RESOURCE_SKIPPEDdeployment event (noerrorfield).
Reach for this ONLY when cdkd genuinely cannot address the resource. If you
know the resource is gone, return void — reporting a skip there would
preserve state and fail the destroy for no reason. The --purge-events flag
and the run-level RUN_FINISHED result both treat a skip as a failed run, so a
new producer changes user-visible exit behavior — decide deliberately.
Reporting a pre-flight guard that could not answer
A guard that runs before the delete and cannot reach a verdict is a THIRD thing,
and it is neither a skip nor a failure: the delete still goes ahead, so the
outcome is normally 'deleted' (a delete that ALSO could not be addressed keeps
'skipped' — the two are independent), and what needs reporting is that a
safety check was not enforced. Issue #2301 added
indeterminateGuards to ResourceDeleteResult for it:
import { withIndeterminateGuard } from '../../deployment/delete-outcome.js';
const guard = await this.confirmSomething(...); // IndeterminateGuard | undefined
// ...perform the delete...
return withIndeterminateGuard(undefined, guard);
withIndeterminateGuard returns its input unchanged when guard is
undefined, so the ordinary path keeps returning the back-compat void, and it
preserves a 'skipped' outcome and its reason when both facts hold at once.
Two rules apply, both because the value is PERSISTED:
- Proceed, do not refuse. These arms exist for probes a least-privilege caller may not be allowed to make. Refusing on an unanswerable probe strands every operator who never granted the permission; the durable record is what makes proceeding acceptable rather than silent.
- Name the GUARD, not the API or the type.
guardis a stable machine-readable id (cc-delete-region-identity, nots3-getbucketlocation) that lands indeployments/*.jsonland is therefore a user contract; a future guard reuses the same event type with its own id rather than minting a second one.reasonis the human half, and distinct causes get distinct text even when they reach the same proceed-anyway outcome, because the remedies differ.
The destroy runner then emits a RESOURCE_GUARD_INDETERMINATE event ALONGSIDE
the resource's own outcome row (never instead of it), counts it into the destroy
summary's N unverified suffix, and prints an aggregate warning. It moves no
outcome counter and does not change the exit code. Note that the deploy engine
and the rollback executor still DISCARD the field
(#2422), so a guard reported from
a delete on those paths is not yet recorded anywhere.
A provider that recurses into another destroy propagates the child's
result. NestedStackProvider.delete returns { outcome: 'skipped' } when the
child runner reports skippedCount > 0 OR interrupted; without that the
parent would print ✓ <Child> (AWS::CloudFormation::Stack) deleted and exit 0
over a child stack that kept a live resource. The child's third
not-gone signal, errorCount > 0, is a THROW rather than a skip (issue
#1777): a resource was attempted
and FAILED, so the parent's row must fail too, and calling it a skip would
assert the opposite of what happened (a skip means no AWS call was issued).
Word such a throw so it contains none of not found / does not exist /
No policy found / NoSuchEntity / NotFoundException (and the deploy
engine's was not found / ResourceNotFoundException) — both callers' catch
blocks read those as an idempotent already-deleted success and DROP the state
record, which is the outcome the throw exists to prevent. Note the asymmetry
between the two return channels: since issue #1762 a { outcome: 'skipped' }
value reaches the deploy-side sites too, but WEAKER — the template-removal
DELETE warns and keeps the record — while a throw FAILS the resource at every
caller, so converting a skip into a throw still changes cdkd deploy's
behavior when the resource is removed from the template.
A provider that DELEGATES its delete to another provider propagates the
delegate's outcome (issue
#1778). CloudControlProvider
hands a protected AWS::AutoScaling::AutoScalingGroup to ASGProvider.delete
under --remove-protection, and discarding that verdict is the nested-stack
hole one layer down: a skip inside the delegate reached the destroy runner as a
plain successful delete. return await delegate.delete(...) is the whole fix.
Type the delegate as ResourceProvider rather than as its concrete class, so
the forwarding stays correct when the delegate widens its own return type.
The alternative contract — ASSERT the delegate cannot skip — was rejected: it
has to be re-verified every time the delegate grows an arm, and it fails
loudly on a case the delegate considers merely unaddressable.
And a provider that pairs create() + delete() inside its own update()
must not swallow the delete's outcome either (same issue). A skip does not
throw, so it slips past the catch that would have warned — precisely when the
old resource IS left behind. What to do depends on the ORDERING, and it is not
the same for all of them:
- create-then-delete (ACM certificate, IAM managed policy, IAM role, and
the API Gateway Resource PathPart replacement): the new resource already
exists, so the replacement cannot be aborted. WARN, in the same orphan wording
the failure arm uses (
The old role may be orphaned and require manual cleanup.), naming the skip'sreason— and since issue #1819 ALSO return{ outcome: 'partial', reason }, so the outcome is counted and recorded rather than living only in the log. - delete-then-create (SNS subscription): ABORT with a
ProvisioningErrorBEFORE creating the replacement. Continuing would leave two subscriptions on the topic and deliver every message twice, and that duplicate is precisely what the CREATE would add — so not making it is the whole remedy.
State the premise as "the resource was not destroyed", never as "no AWS
call was issued". ResourceDeleteResult's own contract says so, and its
warning is not academic: NestedStackProvider reports skipped after
recursing into a child destroy that may already have deleted the child's other
resources. The abort is right under the weaker premise anyway — whatever the
skipping delete did, adding a second live resource on top of it is not an
improvement.
An abort on an update() path must be right for EVERY caller, not only the
template-driven one — cdkd drift --revert and the rollback executor's revert
arms call update() with a state-borne desired bag (see Pre-flight refusal). The SNS abort
clears that bar because a second live subscription serves none of the three;
where a refusal would not, downgrade it to a warning instead.
Two mechanical details any such abort inherits:
Do not interpolate the provider-supplied
reasoninto the thrown message. The rollback executor wrapsupdate()inwithRetry, andretryable-errors.tsclassifies by SUBSTRING — so a reason carryingdoes not exist,Rate exceeded, orbecause it is in usewould flip a deterministic abort to "retryable" and burn the whole backoff schedule before a certain failure. Log the reason; interpolate nothing into the thrown message but the TEMPLATE logical id. The state-borne physical id is the worst candidate of all — the only skip family a REPLACE path meets today is literally "malformed physicalId in state", andcdkd import --resource '<id>=<anything>'puts an arbitrary string there....and then
markNonRetryablethe error, because keeping values out of the message cannot close the hole. The match is a SUBSTRING, not an equality, so an ordinary composite logical id (MyDependencyViolationSub) still carries a pattern — measured — and a message that names nothing is not diagnosable.markNonRetryable(new ProvisioningError(...))(fromsrc/deployment/retryable-errors.ts) sets a non-enumerableSymbol.formarker thatisRetryableTransientErrorconsults BEFORE any name or message heuristic, so the refusal is terminal by DECLARATION and no wording can overturn it. Reach for it only where the error means "this cannot succeed on a retry" as a matter of cdkd's own logic — never for a relayed AWS failure, whose retryability is the classifiers' business.The fence is at the RETRY LOOP, not only in that one classifier:
withRetryrethrows a marked error BEFORE choosing between the default classifier and a caller-suppliedopts.isRetryable, and the destroy runner's delete-retry loop gates its ownToo Many Requestsmessage test the same way. So the message-only classifiers (isNameCollisionError'sAlreadyExists,isNameCooldownError'sQueueDeletedRecently/StateMachineDeleting— all of which match a bare logical id) cannot resurrect a refusal even though they cannot read the marker themselves. Keep the message discipline above regardless: it is what a future classifier added outside those two loops still relies on.Point the remediation at the STATE record, not only at AWS. Both skip families shipping today — the state-borne composite-id arms and
NestedStackProvider's propagation — describe a resource whose repair "remove the old resource in AWS" does not cover on its own, and for the malformed-physicalId one deleting the AWS resource fixes nothing at all. Name both repairs.
Note this is deliberately NOT symmetric with a THROWN delete, which those providers still warn-and-continue past: a throw may mean the delete partially landed or was transient, while a skip is a positive statement that the resource was not destroyed.
Producers today: the malformed-composite-physicalId family
(src/provisioning/composite-id.ts, five arms), the nested-stack propagation
above, and — since issue #1770 —
eight same-class arms outside the composite-id family: both malformed
LayerVersionArn arms in lambda-layer-provider.ts, the missing-FunctionName
arm in lambda-permission-provider.ts, the no-properties / no-ServiceToken
arms in custom-resource-provider.ts (and, since issue
#3938 and
#3960, its masked- and
secret-reference-ServiceToken arms), the empty-policy-name arm in
iam-policy-provider.ts, and both AWS::IAM::UserToGroupAddition arms in
iam-user-group-provider.ts. Issue
#3878 added the
malformed-target-list arm in iam-policy-provider.ts: a recorded Roles /
Groups / Users that is not a list of IAM names. Issue
#3888 added the same arm for an
AWS::IAM::UserToGroupAddition record's Users. Issue
#4150 added one arm to each of
those two providers: a destroy that resolved a secret-derived list and found a
principal the secret named without the grant keeps the record, since the secret
may have rotated. For an inline policy that is NoSuchEntity on the delete
call. For a membership it is a ListGroupsForUser read BEFORE the call, because
RemoveUserFromGroup succeeds for an existing user outside the group. Each exports its reason as
a named constant beside the provider, so the wording is pinned by a test instead
of retyped.
Issue #3952 added one shared
arm, redactedDeleteAddressSkip (src/provisioning/redacted-delete-address.ts),
for a delete whose recorded address property cdkd redacted (the *** mask or a
secret {{resolve:...}} expression): ApiGateway / ApiGatewayV2 children, ECS
services, Glue catalog-scoped resources, Scheduler schedules (unless the record
holds the schedule's creation date, which with the recorded target ARN and role
locates it in any group; a record with no stack region, or a schedule matching
only the date, is still skipped:
#4275), Route 53 record
sets, security-group ingress rules, CloudWatch anomaly detectors, DB proxy
target groups, and the
fallback-less arms of the Lambda permission, IAM policy, UserToGroupAddition and
ECS service deletes.
Three lessons from that issue's code review are worth reusing before you add a skip arm of your own.
Exhaust every addressable source first. A skip is not a free "safe"
default: it preserves the record, warns, and exits 2 — on every re-run, so the
destroy can never go green. Two of the eight arms were skipping although a
second source held the value. RemovePermission accepts a full ARN as
FunctionName, and the physicalId's documented
<functionArn>|<statementId> shape carries one; PolicyName is in
handledProperties and create() uses it verbatim as the real AWS name. Order
the sources by what was DEPLOYED — the physicalId wins over the property, or a
template edit that renamed the policy without a replacement having landed would
send a delete for a name AWS never had.
Validate the fallback, or it is worse than the skip. A truthy NON-string
(an unresolved intrinsic, an array) must not beat a good second source — the SDK
URI-encodes it into a garbage label that 404s, and the idempotent *NotFound
arm then reports DELETED. Check the fallback's REGION too: a provider holds one
client, so a cross-region ARN sends the call to the wrong region and is reported
DELETED the same way. And do not let the gate refuse a genuine source — the CFn
primary identifier [FunctionName, Id] carries a BARE name as often as an ARN.
A guard that admits a record must check the path it opens reaches AWS. The
IAM PolicyName fallback turned an honest skip into a silent deleted: with
the name resolvable, a record naming no Roles / Groups / Users reached a
body where every branch is skipped — zero AWS calls, returning undefined, i.e.
DELETED. Trace the admitted path to an actual call, and use the same truthiness
spelling the branches downstream use so a null-valued list cannot slip past.
"LEFT IN PLACE" is false when the parent is in the same stack. The destroy
keeps going after a skip, and deleteGroup / deleteUser remove exactly those
memberships, a deleted Lambda function drops its whole resource policy, a
deleted IAM role drops its inline policies. Qualify the wording and name the
command that drops the record through stateOrphanRecordRemedy
(src/provisioning/state-orphan-remedy.ts): cdkd state orphan '<stack>' only
on a stack destroy, --resource <logicalId> everywhere else, since on a deployed
stack the whole-stack form drops every live record
(#4602). Do not copy the
qualifier onto an arm where it is false (a layer version and a Custom
Resource's external side effects are undone by nothing).
"Repair state.json and re-run" holds on destroy AND on the deploy engine's
template-removal DELETE — both keep the record (issue
#1762). It does NOT hold for a
deploy-side REPLACEMENT or rollback delete, which fail the resource and can
leave the old one untracked, so every skip warning carries the same caveat
compositeIdFormatMessage does.
Two judgment calls from that issue are worth reusing. The
UserToGroupAddition arms logged at DEBUG, which read as "routine, nothing to
do" — but GroupName and Users are both REQUIRED by the CloudFormation
schema, so a record missing either is CORRUPT rather than empty, and the
memberships AddUserToGroup created survive the destroy; they are skips, and
the level was raised to WARN to match (a skip preserves state and exits
non-zero, so the user needs the explanation at normal verbosity). An EMPTY
Users: [] is the opposite case and stays a deleted: an array is truthy, so
it falls through to the removal loop and correctly does nothing.
The *NotFound idempotent arms in those same files are deliberately
untouched, as is CustomResourceProvider's backing-Lambda-is-gone pre-check —
those mean the resource IS gone. The deploy-engine / rollback-executor callers
consume the return value as of issue
#1762: the template-removal
DELETE warns and keeps the record, a replacement delete fails the resource, and
a rollback delete counts as a per-op failure — except at the arm that deletes
the NEW resource after the old one was already re-created, where the delete is
best-effort and a skip only warns.
A create that mints a server-side id needs an idempotency token
Written for issue #2039.
The deploy engine wraps every provider.create() in its outer transient-error
retry, and HTTP 500 / 502 / 504 are retryable (issue #2026). So a 500 whose
request actually SUCCEEDED server-side re-invokes create() from the top. For a
create whose id comes from AWS rather than from a name, the replay provisions a
SECOND resource: it has no state entry, cdkd destroy never reaches it, and it
bills indefinitely.
When the API takes a ClientToken / ClientRequestToken / IdempotencyToken /
CallerReference, take it from the shared helper:
import { acquireIdempotencyToken } from './idempotency-token.js';
const token = acquireIdempotencyToken({ scope: 'RunInstances', logicalId });
const response = await this.client.send(
new RunInstancesCommand({ ...input, ClientToken: token.value })
);
token.release(); // SUCCESS PATH ONLY
Three rules, each of which has a failure mode behind it:
- Never derive the token per attempt. A
Date.now()/randomUUID()value is worse than no token at all, because the call site then LOOKS idempotent while behaving exactly as before.AWS::Route53::HostedZoneshipped that way. - Release only on success, and also wherever the provider DESTROYS the
resource the token names (the
RunInstanceswiring-failure path terminates the instance). A released token is never handed out again, so releasing on a failure path defeats the mechanism, while NOT releasing after the resource is gone can hand a later create the resource it just destroyed — EC2 keeps aRunInstancestoken for ~24h. - Check what the API does with a repeat, and for how long. Most return the
original resource; Route 53 REFUSES a repeated
CallerReference(HostedZoneAlreadyExists), so that provider recovers by looking the zone up by its caller reference and adopting it. EFS refuses too: a repeatedCreateAccessPointClientTokenis refused withAccessPointAlreadyExists(which names the survivingAccessPointId) whether or not the repeat's parameters match (measured in #2442), soEFSProvider.createOrAdoptAccessPointreads that access point back, confirms that itsClientTokenis the one cdkd minted, that it belongs to the file system cdkd asked for, and that itsPosixUserandRootDirectoryare the ones this create requested (read through EFS's defaults: an absent ornullPOSIX user, a/root), and adopts it. Because the refusal ignores the parameters, the token alone would adopt a mismatched access point. A stable token that turns every retry into a hard failure is only half a fix. Note EFS documents no retirement period for this token -- the "one minute" in the EFS User Guide's "Creation token and idempotency" section is about file-system CREATION tokens -- so do not write a provider whose correctness depends on that duration; see the DERIVATION note increateAccessPoint. - Wrap the rethrow when you decline to adopt.
DeployEngine's replacement path classifies a failed create by testing the TOP-LEVEL message foralready exists/AlreadyExists-- exactly what the raw AWS conflict carries -- and by checking the error chain for a duplicate-name exception name. Rethrowing it bare makes the engine read a TOKEN collision as a physical-NAME collision, and under--replacefall back to delete-first: deleting the OLD resource and re-creating with the same still-unreleased token. Wrap it, keep the AWS error ascause. - Do not assume the SDK's auto-fill is a fix. A field carrying Smithy's
idempotencyTokentrait is auto-populated when the caller omits it -- with a FRESH value per request, which is the per-attempt derivation the first rule forbids. Measured on@aws-sdk/client-efs3.1018.0: two identicalCreateAccessPointsends went out with different UUIDs, and a caller-supplied value went out verbatim.
EFSProvider's FILE SYSTEM CreationToken, FSxFileSystemProvider's
ClientRequestToken and CloudFrontOAIProvider's CallerReference
deliberately do NOT use acquireIdempotencyToken: those APIs bind the token to
the resource for its LIFETIME, so a deterministic hash of the immutable create
inputs is right there (it also lets the new file system coexist with the old one
during a replacement). Derive it with stackScopedCreateToken, which also
hashes the stack name and region, and refuses to run outside a withStackName
scope: without them, two copies of one stack in an account sent the same token,
and FSx and CloudFront handed the second stack the first one's resource (#4428).
Take it through reserveStackCreateToken
(src/provisioning/providers/create-token-ledger.ts), which folds in the
stack's create-token ledger nonce -- replaced whenever cdkd lets go of a
resource that still holds a token (#4438) -- and records the send before the
create is sent, so a re-run after an interruption sends the same token and
reports when that earlier attempt was first sent. Add the type to
LEDGER_TOKEN_RESOURCE_TYPES, and pass any new site that keeps such a resource
alive while dropping it from state to noteRetainedResource(type, logicalId),
which rotates the nonce and drops that logical id's entry. The entry is kept
until the deploy that wrote it succeeds (forgetRecordedCreateTokens after the
final save): a re-run of an interrupted deploy needs it, and a later let-go
must not find it. When the ledger cannot be read or written,
reserveStackCreateToken throws: a create sent with any other token could not
be found again by the re-run, and its resource would leak. A
token stable across runs can still be held by a resource cdkd did not make, so a create takes a
handed-back or named file system only when it was created after that create's
FIRST send (in this process, or recorded by the ledger) -- judged on AWS's clock, never the host's: a host
clock running fast would otherwise refuse every ordinary create.
src/provisioning/providers/server-clock.ts reads the response's HTTP Date
(withServerClock / earliestOwnCreationTime) and falls back to SigV4's
five-minute bound without one. FSx refuses an older one outright, and
EFSProvider.sendCreateFileSystem (EFS refuses a repeated CreationToken with
FileSystemAlreadyExists) adopts the named one only when this process's own
earlier attempt, or the ledger's, may have made it. A holder still deleting, or one whose
read-back fails transiently, rethrows EFS's own error stamped
markReplayMayCollide: a delete-first re-create's message-based retry waits it
out, and no delete-first site acts on a holder cdkd does not own. That reasoning is
per-API, not per-provider --
the same EFSProvider DOES take acquireIdempotencyToken for CreateAccessPoint. That choice
does not rest on knowing how long the token lives, because the process-scoped
derivation is right under both readings: if the token retires quickly,
stability across runs buys nothing (no two deploys are that close together on
one create) while it would put a --replace delete-then-create inside the
replay window of the access point it just deleted; and if it instead binds for
the access point's lifetime, a deterministic hash is worse still, since any
re-create with unchanged inputs would collide with the live access point
instead of creating one.
Where the API has NO token member, the choice is between disableOuterRetry
(which makes the provider single-shot for EVERY transient error, IAM propagation
included) and a pre-create reconcile that detects the previous attempt's orphan.
Say in-code which you chose and what it costs.
A reconcile is not simply "delete what appeared since my baseline" — write it
so the wrong answer LEAVES an orphan rather than destroying a live resource.
"Newer than my snapshot" is equally true of the orphan your own attempt minted
and of a resource something else created in the same window: a sibling resource
in the same stack (the deploy engine dispatches at --concurrency 10 by
default), a second cdkd process, or a human. Deleting the second kind is
strictly worse than the bug being fixed — it destroys something cdkd state still
advertises, and for a type whose readCurrentState cannot tell that the
resource is gone (returning undefined rather than RESOURCE_NOT_FOUND),
cdkd drift reports UNKNOWN rather than drift, so the loss is invisible.
IAMAccessKeyProvider shows the shape:
- Serialize per owning resource in-process. A
Map<owner, Promise>queue, with the baseline read, the create, and the reconcile all inside it, so no sibling create can interleave with them. - Reconcile from the FAILURE path of the attempt that took the baseline, never from the top of the next attempt (this is about a reconcile that DELETES; a report-only lookup, below, runs at the top of the next attempt on purpose). A baseline that spans the whole retry schedule describes a window many seconds wide.
- Send the create through a client that refuses the SDK's 5xx retry
(
withoutServerErrorRetries; #4639). The SDK's replay inside onesendSUCCEEDS with a second resource, so the failure path, and the reconcile in it, never runs. The SDK still replays a socket reset or timeout, so ALSO run the reconcile on a success whose outputreplayFollowedAmbiguousAttemptreports, after recording the returned id in the set below (#4687). Not onreplayedInSendalone: that is also true after a throttle the SDK retried, which minted nothing, and every needless reconcile is one more chance to delete another process's fresh key. - Require a creation timestamp at or after the attempt started, with a small margin for clock skew between you and the service.
- Keep a set of ids this process created successfully and never delete one. This is the condition a baseline structurally cannot express, and it is the one that covers the case the in-process lock cannot: another process.
- Report anything you decline to delete, at
warn, and do not claim an attribution the code did not establish.
Where the API has no token and nothing can be deleted safely,
src/provisioning/providers/ambiguous-create.ts is the shared shape (issue
#2080: CreateKey,
CreateUserPool, CreateGraphqlApi, and API Gateway's CreateAuthorizer,
CreateDeployment, CreateApi and CreateIntegration, EMR's
RunJobFlow, AddInstanceFleet and AddInstanceGroups, Lambda's
PublishLayerVersion and CreateEventSourceMapping, AppSync's
CreateApiKey, DLM's CreateLifecyclePolicy, ECS's
RegisterTaskDefinition, and EC2's CreateVpc, CreateSubnet,
CreateInternetGateway, AllocateAddress and CreateSecurityGroup, whose
shared report is orphan-report.ts):
- Keep the SDK from replaying a 5xx. Send the create through a dedicated
client wrapped by
withoutServerErrorRetries: the SDK's own retry of a 5xx inside onesendis a duplicate nothing can see. The engine's retry covers 5xx and throttles; the SDK keeps its retry of throttles, connection failures, clock skew and socket resets, which the engine does not retry. Every token-less create whose replay COLLIDES instead of duplicating (a name-unique create: a stream, repository, role, function, table, cluster, directory or table bucket and the like — the providers that send one through this client and arm no latch, plus EC2'sCreateSecurityGroup, which collides on its group name in the VPC, and CodeCommit's seedCreateCommit, which collides once the branch exists) needs this client and no lookup: the surfaced 5xx letswithRetrymark the collision as possibly this create's own (#3978), so it is never credited to another holder. The exception is a collision whose error does not name the holder: EC2'sCreateSubnetcollides on its CIDR, butInvalidSubnet.Conflictnames no subnet id, so it keeps the lookup to tell the user which subnet to inspect. S3'sCreateBucketis not one of these either: it keeps its latch and refuses a bucket that exists after its own server error rather than adopting it (#4686). - Look only after an AMBIGUOUS failure.
AmbiguousCreateLatch.noteFailurearms onisAmbiguousOutcomeErrorthrown by the create call itself, and the next attempttakes it before creating again. A definite refusal (a 4xx, a throttle, IAM propagation'snot authorized) costs no lookup. Pass the taken window back tonoteFailure, so two ambiguous attempts in a row are both covered. A create that SUCCEEDS after the SDK replayed it inside itssend(replayedInSend:$metadata.attempts > 1, e.g. after a socket reset) runs the same lookup right after the create, overreplayedSendWindow, never offering the returned id as a candidate: exclude it by id when it joins the provider'sRecentIdSetonly after follow-up calls, so a resource a failed rollback leaves behind stays nameable by a later lookup (#4687). - Bound the window on BOTH ends. The latch records the ambiguous
attempt's start AND end (each widened by a skew margin) and expires after
AMBIGUOUS_LATCH_TTL_MS: where the resource carries a creation date, a same-named resource created later -- by another process, after this one gave up -- is never a candidate. A lookup over a type with no creation date cannot apply the window, and its report says so. Where only a per-candidate read carries the date (GetLifecyclePolicy,DescribeTaskDefinition), narrow by the list's own fields first, then read each remaining candidate; a list ordered newest first stops at the first candidate older than the window. - Adopt only on EXACT attribution, which a lookup never has. The one adoption in this family is a KMS key id that came back in this process's own response before a follow-up call failed, bound to a digest of its inputs. A NAME is not attribution, even a cdkd-generated one: names are scoped to the account and region, the stack lock to the state bucket and prefix, so another run of the same stack name can create an identical resource in the window. REPORT candidates and create anyway.
- Lead the report with a READ command, and offer a delete command only
after it, conditional on confirming the candidate is this deploy's orphan: a
candidate may belong to another deploy. Where there is no window at all
(no creation date:
GraphqlApi, an AppSync API key, an API Gateway authorizer, an API Gateway v2 integration, an EC2 VPC, subnet, internet gateway or Elastic IP) print no delete command. - A lookup that fails, transiently or not, warns and lets the create proceed: nothing adopts, so a failed lookup has no stake worth failing a create over. Say only what was LISTED: list APIs are eventually consistent.
- Tagging the create with a cdkd token was rejected as the attribution
channel: it makes a
TagResourcepermission (or a request-tag condition) a requirement of every create, collides with tag policies, and leaves a tag every read and drift path must strip.
Update removal semantics: clear-on-removal
CloudFormation resets a property removed from the template to its
default. Most AWS Update* / Modify* APIs do the opposite: an absent
input field means "no change" (merge semantics). A provider update() that
maps template properties straight into the SDK input therefore silently
drops removals:
// BROKEN for removal: properties['Timeout'] is undefined when the user
// deletes `timeout` from their template → the SDK omits the field → AWS
// keeps the old value, while CFn would reset it to the default.
Timeout: properties['Timeout'] as number | undefined,
The failure is invisible end-to-end: the diff layer correctly reports
old → undefined, the deploy reports updated + success, state drops the
field — so the very next cdkd diff says "No changes" while AWS still holds
the old value. Permanent, undetectable divergence from CloudFormation
(confirmed live for Lambda in issue #1155).
Rule: every optional, mutable property passed to a merge-semantics update
API needs an explicit reset when it was present before and is absent now.
There are two shapes, both in src/provisioning/update-removal.ts (issue
#1160):
- Declare it when the reset is a CONSTANT, CFn-shaped, top-level value
that
update()forwards verbatim.ResourceProvider.removalDefaultsmaps type → property → value; the update CALLER (the deploy's in-place update, both rollback revert arms) injects it into the bagupdate()receives, and only there: state keeps the template, and a reset echoed back througheffectivePropertiesis taken out again.update()starts withproperties = withRemovalDefaults(this.removalDefaults, resourceType, properties, previousProperties, context), which injects the same values on a direct call and is a no-op once the caller did. - Keep a local
clearOnUpdateRemovalfor anything else: a coerced value (Number(...)), an aliased key, a clear that depends on another property or on the logical id, a nested member, or an SDK-shaped value.
removalDefaults = new Map([
['AWS::Lambda::Function', new Map<string, unknown>([['Timeout', 3], ['MemorySize', 128]])],
]);
// clearOnUpdateRemoval(newValue, previousValue, clearValue):
// present -> pass through; removed -> explicit reset; never set -> stay absent.
MaxInstanceLifetime: clearOnUpdateRemoval(newLifetime, prevLifetime, 0),
A type whose update() has been audited for EVERY property also declares
removalHandledInUpdate: the properties whose removal update() handles
itself (a local reset, a diff, a required or create-only property, or a
value left in place that update() names in its own warning). The caller
then warns once per resource for any removed property in neither map,
naming it as left at its current AWS value AND saying CloudFormation would
reset it to its default — so a property with no such reset (FSx
StorageCapacity cannot shrink), or one whose reset is unmeasured, goes into removalHandledInUpdate with its
own warning, never into the shared line. A type without that entry is never
warned about, because cdkd cannot say what its update() does.
drift --revert injects nothing: its previous side is an AWS readback, not a
template, so UpdateContext.removedProperties is empty there. A local
clearOnUpdateRemoval still reads that readback, which is what reverts a
console change the baseline records as unset (a placeholder).
Checklist when writing or reviewing an update():
Classify the API: merge (absent = unchanged — most
Update*/Modify*calls) vs full-replace (the whole config object is replaced, or removal is handled by a dedicatedDelete*call, like the S3 bucket sub-config pattern). Only merge APIs need clear-on-removal — but a full-replace API has the MIRROR hazard instead, see Full-replace update APIs erase AWS-authored values.For each optional mutable property: what does AWS need to receive to get back to the CFn default? Common shapes: the documented default scalar (
3,128), an empty container ({ Variables: {} },[]), an empty string (''for description/KMS-key fields), or a sentinel ({ ApplyOn: 'None' },{ Mode: 'PassThrough' }).Never synthesize a reset for a field that was never set — that turns every unrelated update into a spurious (and sometimes invalid) write.
The reset SENTINEL can vary per API and per value KIND within one API, so verify it per key rather than picking one for the whole call (issue #1609 item 1). ELBv2's three attribute APIs take an identical
{Key, Value}list and disagree three ways:ModifyTargetGroupAttributesrejects an emptyValuefor EVERY key ("A target group attribute value must be specified"), whileModifyLoadBalancerAttributes/ModifyListenerAttributesaccept''for numeric and free-form-string keys but reject it for BOOLEAN / ENUM ones ("The value of 'deletion_protection.enabled' must be 'true' or 'false', but was ''"). So a single provider needs a per-key resolver — a documented defaults table with a fallback — not one sentinel. Three things make this worth its own bullet. The blast radius is the whole call, not the key: these APIs validate the entire attribute list, so ONE removed boolean fails the deploy — and it can fail the automatic ROLLBACK too, leaving no template-side way out. (The rollback re-runs the same diff with the sides SWAPPED. That turns a removal of key K into an ADDITION of K, so it is not the same refusal mirrored — but any key the two template sides do not share becomes a removal in the other direction, which is exactly how the live run failed on a SECOND key the forward pass never touched.) The previous side is not always a template, which decides what a safe reset value even is:cdkd drift --reverthands the provider thereadCurrentStatesnapshot aspreviousProperties, so an attribute AWS reports but the template never set can look REMOVED. Writing a documented default for each would silently reset a dozen live settings. Skip a key whose CURRENT value already equals the default — an untemplated key is by definition sitting at its default, while a genuinely templated one is not. Since issue #1626 the caller meets you halfway: when the reverted resource has NOobservedProperties(the baseline is the raw template, so cdkd cannot tell AWS-authored from out-of-band),runRevertMERGES every untemplated path into the desired bag it hands you, so those keys arrive with their AWS-current values on BOTH sides and your diff sees no change. (ELBv2 goes further: ondesiredFromAwsReadbackit sends no attribute removal at all, because no ELBv2 call removes a key, so an absence from even an observed baseline only means AWS was not returning the key at capture — issue #4147.) That covers the bulk case and the non-default residual the per-provider skip cannot — and, because it is on the desired side, it also covers a provider that replaces the bag wholesale and never readspreviousProperties. It does NOT cover the observed-capture baseline, where an absence IS an intentional removal — so keep the skip. And mocked unit tests cannot find any of it — they agree with whatever sentinel the provider chose. The ELBv2 arms shipped with two tests literally named "AWS-documented clear" / "empty-string default" that pinned a payload real AWS rejects; only an integ whose removal phase actually drops the key surfaced it. If you add a clear-on-removal arm, add the removal phase to a fixture in the same change.The CFn-reset premise itself is per-type — A/B it before writing the reset, because where CloudFormation does NOT reset, the reset is the bug. Removing a Cognito UserPool
Policiessub-key (SignInPolicy/PasswordPolicy, or the whole container) reaches UPDATE_COMPLETE in CloudFormation with every live value intact (measured us-east-1 2026-09-02, issue #1979) — CFn performs the same merge-semantics no-op the raw API does, so a cdkd-side reset would DIVERGE from template compatibility. The correct shape there is pass-through-nothing plus a WARN naming the sub-key and the explicit-declaration remedy (warnOnUnremovablePoliciesSubKeysincognito-provider.ts): under both A/B answers the only wrong option is silence, because the user who deleted the sub-key believes a change (often a TIGHTENING) landed.Sub-structures can carry the same hazard one level down: an object that is present but missing its inner key (e.g.
Environment: {}withoutVariables) may also read as "no change" — normalize it to the explicit clear shape (issue #1158; live-verified). Conversely some config objects are replaced wholesale (Lambda'sLoggingConfig: sending{LogFormat: 'Text'}alone resets an unspecified customLogGroupto the default — also live-verified), so verify per API, not by analogy. Lambda'sImageConfigis the same whole-replace shape (issue #1225, live A/B 2026-08-11):UpdateFunctionConfigurationwith{EntryPoint}alone CLEARED the previously-setCommand/WorkingDirectory, and dropping those two from a kept CFnImageConfigblock reached the same end state — so the provider passes the kept-partial block through verbatim and synthesizing per-sub-field clears there would DIVERGE from CloudFormation. Measure both halves: "does CFn reset it" and "what does the API do with a partial object" are two different questions, and only the second one tells you whether pass-through is already correct. A MERGING sub-shape is not automatically a cdkd bug — ECSDeploymentConfigurationis the counterexample, and it is the reason the question above must be asked as an A/B rather than answered from the API semantics alone (issue #1225, live A/B us-east-1 2026-08-13).UpdateServicegenuinely MERGES: a kept block carrying onlyMaximumPercentleftMinimumHealthyPercentand the circuit breaker at their live values. That is exactly the silent-drop shape — and a CloudFormation stack given the same partial block reached the IDENTICAL end state on a real resource UPDATE, so the retained value is CFn's behavior too and normalizing it away would DIVERGE. (The END STATE is what was measured; that CFn's handler submits the same partial toUpdateServiceis the inference.) The same run settled the whole-block removal the #1160 audit had left open for this field: CFn issued a real UPDATE and did NOT reset the configuration, so passingundefinedis right and adding a reset would have been the divergence. One level DOWN the answer flips, and the flip is the part worth remembering. A keptDeploymentCircuitBreakermissingRollbackhas the nested struct REPLACED, not merged — the SDK ACCEPTS it and a liverollbackwent true -> false. CloudFormation refuses that template up front (Model validation failed (... required key [Rollback] not found)), because the registry schema marks theDeploymentCircuitBreakerdefinition required [Enable, Rollback] andDeploymentAlarms(theAlarmsproperty) required [AlarmNames, Rollback, Enable]. Until issue #1802 cdkd enforced no nested required-ness, so it DEPLOYED that template and silently flipped the live setting — a divergence in the PERMISSIVE direction, NOT a parity row. The deploy pre-flight (src/provisioning/nested-required.ts) now refuses a PRESENT nested block missing a required member, but only for the types CloudFormation was MEASURED to enforce that on: a schema'srequiredlist alone is not evidence (CFn accepts anAWS::Logs::LogGrouptag with neitherKeynorValue). Do not file the row under the ASGInstanceMaintenancePolicydisposition: there AWS ITSELF rejected the partial, so cdkd failed loudly too and pass-through really was parity. The generalizable check is therefore two questions, not one — does the API merge or replace the nested struct, and does anything on cdkd's side refuse the shape the way CFn does? For a PRESENT block missing a required member, the answer is the pre-flight insrc/provisioning/nested-required.ts, which reads the fixtures' per-PATHnestedRequiredsection (inline nested objects included) and refuses only for its measuredCFN_ENFORCED_TYPES. Answering the second question no longer needs a livedescribe-type: since issue #1800 eachtests/fixtures/cfn-schemas/*.jsoncarries adefinitionRequiredsection (per definition, itsrequiredlist; the top-level list under the reserved#topkey), so a test can assert the actual lists a parity verdict rests on. Read a MISSING key as "nothing is required here" — a definition with norequiredlist, or an empty one, gets no entry. Note the section's ABSENCE does not mean "not captured": seven types legitimately require nothing anywhere and carry no section at all, so usegeneratedAt(>= 2026-08-13) to tell a captured fixture from a stale one. Note also that only a definition's OWN top-levelrequiredarray is captured — required-ness expressed through aoneOfcombinator (AWS::S3::Bucket'sTargetObjectKeyFormat) or on an inline nested object (AWS::WAFv2::WebACL'sFieldToMatch.SingleHeader) is absent, so for those shapes a missing entry means UNKNOWN rather than "CFn requires nothing"; measure against the live schema before concluding a parity verdict. That absence is itself load-bearing:LinearConfigurationandCanaryConfigurationcarry no required list, so a kept-but-partial block there IS reachable from a template CFn accepts — which is issue #1806, and was expected to be strictly worse than #1802 precisely because there is no CFn-side refusal to point at. It MEASURED as parity instead; the next paragraph is that measurement, and it is a good example of why the reachability of a shape does not predict its verdict.And REPLACE does not imply divergence — ask the second question before concluding one (issue #1806, measured us-east-1 2026-08-13 on the same type, SDK and CloudFormation A/B per block). The three DEPTH-1
DeploymentConfigurationblocks whose partials CFn ACCEPTS —LinearConfigurationandCanaryConfiguration(norequiredlist at all) andLifecycleHooks[], whose element requires onlyLifecycleStages(the partial measured there is one element'sTimeoutConfiguration) — are what made this the worrying case. All three replace the nested struct exactly asDeploymentCircuitBreakerdoes, but the absent member is filled with an AWS-side DEFAULT (StepBakeTimeInMinutes-> 6,CanaryBakeTimeInMinutes-> 10, the hook'sAction->ROLLBACK) and CloudFormation handed the same partial template reached the IDENTICAL end state every time — so the verbatim pass-through is PARITY and re-filling the member from the previous side would be the divergence. What separates them from the #1802 row is only the second question: nothing refuses the shape, on either side. Two further findings generalize:- A block reachable only under one MODE has to be probed in that mode, and
the mode may not be the one the issue names. These three are documented
as blue/green machinery, but
LinearConfigurationis gated onStrategy: LINEARandCanaryConfigurationonCANARY— AWS answers aBLUE_GREENservice carrying one withLinear configuration can only be present with LINEAR deployment strategy, so the fixture the issue prescribed could not have reached them at all.LifecycleHooksis not strategy-gated that way (measured underCANARY; its reach underROLLINGwas not probed). - The default-fill is not phantom drift, because the deploy engine
captures
observedPropertiesfromreadCurrentStateand the AWS-filled value lands in the drift baseline. Check that before reaching foreffectiveProperties: a value AWS computes belongs inobservedPropertiesby this file's own rule, and the capture already puts it there. It holds only while the property stays drift-COMPARED, so fence it — declaring the property ingetDriftUnknownPathslater would falsify the parity verdict with every send-side test still green.
The rest of the tree splits TWO ways, plus the array shape below — so neither "the other blocks are refused" nor "the rest is parity" is the thing to carry away, and
DeploymentAlarmskeeps the disposition above (the permissive divergence, now refused at pre-flight):- AWS ITSELF REFUSES (the ASG
InstanceMaintenancePolicydisposition — cdkd fails loudly, no nested required-ness check involved): a hook element missingLifecycleStages; a hook element missingTargetType(absent DEFAULTS toAWS_LAMBDA, which then demands a target ARN and role a PAUSE hook does not carry, so dropping ONE member arms a requirement for TWO others); andDeploymentCircuitBreaker.ThresholdConfigurationmissingValue. - AWS RETAINS it and CLOUDFORMATION RESETS it — the one DIVERGENCE the
sweep found, and the opposite polarity to the permissive one above:
DeploymentCircuitBreaker's OPTIONAL children (ResetOnHealthyTask, and the wholeThresholdConfigurationblock). Dropping them from an otherwise complete parent is a CFn-accepted partial that reaches AWS; from the identical baseline, with a sibling flipped in the same call so the update demonstrably applied, cdkd leftfalse/{COUNT, 7}intact while CloudFormation reset them totrue/{BOUNDED_PERCENT, 50}. So cdkd fails to apply a removal CFn applies — too STICKY rather than too permissive. Filed as issue #1861 and FIXED there: those two members — and only those two, and only while BOTH template sides still declare theDeploymentCircuitBreakerblock — are routed throughclearOnUpdateRemovalinECSProvider.resolveDeploymentConfiguration, so a declared-then-dropped member is now sent as its AWS default. Everything else in the property still ships as the verbatim pass-through the parity rows measured.
The
LifecycleHooksARRAY is replaced WHOLESALE, so a per-element drop is a non-issue. Three things generalize:- Enumerate the tree from the registry definitions, not from the blocks
the issue names. #1806 named three; the live schema has more, and the two
the issue never mentioned (
ThresholdConfiguration,ResetOnHealthyTask) are the ones that produced new behavior. - Check the CHILD's own
requiredlist before calling a nested partial reachable. A satisfied PARENT requirement does not make the child's partial CFn-accepted:ThresholdConfigurationcarriesrequired: [Type, Value], so CFn refuses it up front — the parity there comes from BOTH engines failing, not from agreeing on an end state. What decides the verdict is whether AWS ACCEPTS, not whether CFn refuses:DeploymentCircuitBreakermissingRollbackis refused by CFn and ACCEPTED by AWS, which is precisely why that row was the permissive divergence (#1802) and this one is not. - Members of ONE block can carry DIFFERENT semantics. In
DeploymentCircuitBreaker, the requiredRollbackreads as replaced while the optional children are retained. Measure per member; do not generalize a row to its siblings. - Where the API RETAINS an omitted optional member of a STILL-DECLARED
struct, cdkd diverges from CloudFormation, which treats the omission as a
REMOVAL and resets the member to its default. Scope that sentence
carefully, because two neighbouring shapes behave differently and both were
measured: removing the WHOLE property resets nothing under either engine
(the depth-0 row), and a member the template NEVER declared is left alone by
CFn too — an out-of-band value survived a CFn update that changed an
unrelated member — and it still survived when a SIBLING INSIDE the same
block was flipped, which is what excludes the competing reading "CFn
re-serializes a struct with defaults whenever its declared content
changed". So this is a previous-present / current-absent REMOVAL, the
semantic
clearOnUpdateRemovalimplements one level up, NOT a handler that materializes every default. Treat the member-vs-whole-property boundary as EMPIRICAL rather than derived: a top-level property is previous-present / current-absent too, and removing it resets nothing. TheLinearConfigurationfamily is parity because the API ITSELF default-fills there, so both engines land in the same place. And the two behaviors are indistinguishable until you probe a member whose live value DIFFERS from its default — measuring from a live value that happens to EQUAL the default proves nothing, which is a mistake this measurement made once and had to redo. Capture the resource's identity on both sides too (a replacement trivially shows defaults, and reads as a reset).
A note on HOW to probe, because it changes the answer: the AWS CLI refuses the
ThresholdConfigurationpartial CLIENT-side (botocore validates the required trait before sending), so it never reaches the service. The JS SDK cdkd uses declares the same member required in its TYPE (value: number | undefined, a non-optional key, unlike the optionaltargetType?:) but performs no runtime required check, so it serializes the partial and the refusal arrives from the SERVICE. The models agree; only where the validation runs differs — so probe through@aws-sdk/client-*, or a CLI probe reports the right verdict for the wrong reason, and reports the WRONG verdict entirely if the service would have accepted it.- A block reachable only under one MODE has to be probed in that mode, and
the mode may not be the one the issue names. These three are documented
as blue/green machinery, but
Unit-test three shapes: removed → exact reset value; never-present → stays absent; mixed kept/removed → kept fields pass through unchanged.
A per-key removal test (one key dropped from a still-present map) does NOT cover whole-block removal (the map itself dropped) — test both.
Walk the DEPTH-1 members too, not just the top-level properties. The removal audit above is normally run over a type's top-level properties, and a nested struct forwarded whole (a recursive case-flip, a spread, a raw cast) hides the same bug one level down: the API retains the omitted member while CloudFormation resets it.
ECSProvider.resolveDeploymentConfiguration(issue #1861) is the worked example — reuseclearOnUpdateRemovalwith the PREVIOUS side read raw, since only its PRESENCE is inspected. Two traps that both cost a round there: a member the template NEVER declared must stay absent (sending the default clobbers an out-of-band value CFn preserves), and the reset re-runs with the sides SWAPPED during a rollback replay (issue #1609), so key strictly on previous-present / current-absent rather than on "the two sides differ".Not every removal is a VALUE on the same call.
clearOnUpdateRemovalfits a property that maps to an input FIELD, so a reset is "send the default instead of omitting". A property whose apply is a separate API call needs a different call on removal, and there is no reset value to pass — theroute53-provider.tspair (issue #1160) is both spellings:HostedZoneTagsapplies viaChangeTagsForResourceand its removal is theRemoveTagKeysargument (a previous-minus-desired KEY DIFF, not a value), whileQueryLoggingConfigapplies viaCreateQueryLoggingConfigand its removal isDeleteQueryLoggingConfigon a sub-resource. Both looked like no-ops precisely because the apply helper took only the DESIRED bag and had nothing to diff against; the fix is threadingpreviousPropertiesinto the helper, keeping it optional socreate()keeps its existing REMOVAL behavior (it has no previous side, so nothing is ever removed — though a shape guard you add along the way does change what create accepts, so do not claim it is byte-identical). Gate the removal on the previous side actually having carried the thing, rather than probing AWS on every update — a config cdkd never created is drift, whichcdkd driftowns, not a removal reset. Two things bite specifically in this shape, both found by review on the route53 batch after its first integ had already passed:- A malformed value must not read as a removal. The desired side reaches
the helper through a tolerant reader, and anything the reader cannot use
collapses to the same emptiness a real removal produces — so an unresolved
intrinsic UNTAGS a live zone or DELETES a live config. Refuse a LOSSY read,
not merely a wrong container:
[{ Key: { Ref: 'X' } }]is genuinely an array, so a shape-only check lets the destructive case through one level down. Compare the parsed length against the raw length. And check what your ownreadCurrentStateemits before calling a shape malformed — route53's emitsQueryLoggingConfig: {}for "no live config", whichcdkd drift --revertfeeds straight back as the DESIRED side. - A failed removal is not self-healing, so it must not be swallowed. A
failed ADD is retried by the next deploy because the template still declares
it. A failed REMOVAL is not: the update returns success, state is rewritten
WITHOUT the property, and the next deploy's previous side no longer carries
it — so the value survives on AWS forever with
cdkd diffclean, which is the #1160 failure mode the fix existed to close. Throw on the removal branch even when the surrounding helper is warn-and-continue.
- A malformed value must not read as a removal. The desired side reaches
the helper through a tolerant reader, and anything the reader cannot use
collapses to the same emptiness a real removal produces — so an unresolved
intrinsic UNTAGS a live zone or DELETES a live config. Refuse a LOSSY read,
not merely a wrong container:
Full-replace update APIs erase AWS-authored values
A full-replace update API needs no clear-on-removal (see Update removal semantics) — omitting a key IS the reset. Its hazard runs the other way: whatever the payload omits is erased, including values AWS itself wrote and your template therefore has no representation for. No template-side diff hints at the loss, and no unit test can catch it, because the values only exist on the live resource.
The live case (GlueProvider, issue
#1461): UpdateTable replaces
TableInput wholesale, so an Iceberg table's Parameters.table_type and
Parameters.metadata_location — written by Glue at create time — were erased
by a deploy that changed only TableInput.Description, silently degrading the
table to a plain external table while the deploy reported success.
When a full-replace update sends a general-purpose bag (a Parameters /
Properties / Tags-shaped map AWS can write into), read the live resource
first and merge those entries back. Key the merge on "present in NEITHER
template side", using previousProperties:
| key in desired | key in previous | outcome |
|---|---|---|
| yes | (either) | the user's value wins |
| no | yes | the USER REMOVED it -> stays removed |
| no | no | AWS-authored -> preserved from live |
Restoring anything merely absent from the DESIRED side would make
user-authored entries unremovable — the mirror image of the bug — and would
break cdkd drift --revert's ability to clear a console-side addition (on an
observed-capture baseline that path passes the AWS-current snapshot as
previousProperties, so every live key lands in the previous column and
nothing is added back). On a resource with no observedProperties the same
path MERGES every untemplated live key into the DESIRED bag instead (issue
#1626) — previousProperties
is still the full AWS-current snapshot — so an untemplated live
key arrives in BOTH columns at its AWS-current value, so this table's first
row keeps it — which is the intended outcome there, since the raw template
cannot distinguish an AWS-authored entry from a console-added one.
State the price. The merge cannot distinguish an AWS-written entry from a
console-written one — neither appears on either template side — so every
out-of-band addition to that bag becomes permanent, and invisible to
cdkd drift once the next deploy folds it into the state baseline. That is an
acceptable trade against silently destroying an AWS-authored value, but it is a
real semantic change: document it on the resource type, and give users the
removal path (delete it directly, or declare-then-undeclare it so it becomes a
normal user-authored removal).
Close the TOCTOU window if the API lets you. Reading a value and writing it
back is not atomic; a concurrent writer landing in between is silently undone
by your write-back. Check the update request for an optimistic-concurrency
member (Glue's UpdateTableRequest.VersionId; the ConcurrentModificationException
in a command's documented error list is the tell) and send the version you
read, so a concurrent change fails loudly instead. Send it on EVERY update
rather than only on the ones that write back live values: the read runs
immediately before the write, so the version is never stale unless somebody
else really wrote in between, and an update that merged nothing still ships a
wholesale replace that a concurrent write would lose. (Scoping it was tried on
Glue and was wrong twice over — it left the empty-live-read case unguarded,
and its premise that a pure template push carries a stale version was false.)
When the API has no such member, say so where users will read it rather than
leaving the exposure implicit.
Prove the precondition, do not assume it. An AWS field named like a version
token is not necessarily enforced — Glue documents VersionId only as "the
version ID at which to update the table contents". Pin the semantics with a
real-AWS probe that advances the version out of band and requires the stale
replay to be REFUSED (and refused with a concurrency error, not any error).
Without that, an ignored token makes the whole guard a placebo that reads as
protection in review.
Two placement rules go with it:
Fail closed on a read failure. Only a definitive not-found may degrade to "no live values" (the update's own error is more actionable). Any other failure must throw, naming the required IAM action — silently skipping the merge reinstates the erasure the read exists to prevent. Wrap the original error as
causeso a transient throttle stays retryable.Throw OUTSIDE the update's
try, and order the read AFTER the pre-flight validation, so the typed error is not re-labelled by the catch wrapper and a refused update issues no extra API call. Move ONLY the read — leaving the payload-building code outside thetryas well turns a malformed-template crash into a rawTypeErrorwith no resource context.Check that every
provider.update()call site retries. Adding a read toupdate()makes it newly sensitive to throttling.deploy-engine/update-in-place.tsanddrift.tswrap their calls inwithRetry;rollback-executor.tsdid not until issue #1461, so the new read would have failed a rollback op that previously issued no read at all — and the best-effort catch there counts that as a failure and moves on. A transient failure on a RECOVERY path is the worst place to introduce one. A newwithRetrymust carry the two conventions the surrounding sites use, or it trades one bug for another: honorprovider.disableOuterRetry(CustomResourceProvider/NestedStackProviderset it AND implementupdate()— re-invoking a Custom Resource derives a fresh RequestId + pre-signed URL and strands the first response at an S3 key nobody polls), and threadisInterrupted/onInterrupted(a rollback polls interrupts only BETWEEN ops, so an un-threaded probe leaves Ctrl-C dead for the whole backoff schedule).The second of those is now MECHANICALLY ENFORCED (issue #2053, after a count found 11 sites under
src/provisioning/providers/**and the pair threaded at exactly one of them).vp run audit:withretry-interrupt:check(scripts/check-withretry-interrupt.ts, a CI step) fails on anywithRetryundersrc/provisioning/**whose options object literal does not declare BOTH names. Half a pair is a defect too: withoutonInterrupted,withRetrythrows a bareError('Interrupted')that names no resource, and withoutisInterruptedthe other half is never called. A conditional spread does not count — it reads as threading to a reviewer while being absent whenever its guard is falsy — and an options bag the checker cannot READ (an identifier, a spread) is reported rather than skipped, so inline it.Do not hand-roll the watch. There is exactly one:
startInterruptWatchinsrc/provisioning/interrupt-watch.ts, and the critic above checks PROVENANCE, so a locally-built object with the right property names fails. Four module-local copies existed briefly and could not agree with each other, which is what made two of its properties wrong; all four are gone. Call it per WAIT, threadwatch.isInterrupted/watch.onInterrupted, anddispose()in afinally.Four properties of that module, each load-bearing and each the fix for a concrete bug:
Per WAIT, never on
this. Providers are registered as SINGLETONS serving concurrent resources, so provider-level state is some other resource's.onInterruptedreturns anInterruptedWaitError, not a bareError.deploy-engine/execute.tsdecides whether to ROLL BACK by asking what the failure was, and its ownInterruptedErroris engine-internal (not re-exported, and providers do not import the engine), so a bareErrorfrom a provider read as a genuine resource failure and rolled the whole stack back on Ctrl-C. UseisInterruptedWaitErrorto recognise one. It walks thecausechain — every provider catch re-wraps — with avisitedset and NO depth ceiling, because the chain grows by one per nested-stack level and a ceiling sized against today's nesting silently reinstates the rollback bug.The latch is STICKY, and only a COMMAND clears it. A wait STARTED after the signal begins is already interrupted. Clearing it when the last watch is disposed looks equivalent and is not: sequential waits always empty the live set between them, so a
GlobalTabledelete running the #1521 gate, then the index-busy loop, then the gone-wait left the two multi-minute waits after the gate DEAF.forwardSigtermToSigint()opens and closes the command scope. The process listener is likewise never removed BETWEEN waits — one torn down in that gap cannot record a signal landing in it.It arms only inside a command that OWNS interrupt handling, signalled by that scope rather than by counting SIGINT listeners. Any listener disables Node's default terminate, so a command with none of its own (
cdkd drift --revertreachesprovider.update) must not gain one here. A listener COUNT is the wrong test and was the first cut:driftruns at concurrency 4, and a concurrent CloudFront / ACM / Route53 wait installs a transient SIGINT listener, so a wait starting in that window armed the shared handler permanently.The handler force-quits when it is the LAST SIGINT listener. A command scope is not the same as a live graceful path:
destroy.tsregisters no handler anddestroy-runner.tsremoves its own in afinally, so between two stacks of a multi-stack destroy the watch is alone — and merely latching there SWALLOWS the Ctrl-C, letting the next stack delete on. Alone, it does what Node would have done with no listener at all, plus acdkd force-unlockhint it cannot verify but cannot afford to omit.The corollary is a rule for any command holding a lock: release it BEFORE unregistering your SIGINT handler. Unregister-first leaves a window where the command holds the lock with no handler of its own, and the force-quit turns a Ctrl-C there into a stranded lock for its full 30-minute TTL.
destroy-runner.tshad it backwards in its mainfinallyand correct in its strong-ref refusal path; both now read the same way, anddestroy-runner-lock-release-ordering.test.tspins it by observing which listeners are registered at the momentreleaseLockis entered.rollback.tshad the same defect until #2118, whose fix adds one thing worth copying into any command you write: both unregistrations sit in a nestedfinally, because the.catch()on a release covers a rejection and not a synchronous throw. It is pinned byrollback-lock-release-ordering.test.ts, which measures the SIGINT and the SIGTERM listener sets SEPARATELY — a fix that reordered only the SIGINT half would leaveunforwardSigterm()running before the release, and CI cancellation delivers SIGTERM, so that half strands the lock on its own. What is NOT a rule is the order within the teardown pair, and the first cut of that fix asserted it in four places before a review round measured it: the two calls are adjacent and synchronous, so no signal can be delivered between them, andunforwardSigterm()first does not empty the SIGINT set anyway because the command's own handler is still registered. Order them for readability; the requirement is release-before-unregister.deploy-engine/deploy-flow.tsis now the only site that still unregisters first, and is safe only becausedeploy.tsholds a handler that outlives it — stated at that call site, because a surviving instance of a corrected anti-pattern has to explain itself.And the corollary's own corollary: keeping the handler armed for longer moves the window a signal can arrive in, so re-check anything that READS the interrupt.
destroy-runner.tsassignsresult.interruptedonce, inside itstry— so arming across the teardown made a first Ctrl-C there setlock.interruptedafter the only read, leaving the flag false and lettingdestroy --alldelete the next stack. It re-syncs withresult.interrupted ||= lock.interrupted && statePreservedat the end of thefinally; thestatePreservedgate keeps a stack whose state was already deleted from reportinginterrupted. That is tactical: the real defect was that flag being the only channel, which #2117 closed with a command-scoped interrupt handler the--allloop reads live.
An interrupt is not a failure — but it is not a reason to leave AWS state behind either. Two shapes to check in any wait you add, and they pull in opposite directions:
- A best-effort
catchthat swallows into awarnwill blame AWS for a user abort and then carry on to the next write.applyAutoScalingDiffre-throws instead, which stops the burst. - A partial-create cleanup arm must STILL run its cleanup delete on an
interrupt. "Ctrl-C must not delete what you just made" is the intuitive
answer and it is wrong here, because
create()is throwing: the physical id never reaches state, the rollback journal recordsphysicalId: undefined(unless the arm callsmarkCreatedBeforeFailure, above), and the rollback executor classifies itskip-failed-unknown. Unmarked, nothing holds the id, so the choice is delete vs orphan forever — and the orphan fails every later deploy on a name collision that neither rollback nor destroy can reach. Both the ELBv2 Listener and Cloud Map Service arms clean up, and each prints the physical id plus a manual delete command BEFORE attempting it, so a process killed mid-cleanup still leaves the user a handle.
An interrupt must also never be read as "already deleted". Both delete paths decide a resource is already gone by SUBSTRING-matching the error message, and an interrupt's message embeds a name the user chose — so a logical id containing
NotFoundExceptionused to drop a live resource's state row.destroy-runner.tsand the deploy engine's delete arms (deploy-engine/update-in-place.ts,deploy-engine/delete.ts) all checkisInterruptedWaitErrorahead of that match; any new message-based classifier on a delete path needs the same guard.A hand-rolled poll on the same path owes the same treatment, and issue #1952 is why the rule is stated for the PATH rather than per call: interrupting one wait while the next one still sits out its cap buys nothing. What a poll DOES on the signal is its own call: throw when nothing has been accepted yet, stop waiting when the operation is already in AWS's hands (see
waitForReplicaGone/waitForTableGoneindynamodb-globaltable-provider.tsfor both answers on one path).
Only bags AWS actually writes into qualify. A purely user-authored bag
(Glue's JobUpdate.DefaultArguments, ConnectionInput.ConnectionProperties)
must NOT be merged — preserving a console-side addition there would break
drift --revert. Audit the sibling update APIs on the same provider when you
fix one; the divergence is per-bag, not per-provider.
Returning attributes
Return attributes accessible via Fn::GetAtt:
return {
physicalId: bucketName,
attributes: {
Arn: s3BucketArn(bucketName, region),
DomainName: s3BucketDomainName(bucketName, region),
RegionalDomainName: s3BucketRegionalDomainName(bucketName, region),
},
};
Never hardcode arn:aws: or amazonaws.com in a value you build: outside the
commercial partition (aws-cn, aws-us-gov, ...) both are wrong, and the value
is still structurally valid, so nothing downstream catches it. Derive the
partition and URL suffix from the region with derivePartitionAndUrlSuffix
(src/utils/aws-partition.ts), or call a shared builder such as the
src/utils/s3-endpoints.ts ones above, which the SDK provider and the resolver
both use (and the Cloud Control route for Arn), so every route records the
same value.
getAttribute() for live Fn::GetAtt resolution
Beyond the initial create/update return value, providers should implement
getAttribute(physicalId, resourceType, attributeName, logicalId) so that live
attribute reads succeed when the value is not in cdkd state — specifically
the cdkd orphan per-resource flow, which splices each referenced attribute
into sibling references. It always reads live first. When that read answers,
it substitutes the orphan's recorded attribute instead wherever cdkd's own
Fn::GetAtt resolution would serve that value, and keeps the live answer only
when the record lacks it or holds a value that cannot be spliced (a
redaction mask, a {{resolve:...}} reference, a stale placeholder ARN, a VPC's
Ipv6CidrBlocks, an impossible empty value). A credential-named attribute
(SecretAccessKey), a known secret-valued attribute (an AppSync API key's
ApiKey, an IPAM verification token's TokenValue, an IVS stream key's
Value), a value holding a
credential-named key, and any custom-resource attribute are never taken from
the record, since the value may be a plaintext secret. A live read that fails
or answers nothing leaves the reference unresolvable without --force, as it
always did, so a provider without getAttribute gets nothing from the record
either. --force's cached fallback does not splice those secret classes
either, nor a mask or a {{resolve:...}} reference: the reference stays
unresolved (#4602). A live
read addresses the resource by its recorded name, so after the resource was
deleted and another one took that name it describes the newcomer; the recorded
value is the one cdkd's own Fn::GetAtt resolution would choose. A recorded
attribute AWS changes later (an instance's PublicIp) can be stale.
Conventions:
- Return
undefinedfor unknown attribute names. Do not throw. - Where a provider does throw (several existing ones still refuse an unknown
attribute), its
ProvisioningErrortakeslogicalIdin its logical-id slot andphysicalIdafter it: the retry classifiers anchor on that slot, and a physical id can be built from a secret (go-to-k/cdkd#4222). - Treat
*NotFoundexceptions asundefinedrather than re-throwing — the live fetch is best-effort, and under--forcecdkd orphanfalls back to the cachedstate.attributeswhen the live resolution comes back empty. - Prefer derivation from
physicalIdwhen CFn returns derivable values (S3 Bucket DomainName/Arn, SNS Topic name from ARN tail, SQS QueueName from URL tail) so the call is free.
Known coverage gaps (deliberate)
The following CloudFormation Fn::GetAtt return values are documented but
not implemented in cdkd's getAttribute(). They require a separate AWS
API call beyond what cdkd already makes, are rarely referenced from CDK
code, or both. If a real-world stack hits one of these, file an issue —
the small additional call is reasonable to add.
| Resource | Unsupported attribute | Why deferred |
|---|---|---|
AWS::SQS::Queue |
(none) | All three CFn return values are covered. |
AWS::S3::Bucket |
(none) | All five CFn return values are covered. |
readCurrentState() for drift detection
readCurrentState(physicalId, logicalId, resourceType, properties?, context?) returns the AWS-current snapshot of a resource for cdkd drift and cdkd state refresh-observed. The drift comparator walks state's top-level keys only (intentionally — to avoid surfacing every FunctionArn / RevisionId / LastModified / etc. that AWS auto-attaches to every response). That design has one consequence the provider author MUST account for:
Any user-controllable top-level CFn property
update()can mutate must be emitted with a placeholder when AWS returns the field as undefined / empty.
If the provider omits the key on the empty path (e.g. if (cfg.Environment?.Variables) result['Environment'] = ...), then on a resource that was deployed WITHOUT that key in its template, state.observedProperties never carries the key — and the comparator's state-keys-only walk skips the field forever. A user adding the property in the AWS console after deploy is silently invisible to drift.
Use these placeholders consistently:
| Type | Placeholder | Example |
|---|---|---|
| Array | ?? [] |
result['ManagedPolicyArns'] = arns; (after building the list) |
| Map / object (when AWS returns the whole object as undefined) | ?? {} |
result['Cors'] = cors; (after building, even if cors ended up empty) |
| Optional string | ?? '' |
result['Description'] = resp.Description ?? ''; |
| Boolean / numeric scalar | ?? <semantic-default> |
Status: resp.Status ?? 'Suspended', BlockPublicAcls: cfg?.BlockPublicAcls ?? false |
| Tags map | ?? [] (already covered for Tags by PR #145) |
result['Tags'] = normalizeAwsTagsToCfn(...); |
When the guard is justified — keep it:
- Immutable on create —
BucketName,Lambda Runtime(when create-time-only),IAM RoleName. The field can't change at all; AWS returning undefined is a wire-layer artifact, not a "user could add this." Skip emit. - AWS-managed read-only —
FunctionArn,RevisionId,CodeSha256, timestamps. These are not in the CFn template; cdkd state never carries them. They should NOT be inreadCurrentStateoutput at all. - Write-only —
Code: { S3Bucket, S3Key },SecretString,LoginProfile.Password. AWS does not return these on read. Declare viagetDriftUnknownPaths()so the comparator skips the entire subtree (see "Known coverage gaps" below).
Wire-layer filtering — the drift comparator does NOT apply per-type denylists for SDK provider results (those are reserved for the CC-API fallback path). If your provider's SDK response includes AWS-managed fields you don't want to surface, do NOT assign them in the first place.
Test convention (mandatory for any provider with readCurrentState): every provider test file MUST have an it('emits placeholders for every user-controllable top-level key on AWS minimum response') block that:
- Mocks the SDK to return the resource exists with all optional fields undefined / empty (just required fields like Name / ARN).
- Calls
readCurrentState(physicalId, logicalId, resourceType). - Asserts
Object.keys(result).sort()matches the complete expected key list for that resource type — not a subset. - Spot-checks the placeholder values for the most fragile keys (
?? ''strings,?? []arrays,?? {}objects,?? <semantic-default>scalars).
Example template:
it('emits placeholders for every user-controllable top-level key on AWS minimum response', async () => {
mockSend.mockResolvedValueOnce({
/* SDK response: required fields only, all optionals undefined */
});
const result = await provider.readCurrentState('phys-id', 'L', 'AWS::My::Type');
expect(Object.keys(result ?? {}).sort()).toEqual(
['Key1', 'Key2', /* ... complete list ... */ ].sort()
);
expect(result?.Key1).toBe(''); // string placeholder
expect(result?.Key2).toEqual([]); // array placeholder
expect(result?.Key3).toEqual({}); // object placeholder
});
See tests/unit/provisioning/lambda-function-provider-readcurrentstate.test.ts and tests/unit/provisioning/cognito-provider-readcurrentstate.test.ts for canonical examples.
This is the structural defense against the "provider author forgets to emit a key" regression class. Without it, the bug only surfaces when a user runs drift on a resource configured exactly the way the test missed (and PR review missed). The test makes silent regression mechanically impossible — a refactor that drops a placeholder fails the key-set assertion immediately.
The properties argument is the DESIRED side, and gating an emission on it is a fix with a TRANSITION COST — do not reach for it alone (issues #1742 / #1760). Every caller passes the resource's state-recorded / template-resolved bag (drift.ts, the deploy engine's observed capture, import.ts, state.ts), so a provider CAN ask what the user actually declared, and the temptation is to use that to stop emitting an AWS-COMPUTED value the desired side can never carry.
The half that is easy to miss: observedProperties bags already in S3 were written by a binary that DID emit the member. The moment the new binary stops, the comparison is baseline has it vs aws side undefined, and the first cdkd drift after the upgrade reports a one-sided phantom on every such resource. It does not self-heal — the observed capture runs only on CREATE / UPDATE, and the auto-refresh skips a record that already has a capture (the one it revisits, a baseline holding a *** mask, is re-captured only when the readback reproduces that baseline, which a missing member prevents) — so the window is however long until the user's next deploy, which is exactly when they are running drift instead. Measured live on AWS::DynamoDB::Table (#1760).
So the emission gate needs a companion that removes the path from the COMPARISON too:
- For a TOP-LEVEL key, declare it in
getDriftUnknownPaths(resourceType, properties)— per-resource via the #1602 seam, so it is ignored on BOTH sides when the template declares nothing and still compared when it does. That covers the transition and the steady state with one declaration. - For a PER-ELEMENT key there is no such declaration: an ignore-path never crosses an array (see the divergence note below), and declaring the enclosing array instead switches drift off for the whole subtree. That case needs a BOTH-SIDES normalizer in the
drift-protocol-normalize.tsmould, and a live test seeded with a STALE observed baseline — a fresh-deploy fixture takes both sides from the same readback and structurally cannot exercise the transition.AWS::DynamoDB::GlobalTable'sGlobalSecondaryIndexes[].WarmThroughputis exactly this shape, and is now the WORKED example rather than the open question: PR #1859 closed it through thecanonicalizeDriftPropertiesseam (issue #1784), which IS the both-sides normalizer this bullet asks for — the provider names the member in a closedDRIFT_STRIPPED_INDEX_MEMBERStable and strips it from each bag, so an already-written baseline and a post-fix readback converge on carrying no member at all. Copy the NORMALIZER half's shape — but note it is only half of what this bullet asks for: the live test seeded with a STALE observed baseline was NOT shipped with it and is still outstanding (issue #1939), so #1859 is the worked answer for the mechanism and an open question for the proof. Take the two things that come with it as well: the emission change and the normalizer had to ship in ONE change (landing the readback half alone is precisely the stranding described above), and the accepted cost is that a DECLARED per-element value stops being reported entirely — the hook sees one bag with no desired-side reference, so it cannot express the declared-gategetDriftUnknownPathsgives you for a top-level key. The siblingAWS::DynamoDB::Tabletype adopts the same seam with a narrower key: it strips only members an SDK index DESCRIPTION carries and a CFn value never does (IndexStatus,NumberOfDecreasesToday, aWarmThroughputwithStatus, ...), so a baseline written by an older binary converges while a template or a current readback passes through by identity and keeps every declared value reported.
Contrast an ORDERING fix, which has no such asymmetry: getDriftUnorderedPaths canonicalizes both sides, so an existing baseline and the readback converge rather than diverge, and it can ship on its own.
getDriftUnknownPaths() for unreadable fields
When AWS does not return a field that cdkd state stores (write-only fields, or a CFn property whose round-trip back to the template shape isn't worth implementing yet), declare the path so the comparator skips it instead of firing guaranteed false-positive drift on every clean run:
getDriftUnknownPaths(): string[] {
return ['Code']; // Lambda::Function: pre-signed URL only
// or ['SecretString', 'GenerateSecretString']
// or ['RedshiftDestinationConfiguration.Password'] // Firehose: write-only, AWS never returns it
}
The comparator does exact-match + entry + '.' prefix-match — listing 'Policies' skips Policies, Policies.Foo, Policies[0].PolicyDocument, etc. Pair this with a docstring explaining why the field is unreadable so a future PR can lift the gap.
Scoping a path to a SUBSET of a type's resources. Some fields are unreadable only for resources in a particular configuration, and declaring them unconditionally would switch drift detection off for the resources where AWS does return the value. The method therefore takes an optional second argument — the resource's state-recorded properties, which cdkd drift passes — so the answer can be per-resource:
getDriftUnknownPaths(resourceType: string, properties?: Record<string, unknown>): string[] {
// AWS silently discards TlsConfig on a PUBLIC ApiGatewayV2 integration and
// never returns it; on a VPC_LINK (private) one it does, so keep comparing.
if (
resourceType === 'AWS::ApiGatewayV2::Integration' &&
properties !== undefined &&
properties['ConnectionType'] !== 'VPC_LINK'
) {
return ['TlsConfig'];
}
return [];
}
DynamoDBTableProvider is the other consumer, and the clearer illustration of the absent-bag rule below. It scopes on whether the template declares the property AT ALL rather than on a sibling's value — WarmThroughput is returned as unknown unless the desired bag declares it — and its predicate deliberately answers DECLARED for an absent or uninformative bag, so the path stays COMPARED:
function declaresWarmThroughput(properties?: Record<string, unknown>): boolean {
// A bag that was never populated is NOT evidence the template declared
// nothing, so answer 'declared' and keep comparing. A wrong DROP here is
// unrecoverable phantom drift; the residual is a loud revert failure.
if (!desiredBagIsInformative(properties)) return true;
return properties !== undefined && wasSentWarmThroughput(properties['WarmThroughput']);
}
It also shows what the seam is FOR beyond a configuration split (issues #1760 / #1602): an AWS-computed value that an earlier binary froze into observedProperties would otherwise drift forever, and per-resource scoping is what retires those records without switching the key off for the type.
Two rules for this shape (issue #1602):
- Tolerate an absent bag. Other callers may have no properties to pass, so fall back to the type-level answer rather than assuming a shape. Default to COMPARING when you cannot tell — hiding real drift is the worse failure, exactly as for
getDriftUnorderedPaths()below. - Say so at write time. A value AWS discards is still worth a
logger.warnon create / update. Declaring the drift path only removes the false positive; without the warning the user never learns the field is inert. Warn, never throw — the update path can be a state-record replay (rollback-executor.ts/drift --revert), where a refusal would leave the resource un-revertable.
getDriftUnorderedPaths() for unordered sets (strings AND objects)
The drift comparator compares arrays positionally, and AWS does not guarantee element ordering across reads. The shared normalizer (src/analyzer/drift-normalize.ts) already auto-canonicalizes two shapes for every type — {Key,...}[] tag lists and arrays whose every element is an AWS resource id (subnet-…, rtb-…) or ARN — but plain-string arrays and non-tag object arrays are deliberately left untouched, because either can be order-significant.
When your readCurrentState emits such an array that is semantically an unordered SET, declare its path so the comparator sorts it on both sides:
getDriftUnorderedPaths(resourceType: string): string[] {
if (resourceType !== 'AWS::FSx::FileSystem') return [];
return ['WindowsConfiguration.Aliases']; // DNS alias names
}
Path matching uses the same shared matchesPathPrefix rule as getDriftUnknownPaths() (exact match, or entry followed by .). A plain entry is a subtree declaration — 'WindowsConfiguration.Aliases' covers that path and everything beneath it.
Append [] to declare the path ALONE (issue #1783). Because this walk descends into array elements (see the divergence note below), a subtree entry for an object list also claims every array nested inside its elements — and that is frequently wrong: AWS::DynamoDB::Table.GlobalSecondaryIndexes is an unordered set keyed by IndexName, but each element carries a KeySchema that is order-SIGNIFICANT (HASH before RANGE), so 'GlobalSecondaryIndexes' would sort the per-index key schema too and hide a real key change. 'GlobalSecondaryIndexes[]' sorts the list and stops there. The marker is understood by getDriftUnknownPaths() as well, so the two lists keep one spelling; a bare '[]' matches nothing.
Two element shapes are sorted, each only when the array is homogeneous in it: every element a plain string (sorted lexically), or every element a plain object (sorted by a key-order-independent canonical JSON — issue #1620). Key order deliberately does not participate: AWS's readback order for an object's own keys is no more guaranteed than its order for the list, so sorting on a raw JSON.stringify would reintroduce the phantom drift the pass exists to remove. A MIXED array, and a nested array inside a declared path, are both left alone — so a mis-declared path can never reorder a heterogeneous or array-valued list.
ElasticLoadBalancingV2::TargetGroup.Targets is the object-array case. It also shows the two things a readback of an unordered set has to get right that no amount of sorting fixes.
Transient lifecycle states. DescribeTargetHealth keeps reporting a just-deregistered target as draining for minutes, so including one would freeze it into the deploy-time observedProperties snapshot and produce permanent phantom drift against every later read. The provider excludes exactly draining and includes every other state — initial, unused, unhealthy, unavailable are all REGISTERED targets, and health is not registration.
Who owns the value. A target group fronting an ECS service or an ASG declares no Targets at all: the sibling resource registers them and re-registers as it scales. Comparing that list would report drift on every scale event of an untouched stack, and --revert would deregister the tasks the service just placed (the #1498 class). So the provider keeps Targets in getDriftUnknownPaths() per resource — using the properties-bag argument described above — and only drops the entry when the template actually declares Targets. An explicit Targets: [] counts as a declaration and stays compared. Reach for this scoping whenever a property can legitimately be authored by a different resource rather than by the template.
One semantic divergence from getDriftUnknownPaths(), required for this pass to work: the comparator compares arrays wholesale and never descends into elements, so an ignore-path can never cross an array. This normalizer does descend into array elements, giving each the parent's path. A path like 'Items.Aliases' is therefore meaningful for getDriftUnorderedPaths() (it reaches an Aliases array inside each Items element) while being inert as an ignore-path. The divergence is strictly more permissive.
Only declare a path AWS documents — or you can demonstrate — as order-insensitive. The failure modes are not symmetric: an undeclared unordered set produces a visible false positive the user can see and correct, but declaring an order-significant list silently hides real drift, which is the worse failure for a drift tool. FSx's SelfManagedActiveDirectoryConfiguration.DnsIps is left undeclared for exactly this reason (DNS resolver lists are conventionally preference-ordered and AWS documents no set semantics), as is ElastiCache's PreferredAvailabilityZones (documented as positionally aligned to node index).
Do NOT sort inside the reverse-mapper instead. It looks like a one-liner, but it breaks the properties fallback baseline: runDriftForStack uses observedProperties as the baseline only when present and falls back to the template properties otherwise, so for a resource deployed before observed-capture the baseline would be the user's TEMPLATE order while the read side is sorted — manufacturing drift instead of removing it. The normalizer runs on both comparison sides, which is exactly the property needed.
canonicalizeDriftProperties() for a member INSIDE an array element
The two seams above are declared as PATHS, and a path can never cross an array: calculateResourceDrift compares arrays wholesale via deepEqual and never descends into elements, so isIgnoredPath is never asked about a member of one. The only suppression an ignore-path can express there is the WHOLE array — which for GlobalSecondaryIndexes means never detecting an out-of-band index add / remove / capacity change again, permanently, to remove a one-time report. That trade is not worth making.
canonicalizeDriftProperties(resourceType, properties) (issue #1784) is the seam for exactly that shape. cdkd drift applies it to both comparison sides, last in the normalization chain, so a provider can strip the AWS-managed member from the baseline as well as from the readback — an old observedProperties record and a new readback then converge with no ignore-path and no lost detection:
canonicalizeDriftProperties(
resourceType: string,
properties: Record<string, unknown>
): Record<string, unknown> {
const indexes = properties['GlobalSecondaryIndexes'];
if (!Array.isArray(indexes)) return properties; // identity when nothing applies
return {
...properties,
GlobalSecondaryIndexes: indexes.map(({ ItemCount, IndexArn, ...rest }) => rest),
};
}
Four rules, each load-bearing:
- There is deliberately no
sideparameter. One-sided normalization is the mistakesrc/analyzer/drift-normalize.ts's header already records — it manufactures drift on theproperties-fallback baseline, where the baseline is the user's raw template. The same pure function must see both bags. - Pure, synchronous, no AWS calls, NON-mutating. It runs inside the comparison, and
--revertdiffs against the raw, uncanonicalized AWS bag — mutating the input would corrupt it. --acceptpersists your output. It writes each reported change'sawsValue, and those come from the canonicalized bags — so on a path that still drifts after canonicalization,--acceptfreezes the stripped shape intoobservedProperties. For the intended use (dropping a member the readback no longer reports) that is the right direction, but do not strip anything you would be unwilling to see removed from state.- Position in the chain matters. The hook runs after the principal /
IpProtocolpasses and before the tag-list / id-array / unordered-path passes, which live insidecalculateResourceDrift. That order is required: an element strip must precede the unordered sort, or the two sides' canonical sort keys are computed over different member sets and diverge. - Return the input by identity when nothing applies, so an unaffected resource pays nothing.
- Reach for it only for a per-array-element member. A top-level key a readback stopped emitting is
getDriftUnknownPaths()'s job; a difference that is only about ORDER isgetDriftUnorderedPaths()'.
canonicalizeDriftPair() when the readback's SHAPE changes
When readCurrentState starts emitting something it used to omit, every observedProperties record captured before the change lacks it, so the first cdkd drift after upgrading reports it on every untouched resource. canonicalizeDriftProperties() cannot absorb that: it sees one bag, so it cannot tell a legacy baseline from a current one, and stripping the new member from both sides would hide its drift forever.
canonicalizeDriftPair(resourceType, baseline, aws) (issue #3573) gets both bags, after the per-side hook. Key the rule on the BASELINE's shape: a legacy baseline selects the absorption, and a current-shape one is compared in full. State is not rewritten; the next deploy's capture moves the record to the current shape. The GlobalTable provider is the example: its readback gained the local (deploy-region) replica entry, and a baseline with no local entry is completed with the readback's.
Complete the baseline from the readback; do not drop from the readback. Two write paths consume these bags:
cdkd drift --revertpasses its DESIRED bag, the recorded baseline, through the same hook against the raw readback and sends the returned baseline toupdate(). A member missing there is a REMOVAL to the provider. A legacy GlobalTable record sent without its local entry untagged the local table.--acceptwrites each change'sawsValuefrom the AWS side. Leaving that side intact means an accept stores the current shape and heals the record.
The reverse change, a readback that STOPS emitting something, is the one case where the hook drops from the baseline. AWS::DynamoDB::Table reads back only the WarmThroughput members cdkd sends (issue #3777), so a baseline block is trimmed to the members the recorded declaration sends, read from the properties argument the hook also receives. Key such a trim on the declaration, never on what the readback omitted alone: a member AWS transiently fails to report must stay reported (the ELBv2 case below reads the readback only for a key the template never declared). That is safe only because a warm-throughput member missing from the desired side is not a removal: every send site coerces the block to its usable members, and AWS cannot lower or unset warm throughput. Before copying it, check that the same holds for your member.
The ELBv2 attribute bags (LoadBalancerAttributes, TargetGroupAttributes, ListenerAttributes) are the one trim that reads the readback too (issue #4144). A baseline entry is dropped only when its key is undeclared AND missing from the readback. That is safe because no ModifyAttributes call removes a key, and the key set otherwise moves only with configuration that is compared on its own (a listener's protocol), so a key AWS stopped reporting is the service's own change, never drift. An empty readback bag is treated as a failed read and trims nothing. A declared key stays reported, as does a value change on a key both sides hold and a key only the readback holds. The last one cannot be told apart from an operator's value in one read, so it stays reported, but update() on desiredFromAwsReadback never sends it as a removal (issue #4147): the removal value would be a guess, and a rejected one fails the whole ModifyAttributes call.
It is non-mutating and returns both inputs by identity when nothing applies. It is async only so a provider can resolve the same client region its readback used; it issues no AWS call.
When there is NO observed baseline at all
The paragraph above describes the round-trip when observedProperties exists.
When it does NOT — a resource deployed before observed-capture shipped —
--revert falls back to the raw TEMPLATE as its desired side while the
previous side is still the AWS-current snapshot. Because the revert overlays a
drifted top-level subtree wholesale, every AWS-authored key inside that subtree
which the template does not declare is DROPPED (issue #1478). Glue
Table.Parameters is the discovered case — table_type / metadata_location
are written by AWS, not by the template, so a --revert de-Icebergs the table —
but the exposure is general to any resource where AWS writes into a bag the
template does not fully declare.
The chosen semantic is warn and proceed: cdkd drift --revert lists the
paths it will drop as part of the plan (before the confirmation prompt, and
under --dry-run) and then reverts. This is a drift concern, NOT a provider
one — the baseline choice affects every provider that implements
readCurrentState, so fixing it inside one provider would make that provider
diverge from its siblings. Nothing is required of a provider here; the note
exists so that reading only the round-trip paragraph above does not leave you
believing the desired side is always an observed snapshot.
Two failure modes when an always-emit placeholder round-trips through update()
cdkd drift --revert round-trips observedProperties (the snapshot readCurrentState produced) back through provider.update. That code path is what surfaces every shape-mismatch bug between the read side (readCurrentState output) and the write side (AWS create/update API input). Two failure classes have been observed; both must be designed around BEFORE adding a new readCurrentState.
Class 1 — type-discriminator-dependent fields. A field is only valid on AWS when a sibling discriminator says so. Examples: SQS DeduplicationScope / FifoThroughputLimit (FIFO-only — FifoQueue=true), SNS FifoThroughputScope (FifoTopic=true), AppSync DataSource shape (DynamoDBConfig / LambdaConfig / HttpConfig discriminated by Type). Emitting a '' placeholder for these on a discriminator-false resource means --revert pushes it back and AWS rejects with "You can specify X only when Y is set to true". Fix: guard the emit on the sibling discriminator — only emit when the discriminator is true. Pattern documented in feedback_always_emit_check_type_discriminator.md. Drift detection is not lost: the discriminator-false state cannot legally have the field on AWS, so console-side ADD is impossible. When the discriminator is N-way rather than boolean — a set of mutually-exclusive <Variant>Configuration blocks selected by a type field, as in AppSync's Type or FSx FileSystem's FileSystemType (LustreConfiguration / WindowsConfiguration / OntapConfiguration / OpenZFSConfiguration) — the same rule reads: emit EXACTLY the one block the discriminator selects, and emit it unconditionally (so the always-emit contract still holds for the one legal block, {} included), never the others. The key-set test is then written once per discriminator value, each asserting its own block is present and the rest absent — see tests/unit/provisioning/providers/fsx-filesystem-provider.test.ts.
The "console-side ADD is impossible" clause holds ONLY when the discriminator is an INDEPENDENT sibling (issue #1565). Where the "discriminator" is really the field group's OWN presence — the group is all-or-nothing, and enabling it IS setting the fields — a console-side add is not merely possible, it is the whole drift you want to catch, and guarding the emit hides it FOREVER: the comparator's top-level walk is baseline-keys-only, so a key absent from the snapshot is never compared. AWS::CloudTrail::Trail's CloudWatchLogsLogGroupArn / CloudWatchLogsRoleArn pair is that shape, and it was guarded on a rationale whose both halves later proved false: AWS was said to reject the '' round-trip (the issue #1160 live probe accepts it and nulls the field out), and a console-side enable was said to surface as both fields appearing at once on the next read (it cannot — the walk never reaches an absent key). For this shape emit the group TOGETHER and UNCONDITIONALLY, with '' placeholders, and keep the all-or-nothing invariant on the WRITE side instead — CloudTrail's update path decides both fields TOGETHER and forwards them on PRESENCE, so an explicit '' CLEARS (which is what lets drift --revert undo a console-side enable) while an ABSENT pair is retained, and a half-populated or non-string shape is refused rather than coerced. Ask which one you have before guarding: is there a sibling field whose value makes mine illegal (guard), or is my own presence the switch (always-emit)?
Class 2 — structurally-incomplete-when-empty fields. An empty-object / empty-array placeholder is structurally invalid as AWS input because a sub-field is required. Example: SQS RedrivePolicy: {} rejects with "Redrive policy does not contain mandatory attribute: maxReceiveCount" because deadLetterTargetArn and maxReceiveCount are required. Other Class 2 candidates: Lambda DeadLetterConfig (TargetArn required), Lambda VpcConfig (SubnetIds + SecurityGroupIds required), EventBridge / SNS DeadLetterConfig, ECS NetworkConfiguration (awsvpcConfiguration.subnets required), various LoggingConfiguration shapes. Fix: keep the placeholder on the read side (drift detection requires it), and sanitize at the wire layer in create() / update() by translating the empty placeholder to whatever AWS accepts as "clear this field" — usually empty string. Canonical pattern in serializeRedrivePolicy (src/provisioning/providers/sqs-queue-provider.ts):
function serializeRedrivePolicy(value: unknown): string {
if (value === null || value === undefined) return '';
if (typeof value === 'object' && Object.keys(value as Record<string, unknown>).length === 0) {
return ''; // AWS-documented way to clear RedrivePolicy
}
return JSON.stringify(value);
}
Pattern documented in feedback_class2_placeholder_round_trip.md.
update() must gate optional fields on !== undefined, not truthy
Truthy gates (if (properties['X']) { ... }) silently drop empty string '', numeric 0, boolean false, and empty array [] (where the CFn type allows it). For cdkd deploy this is mostly invisible. For cdkd drift --revert it is a load-bearing bug: when state has no description but AWS does, the desired value to push back is Description: ''. A truthy gate drops it, the AWS update succeeds with no actual change, --revert reports ✓ reverted, and the very next cdkd drift re-detects the same drift — silent fail mode. Fix: use if (properties['X'] !== undefined) so explicit-empty values reach AWS:
// IAM Role example (PR #161 fix):
if (properties['Description'] !== undefined) {
updateParams.Description = properties['Description'] as string;
}
Truthy gates are correct ONLY for fields where the value range excludes the falsy form (e.g. boolean flags where false means "use default", or Path where empty is invalid). Add a code comment when the truthy form is intentional. Pattern documented in feedback_update_optional_field_undefined_check.md.
Read-update round-trip test (mandatory for any provider with readCurrentState)
The above failure modes (Class 1, Class 2, truthy gate) all surface only on the cdkd drift --revert code path, which round-trips observedProperties (= a previous readCurrentState snapshot) back through provider.update. Document review and code grep cannot catch every instance — write a test that exercises the round-trip mechanically:
it('round-trip: readCurrentState placeholders survive update() without AWS-invalid inputs', async () => {
// 1. Mock SDK to return the AWS-minimum response (only required
// fields, optionals undefined). readCurrentState should emit
// every always-emit placeholder.
mockSend.mockResolvedValueOnce({ /* minimum SDK shape */ });
// ...
const observed = await provider.readCurrentState(physicalId, 'L', RESOURCE_TYPE);
// Spot-check the placeholders are present (this is the always-emit
// contract, see "Test convention" earlier in this section).
expect(observed?.RedrivePolicy).toEqual({}); // Class 2 placeholder
// ...
// 2. Reset mocks and set up update() expectations.
vi.clearAllMocks();
mockSend.mockResolvedValueOnce({}); // SDK update call
// ...
// 3. Round-trip: pass observed as both new (desired) and old (previous).
// No drift → update should be a logical no-op on AWS.
await provider.update('L', physicalId, RESOURCE_TYPE, observed!, observed!);
// 4. Assert: no SDK call sent a value AWS would reject.
// Per-provider — list the AWS-rejection-shaped values you know of:
const setAttrsCall = mockSend.mock.calls.find(
(c) => c[0] instanceof SetQueueAttributesCommand
);
if (setAttrsCall) {
const attrs = setAttrsCall[0].input.Attributes;
if (attrs.RedrivePolicy !== undefined) {
// Class 2: '{}' would fail "Redrive policy does not contain
// mandatory attribute: maxReceiveCount"
expect(attrs.RedrivePolicy).not.toBe('{}');
}
// ... other per-provider rejection-shape checks
}
});
The round-trip test catches all three classes mechanically:
- Class 1 — discriminator-false placeholders that AWS rejects when shipped: assert the relevant
SetXxxAttributes/UpdateXxxcall does NOT include the discriminator-only attribute when the discriminator is false in the mock setup. - Class 2 — structurally-incomplete placeholders: assert the AWS API call does NOT contain the empty-object / empty-array shape AWS validates and rejects (e.g.
RedrivePolicy: '{}',VpcConfig: {}). - Truthy gate — assert that empty-string / 0 / false placeholder values DO reach the relevant AWS API call (e.g.
UpdateRoleCommandinput must containDescription: ''whenobservedProperties.Description === '').
See tests/unit/provisioning/sqs-queue-provider-update.test.ts (Class 2 round-trip), tests/unit/provisioning/iam-role-provider.test.ts (truthy-gate round-trip), and tests/unit/provisioning/sns-topic-provider-roundtrip.test.ts (Class 1 round-trip) for canonical examples.
Gone versus no read path. Return RESOURCE_NOT_FOUND (from src/types/resource.ts) when AWS itself says the resource does not exist: the service's not-found error code or fault NAME, an empty describe list for the requested id, or a terminal deleted status the service keeps listing (ECS INACTIVE, EMR TERMINATED). cdkd drift reports it as deleted and exits 1. Return undefined only for a type the provider has no read path for, or an answer that proves nothing (an unparseable id, a successful call with no body). Never return the sentinel on a message-text match alone, an access-denied or a throttle: rethrow those, or keep the old undefined, so a permissions gap is never reported as a deleted resource.
handledProperties against the CFn schema
Every SDK Provider declares a handledProperties: Map<string, ReadonlySet<string>> field naming the CFn template properties it knows how to wire to its AWS API calls. The provider registry's getProviderFor consults that set at routing time — a template carrying a property NOT in the set is auto-routed via Cloud Control API (which forwards the full property map to AWS, closing the silent-drop bug — see #614). --prefer-sdk-route Type:Prop is the per-property opt-out that forces the SDK Provider path and accepts the silent drop.
That's a runtime safety net. It doesn't help during development. A provider author who simply forgets to list a property in handledProperties AND forgets to wire it in create() / update() ships a silent bug — exactly what PR #370 (ApiGateway::Method dropped 15+ fields) demonstrated.
The structural prevention layer lives at tests/unit/provisioning/property-coverage.test.ts. It cross-references every registered provider's handledProperties against the canonical CFn schema (snapshotted to tests/fixtures/cfn-schemas/) and fails when a schema property is unaccounted for.
The four "OK" buckets
For each schema property the test classifies it into one of four buckets (in priority order):
| Bucket | Where declared | When to use |
|---|---|---|
handled |
provider.handledProperties.get(type) |
The provider's create() / update() actually wires the property to the SDK call. |
by-design |
provider.unhandledByDesign.get(type) (with rationale string) |
The provider INTENTIONALLY does not wire it — separate code path, deprecated, immutable post-create, AWS API doesn't accept it, etc. |
backfill |
tests/fixtures/cfn-schemas/_todo-backfill.json under types[<type>] |
Auto-generated catch-all for incremental rollout. Each entry MUST be migrated to handled or by-design eventually. |
read-only |
readOnlyProperties in the schema fixture |
AWS computes the value; cdkd cannot wire it on Create/Update by definition (e.g. Arn). Automatically excluded. |
A property in NONE of the above → test fails with the offending type + property list + the three actions you can take.
Adding unhandledByDesign to a provider
A clean example is AWS::ApiGatewayV2::Api's OpenAPI-import fields (Body / BodyS3Location / FailOnWarnings / DisableSchemaValidation / BasePath): they trigger an entirely separate ImportApi AWS API, not CreateApi. Listing them in handledProperties would be a lie; listing them in unhandledByDesign documents the deliberate skip:
unhandledByDesign = new Map<string, ReadonlyMap<string, string>>([
[
'AWS::ApiGatewayV2::Api',
new Map([
['Body', 'OpenAPI/Swagger inline spec; routed through ImportApi, not the field-by-field CreateApi path.'],
['BodyS3Location', 'OpenAPI/Swagger spec on S3; routed through ImportApi, not the field-by-field CreateApi path.'],
// ...
]),
],
]);
Wired into the provider class:
export class ApiGatewayV2Provider implements ResourceProvider {
handledProperties = new Map<string, ReadonlySet<string>>([...]);
unhandledByDesign = new Map<string, ReadonlyMap<string, string>>([...]);
// ... rest of provider
}
Rationales are free text but should be greppable. Common shapes:
"create-only — AWS rejects on update""AWS-managed read-only attribute"(when not already inreadOnlyProperties)"deprecated — superseded by Y""tags handled via per-resource Tag API, not the create input""covered by separate AWS::Foo::Bar resource type""OpenAPI-import-only flag; meaningful only on the ImportApi code path"
NON_PROVISIONABLE types: list them in SDK_PROVIDER_NON_PROVISIONABLE_TYPES.
A template property in neither handledProperties nor the allow set normally
auto-routes the resource through Cloud Control (issue #614), and so does a key
the schema snapshot does not know (#3713). If your provider covers a
ProvisioningType: NON_PROVISIONABLE type (e.g. AWS::FSx::FileSystem,
AWS::CodeBuild::Project), that route target does not exist — Cloud Control
has no handlers — and the generated Tier 3 set cannot catch it: it excludes
SDK-covered types by design, so isNonProvisionable() returns false once the
audit is regenerated after your provider is registered. Add the type to
SDK_PROVIDER_NON_PROVISIONABLE_TYPES in
src/provisioning/unsupported-types.ts (measured with
aws cloudformation list-types --visibility PUBLIC --type RESOURCE --provisioning-type NON_PROVISIONABLE). The ProviderRegistry then rejects a
silent-drop property pre-flight with a clear error (property rationale +
--prefer-sdk-route escape hatch) instead of failing at provisioning time with
an opaque UnsupportedActionException, and keeps an unknown key on the SDK
provider with a warning. It matters for every such type, fully handled or not,
and it is per TYPE, so a provider class that also serves provisionable types
(EC2Provider for AWS::EC2::NetworkAclEntry) needs nothing else.
property-coverage-cc-fallback-binding.test.ts fails while a registered type is
still in the Tier 3 set and missing from the list.
readonly disableCcApiFallback = true; on a provider class is the other
opt-out: it covers every type the class serves, for a provider that must not
fall back for a reason other than missing handlers (NestedStackProvider).
Workflow when adding a new provider
- Add the provider as usual — see Provider Implementation Examples.
- Re-export the class from
src/provisioning/provider-classes.tsand register the new resource type insrc/provisioning/register-providers.ts. - Refresh the CFn schema fixture:
This fetches only the newly-registered type via
node scripts/refresh-cfn-schemas.mjs --only-missingcloudformation:DescribeTypeand writestests/fixtures/cfn-schemas/<sanitized-type>.json. Requires AWS credentials withcloudformation:DescribeTypepermission. - Run
vp test run property-coverage— it will fail listing the schema properties your provider has not yet accounted for. - For each unaccounted property, EITHER:
- Add it to
handledProperties(ifcreate()/update()already wires it), OR - Add it to
unhandledByDesignwith a one-line rationale.
- Add it to
- Re-run the test — green.
If you really need to ship before classifying every property, you can regenerate the backfill TODO:
CDKD_GENERATE_BACKFILL=true vp test run property-coverage
This dumps every unaccounted property per type into tests/fixtures/cfn-schemas/_todo-backfill.json so the test passes. The intent is short-lived — a follow-up PR must migrate the backfill entries to handled or by-design.
Workflow when AWS publishes new properties
AWS adds properties to existing resource types fairly regularly — measured at roughly 3 writable properties per month across the whole Tier 1 surface. A property not yet in the fixture is unrecognized, and cdkd routes a resource carrying one through Cloud Control, exactly as it routes a known silent drop. So a new property reaches AWS the day AWS publishes it; the fixture refresh is what lets an SDK provider take the property over later. Why, and what is excluded: the design record.
Two mechanisms cover it. The operator's side of the first — what arrives, what
to do with each class, and what to do when NOTHING arrives because the job
failed or never fired — is the
CFn schema refresh runbook. That page is unlisted
and is linked from the pull requests the job opens, so this is the pointer that
still works when there is no pull request to read.
Daily, automatically. .github/workflows/cfn-schema-refresh.yml runs vp run gen:cfn-schemas-from-zip every day and, on drift, opens a PR carrying the mechanical regeneration. A day with no drift opens nothing — the job stops at the drift check — and the open-PR guard bounds concurrency at one, so the cadence costs a short run per quiet day rather than a review cycle. It reads AWS's public schema bundle, so no AWS credentials and no CI IAM role are involved — the prerequisite that kept this manual. Only fixtures that actually changed are rewritten (generatedAt-only churn is excluded from the comparison), so the PR diff is the real drift.
That PR is allowed to land red, and the red is the hand-off rather than a bug. The job runs only the mechanical chain and hand-classifies nothing, so two classes still need you:
- A removed or renamed property turns a matching
handledProperties/unhandledByDesigndeclaration into a bogus entry — retire the declaration or add abogusToleratedrationale (see below). audit:nested-key-coverage:checkreports divergences, not staleness — a new nested key on aNESTED_KEY_TARGETStype needs a provider fix or aNESTED_KEY_ALLOW_LISTentry with a rationale.
Newly unaccounted writable properties land in _todo-backfill.json; do not file an issue for any of them. .github/workflows/backfill-umbrella-sync.yml rewrites the standing backfill campaign's checklist from main's coverage map once the refresh merges — ONE umbrella issue whose body carries a generated block between <!-- backfill-types:start --> and <!-- backfill-types:end -->, one row per resource type, which gains, loses and regains rows as the map moves. The audit provenance lives outside those markers and is never rewritten. (The campaign is the open issue carrying the backfill-umbrella label — #2762 today — which superseded the closed #609; #2949 briefly fanned it into ~44 generated per-type sub-issues, since folded back into the block because a bot-filed slice is indistinguishable from an unfixed defect in a public open-issue count.) A pull request wiring a type writes Refs, and the type's row disappears once the coverage map says it is done. Types the public bundle does not carry are skipped with their fixtures left untouched and still need the authenticated path below.
On demand, by hand — to pull a refresh forward, or for a type the bundle lacks:
- Run
vp run gen:cfn-schemas-from-zip(public bundle, no credentials), ornode scripts/refresh-cfn-schemas.mjs '<AWS::Service::Type>'for one type viacloudformation:DescribeType. git diff tests/fixtures/cfn-schemas/shows what AWS changed.- The next
vp test run property-coveragerun fails naming the newly-unaccounted properties. - Triage each: wire it through, mark
unhandledByDesign, or backfill (with follow-up). Before wiring one,grep -rn '<Property>' tests/integration/*/verify.sh tests/integration/*/lib— a fixture that keys its Cloud Control route on that property loses its premise the moment it is handled, and has to be re-seeded in the same PR (integ-fixture-conventions.md).
There is deliberately no CI staleness check on main. It would go red whenever AWS publishes a property — noise on a schedule nobody controls, the same reasoning gen:aws-cli-removals carries in vite.config.ts. A scheduled job whose red is confined to its own PR is the shape that argument leaves open.
What stays on the SDK route. An unrecognized property does not route in four cases. The first three produce a deploy-time warning that names the property, says it will not reach AWS, and gives the reason; the fourth is silent:
| Case | Why it stays |
|---|---|
| The property is read-only | CloudFormation ignores a read-only property in a template too. |
| The type has no usable Cloud Control route | Its provider declares disableCcApiFallback, the type is NON_PROVISIONABLE, or its Cloud Control handler cannot manage it (a 'cc-broken' sticky-CC exemption). |
| The resource was deployed on the SDK route with the property, and its value is unchanged | cdkd keeps an existing resource on its route. Changing the value routes it. |
--prefer-sdk-route <Type>:<Prop> names it |
You chose the drop; the warning is suppressed. |
A misspelled property on a new resource, or one newly added, is routed like any other and Cloud Control rejects it with Unsupported property, as CloudFormation does.
"Bogus" entries and the tolerance list
A property in handledProperties (or unhandledByDesign) that is NOT in the CFn schema is a "bogus" entry — most often an SDK input field name that diverges from the CFn property name (e.g. SDK DefaultCooldown vs CFn Cooldown on AutoScalingGroup), a typo (PlacementStrategy vs PlacementStrategies on ECS::Service), or a stale alias from before AWS renamed the property.
The test reports these — but fixing each requires per-provider investigation that often touches the safety-net's runtime behavior. As a stopgap, the tolerance list at tests/fixtures/cfn-schemas/_todo-backfill.json under bogusTolerated[<type>][<prop>] accepts a one-line rationale per entry, the test stays green, and follow-up PRs investigate one at a time. Day-1 of issue #391 the test surfaced 10 such entries — see the rationale strings in that file for the canonical examples.
An entry is a claim about a DECLARATION — "the provider declares <prop>, the schema no longer has it, and here is why the declaration stays" — and every reader keys on that premise: the coverage test consults it only while walking handledProperties / unhandledByDesign / the backfill list, and the schema-refresh diagnosis only for a property in the generated handled map. So when the declaration itself is renamed or removed (the usual fix), retire the entry in the same change: the test fails an entry that names a property no declaration on that type carries, because such an entry is inert while reading as protection — and if the declaration was meant to exist, it is masking a silent drop (issue #3034, where a NetworkAclEntry.IcmpTypeCode entry written on 2026-05-16 outlived the 2026-05-27 rename to Icmp until 2026-09-13). The other staleness direction — AWS re-adds the property — is reported by its own case.
Admitting a type to the sticky-CC exemption
Closing the last silent-drop gap for a type is not the end of the story for
resources ALREADY deployed. A resource recorded provisionedBy: 'cc-api'
stays on Cloud Control by default, so the backfill you just landed speeds up
new resources only — every existing one keeps paying Cloud Control latency
until the type is admitted to STICKY_CC_MIGRATION_EXEMPT
(src/provisioning/provider-registry.ts,
issue #2719).
Consider admitting a type in the same PR that empties its silentDrop map, or
in a small follow-up. It is not mandatory and nothing blocks a release without
it: an un-admitted type is merely slower, which is the status quo — unless
Cloud Control also mishandles an update for it, in which case admission is
the fix (issue #4679).
The bar is physicalId parity, and it is EVIDENCE, not an argument. Cloud
Control mints its identifier from the schema's primaryIdentifier; the SDK
provider stores whatever its create() returned. When those differ, a flip
leaves cdkd addressing the resource by a string AWS does not recognise. They
often agree — but "often" is why this needs observing rather than reasoning.
To admit a type:
- Confirm the type's
silentDropmap is empty inproperty-coverage.generated.ts. A type with a real drop must not be admitted; the per-resource condition would refuse every resource anyway, making the entry dead weight. And keep it empty: the daily schema refresh re-adds a drop the moment AWS publishes a property the provider does not write (AWS::SNS::TopicgainedMaximumMessageSizethat way on 2026-09-18, issue #3413), so a refresh landing a drop on an admitted type owes the wiring in the same cycle — otherwisecdkd diffrendersstickyfor a resource thatcdkd deploy --allow-unsupported-propertieswould flip. - Read the CFn schema's
primaryIdentifierand the provider'screate()return, and write what BOTH store intophysicalIdForm— in words a reviewer can check, not "verified". - Write or extend an integ fixture that OBSERVES the flip on a live resource:
deploy, force the resource onto Cloud Control with
--recreate-via-cc-api, redeploy with a real property change, then assert the physical id is unchanged AND the record flipped to'sdk'.tests/integration/cc-to-sdk-reroute/is the model. Comparing physical ids is not always sufficient — a resource with a user-supplied name keeps its id through a destroy + recreate, so that fixture also attaches an out-of-band subscription cdkd does not manage and asserts it survives. Without such a witness the arm cannot tell an in-place update from a replacement, which is the whole claim. - Add the entry naming that fixture.
tests/unit/provisioning/sticky-exempt-registry.test.tsrefuses an entry whose fixture does not exist or has no row in the integ ledger. A fixture that ran before your arm was added still passes that check, so the arm's own run (step 5) is what the PR must wait for. - Run the fixture and commit the ledger row.
Use mode: 'cc-broken' only when Cloud Control genuinely cannot manage the
type (its escape is unconditional and ignores --pin-cc-api). A type that
works on Cloud Control and is only slower is 'sdk-coverage', and so is one
whose Cloud Control defect only a mutating deploy reaches, since that deploy is
the one that flips the record (AWS::ElasticLoadBalancingV2::Listener: Cloud
Control leaves a removed ListenerAttributes key at its old value, issue
#4679).
Logging
info: Successful operationsdebug: Detailed informationwarn: Non-fatal errorserror: Fatal errors
this.logger.info(`Creating ${resourceType} ${logicalId}`);
this.logger.debug(`Using properties:`, properties);
this.logger.warn(`Old resource deletion failed: ${String(error)}`);
this.logger.error(`Failed to create ${logicalId}:`, error);
Resource name constraints
AWS services have length and character constraints on names:
// IAM Role example (64 character limit)
private shortenRoleName(roleName: string): string {
const MAX_LENGTH = 64;
if (roleName.length <= MAX_LENGTH) {
return roleName;
}
const hash = Buffer.from(roleName)
.toString('base64')
.replace(/[^a-zA-Z0-9]/g, '')
.substring(0, 8);
const maxPrefixLength = MAX_LENGTH - hash.length - 1;
const prefix = roleName.substring(0, maxPrefixLength);
return `${prefix}-${hash}`;
}
Related
- Provider Development — the interface, the examples, and the steps to add a provider
- Integration fixture conventions — the rules a fixture exercising a provider follows
- Supported Resources — which types have an SDK provider today