Skip to content
cdkd

Provider implementation rules

The rules an SDK provider follows, and why each exists. Most were written after a defect got past review, so each carries the shape of the failure it prevents rather than only the rule.

This page is for someone writing or reviewing a provider. Start from Provider Development, which walks the interface and the steps end to end; come here for the decisions each step implies.

Error handling

  • Wrap all AWS SDK calls in try-catch
  • Use ProvisioningError to provide detailed context
try {
  await this.client.send(new CreateXxxCommand({ ... }));
} catch (error) {
  throw new ProvisioningError(
    `Failed to create ${logicalId}: ${String(error)}`,
    resourceType,
    logicalId,
    physicalId,
    error instanceof Error ? error : undefined
  );
}

Never infer a default from a possibly-malformed value

Reading a string out of a nested config block with || — or with ?? — looks harmless and is not:

// WRONG — a string / array / unresolved intrinsic container indexes to
// `undefined`, and the `||` silently substitutes the OPPOSITE of the
// declared intent.
const status = (versioningConfig['Status'] as string) || 'Suspended';

// EQUALLY WRONG — `??` defaults on exactly the same `undefined` (issue #1493).
const type = (source['Type'] as string) ?? 'NO_SOURCE';

VersioningConfiguration: 'Enabled' (a hand-written L1 template, or an intrinsic the resolver could not resolve) therefore turned versioning off on a live bucket, with no error anywhere. Use the shared guards instead:

import { readConfigString, requireConfigString } from '../config-shape.js';

// nested container, which may itself be malformed
const status = readConfigString(
  versioningConfig,
  'Status',
  'Suspended',
  'AWS::S3::Bucket VersioningConfiguration'
);

// top-level field — keep the `properties['X']` read at the call site so the
// handled-property-wiring critic can still see the property is consumed
const scope = requireConfigString(properties['Scope'], 'REGIONAL', 'AWS::WAFv2::WebACL Scope');

An absent container and an absent key still take the default ({} legitimately means "defaulted"); a container that is present but not an object, and a key that is present but not a non-blank string, are refused by name.

A container you never read a string out of needs its own guard. The two above can only fire while reading a FIELD, so a block whose members you merely probe for PRESENCE — or hand to .map — slips past both, and a malformed value there reads as an EMPTY block rather than as an error:

import { requireConfigArray, requireConfigObject } from '../config-shape.js';

// LIST block: a truthy non-array reaches `.map` and dies with a raw TypeError,
// a falsy one is silently dropped by the truthiness gate in front of it.
const tagFilters = requireConfigArray(raw, 'AWS::S3::Bucket …TagFilters');

// OBJECT block: every probe of a malformed container indexes to `undefined`,
// so the block reads as empty and the caller proceeds WITHOUT it — an S3
// lifecycle rule losing its whole location scope and applying bucket-wide.
const filter = requireConfigObject(raw, 'AWS::S3::Bucket …Rules[].Filter');

Both leave the ABSENT case to you (raw != null && …), because an omitted block legitimately means "no entries" / "defaulted". Both also take the onUnusable downgrade below, and when you pass it they return undefined instead of throwing — you then decide the skip UNIT, and that decision is the whole point: skipping must never be the misbehavior you were refusing. S3's per-item Puts skip the single configuration item, while its lifecycle Put skips the WHOLE configuration, because that call replaces every rule and applying the valid siblings alone would DELETE the malformed one from AWS.

A per-ITEM STRING read may need that same skip, not readConfigString's default. readConfigString's onUnusable downgrade is warn-and-DEFAULT, which is right for a field whose default is inert — and wrong wherever the default is applied to a LIVE resource and turns something ON. All four of S3's per-item reads are that shape (issue #1595): defaulting a replayed lifecycle / intelligent-tiering / replication Status to Enabled starts expiring, archiving or replicating objects for a rule the template had DISABLED. Ask the question without taking the answer:

import { configStringRefusal } from '../config-shape.js';

// `undefined` when the read would have succeeded; otherwise the refusal
// SENTENCE, with no action clause — you supply the one that is true here.
const refusal = configStringRefusal(rule, 'Status', 'Enabled', '…Rules[]');
if (onUnusable && refusal !== undefined) {
  // Keep the clause PATH-NEUTRAL. The replay-CREATE arm reaches this too
  // (a reverse-replacement revives the resource), so a message asserting a
  // "live" configuration would be false there.
  onUnusable(`${refusal}. Leaving the whole configuration unapplied here; …`);
  return; // ...or `continue`, per the skip UNIT this API implies
}

On the create path you do not probe at all, so the original read still refuses exactly as before. Use this helper rather than a hand-written typeof check: it shares requireConfigString's predicate, and a twin disagrees with the read it fronts on precisely the values that matter — a blank string, an explicit null, a coerced number.

Pass the desired side only. previousProperties comes from cdkd state rather than the user's template, so refusing a malformed value recorded there by an older binary would make the stack undeployable with no way out short of hand-editing the state file.

A list whose removals are the gap between the two sides is the exception (issue #3948): Auto Scaling group attachment and entry lists (tags, metrics, lifecycle hooks, traffic sources, notifications), Firehose tags, an ELBv2 target group's Targets, Budgets NotificationsWithSubscribers / ResourceTags, and CodeCommit Triggers / Tags (issue #3989), and most providers' CloudFormation Tags diffs, read through src/provisioning/tag-list.ts (issue #3994; Glue and log groups: #4073). There, reading a malformed DESIRED side as empty detaches or deletes everything the record holds, so it is refused before any call on every path, a state replay included. A malformed RECORDED side only misses removals: apply it ADD-only where you can — read the live resource (keep the live entries the desired side names, warn about the rest), or, where every add is idempotent (an upsert, or a create whose duplicate is success), send only the adds and warn — otherwise refuse it with a repair that never asks for a secret in state.json. undefined / null stays the empty list.

A top-level read takes two further decisions, both per site (issue #1513), expressed as options on requireConfigString:

// CFn coerces scalars and cdkd does not, so an unquoted YAML `IpProtocol: -1`
// arrives as a NUMBER and deploys fine today — refusing it would break a
// working template. Only for genuinely numeric-looking fields; a number where
// an enum belongs (`InstanceType`) stays a refusal.
//
// The real site wraps this call in `narrowIngressIpProtocol` so the DIFF can
// resolve the protocol the same way — see the `effectiveProperties` section
// above for why a coerced value has to be recorded as well as sent (#1633).
const ipProtocol = requireConfigString(
  properties['IpProtocol'],
  '-1',
  'AWS::EC2::SecurityGroupIngress IpProtocol',
  { coerceNumber: true }
);

// UPDATE-path sites WARN instead of throwing when the desired bag is a state
// record (a rollback revert or `cdkd drift --revert`) — see [Pre-flight refusal](#pre-flight-refusal-when-a-provider-may-reject-what-cloudformation-forwards):
// a refusal there can leave the resource un-rollbackable with no template-side
// remedy. A template-path update refuses the same value, before any call.
// (`context` is `update()`'s `UpdateContext`; `previousStatus` is the recorded
// status, the warn fallback, so a warning never ENABLES a key.)
const stateBorneDesired =
  context?.replayingState === true || context?.desiredFromAwsReadback === true;
const status = requireConfigString(
  properties['Status'],
  previousStatus,
  'AWS::IAM::AccessKey Status',
  stateBorneDesired ? { onUnusable: (message) => this.logger.warn(message) } : {}
);

And check WHERE the read lives before guarding it: a helper the delete() or diff paths also reach is fed state-borne values, so guard the create CALL SITE instead (EC2Provider.buildIpPermission is the tree's example — it is shared with deleteSecurityGroupIngress and with the revoke half of the inline-rule update diff).

Pre-flight refusal: when a provider may reject what CloudFormation forwards

cdkd's compatibility target is CloudFormation, so the default for a property cdkd cannot handle is to forward it and let AWS answer, never to invent a validation CloudFormation does not have. A provider that refuses a template CloudFormation would accept is a parity break, and parity breaks are how a tool that claims template compatibility stops being trustworthy.

There is one narrow exception, and it has a high bar. A provider MAY refuse a property before the AWS call when all of the following hold:

  1. The property is undeployable on cdkd's OWN path, proven by a live probe against the real API the provider calls — not inferred from CloudFormation also rejecting it. This is the load-bearing condition: "CFn rejects it too" is not sufficient, because cdkd calls the service API directly and could in principle succeed where a CFn resource handler fails.
  2. No shape of it works, so no user loses a working deployment. If some combination deploys, forward it.
  3. The refusal names the working alternative, concretely enough to copy. A refusal that only says "no" is worse than AWS's own error.
  4. The rationale is recorded at the check site, as a comment naming the probe (date, region, what was tried, what AWS said) and flagging it as a deliberate parity divergence — so the next reader can re-evaluate it when AWS changes.

GlueProvider's enforceIcebergTableInputAbsent (issue #1454) is the reference implementation. Three details are worth copying:

  • The check runs before the try block, so the typed ProvisioningError is not caught and re-labelled by the provider's own error wrapper. A test that only asserts message CONTENT cannot catch this being moved inside the try, because the wrapper EMBEDS the original message — assert the raw message PREFIX and an absent cause instead.

  • Refuse a TEMPLATE-driven call; WARN on a STATE REPLAY. This asymmetry is the rule, not a Glue quirk, and the axis is the ORIGIN of the properties, not the operation name. cdkd rollback replays from cdkd STATE, not from the template. A resource whose state record already carries the offending value (written by an older cdkd build, or by an import) would become not just un-updatable but UN-RESTORABLE, and unlike the template case the user has no remedy at all — only hand-editing state.json. Concretely:

    • update() — refuse on the template path, warn on a replay. The rollback executor's two revert arms call provider.update(..., op.previousState.properties, ...) with UpdateContext.replayingState, and cdkd drift --revert sets desiredFromAwsReadback; the deploy engine sets neither. Refuse only when both are unset AND the refused value is template-borne on that path — a value the recorded resource already carries, such as an identity segment that only a replacement could change, keeps the warning, with the reason stated at the site (issue #3728). Refuse BEFORE the first write: a throw from a mid-update arm strands whatever the earlier arms applied (a read such as a hosted-zone lookup may precede it). Where the arm itself runs mid-update, ask its OWN predicate from a pre-flight rather than restating it — S3's per-config appliers run on a probe whose client writes nothing (issue #3740). Gate the refusal on the value having changed wherever an unchanged one sends nothing, and do NOT gate it where the value goes out on every update: Glue's UpdateDatabase replaces DatabaseInput wholesale, so its malformed blocks are refused whether or not they changed (#3740). A list whose removals are the gap between the two sides refuses on a replay too (see "A list whose removals are the gap between the two sides is the exception" above, #3948): warning and reading it as empty would remove everything the record holds. Separately, a create-only value such as AWS::RDS::DBProxyTargetGroup TargetGroupName or AWS::Lambda::EventInvokeConfig Qualifier keeps the warning on purpose, for the reason above. On the warn path, pick the FALLBACK per site (issue #1551): warning and then applying the CREATE DEFAULT is frequently worse than the refusal was, because the default lands on a LIVE resource — it flipped an IAM-guarded Lambda function URL to PUBLIC, re-pointed a live DynamoDB stream, and read a malformed GSI block as "delete every index". Keep the PREVIOUS value where one exists (omitting the field entirely when the API has merge semantics), otherwise SKIP the block or SUPPRESS that part of the diff. Whichever you choose, remember the warn path now records the unusable value as the new state, so seed the next comparison from AWS's live value where the provider already holds it (issue #1552).

    • create() — refuse, unless CreateContext.replayingState is set. create() is a replay path only through the rollback executor's reverse-replacement arm, which revives the OLD resource from previousState.properties (issue #1199). That arm — and nothing else in cdkd — passes the optional 4th parameter context?: CreateContext with replayingState: true (issue #1463). The deploy engine's six create sites (CREATE, the property-driven replacement, the --recreate-via-* destroy-then-create, the --replace delete-first fallback, the update-failure replacement, and createFirstThenDeleteOld, which --recreate-via-* and the update-failure replacement take for a renamed resource) are driven by freshly resolved TEMPLATE properties and never set replayingState, so the refusal stands where the user can actually act on it. (They DO pass a context — every provider call now carries maskSecrets, see the maskSecrets bullet below — so the test fences for this read the context's key set rather than the call's arity.)

    • Do not re-create inside update() if you have a create-side refusal. Several providers call this.create(...) from their own update() (ACM certificate, IAM managed policy, IAM role, Lambda permission, SNS subscription). Those internal re-creates never receive replayingState — update()'s own context is an UpdateContext, and the most any of the five builds from it is a CreateContext carrying only maskSecrets (IAM role, IAM managed policy and Lambda permission, issue #2177) — and the properties they forward ARE a state record during a rollback replay. So the refusal would fire on a replay with no way to detect it. None of those providers has a pre-flight refusal today (they validate required fields only, which correctly stays a hard error), so there is no live gap; this is a constraint on the NEXT provider, not a description of the current tree. Since issue #3141 UpdateContext carries its own replayingState — set by the rollback executor's two revert arms — so the next provider CAN build a CreateContext from it instead of relying on the constraint.

    • Retire what your create materialized before the error leaves. A create shaped <one API call that materializes the resource> then <a wait for it to become usable> can fail with the resource already alive in AWS. Before issue #2169 an AWS::CertificateManager::Certificate in that state was simply LOST: the create never returned, so nothing was written to state, and it was invisible to cdkd state show, unreachable by cdkd destroy, and re-requested by the next cdkd deploy — an orphan per attempt. Delete it in your own catch, best-effort, and re-throw the ORIGINAL error.

      Three points that decide whether you get this right:

      • Track the id from the API response, not from the name you computed. Set the local only once AWS has confirmed the resource exists, so a failure BEFORE the call deletes nothing. Do NOT reuse ProvisioningError.physicalId as this signal — 462 call sites across 75 provider files pass one, and on a create path it is usually the intended name computed beforehand (IAMRoleProvider.create hands its catch the roleName it derived, whether or not CreateRole ever ran).
      • Never let the cleanup replace the diagnosis. The create's own error is what the user needs; a cleanup that FAILED appends a line naming the survivor and the manual command to retire it, and a cleanup that threw must never surface instead of the original.
      • Do not record the remnant in state as an alternative. It looks like the tidier answer and it is a trap: the recorded properties ARE the template's, so the next deploy diffs the resource as NO_CHANGE and never touches it again — cdkd deploy prints "No changes detected" and exits 0 over a resource that is still unusable. Turning a loud failure into a silent one is worse than the orphan. Making the diff re-provision it instead needs a marker on the state record, i.e. a schema bump.

      Whether deleting is SAFE is a per-service question you have to answer, not assume. For ACM it is, and the reason is documented rather than inferred: the DNS validation CNAME is derived from the domain and the account, not from the certificate, and AWS states you can "replace a deleted certificate" without repeating validation — so the records a user adds after the failure validate the retry's certificate too. If your service has no such property, say so and choose differently.

      A failing AUXILIARY call in create() — a tag, policy, attribute or other write after the main create — must be marked with markAuxiliaryFailure(error, logicalId) (src/provisioning/auxiliary-failure.ts, issue #3826), or that object's "already exists" can be classified as the resource's own name collision, and --replace or a replacement rollback then deletes the live resource.

      A create() that fails AFTER its main create call returned must also mark the error it throws with markCreatedBeforeFailure(error, ownerLogicalId, resourceType, physicalId) (same module, issue #1710). The resource then exists with no state record; the mark is what lets the failed-CREATE journal name it, so the automatic rollback, cdkd rollback and cdkd destroy can delete it. When the create call's response carries an immutable AWS-generated id for the resource (RDS's DbClusterResourceId / DbiResourceId), pass it as the optional fifth argument, createdResourceIdentity. It is journaled without a later read, and it must be the value the provider's resourceIdentity returns for that resource, since a later successful deploy deletes a name-keyed orphan only when the two match. Set it only behind a flag the create call's success sets (Kinesis's streamCreated): ProvisioningError.physicalId is NOT that proof, since providers put the intended name on refusals and on the create call's own "already exists", which names another owner's resource. Mark the outermost error create() throws, with the id delete() takes (the success path's physicalId). A catch that deletes its own resource marks only when that cleanup FAILED; a resource the create adopted or found held (an idempotent create returning an existing one, heldBefore !== 'free') is never marked; a provider whose delete would reach beyond what the create wrote (AWS::IAM::Policy walks every listed principal) removes its own writes in the catch instead (issue #4583).

    • maskSecrets — mask a resolved property value before you log it. Both CreateContext and UpdateContext extend a shared SecretMaskingContext, whose one optional field is maskSecrets?: (text: string) => string (issue #1932). Any provider log line that interpolates a value from the properties bag MUST run through it: those values arrive RESOLVED, so a {{resolve:secretsmanager:...}} property is already plaintext by the time a provider sees it, and cdkd's other masking boundaries (the deploy engine's error / reason text, the resolver's own debug line) do not cover a provider's own logger.warn. Read it defensively — const mask = context?.maskSecrets ?? ((t: string) => t) — since create() / update() are also called by cdkd drift --revert, by the import path, and by tests. It is per-CALL, so never cache it on this: providers are registered as singletons and serve concurrent resources. A provider whose create() / update() reach many private helpers may instead re-enter itself on a fresh per-call object whose prototype is the singleton and whose logger is masked, so every helper's this.logger line is masked by construction (S3BucketProvider.maskedView, issue #2177). Mask the VALUE before it is stringified or interpolated; the finished message is a FALLBACK, not an equivalent. Two independent reasons: (1) escaping — a masker matches by literal occurrence, and JSON.stringify escapes ", \ and newlines, so a secret containing any of them (i.e. any Secrets Manager JSON document, the commonest real secret shape) no longer occurs in the finished line and passes through verbatim; (2) length — maskSecretsInText masks an exact whole-value match at ANY length but only scans for SUBSTRING needles of at least MIN_NEEDLE_LENGTH (4) characters, and a message is always longer than the value inside it, so a 1-3 character secret survives. Do BOTH: mask each string leaf before interpolating (catches escaped and short secrets), and route the assembled message through the masker too (catches interpolations added later, and text the leaf pass never sees). The mask is idempotent, so the layers compose. createMaskedLogSinks (in masked-retry-logger.ts, below) builds both halves for one operation: a raw value masker plus masked debug / warn sinks. A physical name derived from a secret by generateResourceNameWithFallback (stack prefix, folded characters, truncation) no longer OCCURS literally, so withDerivedNameMasks adds the derived spelling as a needle when its raw value is a secret. The same helper covers a rotated secret on update(): pair the PREVIOUS value with the recorded name, and a previous value still spelling {{resolve:, or persisted as exactly the redaction mask ***, makes that name a needle although its old plaintext is in no bag of this deploy. Use maskDeep from src/provisioning/masked-retry-logger.ts for the leaf pass — do NOT hand-roll one. Issue #2176 found SIX private copies of that walk, FOUR of which had lost the depth cap. A walk that encodes a security contract drifts the moment it is copied: a hardening applied to one copy silently leaves the rest behind. Grep for the walk's SHAPE rather than a name — two of the six were nearly missed twice because they spell it maskLeaf / maskLeafValue and declare no named depth constant to grep for. The leaf pass is now MECHANICALLY ENFORCED for the JSON.stringify case (issue #2178). vp run audit:provider-secret-mask:check (scripts/check-provider-secret-mask.ts, a CI step) fails on any ${JSON.stringify(X)} interpolated into a message under src/provisioning/providers/** (plus composite-id.ts) where X reaches no masker. It is DATAFLOW-aware, so the mask may sit anywhere upstream — through a const binding, a ?: / ?? arm, an array or object literal, a .map() callback, or a helper — and both real upstream-mask sites (asg-provider.ts, cloudfront-distribution-provider.ts) pass unchanged. It became a check rather than another sentence because the rule above was already written and violated anyway: the #2176 sweep found raw sites in files already hardened for this exact contract.

      An IDENTITY masker does NOT satisfy it. A binding such as const M: MaskerFn = maskerOrIdentity(undefined) type-checks, reads as protection, and masks nothing — so the critic classifies a site masked only by one as raw. Issue #2007 is why that is worse than having no masker at all: its presence stops the next author looking. This holds in the ARGUMENT position too, which is the one you are most likely to write: maskDeep(value, maskerOrIdentity(undefined)) walks every leaf and changes nothing, so the critic classifies it raw exactly as it does a bare value. That position is an ALLOW-LIST — it accepts only what resolves to the shared capability and refuses anything else, so an inline (t) => t, a cast, a hand-rolled function noMask(v) { return v; } and an arbitrary object property are all refused without the check having to enumerate them.

      What the check does NOT prove. It is syntactic, so a no-op DECLARED to be the capability is believed: const m: MaskerFn = someNoOp and { maskSecrets: (t) => t } both pass. A green run means the value reached something declared to be the project's masker — not that it was masked. Thread the real capability; do not satisfy the checker. Where a path genuinely has no masker to thread — DeleteContext carries none by the SecretMaskingContext contract — thread the capability first, or record the site in the critic's EXEMPT list, where it is COUNTED and re-audited on every run and fails the moment the capability arrives. Do not reach for an identity default to make the check pass. A parameter's identity DEFAULT is a different thing and stays legitimate: it is the contract's back-compatible absent-means-unmasked answer, and what the critic checks there is that every CALL SITE threads a real masker — but only for an UNTYPED parameter of a non-exported function. A parameter ANNOTATED MaskerFn / SecretMasker is treated as a masker by TYPE, so it joins the file's masker set unconditionally and its call sites are NOT checked; the same is true of any exported function. Prefer the REQUIRED spelling (mask: MaskerFn, no identity default) for new sites: the compiler then enforces threading, which is the only mechanism that does not depend on the critic noticing. Every in-repo call site does thread today — verified, not assumed — so this is a shape to prefer, not a live gap.

      Back to the masker itself, whose scope the paragraphs above interrupted. maskDeep masks what the caller's resolution recorded: the dynamic-reference secrets, and since issue #1998 the value of a NoEcho: true template PARAMETER that a Ref or Fn::Sub variable served, recorded as a LOG-ONLY needle that the masker reads and, since state schema v11, also as a mask-only needle, so state persists *** where the value stood. No provider change was needed for either, which is the point of handing providers a function. A DIFFERENT NoEcho — the custom-resource RESPONSE field of the same name — IS covered since issue #2274, through noEchoAttributes above; the two share only a spelling, so do not read one as closing the other. delete() has no masker — see SecretMaskingContext for why that is a statement about providers rather than about the bag, and why you must thread the capability BEFORE adding any delete-side log line that names a property value. Mask where the message is CONSTRUCTED, not where the arm catches it, when the service reports rejections through an operation poller. For an operation-based API (AWS Cloud Map's Create*Namespace / Update*Namespace return an OperationId, and the rejection arrives later as Operation.ErrorMessage with the offending value quoted back), a shared poller builds the error and every arm's catch opens with if (error instanceof ProvisioningError) throw error; — so the poller's error is re-thrown VERBATIM and masking the arm does nothing for it. Threading the masker into the arm therefore looks like a fix and is inert on the path that actually fires, since for these types a FAILED operation is the NORMAL rejection route rather than an edge case. pollOperation in src/provisioning/providers/servicediscovery-provider.ts is the worked example (issue #2063): it masks the raw ErrorMessage at construction, and it takes the masker from every CREATE / UPDATE caller rather than only the newly-threaded ones — an arm fixed by an earlier pass is not evidence that its poller was. The DELETE callers stay unthreaded, which is the contract above rather than an oversight: DeleteContext carries no masker, and their payload is a physical id rather than a resolved property bag. The reference implementation is buildMfaConfigRequest in src/provisioning/providers/cognito-provider.ts, which routes every warning through one masked sink rather than masking at each call; create() in src/provisioning/providers/ssm-parameter-provider.ts is the same shape for a whole operation (one mask, one warn, one debug, built at the top and used everywhere below). Per-site masking DOES drift, and that is measured rather than predicted. Issue #2176's sweep found raw ${JSON.stringify(...)} sites in files ALREADY hardened for this contract — dynamodb-table-provider.ts masked one argument of a warn call and left the argument beside it raw. A sink is what makes the next line added below it correct by default.

    The parameter is optional, so a provider with no pre-flight needs no change. A provider that HAS one threads context from create() to its check and emits a warning instead of throwing when the flag is set. What the flag means (and, just as importantly, what it does NOT license — nothing about the properties' content, no relaxing of data-safety guards, no skipping the validation that protects the AWS call itself) is spelled out on CreateContext in src/types/resource.ts, next to the ResourceProvider interface that consumes it. Its sibling DeleteContext lives in src/provisioning/region-check.ts instead, because expectedRegion feeds that module's assertRegionMatch helper; CreateContext has no region-checking role, so it is not filed there for symmetry alone. A pointer next to DeleteContext links the two.

    One honest consequence: warning on a replay means the value IS forwarded on the create path, unlike the update path where the SDK command has no member for it. So the re-created resource is degraded in whatever way the original was, and the AWS call may still fail. That is strictly better than refusing — a refusal guarantees the resource is not restored — but the warning must SAY so and name the fix-forward (cdkd deploy with the working shape).

  • Where update genuinely does not forward the property anyway (cdkd does not wire Glue's update-only UpdateOpenTableFormatInput shape), warning costs nothing: no bad value can reach AWS from that path, and the user still gets the full message. Share ONE message builder between the refusal and the warning so they cannot drift.

A pre-flight in a provider only covers the SDK route. A resource whose state records provisionedBy: 'cc-api' is sticky-routed to CloudControlProvider and bypasses it entirely. That is usually acceptable (the deploy still fails, just later and less helpfully) — but say so in the docs rather than letting a reader assume the refusal is total.

If a property is merely unimplemented rather than undeployable, this is the wrong mechanism — move it to unhandledByDesign, which converts the silent drop into the Cloud Control auto-route (see handledProperties against the CFn schema).

Idempotency

  • Handle when create is called on existing resource
  • Handle when delete is called on non-existent resource

Region verification on *NotFound: A *NotFound error during DELETE must NOT be treated as idempotent success without confirming that the AWS client's region matches the region the resource was deployed to. A destroy run pointing at the wrong region would otherwise receive NotFound for every resource and silently strip them all from state, leaving the actual AWS resources orphaned in the real region (this is the silent-failure incident that motivated PR 2 of the region/state refactor).

Providers MUST call assertRegionMatch() from src/provisioning/region-check.ts before returning early on a *NotFound error:

import { assertRegionMatch, type DeleteContext } from '../region-check.js';

async delete(
  logicalId: string,
  physicalId: string,
  resourceType: string,
  _properties?: Record<string, unknown>,
  context?: DeleteContext,
): Promise<void> {
  try {
    await this.client.send(new DeleteXxxCommand({ Id: physicalId }));
  } catch (error) {
    if (error instanceof ResourceNotFoundException) {
      const clientRegion = await this.client.config.region();
      assertRegionMatch(
        clientRegion,
        context?.expectedRegion,
        resourceType,
        logicalId,
        physicalId,
      );
      this.logger.info('Resource not found, skipping deletion');
      return;
    }
    throw error;
  }
}

assertRegionMatch is a no-op when context.expectedRegion is undefined, preserving the existing idempotent semantics for callers that have not been threaded with state region. When set, a region mismatch throws a ProvisioningError that surfaces both regions and a hint to rerun with the correct --region.

The NotFound branch is the REACTIVE use, and it is not the only one (issue #2301). A wrong-region call usually never REACHES that branch: most physical ids are NAMES, and the same name commonly exists in the client's region as well — the same stack deployed twice, or resource-name.ts deriving an identical name from an identical stack plus logical id — so the call succeeds against the WRONG resource instead of erroring. assertRegionMatch therefore takes an optional sixth phase argument: the default 'not-found' keeps the wording above, while 'pre-delete' / 'pre-update' word the same refusal for a call that has not been issued yet. The comparison and the three outcomes are one implementation across all three phases; only the message differs.

CloudControlProvider uses those phases UNCONDITIONALLY, for every Cloud-Control-routed type, at the top of both delete() and update() — before the --remove-protection flips, before the SDK delegations, before the AWS::S3::Bucket identity probe, and before the DescribeType the update path's write-only lookup performs. The update side is why UpdateContext carries an expectedRegion of its own (src/types/resource.ts), threaded by deploy-engine.ts, rollback-executor.ts and drift --revert. An SDK provider MAY read the same field, but nothing requires it to: absent, the guard is a no-op, exactly as on the delete side.

The create side has the same question, and it is NOT the same answer. "Handle when create is called on an existing resource" above is the usual idempotent short-circuit: an *AlreadyExists / *AlreadyOwned error is swallowed and the create reported as success. Before writing one, ask what namespace that error is scoped to and whether it is narrower, equal to, or WIDER than the region the client points at. It is sound only when the error can mean nothing but "the resource I want, where I want it, already exists".

For nearly every type that is automatic — either the namespace is REGIONAL (SNS, CloudWatch Logs, SSM, EC2), so the conflict is by construction in the region you called, or the namespace is global AND so is the resource (IAM, Route 53, CloudFront), so there is no other region for it to be in. AWS::S3::Bucket is the one type where the two come apart, a globally unique name over a regionally located resource, and it broke exactly there: measured on real AWS, a bucket in us-west-2 answers BucketAlreadyOwnedByYou to a CreateBucket in eu-west-1 and in us-east-1 alike, so the short-circuit adopted another region's bucket and applied the whole stack's configuration to it while reporting success (issue #2227).

A conflict that carries no name at all is never sound to swallow: an AWS::EC2::SecurityGroupIngress duplicate means only that an identical rule is on the group, whoever added it. Such a create adopts only on this stack's own evidence — its rollback journal (getPriorAttempts), a rollback replay, or another record of the same stack holding the rule (getStackRecords; two template rules CDK could not dedupe, issue #4492, waiting for such a twin still being created in the same deploy, as CloudFormation accepts both) — and refuses otherwise. Two records then share one rule, so a delete leaves it in place while a record that outlives the operation still holds it, and the last holder revokes it.

S3BucketProvider.assertExistingBucketRegion is the create-side twin of assertRegionMatch. It reads the bucket's region from the x-amz-bucket-region header on the 409 itself — no extra call, no extra IAM — and falls back to GetBucketLocation. Note assertRegionMatch is a no-op without expectedRegion while this one always has a region to compare against: the create knows where it is deploying.

There are FOUR of these guards, not two, and assertRegionMatch is only the generic one. The two above are the SDK route; the Cloud Control route has its own pair, because a bucket declaring a silent-drop property (AccessControl) is auto-routed there and S3BucketProvider.delete never runs at all (issue #2283, fixed in #2309):

Guard Route / phase On an INDETERMINATE answer
S3BucketProvider.assertExistingBucketRegion SDK, create refuses
assertRegionMatch (src/provisioning/region-check.ts) generic, delete; update only where a provider opts in (S3BucketProvider alone today) no-op without expectedRegion
CloudControlProvider.confirmDeleteTargetIdentity CC, delete, per-type (S3 today) warns and PROCEEDS
CloudControlProvider.assertRecordedRegionAgainstClient CC, delete + update refuses

The last two sit in the same file and answer indeterminacy in OPPOSITE directions, which is deliberate rather than an inconsistency. The probe asks a REMOTE service where a globally unique name lives, so a least-privilege role never granted s3:GetBucketLocation would be stranded by a refusal; the comparison is LOCAL and free against a region the caller positively recorded, so a client that cannot say where it points cannot be shown to point there. Two further details of the probe are decisions, not oversights: it is keyed on a per-type set rather than being generic, and it deliberately omits ExpectedBucketOwner — the opposite of the convention state.ts and aws-region-resolver.ts follow — because the hazard is a name resolving to a bucket in ANOTHER ACCOUNT, and the guard has to hear that foreign answer to refuse. Passing the parameter would turn every cross-account collision into a 403, i.e. into the indeterminate arm, i.e. into warn-and-proceed.

Deliberately NOT HeadBucket, which 301s cross-region and which SDK v3 turns into a synthetic UnknownError (src/utils/aws-region-resolver.ts records the mechanism), and deliberately not that module's resolveBucketRegion, which never throws and returns a fallback region — wiring it here would turn a fail-closed guard into a fail-open one. See .claude/rules/provider-resource-identity.md for the full rule, including the two legacy GetBucketLocation spellings the fold must absorb and why the refusal is markNonRetryable. (.claude/rules/providers.md is a routing index now — the rules corpus was split under the corpus-wide byte ceiling that issue #2310 has since retired in favour of a per-paths:-glob budget — so it points on rather than carrying the rule itself.)

That guard is scoped to the BucketAlreadyOwnedByYou catch on the CREATE path, which leaves two neighbouring routes to the wrong bucket that it cannot see — us-east-1's legacy 200 (issue #2241) and a state record written before the guard existed (issue #2245). S3BucketProvider now also pre-flights the bucket NAME before CreateBucket when the target region is us-east-1, and shares one assertStateBucketRegion between update() and delete(). .claude/rules/provider-resource-identity.md carries the mechanism and the three rules worth reusing: the pre-flight informs the partial-create CLEANUP GATE (and, after an ambiguous attempt of the same create, a refusal that adopts nothing, issue #4639) and never licenses an adoption in place of CreateBucket (the ownership oracle), an unanswered probe is kept distinct from a confirmed absence, and the state-record guard PROCEEDS on both rather than stranding every update and destroy for a least-privilege role.

Region is not the only question. A bucket under an EXPLICIT BucketName that this create did not make is refused even in the right region, on both the BucketAlreadyOwnedByYou arm and the us-east-1 pre-flight (before the send, so the legacy 200 never resets its ACLs; when the location lookup cannot answer, the account's own bucket list decides, since the legacy 200 adopts only a bucket the caller owns), as CloudFormation fails the same create (issue #4684): a name is not attribution. Only a generated name's holder is still adopted, presumed this stack's own (#4345). The refusal is marked a name collision, so a create-first replacement and a rollback's reverse replacement take their record-proven collision paths rather than recording the bucket. It is deliberately retryable by text ("already exists"): a delete-first re-create retries that text while its own deleted name is released, and no other caller's classifier retries a collision.

A bucket has no immutable AWS id, so its resourceIdentity (the token a fix-forward settle compares before deleting a failed CREATE's orphan, issue #4606) is its name, its region and the CreationDate this account's own ListBuckets reports. A bucket the list does not show, or shows in another region, has no token. A bucket re-created under the name reports the new create's second (measured in us-east-1 and us-west-2, and asserted by the s3-fix-forward-orphan fixture before it trusts the token). Outside us-east-1 the date also moves to the second of a versioning, tagging, encryption or policy write, so the token is read only after the failed CREATE's last write; a moved date reads as another bucket and keeps the orphan, the safe direction.

Note what that last choice COSTS, because it is a new IAM dependency rather than a free win: the guard works only where the caller can call GetBucketLocation on the target. A principal without that grant is NOT an unrelated population — it is the same population, unprobeable, and for it the guard is inert on every call; a bucket POLICY can Deny the call as well, which is indistinguishable on the wire from a missing grant. That is why the delete-side degrade logs at warn rather than debug: proceeding unverified into an irreversible delete must never look identical to a healthy one.

That rule is also where the probe-first lesson lives, and it cuts both ways here: the mid-delete hazard the issue was FILED for is not reachable, so a guard against it could never have fired — while the guard that WAS written could not fire either, because its unit tests encoded the AWS CLI's wire shape rather than the SDK's. A mutation probe cannot catch that: it perturbs the code and reads the test, and both read the same fixture, so a premise shared by code and mock is invariant under it. Only a recorded real response or a live arm falsifies a fixture.

The S3 pre-flight's partial-create cleanup gate has a sibling for every create API that answers success with a resource already holding the name: ELBv2 CreateLoadBalancer / CreateTargetGroup on identical settings, SNS CreateTopic, and EventBridge PutRule, which overwrites (issue #4403). Before the create, and only when a wiring step that can fail is declared, the provider looks up the name it is about to send (src/provisioning/providers/create-ownership.ts). Only the service's own not-found answer licenses the partial-create cleanup. A name that was held, or a lookup that could not answer, leaves the resource in place with a warning saying why and the manual delete command. A new provider of such a type gates its cleanup the same way. A create that hands back by idempotency TOKEN rather than by name (FSx CreateFileSystem) cannot be looked up beforehand, so FSxFileSystemProvider takes -- records, and so may later clean up -- only a file system whose CreationTime is no earlier than that token's first send in this process, on AWS's clock, and refuses any other (#4428).

Reporting a skipped delete

Written for issue #1752.

delete() returning normally means "the resource is gone" — a delete call succeeded, or the resource was already absent (the *NotFound idempotent arm above, CustomResourceProvider's backing-Lambda-is-gone pre-check). Both are honest, so both count toward N deleted.

A warn-and-continue arm that issues no AWS call at all is different, and returning void from one is a bug. cdkd destroy's only signal used to be "did not throw", so such an arm printed ✓ <id> (<type>) deleted, counted toward N deleted, dropped the state record, and exited 0 — over a resource that may still be alive and billing. Report the skip instead:

const decoded = decodeTableId(physicalId, properties);
if (!decoded) {
  this.logger.warn(
    compositeIdFormatMessage(GLUE_TABLE_ID_FORMAT, logicalId, physicalId, { skipping: true })
  );
  return compositeIdSkipResult(); // { outcome: 'skipped', reason }
}

What the destroy runner then does, and why each half matters:

  • prints ⚠ <id> (<type>) skipped (<reason>) — a yellow warning glyph, so it can be read as neither the ✓ … deleted success line nor the ✗ Failed to delete failure line;
  • counts it in a separate skipped column (2 deleted, 1 skipped, 0 errors) rather than as an error — nothing was attempted, so nothing failed;
  • KEEPS the state record. This is the half that is easy to miss: dropping it leaves the user with neither the AWS resource deleted nor an id to go and delete it with;
  • preserves state.json and exits 2 (PartialFailureError), the same contract a partial / interrupted destroy already carries. Exiting 0 while leaving a resource behind is the mis-report this mechanism exists to prevent;
  • emits a RESOURCE_SKIPPED deployment event (no error field).

Reach for this ONLY when cdkd genuinely cannot address the resource. If you know the resource is gone, return void — reporting a skip there would preserve state and fail the destroy for no reason. The --purge-events flag and the run-level RUN_FINISHED result both treat a skip as a failed run, so a new producer changes user-visible exit behavior — decide deliberately.

Reporting a pre-flight guard that could not answer

A guard that runs before the delete and cannot reach a verdict is a THIRD thing, and it is neither a skip nor a failure: the delete still goes ahead, so the outcome is normally 'deleted' (a delete that ALSO could not be addressed keeps 'skipped' — the two are independent), and what needs reporting is that a safety check was not enforced. Issue #2301 added indeterminateGuards to ResourceDeleteResult for it:

import { withIndeterminateGuard } from '../../deployment/delete-outcome.js';

const guard = await this.confirmSomething(...); // IndeterminateGuard | undefined
// ...perform the delete...
return withIndeterminateGuard(undefined, guard);

withIndeterminateGuard returns its input unchanged when guard is undefined, so the ordinary path keeps returning the back-compat void, and it preserves a 'skipped' outcome and its reason when both facts hold at once.

Two rules apply, both because the value is PERSISTED:

  • Proceed, do not refuse. These arms exist for probes a least-privilege caller may not be allowed to make. Refusing on an unanswerable probe strands every operator who never granted the permission; the durable record is what makes proceeding acceptable rather than silent.
  • Name the GUARD, not the API or the type. guard is a stable machine-readable id (cc-delete-region-identity, not s3-getbucketlocation) that lands in deployments/*.jsonl and is therefore a user contract; a future guard reuses the same event type with its own id rather than minting a second one. reason is the human half, and distinct causes get distinct text even when they reach the same proceed-anyway outcome, because the remedies differ.

The destroy runner then emits a RESOURCE_GUARD_INDETERMINATE event ALONGSIDE the resource's own outcome row (never instead of it), counts it into the destroy summary's N unverified suffix, and prints an aggregate warning. It moves no outcome counter and does not change the exit code. Note that the deploy engine and the rollback executor still DISCARD the field (#2422), so a guard reported from a delete on those paths is not yet recorded anywhere.

A provider that recurses into another destroy propagates the child's result. NestedStackProvider.delete returns { outcome: 'skipped' } when the child runner reports skippedCount > 0 OR interrupted; without that the parent would print ✓ <Child> (AWS::CloudFormation::Stack) deleted and exit 0 over a child stack that kept a live resource. The child's third not-gone signal, errorCount > 0, is a THROW rather than a skip (issue #1777): a resource was attempted and FAILED, so the parent's row must fail too, and calling it a skip would assert the opposite of what happened (a skip means no AWS call was issued). Word such a throw so it contains none of not found / does not exist / No policy found / NoSuchEntity / NotFoundException (and the deploy engine's was not found / ResourceNotFoundException) — both callers' catch blocks read those as an idempotent already-deleted success and DROP the state record, which is the outcome the throw exists to prevent. Note the asymmetry between the two return channels: since issue #1762 a { outcome: 'skipped' } value reaches the deploy-side sites too, but WEAKER — the template-removal DELETE warns and keeps the record — while a throw FAILS the resource at every caller, so converting a skip into a throw still changes cdkd deploy's behavior when the resource is removed from the template.

A provider that DELEGATES its delete to another provider propagates the delegate's outcome (issue #1778). CloudControlProvider hands a protected AWS::AutoScaling::AutoScalingGroup to ASGProvider.delete under --remove-protection, and discarding that verdict is the nested-stack hole one layer down: a skip inside the delegate reached the destroy runner as a plain successful delete. return await delegate.delete(...) is the whole fix. Type the delegate as ResourceProvider rather than as its concrete class, so the forwarding stays correct when the delegate widens its own return type. The alternative contract — ASSERT the delegate cannot skip — was rejected: it has to be re-verified every time the delegate grows an arm, and it fails loudly on a case the delegate considers merely unaddressable.

And a provider that pairs create() + delete() inside its own update() must not swallow the delete's outcome either (same issue). A skip does not throw, so it slips past the catch that would have warned — precisely when the old resource IS left behind. What to do depends on the ORDERING, and it is not the same for all of them:

  • create-then-delete (ACM certificate, IAM managed policy, IAM role, and the API Gateway Resource PathPart replacement): the new resource already exists, so the replacement cannot be aborted. WARN, in the same orphan wording the failure arm uses (The old role may be orphaned and require manual cleanup.), naming the skip's reason — and since issue #1819 ALSO return { outcome: 'partial', reason }, so the outcome is counted and recorded rather than living only in the log.
  • delete-then-create (SNS subscription): ABORT with a ProvisioningError BEFORE creating the replacement. Continuing would leave two subscriptions on the topic and deliver every message twice, and that duplicate is precisely what the CREATE would add — so not making it is the whole remedy.

State the premise as "the resource was not destroyed", never as "no AWS call was issued". ResourceDeleteResult's own contract says so, and its warning is not academic: NestedStackProvider reports skipped after recursing into a child destroy that may already have deleted the child's other resources. The abort is right under the weaker premise anyway — whatever the skipping delete did, adding a second live resource on top of it is not an improvement.

An abort on an update() path must be right for EVERY caller, not only the template-driven one — cdkd drift --revert and the rollback executor's revert arms call update() with a state-borne desired bag (see Pre-flight refusal). The SNS abort clears that bar because a second live subscription serves none of the three; where a refusal would not, downgrade it to a warning instead.

Two mechanical details any such abort inherits:

  • Do not interpolate the provider-supplied reason into the thrown message. The rollback executor wraps update() in withRetry, and retryable-errors.ts classifies by SUBSTRING — so a reason carrying does not exist, Rate exceeded, or because it is in use would flip a deterministic abort to "retryable" and burn the whole backoff schedule before a certain failure. Log the reason; interpolate nothing into the thrown message but the TEMPLATE logical id. The state-borne physical id is the worst candidate of all — the only skip family a REPLACE path meets today is literally "malformed physicalId in state", and cdkd import --resource '<id>=<anything>' puts an arbitrary string there.

  • ...and then markNonRetryable the error, because keeping values out of the message cannot close the hole. The match is a SUBSTRING, not an equality, so an ordinary composite logical id (MyDependencyViolationSub) still carries a pattern — measured — and a message that names nothing is not diagnosable. markNonRetryable(new ProvisioningError(...)) (from src/deployment/retryable-errors.ts) sets a non-enumerable Symbol.for marker that isRetryableTransientError consults BEFORE any name or message heuristic, so the refusal is terminal by DECLARATION and no wording can overturn it. Reach for it only where the error means "this cannot succeed on a retry" as a matter of cdkd's own logic — never for a relayed AWS failure, whose retryability is the classifiers' business.

    The fence is at the RETRY LOOP, not only in that one classifier: withRetry rethrows a marked error BEFORE choosing between the default classifier and a caller-supplied opts.isRetryable, and the destroy runner's delete-retry loop gates its own Too Many Requests message test the same way. So the message-only classifiers (isNameCollisionError's AlreadyExists, isNameCooldownError's QueueDeletedRecently / StateMachineDeleting — all of which match a bare logical id) cannot resurrect a refusal even though they cannot read the marker themselves. Keep the message discipline above regardless: it is what a future classifier added outside those two loops still relies on.

  • Point the remediation at the STATE record, not only at AWS. Both skip families shipping today — the state-borne composite-id arms and NestedStackProvider's propagation — describe a resource whose repair "remove the old resource in AWS" does not cover on its own, and for the malformed-physicalId one deleting the AWS resource fixes nothing at all. Name both repairs.

Note this is deliberately NOT symmetric with a THROWN delete, which those providers still warn-and-continue past: a throw may mean the delete partially landed or was transient, while a skip is a positive statement that the resource was not destroyed.

Producers today: the malformed-composite-physicalId family (src/provisioning/composite-id.ts, five arms), the nested-stack propagation above, and — since issue #1770 — eight same-class arms outside the composite-id family: both malformed LayerVersionArn arms in lambda-layer-provider.ts, the missing-FunctionName arm in lambda-permission-provider.ts, the no-properties / no-ServiceToken arms in custom-resource-provider.ts (and, since issue #3938 and #3960, its masked- and secret-reference-ServiceToken arms), the empty-policy-name arm in iam-policy-provider.ts, and both AWS::IAM::UserToGroupAddition arms in iam-user-group-provider.ts. Issue #3878 added the malformed-target-list arm in iam-policy-provider.ts: a recorded Roles / Groups / Users that is not a list of IAM names. Issue #3888 added the same arm for an AWS::IAM::UserToGroupAddition record's Users. Issue #4150 added one arm to each of those two providers: a destroy that resolved a secret-derived list and found a principal the secret named without the grant keeps the record, since the secret may have rotated. For an inline policy that is NoSuchEntity on the delete call. For a membership it is a ListGroupsForUser read BEFORE the call, because RemoveUserFromGroup succeeds for an existing user outside the group. Each exports its reason as a named constant beside the provider, so the wording is pinned by a test instead of retyped.

Issue #3952 added one shared arm, redactedDeleteAddressSkip (src/provisioning/redacted-delete-address.ts), for a delete whose recorded address property cdkd redacted (the *** mask or a secret {{resolve:...}} expression): ApiGateway / ApiGatewayV2 children, ECS services, Glue catalog-scoped resources, Scheduler schedules (unless the record holds the schedule's creation date, which with the recorded target ARN and role locates it in any group; a record with no stack region, or a schedule matching only the date, is still skipped: #4275), Route 53 record sets, security-group ingress rules, CloudWatch anomaly detectors, DB proxy target groups, and the fallback-less arms of the Lambda permission, IAM policy, UserToGroupAddition and ECS service deletes.

Three lessons from that issue's code review are worth reusing before you add a skip arm of your own.

Exhaust every addressable source first. A skip is not a free "safe" default: it preserves the record, warns, and exits 2 — on every re-run, so the destroy can never go green. Two of the eight arms were skipping although a second source held the value. RemovePermission accepts a full ARN as FunctionName, and the physicalId's documented <functionArn>|<statementId> shape carries one; PolicyName is in handledProperties and create() uses it verbatim as the real AWS name. Order the sources by what was DEPLOYED — the physicalId wins over the property, or a template edit that renamed the policy without a replacement having landed would send a delete for a name AWS never had.

Validate the fallback, or it is worse than the skip. A truthy NON-string (an unresolved intrinsic, an array) must not beat a good second source — the SDK URI-encodes it into a garbage label that 404s, and the idempotent *NotFound arm then reports DELETED. Check the fallback's REGION too: a provider holds one client, so a cross-region ARN sends the call to the wrong region and is reported DELETED the same way. And do not let the gate refuse a genuine source — the CFn primary identifier [FunctionName, Id] carries a BARE name as often as an ARN.

A guard that admits a record must check the path it opens reaches AWS. The IAM PolicyName fallback turned an honest skip into a silent deleted: with the name resolvable, a record naming no Roles / Groups / Users reached a body where every branch is skipped — zero AWS calls, returning undefined, i.e. DELETED. Trace the admitted path to an actual call, and use the same truthiness spelling the branches downstream use so a null-valued list cannot slip past.

"LEFT IN PLACE" is false when the parent is in the same stack. The destroy keeps going after a skip, and deleteGroup / deleteUser remove exactly those memberships, a deleted Lambda function drops its whole resource policy, a deleted IAM role drops its inline policies. Qualify the wording and name the command that drops the record through stateOrphanRecordRemedy (src/provisioning/state-orphan-remedy.ts): cdkd state orphan '<stack>' only on a stack destroy, --resource <logicalId> everywhere else, since on a deployed stack the whole-stack form drops every live record (#4602). Do not copy the qualifier onto an arm where it is false (a layer version and a Custom Resource's external side effects are undone by nothing).

"Repair state.json and re-run" holds on destroy AND on the deploy engine's template-removal DELETE — both keep the record (issue #1762). It does NOT hold for a deploy-side REPLACEMENT or rollback delete, which fail the resource and can leave the old one untracked, so every skip warning carries the same caveat compositeIdFormatMessage does.

Two judgment calls from that issue are worth reusing. The UserToGroupAddition arms logged at DEBUG, which read as "routine, nothing to do" — but GroupName and Users are both REQUIRED by the CloudFormation schema, so a record missing either is CORRUPT rather than empty, and the memberships AddUserToGroup created survive the destroy; they are skips, and the level was raised to WARN to match (a skip preserves state and exits non-zero, so the user needs the explanation at normal verbosity). An EMPTY Users: [] is the opposite case and stays a deleted: an array is truthy, so it falls through to the removal loop and correctly does nothing.

The *NotFound idempotent arms in those same files are deliberately untouched, as is CustomResourceProvider's backing-Lambda-is-gone pre-check — those mean the resource IS gone. The deploy-engine / rollback-executor callers consume the return value as of issue #1762: the template-removal DELETE warns and keeps the record, a replacement delete fails the resource, and a rollback delete counts as a per-op failure — except at the arm that deletes the NEW resource after the old one was already re-created, where the delete is best-effort and a skip only warns.

A create that mints a server-side id needs an idempotency token

Written for issue #2039.

The deploy engine wraps every provider.create() in its outer transient-error retry, and HTTP 500 / 502 / 504 are retryable (issue #2026). So a 500 whose request actually SUCCEEDED server-side re-invokes create() from the top. For a create whose id comes from AWS rather than from a name, the replay provisions a SECOND resource: it has no state entry, cdkd destroy never reaches it, and it bills indefinitely.

When the API takes a ClientToken / ClientRequestToken / IdempotencyToken / CallerReference, take it from the shared helper:

import { acquireIdempotencyToken } from './idempotency-token.js';

const token = acquireIdempotencyToken({ scope: 'RunInstances', logicalId });
const response = await this.client.send(
  new RunInstancesCommand({ ...input, ClientToken: token.value })
);
token.release(); // SUCCESS PATH ONLY

Three rules, each of which has a failure mode behind it:

  • Never derive the token per attempt. A Date.now() / randomUUID() value is worse than no token at all, because the call site then LOOKS idempotent while behaving exactly as before. AWS::Route53::HostedZone shipped that way.
  • Release only on success, and also wherever the provider DESTROYS the resource the token names (the RunInstances wiring-failure path terminates the instance). A released token is never handed out again, so releasing on a failure path defeats the mechanism, while NOT releasing after the resource is gone can hand a later create the resource it just destroyed — EC2 keeps a RunInstances token for ~24h.
  • Check what the API does with a repeat, and for how long. Most return the original resource; Route 53 REFUSES a repeated CallerReference (HostedZoneAlreadyExists), so that provider recovers by looking the zone up by its caller reference and adopting it. EFS refuses too: a repeated CreateAccessPoint ClientToken is refused with AccessPointAlreadyExists (which names the surviving AccessPointId) whether or not the repeat's parameters match (measured in #2442), so EFSProvider.createOrAdoptAccessPoint reads that access point back, confirms that its ClientToken is the one cdkd minted, that it belongs to the file system cdkd asked for, and that its PosixUser and RootDirectory are the ones this create requested (read through EFS's defaults: an absent or null POSIX user, a / root), and adopts it. Because the refusal ignores the parameters, the token alone would adopt a mismatched access point. A stable token that turns every retry into a hard failure is only half a fix. Note EFS documents no retirement period for this token -- the "one minute" in the EFS User Guide's "Creation token and idempotency" section is about file-system CREATION tokens -- so do not write a provider whose correctness depends on that duration; see the DERIVATION note in createAccessPoint.
  • Wrap the rethrow when you decline to adopt. DeployEngine's replacement path classifies a failed create by testing the TOP-LEVEL message for already exists / AlreadyExists -- exactly what the raw AWS conflict carries -- and by checking the error chain for a duplicate-name exception name. Rethrowing it bare makes the engine read a TOKEN collision as a physical-NAME collision, and under --replace fall back to delete-first: deleting the OLD resource and re-creating with the same still-unreleased token. Wrap it, keep the AWS error as cause.
  • Do not assume the SDK's auto-fill is a fix. A field carrying Smithy's idempotencyToken trait is auto-populated when the caller omits it -- with a FRESH value per request, which is the per-attempt derivation the first rule forbids. Measured on @aws-sdk/client-efs 3.1018.0: two identical CreateAccessPoint sends went out with different UUIDs, and a caller-supplied value went out verbatim.

EFSProvider's FILE SYSTEM CreationToken, FSxFileSystemProvider's ClientRequestToken and CloudFrontOAIProvider's CallerReference deliberately do NOT use acquireIdempotencyToken: those APIs bind the token to the resource for its LIFETIME, so a deterministic hash of the immutable create inputs is right there (it also lets the new file system coexist with the old one during a replacement). Derive it with stackScopedCreateToken, which also hashes the stack name and region, and refuses to run outside a withStackName scope: without them, two copies of one stack in an account sent the same token, and FSx and CloudFront handed the second stack the first one's resource (#4428). Take it through reserveStackCreateToken (src/provisioning/providers/create-token-ledger.ts), which folds in the stack's create-token ledger nonce -- replaced whenever cdkd lets go of a resource that still holds a token (#4438) -- and records the send before the create is sent, so a re-run after an interruption sends the same token and reports when that earlier attempt was first sent. Add the type to LEDGER_TOKEN_RESOURCE_TYPES, and pass any new site that keeps such a resource alive while dropping it from state to noteRetainedResource(type, logicalId), which rotates the nonce and drops that logical id's entry. The entry is kept until the deploy that wrote it succeeds (forgetRecordedCreateTokens after the final save): a re-run of an interrupted deploy needs it, and a later let-go must not find it. When the ledger cannot be read or written, reserveStackCreateToken throws: a create sent with any other token could not be found again by the re-run, and its resource would leak. A token stable across runs can still be held by a resource cdkd did not make, so a create takes a handed-back or named file system only when it was created after that create's FIRST send (in this process, or recorded by the ledger) -- judged on AWS's clock, never the host's: a host clock running fast would otherwise refuse every ordinary create. src/provisioning/providers/server-clock.ts reads the response's HTTP Date (withServerClock / earliestOwnCreationTime) and falls back to SigV4's five-minute bound without one. FSx refuses an older one outright, and EFSProvider.sendCreateFileSystem (EFS refuses a repeated CreationToken with FileSystemAlreadyExists) adopts the named one only when this process's own earlier attempt, or the ledger's, may have made it. A holder still deleting, or one whose read-back fails transiently, rethrows EFS's own error stamped markReplayMayCollide: a delete-first re-create's message-based retry waits it out, and no delete-first site acts on a holder cdkd does not own. That reasoning is per-API, not per-provider -- the same EFSProvider DOES take acquireIdempotencyToken for CreateAccessPoint. That choice does not rest on knowing how long the token lives, because the process-scoped derivation is right under both readings: if the token retires quickly, stability across runs buys nothing (no two deploys are that close together on one create) while it would put a --replace delete-then-create inside the replay window of the access point it just deleted; and if it instead binds for the access point's lifetime, a deterministic hash is worse still, since any re-create with unchanged inputs would collide with the live access point instead of creating one.

Where the API has NO token member, the choice is between disableOuterRetry (which makes the provider single-shot for EVERY transient error, IAM propagation included) and a pre-create reconcile that detects the previous attempt's orphan. Say in-code which you chose and what it costs.

A reconcile is not simply "delete what appeared since my baseline" — write it so the wrong answer LEAVES an orphan rather than destroying a live resource. "Newer than my snapshot" is equally true of the orphan your own attempt minted and of a resource something else created in the same window: a sibling resource in the same stack (the deploy engine dispatches at --concurrency 10 by default), a second cdkd process, or a human. Deleting the second kind is strictly worse than the bug being fixed — it destroys something cdkd state still advertises, and for a type whose readCurrentState cannot tell that the resource is gone (returning undefined rather than RESOURCE_NOT_FOUND), cdkd drift reports UNKNOWN rather than drift, so the loss is invisible. IAMAccessKeyProvider shows the shape:

  • Serialize per owning resource in-process. A Map<owner, Promise> queue, with the baseline read, the create, and the reconcile all inside it, so no sibling create can interleave with them.
  • Reconcile from the FAILURE path of the attempt that took the baseline, never from the top of the next attempt (this is about a reconcile that DELETES; a report-only lookup, below, runs at the top of the next attempt on purpose). A baseline that spans the whole retry schedule describes a window many seconds wide.
  • Send the create through a client that refuses the SDK's 5xx retry (withoutServerErrorRetries; #4639). The SDK's replay inside one send SUCCEEDS with a second resource, so the failure path, and the reconcile in it, never runs. The SDK still replays a socket reset or timeout, so ALSO run the reconcile on a success whose output replayFollowedAmbiguousAttempt reports, after recording the returned id in the set below (#4687). Not on replayedInSend alone: that is also true after a throttle the SDK retried, which minted nothing, and every needless reconcile is one more chance to delete another process's fresh key.
  • Require a creation timestamp at or after the attempt started, with a small margin for clock skew between you and the service.
  • Keep a set of ids this process created successfully and never delete one. This is the condition a baseline structurally cannot express, and it is the one that covers the case the in-process lock cannot: another process.
  • Report anything you decline to delete, at warn, and do not claim an attribution the code did not establish.

Where the API has no token and nothing can be deleted safely, src/provisioning/providers/ambiguous-create.ts is the shared shape (issue #2080: CreateKey, CreateUserPool, CreateGraphqlApi, and API Gateway's CreateAuthorizer, CreateDeployment, CreateApi and CreateIntegration, EMR's RunJobFlow, AddInstanceFleet and AddInstanceGroups, Lambda's PublishLayerVersion and CreateEventSourceMapping, AppSync's CreateApiKey, DLM's CreateLifecyclePolicy, ECS's RegisterTaskDefinition, and EC2's CreateVpc, CreateSubnet, CreateInternetGateway, AllocateAddress and CreateSecurityGroup, whose shared report is orphan-report.ts):

  • Keep the SDK from replaying a 5xx. Send the create through a dedicated client wrapped by withoutServerErrorRetries: the SDK's own retry of a 5xx inside one send is a duplicate nothing can see. The engine's retry covers 5xx and throttles; the SDK keeps its retry of throttles, connection failures, clock skew and socket resets, which the engine does not retry. Every token-less create whose replay COLLIDES instead of duplicating (a name-unique create: a stream, repository, role, function, table, cluster, directory or table bucket and the like — the providers that send one through this client and arm no latch, plus EC2's CreateSecurityGroup, which collides on its group name in the VPC, and CodeCommit's seed CreateCommit, which collides once the branch exists) needs this client and no lookup: the surfaced 5xx lets withRetry mark the collision as possibly this create's own (#3978), so it is never credited to another holder. The exception is a collision whose error does not name the holder: EC2's CreateSubnet collides on its CIDR, but InvalidSubnet.Conflict names no subnet id, so it keeps the lookup to tell the user which subnet to inspect. S3's CreateBucket is not one of these either: it keeps its latch and refuses a bucket that exists after its own server error rather than adopting it (#4686).
  • Look only after an AMBIGUOUS failure. AmbiguousCreateLatch.noteFailure arms on isAmbiguousOutcomeError thrown by the create call itself, and the next attempt takes it before creating again. A definite refusal (a 4xx, a throttle, IAM propagation's not authorized) costs no lookup. Pass the taken window back to noteFailure, so two ambiguous attempts in a row are both covered. A create that SUCCEEDS after the SDK replayed it inside its send (replayedInSend: $metadata.attempts > 1, e.g. after a socket reset) runs the same lookup right after the create, over replayedSendWindow, never offering the returned id as a candidate: exclude it by id when it joins the provider's RecentIdSet only after follow-up calls, so a resource a failed rollback leaves behind stays nameable by a later lookup (#4687).
  • Bound the window on BOTH ends. The latch records the ambiguous attempt's start AND end (each widened by a skew margin) and expires after AMBIGUOUS_LATCH_TTL_MS: where the resource carries a creation date, a same-named resource created later -- by another process, after this one gave up -- is never a candidate. A lookup over a type with no creation date cannot apply the window, and its report says so. Where only a per-candidate read carries the date (GetLifecyclePolicy, DescribeTaskDefinition), narrow by the list's own fields first, then read each remaining candidate; a list ordered newest first stops at the first candidate older than the window.
  • Adopt only on EXACT attribution, which a lookup never has. The one adoption in this family is a KMS key id that came back in this process's own response before a follow-up call failed, bound to a digest of its inputs. A NAME is not attribution, even a cdkd-generated one: names are scoped to the account and region, the stack lock to the state bucket and prefix, so another run of the same stack name can create an identical resource in the window. REPORT candidates and create anyway.
  • Lead the report with a READ command, and offer a delete command only after it, conditional on confirming the candidate is this deploy's orphan: a candidate may belong to another deploy. Where there is no window at all (no creation date: GraphqlApi, an AppSync API key, an API Gateway authorizer, an API Gateway v2 integration, an EC2 VPC, subnet, internet gateway or Elastic IP) print no delete command.
  • A lookup that fails, transiently or not, warns and lets the create proceed: nothing adopts, so a failed lookup has no stake worth failing a create over. Say only what was LISTED: list APIs are eventually consistent.
  • Tagging the create with a cdkd token was rejected as the attribution channel: it makes a TagResource permission (or a request-tag condition) a requirement of every create, collides with tag policies, and leaves a tag every read and drift path must strip.

Update removal semantics: clear-on-removal

CloudFormation resets a property removed from the template to its default. Most AWS Update* / Modify* APIs do the opposite: an absent input field means "no change" (merge semantics). A provider update() that maps template properties straight into the SDK input therefore silently drops removals:

// BROKEN for removal: properties['Timeout'] is undefined when the user
// deletes `timeout` from their template → the SDK omits the field → AWS
// keeps the old value, while CFn would reset it to the default.
Timeout: properties['Timeout'] as number | undefined,

The failure is invisible end-to-end: the diff layer correctly reports old → undefined, the deploy reports updated + success, state drops the field — so the very next cdkd diff says "No changes" while AWS still holds the old value. Permanent, undetectable divergence from CloudFormation (confirmed live for Lambda in issue #1155).

Rule: every optional, mutable property passed to a merge-semantics update API needs an explicit reset when it was present before and is absent now. There are two shapes, both in src/provisioning/update-removal.ts (issue #1160):

  • Declare it when the reset is a CONSTANT, CFn-shaped, top-level value that update() forwards verbatim. ResourceProvider.removalDefaults maps type → property → value; the update CALLER (the deploy's in-place update, both rollback revert arms) injects it into the bag update() receives, and only there: state keeps the template, and a reset echoed back through effectiveProperties is taken out again. update() starts with properties = withRemovalDefaults(this.removalDefaults, resourceType, properties, previousProperties, context), which injects the same values on a direct call and is a no-op once the caller did.
  • Keep a local clearOnUpdateRemoval for anything else: a coerced value (Number(...)), an aliased key, a clear that depends on another property or on the logical id, a nested member, or an SDK-shaped value.
removalDefaults = new Map([
  ['AWS::Lambda::Function', new Map<string, unknown>([['Timeout', 3], ['MemorySize', 128]])],
]);

// clearOnUpdateRemoval(newValue, previousValue, clearValue):
//   present -> pass through; removed -> explicit reset; never set -> stay absent.
MaxInstanceLifetime: clearOnUpdateRemoval(newLifetime, prevLifetime, 0),

A type whose update() has been audited for EVERY property also declares removalHandledInUpdate: the properties whose removal update() handles itself (a local reset, a diff, a required or create-only property, or a value left in place that update() names in its own warning). The caller then warns once per resource for any removed property in neither map, naming it as left at its current AWS value AND saying CloudFormation would reset it to its default — so a property with no such reset (FSx StorageCapacity cannot shrink), or one whose reset is unmeasured, goes into removalHandledInUpdate with its own warning, never into the shared line. A type without that entry is never warned about, because cdkd cannot say what its update() does.

drift --revert injects nothing: its previous side is an AWS readback, not a template, so UpdateContext.removedProperties is empty there. A local clearOnUpdateRemoval still reads that readback, which is what reverts a console change the baseline records as unset (a placeholder).

Checklist when writing or reviewing an update():

  • Classify the API: merge (absent = unchanged — most Update* / Modify* calls) vs full-replace (the whole config object is replaced, or removal is handled by a dedicated Delete* call, like the S3 bucket sub-config pattern). Only merge APIs need clear-on-removal — but a full-replace API has the MIRROR hazard instead, see Full-replace update APIs erase AWS-authored values.

  • For each optional mutable property: what does AWS need to receive to get back to the CFn default? Common shapes: the documented default scalar (3, 128), an empty container ({ Variables: {} }, []), an empty string ('' for description/KMS-key fields), or a sentinel ({ ApplyOn: 'None' }, { Mode: 'PassThrough' }).

  • Never synthesize a reset for a field that was never set — that turns every unrelated update into a spurious (and sometimes invalid) write.

  • The reset SENTINEL can vary per API and per value KIND within one API, so verify it per key rather than picking one for the whole call (issue #1609 item 1). ELBv2's three attribute APIs take an identical {Key, Value} list and disagree three ways: ModifyTargetGroupAttributes rejects an empty Value for EVERY key ("A target group attribute value must be specified"), while ModifyLoadBalancerAttributes / ModifyListenerAttributes accept '' for numeric and free-form-string keys but reject it for BOOLEAN / ENUM ones ("The value of 'deletion_protection.enabled' must be 'true' or 'false', but was ''"). So a single provider needs a per-key resolver — a documented defaults table with a fallback — not one sentinel. Three things make this worth its own bullet. The blast radius is the whole call, not the key: these APIs validate the entire attribute list, so ONE removed boolean fails the deploy — and it can fail the automatic ROLLBACK too, leaving no template-side way out. (The rollback re-runs the same diff with the sides SWAPPED. That turns a removal of key K into an ADDITION of K, so it is not the same refusal mirrored — but any key the two template sides do not share becomes a removal in the other direction, which is exactly how the live run failed on a SECOND key the forward pass never touched.) The previous side is not always a template, which decides what a safe reset value even is: cdkd drift --revert hands the provider the readCurrentState snapshot as previousProperties, so an attribute AWS reports but the template never set can look REMOVED. Writing a documented default for each would silently reset a dozen live settings. Skip a key whose CURRENT value already equals the default — an untemplated key is by definition sitting at its default, while a genuinely templated one is not. Since issue #1626 the caller meets you halfway: when the reverted resource has NO observedProperties (the baseline is the raw template, so cdkd cannot tell AWS-authored from out-of-band), runRevert MERGES every untemplated path into the desired bag it hands you, so those keys arrive with their AWS-current values on BOTH sides and your diff sees no change. (ELBv2 goes further: on desiredFromAwsReadback it sends no attribute removal at all, because no ELBv2 call removes a key, so an absence from even an observed baseline only means AWS was not returning the key at capture — issue #4147.) That covers the bulk case and the non-default residual the per-provider skip cannot — and, because it is on the desired side, it also covers a provider that replaces the bag wholesale and never reads previousProperties. It does NOT cover the observed-capture baseline, where an absence IS an intentional removal — so keep the skip. And mocked unit tests cannot find any of it — they agree with whatever sentinel the provider chose. The ELBv2 arms shipped with two tests literally named "AWS-documented clear" / "empty-string default" that pinned a payload real AWS rejects; only an integ whose removal phase actually drops the key surfaced it. If you add a clear-on-removal arm, add the removal phase to a fixture in the same change.

  • The CFn-reset premise itself is per-type — A/B it before writing the reset, because where CloudFormation does NOT reset, the reset is the bug. Removing a Cognito UserPool Policies sub-key (SignInPolicy / PasswordPolicy, or the whole container) reaches UPDATE_COMPLETE in CloudFormation with every live value intact (measured us-east-1 2026-09-02, issue #1979) — CFn performs the same merge-semantics no-op the raw API does, so a cdkd-side reset would DIVERGE from template compatibility. The correct shape there is pass-through-nothing plus a WARN naming the sub-key and the explicit-declaration remedy (warnOnUnremovablePoliciesSubKeys in cognito-provider.ts): under both A/B answers the only wrong option is silence, because the user who deleted the sub-key believes a change (often a TIGHTENING) landed.

  • Sub-structures can carry the same hazard one level down: an object that is present but missing its inner key (e.g. Environment: {} without Variables) may also read as "no change" — normalize it to the explicit clear shape (issue #1158; live-verified). Conversely some config objects are replaced wholesale (Lambda's LoggingConfig: sending {LogFormat: 'Text'} alone resets an unspecified custom LogGroup to the default — also live-verified), so verify per API, not by analogy. Lambda's ImageConfig is the same whole-replace shape (issue #1225, live A/B 2026-08-11): UpdateFunctionConfiguration with {EntryPoint} alone CLEARED the previously-set Command / WorkingDirectory, and dropping those two from a kept CFn ImageConfig block reached the same end state — so the provider passes the kept-partial block through verbatim and synthesizing per-sub-field clears there would DIVERGE from CloudFormation. Measure both halves: "does CFn reset it" and "what does the API do with a partial object" are two different questions, and only the second one tells you whether pass-through is already correct. A MERGING sub-shape is not automatically a cdkd bug — ECS DeploymentConfiguration is the counterexample, and it is the reason the question above must be asked as an A/B rather than answered from the API semantics alone (issue #1225, live A/B us-east-1 2026-08-13). UpdateService genuinely MERGES: a kept block carrying only MaximumPercent left MinimumHealthyPercent and the circuit breaker at their live values. That is exactly the silent-drop shape — and a CloudFormation stack given the same partial block reached the IDENTICAL end state on a real resource UPDATE, so the retained value is CFn's behavior too and normalizing it away would DIVERGE. (The END STATE is what was measured; that CFn's handler submits the same partial to UpdateService is the inference.) The same run settled the whole-block removal the #1160 audit had left open for this field: CFn issued a real UPDATE and did NOT reset the configuration, so passing undefined is right and adding a reset would have been the divergence. One level DOWN the answer flips, and the flip is the part worth remembering. A kept DeploymentCircuitBreaker missing Rollback has the nested struct REPLACED, not merged — the SDK ACCEPTS it and a live rollback went true -> false. CloudFormation refuses that template up front (Model validation failed (... required key [Rollback] not found)), because the registry schema marks the DeploymentCircuitBreaker definition required [Enable, Rollback] and DeploymentAlarms (the Alarms property) required [AlarmNames, Rollback, Enable]. Until issue #1802 cdkd enforced no nested required-ness, so it DEPLOYED that template and silently flipped the live setting — a divergence in the PERMISSIVE direction, NOT a parity row. The deploy pre-flight (src/provisioning/nested-required.ts) now refuses a PRESENT nested block missing a required member, but only for the types CloudFormation was MEASURED to enforce that on: a schema's required list alone is not evidence (CFn accepts an AWS::Logs::LogGroup tag with neither Key nor Value). Do not file the row under the ASG InstanceMaintenancePolicy disposition: there AWS ITSELF rejected the partial, so cdkd failed loudly too and pass-through really was parity. The generalizable check is therefore two questions, not one — does the API merge or replace the nested struct, and does anything on cdkd's side refuse the shape the way CFn does? For a PRESENT block missing a required member, the answer is the pre-flight in src/provisioning/nested-required.ts, which reads the fixtures' per-PATH nestedRequired section (inline nested objects included) and refuses only for its measured CFN_ENFORCED_TYPES. Answering the second question no longer needs a live describe-type: since issue #1800 each tests/fixtures/cfn-schemas/*.json carries a definitionRequired section (per definition, its required list; the top-level list under the reserved #top key), so a test can assert the actual lists a parity verdict rests on. Read a MISSING key as "nothing is required here" — a definition with no required list, or an empty one, gets no entry. Note the section's ABSENCE does not mean "not captured": seven types legitimately require nothing anywhere and carry no section at all, so use generatedAt (>= 2026-08-13) to tell a captured fixture from a stale one. Note also that only a definition's OWN top-level required array is captured — required-ness expressed through a oneOf combinator (AWS::S3::Bucket's TargetObjectKeyFormat) or on an inline nested object (AWS::WAFv2::WebACL's FieldToMatch.SingleHeader) is absent, so for those shapes a missing entry means UNKNOWN rather than "CFn requires nothing"; measure against the live schema before concluding a parity verdict. That absence is itself load-bearing: LinearConfiguration and CanaryConfiguration carry no required list, so a kept-but-partial block there IS reachable from a template CFn accepts — which is issue #1806, and was expected to be strictly worse than #1802 precisely because there is no CFn-side refusal to point at. It MEASURED as parity instead; the next paragraph is that measurement, and it is a good example of why the reachability of a shape does not predict its verdict.

    And REPLACE does not imply divergence — ask the second question before concluding one (issue #1806, measured us-east-1 2026-08-13 on the same type, SDK and CloudFormation A/B per block). The three DEPTH-1 DeploymentConfiguration blocks whose partials CFn ACCEPTS — LinearConfiguration and CanaryConfiguration (no required list at all) and LifecycleHooks[], whose element requires only LifecycleStages (the partial measured there is one element's TimeoutConfiguration) — are what made this the worrying case. All three replace the nested struct exactly as DeploymentCircuitBreaker does, but the absent member is filled with an AWS-side DEFAULT (StepBakeTimeInMinutes -> 6, CanaryBakeTimeInMinutes -> 10, the hook's Action -> ROLLBACK) and CloudFormation handed the same partial template reached the IDENTICAL end state every time — so the verbatim pass-through is PARITY and re-filling the member from the previous side would be the divergence. What separates them from the #1802 row is only the second question: nothing refuses the shape, on either side. Two further findings generalize:

    • A block reachable only under one MODE has to be probed in that mode, and the mode may not be the one the issue names. These three are documented as blue/green machinery, but LinearConfiguration is gated on Strategy: LINEAR and CanaryConfiguration on CANARY — AWS answers a BLUE_GREEN service carrying one with Linear configuration can only be present with LINEAR deployment strategy, so the fixture the issue prescribed could not have reached them at all. LifecycleHooks is not strategy-gated that way (measured under CANARY; its reach under ROLLING was not probed).
    • The default-fill is not phantom drift, because the deploy engine captures observedProperties from readCurrentState and the AWS-filled value lands in the drift baseline. Check that before reaching for effectiveProperties: a value AWS computes belongs in observedProperties by this file's own rule, and the capture already puts it there. It holds only while the property stays drift-COMPARED, so fence it — declaring the property in getDriftUnknownPaths later would falsify the parity verdict with every send-side test still green.

    The rest of the tree splits TWO ways, plus the array shape below — so neither "the other blocks are refused" nor "the rest is parity" is the thing to carry away, and DeploymentAlarms keeps the disposition above (the permissive divergence, now refused at pre-flight):

    • AWS ITSELF REFUSES (the ASG InstanceMaintenancePolicy disposition — cdkd fails loudly, no nested required-ness check involved): a hook element missing LifecycleStages; a hook element missing TargetType (absent DEFAULTS to AWS_LAMBDA, which then demands a target ARN and role a PAUSE hook does not carry, so dropping ONE member arms a requirement for TWO others); and DeploymentCircuitBreaker.ThresholdConfiguration missing Value.
    • AWS RETAINS it and CLOUDFORMATION RESETS it — the one DIVERGENCE the sweep found, and the opposite polarity to the permissive one above: DeploymentCircuitBreaker's OPTIONAL children (ResetOnHealthyTask, and the whole ThresholdConfiguration block). Dropping them from an otherwise complete parent is a CFn-accepted partial that reaches AWS; from the identical baseline, with a sibling flipped in the same call so the update demonstrably applied, cdkd left false / {COUNT, 7} intact while CloudFormation reset them to true / {BOUNDED_PERCENT, 50}. So cdkd fails to apply a removal CFn applies — too STICKY rather than too permissive. Filed as issue #1861 and FIXED there: those two members — and only those two, and only while BOTH template sides still declare the DeploymentCircuitBreaker block — are routed through clearOnUpdateRemoval in ECSProvider.resolveDeploymentConfiguration, so a declared-then-dropped member is now sent as its AWS default. Everything else in the property still ships as the verbatim pass-through the parity rows measured.

    The LifecycleHooks ARRAY is replaced WHOLESALE, so a per-element drop is a non-issue. Three things generalize:

    • Enumerate the tree from the registry definitions, not from the blocks the issue names. #1806 named three; the live schema has more, and the two the issue never mentioned (ThresholdConfiguration, ResetOnHealthyTask) are the ones that produced new behavior.
    • Check the CHILD's own required list before calling a nested partial reachable. A satisfied PARENT requirement does not make the child's partial CFn-accepted: ThresholdConfiguration carries required: [Type, Value], so CFn refuses it up front — the parity there comes from BOTH engines failing, not from agreeing on an end state. What decides the verdict is whether AWS ACCEPTS, not whether CFn refuses: DeploymentCircuitBreaker missing Rollback is refused by CFn and ACCEPTED by AWS, which is precisely why that row was the permissive divergence (#1802) and this one is not.
    • Members of ONE block can carry DIFFERENT semantics. In DeploymentCircuitBreaker, the required Rollback reads as replaced while the optional children are retained. Measure per member; do not generalize a row to its siblings.
    • Where the API RETAINS an omitted optional member of a STILL-DECLARED struct, cdkd diverges from CloudFormation, which treats the omission as a REMOVAL and resets the member to its default. Scope that sentence carefully, because two neighbouring shapes behave differently and both were measured: removing the WHOLE property resets nothing under either engine (the depth-0 row), and a member the template NEVER declared is left alone by CFn too — an out-of-band value survived a CFn update that changed an unrelated member — and it still survived when a SIBLING INSIDE the same block was flipped, which is what excludes the competing reading "CFn re-serializes a struct with defaults whenever its declared content changed". So this is a previous-present / current-absent REMOVAL, the semantic clearOnUpdateRemoval implements one level up, NOT a handler that materializes every default. Treat the member-vs-whole-property boundary as EMPIRICAL rather than derived: a top-level property is previous-present / current-absent too, and removing it resets nothing. The LinearConfiguration family is parity because the API ITSELF default-fills there, so both engines land in the same place. And the two behaviors are indistinguishable until you probe a member whose live value DIFFERS from its default — measuring from a live value that happens to EQUAL the default proves nothing, which is a mistake this measurement made once and had to redo. Capture the resource's identity on both sides too (a replacement trivially shows defaults, and reads as a reset).

    A note on HOW to probe, because it changes the answer: the AWS CLI refuses the ThresholdConfiguration partial CLIENT-side (botocore validates the required trait before sending), so it never reaches the service. The JS SDK cdkd uses declares the same member required in its TYPE (value: number | undefined, a non-optional key, unlike the optional targetType?:) but performs no runtime required check, so it serializes the partial and the refusal arrives from the SERVICE. The models agree; only where the validation runs differs — so probe through @aws-sdk/client-*, or a CLI probe reports the right verdict for the wrong reason, and reports the WRONG verdict entirely if the service would have accepted it.

  • Unit-test three shapes: removed → exact reset value; never-present → stays absent; mixed kept/removed → kept fields pass through unchanged.

  • A per-key removal test (one key dropped from a still-present map) does NOT cover whole-block removal (the map itself dropped) — test both.

  • Walk the DEPTH-1 members too, not just the top-level properties. The removal audit above is normally run over a type's top-level properties, and a nested struct forwarded whole (a recursive case-flip, a spread, a raw cast) hides the same bug one level down: the API retains the omitted member while CloudFormation resets it. ECSProvider.resolveDeploymentConfiguration (issue #1861) is the worked example — reuse clearOnUpdateRemoval with the PREVIOUS side read raw, since only its PRESENCE is inspected. Two traps that both cost a round there: a member the template NEVER declared must stay absent (sending the default clobbers an out-of-band value CFn preserves), and the reset re-runs with the sides SWAPPED during a rollback replay (issue #1609), so key strictly on previous-present / current-absent rather than on "the two sides differ".

  • Not every removal is a VALUE on the same call. clearOnUpdateRemoval fits a property that maps to an input FIELD, so a reset is "send the default instead of omitting". A property whose apply is a separate API call needs a different call on removal, and there is no reset value to pass — the route53-provider.ts pair (issue #1160) is both spellings: HostedZoneTags applies via ChangeTagsForResource and its removal is the RemoveTagKeys argument (a previous-minus-desired KEY DIFF, not a value), while QueryLoggingConfig applies via CreateQueryLoggingConfig and its removal is DeleteQueryLoggingConfig on a sub-resource. Both looked like no-ops precisely because the apply helper took only the DESIRED bag and had nothing to diff against; the fix is threading previousProperties into the helper, keeping it optional so create() keeps its existing REMOVAL behavior (it has no previous side, so nothing is ever removed — though a shape guard you add along the way does change what create accepts, so do not claim it is byte-identical). Gate the removal on the previous side actually having carried the thing, rather than probing AWS on every update — a config cdkd never created is drift, which cdkd drift owns, not a removal reset. Two things bite specifically in this shape, both found by review on the route53 batch after its first integ had already passed:

    • A malformed value must not read as a removal. The desired side reaches the helper through a tolerant reader, and anything the reader cannot use collapses to the same emptiness a real removal produces — so an unresolved intrinsic UNTAGS a live zone or DELETES a live config. Refuse a LOSSY read, not merely a wrong container: [{ Key: { Ref: 'X' } }] is genuinely an array, so a shape-only check lets the destructive case through one level down. Compare the parsed length against the raw length. And check what your own readCurrentState emits before calling a shape malformed — route53's emits QueryLoggingConfig: {} for "no live config", which cdkd drift --revert feeds straight back as the DESIRED side.
    • A failed removal is not self-healing, so it must not be swallowed. A failed ADD is retried by the next deploy because the template still declares it. A failed REMOVAL is not: the update returns success, state is rewritten WITHOUT the property, and the next deploy's previous side no longer carries it — so the value survives on AWS forever with cdkd diff clean, which is the #1160 failure mode the fix existed to close. Throw on the removal branch even when the surrounding helper is warn-and-continue.

Full-replace update APIs erase AWS-authored values

A full-replace update API needs no clear-on-removal (see Update removal semantics) — omitting a key IS the reset. Its hazard runs the other way: whatever the payload omits is erased, including values AWS itself wrote and your template therefore has no representation for. No template-side diff hints at the loss, and no unit test can catch it, because the values only exist on the live resource.

The live case (GlueProvider, issue #1461): UpdateTable replaces TableInput wholesale, so an Iceberg table's Parameters.table_type and Parameters.metadata_location — written by Glue at create time — were erased by a deploy that changed only TableInput.Description, silently degrading the table to a plain external table while the deploy reported success.

When a full-replace update sends a general-purpose bag (a Parameters / Properties / Tags-shaped map AWS can write into), read the live resource first and merge those entries back. Key the merge on "present in NEITHER template side", using previousProperties:

key in desired key in previous outcome
yes (either) the user's value wins
no yes the USER REMOVED it -> stays removed
no no AWS-authored -> preserved from live

Restoring anything merely absent from the DESIRED side would make user-authored entries unremovable — the mirror image of the bug — and would break cdkd drift --revert's ability to clear a console-side addition (on an observed-capture baseline that path passes the AWS-current snapshot as previousProperties, so every live key lands in the previous column and nothing is added back). On a resource with no observedProperties the same path MERGES every untemplated live key into the DESIRED bag instead (issue #1626) — previousProperties is still the full AWS-current snapshot — so an untemplated live key arrives in BOTH columns at its AWS-current value, so this table's first row keeps it — which is the intended outcome there, since the raw template cannot distinguish an AWS-authored entry from a console-added one.

State the price. The merge cannot distinguish an AWS-written entry from a console-written one — neither appears on either template side — so every out-of-band addition to that bag becomes permanent, and invisible to cdkd drift once the next deploy folds it into the state baseline. That is an acceptable trade against silently destroying an AWS-authored value, but it is a real semantic change: document it on the resource type, and give users the removal path (delete it directly, or declare-then-undeclare it so it becomes a normal user-authored removal).

Close the TOCTOU window if the API lets you. Reading a value and writing it back is not atomic; a concurrent writer landing in between is silently undone by your write-back. Check the update request for an optimistic-concurrency member (Glue's UpdateTableRequest.VersionId; the ConcurrentModificationException in a command's documented error list is the tell) and send the version you read, so a concurrent change fails loudly instead. Send it on EVERY update rather than only on the ones that write back live values: the read runs immediately before the write, so the version is never stale unless somebody else really wrote in between, and an update that merged nothing still ships a wholesale replace that a concurrent write would lose. (Scoping it was tried on Glue and was wrong twice over — it left the empty-live-read case unguarded, and its premise that a pure template push carries a stale version was false.) When the API has no such member, say so where users will read it rather than leaving the exposure implicit.

Prove the precondition, do not assume it. An AWS field named like a version token is not necessarily enforced — Glue documents VersionId only as "the version ID at which to update the table contents". Pin the semantics with a real-AWS probe that advances the version out of band and requires the stale replay to be REFUSED (and refused with a concurrency error, not any error). Without that, an ignored token makes the whole guard a placebo that reads as protection in review.

Two placement rules go with it:

  • Fail closed on a read failure. Only a definitive not-found may degrade to "no live values" (the update's own error is more actionable). Any other failure must throw, naming the required IAM action — silently skipping the merge reinstates the erasure the read exists to prevent. Wrap the original error as cause so a transient throttle stays retryable.

  • Throw OUTSIDE the update's try, and order the read AFTER the pre-flight validation, so the typed error is not re-labelled by the catch wrapper and a refused update issues no extra API call. Move ONLY the read — leaving the payload-building code outside the try as well turns a malformed-template crash into a raw TypeError with no resource context.

  • Check that every provider.update() call site retries. Adding a read to update() makes it newly sensitive to throttling. deploy-engine/update-in-place.ts and drift.ts wrap their calls in withRetry; rollback-executor.ts did not until issue #1461, so the new read would have failed a rollback op that previously issued no read at all — and the best-effort catch there counts that as a failure and moves on. A transient failure on a RECOVERY path is the worst place to introduce one. A new withRetry must carry the two conventions the surrounding sites use, or it trades one bug for another: honor provider.disableOuterRetry (CustomResourceProvider / NestedStackProvider set it AND implement update() — re-invoking a Custom Resource derives a fresh RequestId + pre-signed URL and strands the first response at an S3 key nobody polls), and thread isInterrupted / onInterrupted (a rollback polls interrupts only BETWEEN ops, so an un-threaded probe leaves Ctrl-C dead for the whole backoff schedule).

    The second of those is now MECHANICALLY ENFORCED (issue #2053, after a count found 11 sites under src/provisioning/providers/** and the pair threaded at exactly one of them). vp run audit:withretry-interrupt:check (scripts/check-withretry-interrupt.ts, a CI step) fails on any withRetry under src/provisioning/** whose options object literal does not declare BOTH names. Half a pair is a defect too: without onInterrupted, withRetry throws a bare Error('Interrupted') that names no resource, and without isInterrupted the other half is never called. A conditional spread does not count — it reads as threading to a reviewer while being absent whenever its guard is falsy — and an options bag the checker cannot READ (an identifier, a spread) is reported rather than skipped, so inline it.

    Do not hand-roll the watch. There is exactly one: startInterruptWatch in src/provisioning/interrupt-watch.ts, and the critic above checks PROVENANCE, so a locally-built object with the right property names fails. Four module-local copies existed briefly and could not agree with each other, which is what made two of its properties wrong; all four are gone. Call it per WAIT, thread watch.isInterrupted / watch.onInterrupted, and dispose() in a finally.

    Four properties of that module, each load-bearing and each the fix for a concrete bug:

    • Per WAIT, never on this. Providers are registered as SINGLETONS serving concurrent resources, so provider-level state is some other resource's.

    • onInterrupted returns an InterruptedWaitError, not a bare Error. deploy-engine/execute.ts decides whether to ROLL BACK by asking what the failure was, and its own InterruptedError is engine-internal (not re-exported, and providers do not import the engine), so a bare Error from a provider read as a genuine resource failure and rolled the whole stack back on Ctrl-C. Use isInterruptedWaitError to recognise one. It walks the cause chain — every provider catch re-wraps — with a visited set and NO depth ceiling, because the chain grows by one per nested-stack level and a ceiling sized against today's nesting silently reinstates the rollback bug.

    • The latch is STICKY, and only a COMMAND clears it. A wait STARTED after the signal begins is already interrupted. Clearing it when the last watch is disposed looks equivalent and is not: sequential waits always empty the live set between them, so a GlobalTable delete running the #1521 gate, then the index-busy loop, then the gone-wait left the two multi-minute waits after the gate DEAF. forwardSigtermToSigint() opens and closes the command scope. The process listener is likewise never removed BETWEEN waits — one torn down in that gap cannot record a signal landing in it.

    • It arms only inside a command that OWNS interrupt handling, signalled by that scope rather than by counting SIGINT listeners. Any listener disables Node's default terminate, so a command with none of its own (cdkd drift --revert reaches provider.update) must not gain one here. A listener COUNT is the wrong test and was the first cut: drift runs at concurrency 4, and a concurrent CloudFront / ACM / Route53 wait installs a transient SIGINT listener, so a wait starting in that window armed the shared handler permanently.

    • The handler force-quits when it is the LAST SIGINT listener. A command scope is not the same as a live graceful path: destroy.ts registers no handler and destroy-runner.ts removes its own in a finally, so between two stacks of a multi-stack destroy the watch is alone — and merely latching there SWALLOWS the Ctrl-C, letting the next stack delete on. Alone, it does what Node would have done with no listener at all, plus a cdkd force-unlock hint it cannot verify but cannot afford to omit.

      The corollary is a rule for any command holding a lock: release it BEFORE unregistering your SIGINT handler. Unregister-first leaves a window where the command holds the lock with no handler of its own, and the force-quit turns a Ctrl-C there into a stranded lock for its full 30-minute TTL. destroy-runner.ts had it backwards in its main finally and correct in its strong-ref refusal path; both now read the same way, and destroy-runner-lock-release-ordering.test.ts pins it by observing which listeners are registered at the moment releaseLock is entered. rollback.ts had the same defect until #2118, whose fix adds one thing worth copying into any command you write: both unregistrations sit in a nested finally, because the .catch() on a release covers a rejection and not a synchronous throw. It is pinned by rollback-lock-release-ordering.test.ts, which measures the SIGINT and the SIGTERM listener sets SEPARATELY — a fix that reordered only the SIGINT half would leave unforwardSigterm() running before the release, and CI cancellation delivers SIGTERM, so that half strands the lock on its own. What is NOT a rule is the order within the teardown pair, and the first cut of that fix asserted it in four places before a review round measured it: the two calls are adjacent and synchronous, so no signal can be delivered between them, and unforwardSigterm() first does not empty the SIGINT set anyway because the command's own handler is still registered. Order them for readability; the requirement is release-before-unregister. deploy-engine/deploy-flow.ts is now the only site that still unregisters first, and is safe only because deploy.ts holds a handler that outlives it — stated at that call site, because a surviving instance of a corrected anti-pattern has to explain itself.

      And the corollary's own corollary: keeping the handler armed for longer moves the window a signal can arrive in, so re-check anything that READS the interrupt. destroy-runner.ts assigns result.interrupted once, inside its try — so arming across the teardown made a first Ctrl-C there set lock.interrupted after the only read, leaving the flag false and letting destroy --all delete the next stack. It re-syncs with result.interrupted ||= lock.interrupted && statePreserved at the end of the finally; the statePreserved gate keeps a stack whose state was already deleted from reporting interrupted. That is tactical: the real defect was that flag being the only channel, which #2117 closed with a command-scoped interrupt handler the --all loop reads live.

    An interrupt is not a failure — but it is not a reason to leave AWS state behind either. Two shapes to check in any wait you add, and they pull in opposite directions:

    • A best-effort catch that swallows into a warn will blame AWS for a user abort and then carry on to the next write. applyAutoScalingDiff re-throws instead, which stops the burst.
    • A partial-create cleanup arm must STILL run its cleanup delete on an interrupt. "Ctrl-C must not delete what you just made" is the intuitive answer and it is wrong here, because create() is throwing: the physical id never reaches state, the rollback journal records physicalId: undefined (unless the arm calls markCreatedBeforeFailure, above), and the rollback executor classifies it skip-failed-unknown. Unmarked, nothing holds the id, so the choice is delete vs orphan forever — and the orphan fails every later deploy on a name collision that neither rollback nor destroy can reach. Both the ELBv2 Listener and Cloud Map Service arms clean up, and each prints the physical id plus a manual delete command BEFORE attempting it, so a process killed mid-cleanup still leaves the user a handle.

    An interrupt must also never be read as "already deleted". Both delete paths decide a resource is already gone by SUBSTRING-matching the error message, and an interrupt's message embeds a name the user chose — so a logical id containing NotFoundException used to drop a live resource's state row. destroy-runner.ts and the deploy engine's delete arms (deploy-engine/update-in-place.ts, deploy-engine/delete.ts) all check isInterruptedWaitError ahead of that match; any new message-based classifier on a delete path needs the same guard.

    A hand-rolled poll on the same path owes the same treatment, and issue #1952 is why the rule is stated for the PATH rather than per call: interrupting one wait while the next one still sits out its cap buys nothing. What a poll DOES on the signal is its own call: throw when nothing has been accepted yet, stop waiting when the operation is already in AWS's hands (see waitForReplicaGone / waitForTableGone in dynamodb-globaltable-provider.ts for both answers on one path).

Only bags AWS actually writes into qualify. A purely user-authored bag (Glue's JobUpdate.DefaultArguments, ConnectionInput.ConnectionProperties) must NOT be merged — preserving a console-side addition there would break drift --revert. Audit the sibling update APIs on the same provider when you fix one; the divergence is per-bag, not per-provider.

Returning attributes

Return attributes accessible via Fn::GetAtt:

return {
  physicalId: bucketName,
  attributes: {
    Arn: s3BucketArn(bucketName, region),
    DomainName: s3BucketDomainName(bucketName, region),
    RegionalDomainName: s3BucketRegionalDomainName(bucketName, region),
  },
};

Never hardcode arn:aws: or amazonaws.com in a value you build: outside the commercial partition (aws-cn, aws-us-gov, ...) both are wrong, and the value is still structurally valid, so nothing downstream catches it. Derive the partition and URL suffix from the region with derivePartitionAndUrlSuffix (src/utils/aws-partition.ts), or call a shared builder such as the src/utils/s3-endpoints.ts ones above, which the SDK provider and the resolver both use (and the Cloud Control route for Arn), so every route records the same value.

getAttribute() for live Fn::GetAtt resolution

Beyond the initial create/update return value, providers should implement getAttribute(physicalId, resourceType, attributeName, logicalId) so that live attribute reads succeed when the value is not in cdkd state — specifically the cdkd orphan per-resource flow, which splices each referenced attribute into sibling references. It always reads live first. When that read answers, it substitutes the orphan's recorded attribute instead wherever cdkd's own Fn::GetAtt resolution would serve that value, and keeps the live answer only when the record lacks it or holds a value that cannot be spliced (a redaction mask, a {{resolve:...}} reference, a stale placeholder ARN, a VPC's Ipv6CidrBlocks, an impossible empty value). A credential-named attribute (SecretAccessKey), a known secret-valued attribute (an AppSync API key's ApiKey, an IPAM verification token's TokenValue, an IVS stream key's Value), a value holding a credential-named key, and any custom-resource attribute are never taken from the record, since the value may be a plaintext secret. A live read that fails or answers nothing leaves the reference unresolvable without --force, as it always did, so a provider without getAttribute gets nothing from the record either. --force's cached fallback does not splice those secret classes either, nor a mask or a {{resolve:...}} reference: the reference stays unresolved (#4602). A live read addresses the resource by its recorded name, so after the resource was deleted and another one took that name it describes the newcomer; the recorded value is the one cdkd's own Fn::GetAtt resolution would choose. A recorded attribute AWS changes later (an instance's PublicIp) can be stale.

Conventions:

  • Return undefined for unknown attribute names. Do not throw.
  • Where a provider does throw (several existing ones still refuse an unknown attribute), its ProvisioningError takes logicalId in its logical-id slot and physicalId after it: the retry classifiers anchor on that slot, and a physical id can be built from a secret (go-to-k/cdkd#4222).
  • Treat *NotFound exceptions as undefined rather than re-throwing — the live fetch is best-effort, and under --force cdkd orphan falls back to the cached state.attributes when the live resolution comes back empty.
  • Prefer derivation from physicalId when CFn returns derivable values (S3 Bucket DomainName/Arn, SNS Topic name from ARN tail, SQS QueueName from URL tail) so the call is free.

Known coverage gaps (deliberate)

The following CloudFormation Fn::GetAtt return values are documented but not implemented in cdkd's getAttribute(). They require a separate AWS API call beyond what cdkd already makes, are rarely referenced from CDK code, or both. If a real-world stack hits one of these, file an issue — the small additional call is reasonable to add.

Resource Unsupported attribute Why deferred
AWS::SQS::Queue (none) All three CFn return values are covered.
AWS::S3::Bucket (none) All five CFn return values are covered.

readCurrentState() for drift detection

readCurrentState(physicalId, logicalId, resourceType, properties?, context?) returns the AWS-current snapshot of a resource for cdkd drift and cdkd state refresh-observed. The drift comparator walks state's top-level keys only (intentionally — to avoid surfacing every FunctionArn / RevisionId / LastModified / etc. that AWS auto-attaches to every response). That design has one consequence the provider author MUST account for:

Any user-controllable top-level CFn property update() can mutate must be emitted with a placeholder when AWS returns the field as undefined / empty.

If the provider omits the key on the empty path (e.g. if (cfg.Environment?.Variables) result['Environment'] = ...), then on a resource that was deployed WITHOUT that key in its template, state.observedProperties never carries the key — and the comparator's state-keys-only walk skips the field forever. A user adding the property in the AWS console after deploy is silently invisible to drift.

Use these placeholders consistently:

Type Placeholder Example
Array ?? [] result['ManagedPolicyArns'] = arns; (after building the list)
Map / object (when AWS returns the whole object as undefined) ?? {} result['Cors'] = cors; (after building, even if cors ended up empty)
Optional string ?? '' result['Description'] = resp.Description ?? '';
Boolean / numeric scalar ?? <semantic-default> Status: resp.Status ?? 'Suspended', BlockPublicAcls: cfg?.BlockPublicAcls ?? false
Tags map ?? [] (already covered for Tags by PR #145) result['Tags'] = normalizeAwsTagsToCfn(...);

When the guard is justified — keep it:

  • Immutable on create — BucketName, Lambda Runtime (when create-time-only), IAM RoleName. The field can't change at all; AWS returning undefined is a wire-layer artifact, not a "user could add this." Skip emit.
  • AWS-managed read-only — FunctionArn, RevisionId, CodeSha256, timestamps. These are not in the CFn template; cdkd state never carries them. They should NOT be in readCurrentState output at all.
  • Write-only — Code: { S3Bucket, S3Key }, SecretString, LoginProfile.Password. AWS does not return these on read. Declare via getDriftUnknownPaths() so the comparator skips the entire subtree (see "Known coverage gaps" below).

Wire-layer filtering — the drift comparator does NOT apply per-type denylists for SDK provider results (those are reserved for the CC-API fallback path). If your provider's SDK response includes AWS-managed fields you don't want to surface, do NOT assign them in the first place.

Test convention (mandatory for any provider with readCurrentState): every provider test file MUST have an it('emits placeholders for every user-controllable top-level key on AWS minimum response') block that:

  1. Mocks the SDK to return the resource exists with all optional fields undefined / empty (just required fields like Name / ARN).
  2. Calls readCurrentState(physicalId, logicalId, resourceType).
  3. Asserts Object.keys(result).sort() matches the complete expected key list for that resource type — not a subset.
  4. Spot-checks the placeholder values for the most fragile keys (?? '' strings, ?? [] arrays, ?? {} objects, ?? <semantic-default> scalars).

Example template:

it('emits placeholders for every user-controllable top-level key on AWS minimum response', async () => {
  mockSend.mockResolvedValueOnce({
    /* SDK response: required fields only, all optionals undefined */
  });
  const result = await provider.readCurrentState('phys-id', 'L', 'AWS::My::Type');
  expect(Object.keys(result ?? {}).sort()).toEqual(
    ['Key1', 'Key2', /* ... complete list ... */ ].sort()
  );
  expect(result?.Key1).toBe('');           // string placeholder
  expect(result?.Key2).toEqual([]);        // array placeholder
  expect(result?.Key3).toEqual({});        // object placeholder
});

See tests/unit/provisioning/lambda-function-provider-readcurrentstate.test.ts and tests/unit/provisioning/cognito-provider-readcurrentstate.test.ts for canonical examples.

This is the structural defense against the "provider author forgets to emit a key" regression class. Without it, the bug only surfaces when a user runs drift on a resource configured exactly the way the test missed (and PR review missed). The test makes silent regression mechanically impossible — a refactor that drops a placeholder fails the key-set assertion immediately.

The properties argument is the DESIRED side, and gating an emission on it is a fix with a TRANSITION COST — do not reach for it alone (issues #1742 / #1760). Every caller passes the resource's state-recorded / template-resolved bag (drift.ts, the deploy engine's observed capture, import.ts, state.ts), so a provider CAN ask what the user actually declared, and the temptation is to use that to stop emitting an AWS-COMPUTED value the desired side can never carry.

The half that is easy to miss: observedProperties bags already in S3 were written by a binary that DID emit the member. The moment the new binary stops, the comparison is baseline has it vs aws side undefined, and the first cdkd drift after the upgrade reports a one-sided phantom on every such resource. It does not self-heal — the observed capture runs only on CREATE / UPDATE, and the auto-refresh skips a record that already has a capture (the one it revisits, a baseline holding a *** mask, is re-captured only when the readback reproduces that baseline, which a missing member prevents) — so the window is however long until the user's next deploy, which is exactly when they are running drift instead. Measured live on AWS::DynamoDB::Table (#1760).

So the emission gate needs a companion that removes the path from the COMPARISON too:

  • For a TOP-LEVEL key, declare it in getDriftUnknownPaths(resourceType, properties) — per-resource via the #1602 seam, so it is ignored on BOTH sides when the template declares nothing and still compared when it does. That covers the transition and the steady state with one declaration.
  • For a PER-ELEMENT key there is no such declaration: an ignore-path never crosses an array (see the divergence note below), and declaring the enclosing array instead switches drift off for the whole subtree. That case needs a BOTH-SIDES normalizer in the drift-protocol-normalize.ts mould, and a live test seeded with a STALE observed baseline — a fresh-deploy fixture takes both sides from the same readback and structurally cannot exercise the transition. AWS::DynamoDB::GlobalTable's GlobalSecondaryIndexes[].WarmThroughput is exactly this shape, and is now the WORKED example rather than the open question: PR #1859 closed it through the canonicalizeDriftProperties seam (issue #1784), which IS the both-sides normalizer this bullet asks for — the provider names the member in a closed DRIFT_STRIPPED_INDEX_MEMBERS table and strips it from each bag, so an already-written baseline and a post-fix readback converge on carrying no member at all. Copy the NORMALIZER half's shape — but note it is only half of what this bullet asks for: the live test seeded with a STALE observed baseline was NOT shipped with it and is still outstanding (issue #1939), so #1859 is the worked answer for the mechanism and an open question for the proof. Take the two things that come with it as well: the emission change and the normalizer had to ship in ONE change (landing the readback half alone is precisely the stranding described above), and the accepted cost is that a DECLARED per-element value stops being reported entirely — the hook sees one bag with no desired-side reference, so it cannot express the declared-gate getDriftUnknownPaths gives you for a top-level key. The sibling AWS::DynamoDB::Table type adopts the same seam with a narrower key: it strips only members an SDK index DESCRIPTION carries and a CFn value never does (IndexStatus, NumberOfDecreasesToday, a WarmThroughput with Status, ...), so a baseline written by an older binary converges while a template or a current readback passes through by identity and keeps every declared value reported.

Contrast an ORDERING fix, which has no such asymmetry: getDriftUnorderedPaths canonicalizes both sides, so an existing baseline and the readback converge rather than diverge, and it can ship on its own.

getDriftUnknownPaths() for unreadable fields

When AWS does not return a field that cdkd state stores (write-only fields, or a CFn property whose round-trip back to the template shape isn't worth implementing yet), declare the path so the comparator skips it instead of firing guaranteed false-positive drift on every clean run:

getDriftUnknownPaths(): string[] {
  return ['Code'];                              // Lambda::Function: pre-signed URL only
  // or ['SecretString', 'GenerateSecretString']
  // or ['RedshiftDestinationConfiguration.Password']  // Firehose: write-only, AWS never returns it
}

The comparator does exact-match + entry + '.' prefix-match — listing 'Policies' skips Policies, Policies.Foo, Policies[0].PolicyDocument, etc. Pair this with a docstring explaining why the field is unreadable so a future PR can lift the gap.

Scoping a path to a SUBSET of a type's resources. Some fields are unreadable only for resources in a particular configuration, and declaring them unconditionally would switch drift detection off for the resources where AWS does return the value. The method therefore takes an optional second argument — the resource's state-recorded properties, which cdkd drift passes — so the answer can be per-resource:

getDriftUnknownPaths(resourceType: string, properties?: Record<string, unknown>): string[] {
  // AWS silently discards TlsConfig on a PUBLIC ApiGatewayV2 integration and
  // never returns it; on a VPC_LINK (private) one it does, so keep comparing.
  if (
    resourceType === 'AWS::ApiGatewayV2::Integration' &&
    properties !== undefined &&
    properties['ConnectionType'] !== 'VPC_LINK'
  ) {
    return ['TlsConfig'];
  }
  return [];
}

DynamoDBTableProvider is the other consumer, and the clearer illustration of the absent-bag rule below. It scopes on whether the template declares the property AT ALL rather than on a sibling's value — WarmThroughput is returned as unknown unless the desired bag declares it — and its predicate deliberately answers DECLARED for an absent or uninformative bag, so the path stays COMPARED:

function declaresWarmThroughput(properties?: Record<string, unknown>): boolean {
  // A bag that was never populated is NOT evidence the template declared
  // nothing, so answer 'declared' and keep comparing. A wrong DROP here is
  // unrecoverable phantom drift; the residual is a loud revert failure.
  if (!desiredBagIsInformative(properties)) return true;
  return properties !== undefined && wasSentWarmThroughput(properties['WarmThroughput']);
}

It also shows what the seam is FOR beyond a configuration split (issues #1760 / #1602): an AWS-computed value that an earlier binary froze into observedProperties would otherwise drift forever, and per-resource scoping is what retires those records without switching the key off for the type.

Two rules for this shape (issue #1602):

  • Tolerate an absent bag. Other callers may have no properties to pass, so fall back to the type-level answer rather than assuming a shape. Default to COMPARING when you cannot tell — hiding real drift is the worse failure, exactly as for getDriftUnorderedPaths() below.
  • Say so at write time. A value AWS discards is still worth a logger.warn on create / update. Declaring the drift path only removes the false positive; without the warning the user never learns the field is inert. Warn, never throw — the update path can be a state-record replay (rollback-executor.ts / drift --revert), where a refusal would leave the resource un-revertable.

getDriftUnorderedPaths() for unordered sets (strings AND objects)

The drift comparator compares arrays positionally, and AWS does not guarantee element ordering across reads. The shared normalizer (src/analyzer/drift-normalize.ts) already auto-canonicalizes two shapes for every type — {Key,...}[] tag lists and arrays whose every element is an AWS resource id (subnet-…, rtb-…) or ARN — but plain-string arrays and non-tag object arrays are deliberately left untouched, because either can be order-significant.

When your readCurrentState emits such an array that is semantically an unordered SET, declare its path so the comparator sorts it on both sides:

getDriftUnorderedPaths(resourceType: string): string[] {
  if (resourceType !== 'AWS::FSx::FileSystem') return [];
  return ['WindowsConfiguration.Aliases'];   // DNS alias names
}

Path matching uses the same shared matchesPathPrefix rule as getDriftUnknownPaths() (exact match, or entry followed by .). A plain entry is a subtree declaration — 'WindowsConfiguration.Aliases' covers that path and everything beneath it.

Append [] to declare the path ALONE (issue #1783). Because this walk descends into array elements (see the divergence note below), a subtree entry for an object list also claims every array nested inside its elements — and that is frequently wrong: AWS::DynamoDB::Table.GlobalSecondaryIndexes is an unordered set keyed by IndexName, but each element carries a KeySchema that is order-SIGNIFICANT (HASH before RANGE), so 'GlobalSecondaryIndexes' would sort the per-index key schema too and hide a real key change. 'GlobalSecondaryIndexes[]' sorts the list and stops there. The marker is understood by getDriftUnknownPaths() as well, so the two lists keep one spelling; a bare '[]' matches nothing.

Two element shapes are sorted, each only when the array is homogeneous in it: every element a plain string (sorted lexically), or every element a plain object (sorted by a key-order-independent canonical JSON — issue #1620). Key order deliberately does not participate: AWS's readback order for an object's own keys is no more guaranteed than its order for the list, so sorting on a raw JSON.stringify would reintroduce the phantom drift the pass exists to remove. A MIXED array, and a nested array inside a declared path, are both left alone — so a mis-declared path can never reorder a heterogeneous or array-valued list.

ElasticLoadBalancingV2::TargetGroup.Targets is the object-array case. It also shows the two things a readback of an unordered set has to get right that no amount of sorting fixes.

Transient lifecycle states. DescribeTargetHealth keeps reporting a just-deregistered target as draining for minutes, so including one would freeze it into the deploy-time observedProperties snapshot and produce permanent phantom drift against every later read. The provider excludes exactly draining and includes every other state — initial, unused, unhealthy, unavailable are all REGISTERED targets, and health is not registration.

Who owns the value. A target group fronting an ECS service or an ASG declares no Targets at all: the sibling resource registers them and re-registers as it scales. Comparing that list would report drift on every scale event of an untouched stack, and --revert would deregister the tasks the service just placed (the #1498 class). So the provider keeps Targets in getDriftUnknownPaths() per resource — using the properties-bag argument described above — and only drops the entry when the template actually declares Targets. An explicit Targets: [] counts as a declaration and stays compared. Reach for this scoping whenever a property can legitimately be authored by a different resource rather than by the template.

One semantic divergence from getDriftUnknownPaths(), required for this pass to work: the comparator compares arrays wholesale and never descends into elements, so an ignore-path can never cross an array. This normalizer does descend into array elements, giving each the parent's path. A path like 'Items.Aliases' is therefore meaningful for getDriftUnorderedPaths() (it reaches an Aliases array inside each Items element) while being inert as an ignore-path. The divergence is strictly more permissive.

Only declare a path AWS documents — or you can demonstrate — as order-insensitive. The failure modes are not symmetric: an undeclared unordered set produces a visible false positive the user can see and correct, but declaring an order-significant list silently hides real drift, which is the worse failure for a drift tool. FSx's SelfManagedActiveDirectoryConfiguration.DnsIps is left undeclared for exactly this reason (DNS resolver lists are conventionally preference-ordered and AWS documents no set semantics), as is ElastiCache's PreferredAvailabilityZones (documented as positionally aligned to node index).

Do NOT sort inside the reverse-mapper instead. It looks like a one-liner, but it breaks the properties fallback baseline: runDriftForStack uses observedProperties as the baseline only when present and falls back to the template properties otherwise, so for a resource deployed before observed-capture the baseline would be the user's TEMPLATE order while the read side is sorted — manufacturing drift instead of removing it. The normalizer runs on both comparison sides, which is exactly the property needed.

canonicalizeDriftProperties() for a member INSIDE an array element

The two seams above are declared as PATHS, and a path can never cross an array: calculateResourceDrift compares arrays wholesale via deepEqual and never descends into elements, so isIgnoredPath is never asked about a member of one. The only suppression an ignore-path can express there is the WHOLE array — which for GlobalSecondaryIndexes means never detecting an out-of-band index add / remove / capacity change again, permanently, to remove a one-time report. That trade is not worth making.

canonicalizeDriftProperties(resourceType, properties) (issue #1784) is the seam for exactly that shape. cdkd drift applies it to both comparison sides, last in the normalization chain, so a provider can strip the AWS-managed member from the baseline as well as from the readback — an old observedProperties record and a new readback then converge with no ignore-path and no lost detection:

canonicalizeDriftProperties(
  resourceType: string,
  properties: Record<string, unknown>
): Record<string, unknown> {
  const indexes = properties['GlobalSecondaryIndexes'];
  if (!Array.isArray(indexes)) return properties;   // identity when nothing applies
  return {
    ...properties,
    GlobalSecondaryIndexes: indexes.map(({ ItemCount, IndexArn, ...rest }) => rest),
  };
}

Four rules, each load-bearing:

  • There is deliberately no side parameter. One-sided normalization is the mistake src/analyzer/drift-normalize.ts's header already records — it manufactures drift on the properties-fallback baseline, where the baseline is the user's raw template. The same pure function must see both bags.
  • Pure, synchronous, no AWS calls, NON-mutating. It runs inside the comparison, and --revert diffs against the raw, uncanonicalized AWS bag — mutating the input would corrupt it.
  • --accept persists your output. It writes each reported change's awsValue, and those come from the canonicalized bags — so on a path that still drifts after canonicalization, --accept freezes the stripped shape into observedProperties. For the intended use (dropping a member the readback no longer reports) that is the right direction, but do not strip anything you would be unwilling to see removed from state.
  • Position in the chain matters. The hook runs after the principal / IpProtocol passes and before the tag-list / id-array / unordered-path passes, which live inside calculateResourceDrift. That order is required: an element strip must precede the unordered sort, or the two sides' canonical sort keys are computed over different member sets and diverge.
  • Return the input by identity when nothing applies, so an unaffected resource pays nothing.
  • Reach for it only for a per-array-element member. A top-level key a readback stopped emitting is getDriftUnknownPaths()'s job; a difference that is only about ORDER is getDriftUnorderedPaths()'.

canonicalizeDriftPair() when the readback's SHAPE changes

When readCurrentState starts emitting something it used to omit, every observedProperties record captured before the change lacks it, so the first cdkd drift after upgrading reports it on every untouched resource. canonicalizeDriftProperties() cannot absorb that: it sees one bag, so it cannot tell a legacy baseline from a current one, and stripping the new member from both sides would hide its drift forever.

canonicalizeDriftPair(resourceType, baseline, aws) (issue #3573) gets both bags, after the per-side hook. Key the rule on the BASELINE's shape: a legacy baseline selects the absorption, and a current-shape one is compared in full. State is not rewritten; the next deploy's capture moves the record to the current shape. The GlobalTable provider is the example: its readback gained the local (deploy-region) replica entry, and a baseline with no local entry is completed with the readback's.

Complete the baseline from the readback; do not drop from the readback. Two write paths consume these bags:

  • cdkd drift --revert passes its DESIRED bag, the recorded baseline, through the same hook against the raw readback and sends the returned baseline to update(). A member missing there is a REMOVAL to the provider. A legacy GlobalTable record sent without its local entry untagged the local table.
  • --accept writes each change's awsValue from the AWS side. Leaving that side intact means an accept stores the current shape and heals the record.

The reverse change, a readback that STOPS emitting something, is the one case where the hook drops from the baseline. AWS::DynamoDB::Table reads back only the WarmThroughput members cdkd sends (issue #3777), so a baseline block is trimmed to the members the recorded declaration sends, read from the properties argument the hook also receives. Key such a trim on the declaration, never on what the readback omitted alone: a member AWS transiently fails to report must stay reported (the ELBv2 case below reads the readback only for a key the template never declared). That is safe only because a warm-throughput member missing from the desired side is not a removal: every send site coerces the block to its usable members, and AWS cannot lower or unset warm throughput. Before copying it, check that the same holds for your member.

The ELBv2 attribute bags (LoadBalancerAttributes, TargetGroupAttributes, ListenerAttributes) are the one trim that reads the readback too (issue #4144). A baseline entry is dropped only when its key is undeclared AND missing from the readback. That is safe because no ModifyAttributes call removes a key, and the key set otherwise moves only with configuration that is compared on its own (a listener's protocol), so a key AWS stopped reporting is the service's own change, never drift. An empty readback bag is treated as a failed read and trims nothing. A declared key stays reported, as does a value change on a key both sides hold and a key only the readback holds. The last one cannot be told apart from an operator's value in one read, so it stays reported, but update() on desiredFromAwsReadback never sends it as a removal (issue #4147): the removal value would be a guess, and a rejected one fails the whole ModifyAttributes call.

It is non-mutating and returns both inputs by identity when nothing applies. It is async only so a provider can resolve the same client region its readback used; it issues no AWS call.

When there is NO observed baseline at all

The paragraph above describes the round-trip when observedProperties exists. When it does NOT — a resource deployed before observed-capture shipped — --revert falls back to the raw TEMPLATE as its desired side while the previous side is still the AWS-current snapshot. Because the revert overlays a drifted top-level subtree wholesale, every AWS-authored key inside that subtree which the template does not declare is DROPPED (issue #1478). Glue Table.Parameters is the discovered case — table_type / metadata_location are written by AWS, not by the template, so a --revert de-Icebergs the table — but the exposure is general to any resource where AWS writes into a bag the template does not fully declare.

The chosen semantic is warn and proceed: cdkd drift --revert lists the paths it will drop as part of the plan (before the confirmation prompt, and under --dry-run) and then reverts. This is a drift concern, NOT a provider one — the baseline choice affects every provider that implements readCurrentState, so fixing it inside one provider would make that provider diverge from its siblings. Nothing is required of a provider here; the note exists so that reading only the round-trip paragraph above does not leave you believing the desired side is always an observed snapshot.

Two failure modes when an always-emit placeholder round-trips through update()

cdkd drift --revert round-trips observedProperties (the snapshot readCurrentState produced) back through provider.update. That code path is what surfaces every shape-mismatch bug between the read side (readCurrentState output) and the write side (AWS create/update API input). Two failure classes have been observed; both must be designed around BEFORE adding a new readCurrentState.

Class 1 — type-discriminator-dependent fields. A field is only valid on AWS when a sibling discriminator says so. Examples: SQS DeduplicationScope / FifoThroughputLimit (FIFO-only — FifoQueue=true), SNS FifoThroughputScope (FifoTopic=true), AppSync DataSource shape (DynamoDBConfig / LambdaConfig / HttpConfig discriminated by Type). Emitting a '' placeholder for these on a discriminator-false resource means --revert pushes it back and AWS rejects with "You can specify X only when Y is set to true". Fix: guard the emit on the sibling discriminator — only emit when the discriminator is true. Pattern documented in feedback_always_emit_check_type_discriminator.md. Drift detection is not lost: the discriminator-false state cannot legally have the field on AWS, so console-side ADD is impossible. When the discriminator is N-way rather than boolean — a set of mutually-exclusive <Variant>Configuration blocks selected by a type field, as in AppSync's Type or FSx FileSystem's FileSystemType (LustreConfiguration / WindowsConfiguration / OntapConfiguration / OpenZFSConfiguration) — the same rule reads: emit EXACTLY the one block the discriminator selects, and emit it unconditionally (so the always-emit contract still holds for the one legal block, {} included), never the others. The key-set test is then written once per discriminator value, each asserting its own block is present and the rest absent — see tests/unit/provisioning/providers/fsx-filesystem-provider.test.ts.

The "console-side ADD is impossible" clause holds ONLY when the discriminator is an INDEPENDENT sibling (issue #1565). Where the "discriminator" is really the field group's OWN presence — the group is all-or-nothing, and enabling it IS setting the fields — a console-side add is not merely possible, it is the whole drift you want to catch, and guarding the emit hides it FOREVER: the comparator's top-level walk is baseline-keys-only, so a key absent from the snapshot is never compared. AWS::CloudTrail::Trail's CloudWatchLogsLogGroupArn / CloudWatchLogsRoleArn pair is that shape, and it was guarded on a rationale whose both halves later proved false: AWS was said to reject the '' round-trip (the issue #1160 live probe accepts it and nulls the field out), and a console-side enable was said to surface as both fields appearing at once on the next read (it cannot — the walk never reaches an absent key). For this shape emit the group TOGETHER and UNCONDITIONALLY, with '' placeholders, and keep the all-or-nothing invariant on the WRITE side instead — CloudTrail's update path decides both fields TOGETHER and forwards them on PRESENCE, so an explicit '' CLEARS (which is what lets drift --revert undo a console-side enable) while an ABSENT pair is retained, and a half-populated or non-string shape is refused rather than coerced. Ask which one you have before guarding: is there a sibling field whose value makes mine illegal (guard), or is my own presence the switch (always-emit)?

Class 2 — structurally-incomplete-when-empty fields. An empty-object / empty-array placeholder is structurally invalid as AWS input because a sub-field is required. Example: SQS RedrivePolicy: {} rejects with "Redrive policy does not contain mandatory attribute: maxReceiveCount" because deadLetterTargetArn and maxReceiveCount are required. Other Class 2 candidates: Lambda DeadLetterConfig (TargetArn required), Lambda VpcConfig (SubnetIds + SecurityGroupIds required), EventBridge / SNS DeadLetterConfig, ECS NetworkConfiguration (awsvpcConfiguration.subnets required), various LoggingConfiguration shapes. Fix: keep the placeholder on the read side (drift detection requires it), and sanitize at the wire layer in create() / update() by translating the empty placeholder to whatever AWS accepts as "clear this field" — usually empty string. Canonical pattern in serializeRedrivePolicy (src/provisioning/providers/sqs-queue-provider.ts):

function serializeRedrivePolicy(value: unknown): string {
  if (value === null || value === undefined) return '';
  if (typeof value === 'object' && Object.keys(value as Record<string, unknown>).length === 0) {
    return '';  // AWS-documented way to clear RedrivePolicy
  }
  return JSON.stringify(value);
}

Pattern documented in feedback_class2_placeholder_round_trip.md.

update() must gate optional fields on !== undefined, not truthy

Truthy gates (if (properties['X']) { ... }) silently drop empty string '', numeric 0, boolean false, and empty array [] (where the CFn type allows it). For cdkd deploy this is mostly invisible. For cdkd drift --revert it is a load-bearing bug: when state has no description but AWS does, the desired value to push back is Description: ''. A truthy gate drops it, the AWS update succeeds with no actual change, --revert reports ✓ reverted, and the very next cdkd drift re-detects the same drift — silent fail mode. Fix: use if (properties['X'] !== undefined) so explicit-empty values reach AWS:

// IAM Role example (PR #161 fix):
if (properties['Description'] !== undefined) {
  updateParams.Description = properties['Description'] as string;
}

Truthy gates are correct ONLY for fields where the value range excludes the falsy form (e.g. boolean flags where false means "use default", or Path where empty is invalid). Add a code comment when the truthy form is intentional. Pattern documented in feedback_update_optional_field_undefined_check.md.

Read-update round-trip test (mandatory for any provider with readCurrentState)

The above failure modes (Class 1, Class 2, truthy gate) all surface only on the cdkd drift --revert code path, which round-trips observedProperties (= a previous readCurrentState snapshot) back through provider.update. Document review and code grep cannot catch every instance — write a test that exercises the round-trip mechanically:

it('round-trip: readCurrentState placeholders survive update() without AWS-invalid inputs', async () => {
  // 1. Mock SDK to return the AWS-minimum response (only required
  //    fields, optionals undefined). readCurrentState should emit
  //    every always-emit placeholder.
  mockSend.mockResolvedValueOnce({ /* minimum SDK shape */ });
  // ...

  const observed = await provider.readCurrentState(physicalId, 'L', RESOURCE_TYPE);
  // Spot-check the placeholders are present (this is the always-emit
  // contract, see "Test convention" earlier in this section).
  expect(observed?.RedrivePolicy).toEqual({});  // Class 2 placeholder
  // ...

  // 2. Reset mocks and set up update() expectations.
  vi.clearAllMocks();
  mockSend.mockResolvedValueOnce({});  // SDK update call
  // ...

  // 3. Round-trip: pass observed as both new (desired) and old (previous).
  //    No drift → update should be a logical no-op on AWS.
  await provider.update('L', physicalId, RESOURCE_TYPE, observed!, observed!);

  // 4. Assert: no SDK call sent a value AWS would reject.
  //    Per-provider — list the AWS-rejection-shaped values you know of:
  const setAttrsCall = mockSend.mock.calls.find(
    (c) => c[0] instanceof SetQueueAttributesCommand
  );
  if (setAttrsCall) {
    const attrs = setAttrsCall[0].input.Attributes;
    if (attrs.RedrivePolicy !== undefined) {
      // Class 2: '{}' would fail "Redrive policy does not contain
      // mandatory attribute: maxReceiveCount"
      expect(attrs.RedrivePolicy).not.toBe('{}');
    }
    // ... other per-provider rejection-shape checks
  }
});

The round-trip test catches all three classes mechanically:

  • Class 1 — discriminator-false placeholders that AWS rejects when shipped: assert the relevant SetXxxAttributes / UpdateXxx call does NOT include the discriminator-only attribute when the discriminator is false in the mock setup.
  • Class 2 — structurally-incomplete placeholders: assert the AWS API call does NOT contain the empty-object / empty-array shape AWS validates and rejects (e.g. RedrivePolicy: '{}', VpcConfig: {}).
  • Truthy gate — assert that empty-string / 0 / false placeholder values DO reach the relevant AWS API call (e.g. UpdateRoleCommand input must contain Description: '' when observedProperties.Description === '').

See tests/unit/provisioning/sqs-queue-provider-update.test.ts (Class 2 round-trip), tests/unit/provisioning/iam-role-provider.test.ts (truthy-gate round-trip), and tests/unit/provisioning/sns-topic-provider-roundtrip.test.ts (Class 1 round-trip) for canonical examples.

Gone versus no read path. Return RESOURCE_NOT_FOUND (from src/types/resource.ts) when AWS itself says the resource does not exist: the service's not-found error code or fault NAME, an empty describe list for the requested id, or a terminal deleted status the service keeps listing (ECS INACTIVE, EMR TERMINATED). cdkd drift reports it as deleted and exits 1. Return undefined only for a type the provider has no read path for, or an answer that proves nothing (an unparseable id, a successful call with no body). Never return the sentinel on a message-text match alone, an access-denied or a throttle: rethrow those, or keep the old undefined, so a permissions gap is never reported as a deleted resource.

handledProperties against the CFn schema

Every SDK Provider declares a handledProperties: Map<string, ReadonlySet<string>> field naming the CFn template properties it knows how to wire to its AWS API calls. The provider registry's getProviderFor consults that set at routing time — a template carrying a property NOT in the set is auto-routed via Cloud Control API (which forwards the full property map to AWS, closing the silent-drop bug — see #614). --prefer-sdk-route Type:Prop is the per-property opt-out that forces the SDK Provider path and accepts the silent drop.

That's a runtime safety net. It doesn't help during development. A provider author who simply forgets to list a property in handledProperties AND forgets to wire it in create() / update() ships a silent bug — exactly what PR #370 (ApiGateway::Method dropped 15+ fields) demonstrated.

The structural prevention layer lives at tests/unit/provisioning/property-coverage.test.ts. It cross-references every registered provider's handledProperties against the canonical CFn schema (snapshotted to tests/fixtures/cfn-schemas/) and fails when a schema property is unaccounted for.

The four "OK" buckets

For each schema property the test classifies it into one of four buckets (in priority order):

Bucket Where declared When to use
handled provider.handledProperties.get(type) The provider's create() / update() actually wires the property to the SDK call.
by-design provider.unhandledByDesign.get(type) (with rationale string) The provider INTENTIONALLY does not wire it — separate code path, deprecated, immutable post-create, AWS API doesn't accept it, etc.
backfill tests/fixtures/cfn-schemas/_todo-backfill.json under types[<type>] Auto-generated catch-all for incremental rollout. Each entry MUST be migrated to handled or by-design eventually.
read-only readOnlyProperties in the schema fixture AWS computes the value; cdkd cannot wire it on Create/Update by definition (e.g. Arn). Automatically excluded.

A property in NONE of the above → test fails with the offending type + property list + the three actions you can take.

Adding unhandledByDesign to a provider

A clean example is AWS::ApiGatewayV2::Api's OpenAPI-import fields (Body / BodyS3Location / FailOnWarnings / DisableSchemaValidation / BasePath): they trigger an entirely separate ImportApi AWS API, not CreateApi. Listing them in handledProperties would be a lie; listing them in unhandledByDesign documents the deliberate skip:

unhandledByDesign = new Map<string, ReadonlyMap<string, string>>([
  [
    'AWS::ApiGatewayV2::Api',
    new Map([
      ['Body', 'OpenAPI/Swagger inline spec; routed through ImportApi, not the field-by-field CreateApi path.'],
      ['BodyS3Location', 'OpenAPI/Swagger spec on S3; routed through ImportApi, not the field-by-field CreateApi path.'],
      // ...
    ]),
  ],
]);

Wired into the provider class:

export class ApiGatewayV2Provider implements ResourceProvider {
  handledProperties = new Map<string, ReadonlySet<string>>([...]);
  unhandledByDesign = new Map<string, ReadonlyMap<string, string>>([...]);
  // ... rest of provider
}

Rationales are free text but should be greppable. Common shapes:

  • "create-only — AWS rejects on update"
  • "AWS-managed read-only attribute" (when not already in readOnlyProperties)
  • "deprecated — superseded by Y"
  • "tags handled via per-resource Tag API, not the create input"
  • "covered by separate AWS::Foo::Bar resource type"
  • "OpenAPI-import-only flag; meaningful only on the ImportApi code path"

NON_PROVISIONABLE types: list them in SDK_PROVIDER_NON_PROVISIONABLE_TYPES. A template property in neither handledProperties nor the allow set normally auto-routes the resource through Cloud Control (issue #614), and so does a key the schema snapshot does not know (#3713). If your provider covers a ProvisioningType: NON_PROVISIONABLE type (e.g. AWS::FSx::FileSystem, AWS::CodeBuild::Project), that route target does not exist — Cloud Control has no handlers — and the generated Tier 3 set cannot catch it: it excludes SDK-covered types by design, so isNonProvisionable() returns false once the audit is regenerated after your provider is registered. Add the type to SDK_PROVIDER_NON_PROVISIONABLE_TYPES in src/provisioning/unsupported-types.ts (measured with aws cloudformation list-types --visibility PUBLIC --type RESOURCE --provisioning-type NON_PROVISIONABLE). The ProviderRegistry then rejects a silent-drop property pre-flight with a clear error (property rationale + --prefer-sdk-route escape hatch) instead of failing at provisioning time with an opaque UnsupportedActionException, and keeps an unknown key on the SDK provider with a warning. It matters for every such type, fully handled or not, and it is per TYPE, so a provider class that also serves provisionable types (EC2Provider for AWS::EC2::NetworkAclEntry) needs nothing else. property-coverage-cc-fallback-binding.test.ts fails while a registered type is still in the Tier 3 set and missing from the list.

readonly disableCcApiFallback = true; on a provider class is the other opt-out: it covers every type the class serves, for a provider that must not fall back for a reason other than missing handlers (NestedStackProvider).

Workflow when adding a new provider

  1. Add the provider as usual — see Provider Implementation Examples.
  2. Re-export the class from src/provisioning/provider-classes.ts and register the new resource type in src/provisioning/register-providers.ts.
  3. Refresh the CFn schema fixture:
    node scripts/refresh-cfn-schemas.mjs --only-missing
    
    This fetches only the newly-registered type via cloudformation:DescribeType and writes tests/fixtures/cfn-schemas/<sanitized-type>.json. Requires AWS credentials with cloudformation:DescribeType permission.
  4. Run vp test run property-coverage — it will fail listing the schema properties your provider has not yet accounted for.
  5. For each unaccounted property, EITHER:
    • Add it to handledProperties (if create() / update() already wires it), OR
    • Add it to unhandledByDesign with a one-line rationale.
  6. Re-run the test — green.

If you really need to ship before classifying every property, you can regenerate the backfill TODO:

CDKD_GENERATE_BACKFILL=true vp test run property-coverage

This dumps every unaccounted property per type into tests/fixtures/cfn-schemas/_todo-backfill.json so the test passes. The intent is short-lived — a follow-up PR must migrate the backfill entries to handled or by-design.

Workflow when AWS publishes new properties

AWS adds properties to existing resource types fairly regularly — measured at roughly 3 writable properties per month across the whole Tier 1 surface. A property not yet in the fixture is unrecognized, and cdkd routes a resource carrying one through Cloud Control, exactly as it routes a known silent drop. So a new property reaches AWS the day AWS publishes it; the fixture refresh is what lets an SDK provider take the property over later. Why, and what is excluded: the design record.

Two mechanisms cover it. The operator's side of the first — what arrives, what to do with each class, and what to do when NOTHING arrives because the job failed or never fired — is the CFn schema refresh runbook. That page is unlisted and is linked from the pull requests the job opens, so this is the pointer that still works when there is no pull request to read.

Daily, automatically. .github/workflows/cfn-schema-refresh.yml runs vp run gen:cfn-schemas-from-zip every day and, on drift, opens a PR carrying the mechanical regeneration. A day with no drift opens nothing — the job stops at the drift check — and the open-PR guard bounds concurrency at one, so the cadence costs a short run per quiet day rather than a review cycle. It reads AWS's public schema bundle, so no AWS credentials and no CI IAM role are involved — the prerequisite that kept this manual. Only fixtures that actually changed are rewritten (generatedAt-only churn is excluded from the comparison), so the PR diff is the real drift.

That PR is allowed to land red, and the red is the hand-off rather than a bug. The job runs only the mechanical chain and hand-classifies nothing, so two classes still need you:

  1. A removed or renamed property turns a matching handledProperties / unhandledByDesign declaration into a bogus entry — retire the declaration or add a bogusTolerated rationale (see below).
  2. audit:nested-key-coverage:check reports divergences, not staleness — a new nested key on a NESTED_KEY_TARGETS type needs a provider fix or a NESTED_KEY_ALLOW_LIST entry with a rationale.

Newly unaccounted writable properties land in _todo-backfill.json; do not file an issue for any of them. .github/workflows/backfill-umbrella-sync.yml rewrites the standing backfill campaign's checklist from main's coverage map once the refresh merges — ONE umbrella issue whose body carries a generated block between <!-- backfill-types:start --> and <!-- backfill-types:end -->, one row per resource type, which gains, loses and regains rows as the map moves. The audit provenance lives outside those markers and is never rewritten. (The campaign is the open issue carrying the backfill-umbrella label — #2762 today — which superseded the closed #609; #2949 briefly fanned it into ~44 generated per-type sub-issues, since folded back into the block because a bot-filed slice is indistinguishable from an unfixed defect in a public open-issue count.) A pull request wiring a type writes Refs, and the type's row disappears once the coverage map says it is done. Types the public bundle does not carry are skipped with their fixtures left untouched and still need the authenticated path below.

On demand, by hand — to pull a refresh forward, or for a type the bundle lacks:

  1. Run vp run gen:cfn-schemas-from-zip (public bundle, no credentials), or node scripts/refresh-cfn-schemas.mjs '<AWS::Service::Type>' for one type via cloudformation:DescribeType.
  2. git diff tests/fixtures/cfn-schemas/ shows what AWS changed.
  3. The next vp test run property-coverage run fails naming the newly-unaccounted properties.
  4. Triage each: wire it through, mark unhandledByDesign, or backfill (with follow-up). Before wiring one, grep -rn '<Property>' tests/integration/*/verify.sh tests/integration/*/lib — a fixture that keys its Cloud Control route on that property loses its premise the moment it is handled, and has to be re-seeded in the same PR (integ-fixture-conventions.md).

There is deliberately no CI staleness check on main. It would go red whenever AWS publishes a property — noise on a schedule nobody controls, the same reasoning gen:aws-cli-removals carries in vite.config.ts. A scheduled job whose red is confined to its own PR is the shape that argument leaves open.

What stays on the SDK route. An unrecognized property does not route in four cases. The first three produce a deploy-time warning that names the property, says it will not reach AWS, and gives the reason; the fourth is silent:

Case Why it stays
The property is read-only CloudFormation ignores a read-only property in a template too.
The type has no usable Cloud Control route Its provider declares disableCcApiFallback, the type is NON_PROVISIONABLE, or its Cloud Control handler cannot manage it (a 'cc-broken' sticky-CC exemption).
The resource was deployed on the SDK route with the property, and its value is unchanged cdkd keeps an existing resource on its route. Changing the value routes it.
--prefer-sdk-route <Type>:<Prop> names it You chose the drop; the warning is suppressed.

A misspelled property on a new resource, or one newly added, is routed like any other and Cloud Control rejects it with Unsupported property, as CloudFormation does.

"Bogus" entries and the tolerance list

A property in handledProperties (or unhandledByDesign) that is NOT in the CFn schema is a "bogus" entry — most often an SDK input field name that diverges from the CFn property name (e.g. SDK DefaultCooldown vs CFn Cooldown on AutoScalingGroup), a typo (PlacementStrategy vs PlacementStrategies on ECS::Service), or a stale alias from before AWS renamed the property.

The test reports these — but fixing each requires per-provider investigation that often touches the safety-net's runtime behavior. As a stopgap, the tolerance list at tests/fixtures/cfn-schemas/_todo-backfill.json under bogusTolerated[<type>][<prop>] accepts a one-line rationale per entry, the test stays green, and follow-up PRs investigate one at a time. Day-1 of issue #391 the test surfaced 10 such entries — see the rationale strings in that file for the canonical examples.

An entry is a claim about a DECLARATION — "the provider declares <prop>, the schema no longer has it, and here is why the declaration stays" — and every reader keys on that premise: the coverage test consults it only while walking handledProperties / unhandledByDesign / the backfill list, and the schema-refresh diagnosis only for a property in the generated handled map. So when the declaration itself is renamed or removed (the usual fix), retire the entry in the same change: the test fails an entry that names a property no declaration on that type carries, because such an entry is inert while reading as protection — and if the declaration was meant to exist, it is masking a silent drop (issue #3034, where a NetworkAclEntry.IcmpTypeCode entry written on 2026-05-16 outlived the 2026-05-27 rename to Icmp until 2026-09-13). The other staleness direction — AWS re-adds the property — is reported by its own case.

Admitting a type to the sticky-CC exemption

Closing the last silent-drop gap for a type is not the end of the story for resources ALREADY deployed. A resource recorded provisionedBy: 'cc-api' stays on Cloud Control by default, so the backfill you just landed speeds up new resources only — every existing one keeps paying Cloud Control latency until the type is admitted to STICKY_CC_MIGRATION_EXEMPT (src/provisioning/provider-registry.ts, issue #2719).

Consider admitting a type in the same PR that empties its silentDrop map, or in a small follow-up. It is not mandatory and nothing blocks a release without it: an un-admitted type is merely slower, which is the status quo — unless Cloud Control also mishandles an update for it, in which case admission is the fix (issue #4679).

The bar is physicalId parity, and it is EVIDENCE, not an argument. Cloud Control mints its identifier from the schema's primaryIdentifier; the SDK provider stores whatever its create() returned. When those differ, a flip leaves cdkd addressing the resource by a string AWS does not recognise. They often agree — but "often" is why this needs observing rather than reasoning.

To admit a type:

  1. Confirm the type's silentDrop map is empty in property-coverage.generated.ts. A type with a real drop must not be admitted; the per-resource condition would refuse every resource anyway, making the entry dead weight. And keep it empty: the daily schema refresh re-adds a drop the moment AWS publishes a property the provider does not write (AWS::SNS::Topic gained MaximumMessageSize that way on 2026-09-18, issue #3413), so a refresh landing a drop on an admitted type owes the wiring in the same cycle — otherwise cdkd diff renders sticky for a resource that cdkd deploy --allow-unsupported-properties would flip.
  2. Read the CFn schema's primaryIdentifier and the provider's create() return, and write what BOTH store into physicalIdForm — in words a reviewer can check, not "verified".
  3. Write or extend an integ fixture that OBSERVES the flip on a live resource: deploy, force the resource onto Cloud Control with --recreate-via-cc-api, redeploy with a real property change, then assert the physical id is unchanged AND the record flipped to 'sdk'. tests/integration/cc-to-sdk-reroute/ is the model. Comparing physical ids is not always sufficient — a resource with a user-supplied name keeps its id through a destroy + recreate, so that fixture also attaches an out-of-band subscription cdkd does not manage and asserts it survives. Without such a witness the arm cannot tell an in-place update from a replacement, which is the whole claim.
  4. Add the entry naming that fixture. tests/unit/provisioning/sticky-exempt-registry.test.ts refuses an entry whose fixture does not exist or has no row in the integ ledger. A fixture that ran before your arm was added still passes that check, so the arm's own run (step 5) is what the PR must wait for.
  5. Run the fixture and commit the ledger row.

Use mode: 'cc-broken' only when Cloud Control genuinely cannot manage the type (its escape is unconditional and ignores --pin-cc-api). A type that works on Cloud Control and is only slower is 'sdk-coverage', and so is one whose Cloud Control defect only a mutating deploy reaches, since that deploy is the one that flips the record (AWS::ElasticLoadBalancingV2::Listener: Cloud Control leaves a removed ListenerAttributes key at its old value, issue #4679).

Logging

  • info: Successful operations
  • debug: Detailed information
  • warn: Non-fatal errors
  • error: Fatal errors
this.logger.info(`Creating ${resourceType} ${logicalId}`);
this.logger.debug(`Using properties:`, properties);
this.logger.warn(`Old resource deletion failed: ${String(error)}`);
this.logger.error(`Failed to create ${logicalId}:`, error);

Resource name constraints

AWS services have length and character constraints on names:

// IAM Role example (64 character limit)
private shortenRoleName(roleName: string): string {
  const MAX_LENGTH = 64;

  if (roleName.length <= MAX_LENGTH) {
    return roleName;
  }

  const hash = Buffer.from(roleName)
    .toString('base64')
    .replace(/[^a-zA-Z0-9]/g, '')
    .substring(0, 8);

  const maxPrefixLength = MAX_LENGTH - hash.length - 1;
  const prefix = roleName.substring(0, maxPrefixLength);

  return `${prefix}-${hash}`;
}

Last updated: