Slow deploys, throttling and retries
cdkd deploys independent resources in parallel and retries the errors that usually clear by themselves. This page explains what limits a deploy's speed, what cdkd retries and for how long, and how to read the line it prints when it stops retrying.
On this page:
- Deployment is Slow
- Cloud Control API Rate Limit
- A name is still held by the resource you just deleted
- "cdkd stopped waiting for it" — a network outage during a Cloud Control operation
Deployment is Slow
A deploy is slower than it should be when a handful of independent resources takes well over a minute, or when adding resources that do not depend on each other lengthens the deploy roughly in proportion. Benchmarks has the figures to compare against.
There are four usual causes:
| Cause | How to check |
|---|---|
| A long dependency chain | Inspect the dependency graph |
| A resource on the slower Cloud Control route | Resources on the Cloud Control route |
| Cloud Control throttling | Rate exceeded under --verbose; see Cloud Control API Rate Limit |
| Asset publishing | Slow asset steps before provisioning starts; see the asset flags under Cloud Control API Rate Limit |
Inspect the dependency graph
A dry-run deploy with --verbose builds the graph and stops before touching
AWS:
cdkd deploy MyStack --dry-run --verbose
... DEBUG [DagBuilder] Dependency graph built: 4 nodes, 3 edges
... DEBUG [DagBuilder] Level 0: 2 resources - Bucket, Table
... DEBUG [DagBuilder] Level 1: 1 resources - Role
... DEBUG [DagBuilder] Level 2: 1 resources - Function
... DEBUG [DagBuilder] Execution levels computed: 3 levels
✓ Dry run completed - no actual changes made
A level is a group of resources that do not depend on each other. Resources
in a chain cannot run at the same time, so a graph with many levels and few
resources per level is what slows a deploy down. A real deploy prints the same
figure in one line:
Deploying 4 resource(s) (DAG: 3 levels, max parallel: 10).
During a deploy each resource starts as soon as its own dependencies complete, so cdkd does not wait for a whole level to finish. A destroy deletes level by level, in reverse order.
Remove dependencies the template does not need
cdkd works out the order from the Ref and Fn::GetAtt references in the
template. An explicit addDependency between resources that do not use each
other adds nothing except a wait:
const bucket = new s3.Bucket(this, 'Bucket');
const role = new iam.Role(this, 'Role', {
assumedBy: new iam.ServicePrincipal('lambda.amazonaws.com'),
});
// delete this line: the role does not use the bucket
role.node.addDependency(bucket);
Resources on the Cloud Control route
cdkd state show MyStack
MyTopic
Type: AWS::SNS::Topic
PhysicalID: arn:aws:sns:us-east-1:123456789012:MyTopic
ProvisionedBy: cc-api
ProvisionedBy: cc-api means cdkd manages the resource through the Cloud
Control API, which is slower than one of cdkd's own SDK providers. cdkd sends
a resource there when its template carries a top-level property the SDK
provider does not send, and the resource stays there on later deploys.
Whether you need to act depends on the type:
- For some types cdkd moves the resource back by itself. Once cdkd covers
every property the resource uses, the next deploy that changes the resource
moves it back in place. See
cdkd diffshows[returning to SDK provider]. - For any other type, move it yourself with
--recreate-via-sdk-provider <LogicalId>. This flag destroys and recreates the resource.
--recreate-via-sdk-provider refuses while the template still carries the
property that sent the resource to Cloud Control. Remove the property, or
pass --prefer-sdk-route <Type>:<Prop> to accept that the property is
dropped.
Provisioning Layers explains the two layers, and
cdkd deploy: safety & compatibility flags the flags.
Cloud Control API Rate Limit
ProvisioningError: Failed to create resource MyTopic
Caused by: CREATE failed for MyTopic: Rate exceeded
The Cloud Control API throttled the request (ThrottlingException). AWS
publishes no rate quota for Cloud Control, but it throttles when requests
arrive too fast, and a stack with many independent resources issues many
calls at once.
cdkd retries a throttled request by itself, so you normally see this message
only under --verbose or after the retries are used up. If throttling fails
the deploy, lower the concurrency:
# default 10, concurrent resource operations
cdkd deploy MyStack --concurrency 4
# default 4, concurrent stacks
cdkd deploy --all --stack-concurrency 2
Asset publishing has its own limits: --asset-publish-concurrency (default 8)
and --image-build-concurrency (default 4).
What cdkd retries
cdkd retries create, update and delete operations during cdkd deploy, and
the replays cdkd rollback and cdkd drift --revert perform. The schedule
depends on the kind of error:
| Kind of error | Examples | Retries | Total wait |
|---|---|---|---|
| Throttling and transient | Rate limits, a resource still leaving Pending, HTTP 500 / 502 / 503 / 504 |
8, backing off 1s to 8s | 47s |
| IAM propagation | cannot be assumed, not authorized to perform, Invalid IAM Instance Profile, not authorized to access the Log Destination |
26, every 0.25s to 2s | 47.75s |
| A name still held by a deleted resource | QueueDeletedRecently, Step Functions StateMachineDeleting |
8, backing off 2s to 10s | 64s |
Two kinds of resource are outside this retry: custom resources (Custom::*
and AWS::CloudFormation::CustomResource) and nested stacks
(AWS::CloudFormation::Stack).
A custom resource has its own retry. When your handler returns FAILED with a
reason that looks like a missing permission, cdkd invokes the handler again.
CDKD_CR_AUTHZ_MAX_RETRIES sets how many times (default 2, at most 10), and
0 turns this off.
Reading the "gave up after" line
When the retries run out, cdkd prints one line at the default log level:
MyFunction: gave up after 26 IAM-propagation retries over 47.75s of propagation backoff (the full propagation budget) - The role defined for the function cannot be assumed by Lambda.
When cdkd can see how the AWS SDK parsed the failing response, the line ends with a bracketed summary:
MyQueuePolicy: gave up after 5 IAM-propagation retries over 5.75s of propagation backoff - Failed to create SQS queue policy MyQueuePolicy: UnknownError [name=InternalFailure http=500 requestId=ebf581cc-6072-5ffc-943a-e33312488615]
What the line tells you:
- It ends with
(the full propagation budget). cdkd retried to the end of its schedule. Usually IAM took longer than 47.75s to propagate. Re-run the deploy. If it happens again, open an issue and include the line. - It ends the same way, but AWS also uses the message for permanent problems. Then the cause is a real misconfiguration, such as the Step Functions log-destination error. Fix what the AWS message names. Re-running never succeeds.
- It has no budget note and a low retry count. Something ended the retries
early: an explicit deny, or an error cdkd could not classify. Read the
bracket. Report it if the status is a 5xx other than 500 / 502 / 503 / 504,
or if the bracket shows
no-$metadata. - It says
transient server-error retries (HTTP 5xx). The retries were spent on transient server errors and not on IAM propagation. Re-run. - There is no such line at all. The retry never started. Either cdkd did not recognise the error, or the resource is a custom resource or a nested stack. Report the unrecognised wording, unless the resource is one of those two types.
Details of the line:
UnknownErroris the AWS SDK's placeholder for a response with no message text. The bracket is then the whole diagnosis:nameandhttpare what cdkd classifies on, andrequestIdis what AWS support needs.no-$metadatain place ofhttp=means the failure never reached the SDK's error parsing: a network or protocol failure.- The seconds count propagation backoff only. A throttle in between uses an attempt without counting, so the retry count on a full-budget line can be below 26.
--verboseprints each attempt with its running total (attempt 15/26, 25.75s backoff through this attempt).- The line can appear during
cdkd rollbackand an automatic rollback as well as duringcdkd deploy.
Troubleshooting internals has the exact schedules and the custom-resource retry rules.
A name is still held by the resource you just deleted
A destroy followed straight away by a deploy fails on a name that looks free:
ProvisioningError: Failed to create resource MyQueue
Caused by: You must wait 60 seconds after deleting a queue before you can create another with the same name.
Some services keep a name reserved for a while after the delete returns. cdkd retries the create for 64 seconds, so usually you never see this error.
| Type | Window |
|---|---|
AWS::SQS::Queue |
60s, stated by the service |
AWS::StepFunctions::StateMachine |
About 23s of status: DELETING on an idle machine |
AWS::S3::Bucket |
Bucket names are global and released asynchronously |
Wait and re-run. When the retries run out, the line says so:
MyQueue: gave up after 8 name-cooldown retries over 64.00s waiting for the name to be released (the full name-cooldown budget) - <the AWS message>
You can hit the Step Functions case without having declared a state machine.
Any CDK app that uses custom_resources.Provider with an isCompleteHandler
contains one.
Name errors that waiting does not fix
cdkd does not retry these, so do not wait for them:
- ELBv2
DuplicateLoadBalancerNameand DynamoDB's create-side refusal, which mean the resource still exists. - A Secrets Manager name scheduled for deletion, which can be held for 7 to 30 days.
During a --replace
During a --replace the retry runs once per replacement attempt, so a
custom-named SQS queue can appear to hang for up to about 11 minutes.
"cdkd stopped waiting for it" — a network outage during a Cloud Control operation
CREATE of Dbwriter9B286E50 was accepted by Cloud Control API, but cdkd stopped
waiting for it (cdkd could not reach Cloud Control API for 121s: connect
ECONNREFUSED 100.72.0.178:443). The operation may still be running in AWS, and
cdkd has NO state record for it, so any resource it creates is untracked by
rollback and by cdkd destroy. Check what it did with:
aws cloudcontrol get-resource-request-status --request-token <token> --region ap-northeast-1
Your connection dropped while a Cloud Control operation was running. A VPN that reconnects is the usual cause.
Cloud Control runs an operation inside AWS and gives cdkd a request token. cdkd asks for the result with that token until the operation finishes, and the operation continues whether or not cdkd can reach AWS. cdkd keeps trying for about two minutes without a connection. After that it gives up, and it does not know how the operation ended.
Run the command the error prints and read OperationStatus:
OperationStatus |
Meaning | Do |
|---|---|---|
SUCCESS |
AWS did the operation; for a create, cdkd has no record of the resource | Adopt it with cdkd import using the response's Identifier, or delete it in AWS |
FAILED |
Nothing was created | Re-run cdkd deploy |
IN_PROGRESS |
Still running | Run the status command again until it settles |
A RequestTokenNotFoundException means the request has aged out of Cloud
Control's status history. Look for the resource in the console or through the
service's API, then import or delete it.
After an update or a delete
Only a create can leave a resource without a state record. After an update or
a delete the resource still has its record, so re-running cdkd deploy or
cdkd destroy finishes the job once the network is back.
If the same deploy also reported Failed to save partial state before rollback or Failed to write rollback journal, the outage took those too. See
After a failed or interrupted deploy.
Related
- Troubleshooting: the common problems and the list of every troubleshooting page