Skip to main content
Comfy API runs workflows on managed, autoscaling GPU endpoints. A Build describes the ComfyUI environment, including its models, custom nodes, and settings. The CLI creates releases from a Build and deploys them as endpoints.

Comfy API Quickstart

Start with an agent or choose another way to provide your environment or workflow.
For Desktop snapshots and workflow JSON imports, see Build Sources.

The Build file

comfy-build.yaml stores the Build definition and last known remote state. The CLI uses that state to detect when the remote Build has changed. Keep the file with the project. It describes the environment but does not contain model files. Check how the local Build compares with the install and remote version before pushing changes:

Update and release

After changing the local install, refresh the Build definition:
To create another release from an existing Build, check the supported targets and select one:
Follow one release’s build log with:

Regions and GPU availability

Region capacity changes, so check the platform catalog when choosing a deployment target. Filter the results with --region <region>:
The catalog is the source of truth for which GPU classes are available in each region at deployment time.

Workspace limits

See Limits for your workspace’s usage and plan limits. Keep --max within the worker limit per deployment.

Monitor and call a deployment

The --min and --max options set the worker bounds. Run comfy deploy status --watch to follow deployment health, release freshness, and serving activity. You can also call the endpoint with the Comfy SDKs. Set COMFY_BASE_URL to the deployment URL and provide an API key. See Choosing a base URL.

Operate a deployment

Inspect and clean up

Deleting a deployment and deleting a Build are separate irreversible operations. Confirm the target before using comfy deploy delete --yes or comfy build delete --id bld_123456 --yes.

FAQ

Yes. In the Builder, upload one or more workflows to preselect their models and custom nodes. Include the dependencies needed by all the workflows you plan to run, then submit each API-format workflow to the deployed endpoint.Each deployment uses one GPU type. To run workflows on different GPU types, create separate deployments of the same Build.
No. Builds and their releases are stored on your account at no charge. You are not billed for the storage footprint of a Build until you deploy it.Storage billing starts with a deployment:
  • When you deploy a release, its models are staged onto network storage shared by the deployment’s workers. The storage is read-only after deployment and billed per GB-month for as long as any deployment of the Build exists in that region, including while a deployment is paused with comfy deploy stop.
  • Each worker also gets a fixed 50 GB container disk. It is ephemeral and is only billed while the worker is running, as part of the worker’s compute cost.
  • Deleting a deployment releases its compute. Staged network storage is cleaned up shortly after the last deployment using it in that region is deleted, and billing for it ends.
For current storage rates, see the Comfy pricing page. The compute catalog (comfy deploy refs compute) and the deploy dialog also show the rates that apply to a deployment.
--min sets the number of active workers: workers that stay running at all times so requests never wait for a cold start. An active worker is billed per second for the entire time it is running, whether or not it is processing jobs.Workers above --min, up to --max, are flex workers. A flex worker is billed per second from the moment it starts (including startup and model loading), through job processing, plus a short idle window (currently 30 seconds) before it scales back down. When flex workers are scaled down, they cost nothing. With --min 0 the whole deployment scales to zero and bills no compute while idle, at the cost of a cold start on the first request.Billing meters actual per-second worker usage, multiplied by the number of workers running. If your workspace runs out of credits, deployments are stopped automatically.For current per-worker GPU rates, see the Comfy pricing page. Rates are quoted per worker-hour and billed per second. GPU availability per region comes from the compute catalog: run comfy deploy refs compute for the current list.
A deployment automatically distributes requests across its workers and scales within the configured --min and --max bounds. Requests that cannot run immediately are queued and processed as worker capacity becomes available.
Follow the deployment state, then inspect its logs and events for the cause:
If the selected GPU or region has no capacity, retry later or run comfy deploy refs compute and choose an available region and GPU pair. Include the Build, release, and deployment IDs when asking for help.
Deleting a deployment removes its endpoint and releases its compute. It does not delete the Build or its releases. After you delete the last deployment using that Build in a region, its staged network storage is cleaned up and storage billing ends shortly afterward.Deleting the Build is a separate operation: comfy build delete --id <build-id> --yes.

Next steps

Comfy SDKs

Use from a Python or TypeScript app to submit workflows and handle inputs, progress, and outputs.

Comfy API v2

Use from any language when calling the HTTP API directly or checking its request and response formats.

Workflow API Format

Use when preparing a workflow JSON for an API request or troubleshooting its inputs and outputs.

Comfy CLI Reference

Use to automate Build, release, deployment, and workflow commands from a terminal.