x-litellm-response-cost response header, for the models that report a Comfy credit price per run (see Runs that don’t report a cost). LiteLLM reads that header on every pass-through response and records it as the request’s spend, so you don’t maintain a price table for those models.
LiteLLM’s built-in
fal_ai/ provider covers image generation only. To route a general workload through LiteLLM, use a pass-through endpoint as shown here, not a native provider.Configure the pass-through
1
Create a Comfy API key
Create a key in your Comfy workspace and make it available to the LiteLLM Proxy process:
2
Add Router to config.yaml
config.yaml
include_subpath: trueforwards/comfy/v2/models/{provider}/{model}tohttps://api.comfy.org/v2/models/{provider}/{model}, so one pair of entries covers every Router model.- Keep
pathandtargetscoped to/v2/models. Pointing the pass-through at the API root would let any caller with a LiteLLM key use your Comfy key on every other Comfy API route, outside LiteLLM’s spend tracking. methodssplits the two entries on the same path. It needs a LiteLLM Proxy release whose pass-through endpoints supportmethods; this page was checked against v1.103.0.X-API-Keycarries yourcomfyui-key. Router also accepts the same key asAuthorization: Bearer comfyui-....- Leave
forward_headersoff. When it’s on, LiteLLM forwards every incoming header to Router, including the caller’s LiteLLM key. With it off, callers pass anIdempotency-Keyasx-pass-Idempotency-Key: LiteLLM strips thex-pass-prefix and forwards it. A plainIdempotency-Keyheader doesn’t reach Router, so a retry would run and bill a second generation. - Set
timeoutabove Router’s run deadline, so Router can return its own504and request ID before LiteLLM gives up. The deadline is 10 minutes by default, the same as LiteLLM’s default pass-through timeout of 600 seconds.
3
Call Router through the proxy
Callers authenticate with their LiteLLM key. The body is the model’s own input, exactly as in the Comfy Router reference:
How spend is recorded
A synchronous run that reports a cost carries these response headers:X-Comfy-Credits-Used is the run’s price in Comfy credits, rounded to two decimal places. X-Litellm-Response-Cost is the same price converted to USD from the unrounded credit amount, so dividing the displayed credits by the conversion rate can differ from it in the last digits. Header names are case-insensitive, so LiteLLM reads it as x-litellm-response-cost. LiteLLM records it exactly as sent and uses it instead of your cost_per_request. x-litellm-total-tokens is always 0, because media runs aren’t metered in tokens.
As with X-Comfy-Credits-Used, the header reports a price rather than a settled charge. Reconcile against workspace billing, not against LiteLLM’s totals.
Runs that don’t report a cost
Some runs don’t carryx-litellm-response-cost. It’s absent in the same cases as X-Comfy-Credits-Used:
- a run billed to your own provider key rather than to your Comfy credits
- a model that isn’t yet on Router’s list of models that report a per-run credit price. That list is conservative, and some providers aren’t on it at all
- a charge that didn’t reach Comfy billing
- an error response
Idempotency-Key (sent through the proxy as x-pass-Idempotency-Key). That retry isn’t charged again, so the replay doesn’t report the run’s exact cost a second time.
When the header is absent, LiteLLM falls back to the endpoint’s cost_per_request. That applies to every case above. Error responses, runs billed to your own provider key, and idempotent replays are therefore recorded at your estimate even though Comfy didn’t charge your credits for them. Set it to a sensible estimate of your typical run rather than 0, so a run that doesn’t report a cost doesn’t appear free in your spend reports, and expect those rows to push LiteLLM’s totals above your Comfy bill. A run that was priced and genuinely cost nothing reports x-litellm-response-cost: 0, which LiteLLM records as zero.
Queued delivery
With queued delivery, one generation reaches LiteLLM as several requests, and each is its own spend row:
The submit response’s
status_url, response_url and cancel_url are absolute https://api.comfy.org/... URLs, and LiteLLM doesn’t rewrite response bodies. Call them through the proxy by replacing https://api.comfy.org with your proxy’s /comfy prefix, for example http://localhost:4000/comfy/v2/models/.... Sent straight to Comfy, your LiteLLM key is rejected.
Queue responses don’t carry x-litellm-response-cost yet, so a queued generation is recorded at your estimate, not at its exact price. Splitting the endpoint by method books that estimate once, on the submit, instead of once per poll. For exact per-run spend in LiteLLM, use the synchronous route.
Notes
- LiteLLM records spend but doesn’t change billing. Comfy bills your workspace exactly as it would without the proxy.
- Full request and response schemas for every model are in the Comfy Router reference. Deadlines, retries, and other limits are in Capabilities and limits.