fal.ai Models Pricing Explained: API & GPU Costs
fal.ai pricingfal.aifal gpu pricingseedance pricingai api pricing

fal.ai Models Pricing Explained: API & GPU Costs

2026-07-23

TL;DR: fal.ai pricing splits into two products: hosted Model APIs (usually billed per successful output — image, megapixel, video second, video, request or token) and Serverless / fal GPU (billed over total runner lifetime). The biggest mistake is calling all fal usage "GPU-second billing." For Seedance 2.0, fal lists $14.00 per million tokens vs LinkModel's $6.30; for GPT Image 2 token categories LinkModel is 25% lower. fal GPU list rates run from $2.99/hr (RTX PRO 6000) to $8.50/hr (B300). The full fal.ai review covers model access and concurrency.

fal.ai pricing becomes clear once you separate two products. The broader fal.ai review covers model access, concurrency and production fit:

  • fal Model APIs: call a hosted model and usually pay for successful output.
  • fal Serverless / fal GPU: deploy custom code and pay for runner lifetime on selected hardware.

The biggest pricing mistake is to describe all fal.ai usage as GPU-second billing. That is not what the current documentation says. Many hosted Model APIs are billed per image, megapixel, generated video second, video, request or token. GPU lifecycle billing belongs to the separate Serverless path.

fal AI Model Pricing: How Hosted Model APIs Are Billed

Model API outputCommon fal billing unit
ImagePer image or megapixel
VideoPer generated second or complete video
Text/selected multimodal modelsTokens
Other hosted endpointsPer request or compute time when no fixed rate exists
Server error/queue timefal says these are not charged

fal states that Model APIs charge for successful outputs. Server errors and queue time are not charged, and Model API cold starts are not charged. A 422 client error may still be billable if compute occurred before the request was rejected.

That is different from fal GPU billing, which appears later in this article.

Seedance Pricing on fal.ai: Native Tokens and Displayed Seconds

fal’s Seedance 2.0 model page publishes:

  • $0.014 per 1,000 tokens
  • equivalent native rate: $14.00 per million tokens
  • page display around $0.3034/sec at 720p
  • page display around $0.682/sec at 1080p

The official token formula shown by fal is:

Video tokens = height × width × duration × 24 ÷ 1,024

720p calculation

Assume 1280×720, 24 fps and one second:

720 × 1,280 × 1 × 24 ÷ 1,024 = 21,600 tokens/sec

21,600 ÷ 1,000,000 × $14.00 = $0.3024/sec

fal’s page displays $0.3034/sec. Under the stated 1280×720 assumption, that display is not mathematically identical to the page’s own $14.00/M token rate and formula. For an auditable comparison, preserve both numbers: treat $0.3034 as fal’s displayed estimate and $0.3024 as the calculation from its published native token price. Do not reverse the displayed estimate into a new “official” token quote.

1080p calculation

Assume 1920×1080, 24 fps and one second:

1,080 × 1,920 × 1 × 24 ÷ 1,024 = 48,600 tokens/sec

48,600 ÷ 1,000,000 × $14.00 = $0.6804/sec

fal displays about $0.682/sec. Under the stated 1920×1080 assumption, the published formula produces $0.6804/sec. Again, keep the displayed estimate and calculated value separate; the native token rate is the stable comparison unit.

fal.ai vs LinkModel Seedance Pricing

BytePlus lists the official no-video-input Seedance rates at $7.00/M for 720p and $7.70/M for 1080p. LinkModel lists $6.30/M and $6.93/M; fal lists $14.00/M.

Native million-token prices:

Seedance workloadBytePlus officialfal nativeLinkModel
720p, no video input$7.00/M$14.00/M$6.30/M
1080p, no video input$7.70/M$14.00/M$6.93/M

Converted per-second estimates:

Seedance workloadOfficial calculation*fal calculation*LinkModel calculation*
720p, no video input$0.15120$0.30240$0.13608
1080p, no video input$0.37422$0.68040$0.336798

Sources: BytePlus pricing, fal Seedance page, and LinkModel models. Derived from the published token rates at 24 fps; these are calculations, not independent provider quotes.

Cost for 1,000 generated seconds

WorkloadBytePlus official calculationfal native-token calculationLinkModel calculation
1,000 sec at 720p$151.20$302.40$136.08
1,000 sec at 1080p$374.22$680.40$336.798

The per-second view is useful for planning. It becomes a weak comparison when the token formula disappears and only the rounded second remains. LinkModel keeps the official-style million-token row visible, so the discount or markup is easier to see.

fal.ai Models Pricing for GPT Image 2 Tokens

fal’s GPT Image 2 page publishes the upstream token categories:

GPT Image 2 categoryOpenAI official / 1M tokensfal.ai / 1M tokensLinkModel / 1M tokens
Text input$5.00$5.00$3.75
Cached text input$1.25$1.25$0.9375
Image input$8.00$8.00$6.00
Cached image input$2.00$2.00$1.50
Image output$30.00$30.00$22.50

Sources: OpenAI API pricing, fal image model page, and the LinkModel models; verified July 22, 2026.

fal also publishes canonical per-image examples because quality and size change the number of image tokens. For example, its 1024×1024 table lists $0.006 low, $0.053 medium and $0.211 high.

Per-image presets are convenient. Token categories are better when your application mixes text, cached context, image input and output. LinkModel’s rate is 25% lower in every same token category listed above.

fal.ai Image Pricing by Number of Images

fal prices Gemini 3.1 Flash Image at $0.08 for 1K, with 0.75× for 512, 1.5× for 2K and 2× for 4K. LinkModel publishes direct call prices for 0.5K, 1K, 2K and 4K.

Per-image prices:

ResolutionGoogle official / imagefal.ai / imageLinkModel / image
0.5K$0.04500$0.06000$0.03375
1K$0.06700$0.08000$0.05025
2K$0.10100$0.12000$0.07575
4K$0.15100$0.16000$0.11325

Cost for 10,000 images:

ResolutionGoogle officialfal.aiLinkModel
0.5K$450.00$600.00$337.50
1K$670.00$800.00$502.50
2K$1,010.00$1,200.00$757.50
4K$1,510.00$1,600.00$1,132.50

Sources: Google API pricing, fal Gemini image page, and the LinkModel models. fal adds $0.015 when web search is used; that optional feature is excluded from the table.

fal GPU Pricing: List Rate vs “As Low As”

fal’s pricing page publishes both list and lower “as low as” rates.

fal GPUMemoryList price/hour“As low as”/hour
B300288 GB$8.50$4.49
B200180 GB$6.25$3.49
H200141 GB$4.50$2.10
H10080 GB$3.99$1.89
RTX PRO 600096 GB$2.99$1.10

Source: fal.ai pricing, verified July 22, 2026.

“As low as” is not the same as a universal on-demand rate. Use list price for conservative budgeting until your workload qualifies for the lower number.

How fal Serverless GPU Cost Is Calculated

fal Serverless documentation describes billing over total runner lifetime. Depending on configuration, that can include:

  • setup;
  • idle time;
  • active processing;
  • draining;
  • teardown.

Pending tasks and image pulls are described separately from active runner billing. Settings such as keep_alive, minimum concurrency and the number of GPUs can materially change cost.

Basic fal GPU formula

Estimated compute cost = GPU hourly rate × billable runner seconds ÷ 3,600 × GPU count

Example using the H100 list rate and one GPU for a 20-minute billable runner lifetime:

$3.99 × 1,200 ÷ 3,600 × 1 = $1.33

At the published “as low as” H100 rate:

$1.89 × 1,200 ÷ 3,600 = $0.63

These are pure rate calculations, not guarantees of availability, performance or eligibility. Your real job cost also depends on startup, idle capacity, throughput and retries.

When fal GPU Is Cheaper—and When It Is Not

fal GPU can be more economical when:

  • utilization is consistently high;
  • you batch requests efficiently;
  • custom code or model control removes other tooling;
  • a fixed hosted output price includes a margin you can avoid;
  • the lower GPU rate is actually available.

A hosted Model API can be cheaper when:

  • traffic is sporadic;
  • cold and idle capacity would dominate runner time;
  • you value a fixed output unit;
  • engineering time to package, observe and optimize the model is significant.

Do not compare “$1.89/hour H100” directly with “$0.08/image” without measuring images per billable runner hour at the same quality.

fal Credits, Concurrency and Failed Requests

  • Purchased credits expire after 365 days.
  • New Model API accounts start with 2 concurrent requests.
  • Limits can rise with paid invoices over the last four weeks, self-service up to 40.
  • Requests above the limit queue.
  • Server errors and queue time are not charged.
  • Some 422 requests can be charged if GPU work occurred before validation.

These rules belong in your cost model because throughput and error handling affect how quickly credits are consumed.

fal.ai Models Pricing FAQs

Is fal.ai priced per second?

Some video Model APIs are priced per generated second, but not every fal product uses that unit. Seedance also publishes a native token price, and Serverless uses GPU runner lifetime.

Does fal.ai price every model by GPU time?

No. Hosted Model APIs usually use an output unit. GPU time is central to the separate Serverless product and some endpoints without predefined output pricing.

What does fal GPU “as low as” mean?

It is the lowest displayed rate, not necessarily the rate for every account or job. Budget with list price until eligibility is confirmed.

Is fal.ai cheaper than LinkModel?

It varies by model. In the verified Seedance, GPT Image 2 token-category and Gemini 3.1 Flash Image comparisons above, LinkModel is lower. The fal.ai alternatives guide compares other providers for teams considering a switch.

How do I estimate 10,000 fal image generations?

Multiply the exact per-image rate by 10,000, then add optional features such as web search and account-specific discounts. Do not use a 1K price for 4K output.

Can fal charge a failed request?

fal says 500-series errors are not charged. A 422 client error may be billable when GPU work occurred before the error was detected.

Final fal Pricing Takeaway

fal.ai has a transparent pricing system once you separate hosted outputs from custom GPU runners. The confusion comes from compressing all of those units into one platform-level slogan.

Use fal Model APIs when the hosted endpoint and output unit fit. Use fal GPU when custom control and utilization justify runner billing. Use LinkModel when the same popular commercial model is available at a lower rate and you want the price to stay aligned with the official token or image unit. The LinkModel and fal.ai comparison turns those pricing differences into a production decision.

Official-unit pricing

See the commercial-model price before you call

LinkModel keeps the official token and per-image unit visible, so the discount is auditable — one key across text, image and video.

Related Posts