TL;DR: fal.ai pricing splits into two products: hosted Model APIs (usually billed per successful output — image, megapixel, video second, video, request or token) and Serverless / fal GPU (billed over total runner lifetime). The biggest mistake is calling all fal usage "GPU-second billing." For Seedance 2.0, fal lists $14.00 per million tokens vs LinkModel's $6.30; for GPT Image 2 token categories LinkModel is 25% lower. fal GPU list rates run from $2.99/hr (RTX PRO 6000) to $8.50/hr (B300). The full fal.ai review covers model access and concurrency.
fal.ai pricing becomes clear once you separate two products. The broader fal.ai review covers model access, concurrency and production fit:
- fal Model APIs: call a hosted model and usually pay for successful output.
- fal Serverless / fal GPU: deploy custom code and pay for runner lifetime on selected hardware.
The biggest pricing mistake is to describe all fal.ai usage as GPU-second billing. That is not what the current documentation says. Many hosted Model APIs are billed per image, megapixel, generated video second, video, request or token. GPU lifecycle billing belongs to the separate Serverless path.
fal AI Model Pricing: How Hosted Model APIs Are Billed
| Model API output | Common fal billing unit |
|---|---|
| Image | Per image or megapixel |
| Video | Per generated second or complete video |
| Text/selected multimodal models | Tokens |
| Other hosted endpoints | Per request or compute time when no fixed rate exists |
| Server error/queue time | fal says these are not charged |
fal states that Model APIs charge for successful outputs. Server errors and queue time are not charged, and Model API cold starts are not charged. A 422 client error may still be billable if compute occurred before the request was rejected.
That is different from fal GPU billing, which appears later in this article.
Seedance Pricing on fal.ai: Native Tokens and Displayed Seconds
fal’s Seedance 2.0 model page publishes:
- $0.014 per 1,000 tokens
- equivalent native rate: $14.00 per million tokens
- page display around $0.3034/sec at 720p
- page display around $0.682/sec at 1080p
The official token formula shown by fal is:
Video tokens = height × width × duration × 24 ÷ 1,024
720p calculation
Assume 1280×720, 24 fps and one second:
720 × 1,280 × 1 × 24 ÷ 1,024 = 21,600 tokens/sec
21,600 ÷ 1,000,000 × $14.00 = $0.3024/sec
fal’s page displays $0.3034/sec. Under the stated 1280×720 assumption, that display is not mathematically identical to the page’s own $14.00/M token rate and formula. For an auditable comparison, preserve both numbers: treat $0.3034 as fal’s displayed estimate and $0.3024 as the calculation from its published native token price. Do not reverse the displayed estimate into a new “official” token quote.
1080p calculation
Assume 1920×1080, 24 fps and one second:
1,080 × 1,920 × 1 × 24 ÷ 1,024 = 48,600 tokens/sec
48,600 ÷ 1,000,000 × $14.00 = $0.6804/sec
fal displays about $0.682/sec. Under the stated 1920×1080 assumption, the published formula produces $0.6804/sec. Again, keep the displayed estimate and calculated value separate; the native token rate is the stable comparison unit.
fal.ai vs LinkModel Seedance Pricing
BytePlus lists the official no-video-input Seedance rates at $7.00/M for 720p and $7.70/M for 1080p. LinkModel lists $6.30/M and $6.93/M; fal lists $14.00/M.
Native million-token prices:
| Seedance workload | BytePlus official | fal native | LinkModel |
|---|---|---|---|
| 720p, no video input | $7.00/M | $14.00/M | $6.30/M |
| 1080p, no video input | $7.70/M | $14.00/M | $6.93/M |
Converted per-second estimates:
| Seedance workload | Official calculation* | fal calculation* | LinkModel calculation* |
|---|---|---|---|
| 720p, no video input | $0.15120 | $0.30240 | $0.13608 |
| 1080p, no video input | $0.37422 | $0.68040 | $0.336798 |
Sources: BytePlus pricing, fal Seedance page, and LinkModel models. Derived from the published token rates at 24 fps; these are calculations, not independent provider quotes.
Cost for 1,000 generated seconds
| Workload | BytePlus official calculation | fal native-token calculation | LinkModel calculation |
|---|---|---|---|
| 1,000 sec at 720p | $151.20 | $302.40 | $136.08 |
| 1,000 sec at 1080p | $374.22 | $680.40 | $336.798 |
The per-second view is useful for planning. It becomes a weak comparison when the token formula disappears and only the rounded second remains. LinkModel keeps the official-style million-token row visible, so the discount or markup is easier to see.
fal.ai Models Pricing for GPT Image 2 Tokens
fal’s GPT Image 2 page publishes the upstream token categories:
| GPT Image 2 category | OpenAI official / 1M tokens | fal.ai / 1M tokens | LinkModel / 1M tokens |
|---|---|---|---|
| Text input | $5.00 | $5.00 | $3.75 |
| Cached text input | $1.25 | $1.25 | $0.9375 |
| Image input | $8.00 | $8.00 | $6.00 |
| Cached image input | $2.00 | $2.00 | $1.50 |
| Image output | $30.00 | $30.00 | $22.50 |
Sources: OpenAI API pricing, fal image model page, and the LinkModel models; verified July 22, 2026.
fal also publishes canonical per-image examples because quality and size change the number of image tokens. For example, its 1024×1024 table lists $0.006 low, $0.053 medium and $0.211 high.
Per-image presets are convenient. Token categories are better when your application mixes text, cached context, image input and output. LinkModel’s rate is 25% lower in every same token category listed above.
fal.ai Image Pricing by Number of Images
fal prices Gemini 3.1 Flash Image at $0.08 for 1K, with 0.75× for 512, 1.5× for 2K and 2× for 4K. LinkModel publishes direct call prices for 0.5K, 1K, 2K and 4K.
Per-image prices:
| Resolution | Google official / image | fal.ai / image | LinkModel / image |
|---|---|---|---|
| 0.5K | $0.04500 | $0.06000 | $0.03375 |
| 1K | $0.06700 | $0.08000 | $0.05025 |
| 2K | $0.10100 | $0.12000 | $0.07575 |
| 4K | $0.15100 | $0.16000 | $0.11325 |
Cost for 10,000 images:
| Resolution | Google official | fal.ai | LinkModel |
|---|---|---|---|
| 0.5K | $450.00 | $600.00 | $337.50 |
| 1K | $670.00 | $800.00 | $502.50 |
| 2K | $1,010.00 | $1,200.00 | $757.50 |
| 4K | $1,510.00 | $1,600.00 | $1,132.50 |
Sources: Google API pricing, fal Gemini image page, and the LinkModel models. fal adds $0.015 when web search is used; that optional feature is excluded from the table.
fal GPU Pricing: List Rate vs “As Low As”
fal’s pricing page publishes both list and lower “as low as” rates.
| fal GPU | Memory | List price/hour | “As low as”/hour |
|---|---|---|---|
| B300 | 288 GB | $8.50 | $4.49 |
| B200 | 180 GB | $6.25 | $3.49 |
| H200 | 141 GB | $4.50 | $2.10 |
| H100 | 80 GB | $3.99 | $1.89 |
| RTX PRO 6000 | 96 GB | $2.99 | $1.10 |
Source: fal.ai pricing, verified July 22, 2026.
“As low as” is not the same as a universal on-demand rate. Use list price for conservative budgeting until your workload qualifies for the lower number.
How fal Serverless GPU Cost Is Calculated
fal Serverless documentation describes billing over total runner lifetime. Depending on configuration, that can include:
- setup;
- idle time;
- active processing;
- draining;
- teardown.
Pending tasks and image pulls are described separately from active runner billing. Settings such as keep_alive, minimum concurrency and the number of GPUs can materially change cost.
Basic fal GPU formula
Estimated compute cost = GPU hourly rate × billable runner seconds ÷ 3,600 × GPU count
Example using the H100 list rate and one GPU for a 20-minute billable runner lifetime:
$3.99 × 1,200 ÷ 3,600 × 1 = $1.33
At the published “as low as” H100 rate:
$1.89 × 1,200 ÷ 3,600 = $0.63
These are pure rate calculations, not guarantees of availability, performance or eligibility. Your real job cost also depends on startup, idle capacity, throughput and retries.
When fal GPU Is Cheaper—and When It Is Not
fal GPU can be more economical when:
- utilization is consistently high;
- you batch requests efficiently;
- custom code or model control removes other tooling;
- a fixed hosted output price includes a margin you can avoid;
- the lower GPU rate is actually available.
A hosted Model API can be cheaper when:
- traffic is sporadic;
- cold and idle capacity would dominate runner time;
- you value a fixed output unit;
- engineering time to package, observe and optimize the model is significant.
Do not compare “$1.89/hour H100” directly with “$0.08/image” without measuring images per billable runner hour at the same quality.
fal Credits, Concurrency and Failed Requests
- Purchased credits expire after 365 days.
- New Model API accounts start with 2 concurrent requests.
- Limits can rise with paid invoices over the last four weeks, self-service up to 40.
- Requests above the limit queue.
- Server errors and queue time are not charged.
- Some 422 requests can be charged if GPU work occurred before validation.
These rules belong in your cost model because throughput and error handling affect how quickly credits are consumed.
fal.ai Models Pricing FAQs
Is fal.ai priced per second?
Some video Model APIs are priced per generated second, but not every fal product uses that unit. Seedance also publishes a native token price, and Serverless uses GPU runner lifetime.
Does fal.ai price every model by GPU time?
No. Hosted Model APIs usually use an output unit. GPU time is central to the separate Serverless product and some endpoints without predefined output pricing.
What does fal GPU “as low as” mean?
It is the lowest displayed rate, not necessarily the rate for every account or job. Budget with list price until eligibility is confirmed.
Is fal.ai cheaper than LinkModel?
It varies by model. In the verified Seedance, GPT Image 2 token-category and Gemini 3.1 Flash Image comparisons above, LinkModel is lower. The fal.ai alternatives guide compares other providers for teams considering a switch.
How do I estimate 10,000 fal image generations?
Multiply the exact per-image rate by 10,000, then add optional features such as web search and account-specific discounts. Do not use a 1K price for 4K output.
Can fal charge a failed request?
fal says 500-series errors are not charged. A 422 client error may be billable when GPU work occurred before the error was detected.
Final fal Pricing Takeaway
fal.ai has a transparent pricing system once you separate hosted outputs from custom GPU runners. The confusion comes from compressing all of those units into one platform-level slogan.
Use fal Model APIs when the hosted endpoint and output unit fit. Use fal GPU when custom control and utilization justify runner billing. Use LinkModel when the same popular commercial model is available at a lower rate and you want the price to stay aligned with the official token or image unit. The LinkModel and fal.ai comparison turns those pricing differences into a production decision.
See the commercial-model price before you call
LinkModel keeps the official token and per-image unit visible, so the discount is auditable — one key across text, image and video.
