- Context
- 1M
- Input / 1M
- $3.00
- Output / 1M
- $15.00
Model catalog
Frontier, open-weight, embedding, and private models — all reachable through the same endpoint. Rates are per 1M tokens and are passed through at provider list price.
- Context
- 400K
- Input / 1M
- $0.80
- Output / 1M
- $4.00
- Context
- 200K
- Input / 1M
- $0.18
- Output / 1M
- $0.72
- Context
- 500K
- Input / 1M
- $4.20
- Output / 1M
- $18.00
- Context
- 250K
- Input / 1M
- $0.95
- Output / 1M
- $3.80
- Context
- 300K
- Input / 1M
- $2.40
- Output / 1M
- $9.60
- Context
- 128K
- Input / 1M
- $0.60
- Output / 1M
- $2.40
- Context
- 128K
- Input / 1M
- $0.35
- Output / 1M
- $0.90
- Context
- 128K
- Input / 1M
- $0.16
- Output / 1M
- $0.48
- Context
- 64K
- Input / 1M
- $0.05
- Output / 1M
- $0.15
- Context
- 256K
- Input / 1M
- $0.28
- Output / 1M
- $0.84
- Context
- 128K
- Input / 1M
- $0.07
- Output / 1M
- $0.21
- Context
- 64K
- Input / 1M
- $0.06
- Output / 1M
- $0.18
- Context
- 32K
- Input / 1M
- $0.02
- Output / 1M
- $0.06
- Context
- 2M
- Input / 1M
- $0.55
- Output / 1M
- $1.65
- Context
- 32K
- Input / 1M
- $0.02
- Output / 1M
- —
- Context
- 32K
- Input / 1M
- $0.008
- Output / 1M
- —
- Context
- 16K
- Input / 1M
- $0.04
- Output / 1M
- —
- Context
- 32K
- Input / 1M
- $0.03
- Output / 1M
- $0.09
- Context
- 96K
- Input / 1M
- $0.12
- Output / 1M
- $0.36
- Context
- 64K
- Input / 1M
- $0.90
- Output / 1M
- $2.70
- Context
- Base model
- Input / 1M
- Base rate
- Output / 1M
- Base rate
No models match those filters. Clear one and try again.
What the numbers mean.
Context
Maximum combined input and output tokens for a single request. Prompt prefix caching does not change the window, only what you pay for the cached portion.
Rates
USD per 1M tokens, input and output, at provider list price. Vantafold adds no markup; the platform fee is separate and is on the pricing page.
Capabilities
What the model supports
through the unified schema. A model without a tools tag will reject a request
containing tool definitions rather than silently ignoring them.
Pin a model, or let the router pick.
Both go through the same endpoint. Change your mind whenever you like.