Free
For prototypes and side projects.
- 1M tokens of routed inference
- All 200+ models
- 7-day trace retention
- Community support
- 1 project, 3 members
No seat minimums, no markup on model tokens. The platform fee covers routing, traces, and support — everything that is actually ours to run.
For prototypes and side projects.
For teams running AI in production.
For regulated and high-volume workloads.
| Free | Pro | Enterprise | |
|---|---|---|---|
| Included routed inference | 1M tokens | 25M tokens | Committed volume |
| Model token rates | Provider list | Provider list | Provider list or your own contracts |
| Models available | All 200+ | All 200+ | All 200+, plus private |
| Smart routing | Not included | Included | Included |
| Per-request budgets | Not included | Included | Included |
| Shadow traffic and replays | Not included | Included | Included |
| Trace retention | 7 days | 90 days | Up to 2 years |
| OpenTelemetry export | Not included | Included | Included |
| Projects | 1 | Unlimited | Unlimited |
| Members | 3 | Unlimited | Unlimited |
| SSO and SAML | Not included | Not included | Included |
| SCIM provisioning | Not included | Not included | Included |
| Audit logs | Not included | 30 days | 2 years, exportable |
| Data residency | Not included | US or EU | US, EU, UK, or single-tenant |
| Zero-retention mode | Not included | Included | Included |
| Private or VPC deployment | Not included | Not included | Included |
| Uptime SLA | None | 99.9% | 99.99% |
| Support | Community | Email and Slack | Dedicated channel, response SLAs |
| Billing | Card | Card | Invoice, annual |
Two ways, billed together. Routed inference is metered per token at the provider's published rate, passed through without markup. The platform fee covers routing, traces, and support. Every request returns its own cost, so your accounting matches ours.
Requests over the cap return a 429 with a machine-readable reason rather than silently costing money. Caps are set per project and per key, and alerts fire at 50%, 80%, and 100%.
No. Model tokens are billed at provider list rates. Vantafold makes money on the platform fee, which means routing you to a cheaper model is aligned with your interests, not against them.
Yes, on Pro and Enterprise. Attach your own provider keys and Vantafold routes through them, billing you only the platform fee. Useful if you hold committed-spend agreements.
The free tier is intended for prototypes and evaluation. It has no SLA and traces are kept for seven days. Production workloads should be on Pro or Enterprise.
Annual, based on committed volume and deployment model. Private and VPC deployments, data residency, and dedicated support are included. Talk to us for a quote.
Yes. Upgrades take effect immediately and are prorated. Downgrades take effect at the end of the current billing period so you keep what you have paid for.
The free tier needs no card, and the numbers you see in the dashboard are the numbers we bill.