Three flat rates on the API and gateway: no subscription tiers, no monthly minimums, no per-seat charges — you pay a markup on top of provider cost. (The standalone Cloak apps are sold separately, on their own monthly plans.)
Pick a mode + model + monthly token volume. The numbers below are all live; we maintain the price book at docs/pricing/price-book.json.
Provider list prices in USD. Local models bill against a cloud-equivalent basis at 5% markup. Move the slider to see how the typical-chat blended rate moves with your input/output ratio.
| Model | Input / 1M | Output / 1M | Typical chat / 1M |
|---|---|---|---|
| Anthropic · Claude Opus 4.8 | $5.00 | $25.00 | — |
| Anthropic · Claude Sonnet 5 | $3.00 | $15.00 | — |
| OpenAI · GPT-5.5 | $5.00 | $30.00 | — |
| OpenAI · GPT-5.4 | $2.50 | $15.00 | — |
| xAI · Grok 4.3 | $1.25 | $2.50 | — |
| Google · Gemini 3.5 Flash | $1.50 | $9.00 | — |
| DeepSeek · V4 Flash | $0.14 | $0.28 | — |
| Local · Llama 3.2 8B (Ollama, 16+ GB VRAM) | $0.00 (own GPU) | $0.00 (own GPU) | — |
| Local · Llama 3 70B (Ollama, 80 GB VRAM) | $0.00 (own GPU) | $0.00 (own GPU) | — |
| Local · Qwen 2.5 7B (Ollama, 16+ GB VRAM) | $0.00 (own GPU) | $0.00 (own GPU) | — |
Prices in cents per 1M tokens stored as data-in / data-out. "Typical chat" = blended rate at the selected ratio (default 1 : 4 — 1 input token per 4 output tokens — matching most assistant-style traffic). Local rows show the 5% CloakAPI markup against a cloud-equivalent basis, since the model itself runs free on your hardware.
No tier comparison columns. There aren't any. Every account gets all of these.
Every gateway promises to handle your data carefully. CloakAPI promises something stronger — three claims, each one you can open.
docs/pricing/price-book.json is source-of-truth and gets a new version every time a provider posts a change; subscribe to its RSS for advance notice.