Three flat rates on the API and gateway: no subscription tiers, no monthly minimums, no per-seat charges — you pay a markup on top of provider cost. (The standalone Cloak apps are sold separately, on their own monthly plans.)
Pick a mode + model + monthly token volume. The provider list prices below are a static snapshot last checked 2026-07-05; they may have changed. Check each provider's current published rates.
Provider list prices in USD. Local models bill against a cloud-equivalent basis at 5% markup. Move the slider to see how the typical-chat blended rate moves with your input/output ratio.
| Model | Input / 1M | Output / 1M | Typical chat / 1M |
|---|---|---|---|
| Anthropic · Claude Opus 4.8 | $5.00 | $25.00 | — |
| Anthropic · Claude Sonnet 5 | $3.00 | $15.00 | — |
| OpenAI · GPT-5.5 | $5.00 | $30.00 | — |
| OpenAI · GPT-5.4 | $2.50 | $15.00 | — |
| xAI · Grok 4.3 | $1.25 | $2.50 | — |
| Google · Gemini 3.5 Flash | $1.50 | $9.00 | — |
| DeepSeek · V4 Flash | $0.14 | $0.28 | — |
| Local · Llama 3.2 8B (Ollama, 16+ GB VRAM) | $0.00 (own GPU) | $0.00 (own GPU) | — |
| Local · Llama 3 70B (Ollama, 80 GB VRAM) | $0.00 (own GPU) | $0.00 (own GPU) | — |
| Local · Qwen 2.5 7B (Ollama, 16+ GB VRAM) | $0.00 (own GPU) | $0.00 (own GPU) | — |
Prices in cents per 1M tokens stored as data-in / data-out. "Typical chat" = blended rate at the selected ratio (default 1 : 4 — 1 input token per 4 output tokens — matching most assistant-style traffic). Local rows show the 5% CloakAPI markup against a cloud-equivalent basis, since the model itself runs free on your hardware.
No tier comparison columns. There aren't any. Every account gets all of these.
Every gateway promises to handle your data carefully. CloakAPI promises something stronger — three claims, each one you can open.