Token Factory delivers sub-50ms TTFT on dedicated FluidCore GPU clusters. Pay per million tokens processed — input and output billed separately for full cost transparency.
Context, prompts, system
First-write prompt caching
Generated completions
| Model tier | Input | Output | Context window | Latency SLA |
|---|---|---|---|---|
| Nano — fast, lightweight tasks | ₹4 | ₹12 | 8K | ≤ 25ms TTFT |
| Compact — balanced performance | ₹12 | ₹12 | 32K | ≤ 50ms TTFT |
| Pro — complex reasoning | ₹40 | ₹12 | 128K | ≤ 100ms TTFT |
| Max — frontier, 200K context | ₹120 | ₹12 | 200K | Dedicated GPU SLA |
| Monthly spend commitment | Discount | Context window |
|---|---|---|
| ₹50,000 – ₹2,00,000 / month | 10% off | Day 1 |
| ₹2,00,000 – ₹10,00,000 / month | 20% off | Day 1 |
| ₹10,00,000 – ₹50,00,000 / month | 30% off | Day 1 |
| ₹50,00,000+ / month | Custom | -- |
Data sovereignty, At-rest encryption, Real-time usage API, S3-compatible SDK are included with all Storage Plans
No noisy-neighbour interference
SSE & WebSocket streaming APIs
Migrate existing integrations in minutes
Automatic 75% cost reduction on repeated context
Zero training on your inference data
Per-tier uptime guarantees with credits