High-throughput LLM inference

Token Factory delivers sub-50ms TTFT on dedicated FluidCore GPU clusters. Pay per million tokens processed — input and output billed separately for full cost transparency.

Input tokens

Context, prompts, system

₹12/ M tokens

Output tokens

Generated completions

₹36/ M tokens

Model tiers — pricing by capability class

Model tierInputOutputContext windowLatency SLA
Nano — fast, lightweight tasks₹4₹128K≤ 25ms TTFT
Compact — balanced performance₹12₹1232K≤ 50ms TTFT
Pro — complex reasoning₹40₹12128K≤ 100ms TTFT
Max — frontier, 200K context₹120₹12200KDedicated GPU SLA

Committed usage — discounts

Monthly spend commitmentDiscountContext window
₹50,000 – ₹2,00,000 / month10% offDay 1
₹2,00,000 – ₹10,00,000 / month20% offDay 1
₹10,00,000 – ₹50,00,000 / month30% offDay 1
₹50,00,000+ / monthCustom--

Included with Token Factory

Data sovereignty, At-rest encryption, Real-time usage API, S3-compatible SDK are included with all Storage Plans

Dedicated GPU clusters

No noisy-neighbour interference

Streaming output

SSE & WebSocket streaming APIs

OpenAI-compatible API

Migrate existing integrations in minutes

Prompt caching

Automatic 75% cost reduction on repeated context

Data isolation

Zero training on your inference data

99.9% availability SLA

Per-tier uptime guarantees with credits