Model catalog

Everything we serve, with the numbers most providers leave out. Quantization is a spec, not a footnote — every model runs with TextCLF Quant (TQ), and you can see the exact precision, latency, and price before you spend a cent.

3 models serving now

Me

Llama 3.3 70B Instruct

Live
meta-llama/Llama-3.3-70B-Instruct

Parameters

70B

Max context

128K

Meta’s flagship dense model, tuned for instruction following, tool use, and multilingual chat. TQ 4-bit keeps it within a fraction of a point of full precision on reasoning benchmarks while cutting the memory it takes to serve.

ProviderTextCLF
Context128K
P50 TTFT544ms
Throughput6.6 tok/s
Uptime99.99%
QuantizationTQ 4-bit
Input /M$0.08
Output /M$0.30
generaltool-usemultilingual
meta-llama/Llama-3.3-70B-Instruct
Me

Llama 3.1 8B Instruct

Live
meta-llama/Llama-3.1-8B-Instruct

Parameters

8B

Max context

128K

The fastest, cheapest model in the catalog. Ideal for classification, extraction, and high-volume routing where latency and price matter more than frontier reasoning.

ProviderTextCLF
Context128K
P50 TTFT140ms
Throughput21.2 tok/s
Uptime99.99%
QuantizationTQ 4-bit
Input /M$0.018
Output /M$0.038
fastcheapclassification
meta-llama/Llama-3.1-8B-Instruct
Al

Qwen3.8 27B

Live
Qwen/Qwen3.8-27B

Parameters

27B

Max context

262K

A balanced mid-tier model with strong coding and math performance and a long context window. The right default when 8B is too small and 70B is more than the task needs.

ProviderTextCLF
Context262K
P50 TTFT393ms
Throughput21.3 tok/s
Uptime99.98%
QuantizationTQ 4-bit
Input /M$0.38
Output /M$2.98
balancedcodinglong-context
Qwen/Qwen3.8-27B
De

DeepSeek V4 Flash

Coming soon
deepseek-ai/DeepSeek-V4-Flash-0731

Parameters

MoE (sparse)

Max context

1M

A fast mixture-of-experts model tuned for reasoning and coding, activating only a fraction of its parameters per token. TQ 4-bit makes frontier-class reasoning practical to serve at open-weight prices. Coming soon — join the waitlist from your dashboard.

ProviderTextCLF
Context1M
P50 TTFT
Throughput
Uptime
QuantizationTQ 4-bit
Input /M$0.05
Output /M$0.12

Prices per 1M tokens (USD). 100+ additional models available on request.