Powered by TextCLF Quant

Low-cost LLM inference, without the accuracy tradeoff.

TextCLF is an inference provider: we run open and frontier models for you behind one OpenAI-compatible API. Our TextCLF Quant compression serves them without losing accuracy, so you get sub-second responses at a fraction of the price — no GPUs to manage, no lock-in.

$5 free credit · no credit card required

Llama 3.3 70B · input

USD / 1M tokens

Closed frontier API$2.50
Other open providers$0.59
TextCLF$0.08

up to 96% lower than leading closed APIs — while preserving full model accuracy.

Benchmarked in production

$0.018
per 1M tokens

Starting price, Llama 3.1 8B

81%
cheaper

vs. leading closed APIs

<200ms
time to first token

p50 across all models

99.99%
uptime

Rolling 90-day average

The technology

The same models, for a fraction of the price.

To run big models cheaply, providers quantize them — and most quantize so aggressively that they quietly trade away accuracy. Our TextCLF Quant doesn't: it makes models far cheaper to run while keeping their answers just as good — and we hand every bit of that saving straight back to you as a lower price per token.

4.1x

smaller footprint

99.6%

accuracy retained vs. FP16

  • The most accuracy per bit

    At any given size — say 4 bits per weight — there’s a mathematical best for how much accuracy you can keep. Most methods leave a lot on the table; our TextCLF Quant gets remarkably close to that best, so at the same size our models stay noticeably sharper.

  • Same answers, proven

    We test every model against the untouched original and keep it within a fraction of a percent on standard benchmarks — no silent quality drops.

  • The savings go straight to you

    Smaller models are cheaper for us to run — and we don’t pocket the difference. Every bit of efficiency TextCLF Quant unlocks is passed directly to you as a lower price per token.

One API. Every model worth running.

Switch models with a single string. Every model on TextCLF is served from the same OpenAI-compatible endpoint at the same low, transparent per-token price.

  • Llama 3.3 70B Instruct

    128K context

    Meta

    Input $0.08
    Output $0.30
    6.6 tok/s
  • Llama 3.1 8B Instruct

    128K context

    Meta

    Input $0.018
    Output $0.038
    21.2 tok/s
  • Qwen3.8 27B

    262K context

    Alibaba

    Input $0.38
    Output $2.98
    21.3 tok/s
  • DeepSeek V4 FlashSoon

    1M context

    DeepSeek

    Input $0.05
    Output $0.12

Prices shown per 1M tokens (USD). 100+ additional models available.

View full model catalog →

Change one line. Cut your bill.

Already using the OpenAI SDK? Point the base URL at TextCLF and keep the rest of your code exactly as it is.

  • Drop-in replacement for the OpenAI SDK
  • Same request and response schema
  • Streaming, tools, and JSON mode supported
quickstart.sh
curl https://api.textclf.com/v1/chat/completions \
  -H "Authorization: Bearer YOUR_TEXTCLF_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3.3-70B-Instruct",
    "messages": [
      { "role": "user", "content": "Explain inference cost in one line." }
    ]
  }'

Prepaid credit. Zero surprises.

Get $5 in free credit to start — no credit card needed. When you're ready for more, top up your balance and spend it at each model's per-token rate. You can never be charged more than you've added.

Free credit

$5on us

Start building the moment you sign up — no credit card required.

  • $5 in credit, free
  • No credit card to start
  • Access to every model
  • OpenAI-compatible API
Claim free credit

Prepaid credits

Most popular
Top upfrom $10

Add credit and draw it down as you call models. Each model bills at its own per-token rate — you only spend what you use.

  • Per-token pricing, set per model
  • Credits never expire
  • Auto-reload so you never run dry
  • Real-time usage & spend dashboard
Buy credits

Enterprise

Customvolume

Committed-use rates, dedicated capacity, and invoicing.

  • Volume discounts on credit
  • Dedicated GPU capacity
  • SOC 2 Type II, SSO/SAML
  • Invoicing & dedicated support
Contact sales

Every model has its own input and output rate. See per-token pricing for all models.