14 models live

Llama, behind one base URL.

A Llama proxy behind one endpoint. Keep the SDK you already use, add one key, and call every Llama model in the live catalog — free to start, at the real upstream rate.

14 models
ModelKindContextText inputText outputCache read
aion-rp-llama-3-1-8bchat33k$0.8$1.6
deepseek-r1-distill-llama-70bchat8k$0.8$0.8
hermes-3-llama-3-1-405bchat131k$1$1
hermes-3-llama-3-1-70bchat131k$0.7$0.7
llama-3-1-70b-instructchat131k$0.4$0.4
llama-3-1-8b-instructchat131k$0.05$0.08$0.025
llama-3-2-1b-instructchat60k$0.027$0.2
llama-3-2-3b-instructchat131k$0.05$0.33
llama-3-3-70b-instructchat131k$0.1$0.32
llama-4-maverickchat1M$0.2$0.7
llama-4-scoutchat1.3M$0.1$0.3
llama-guard-4-12bchat1M$0.18$0.18
llama-nemotron-embed-vl-1b-v2embedding131k
llama-nemotron-rerank-vl-1b-v2rerank10k

Questions

Straight answers.

Llama keys, pricing, and the basics.

01Is this a Llama reverse proxy?

In effect, yes. Your request hits our gateway, we route it upstream to a provider serving Llama, and the response comes back the way your client expects.

02Do I need my own provider account?

No. We run the upstream accounts. One SilvR key covers Llama and the rest of the catalog.

03Can I use Llama for free?

Yes. Every account starts on the free plan with a daily allowance you can spend on any model in the catalog, Llama included.