14 models live
Llama, behind one base URL.
A Llama proxy behind one endpoint. Keep the SDK you already use, add one key, and call every Llama model in the live catalog — free to start, at the real upstream rate.
| Model | Kind | Context | Text input | Text output | Cache read |
|---|---|---|---|---|---|
| aion-rp-llama-3-1-8b | chat | 33k | $0.8 | $1.6 | — |
| deepseek-r1-distill-llama-70b | chat | 8k | $0.8 | $0.8 | — |
| hermes-3-llama-3-1-405b | chat | 131k | $1 | $1 | — |
| hermes-3-llama-3-1-70b | chat | 131k | $0.7 | $0.7 | — |
| llama-3-1-70b-instruct | chat | 131k | $0.4 | $0.4 | — |
| llama-3-1-8b-instruct | chat | 131k | $0.05 | $0.08 | $0.025 |
| llama-3-2-1b-instruct | chat | 60k | $0.027 | $0.2 | — |
| llama-3-2-3b-instruct | chat | 131k | $0.05 | $0.33 | — |
| llama-3-3-70b-instruct | chat | 131k | $0.1 | $0.32 | — |
| llama-4-maverick | chat | 1M | $0.2 | $0.7 | — |
| llama-4-scout | chat | 1.3M | $0.1 | $0.3 | — |
| llama-guard-4-12b | chat | 1M | $0.18 | $0.18 | — |
| llama-nemotron-embed-vl-1b-v2 | embedding | 131k | — | — | — |
| llama-nemotron-rerank-vl-1b-v2 | rerank | 10k | — | — | — |
Questions
Straight answers.
Llama keys, pricing, and the basics.
01Is this a Llama reverse proxy?
In effect, yes. Your request hits our gateway, we route it upstream to a provider serving Llama, and the response comes back the way your client expects.
02Do I need my own provider account?
No. We run the upstream accounts. One SilvR key covers Llama and the rest of the catalog.
03Can I use Llama for free?
Yes. Every account starts on the free plan with a daily allowance you can spend on any model in the catalog, Llama included.