Meta / meta/llama-3.1-8b-instruct-16k
Llama 3.1 8B Instruct - access through LLMTR
Llama 3.1 8B Instruct is Meta's fast, efficient, small open-weight instruction model. It suits short chat, classification, data transformation, and light automation tasks that need low latency, and it supports JSON-formatted output. Both the context window and the maximum single-response output are capped at 16,384 tokens. This model does not support function calling.
Technical specifications
| Canonical ID | meta/llama-3.1-8b-instruct-16k |
|---|---|
| Provider | Meta |
| Context window | 16,384 tokens |
| Operations | CHAT_COMPLETIONS |
| Modalities | text |
Pricing
An 8% platform margin applies to credit top-ups; model usage prices are not separately marked up.
| Operation | Metric | Unit | Price |
|---|---|---|---|
| CHAT_COMPLETIONS | INPUT_TEXT | PER_1M_TOKENS | $0.020000 |
| CHAT_COMPLETIONS | OUTPUT_TEXT | PER_1M_TOKENS | $0.050000 |
Example usage
With existing OpenAI SDK flows, change only the base URL and model identifier.
curl https://llmtr.com/v1/chat/completions \
-H "Authorization: Bearer llmtr-your_key" \
-H "Content-Type: application/json" \
-d '{"model":"meta/llama-3.1-8b-instruct-16k","messages":[{"role":"user","content":"Hello"}]}'
Related models
- Muse Spark 1.2 meta/muse-spark-1.2
- Muse Glimmer 30B meta/muse-glimmer-30b
- Muse Spark 1.1 meta/muse-spark-1.1
- Muse Spark 1.2 Contributor meta/muse-spark-1.2-contributor
- Llama 3.3 70B Instruct (12K) meta/llama-3.3-70b-instruct-12k
- Gemma 4 llmtr/gemma-4
- Qwen 3.6 35B-A3B llmtr/qwen3-6-35b
- Trendyol Asure 12B llmtr/trendyol-asure-12b