Together AI
Fast inference API for open-source models — Llama, Mixtral, Flux, at low cost.
Add Together AI to your hut →Together AI is one of the leading inference clouds for open models — a single API serving hundreds of open-weight models (Llama, DeepSeek, Qwen, Flux, and more) with serverless pay-per-token pricing, dedicated GPU endpoints, and fine-tuning. It runs its own GPU fleet, so the pitch is raw speed and cost at scale rather than aggregation.
Most often compared to OpenRouter and Fireworks — OpenRouter routes one key across many providers (including closed models), while Together hosts the models itself. Choose Together when you've settled on open models and want throughput, fine-tuning, and predictable per-token cost from a single provider.
| Made by | Together AI |
|---|---|
| Pricing | Pay-per-token serverless · dedicated endpoints · fine-tuning per token |
| Best for | Open-model inference at scale, fine-tuning, dedicated endpoints |
Alternatives to Together AI
- Hugging Face
The GitHub of AI — browse, download, and deploy 500,000+ open-source models and datasets.
- OpenRouter
Single API for 200+ LLMs — route between Claude, GPT, Gemini, Llama, and more.
- Replicate
Run open-source ML models via API — image, video, audio, LLMs, no infra needed.
- Groq
Extremely fast LLM inference on custom LPU hardware.
- Fireworks AI
Fast, low-cost inference and fine-tuning for open models.
- fal.ai
Fast generative media inference — image, video, and audio APIs.