A cloud inference platform that hosts open-weight LLMs with fast, low-cost serving.
It offers major open models such as Llama and Mixtral through an OpenAI-compatible API, and also provides fine-tuning and dedicated GPU clusters.
© 2026 ITBGM