An AI inference platform specialized in serving open-weight models with low latency.
It runs open models through a custom-optimized inference engine using techniques like continuous batching and speculative decoding, making it well suited for latency-sensitive production workloads.
© 2026 ITBGM