Together AI is a San Francisco-based "neocloud" that gives developers and enterprises a single platform to run, fine-tune, and deploy open-source AI models — without touching a hyperscaler. Its core products are: (1) a high-performance LLM inference API that hosts hundreds of open-weight models; (2) Together GPU Clusters — self-serve, API-first Nvidia GPU clusters scaling from 8 to 100,000+ GPUs (H100/H200/B200/GB200 NVL72); (3) a Model Shaping suite for fine-tuning and post-training; and (4) a real-time Voice AI platform that co-locates speech-to-text, LLM, and text-to-speech on one cloud with sub-500ms latency. The company builds its own inference engine research (FlashAttention-4, ATLAS speculative decoding) rather than just reselling compute. Buyers are AI-native companies, frontier labs, and enterprises that want open-source model flexibility at far lower cost than closed-model APIs like OpenAI or Anthropic. Founded 2022.