Founder Stack

Together AI is betting the open-weight ecosystem n

Together AI is betting the open-weight ecosystem needs its own AWS

· AI · The Information

Together AI has positioned itself as the infrastructure layer for a corner of the AI market that the biggest cloud providers have been slower to serve well: companies that want to run open-weight models like Llama, Mixtral, DeepSeek, and Qwen at scale without building their own GPU clusters or navigating Nvidia's allocation queues directly. Founded by Vipul Ved Prakash, Ce Zhang, Chris Re, and Percy Liang, the company built its early reputation on research contributions to distributed training and inference optimization before turning those techniques into a commercial cloud platform that now serves both fine-tuning and high-throughput inference workloads. The business model sits between a traditional cloud provider and a model API company. Together AI does not train its own frontier models in the way OpenAI or Anthropic do, but it also does not merely resell raw GPU capacity like a bare-metal provider; instead it optimizes inference serving for open-weight architectures, often achieving latency and cost figures that beat what customers could get running the same models on general-purpose cloud instances. That positioning attracted a $305 million Series B in early 2024 at a $3.3 billion valuation, backed by investors including General Catalyst, Prosperity7, and Salesforce Ventures, and the company has continued raising and deploying capital into GPU capacity commitments through 2025 and into 2026. The competitive set has grown crowded quickly. Fireworks AI and Groq both compete for inference workloads on open-weight models, CoreWeave and Lambda compete on raw GPU cloud capacity, and the hyperscalers themselves, Amazon, Microsoft, and Google, have all built managed services for hosting open-weight models that erode some of Together's differentiation for customers already committed to a single cloud. Together's counter has been to lean harder into research-driven efficiency gains, publishing techniques that reduce serving costs per token, and into enterprise fine-tuning services that require more hand-holding than a commodity API can provide. The compute economics underlying the whole open-weight inference category remain fragile. GPU costs, chip allocation constraints, and the capital intensity of building or leasing data center capacity mean that inference-layer companies without their own chip advantage, unlike Groq with its LPU, are effectively arbitraging Nvidia's supply chain and margin structure. That works as long as demand for open-weight inference keeps growing faster than hyperscaler capacity expansion, but it leaves Together AI structurally exposed to any pricing war the hyperscalers choose to start. What to watch: whether Together AI raises another growth round in 2026 at a valuation that reflects inference-market growth rather than compressed margins, whether its efficiency research keeps it ahead of hyperscaler-native open-weight hosting, and whether it expands into training infrastructure for mid-size labs rather than remaining purely an inference layer.

Original source: The Information
Read more on Founder Stack