Baseten is the infrastructure layer that lets AI-native companies run their models in production. Think "AWS for inference": a company packages its custom or open-source model (Llama, DeepSeek, etc.), Baseten serves it at scale across 20+ cloud providers, handling GPU orchestration, autoscaling, observability, and billing. Customers pay per GPU-minute for dedicated deployments or per token via a Model APIs catalog. Beyond pure serving, Baseten now also does post-training — fine-tuning and optimizing models so customers compound value from proprietary data. Founded in 2019 and headquartered in San Francisco, Baseten targets AI-native builders (Cursor, Notion, Abridge, Clay, Gamma, Writer, Harvey, HubSpot) who need mission-critical inference uptime and performance but don't want to build GPU infrastructure themselves. Revenue is usage-based and scales directly with customer AI compute consumption — no seat licenses.