Skip to main content
F

fal

fal runs a generative media cloud offering 600+ image, video, audio and 3D models via a unified API, backed by a fast inference engine and serverless NVIDIA GPU compute.

fal is a generative media cloud, founded in 2021, that gives developers access to more than 600 production-ready image, video, audio and 3D models through a unified API. The company builds both the inference engine behind those models and the GPU infrastructure that runs them, covering serverless on-demand compute and dedicated clusters. Its technical work spans inference optimisation, large-scale model serving and NVIDIA GPU infrastructure, including H100, H200 and B200 chips.

The platform's inference engine is described as the fastest available for generative media, with speed advantages of up to 10x over alternatives on models such as SDXL and Whisper. On-demand serverless GPUs are priced from $1.20 per hour for premium hardware, while dedicated compute clusters are offered for training and deploying custom AI models. The service reports 99.99% uptime and support for more than 100m daily inference calls.

More than 1,000,000 developers use the platform, and the fastest inference models built on it serve hundreds of millions of end customers. fal operates globally and holds SOC 2 compliance, with an emphasis on enterprise-grade security. Its software is used across generative AI, creative and design tools, enterprise software, AI assistants and search, and media and content generation.

Open jobs

No open jobs right now.