Underneath every box above sits the one nobody thinks about until the bill arrives: inference, the compute layer that actually runs the models. When a creator hits generate on Krea, or an app quietly calls a model behind the scenes, someone is serving that model on a GPU, and a whole business exists to do it faster and cheaper than the labs do themselves. fal is the developer favorite for fast generative-media inference, and in 2026 it named AWS its preferred cloud and said it was serving roughly 2.5 million developers; Replicate, now part of Cloudflare, made running and fine-tuning open models a one-line API; Baseten and Together AI sell production infrastructure and autoscaling, and both raised enormous rounds in mid-2026, Baseten near a 13 billion dollar valuation and Together AI at 8.3 billion, on inference demand that each said was growing many times over year on year; Hugging Face remains the hub where the open models, datasets, and demos live. The scale of that funding is the tell: whoever serves the picks and shovels at the lowest price is suddenly worth a great deal.
This layer matters more than its invisibility suggests. It is where the race to the bottom on price per image and per second actually happens, which is precisely why value keeps migrating away from it, toward the products above and the audiences beyond. The inference layer sells the picks and shovels, and like every picks-and-shovels business it is essential, high-volume, and structurally low-margin.
The compute layer that serves every model
Every model in this guide has to run somewhere. A lab can ship the best image or video weights in the world, but a creator hitting generate needs those weights loaded on a GPU, warmed up, and returning frames in seconds, not minutes. Inference and hosting is the layer that does this: it takes open and closed models and turns them into an API endpoint that scales, stays up, and bills by the call. It is the picks-and-shovels business of generative media, structurally low-margin because the underlying GPUs are a commodity and customers can switch providers with a config change. That did not stop it from drawing some of the largest 2026 funding rounds in the space, as investors bet that whoever owns the serving layer sits in the toll booth for the entire creative economy.
fal is the developer favorite for generative media, having narrowed its focus to fast inference on image, video, and audio models rather than trying to serve every large language model on earth. That specialization shows up as lower latency and a catalog tuned to what creative teams actually call. Runware competes on the same axis from the low end, offering fast image inference at aggressive prices, while Prodia plays the cheap-and-fast image API role for high-volume, cost-sensitive workloads. The pattern across all three is the same: media inference is a different problem from text inference, and the providers who win it optimize hard for throughput per dollar.
Baseten and Together AI sit closer to production infrastructure. Baseten sells autoscaling inference for teams that need reliability and predictable behavior under load, the unglamorous work of keeping an endpoint healthy when traffic spikes. Together AI runs both inference and training for open models, positioning itself as a full cloud for teams building on open weights rather than a single-purpose API. Both raised heavily in 2026 on the thesis that serious companies will pay for someone else to run the GPU fleet.
The two anchors of the open ecosystem round out the category. Hugging Face is the hub where open models, datasets, and their inference endpoints live, the default place a new open release gets published and pulled from. Replicate made running and fine-tuning open models as simple as an API call and was acquired by Cloudflare, folding model hosting into a large edge network. Alongside them sit the GPU-supply plays, GMI Cloud and Inference.ai as marketplaces and clouds for raw compute, and Oxen.ai handling the data and model versioning that ML teams need underneath all of it.