Pricing
When using Modal Labs (the default cloud provider), compute is billed per second of GPU usage. Relay automatically selects the cheapest GPU that meets the model’s VRAM requirements.GPU pricing
Cost estimation
Before running a model, Relay estimates the compute cost:Typical costs per model
Per-user limits
These limits can be adjusted by workspace administrators.
Other providers
- Local Docker — free (uses your own hardware)
- HPC/SLURM — costs determined by your institution
- User GPU Server — costs determined by your infrastructure
- HuggingFace Spaces — free for public spaces, paid for private