Skip to main content

Pricing

When using Modal Labs (the default cloud provider), compute is billed per second of GPU usage. Relay automatically selects the cheapest GPU that meets the model’s VRAM requirements.

GPU pricing

Cost estimation

Before running a model, Relay estimates the compute cost:
The estimate appears in the job confirmation dialog so you know the cost before committing.

Typical costs per model

Per-user limits

These limits can be adjusted by workspace administrators.

Other providers

  • Local Docker — free (uses your own hardware)
  • HPC/SLURM — costs determined by your institution
  • User GPU Server — costs determined by your infrastructure
  • HuggingFace Spaces — free for public spaces, paid for private