Run an inference server yourself for low latency, on-prem data, or offline use; use a hosted API otherwise. Start one locally with Docker.