Scheduling every job on a GPU that can only hold one model - DEV Community
I run this homelab on two small boxes. One never turns off and just serves DNS and background services; the other has a GPU with a memory pool that can hold exactly one loaded model at a time, nothing more. This post covers the queue and gateway that constraint forced me to build, and the network shape wrapped around both boxes.