Self-hosting

·

Running models on your own hardware: VRAM math, quantization trade-offs, and when local actually beats the API.