Running models on hardware you own, on a network you control. No API keys, no per-token billing, no data leaving the building. These guides cover the builds, the tuning, and the limits of what consumer GPUs can actually do.
70B models need around 48 GB of VRAM. Every route to that compared: two used cards, unified memory, CPU offload, harder quantisation, or renting.
Which graphics card should you buy to run AI models locally? A guide organised by VRAM tier rather than price, covering NVIDIA, AMD and Intel, new and used.
Build a headless Ubuntu LLM inference server with Ollama and Open WebUI on an 8 GB NVIDIA GPU: VRAM budgeting, nginx, and a test checklist.
Tokens, parameters, quantisation, KV cache, CUDA and context windows: the vocabulary of running LLMs locally, explained through one worked example.
What a language model really is, why graphics cards turned out to run them, and how one 2017 research paper set off everything that followed. No maths required.