ollama

Running Local LLMs on Unraid

The first local model I ran took about 45 seconds to respond to a simple question. I almost gave up right there. The model was too big for my VRAM, the quantization was wrong, and I hadn’t understood yet that “running a local LLM” and “running a usable local LLM” are meaningfully different things.

I run Ollama inside a Docker container on Unraid. The container gets passed through to an NVIDIA GPU, which handles all the inference. The host is a Ryzen 9 5900X on an ASUS TUF Gaming X570-Plus (Wi-Fi) board. Getting GPU passthrough working was the first real obstacle. Unraid handles it through its Docker GPU assignment settings, but you have to make sure the NVIDIA container toolkit is installed and the container is configured to use it. There are Unraid forum threads on this; I won’t reproduce the whole process here, but it took me longer than I expected because the error messages weren’t obvious about what was actually wrong.

The hardware question matters more than the model question, at first. Most consumer GPUs in the 8-12GB VRAM range can run 7B or 8B parameter models fully in memory and respond in a few seconds. That’s where I’d start. The models in that range, Llama 3, Mistral, Gemma, Qwen, have gotten good enough over the past year that the quality is genuinely useful for most everyday tasks. Quantized versions, typically Q4 or Q5, cut memory requirements significantly with a modest quality tradeoff. For anything that’s not critical reasoning, the tradeoff is worth it.

What I actually run daily is a quantized Llama 3 8B for quick, low-stakes queries and a mid-size model in the 13-20B range for tasks where quality matters more than speed. The 8B model responds in 2-4 seconds on my hardware, which feels fast enough to be comfortable. The larger model takes 15-25 seconds per response, which I find acceptable for deliberate tasks but too slow for conversational back-and-forth.

The models I don’t run locally are the ones where I genuinely need top-tier reasoning or writing quality. For those I send API calls to Anthropic. The local models are good for information retrieval, summarization, drafting, code explanation, and routine agent tasks. They’re not good enough to replace a frontier model for complex multi-step reasoning or polished writing that goes on the internet. Knowing that line has saved me a lot of frustration.

Unraid’s container management makes it fairly easy to keep Ollama updated and to pull new models. The Ollama API makes it straightforward to switch which model a request goes to. The practical challenge is storage: model weights pile up fast. A handful of models can eat 50-100GB without trying hard. I keep the ones I actually use and delete the rest. The Unraid array handles the storage; I’ve dedicated a Samsung 860 EVO 1TB as a cache SSD specifically for models so loading times stay fast.

One thing I wish I’d known earlier: the model name matters, but the quantization label matters almost as much. “Llama 3 8B” can mean very different performance depending on whether it’s Q2, Q4, Q5, or Q8. I run Q4 as a default, Q5 when I want a bit more quality on a task that warrants it. Q8 is overkill for most things and Q2 degrades quality more than I find acceptable.

If you’re curious about specific model recommendations for your VRAM budget, feel free to drop the specs in a comment and I’ll give you my current take.

Hardware linked in this post:


Affiliate disclosure: Some links in this post are Amazon affiliate links. If you buy through them, I get a small commission at no cost to you. It helps keep the lights on here.

2026-06-25T14:29:17-07:00July 9th, 2026|Categories: Blog|Tags: , , , , , , , , , |0 Comments

The Self-Hosted AI Stack I’d Build If I Were Starting Over

I took a winding road to get my current AI homelab working. I’d make different choices if I were starting from scratch, and most of them would come down to doing less sooner rather than more.

The first thing I’d do is separate the model serving layer from everything else. Ollama as a standalone container, exposed on a private network, nothing else bundled in. A lot of guides will tell you to start with a full OpenWebUI stack, and OpenWebUI is fine, but it creates a coupling that makes things harder to reason about later. If your UI and your model server are the same deployment, you end up with friction when you want to swap one out or add a second frontend. Keep them separate from the start.

For hardware, I’d be more honest with myself about the model size tradeoff. My current build is a Ryzen 9 5900X on an ASUS TUF Gaming X570-Plus (Wi-Fi) board, and it handles everything I throw at it for inference routing and container management. A 7B parameter model runs well on consumer GPU memory, responds quickly, and handles most practical tasks. I spent too long trying to run 34B models on hardware that wasn’t really right for them, getting slow responses, and convincing myself the capability justified the latency. It usually didn’t. For day-to-day assistant work, a well-quantized 7B or 8B model is more useful than a sluggish 34B. Save the bigger models for tasks where reasoning quality actually matters.

The gateway layer is where I’d invest more early effort. This is the piece that connects LLM inference to real tools: file system access, APIs, shell commands, memory. I’m running OpenClaw for this. If I were starting fresh, I’d still choose a purpose-built gateway over trying to wire this together myself with n8n or LangChain. The operational overhead of maintaining custom orchestration code is real. A gateway that’s designed to manage agent lifecycles, credential handling, and tool permissions out of the box is worth the setup time.

Memory is something I’d take seriously from day one. The difference between an AI that knows the state of your environment and one that starts fresh every session is enormous in practice. That means deciding early on where state lives, how agents read and write it, and what format it’s in. Markdown files on a shared volume have worked well for me: human-readable, easy to edit when something’s wrong, git-friendly if you want version history.

For API keys and credentials, I’d use a secrets directory with tight permissions from the start rather than environment variables scattered across docker-compose files. It’s easier to audit, easier to rotate, and easier to scope to specific containers when something needs to change. This sounds like overkill when you’re standing up one container. It pays off when you have eight.

The thing I’d skip entirely on a first build is trying to run everything locally. Ollama handles local inference well. But for tasks that genuinely need a frontier model, the cost of API calls is low and the capability gap is large enough to matter. Don’t try to replace Claude with a local model for complex reasoning. Use local models where they’re good enough and cloud APIs where they’re not. That hybrid approach is cheaper and more capable than either extreme.

Finally, I’d document my container layout before it gets complicated. Which container serves which purpose, which ports are mapped, what credentials it needs. This sounds tedious and it is. Three months later when you’re trying to figure out why something stopped working, you’ll be glad you did it.

Hardware linked in this post:


Affiliate disclosure: Some links in this post are Amazon affiliate links. If you buy through them, I get a small commission at no cost to you. It helps keep the lights on here.

2026-06-18T11:46:47-07:00July 3rd, 2026|Categories: Blog|Tags: , , , , , , , , , |0 Comments