Self-hosted LLM & MLOps — DevOps × AI POC
The angle. Cross my DevOps foundation with AI: serve language models privately, without depending on a third-party API, keeping control of data and cost.
What I prototype
- Serving open models via Ollama / vLLM, exposed as an OpenAI-compatible API.
- Deployment on Kubernetes/k3s, GPU management and scaling.
- Inference observability (latency, tokens/s, cost per request).
- Caching, quantization and cost/performance trade-offs.
Stack
Ollama · vLLM · Kubernetes · Docker · GPU · API-first
POC — AI/MLOps test bed, not deployed to production.