S> initializing
SLAEGA·19
poc-prototypedevops-infrastructureapi-webservice

Self-hosted LLM & MLOps — private inference

Mar 20261 min read
Self-hosted LLM &…
poc-prototype
devops-infras…
api-webservice
IA
LLM
Ollama
vLLM
MLOps
Kubernetes
GPU
DevOps
POC
drag

Self-hosted LLM & MLOps — DevOps × AI POC

The angle. Cross my DevOps foundation with AI: serve language models privately, without depending on a third-party API, keeping control of data and cost.

What I prototype

  • Serving open models via Ollama / vLLM, exposed as an OpenAI-compatible API.
  • Deployment on Kubernetes/k3s, GPU management and scaling.
  • Inference observability (latency, tokens/s, cost per request).
  • Caching, quantization and cost/performance trade-offs.

Stack

Ollama · vLLM · Kubernetes · Docker · GPU · API-first

POC — AI/MLOps test bed, not deployed to production.

Technologies

Project stack

IALLMOllamavLLMMLOpsKubernetesGPUDevOpsPOC

Also explore

Similar projects