Notes from running private AI in production.
Method choices, GPU economics, retrieval design and the compliance detail that decides architectures. Written by the people on call.
Open-source model selection in 2026: a practical framework
How to evaluate open-weight LLMs for production on license, context length, tooling, and hosting cost rather than benchmark scores.
Read articleKeeping models private: isolation, encryption, and deletion
Practical controls for protecting proprietary models and fine-tuned weights: tenant isolation, encryption at rest and in transit, and verifiable deletion.
Observability for LLM endpoints: what actually matters
Metrics, logging, and tracing practices for production LLM endpoints that go beyond basic uptime monitoring.
Pilot to production: a checklist for AI deployments
The gaps that typically block an AI pilot from reaching production, covering data, security, cost, and operational ownership.
Chunking and embedding strategies that actually improve RAG
Practical guidance on chunk size, overlap, and embedding model choice for retrieval quality, with tradeoffs explained.
Evaluating a private LLM before launch: what to test
A structured evaluation approach for private LLM deployments covering accuracy, safety, latency, and cost before go-live.
Bare metal vs virtualised GPU: choosing the right layer
When dedicated bare metal GPU access is worth the operational overhead compared to virtualised or shared GPU instances.
Cutting inference cost per token without cutting quality
Practical levers for reducing LLM inference cost: quantization, batching, caching, and model routing, with realistic tradeoffs.
EU data residency and GDPR for AI workloads: a practical guide
What GDPR and EU data residency actually require for RAG and LLM deployments, and how infrastructure choices affect compliance.
Choosing a GPU: L4, A100, or H100 for your workload
A practical comparison of L4, A100, and H100 GPUs for inference, fine-tuning, and training, with guidance on when each fits.
LoRA fine-tuning in practice: what works and what does not
Practical guidance on LoRA fine-tuning, covering rank selection, target modules, dataset size, and common failure modes.
Building a production RAG pipeline: the components that matter
A component-by-component walkthrough of a production-grade RAG pipeline, from ingestion through generation and monitoring.
RAG vs fine-tuning: a decision guide for enterprise AI
How to decide between retrieval-augmented generation and fine-tuning based on data volatility, cost, and the type of knowledge involved.
Make room for your next idea.
Bring your data, models, and compute into one workspace. Start with a project and build from there.