Blog

Notes from running private AI in production.

Method choices, GPU economics, retrieval design and the compliance detail that decides architectures. Written by the people on call.

Inference5 Aug 2026 · 8 min read

Open-source model selection in 2026: a practical framework

How to evaluate open-weight LLMs for production on license, context length, tooling, and hosting cost rather than benchmark scores.

Read article
Compliance7 min

Keeping models private: isolation, encryption, and deletion

Practical controls for protecting proprietary models and fine-tuned weights: tenant isolation, encryption at rest and in transit, and verifiable deletion.

22 Jul 2026
Infrastructure7 min

Observability for LLM endpoints: what actually matters

Metrics, logging, and tracing practices for production LLM endpoints that go beyond basic uptime monitoring.

8 Jul 2026
Infrastructure7 min

Pilot to production: a checklist for AI deployments

The gaps that typically block an AI pilot from reaching production, covering data, security, cost, and operational ownership.

24 Jun 2026
RAG8 min

Chunking and embedding strategies that actually improve RAG

Practical guidance on chunk size, overlap, and embedding model choice for retrieval quality, with tradeoffs explained.

10 Jun 2026
Inference7 min

Evaluating a private LLM before launch: what to test

A structured evaluation approach for private LLM deployments covering accuracy, safety, latency, and cost before go-live.

27 May 2026
Infrastructure7 min

Bare metal vs virtualised GPU: choosing the right layer

When dedicated bare metal GPU access is worth the operational overhead compared to virtualised or shared GPU instances.

13 May 2026
Cost8 min

Cutting inference cost per token without cutting quality

Practical levers for reducing LLM inference cost: quantization, batching, caching, and model routing, with realistic tradeoffs.

29 Apr 2026
Compliance8 min

EU data residency and GDPR for AI workloads: a practical guide

What GDPR and EU data residency actually require for RAG and LLM deployments, and how infrastructure choices affect compliance.

15 Apr 2026
Infrastructure7 min

Choosing a GPU: L4, A100, or H100 for your workload

A practical comparison of L4, A100, and H100 GPUs for inference, fine-tuning, and training, with guidance on when each fits.

1 Apr 2026
Fine-tuning8 min

LoRA fine-tuning in practice: what works and what does not

Practical guidance on LoRA fine-tuning, covering rank selection, target modules, dataset size, and common failure modes.

18 Mar 2026
RAG9 min

Building a production RAG pipeline: the components that matter

A component-by-component walkthrough of a production-grade RAG pipeline, from ingestion through generation and monitoring.

4 Mar 2026
RAG8 min

RAG vs fine-tuning: a decision guide for enterprise AI

How to decide between retrieval-augmented generation and fine-tuning based on data volatility, cost, and the type of knowledge involved.

18 Feb 2026

Make room for your next idea.

Bring your data, models, and compute into one workspace. Start with a project and build from there.