All articles
Inference27 May 2026 · 7 min read · Nevastack Engineering

Evaluating a private LLM before launch: what to test

Launching a private LLM deployment without a structured evaluation plan usually means discovering quality and safety problems from production users instead of from a test suite. This post outlines the evaluation dimensions we consider necessary before a private LLM goes live, beyond the informal spot-checking that most teams start with.

Build a representative test set, not a generic one

Public benchmarks measure general capability but rarely match the specific distribution of inputs a deployed system will see. Collect real or realistic examples of the queries, documents, and edge cases your users will actually produce, and use that set as the primary evaluation basis.

Include difficult and ambiguous cases deliberately, since these are where models most often fail in ways that matter to users. A test set composed only of easy, clear-cut examples will pass evaluation and still disappoint in production.

Accuracy and grounding for RAG systems

For retrieval-augmented systems, evaluate accuracy and grounding separately. Accuracy measures whether the final answer is correct; grounding measures whether the answer is actually supported by the retrieved context rather than the model's own prior knowledge, which may be outdated or wrong for your domain.

A model can produce a correct answer that happens to align with its training data while ignoring the retrieved context entirely, which will fail the moment your documents diverge from general knowledge. Test with documents that intentionally contradict common knowledge to catch this failure mode.

Safety and refusal behavior

Test how the model handles prompts designed to extract system instructions, bypass content restrictions, or produce outputs outside the intended scope of the application. Private deployment does not remove the need for this testing; it shifts responsibility for it entirely onto the deploying team.

Define acceptable refusal behavior explicitly. A model that refuses too aggressively frustrates legitimate use, while one that refuses too rarely creates risk. Tune this balance against your specific use case rather than accepting a model's default behavior.

  • Test prompt injection attempts specific to your application's tool or retrieval integrations
  • Check behavior on out-of-scope requests unrelated to the application's purpose
  • Verify the model does not leak system prompts or internal configuration

Latency and cost under realistic load

Evaluate latency and cost at the concurrency level expected in production, not a single-request test. Batch behavior, queueing, and GPU memory pressure all change system behavior under load in ways a single test call cannot reveal.

Measure cost per resolved task, not only cost per token, since a model that requires several follow-up turns to reach a correct answer may cost more overall than a slightly more expensive model that answers correctly on the first attempt.

A go-live gate, not a one-time check

Treat evaluation as a gate that must pass before each significant change, including model updates, prompt changes, and retrieval pipeline changes, not only before the initial launch. Automating the evaluation suite so it can run against any candidate change keeps this discipline sustainable.

Set explicit pass thresholds for each dimension before running the evaluation, so the decision to launch is based on a predefined bar rather than a subjective read of results after the fact.

Want this applied to your data?

We scope private AI projects in one call and start with a pilot you can evaluate.

Make room for your next idea.

Bring your data, models, and compute into one workspace. Start with a project and build from there.