Pilot to production: a checklist for AI deployments
Most AI pilots demonstrate feasibility, not production readiness. A prototype that answers questions correctly in a demo often fails on the requirements that only appear once real users, real data volumes, and real security review enter the picture. This checklist captures the gaps we see most often when helping teams move from pilot to production.
Data access and freshness
Pilots frequently run against a static export of documents pulled once and never updated. Production systems need a defined ingestion pipeline that keeps the knowledge base current, with a clear owner responsible for source connections when an upstream system changes its schema or access method.
Confirm who owns each data source in production, what happens when a source becomes unavailable, and how stale content is identified and removed rather than silently persisting in a retrieval index.
Security review that a pilot skipped
Pilots often run with broad permissions and no formal security review because the stakes seemed low during evaluation. Before production, confirm authentication on every endpoint, verify that retrieval respects the same access controls as the source documents, and check that logs do not capture sensitive content unnecessarily.
Document data residency explicitly if the organization has EU data protection obligations. A pilot hosted on a convenient regional endpoint does not automatically satisfy a legal or contractual residency requirement.
- Confirm document-level access control is enforced in retrieval, not only in the source system
- Review what gets logged and for how long
- Verify the hosting region matches contractual and regulatory requirements
Cost at expected scale
A pilot serving a handful of internal users tells you almost nothing about cost at hundreds or thousands of concurrent users. Model a realistic load scenario, including peak usage patterns, and estimate token consumption and GPU hours at that scale before committing to a production budget.
Identify which parts of the cost structure scale linearly with usage and which are fixed. This distinction determines whether growth in usage is a manageable cost increase or a structural problem requiring a different serving architecture.
Operational ownership
A pilot is often owned informally by whoever built it. Production requires a named on-call owner, a runbook for common failure modes, and an escalation path when the model provider or infrastructure vendor has an incident.
Decide in advance how model updates will be evaluated and rolled out. Unannounced upstream model changes, whether from an API provider or an internal retraining schedule, should not reach production users without a review step.
Failure modes and fallback behavior
Define what the system does when retrieval returns nothing relevant, when the model times out, or when confidence in an answer is low. A pilot can get away with an occasional bad answer; production users need a defined fallback such as a clear no-answer response or a handoff to a human.
Test these failure paths explicitly before launch rather than discovering them from user complaints. Building a small set of adversarial and edge-case prompts as part of the release process catches most of these gaps early.
Want this applied to your data?
We scope private AI projects in one call and start with a pilot you can evaluate.