All articlesLLM Ops

Prompt Versioning and CI/CD: Treating Prompts Like Code

Prompts change more often than application code. Without versioning, testing, and rollback, every change is a gamble. Here is the system.

Sri Raman17 August 20268 min read
Prompt Versioning and CI/CD: Treating Prompts Like Code

Prompts are code. They change, they regress, and they ship to production. Yet most teams manage prompts the way they managed source code in 2005: edit in production, hope it works, roll back by memory. The result is silent quality regressions that no one notices until a customer complains.

Version every prompt

Every prompt that ships needs a version tag, stored alongside its change history. When you change a prompt, you create a new version, not an overwrite. This gives you the ability to compare versions, roll back, and attribute quality changes to specific prompt edits. Without versioning, 'the output got worse last week' is un-debuggable.

Run evals on every change

Before a new prompt version ships, run it against your eval set and compare to the current version. If any metric regresses beyond a threshold, block the deploy. This is CI for prompts. The eval set does not need to be huge — 20-50 representative inputs — but it must cover the failure modes you have seen in production.

The rollback button

Even with evals, a prompt can regress in production on inputs the eval set did not cover. You need a one-click rollback to the previous version. Not a git revert and redeploy — a config change that swaps the active prompt version without a release. This requires the prompt to be loaded at runtime from a store, not hardcoded in source.

Shadow testing

Before a new prompt version goes live, run it in shadow: the old prompt handles the request, the new prompt runs in parallel, and you compare outputs. If the new version diverges in concerning ways, you do not ship. This catches regressions the eval set missed because it tests on real traffic.

Share this article