Once I started treating prompts like code, prompt bugs got easier to debug. I could compare versions and experiments, then roll back a bad change.
Every prompt change in the offer-bundling assistant goes through the same five steps. I care about running them every time, even when the prompt that comes out isn't perfect.
In 30 seconds
- Fixing a pricing bug in one prompt once removed a compliance clause by accident.
- Use a lifecycle: ideate -> prototype -> evaluate -> deploy -> monitor.
- Version prompts and run evals before every release.
The problem with prompt chaos
If you change prompts in a rush, you lose track of what improved the system and what broke it. In the offer-bundling assistant, I once fixed a pricing bug and accidentally removed a compliance clause. That happened because we had no lifecycle.
The lifecycle I use
-
Ideate
- Write the intent in one sentence.
- Decide what the output must include.
-
Prototype
- Try 3 to 5 variants quickly.
- Keep only the best candidate.
-
Evaluate
- Run against the golden dataset.
- Compare metrics to the current version.
- Most prompt variants fail, and that’s expected.
-
Deploy
- Ship the new prompt with a version tag.
- Store it in a prompt registry or file with a unique ID.
-
Monitor
- Watch traces and online evals for drift.
- Roll back if metrics regress.
Versioning prompts (minimal version)
I store prompts like code. Each version has:
- An ID (for example:
bundle-prompt-v3) - A short changelog
- The eval results that justified the change
Even a simple folder structure works:
/prompts
bundle_prompt_v1.txt
bundle_prompt_v2.txt
bundle_prompt_v3.txtEvals as gates
If a prompt change does not pass your evals, it should not ship. In LangSmith, I run a small golden set and compare:
- accuracy of SKU selection
- missing compliance clauses
- price threshold violations
A regression on any of them blocks the release.
Rollback strategy
Prompt changes should be reversible in minutes. I keep:
- the previous prompt in version control
- a rollback toggle (just a config flag)
This saved me more than once.
Why the evals live next to the prompts
In the offer-bundling assistant, each prompt change can shift the structure of the output or the content. That is why each prompt is kept with its schema and evals. Changing one means checking the other two.