devtools.codes

Why did my prompt stop working?

Your tool input is processed locally in your browser and is not intentionally uploaded to our servers. Advertising and analytics providers may still process normal page, device, cookie and network information.

Prompts degrade quietly. There is rarely a commit to point at, because the text often lives in a database row, a spreadsheet, a config file edited in production or somebody's notes. By the time the output is visibly worse, several people may have touched it and none of them changed anything they considered significant. The first useful step is almost always mechanical: put the version that worked next to the version that does not, and look at what actually differs.

Character-level comparison matters more here than it does in code review, because the edits that break prompts are frequently invisible in a rendered view. A curly quote pasted from a document in place of a straight one. A non-breaking space from a web page. A trailing newline removed. An em dash where a hyphen was. None of these change what a human reads, and all of them change the token sequence the model receives — which is the only thing it responds to.

Structural changes are the other common cause, and they are easier to spot once you are looking. Reordering instructions changes their weight, because material near the start and end of a long prompt carries more influence than material buried in the middle. Deleting an example removes a constraint you may not have realised the example was enforcing. Adding a new rule can contradict an older one further up, and the model will not tell you it noticed the contradiction — it will simply pick one.

Not every difference is worth chasing. Whitespace changes inside a paragraph rarely matter. Rewording a sentence to mean the same thing usually does not. What repays attention is anything that changes the structure, the order, the examples or the explicit constraints — and any change to the very beginning or the very end. Diff first, then reason about the subset that could plausibly matter, rather than reading the whole prompt again from the top.

If neither version differs meaningfully, the prompt is probably not the variable. Model versions get deprecated and silently reassigned, providers adjust defaults, temperature and sampling settings drift between environments, and retrieval steps start returning different context. Ruling the prompt out is genuinely useful: it is the cheapest thing to check and the most commonly blamed.

More on prompt changes and model behaviour

Can invisible characters really change a model's output?

Yes. Tokenisers operate on the exact byte sequence, so a non-breaking space, a zero-width joiner or a curly quote produces a different token from its plain equivalent. That shifts the tokenisation of the surrounding text as well. The effect is usually small, but on a prompt already near a decision boundary it is enough to flip behaviour.

Does the order of instructions matter?

It does, and more than most people expect. Attention is not uniform across a long context: instructions near the beginning and the end carry more weight than those in the middle, an effect widely reported as "lost in the middle". Moving a constraint into the centre of a long prompt can weaken it without deleting a word.

Should I version-control my prompts?

Yes, and the reason is this page. A prompt is production logic, and treating it as content rather than code is why the failing version and the working version are so often both unavailable. Plain files in the repository beside the code that sends them is enough — the mechanism matters far less than having any history at all.

How much of a prompt change is too much to review at once?

If the diff is large enough that you cannot hold the changes in your head, the honest answer is that you cannot attribute a behaviour change to any one of them. Change prompts in small increments and evaluate between them, for the same reason you would not ship fifty code changes and then bisect by reading.

Related