Ask for the criteria before the verdict, and the verdict has to live with them

A verdict on its own cannot be examined
Ask whether a plan is any good and you get an assessment. It will be fluent, it will mention some strengths and some risks, and there is almost nothing you can do with it, because you have no idea what standard it was measured against. Agreeing or disagreeing with it is the only available response, and neither is informative.
The problem is not that the assessment is wrong. It is that it is unexaminable, and an unexaminable judgement is a poor input to a decision you are accountable for.
The fix is a matter of ordering rather than of wording. Get the standard written down first, look at it, and only then ask for the item to be measured against it.
Criteria first, and preferably before the item is in view
Ask what would make a plan of this kind good, what the common weaknesses are, and what evidence would distinguish a strong version from a weak one — as a first, separate request, before the specific plan is supplied. Criteria written without the candidate in front of them tend to be more general and less shaped to flatter or to condemn what happens to be there.
Then supply the plan and ask for it to be assessed against those criteria, one at a time, with a stated position on each. The output stops being a general impression and becomes a set of specific claims, and specific claims can be checked.
This is also the point at which disagreements become productive. If you think criterion four is irrelevant to your situation, you can strike it and the assessment changes accordingly, which is not something you can do to a paragraph of overall opinion.
The criteria are yours to reject
Half the value here is in reading the criteria and finding them wrong. A list that omits the constraint that dominates your situation tells you the assessment that follows will be about a different problem, and you have learned that before spending any time on the verdict rather than after.
Edit them, then. Delete the ones that do not apply, add the two that matter most in your case, and say which are decisive and which are secondary. A weighted, edited list is a description of what you actually care about, and it is worth keeping for the next item of the same kind.
Over time these lists become the more valuable artefact. The assessments are disposable; the standard you use to judge a category of thing is reusable, improves with each round, and can be handed to a colleague.
This is not a window into anything
Be careful about what the arrangement buys. Criteria written before a judgement constrain the text that follows them, which is a real and useful effect on the output. They are not a record of how the judgement was reached, and treating a stated rationale as an account of process is a separate mistake with its own costs.
What you have is a claim you can test. Criterion three says the plan lacks a rollback path; you can look at the plan and see whether it does. The value comes from your ability to check the individual assertions, not from any confidence that the assertions describe an inner method.
So read the criteria as a proposal about what matters and the assessment as a set of testable statements. Both are useful in that form. Neither becomes more reliable because it was presented in the right order.
Where scoring against criteria is the wrong shape
Some judgements do not decompose. Whether a piece of writing is any good, whether a person is right for a role, whether a design has the quality that makes people want to use it — these resist being split into weighted components, and the decomposition can actively mislead by making a thing that scores well on every axis look like a good answer.
Where the question is genuinely holistic, a better request is a comparison. Ranking two or three candidate options against each other forces a position without pretending to a precision that is not there.
And where the criteria are established externally — a standard, a rubric, a specification, a regulation — do not ask for them to be reconstructed. Supply the real ones. A plausible approximation of a published standard is exactly the kind of near-miss that survives review, and the actual document is usually a search away.
Common questions
Why request the criteria before supplying the item?
Criteria written with the candidate in view tend to be shaped by it. Asking first produces a more general standard, and it lets you notice that the standard is wrong before you have read a verdict that would be hard to unread.
Does stating criteria first make the assessment more accurate?
It makes the assessment checkable, which is not the same thing. Each claim can be tested against the item. Nothing about the ordering gives the output access to a method, and a stated rationale should not be read as a record of one.
When is this approach unsuitable?
When the judgement is holistic and resists decomposition, where a forced comparison between candidates works better. And when a real rubric or standard exists, in which case supply it rather than accepting a reconstruction of it.
Deputy editor, Prompt After Prompt
Adrian has written about prompt craft, writing with ai, images & audio for most of the last decade and prefers a plain explanation to a clever one.