Compare AI tools using the same job and review standard
Use the same permitted task pack and acceptance checks for each tool. Record the actual setup and the effort needed to produce a checked working output, then describe the limits of what the small test establishes.
Write the acceptance checks first
Choose a recurring task whose result a reviewer understands. Write down what must be correct before looking at either tool’s answer. Include the important facts, unresolved points and required output format. This prevents an attractive response from changing the standard halfway through the comparison.
Use tools approved for the information involved. Record the account type, available features, test date and any relevant configuration visible to the user. This guide does not rank current products. It provides a way to evaluate the setups your team can actually use on a specific job.
Score observable results
Use a simple row for each acceptance check: met, partly met or missed, with a source-linked explanation. An unsupported conclusion is a meaningful failure even if the rest of the prose is excellent. Keep critical errors visible instead of allowing them to disappear inside an average score.
Measure the work the person actually did: preparation, interaction, verification and correction. Record whether the output opened correctly in its destination. Where practical, have the reviewer assess labelled outputs without seeing the tool names first. This can reduce the effect of an existing brand preference on the content review.
Editorial guidance informed by Acuity's work with teams. Examples are fictional and client materials stay private. To record who owns an AI use and when it is reviewed, see AI Register.
