A coding assistant can produce a convincing patch before you know whether it solved the problem. For a useful comparison, make the unit of work a finished, reviewed change. This guide gives you a small evaluation you can repeat with different tools. It does not name a winning product or claim measured time savings.
1. Choose a task with an observable failure
Pick a bug you understand in a repository you are allowed to use. A form accepting whitespace-only input is a better first trial than “improve the architecture.” Write the reproduction steps, the current result, and the expected result. Include a nearby case that must keep working, such as a valid input with surrounding spaces. Decide these cases before asking the assistant to write tests.
Use a separate working copy for each trial and record the starting revision. Run the existing checks first. If the baseline already fails, record that failure separately so you do not credit or blame the assistant for it. Keep credentials and unrelated private files out of the trial.
2. Give every tool the same brief
Provide the repository instructions, reproduction steps, and acceptance cases. Ask for an explanation before edits. This reveals whether the assistant found the actual validation path or merely guessed from a filename. Avoid supplying a detailed implementation to one tool and an ambiguous request to another.
Investigate why this form accepts whitespace-only input. First identify the validation path and explain the cause. Make the smallest fix that rejects empty or whitespace-only input while preserving valid input. Add a regression check, run the relevant existing checks, and list anything you could not verify. Do not refactor unrelated code.
3. Keep a trial log
- Setup: repository revision, tool version, selected model, account plan, and task date.
- Time: minutes to a proposed fix, minutes reviewing it, and minutes correcting it.
- Interventions: every extra instruction you gave and every edit you made yourself.
- Evidence: original failure, final check output, and the reviewed diff.
Start the timer when you give the task and stop after review, not when the tool says it is done. Record setup time separately if it is a one-time cost. If a usage limit interrupts the run, keep that result in the log: it affects whether the workflow fits your normal workday.
4. Check the patch independently
Read every changed file. Look for deleted assertions, broad exception handling, unexpected dependency changes, or a fix that only handles the supplied example. Run your acceptance cases yourself. Where practical, apply the new regression test to the original code and confirm that it detects the defect. A passing test that never fails on the old behavior gives weak evidence.
For a user interface change, exercise the actual form as well as its validation function. Check the error message, keyboard submission, and recovery after correcting the value. A function can be correct while the page still submits stale state. Record a screenshot or short reproduction note when that helps another reviewer verify the result.
5. Decide using the costs you actually observed
Use four judgments: did it find the cause, did it preserve the surrounding behavior, could you understand the patch, and how much work remained? Keep correctness as a prerequisite rather than averaging it away with speed. If two trials both succeed, compare their total completion time and the number of interventions.
One bug is a screening exercise, not a general ranking. Repeat with another task that represents your workload before paying for a team-wide rollout. A tool that helps with isolated edits may behave differently on unfamiliar code or multi-file changes. Keep the evidence and describe those limits when sharing your conclusion.
Where to start
See our Cursor trial checklist for an example brief. Cursor’s official Agent documentation describes its workflow. Check the documentation and current plan limits of whichever product you test; this exercise assumes no particular subscription or benchmark score.