No universal stars
A single number hides task fit, environment, trade-offs, and failure modes.
The method makes disagreement useful. You should be able to rerun the work, find the assumption, and see where your result diverges.
Record the tool, version, release channel, environment, date, region, plan or price context, task, input, and expected result. Separate vendor claims from the test hypothesis.
Capture commands, configuration, prompts where relevant, elapsed time, direct cost, retries, outputs, and observable behavior. Redact secrets and personal data before publication.
Test a relevant constraint, malformed input, partial failure, timeout, cancellation, permission boundary, or recovery path. The point is to find the edge of the claim.
Publish observed facts, supported interpretation, and the recommendation as distinct parts. Name the editor, technical reviewer, conflicts, review copies, affiliate relationships, common ownership, and limitations beside the verdict. When no independent reviewer exists, say so plainly instead of implying one.
Provide the smallest safe fixture and instructions another developer can run. If the real system cannot be shared, document the missing layer and reduce the strength of the claim.
A tool test is a snapshot, not a permanent ranking. Material product changes trigger a retest, a visible stale notice, or retirement. Corrections preserve the original claim and explain what changed.
A single number hides task fit, environment, trade-offs, and failure modes.
Sponsorship, access, affiliate links, and common ownership are disclosed and never determine the result.
Material edits receive a date and reason. A challenged result is checked against the saved rig.
A versioned product can outgrow a test. Stale evidence is labeled, rerun, or retired.