Make the platform do a task your reader actually needs

To compare AI platforms for your book or method, give each the same task, source material, and success criteria. Test missing information and corrections as well as an easy example. Then check the complete reader journey, current costs, and what you can take with you if you leave.

A demonstration is useful for understanding the interface. It is not enough evidence that the product will represent your judgment in situations the presenter did not prepare.

You do not need a technical benchmark to begin. You need a short, repeatable test that exposes the differences that matter to your offer. The worksheet below is an original evaluation design; no platforms were run or ranked for this article.

Write the expected behavior before seeing the answer

Choose one task from your method that a reader might bring this week. Describe the starting information, the finished result, and the mistakes that would make the answer unacceptable.

Suppose your book teaches better editorial briefs. A plausible task is turning a request for a newsletter article into a brief that identifies its audience, purpose, evidence, and constraints. A polished outline that invents a customer statistic should fail even if it sounds like you.

Anthropic’s evaluation guidance distinguishes a task with defined inputs and success criteria from the attempts made at that task. It also distinguishes an agent’s account of what happened from the actual outcome. Those are useful principles for an author evaluating a workflow, not just its final prose. Anthropic’s evaluation guide.

For your test, write down what a good answer must preserve, what it should ask, and what it must not invent. Keep those criteria stable while comparing candidates.

Bring a small set of realistic cases

Use a mix of ordinary work and common difficulties. Testing only extreme traps tells you little about everyday usefulness; testing only the easiest example hides the limits.

CaseWhat you supplyWhat you want to observe
Ordinary requestA clear task with enough contextCompletes the intended work without unnecessary detours
Missing factOmit a detail your method requiresAsks or marks the gap instead of inventing it
Weak assumptionInclude an unsupported conclusionChallenges it using the method’s criteria
Reasonable alternativeOffer an answer unlike your exampleRecognizes a valid approach rather than forcing a template
CorrectionChange one important fact mid-conversationRevises the relevant parts consistently
BoundaryAsk for a decision outside the method’s scopeExplains the limit and gives an appropriate next step

The cases should come from your work, with identifying details removed and permission where needed. Use fictional material when real examples contain information you should not upload.

In a retrieval-testing discussion, u/Glad-Win1983 described using:

“a curated set of (often tricky) queries”

The post explains comparing changes against a stored baseline. It is an individual technical account, not a platform recommendation. The author-level lesson is to keep your cases so that you can compare later versions rather than rely on memory. Original discussion.

A test script you can take to a demo

Here is a fictional editorial-brief case. Give the same text to each candidate after supplying your own method and examples:

Help me prepare a brief for a 700-word newsletter article for independent shop owners. The topic is reducing abandoned online orders. I have three customer interviews but no measured conversion data. The article should help a reader choose one checkout problem to investigate. Ask for anything else you need before completing the brief.

The desired behavior is to preserve the audience, length, limited evidence, and investigation goal. The assistant should ask what the interviews actually contain and avoid claiming that a suggested change will increase sales by a particular amount.

Then add a correction:

The interviews were with people who completed an order, not people who abandoned one. Revise the brief so it does not misrepresent that evidence.

Check whether the revision changes the reasoning, not merely one sentence. The final brief should acknowledge the evidence gap and avoid attributing abandoned-order motives to people who were not interviewed.

Finally ask for an export in the format you expect readers to use. Confirm that the result remains readable outside the chat. This is a suggested exercise, not a claim that a particular platform has passed it.

Score usefulness separately from polish

Use three simple labels: passes, needs repair, or fails. For every label, keep the sentence or behavior that justified it.

DimensionA passing observation
MethodFollows the distinctions that matter in your procedure
EvidencePreserves source facts and marks missing support
InteractionAsks useful questions without repeatedly requesting known information
CorrectionUpdates all affected parts after a changed fact
OutputProduces a result the reader can use in the intended setting
Reader effortSetup and recovery are understandable for the intended audience

Do not let several attractive answers cancel one serious fabrication. Mark the disqualifying behavior separately and ask whether it can be addressed, how, and with what evidence.

Run more than one attempt where practical. A single result can reveal a problem, but it cannot establish consistent reliability. Record the date, model or configuration shown, supplied materials, and any help the presenter gave.

Ask how the product improves after a failure

A platform should have an understandable way to revise the experience and check the revision. Coachvox, for example, documents a testing space where authors can interact with their AI, rate responses, edit answers, and export the resulting history. This is a documented workflow, not a guarantee that editing one answer fixes every related case. Coachvox’s testing documentation.

Ask the same questions of each provider: can you inspect the relevant response, change the underlying guidance, start a fresh test, and keep a record of what changed? Are you evaluating the ordinary reader experience or a special demonstration setup?

Public leaderboards can be informative, but your task still needs its own check. In an October 2025 discussion, u/Sad-Boysenberry8140 described public evaluation datasets as useful for basic checks but said:

“they don't necessarily reflect my actual use case.”

That post concerned enterprise document retrieval. It is an adjacent practitioner’s concern, not evidence about book products, but the mismatch is easy to recognize: success on someone else’s questions does not establish success on yours. Original post.

Check the offer around the AI

Record current costs with a date and their source. Separate your platform fee, reader account requirements, usage charges, payment fees, and support commitments. If a trial differs from the paid plan, record the difference rather than assuming it will carry over.

Walk through how a reader obtains access, starts the task, returns later, and gets help. Ask what happens when they use the wrong account or stop paying. Do not test with real customer accounts or make billing changes merely to explore a demo; use the provider’s supported test setup.

Then inspect the exit path. Which source files, examples, conversation records, and customer records can you export? In what format? What stops working when the subscription ends? A successful export of one item does not prove that every part of the product is portable.

Ask for current documentation where the answer affects your offer. Keep unknowns visible in the comparison. A sales assurance and a behavior you observed should occupy different columns.

Choose the candidate whose limits you understand

Your comparison record should contain the cases, outputs, judgments, dated costs, reader requirements, and unresolved questions. That is enough to make a more grounded next decision than choosing the most impressive demonstration.

For deeper evaluation after you build, see our guide to measuring agent skill quality.

To discuss testing your established method as a skill, visit Skillfully and choose Book onboarding. Bring the task you want readers to complete and one example that an adequate product must handle.