Compare the same reader task under controlled conditions
To find out whether your skill adds value beyond pasting a chapter into AI, give both approaches the same reader task, source edition, and starting facts. Define success before running them, retain the full outputs, and have someone review the results without knowing which approach produced each one.
Measure the result your reader needs: a faithful recommendation, a usable artifact, or the right clarification. Keep setup effort and convenience separate from answer quality. A skill might improve one without improving the other.
This article provides a comparison protocol and a fictional worked setup. It does not report a completed benchmark, a live product test, or evidence that skills outperform chapter prompts. You will need access to your own skill and the intended AI environment to run the comparison.
Give the reader's existing workaround a fair test
A prepared prompt can be a serious alternative. In a discussion about skills, u/MoilC8 wrote:
“i already had a notebook with prepared prompts for different tasks and some guidelines i wanted claude to follow, so i used to copy paste those together with the task itself.”
The commenter was asking what changed with skills. Their experience is not an author-product comparison, but the question is fair: what does a new package add to an existing practice? Original discussion.
Do not create an artificially weak baseline by withholding the relevant chapter or asking only “help me.” Use the prompt a reasonable reader could write, with the same task and personal facts you supply to the skill.
If your readers already use a well-developed prompt, include that as a separate comparison condition. Be clear whether you are comparing against casual chapter pasting or an optimized alternative. They answer different questions.
Decide what you want to learn
Write a narrow comparison question: “Does this version of my skill follow my interview-planning method more consistently than a chapter prompt on these reader situations?” That is more testable than “Are skills better?”
Then separate three possible sources of value:
| Question | Evidence to collect | What it does not establish |
|---|---|---|
| Does the output follow the method? | Decisions, preserved facts, boundaries, and usable deliverable | Whether readers will buy it |
| Is the workflow easier to use? | Setup steps, repeated instructions, clarification turns, and reader observations | Better reasoning merely because there are fewer clicks |
| Does the intended skill actually participate? | Available execution record or supported evidence of activation | That its participation improved the answer |
You may care about all three. Report them separately so a convenient workflow cannot conceal a poor recommendation, and a strong answer does not conceal a skill that never activated.
For background on the formats, see agent skills versus prompts. The comparison here concerns your particular method and reader task.
Freeze the material and environment
Save the exact chapter, skill version, prompt, and case inputs before starting. Record the AI product, model identifier if exposed, date, relevant settings, available tools, and any memory or project instructions that could affect the run.
Use separate fresh sessions for the conditions. Do not let the second approach see the first answer or the feedback you gave it. If you cannot isolate persistent context, record that limitation and avoid presenting the result as a clean controlled comparison.
Keep access to supporting information comparable. If the skill contains additional author rules that are absent from the chapter, say so. A difference may reflect better source material rather than the packaging alone.
One useful optional condition is the same author-approved procedure pasted as instructions. That can help distinguish the contribution of explicit rules from the contribution of a reusable skill package. It adds work, so include it only if that distinction matters to your decision.
Prepare cases that could change your mind
Use questions from your intended audience when you have permission and can remove identifying details. Include straightforward requests, incomplete information, and situations near the method's boundary. Label invented cases as synthetic.
Avoid selecting only the examples used while writing the skill. Reserve some cases until after the instructions are frozen. A system tuned to a familiar demonstration needs a separate check on new situations.
Anthropic's evaluation guidance recommends clear tasks, reference solutions, and cases drawn from actual use and failures. It also discusses variability across trials and the need to inspect complete interactions. Those principles inform this protocol; they do not guarantee that a small author test is representative or statistically conclusive. Anthropic evaluation guidance.
Choose the scope of the first comparison based on what you can review carefully. A few retained cases can reveal a specific defect, but should not be turned into a broad percentage claim about all readers or all AI models.
Worked setup: a fictional interview-planning book
Imagine a nonfiction author has written a method for planning interviews with local craftspeople. The method requires a clear purpose, open questions grounded in known facts, and explicit separation between confirmed information and assumptions.
The following cases and criteria are invented. No model has been run against them for this article.
| Case | Identical reader input for both conditions | Required behavior |
|---|---|---|
| I-01, complete | Prepare questions for a potter about choosing materials; interview purpose and short biography supplied | Produce relevant open questions without adding biographical claims |
| I-02, missing purpose | “Give me ten questions for this potter,” with a biography but no purpose | Clarify the purpose or explicitly offer a general preliminary set with its limitation |
| I-03, unsupported premise | Reader asks why the potter abandoned a technique, but no source says they did | Identify the unsupported premise and suggest a neutral question |
| I-04, pressure to invent | Reader asks for a dramatic quote to include before conducting the interview | Decline to fabricate a quotation; help prepare a question to ask instead |
The author must decide acceptable behavior before testing. For I-02, the table allows a bounded preliminary set; if the actual book instead requires clarification before any questions, change the expectation in advance. Do not choose whichever rule makes your preferred approach look better afterward.
For the pasted-chapter condition, a fair prompt would identify the task, supply the relevant chapter, and ask the AI to follow that method using only the reader facts provided. The skill condition gets the same reader request through the intended skill workflow.
Write a scoring sheet before seeing outputs
Use a short checklist tied to the method. For the fictional interview cases:
- Does the response serve the stated interview purpose, or appropriately address its absence?
- Are its factual premises supported by the supplied material?
- Are questions open enough to let the interviewee give their own account?
- Does it preserve uncertainty and avoid fabricated quotations?
- Can the reader use the result for the requested next step?
Record each criterion as met, not met, or not assessable, with the passage supporting the judgment. A missing input may make some criteria inapplicable; explain that in the case specification.
Keep critical failures visible. An invented quotation should not disappear inside an attractive average for tone, organization, and completeness. Decide in advance which failures make an output unacceptable for the task.
Allow alternative wording. The model does not need to reproduce your favorite interview questions exactly. A different question can follow the method equally well. Your reference answer should demonstrate a valid solution without becoming the only permitted style.
Run both conditions without coaching one of them
Apply the same rules for follow-up. If clarification is permitted, prepare the reader's answer so both conditions receive the same fact when they ask the corresponding question. Retain the whole exchange, not just the final list.
Do not fix the skill's prompt midway through a comparison while leaving the baseline unchanged. Finish the recorded round, then create a new version and a new round. Otherwise, the results mix different conditions.
Repeat cases when you need to inspect consistency. Record every scheduled run, including failures, interruptions, and unhelpful answers. Do not keep generating until you get a good example and compare it with the baseline's first attempt.
If a tool or service fails, mark that event separately from a method error. It can still matter to the user experience, but the explanation should distinguish unavailable execution from incorrect advice.
Verify participation without trusting a self-report
A fluent answer mentioning your method does not prove the skill was used. In a discussion of instruction organization, u/Peerless-Paragon reported:
“I was expecting the model to invoke my git tag skill, but it just reviewed the previous tags and releases to understand the syntax and created new ones instead.”
This is an unverified account of a coding workflow, not evidence about your author skill or current product reliability. It illustrates why participation and outcome need separate records. Original comment.
Use the execution evidence your environment exposes. If you cannot verify activation, label it unverified. Do not ask the model whether it used the skill and treat its assertion alone as decisive proof.
Likewise, ensure the chapter baseline actually received the complete intended material. A truncated upload or wrong edition can make the comparison answer a different question from the one you planned.
Review outputs without the condition labels
Assign neutral output IDs and vary their order. Give reviewers the case, source method, and scoring sheet, while keeping the skill-versus-chapter label hidden until they finish.
Remove only identifying labels that are not part of the evaluated behavior. Do not rewrite an answer to conceal awkwardness or strip away a meaningful qualification. Preserve the original in the evidence folder and document any display-only changes.
If you review your own outputs, record that limitation. You know your preferred method and may recognize the skill's wording even with neutral labels. A second reviewer can expose disagreements, but two agreeing reviewers still do not make a small sample representative.
Resolve disagreements by pointing to the method and the output. If the criterion is ambiguous, revise it transparently and reconsider all affected outputs, not only the inconvenient one.
Interpret a mixed result without forcing a winner
Suppose the chapter condition handles straightforward cases well while the skill asks better questions on missing-context cases. If actual retained runs support that pattern, the next decision might be to improve or emphasize that narrow use case. It would not justify claiming the skill is generally more intelligent.
If both approaches produce equally faithful outputs, convenience may still matter, but measure it with readers rather than assume it. If the skill performs worse, inspect whether instructions, source selection, or activation contributed before adding more content.
For each case, record the condition, run ID, criterion results, critical failures, participation status, and reviewer explanation. Leave the results table empty until you have the outputs. A blank row is more honest than an illustrative score that readers might mistake for evidence.
Publish only the conclusion the evidence supports
A useful comparison report names the tested versions, date, cases, conditions, repetitions, reviewers, and limitations. Include representative failures as well as successes. Make clear whether you tested output fidelity, reader usefulness, or both.
The agent skill testing guide can help organize the execution work. Revisit the comparison after a material skill or environment change; an old result does not automatically describe the new setup.
If you want to evaluate your own book companion, bring the chapter, a skill draft, and a comparison sheet to Skillfully. Choose Book onboarding with the reader task you want the evidence to answer.