Find the knowledge your skill assumes
Your most experienced reader may finish an exercise by supplying an explanation your skill never gave. Your newest reader may stop at that exact point. If you test only with people who already use your framework, you can mistake their knowledge for your companion's effectiveness.
To test an author skill across different reader experience levels, give both groups the same realistic task. Record their prior knowledge separately from their tool experience, observe where they need help, and judge the resulting work against the same quality criteria. Keep assistance visible in your notes.
The aim is to find which parts of your method the skill makes usable, and which parts the reader still has to bring. You do not need to label someone an expert for life to answer that question.
Define experience in relation to this task
A reader can know your subject well without knowing your vocabulary. They can also know your book by heart without ever having applied its method.
In a request for accessible nonfiction, u/Busy_Antelope3507 described an academic background in economics while seeking a starting point in other subjects:
For philosophy, I'd eventually like to understand thinkers like Socrates, Plato, Aristotle and others, but I have essentially no background in philosophy.
This is one reader's account, not an assessment of their ability. It makes the recruitment problem concrete: an experienced reader can still be a beginner in the particular subject your exercise requires. Original reader request
Kate Kaplan makes a related distinction in her discussion of complex software: subject expertise and fluency with a system do not necessarily travel together. Someone can understand the work deeply and still be new to the interface. That distinction is useful when planning a book companion test, even though her article concerns complex applications. Nielsen Norman Group's user profiles
For your participants, record three things separately:
- Subject experience: What relevant work have they done before?
- Method familiarity: Which parts of your framework have they read and actually applied?
- Tool fluency: Have they used the kind of AI interface in which this skill runs?
Avoid using a job title or years subscribed to your newsletter as a substitute for those answers. Ask for a recent example of doing the relevant task. Someone who describes a concrete attempt gives you more useful recruitment information than someone who selects “advanced” on a form.
Build two groups around one clear difference
Suppose Linnea, a fictional author of a book about interpreting everyday data, wants to test a skill for checking claims made from charts. Her method asks readers to inspect what is counted, identify the comparison, and limit the conclusion to what the data supports.
The following plan and dataset are original illustrations. No participants have completed this test, and there are no results to report.
Linnea's question is specific: can people who are new to her method use the skill to check a claim as successfully as readers who have previously applied it?
She is not trying to compare all beginners with all experts. She wants both groups to have basic familiarity with tables and percentages. Otherwise a difference in arithmetic knowledge could obscure a missing explanation in her method.
| Recruitment detail | New to the method | Experienced with the method |
|---|---|---|
| Subject background | Has worked with simple count and percentage tables | Has worked with simple count and percentage tables |
| Framework exposure | Has not read or practiced Linnea's framework | Has applied the framework to a previous, different example |
| Evidence collected | Describes a recent use of a basic table | Describes a recent table and a previous framework application |
| Tool experience | Recorded separately; seek a mix comparable to the other group | Recorded separately; seek a mix comparable to the other group |
| Relationship to author | Record whether personally known | Record whether personally known |
| Task and skill version | Identical task, instructions and version | Identical task, instructions and version |
For a manageable first round, she could invite three people in each group. That is a proposed learning exercise, not a statistically justified sample size. It can reveal concrete failures worth investigating; it cannot establish a population success rate or prove that familiarity caused a difference.
If she can recruit only professional analysts for one group and occasional spreadsheet users for the other, she should write that mismatch down. Calling them two cohorts does not remove the confounding difference.
Give both groups exactly the same assignment
Here is Linnea's shared task card:
A community center is preparing a short report. Its draft says: “More visitors chose the evening program this month, so the evening program became more popular relative to our other programs.” Use the skill to check that claim. Produce a corrected statement and explain what the table does and does not establish. Treat these as complete attendance counts for each month, not a sample survey.
| Attendance measure | April | May |
|---|---|---|
| Evening program visits | 100 | 120 |
| All program visits, including evening | 200 | 300 |
Visits are the unit counted. The table does not identify unique visitors, reasons for attendance, or satisfaction.
The evaluation key, kept out of the participant's task card, is straightforward: evening visits rose by 20, while their share of all visits fell from 50% to 40%. The table supports an increase in evening visits. It does not support the draft's claim of increased relative popularity, and it cannot establish why attendance changed.
One acceptable corrected statement would be: “Evening program visits increased from 100 to 120, while their share of total program visits fell from 50% to 40%.” There are other valid phrasings. The test should reward the reasoning, rather than require the participant to reproduce that sentence.
Linnea gives neither group an extra hint about denominators. An experienced reader may recognize the issue immediately; a new reader may rely on the skill to surface it. That difference is part of what she wants to observe.
A pretest can accidentally teach the answer. If recruitment requires checking basic percentage knowledge, use a different example and avoid previewing the specific count-versus-share trap. Record the check so it remains part of the test conditions.
Record the route, not just the final answer
A correct sentence at the end can conceal several different experiences. The reader might have reasoned through it, copied an answer without understanding it, or received a decisive explanation from the moderator.
Use one observation sheet for both groups:
| Field | What to record |
|---|---|
| Participant profile | Method familiarity, subject background and tool experience |
| First uncertain point | The step or phrase where uncertainty became visible |
| Skill response | What explanation, question or answer the skill supplied |
| Participant action | What the person did next, without guessing their motive |
| Human assistance | Exact help given and when it occurred |
| Output before help | Saved separately from any revised answer |
| Final output | Whether the required distinctions are present |
| Explanation afterward | Whether the reader can explain the conclusion in their own words |
| Unresolved question | What you would need another session to establish |
Before testing, agree on a help policy. You might allow clarification of technical access problems while withholding methodological hints during the initial attempt. If someone becomes stuck, you can move to an assisted phase instead of leaving them frustrated. Mark that transition and preserve the earlier work.
The distinction matters in the results. “Finished after the moderator explained the denominator” is useful evidence about a missing explanation. It is not independent completion.
Similarly, separate trouble opening the skill from trouble interpreting its instructions. Both affect the reader's experience, but they call for different fixes. Do not attribute an unfamiliar interface to a flaw in the author's conceptual method without observing where the failure occurred.
Judge both groups against the same outcome
For this task, Linnea can use four checks:
- The output identifies visits as the unit and avoids calling them unique people.
- It distinguishes the increase in evening visits from the decrease in their share.
- It corrects the original claim without inventing a cause or a satisfaction measure.
- The participant can explain the count-versus-share distinction after using the skill.
Record each check separately rather than compressing everything into a score. A polished answer with unsupported causal language needs a different response from an accurate answer the participant cannot explain.
The final explanation is a limited check of understanding in that moment. It does not prove durable learning or independent ability on a future task. To investigate transfer, plan a later task with a different dataset. Repeating this exact table would also test memory of the answer.
Keep the observations individual at first. For a small round, a row for each participant is often more honest and useful than announcing that one cohort achieved a particular percentage improvement.
Turn differences into specific revisions
The following are possible interpretations, not findings from Linnea's fictional test:
| Observation you might encounter | Revision to investigate |
|---|---|
| New readers stop at an unexplained term; experienced readers supply its meaning | Add a short explanation where the term first matters |
| Both groups accept a count as evidence of relative share | Rework the comparison check, then test it again |
| Experienced readers repeatedly skip a lengthy introduction | Make background available without forcing it before every attempt |
| New readers produce correct answers but cannot explain them | Ask for a brief interpretation before presenting the finished statement |
| Tool newcomers cannot find where to begin | Address entry instructions separately from method explanations |
Do not infer the cause from the cohort label alone. Revisit the actual interaction and ask what evidence supports your interpretation. A reader who pauses may be confused, checking the numbers carefully, or dealing with an interruption.
A revision that helps new readers can also burden experienced readers. Retest with both profiles. Optional explanations may be worth trying, but they are a design proposal until you see whether people notice and use them appropriately.
When you revise the skill, preserve its task, decision rules and output requirements so you know what changed. The guide to writing an agent skill can help you turn the observed gap into a clearer instruction rather than adding an entire chapter of background.
Your eventual claim should match the evidence: these participants attempted this task with this version under these conditions. Their work can justify the next revision. It cannot yet justify saying the skill works for every beginner or expert who buys your book.
If you have a framework readers understand in theory but struggle to use, bring one exercise and this two-group test plan to Skillfully. Choose Book onboarding to discuss turning that exercise into an author skill readers can try.