Find the knowledge your skill assumes

Your most experienced reader may finish an exercise by supplying an explanation your skill never gave. Your newest reader may stop at that exact point. If you test only with people who already use your framework, you can mistake their knowledge for your companion's effectiveness.

To test an author skill across different reader experience levels, give both groups the same realistic task. Record their prior knowledge separately from their tool experience, observe where they need help, and judge the resulting work against the same quality criteria. Keep assistance visible in your notes.

The aim is to find which parts of your method the skill makes usable, and which parts the reader still has to bring. You do not need to label someone an expert for life to answer that question.

Define experience in relation to this task

A reader can know your subject well without knowing your vocabulary. They can also know your book by heart without ever having applied its method.

In a request for accessible nonfiction, u/Busy_Antelope3507 described an academic background in economics while seeking a starting point in other subjects:

For philosophy, I'd eventually like to understand thinkers like Socrates, Plato, Aristotle and others, but I have essentially no background in philosophy.

This is one reader's account, not an assessment of their ability. It makes the recruitment problem concrete: an experienced reader can still be a beginner in the particular subject your exercise requires. Original reader request

Kate Kaplan makes a related distinction in her discussion of complex software: subject expertise and fluency with a system do not necessarily travel together. Someone can understand the work deeply and still be new to the interface. That distinction is useful when planning a book companion test, even though her article concerns complex applications. Nielsen Norman Group's user profiles

For your participants, record three things separately:

  • Subject experience: What relevant work have they done before?
  • Method familiarity: Which parts of your framework have they read and actually applied?
  • Tool fluency: Have they used the kind of AI interface in which this skill runs?

Avoid using a job title or years subscribed to your newsletter as a substitute for those answers. Ask for a recent example of doing the relevant task. Someone who describes a concrete attempt gives you more useful recruitment information than someone who selects “advanced” on a form.

Build two groups around one clear difference

Suppose Linnea, a fictional author of a book about interpreting everyday data, wants to test a skill for checking claims made from charts. Her method asks readers to inspect what is counted, identify the comparison, and limit the conclusion to what the data supports.

The following plan and dataset are original illustrations. No participants have completed this test, and there are no results to report.

Linnea's question is specific: can people who are new to her method use the skill to check a claim as successfully as readers who have previously applied it?

She is not trying to compare all beginners with all experts. She wants both groups to have basic familiarity with tables and percentages. Otherwise a difference in arithmetic knowledge could obscure a missing explanation in her method.

Recruitment detailNew to the methodExperienced with the method
Subject backgroundHas worked with simple count and percentage tablesHas worked with simple count and percentage tables
Framework exposureHas not read or practiced Linnea's frameworkHas applied the framework to a previous, different example
Evidence collectedDescribes a recent use of a basic tableDescribes a recent table and a previous framework application
Tool experienceRecorded separately; seek a mix comparable to the other groupRecorded separately; seek a mix comparable to the other group
Relationship to authorRecord whether personally knownRecord whether personally known
Task and skill versionIdentical task, instructions and versionIdentical task, instructions and version

For a manageable first round, she could invite three people in each group. That is a proposed learning exercise, not a statistically justified sample size. It can reveal concrete failures worth investigating; it cannot establish a population success rate or prove that familiarity caused a difference.

If she can recruit only professional analysts for one group and occasional spreadsheet users for the other, she should write that mismatch down. Calling them two cohorts does not remove the confounding difference.

Give both groups exactly the same assignment

Here is Linnea's shared task card:

A community center is preparing a short report. Its draft says: “More visitors chose the evening program this month, so the evening program became more popular relative to our other programs.” Use the skill to check that claim. Produce a corrected statement and explain what the table does and does not establish. Treat these as complete attendance counts for each month, not a sample survey.

Attendance measureAprilMay
Evening program visits100120
All program visits, including evening200300

Visits are the unit counted. The table does not identify unique visitors, reasons for attendance, or satisfaction.

The evaluation key, kept out of the participant's task card, is straightforward: evening visits rose by 20, while their share of all visits fell from 50% to 40%. The table supports an increase in evening visits. It does not support the draft's claim of increased relative popularity, and it cannot establish why attendance changed.

One acceptable corrected statement would be: “Evening program visits increased from 100 to 120, while their share of total program visits fell from 50% to 40%.” There are other valid phrasings. The test should reward the reasoning, rather than require the participant to reproduce that sentence.

Linnea gives neither group an extra hint about denominators. An experienced reader may recognize the issue immediately; a new reader may rely on the skill to surface it. That difference is part of what she wants to observe.

A pretest can accidentally teach the answer. If recruitment requires checking basic percentage knowledge, use a different example and avoid previewing the specific count-versus-share trap. Record the check so it remains part of the test conditions.

Record the route, not just the final answer

A correct sentence at the end can conceal several different experiences. The reader might have reasoned through it, copied an answer without understanding it, or received a decisive explanation from the moderator.

Use one observation sheet for both groups:

FieldWhat to record
Participant profileMethod familiarity, subject background and tool experience
First uncertain pointThe step or phrase where uncertainty became visible
Skill responseWhat explanation, question or answer the skill supplied
Participant actionWhat the person did next, without guessing their motive
Human assistanceExact help given and when it occurred
Output before helpSaved separately from any revised answer
Final outputWhether the required distinctions are present
Explanation afterwardWhether the reader can explain the conclusion in their own words
Unresolved questionWhat you would need another session to establish

Before testing, agree on a help policy. You might allow clarification of technical access problems while withholding methodological hints during the initial attempt. If someone becomes stuck, you can move to an assisted phase instead of leaving them frustrated. Mark that transition and preserve the earlier work.

The distinction matters in the results. “Finished after the moderator explained the denominator” is useful evidence about a missing explanation. It is not independent completion.

Similarly, separate trouble opening the skill from trouble interpreting its instructions. Both affect the reader's experience, but they call for different fixes. Do not attribute an unfamiliar interface to a flaw in the author's conceptual method without observing where the failure occurred.

Judge both groups against the same outcome

For this task, Linnea can use four checks:

  1. The output identifies visits as the unit and avoids calling them unique people.
  2. It distinguishes the increase in evening visits from the decrease in their share.
  3. It corrects the original claim without inventing a cause or a satisfaction measure.
  4. The participant can explain the count-versus-share distinction after using the skill.

Record each check separately rather than compressing everything into a score. A polished answer with unsupported causal language needs a different response from an accurate answer the participant cannot explain.

The final explanation is a limited check of understanding in that moment. It does not prove durable learning or independent ability on a future task. To investigate transfer, plan a later task with a different dataset. Repeating this exact table would also test memory of the answer.

Keep the observations individual at first. For a small round, a row for each participant is often more honest and useful than announcing that one cohort achieved a particular percentage improvement.

Turn differences into specific revisions

The following are possible interpretations, not findings from Linnea's fictional test:

Observation you might encounterRevision to investigate
New readers stop at an unexplained term; experienced readers supply its meaningAdd a short explanation where the term first matters
Both groups accept a count as evidence of relative shareRework the comparison check, then test it again
Experienced readers repeatedly skip a lengthy introductionMake background available without forcing it before every attempt
New readers produce correct answers but cannot explain themAsk for a brief interpretation before presenting the finished statement
Tool newcomers cannot find where to beginAddress entry instructions separately from method explanations

Do not infer the cause from the cohort label alone. Revisit the actual interaction and ask what evidence supports your interpretation. A reader who pauses may be confused, checking the numbers carefully, or dealing with an interruption.

A revision that helps new readers can also burden experienced readers. Retest with both profiles. Optional explanations may be worth trying, but they are a design proposal until you see whether people notice and use them appropriately.

When you revise the skill, preserve its task, decision rules and output requirements so you know what changed. The guide to writing an agent skill can help you turn the observed gap into a clearer instruction rather than adding an entire chapter of background.

Your eventual claim should match the evidence: these participants attempted this task with this version under these conditions. Their work can justify the next revision. It cannot yet justify saying the skill works for every beginner or expert who buys your book.

If you have a framework readers understand in theory but struggle to use, bring one exercise and this two-group test plan to Skillfully. Choose Book onboarding to discuss turning that exercise into an author skill readers can try.