Give a task, observe the attempt, and save the lesson for later
To see whether readers can use your framework independently, give them a realistic task and the materials your intended product provides. Let them choose their next steps. Record their actions, questions, and output before explaining what you meant.
You can ask neutral questions and help someone who needs to stop or recover. The important boundary is that assistance which supplies part of the solution changes what the attempt demonstrates. Mark it as assisted rather than counting it as independent completion.
This is different from teaching a workshop. In a workshop, your explanation is part of the experience. Here you want to discover which explanations the companion needs to carry when you are absent.
Define the question before you watch
Consider a fictional author, Petra, whose book teaches readers to compare historical accounts. Her proposed companion helps someone distinguish a source's observation from a later interpretation and identify what remains uncertain.
Petra wants to know whether a reader can use the companion to compare two short accounts without treating the more detailed account as automatically more reliable. She prepares two invented accounts of a neighborhood festival. One was written by an attendee that evening; another was written decades later by someone recalling family stories. Both contain gaps.
The reader's task is to write a short account of what can reasonably be said about the festival, with uncertainty visible. Petra is testing the process for assessing evidence, not the reader's knowledge of local history.
Everything in this example is a proposed study design. The accounts, observations, and note entries below are illustrative; no sessions or product tests have been run.
Before recruiting, Petra writes three observable criteria: the output distinguishes the two accounts, supports factual statements with their source, and leaves unresolved differences unresolved. Beautiful prose and agreement with Petra are not criteria.
Use instructions that do not reveal the route
A task should tell participants what they are trying to accomplish without directing every action. GOV.UK's moderated usability guidance recommends relevant, believable tasks with a clear goal that do not disclose the solution.
For Petra, compare these task introductions:
| Leading introduction | Task-focused introduction |
|---|---|
| Use the source comparison step to check which account was written closest to the event. | Prepare a short account of the event that you would be comfortable sharing, using the supplied material and companion. |
| Make sure you include the uncertainty section. | Keep clear what the material establishes and what it leaves open. |
| Click the evidence button if you need help. | Use the available materials as you normally would. |
The second column still defines the desired output. It does not reveal where the companion places its guidance or which step to follow first.
Distinguish a prerequisite from a hint. If the product is only for readers who have completed a particular exercise, recruit accordingly. If it claims to serve newcomers, do not teach that exercise immediately before the test and then conclude that newcomers can use it.
Prepare the materials and try the session logistics with a colleague. That rehearsal checks whether the files open and the task makes sense. It does not substitute for observing intended readers.
Explain why you may stay quiet
At the beginning, explain who will observe, whether anything will be recorded, what you will keep, and how the participant can stop. Make it clear that the work being evaluated is the companion, not their intelligence or historical expertise.
Use a short introduction such as this original draft:
I want to see how these instructions work when I am not explaining them. Please approach the task in your own way. You can tell me what you are looking for or thinking about, but you do not need to narrate every thought. I may save some questions until afterward. If you need a break or want to stop, tell me.
Check that the setup allows the participant to work in a comfortable format. If speaking while reading is disruptive, agree to discuss key moments after the task instead. Record that choice so you understand how the evidence was collected.
Thinking aloud is not direct access to someone's mind. In his guide to think-aloud testing, Jakob Nielsen notes that the situation can be unnatural, statements may be filtered, and facilitator interruptions can influence behavior. Treat speech as one source of evidence alongside actions and artifacts.
Keep three kinds of moderator speech separate
Make a prompt card before the session. When you are tempted to explain, it helps to have a neutral alternative ready.
| Situation | A possible response | How to record it |
|---|---|---|
| Participant pauses while reading | Wait without supplying a theory about the pause | Silence observed; meaning unknown. |
| Participant appears to be searching | “What are you looking for?” | Neutral probe; record the answer. |
| Participant asks what a term means | “What does it mean to you here?” | Interpretation probe, before an explanation. |
| Participant cannot proceed and wants help | Offer to stop, skip, or receive a hint | Record their choice and any hint verbatim. |
| Participant asks about a later consequence | “Let's return to that after this attempt.” | Deferred question; add it to the debrief list. |
| Private material appears unexpectedly | Pause the session and help remove it from view | Safety/privacy interruption; not a usability failure. |
These are options, not lines to repeat mechanically. If a participant asks whether recording has stopped, answer directly. If someone is distressed, attend to them. A research protocol is not a reason to withhold necessary help.
At the same time, notice how easily ordinary conversation can supply an answer. A student preparing a first usability study, u/PoloceTime, reflected:
I will need to keep mindful of the non-leading questions as I typically will say what first comes to mind to keep the conversation flowing rather than considering the bias.
This is the student's account of their own facilitation habit, not a measured claim about every moderator. It captures a practical reason to prepare prompts before a session. Read the original comment.
Observe a pause before interpreting it
If Petra sees someone reread the same passage, several explanations remain possible. They may be comparing wording, checking a detail, losing their place, or struggling with the companion. “Confused by the framework” is an interpretation, not an observation.
Her note should first say what happened: the participant returned to account B twice before writing. She can then ask what they were checking, either during a natural pause or afterward.
Kate Kaplan's guidance on intentional silence recommends allowing space for reading and reflection, while also cautioning against leaving a frustrated participant in prolonged silence. You do not need to turn that advice into a fixed waiting period. Attend to the task and person in front of you.
Avoid celebrating correct moves with “Exactly!” or reacting to mistakes with a worried expression. Either can tell the participant what you hope to see. Acknowledge their effort without grading the answer during the attempt.
If you accidentally steer them, write it down. A contaminated moment does not make the whole session worthless, but it limits what you can conclude about that part of the task.
Use a four-column observation sheet
Keep observed behavior, spoken words, interpretation, and help separate. This fictional example illustrates the distinction:
| Observed action | Participant's words | Tentative interpretation | Moderator action |
|---|---|---|---|
| Copies a vivid detail from account B | “This one gives more information.” | May be treating detail as credibility; unconfirmed | No intervention; flag for debrief. |
| Searches the opening screen twice | “Where was the comparison?” | May not recognize the step label | Asks what they are looking for. |
| Revises the account after a hint | “Oh, I should say who wrote it.” | Attribution instruction may be insufficient | Author supplied a hint; mark assisted. |
| Leaves a discrepancy unresolved | “I can't tell which is right.” | Appropriate uncertainty may have been understood | Save output for later assessment. |
Add the material version and a time or step reference to each row. If you are using an AI companion, retain the relevant generated output with permission; different responses could explain different experiences. Do not silently assume every participant saw identical guidance.
A second observer can take notes while you facilitate, if the participant agrees to their presence. If you work alone, short factual notes are better than a detailed narrative written from memory hours later. Record only what you have permission and a reason to keep.
Debrief without rewriting the attempt
After the attempt ends, return to specific moments. Ask what the reader expected when they searched for the comparison, what made a particular detail seem usable, and whether the final account expresses their own judgment.
Let them correct your interpretation. A search you coded as confusion may have been deliberate checking. A seemingly confident answer may turn out to have been a guess.
You can now explain the intended process and offer teaching if they want it. Keep anything produced after that explanation separate from the independent attempt. A corrected answer demonstrates what happened with clarification; it does not erase the original obstacle.
Evaluate the saved output against the criteria you wrote beforehand. Mark each criterion as met independently, met after assistance, not met, or not assessable. Add a reason. Do not collapse a partly successful attempt into a cheerful pass because the participant liked the book.
Turn an observation into a focused revision
For each issue, preserve the evidence and name the smallest plausible change. Petra might replace an unclear step label, add an example showing that detail does not establish reliability, or make the unresolved-difference field easier to find. These are hypotheses about improvement until another reader tries them.
Use measuring agent skill quality to connect those observations to explicit output criteria. Recheck with a new participant where possible: the first reader has now learned the intended path.
You are ready for the next version when you can explain what broke, what you changed, and what remains untested. If that version is becoming a publishable companion, visit Skillfully and choose Book onboarding. Bring the exercise and the observation notes that show where its instructions need to stand on their own.