Start with the decision you need to make

You do not need every reader's conversation to answer every product question. To investigate whether setup works, you may need a count of successful starts and a few voluntarily reported problems. To judge output quality, you may need a small set of permitted examples. Those are different purposes with different evidence requirements.

Write the decision first, then identify the least intrusive observation that can inform it. Leave fields out when they do not help answer the question. If the remaining evidence cannot support a conclusion, narrow the conclusion rather than silently collecting more.

This resource is an original measurement plan for authors. It is not a privacy audit, a compliance certification, or a description of the data Skillfully currently exposes to authors. No reader conversations or analytics accounts were accessed for it.

Improvement and privacy can pull in different directions

A 2025 CHI paper by Ma and colleagues interviewed 23 users and creators about custom GPT privacy perceptions. Some creators wanted more information to improve their products; others valued limits on conversation access. Participant P1 said:

“I’m not allowed to see your conversation, which is a good thing, right?”

The quotation describes a participant's view in that study, not a current platform guarantee. It captures a useful design question for authors: what do you need to learn, and can you learn it without reading everything? Original research, section 5.2.1

The UK's Information Commissioner's Office frames data minimisation around a stated purpose and the information necessary for it. Its guidance is currently marked as under review. Here, that purpose-first idea informs the worksheet; the worksheet does not establish your legal obligations or satisfy them on its own. ICO guidance

Use a minimal event plan

Consider a fictional author, Iris, whose book teaches readers to revise an artist statement. Her proposed companion helps a reader turn a vague description of their work into a clearer statement of subject, approach, and intent.

Iris wants to know whether readers can start, whether the questions are understandable, and whether the exercise is useful. This proposed plan has not been implemented or tested.

DecisionMinimal proposed observationDeliberately excluded
Does the starting guidance work?Count of confirmed successful starts, with its exact definitionFull prompts, uploaded statements, names of artworks
Where do people ask for help?Support category such as access, setup, or unclear questionFull support message in the analytics table
Do readers report finishing?Optional answer: finished, stopped, or not tried yetThe finished statement unless separately volunteered
Which revision needs attention?Verified revision label, where availableAn assumed revision inferred only from today's date
Is a particular question confusing?Optional selection of the question number and a short explanationAutomatic copying of the surrounding conversation

For each row, add an owner, evidence source, access permissions, and review date for whether the information is still needed. A field is not minimal merely because it is short: a user identifier or a revealing task label can still identify someone.

Keep operational records where they belong. A support team may need contact details to resolve a request, while the author’s weekly improvement sheet needs only the category and time spent. Do not duplicate the full support record into every report.

Inspect what your analytics tool actually receives

An event name such as “exercise started” can be simple while its attached data is not. Check URLs, page titles, extra properties, and free-text fields. A supposedly harmless label can contain a reader's name or confidential project title if it is generated from their input.

Plausible's custom-property documentation explicitly excludes personally identifying information, including pseudonymous end-user identifiers. Choosing a tool with a privacy-oriented design does not remove your responsibility to check the information you send to it. Plausible's custom-property rules

For Iris's proposed plan, use fixed categories such as “question 2 unclear,” not the title of a reader's exhibition. Keep raw inputs and generated statements out of the event payload. Verify the actual fields through an authorized test before describing the implementation as following this plan.

Also distinguish your own measurement from other parties' processing. A decision not to collect transcripts for author analytics says nothing by itself about the AI provider, hosting service, integrations, or support systems. Any reader-facing explanation should cover the actual arrangement, with uncertainties resolved rather than replaced by a broad privacy promise.

Read aggregate counts within their limits

Suppose Iris's fictional weekly sheet shows 30 successful starts and twelve voluntary completion replies: nine finished and three stopped. She can report 30 recorded starts and the twelve replies. She cannot infer a completion rate for all starters from those replies, or assume that missing replies mean failure.

Nor can she necessarily say that thirty different people started. If the observation counts events and one person can trigger several, label it as starts. A distinct-reader count requires an appropriate way to distinguish readers, which may be unavailable or unnecessary for the decision.

Aggregated presentation does not automatically make the underlying data anonymous. Avoid publishing tiny breakdowns that allow someone to recognize an individual. A category containing one identifiable workshop participant may reveal more than a large total suggests.

Keep the interpretation modest: these observations indicate where to investigate. For quality, use a separate review process and the criteria in the guide to measuring agent skill quality.

Ask for deeper evidence separately

When Iris needs to understand why a question confused readers, she can invite a few people to an optional research conversation. Explain the topic, what she wants them to share, whether notes or a recording are proposed, who will use the material, and how long she intends to keep it. Participation should not be a hidden condition of receiving the product they bought.

Offer a fictional practice statement for readers who prefer not to share their own work. If someone volunteers a real example, agree on its use and remove unnecessary details. Permission to discuss an example for product improvement is not automatically permission to publish it as a testimonial.

The resulting notes should distinguish the reader's account, the author's interpretation, and any inspected artifact. Keep a useful finding such as “question two assumes the reader has chosen an audience” without retaining an entire personal narrative when it is unnecessary.

Remove fields that stop earning their place

Review the plan after you make a product decision. If a temporary field no longer answers an active question, decide whether it should be removed and how existing records should be handled under your actual obligations and policies. Do not let an experimental measurement become permanent by neglect.

The goal is a small body of evidence you understand well enough to act on. Fewer fields can make the connection between a reader problem and an author decision easier to see, provided you preserve the limits of what was observed.

To plan a useful first skill around your established method, visit Skillfully and choose Book onboarding. Bring the reader task and the few decisions your measurement plan needs to support.