Define what faithful advice looks like before you score it
An evaluation rubric for an author skill should measure whether the answer follows your method, uses the reader’s actual circumstances, produces useful work, respects your limits, and makes its reasoning inspectable. Give each dimension concrete examples. Identify mistakes that fail the answer regardless of its total score.
You are checking something more specific than whether the assistant sounds intelligent. If your book teaches readers to investigate a problem before proposing a solution, a beautifully written sales pitch can still be the wrong output.
Start with one reader task. Write down the rules your method requires, then score several possible answers against those rules. This article provides a complete fictional example and a worksheet you can adapt. The scores illustrate editorial judgment; they are not measured results from a deployed skill.
Why a convincing answer needs a separate check
Authors already encounter the problem when they ask AI to critique their writing. In a public discussion, u/TecBrat2 described using AI for feedback while keeping prose writing entirely their own:
“It's so positive, I'm concerned that it is trying too hard to please me and that it might miss opportunities to offer correction.”
That is one writer’s concern about their experience, not a performance claim about every model. It captures the uncertainty a rubric should resolve: which specific correction should have been made? Read the original post.
Another contributor in the same thread, u/phototransformations, reported a more useful experience when asking for strengths and weaknesses of individual scenes:
“Most of the time, it comes up with a balanced assessment that align with critiques from human readers.”
The contributor also said broader submissions were less effective. Their experience does not validate a particular tool, but it gives a practical comparison to pursue: does the AI’s assessment agree with informed human review on a bounded task? Read the comment.
Your rubric makes that comparison possible. “I like this answer” becomes “it identified the missing evidence and asked the question my method requires.” You can disagree about the latter and point to the exact sentence causing the disagreement.
Write the rules of your method in observable terms
Suppose an author’s skill helps founders prepare a short customer interview. For this fictional exercise, the author defines four rules:
- Ask about a recent event the customer actually experienced.
- Explore their current approach before discussing a proposed product.
- Keep unverified assumptions visible; do not present them as customer facts.
- Return three opening questions and one follow-up, without claiming that an interview plan validates demand.
These are invented rules for the example, not an attribution to an existing book or a universal interview standard. Your own rubric should use the decisions your method actually teaches.
Now give the skill this input:
I am talking to a freelance designer tomorrow. I think chasing overdue invoices might be frustrating, but I have no evidence yet. I want to learn what happened with their most recent overdue invoice and how they handled it. Help me prepare the opening questions. I have not interviewed them before.
A reviewer can now assess the answer without guessing the author’s intent. The desired task, factual context, required output, and scope boundary are all visible.
Avoid criteria such as “shows expertise” or “understands the reader.” Replace them with what a reviewer can observe. Here, “keeps the overdue-invoice problem hypothetical until the customer describes it” is something two people can check.
Use five dimensions, with a reason for every score
The following is an original starting rubric for author skills. Score each dimension from zero to two. Adapt the anchors to your task rather than applying the wording mechanically to every skill.
| Dimension | 0: fails | 1: needs repair | 2: usable |
|---|---|---|---|
| Method fidelity | Contradicts an essential rule | Uses part of the method but omits a required step | Preserves the required decisions and sequence |
| Context and evidence | Invents or contradicts a consequential fact | Uses the input but leaves an important assumption unclear | Distinguishes supplied facts, assumptions, and missing information |
| Practical usefulness | Gives no usable next step | Gives something concrete that needs substantial repair | Delivers the requested artifact ready for the reader’s review |
| Scope and uncertainty | Claims an unsupported result or crosses a stated boundary | States a limit but applies it inconsistently | Handles limits and uncertainty appropriately |
| Inspectability | Reader cannot see the basis of the advice | Basis is partly visible but an important connection is missing | Reader can connect the advice to the input and method |
Require a brief reason and a supporting excerpt for each score. A bare number cannot tell you whether to fix an instruction, an example, or your expectations.
Microsoft’s rubric reference distinguishes dimensions such as completeness and groundedness and describes comparing AI grades with human judgments during refinement. That supports separating qualities instead of treating “good” as one undifferentiated score. Our five dimensions and score anchors are an adaptation for the author task, not Microsoft’s prescribed rubric. Microsoft’s rubric reference.
There is intentional overlap between some dimensions. A fabricated fact may damage both grounding and scope. Record both consequences, but do not pretend the resulting total is a precise statistical measure of quality. Its purpose is to expose repairable failures.
Set critical failures before adding the scores
For this interview example, three failures override the total:
- Inventing a customer finding or quotation.
- Replacing past-event research with a product pitch.
- Claiming that the plan proves customer demand.
You could give an answer two points for being concise and still refuse to release it because it invented a finding. Critical failures protect the parts of your method that cannot be compensated for elsewhere.
Set a provisional acceptance rule for the exercise: no critical failures, no zero in any dimension, and at least eight out of ten. This is a working editorial threshold, not an industry standard or proof of reliability. Tighten it if the consequence of a wrong answer warrants stricter review.
Do not decide the threshold after seeing your favorite answer. That makes it too easy to redefine success around a draft you already want to publish.
Score three answers to the same reader request
The following answers were written for this article. They are not captured outputs from a named AI model.
Answer A: faithful and useful
You have a hypothesis about overdue invoices, not evidence that this designer finds them frustrating. Start with a recent event:
- Tell me about the most recent invoice that was paid later than you expected.
- What happened between sending the invoice and receiving payment?
- What, if anything, did you do to follow it up?
Follow-up: You mentioned [their action]. What happened next?
These questions explore an actual experience and the designer’s existing approach. If they have not had an overdue invoice, do not assume a problem; ask about their usual payment process instead. Their answers may change your hypothesis. This plan does not establish demand for a product.
Score: 10/10; no critical failure.
It preserves past-event inquiry and the current approach, identifies the founder’s hypothesis, gives the requested questions, and explains their purpose. The conditional follow-up depends on what the customer actually says; it does not supply an invented answer.
Answer B: polished but contrary to the method
This designer is losing three hours a week chasing invoices, which confirms a strong need for automation. Ask: “Would a tool that recovered those hours be valuable?” Then ask which premium features they would pay for. This gives you a compelling opening and validates your product direction.
Score: 2/10; critical failures present.
The answer introduces a numerical finding that was never supplied, moves directly to a proposed solution, and claims validation. It has concrete questions, but they require substantial repair to serve the requested task. Its explanation is visible enough to inspect, although it is based on invented information.
Answer C: cautious but incomplete
Ask about their recent experience rather than assuming they have a problem. Keep an open mind and listen carefully. You do not yet know whether your idea is useful.
Score: 6/10; no critical failure, but not acceptable.
This preserves uncertainty and part of the method. It does not provide the three opening questions and follow-up the reader asked for, or explicitly explore the customer’s existing approach. The answer needs work even though it avoids the dramatic mistakes in Answer B.
| Dimension | A | B | C |
|---|---|---|---|
| Method fidelity | 2 | 0 | 1 |
| Context and evidence | 2 | 0 | 2 |
| Practical usefulness | 2 | 1 | 0 |
| Scope and uncertainty | 2 | 0 | 2 |
| Inspectability | 2 | 1 | 1 |
| Total | 10 | 2 | 6 |
This comparison gives the rubric a useful test of its own. It distinguishes an answer that follows the method, an answer that violates it, and an answer that avoids mistakes but leaves the reader stranded.
If your rubric gives all three similar scores, revise the criteria before using it to judge real work.
Check whether two reviewers mean the same thing by “good”
Ask another person familiar with the method to score a few examples independently. Hide your scores until both of you have finished. Give them the input, the relevant method rules, the full output, and the rubric.
Discuss disagreements at the criterion level. If you give a response a two for usefulness and they give it a zero, identify the artifact each of you expected. Perhaps you accepted general advice while they expected a completed worksheet. The disagreement may reveal an unclear brief rather than a weak reviewer.
Save a good, borderline, and unacceptable example with explanations. Those examples give future reviewers something more concrete than adjectives. Do not require identical wording; a different answer can apply the same method correctly.
Google’s guide to building an expert judge similarly uses expert-labeled examples and compares automated judgments with human assessments. It also warns against evaluating a judge on the same examples already supplied in its instructions. Google’s expert-judge guidance.
For an author working manually, the practical lesson is to keep some examples back. Agreement on the demonstrations you have already discussed is less informative than agreement on a new reader situation.
Use AI grading as assistance you still have to evaluate
An AI reviewer can help organize a growing set of outputs. Ask it to apply each criterion separately, quote the relevant evidence, and identify uncertainty. Give it the actual source rules and reader input.
Then check its judgments against yours. Does it miss the invented “three hours” in Answer B? Does it reward Answer C for being cautious while overlooking the missing questions? Those are grader failures, even if the written explanation sounds reasonable.
Anthropic’s evaluation guidance distinguishes code-based, model-based, and human graders. It notes that model graders require calibration against human judgments and recommends allowing an “Unknown” response when information is insufficient. Anthropic’s evaluation guidance.
Keep “not enough evidence to score” separate from zero. Zero means a criterion failed; uncertainty may mean the evaluator lacks the necessary input. Do not count an unscored case as a pass. Resolve the missing evidence or retain the gap.
You do not need automated grading to start. A spreadsheet containing the input, output, five scores, supporting excerpts, and critical failures can make your first review useful.
Keep the rubric useful after the first draft
Use the same examples to compare revisions, then add cases from real reader problems. Record the skill version, tool/model, date, and any changes to the scoring rules.
If you change the rubric itself, rescore the old answers before announcing an improvement. A higher score under easier criteria does not show that the skill improved.
Review patterns rather than only totals. Repeated grounding failures suggest work on how the skill uses inputs. Repeated usefulness failures suggest the promised output is underspecified. Disagreement about method fidelity may mean you need to explain an exception that exists in your head but not in the instructions.
For the broader operational workflow, see how to test an agent skill. For combining scored reviews with usage and feedback, see measuring agent skill quality. This rubric supplies the author-specific judgment those processes need.
Copy this review record for your own method
- Reader task and supplied input:
- Method rules and relevant source passages:
- Required output:
- Critical failures:
- Output being reviewed:
- Five scores, each with an excerpt and reason:
- Missing evidence or reviewer disagreement:
- Decision: accept, revise, or investigate:
- Instruction or example to change next:
- Skill version, reviewer, tool/model, and date:
Choose one difficult reader question and fill in this record before reviewing a larger batch. Your first useful rubric is the one that helps you explain a mistake and decide what to change.
When you are ready to review a skill built from your method, visit Skillfully and choose Book onboarding. Bring one acceptable output, one failure, and the rules that explain the difference.