Human evaluation
Also known as Human preference evaluation, Expert review
Definition
Human evaluation is a structured review in which people rate or compare model outputs using defined criteria. It captures qualities such as usefulness, factuality, tone, and task fit that may be difficult to reduce to one automated score.
Make the rubric concrete
Reviewers need observable criteria and examples. A rubric might ask whether an answer contains the required facts, follows instructions, avoids unsupported claims, and helps the user complete the task. Define what each score means before collecting ratings.
Control reviewer effects
Use more than one reviewer for important tasks, randomize comparison order, and separate model identity from the output when possible. Measure agreement and investigate disagreements instead of averaging them away. Confidence intervals help communicate uncertainty in the result.
Connect preference to outcomes
A preferred answer is not always a completed task. Pair human ratings with acceptance, correction, escalation, and cost signals. Weave's evaluation context helps teams decide whether a model change improved the product experience or only the review score.
Provide reviewers with enough context to judge the task, while hiding details that could bias them toward one system. Store the rubric version and sampling method with every result. These controls make later comparisons more trustworthy.
How this relates to Weave
Weave can connect human ratings to prompt versions, model routes, latency, and cost. That lets teams investigate whether a preference change came from better quality, a different task mix, or a routing shift.
Explore Token intelligence