Measurement and experimentation

Human evaluation

Also known as Human preference evaluation, Expert review

By WeavePublished 1 min read

Definition

Human evaluation is a structured review in which people rate or compare model outputs using defined criteria. It captures qualities such as usefulness, factuality, tone, and task fit that may be difficult to reduce to one automated score.

Make the rubric concrete

Reviewers need observable criteria and examples. A rubric might ask whether an answer contains the required facts, follows instructions, avoids unsupported claims, and helps the user complete the task. Define what each score means before collecting ratings.

Control reviewer effects

Use more than one reviewer for important tasks, randomize comparison order, and separate model identity from the output when possible. Measure agreement and investigate disagreements instead of averaging them away. Confidence intervals help communicate uncertainty in the result.

Connect preference to outcomes

A preferred answer is not always a completed task. Pair human ratings with acceptance, correction, escalation, and cost signals. Weave's evaluation context helps teams decide whether a model change improved the product experience or only the review score.

Provide reviewers with enough context to judge the task, while hiding details that could bias them toward one system. Store the rubric version and sampling method with every result. These controls make later comparisons more trustworthy.

How this relates to Weave

Weave can connect human ratings to prompt versions, model routes, latency, and cost. That lets teams investigate whether a preference change came from better quality, a different task mix, or a routing shift.

Explore Token intelligence

Sources and further reading

  1. Best Practices for Human Evaluation of Large Language Models