Statistical power
Also known as AI Statistical power, Statistical power measure
Definition
Statistical power is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.
How to use Statistical power
Statistical power gives an evaluation a defined lens. The team states what behavior matters, identifies the unit being judged, and records the evidence supporting each result. The same label can produce different conclusions when task population, prompt, model version, sampling rule, or judge instructions change. Write those conditions down before comparing systems. A useful report includes the denominator, exclusions, scoring rule, and representative failures so a convenient proxy does not become the definition of quality.
Concrete example
Imagine a team comparing two coding assistants on repository maintenance tasks. It freezes the task set, records tool permissions and runtime conditions, and stores each output with its score and review notes. It examines successes and failures, including cases where a high score concealed a missing requirement. If the result guides a release or routing decision, it also checks latency, cost, and correction work. That context turns a number into a traceable decision aid.
Limitations
No single evaluation captures every production condition. Small or familiar samples can be noisy, automated judges can inherit systematic preferences, and aggregate results can hide important slices. Use this term to answer a stated question, not to prove universal superiority. Preserve the protocol, report uncertainty, and rerun when the workload or configuration changes.
How this relates to Weave
For Weave Token Intelligence users, this concept helps frame evaluation evidence around real AI usage. Connect the result to task, model or route, latency, and cost context available in the product, then inspect representative cases before drawing a conclusion.
Explore Token intelligence