A reference from Weave

Engineering & AI glossary

Understand the metrics, models, and methods behind modern engineering. Clear definitions, practical examples, and a closer look at what the numbers actually mean.

All terms

2,008 terms
  • Outcome-oriented roadmap

    Outcome-oriented roadmap is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.

    Developer productivity
  • Outer loop efficiency

    Outer loop efficiency is a developer productivity concept that helps teams understand outer loop efficiency in the context of software delivery.

    Developer productivity
  • Output constraint

    Output constraint is a language-model concept about instruction design and control. It names a mechanism, representation, prompting pattern, decoding control, or context behavior that can change how an AI system processes input and produces output.

    LLM fundamentals
  • Output context

    Output context is a language-model concept about instruction design and control. It names a mechanism, representation, prompting pattern, decoding control, or context behavior that can change how an AI system processes input and produces output.

    LLM fundamentals
  • Output length

    Output length is the serving concept concerned with output length during AI inference.

    Inference performance
  • Output metric

    Output metric describes what a workflow immediately produced.

    Engineering analytics
  • Output quality

    Output quality is a developer productivity concept that helps teams understand output quality in the context of software delivery.

    Developer productivity
  • Output token

    An output token is a unit generated by a language model in its response. Output tokens include visible text and, depending on the API, structured fields or reasoning content returned for the application to process.

    Token costs and AI ROI
  • Output verbosity cost

    Output verbosity cost is expense associated with generating longer model responses. It gives teams a way to name, measure, or reason about an economic property of an AI workload without treating raw usage as proof of value.

    Token costs and AI ROI
  • Overfitting

    Overfitting is a language-model concept about evaluation design and failure analysis. It names a mechanism, representation, training practice, or operational behavior that can change how an AI system processes input and produces output.

    LLM fundamentals
  • Overload control

    Overload control is a reliability concept used to describe a specific condition, control, or decision in the operation of software services.

    Reliability and observability
  • Ownership boundary

    Ownership boundary is a practical concept in a pull-request workflow that shapes how people examine, discuss, own, or integrate a proposed change.

    Code review
  • Ownership debt

    Ownership debt is a software maintenance concern describing a condition that can make future changes, verification, operation, or ownership harder. Its practical importance depends on supported behavior, rate of change, and the consequences of delay.

    Code quality and technical debt
  • Ownership transfer

    Ownership transfer is a practical concept in a pull-request workflow that shapes how people examine, discuss, own, or integrate a proposed change.

    Code review
  • P-hacking

    P-hacking is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.

    Measurement and experimentation
  • P50 latency

    P50 latency is the serving concept concerned with p50 latency during AI inference.

    Inference performance
  • P95 latency

    P95 latency is the serving concept concerned with p95 latency during AI inference.

    Inference performance
  • P99 latency

    P99 latency is the serving concept concerned with p99 latency during AI inference.

    Inference performance
  • Paged attention

    Paged attention is the serving concept concerned with paged attention during AI inference.

    Inference performance
  • Paging policy

    Paging policy is the rules defining which conditions warrant immediate human notification.

    Reliability and observability
  • Pair programming

    Pair programming is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.

    Developer productivity
  • Pair review

    Pair review is a practical concept in a pull-request workflow that shapes how people examine, discuss, own, or integrate a proposed change.

    Code review
  • Paired t-test

    Paired t-test is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.

    Measurement and experimentation
  • Pairing time

    Pairing time is a developer productivity concept that helps teams understand pairing time in the context of software delivery.

    Developer productivity
  • Pairwise ranking

    Pairwise ranking is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Pairwise testing

    Pairwise testing is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.

    Code quality and technical debt
  • Pairwise win rate

    Pairwise win rate is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Panel data

    Panel data is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.

    Measurement and experimentation
  • Parallel inheritance hierarchies

    A code smell in which adding one subtype in one hierarchy requires adding a matching subtype in another.

    Code quality and technical debt
  • Parameter count

    The number of arguments accepted by a function, method, or constructor.

    Code quality and technical debt
  • Pass at k

    Pass at k is the probability that at least one of k generated samples solves an evaluation task. It measures the benefit of giving a model several attempts, rather than the reliability of its first answer.

    Evaluations and benchmarks
  • Pass at one

    Pass at one is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Pass-fail evaluation

    Pass-fail evaluation is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Paved road

    Paved road is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.

    Developer productivity
  • Peer benchmark

    Peer benchmark is an analytical concept for using reference values to understand engineering performance, variation, or capability.

    Engineering analytics
  • Peer Group

    Peer Group is an analytical concept for separating engineering observations into populations whose differences may matter to a decision.

    Engineering analytics
  • Peer mentoring

    Peer mentoring is a way to organize, support, or evaluate software work so that teams can make useful progress with less avoidable friction. It is most valuable when connected to a concrete outcome and the local conditions of the team using it.

    Developer productivity
  • Penetration testing

    Penetration testing is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.

    Code quality and technical debt
  • Per protocol analysis

    Per protocol analysis is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.

    Measurement and experimentation
  • Per-unit metric

    Per-unit metric expresses an amount relative to a defined unit of exposure.

    Engineering analytics
  • Percentage metric

    Percentage metric expresses a part-to-whole relationship out of one hundred.

    Engineering analytics
  • Percentile forecast

    Percentile forecast is a concept used in software delivery planning to describe a condition, relationship, estimate, or decision about engineering work.

    Flow and capacity planning
  • Performance debt

    Performance debt is a software maintenance concern describing a condition that can make future changes, verification, operation, or ownership harder. Its practical importance depends on supported behavior, rate of change, and the consequences of delay.

    Code quality and technical debt
  • Performance testing

    Performance testing is a software testing or test-design practice used to gather evidence about a defined risk, behavior, boundary, or operating condition. It makes the question under test explicit, identifies the inputs and observations that matter, and gives a team a repeatable basis for deciding whether the result is acceptable.

    Code quality and technical debt
  • Perplexity

    Perplexity is a defined lens for examining AI system behavior with a stated task, evidence, and interpretation rule.

    Evaluations and benchmarks
  • Pipeline as code

    Pipeline as code is a software delivery concept used to describe a specific event, interval, control, or operating condition in the path from source change to production behavior. A useful definition names the boundary, unit, and decision the measure supports.

    DORA and DevOps
  • Pipeline failure rate

    Pipeline failure rate is a software delivery concept used to describe a specific event, interval, control, or operating condition in the path from source change to production behavior. A useful definition names the boundary, unit, and decision the measure supports.

    DORA and DevOps
  • Pipeline flakiness

    Pipeline flakiness is a software delivery concept used to describe a specific event, interval, control, or operating condition in the path from source change to production behavior. A useful definition names the boundary, unit, and decision the measure supports.

    DORA and DevOps
  • Pipeline parallelism

    Pipeline parallelism is the serving concept concerned with pipeline parallelism during AI inference.

    Inference performance
  • Placebo effect

    Placebo effect is a statistical or measurement concept used to describe, compare, or interpret engineering data. Its meaning depends on the unit of analysis, data-generating process, and question being asked.

    Measurement and experimentation