← All Blogs

Code Quality Metrics That Reveal Technical Debt

Published Read Time: 16 min

Code quality metrics reveal possible technical debt when complexity, structural warnings, and repeated changes worsen together in the same files. Measure four signals—cyclomatic complexity, code smells, lines of code, and code churn—against each repository's history. Code rot describes the longer-term deterioration those signals can help you investigate; it is not a fifth calculated metric.

TL;DR

  • Track cyclomatic complexity against each repository's baseline to find logic that grows harder to test and change.
  • Monitor code smell density over time, and use increases to direct human review toward structural problems.
  • Use lines of code as a denominator for other metrics, not as a measure of output or quality.
  • Measure code churn by file, and investigate modules with repeated modification or rework.
  • Investigate possible code rot when complexity, smells, and concentrated churn deteriorate together. Treat these as maintainability signals rather than automated review verdicts or individual developer rankings.

Why code quality metrics matter now

Technical debt often becomes visible only after it starts affecting delivery. Developers encounter the accumulated cost through slower changes, repeated fixes, fragile tests, and longer reviews. Code quality metrics provide earlier signals by tracking structural conditions associated with difficult maintenance. They do not measure every form of debt: missing documentation, obsolete dependencies, and architectural constraints may need separate assessment.

AI coding tools can increase the volume of code that enters a repository faster than manual review capacity grows. Tools such as Cursor, Claude Code, and Codex can produce large changes quickly, but reviewers still need to understand the resulting logic and maintenance burden. Generated code may add duplication, inflate file size, concentrate complex logic, or trigger repeated revisions. These are patterns to investigate, not evidence that AI-generated code is inherently worse.

A repository's own history provides the most useful reference point. Languages, frameworks, generated files, and architectural styles produce different normal ranges, so one threshold rarely works across codebases. Track changes against a stable baseline instead. A sustained rise in complexity, smell density, or repeated churn deserves investigation even when every value remains below a generic threshold.

Code quality metrics should guide investigation rather than produce pass or fail judgments. When several signals worsen in the same files or services, you have a more useful reason to investigate maintenance risk than any single score provides.

Cyclomatic complexity

Cyclomatic complexity measures the number of linearly independent paths through a unit of code. For a simple function, start at one and add one for each decision that creates another path, such as an if, loop, or case. Analyzers can also count compound Boolean conditions, so compare results produced by the same analyzer and configuration. Microsoft's cyclomatic complexity documentation explains this calculation and its configurable rule threshold.

Higher complexity generally increases the effort needed to test control flow. A function with a score of 12 has 12 linearly independent control-flow paths, though it may have far more possible execution paths at runtime. The score can point to code that demands more review or may resist safe modification. It cannot predict defects, readability, business importance, or maintainability by itself.

Some analysis tools group scores into risk bands, but the boundaries vary by tool and configuration. Treat these bands as review triggers rather than universal pass or fail rules. Functional code, exception-heavy code, generated parsers, and languages with pattern matching can produce scores that do not compare cleanly.

Context determines whether a high score signals technical debt. Generated boilerplate may contain many branches but change rarely and receive extensive automated testing. Table-driven code can represent many explicit cases without creating maintenance problems. Meanwhile, a low-scoring function can still hide confusing state changes or unsafe external calls.

Track complexity by function and summarize its distribution at the repository or service level. Watch for rising median complexity, growth in the high-complexity tail, or repeated increases in frequently changed files. Those trends reveal more than an isolated score because they show where control flow is changing over time.

Code smells

Code smells are structural warning signs that suggest code may become harder to understand, change, or test. Common examples include duplicated logic, long methods, feature envy, and god classes. Feature envy describes code that depends more on another class than its own. A god class concentrates too many responsibilities in one place. Unlike bugs, code smells do not prove that software behaves incorrectly. Martin Fowler's explanation of code smells emphasizes that a visible warning needs deeper investigation before concluding there is a problem.

Static analysis tools and linters find some smells by matching source code against configurable structural rules. A tool might flag repeated blocks, methods above a length limit, or classes with unusually broad responsibilities. Rule sets differ across languages and tools, and not every architectural smell is mechanically detectable. Keep configurations stable when comparing results over time.

Track each smell category by repository or module, then normalize counts by lines of code when codebase size changes materially. A rising density of duplicated logic signals a different maintenance problem than a rising count of long methods. File-level concentration also matters: 20 smells in one frequently changed module may deserve more attention than 20 scattered across stable utility files.

Each smell identifies where a reviewer should investigate, not how much technical debt exists. Duplicated test setup may reflect a reasonable readability tradeoff, while duplicated authorization logic can create inconsistent security behavior. Engineers must judge severity using business importance, change frequency, test coverage, and the cost of correction. Preserve smell types and locations instead of compressing every finding into one score.

Lines of code

Lines of code, or LOC, measures source-code size. Tools may count physical lines or logical statements, with different treatment of comments, blank lines, generated files, and vendored dependencies. Raw LOC says little about maintainability because ten clear lines can replace one dense line without increasing technical debt.

AI-generated code makes raw LOC even less reliable as an output measure. A coding tool may produce verbose scaffolding with simple control flow, which raises LOC without adding proportional maintenance risk. Another tool may compress logic into a small abstraction that hides several decision paths. Language conventions, formatting rules, and code generators create similar distortions.

LOC works best here as a denominator for other code quality metrics. You can track churn and smell counts per thousand lines of code, often abbreviated as KLOC, while accounting for differences in language and repository structure. A sharp increase in LOC becomes more useful when complexity, smell density, or repeated churn rises with it.

Use consistent counting rules and exclude artifacts your engineers do not maintain, such as vendored libraries and generated build output. Review trends at the repository or service level rather than treating line production as developer output. For more on that distinction, see why lines of code are a poor measure of developer productivity.

Code churn

Code churn measures how much a codebase changes by counting lines added and deleted during a defined period. Many diff tools represent a modified line as one deletion plus one addition. Some teams use narrower definitions, such as code rewritten shortly after merge; document which measure you mean before comparing results.

Track churn weekly or monthly, and normalize it by repository size when comparing periods. In their study of Windows Server 2003, Nagappan and Ball found relative churn measures more useful for predicting defect density than absolute churn. That result supports examining context and normalization; it does not establish a universal threshold for your repository.

High churn can reflect healthy development. A new feature, migration, or planned refactor may change thousands of lines without creating technical debt. Churn becomes more concerning when the same files change repeatedly across several pull requests. Repeated fixes, reversions, and rewrites can suggest difficulty stabilizing the design, but they can also reflect changing requirements.

Churn concentration provides a stronger investigation prompt than total volume. For example, 10,000 changed lines spread across a planned migration carry a different meaning than 2,000 changed lines concentrated in one authentication module over six weeks. Review the affected files alongside defect history, code smells, and complexity to distinguish active work from rework.

AI-assisted commits require the same context. A rising merge rate may accompany more completed work, while repeated edits may reflect experimentation, generated code that needs cleanup, or abstractions that fail during integration. Compare AI-assisted and other changes within similar work types and repositories. Keep unknown AI attribution separate rather than assuming untagged commits are entirely human-written.

Code rot

Code rot is the gradual decline of a module's maintainability as the codebase and its operating environment change. A module can deteriorate without a single damaging commit. New dependencies, updated interfaces, changing requirements, and surrounding refactors can make previously sound code harder to modify or test.

No single metric measures code rot directly. Look for sustained deterioration across complexity, smell density, and repeated churn. Rising complexity may show that developers keep adding branches to preserve old behavior. Growing smell density can reveal accumulated duplication or oversized classes. Repeated churn in the same aging modules may indicate that developers must revisit fragile code whenever nearby features change.

Historical baselines make these patterns visible. Module age alone does not indicate decay, and a temporary churn spike may reflect a healthy refactor. Compare each module with its own earlier measurements. Deterioration across several signals creates a reason to investigate; stable scores also do not rule out problems such as unsupported dependencies or changing external APIs.

How the signals work together

Combine code quality metrics at the file or module level because each answers a different question. Complexity measures branching paths, smells identify structural concerns, and churn shows where developers repeatedly modify code. LOC provides a denominator for comparing changes in codebases of different sizes. Code rot describes the longer-term pattern you are investigating.

Worked example: a payment module that needs review

The following is an illustrative example, not Weave customer data or an industry benchmark. Compare the same maintained source files, analyzer rules, and 30-day measurement windows. The baseline snapshot is from six months earlier.

Scroll horizontally if needed →

SignalBaseline windowCurrent windowInterpretation
Maintained source lines10,00012,000The module grew 20%; raw counts need context.
Code smell findings4096Density increased from 4 to 8 findings per KLOC.
Function complexity, 90th percentile814Complexity increased among the most branch-heavy functions.
Added plus deleted lines in 30 days1,0002,400Churn rose from 100 to 200 changed lines per KLOC.
Share of churn in the same five files30%65%More change is concentrated in a small part of the module.

Here, smell density = smell count ÷ maintained LOC × 1,000. For churn, use (added lines + deleted lines) ÷ maintained LOC × 1,000, with LOC measured at the end of each window. A rewritten line can contribute twice to this churn measure. Churn is not the percentage of unique code replaced and can exceed the size of the module.

Smell density and normalized churn both doubled while complexity and concentration increased. Inspect those five files for repeated fixes, dependency problems, missing tests, or unclear responsibilities. Check whether a planned migration explains the changes and whether adjacent periods show the same trend. Two snapshots alone do not establish persistent deterioration or prove that AI caused it.

Avoid collapsing every signal into a universal weighted score. Instead, flag modules when multiple measures depart materially from their historical ranges and remain elevated across several periods. Human review supplies the context needed to decide whether to refactor, improve tests, simplify requirements, or leave the code as it is.

Code quality metrics comparison

Use this table to compare the four calculated metrics and the pattern used to investigate code rot.

Scroll horizontally if needed →

MetricWhat it measuresStrengthLimitationRecommended use
Cyclomatic complexityIndependent control-flow pathsFinds logic-heavy functionsLanguage, analyzer rules, and generated code affect comparisonsTrack distributions and changes against the repository baseline
Code smellsStructural warnings such as duplication and long methodsSurfaces recurring maintainability concernsA smell does not prove a defect or quantify debt severityMonitor findings per KLOC by category and review concentrations
Lines of codeSource-code size under a defined counting conventionProvides a denominator for size-sensitive countsVerbosity, generated code, and language differences weaken comparisonsNormalize smell counts and churn; never treat LOC as developer value
Code churnAdded and deleted lines over a defined periodShows where changes concentrateFeature development and refactoring can produce healthy churnTrack recurring changes by file, module, or service
Code rotA pattern of maintainability decline, not a standalone calculationConnects warning signs across multiple changesNeeds historical and human context; static metrics miss some deteriorationInvestigate long-lived modules against their own history

Turning metrics into team-level trends without ranking developers

Aggregate code quality metrics by repository or service, then map those trends to the team responsible for maintaining that code during the measurement period. Track smell density and churn using consistent denominators. For complexity, monitor the median and upper range rather than relying on one average that can hide a few difficult modules.

Individual scores distort behavior because developers can influence measurements without improving maintainability. A developer may split functions into unnecessary abstractions to lower complexity readings or avoid a valuable refactor because it creates churn. Individual attribution also assigns debt to the person touching old code, even when the underlying decisions predate their work.

The SPACE framework research explains why developer productivity cannot be reduced to one activity metric or dimension. Team-level reporting helps reduce the incentive to optimize personal scores, but still needs release plans, migrations, and incident history to interpret each change.

Use a simple measurement cadence:

  1. Record a baseline. Save metric definitions, analyzer versions, exclusions, and the measurement window. Annotate major rewrites and ownership changes; preserve earlier history rather than resetting away a deteriorating trend.
  2. Review monthly deltas. Look for sustained changes in complexity, smell density, codebase size, and concentrated churn.
  3. Review longer trends quarterly. Compare them with delivery outcomes and major technical initiatives.
  4. Set local review triggers. Investigate departures from the repository's normal range instead of imposing one universal limit across languages and services.

Use commit and pull request history for diagnosis, not a developer leaderboard. Evaluate the technical context and resulting delivery impact instead of converting the signal into a performance score.

Connecting code quality to delivery risk and AI-generated code

Deteriorating maintainability signals can accompany slower delivery when developers spend more time understanding existing code, expanding tests, and reworking changes. Complexity points to control flow they must consider. Concentrated churn points to repeated changes, while smell density surfaces structural concerns. None of these signals establishes delivery impact without evidence from the work itself.

AI coding tools can generate and revise code quickly, while pull-request review may not reveal the resulting repository-level trend. Reviewers usually judge one change at a time, so they may miss gradual movement across dozens of acceptable pull requests.

Compare AI-assisted work with each repository's historical baseline, accounting for work type, team changes, and release plans. A churn spike may reflect productive experimentation, while recurring churn alongside rising complexity suggests a useful place to investigate unstable design or rework. Before-and-after comparisons show association, not causal proof of AI impact. The AI impact measurement framework explains how to connect adoption, output, quality, and delivery evidence.

Maintainability tracking serves a different purpose than automated code review or linting. Review tools evaluate individual changes against rules and may block a pull request. Trend metrics aggregate changes over weeks or months to reveal whether a repository is becoming harder to modify. A series of individually acceptable pull requests can still produce an unhealthy long-term trend.

Where Weave Code Intelligence fits

Weave Code Intelligence uses Silk 1 to score engineering output from code changes, considering intent, complexity, and dependencies across files. It provides explanations and calibration through team feedback. That modeled output measure is different from cyclomatic complexity, a static-analysis smell count, or a direct measurement of technical debt.

Weave Token Intelligence connects AI adoption and costs with engineering output. Its attribution methods include native integrations, commit metadata, and estimated activity-based signals; the available evidence depends on the connected tool. Those views can provide context when investigating whether changes in AI usage coincide with output or maintenance trends.

Use your static analyzer and version-control measurements for the complexity, smell-density, and churn calculations in this guide, then compare the same repositories and time windows with Weave's output and AI-usage evidence. Confirm tool coverage and attribution before drawing conclusions about Cursor, Claude Code, Codex, or a particular model. This workflow does not assume Weave natively calculates every metric described here.

Weave complements code review and static analysis. Evaluate AI ROI using shipped work, maintainability, and cost together; adoption or token volume alone cannot establish value. See how Weave measures engineering output or book a demo to review the measurements available for your team.

FAQs

Is there a universal good cyclomatic complexity number?

No single threshold works across languages, programming styles, and repositories. Use established thresholds as review prompts, then compare each function or module with similar code and the repository's historical baseline. Keep the analyzer and configuration consistent.

How often should you measure code churn?

Collect churn data continuously and review repository or team trends at least monthly. Investigate sudden increases and repeated changes to the same files, since feature development and planned refactoring can produce high churn without indicating technical debt.

Do code smells always indicate technical debt?

No. Code smells identify structural patterns that may make code harder to maintain, but generated code, framework conventions, and deliberate tradeoffs can trigger them. A reviewer should assess context and severity before classifying a smell as debt.

Can code quality metrics be gamed?

Yes. Developers can split functions to reduce complexity, avoid necessary refactoring to suppress churn, or delete useful code to improve LOC trends. Repository and team-level reporting reduces the incentive to optimize personal scores, although metric definitions and review practices must still discourage gaming.

Should lines of code influence engineering performance reviews?

No. Lines of code measure codebase size, not developer value, difficulty, or quality. Use LOC as a denominator for metrics such as smell density or churn per thousand lines, with consistent counting rules and exclusions.

How should you evaluate AI-generated code quality?

Compare maintainability trends before and after adoption of tools such as Cursor, Claude Code, or Codex, and compare similar work within a repository. Track whether AI-assisted changes coincide with rising churn, complexity, or smell density. Review whether those trends persist in the same modules, examine delivery outcomes, and avoid treating correlation as proof of causation.