Skip to main content
The history.* family reports process metrics — signals derived from how a file changed over time rather than from its current content. The empirical literature is consistent that process metrics out-predict static code metrics for defects while costing far less to compute; mehen pairs them with its real complexity metrics to get composites (like the hotspot) that neither side can produce alone.

What mehen emits

Two of these — history.hotspot and history.churn.relative — are part of the default mehen diff columns for source-code files.

How the history walk works

History metrics cannot come from a language analyzer (which sees one file’s content at one revision). Instead, mehen walks the repository history reachable from the analyzed revision once per revision and folds per-file statistics into each file’s metric set after static analysis:
  • Deterministic by construction. The walk is a pure function of the repository state at the analyzed revision: commits are visited in --date-order topological order, rename matching uses pinned gix rewrite options over raw object bytes (ignoring local diff config and attributes), and “now” for code age is the analyzed revision’s committer timestamp — never wall-clock time. Two machines analyzing the same commit always report identical values.
  • Rename-aware identity. A file renamed along the way accumulates one history: statistics follow the file across renames instead of resetting. Delete-then-recreate sequences split identity, so a new file that reuses an old path does not inherit the dead file’s history — including through merges, parallel branches, and path reuse.
  • Merges are identity-only. Merge commits contribute no churn (matching git log --no-merges and code-maat), but renames performed by conflict resolution still establish file identity.
  • Binary-safe churn. Blobs that are binary (NUL sniff) or larger than 8 MiB churn zero lines, mirroring git log --numstat reporting - for binary files.

Requirements

The walk needs the actual history: run against a full clone. In GitHub Actions use actions/checkout with fetch-depth: 0. Because the walk costs one tree diff per commit, mehen only runs it when a history.* metric is actually requested — a SQL-only or docs-only diff never pays for it.

In diffs

mehen diff walks the history of both revisions, so history columns carry real deltas — the commits and churn a file gained between base and head — rather than comparing against a phantom zero baseline.

References

  • Rahman, F. & Devanbu, P. (2013). How, and why, process metrics are better. ICSE 2013.
  • Tornhill, A. (2015). Your Code as a Crime Scene. Pragmatic Bookshelf.
  • PyDriller process metrics — the reference implementation for churn and ownership semantics.
  • code-maat — Tornhill’s original analyses (age, authors, coupling, churn).