This metric addresses the AI-era documentation problem:
The document is just filler: structure is lazy, there are no references, it is large but useless.
This is not AI-authorship detection. It reports structural evidence — unanchored prose, low artifact
density, weak repository grounding, lazy sectioning, repetition, specificity scarcity, hollow
references, and placeholder density — without making any claim about how the text was written.
Sub-scores (.1–17.8)
Bands
Diagnostic labels
High scores attach stable string labels reviewers can act on:
large-unanchored-prose
low-repository-grounding
lazy-sectioning
low-artifact-density
near-duplicate-paragraphs
specificity-scarcity
hollow-references
placeholder-heavy
The PR comment quotes these labels verbatim instead of paraphrasing.
Example output
References
- Pirolli, P. & Card, S. (1999). Information Foraging. Psychological Review 106(4): 643–675 —
motivates the evidence-anchor and specificity-scarcity sub-scores.
DOI.
- Halliday, M. A. K. (1985). Spoken and Written Language. Oxford University Press — lexical-density
basis used by
SpecificityScarcity.
- Manning, C. D., Raghavan, P. & Schütze, H. (2008). Introduction to Information Retrieval, ch. 6.
Cambridge University Press — Jaccard / token-shingle methods used by
RepetitionDensity.
Stanford online edition.
See also