essay · education learning systems

What the Formula Counts

Use readability scores to locate possible friction, and reserve claims about understanding for evidence about readers.

By Zain Dana Harper · Published 2026-09-28 · Updated 2026-09-28

In this article
Research summary and limits

Question. What changes when a document receives a lower reading-grade score?

Finding. The formula’s counted features change. That result does not directly show what a reader understood or could do.

Evidence. Formula calibration, official communication guidance and a randomized trial distinguish scoring from reader outcomes.

Limit. A null result across two health topics and grades 8 to 14 does not settle the value of simpler wording, organization, translation or visual explanation.

A score and a reader

After an edit, a document gets a lower reading-grade score. Something has changed: the features counted by the formula. Whether a reader now understands the document needs other evidence. Writers should use the score to locate possible friction and reserve claims about understanding for evidence about readers.

Readability formulas usually count features such as word length and sentence length. Those features can help identify passages worth revising. They miss whether a term is familiar to the intended audience, whether an explanation comes in the right order, and whether the reader knows what to do next. AHRQ's guidance recommends using formulas for diagnosis and warns that a grade score cannot show that a report is suitable overall. AHRQ's formula guidance

A formula's history matters too. The Flesch-Kincaid recalibration drew on Navy enlisted personnel and technical-training passages. That origin does not make the calculation useless today. Putting its grade label on a new audience, though, takes judgment about that audience and their task. A label from one calibration cannot stand in for a fresh comprehension test. Kincaid and colleagues' report

A recent randomized trial puts the distinction to work. Researchers compared health information targeted at grades 8, 10, 12 and 14, analyzing 2,235 complete, valid responses from Australian adults. They detected no difference in knowledge between the versions at the reported threshold. The result covers two health topics, immediate outcomes and a lowest target of grade 8. It leaves open the value of simpler wording below that range, of better organization, of translation and of visual explanation. The Readability Study

The strongest objection is practical: testing with the audience takes time, and a score is cheap. That argues for using the score to triage. The score still observes only what it counts. A team with limited resources can inspect difficult passages, review the message and action instructions, and disclose that reader understanding remains untested. Cost explains why a team accepts the uncertainty; the uncertainty remains.

Broader review tools help, though none replaces a reader. The CDC Clear Communication Index asks about audience, message, design, numbers and recommendations. PEMAT separates understandability from actionability and states that accuracy and completeness fall outside its scope. Each tool's limits belong beside its results. CDC Index, PEMAT guide

Accessibility needs its own evidence as well. WCAG's Level AAA reading-level criterion calls for supplemental or less demanding content under specified conditions. A document can pass a formula check and still fall short of the wider standard. W3C's explanation of SC 3.1.5

A workable editorial sequence has three steps: inspect what the formula flags, judge the whole document against its purpose, then test the important claims about readers with the intended audience. This workflow is a proposal and has not been validated as a general protocol. Its value is that each result keeps its own meaning. A document can become easier to score before anyone has shown it became easier to use.

Sources

  1. Tip 6. Use Caution With Readability Formulas for Quality Reports. Agency for Healthcare Research and Quality. Official formula guidance; created February 2015 and last reviewed May 2015.. Publication date unavailable; observed 2026-09-28.
  2. Derivation of New Readability Formulas (Automated Readability Index, Fog Count and Flesch Reading Ease Formula) for Navy Enlisted Personnel. Naval Technical Training Command; University of Central Florida repository. Historical primary calibration report, 1975; Navy personnel and technical-training passages.. Publication date unavailable; observed 2026-09-28.
  3. The Readability Study: A Randomised Trial of Health Information Written at Different Grade Reading Levels. Journal of General Internal Medicine. Randomized reader-outcome trial; first published online, with June 2025 issue publication.. Published 2024-12-20; observed 2026-09-28.
  4. CDC Clear Communication Index User Guide. Centers for Disease Control and Prevention. Official material-review guide, August 2019; audience, message and action criteria.. Publication date unavailable; observed 2026-09-28.
  5. The Patient Education Materials Assessment Tool and User’s Guide. Agency for Healthcare Research and Quality. Official understandability and actionability guide, November 2013; updated August 2014.. Publication date unavailable; observed 2026-09-28.
  6. Understanding Success Criterion 3.1.5: Reading Level. World Wide Web Consortium. Informative explanation of WCAG 2.1 Level AAA reading-level criterion; living guidance.. Publication date unavailable; observed 2026-09-28.
Claim notes and limitations
  1. Readability formulas commonly count word and sentence length without directly observing understanding.

    formula-scope · verified

    Sources: ahrq-formulas

    Scope
    Formula features and diagnostic use described by AHRQ.
    Uncertainty
    Different formulas and readers can yield different interpretations.
    Does not prove
    Overall suitability, comprehension, audience fit or accessibility.
  2. The Flesch-Kincaid recalibration used Navy enlisted personnel and technical-training passages.

    calibration · verified

    Sources: kincaid

    Scope
    1975 report; 531 personnel, four technical schools, two bases and 18 training-manual passages.
    Uncertainty
    A calibration population and task do not represent every later audience.
    Does not prove
    An equivalent comprehension result for a present-day reader.
  3. The trial did not detect a knowledge difference between the targeted grade conditions at its reported threshold.

    trial-null · verified

    Sources: mac

    Scope
    2,235 complete, valid analyzed responses from 2,639 randomized Australian adults; target grades 8, 10, 12 and 14; two health topics.
    Uncertainty
    Immediate outcomes, online recruitment and the grade-eight lower boundary limit inference.
    Does not prove
    Equivalence, failure of all plain-language work, or the value of lower grades, visuals, translation and organization.
  4. The CDC Index and PEMAT examine material features beyond a readability formula.

    review-tools · verified

    Sources: cdc-index, pemat

    Scope
    CDC audience, message, design, numbers and recommendations; PEMAT understandability and actionability.
    Uncertainty
    PEMAT leaves accuracy and completeness outside its scope.
    Does not prove
    Reader comprehension, behavior change or overall material quality.
  5. WCAG’s Level AAA reading-level criterion is one part of a broader accessibility standard.

    accessibility · verified

    Sources: wcag-reading

    Scope
    SC 3.1.5 supplemental or less demanding content under its stated conditions.
    Uncertainty
    The cited Understanding page is informative guidance.
    Does not prove
    Full WCAG conformance from a formula result.
  6. Inspect formula flags, review the whole document and test important reader claims with the intended audience.

    editorial-sequence · inferred

    Sources: ahrq-formulas, cdc-index, pemat, mac

    Scope
    Proposed editorial workflow.
    Uncertainty
    Resources constrain testing, and the sequence is not a universal validated protocol.
    Does not prove
    That a lower score makes a document easier to use.

Corrections

No corrections recorded.