HarperZ9/chorusExplainer, built from commit 29cefb2All repository explainers

chorus

Turn a comment section into a ranked digest you can re-check.

What it does for you

chorus reads a pile of comments and tells you what people are saying: the themes, ranked by how much the crowd engaged and how strongly it felt, the sharpest dissent, and the topics the crowd is split on. Every digest carries a receipt a stranger can re-run to get the same answer. It works on a corpus that gather captured, or on a plain JSON list.

Source: README.md at 29cefb2 (version 0.3.0)

Watch

No concept film fits this tool closely yet. The walkthrough below covers it in text, with real commands and output.

Video walkthrough: coming with the next release.

How it works, one step at a time

Scroll, or use the step buttons. The panel follows the bundled sample, examples/discourse-sample.json: twelve comments under one phone review. Every number is output from chorus at commit 29cefb2, with no model.

  1. 01

    Twelve comments under one review

    Five comments talk about the battery, four about the display and three about the camera. Each carries its like count. The battery comments disagree: two love it and three hate it.

    Source: examples/discourse-sample.json

  2. 02

    Score each comment with a small lexicon

    Sentiment comes from a lexicon of thirty words with rules for negation, intensifiers, capitals and punctuation. The same text always scores the same. c02 scores 0.79 on amazing and best; c03 scores -0.75 on terrible and worst.

    c01 says the battery is incredible, a word the lexicon does not hold, so it scores 0. The digest states this coarseness about itself: English-only and literal.

    Source: src/chorus/sentiment.py, score_text

  3. 03

    Weight by engagement, then by feeling

    A comment's weight is log(1 + likes) multiplied by (1 + 0.5 times the size of its sentiment). A loud comment nobody engaged with stays small, and a neutral comment with many likes still counts.

    c04 has 260 likes and a sentiment of -0.5574, so its weight is log(261) x 1.2787 = 7.115.

    Source: src/chorus/synthesize.py, item_weight

  4. 04

    Cluster into themes

    Comments are clustered by hashed TF-IDF cosine over 512 dimensions, seeded most-engaged first. Five themes come out, ranked by weight. The display leads; the battery splits into a praise cluster and a complaint cluster; c04 stands alone and is labelled a singleton, so it is not presented as a broad crowd theme.

    Source: src/chorus/synthesize.py, cluster

  5. 05

    Name the contested topic

    Clustering filed the battery praise and the battery complaints under different themes, which would hide the disagreement. The contested lens measures every comment that mentions a term. The battery appears in five comments, 60% negative and 20% positive: contested at 0.5386.

    One-sided praise and neutral chatter are left out of this list. The display and the camera are not contested.

    Source: src/chorus/synthesize.py, contested_aspects

  6. 06

    A receipt that re-derives

    The receipt hashes the inputs, the parameters and the digest body. Verify re-scores every comment from its text, re-clusters and re-weights, then compares. The unchanged digest verifies.

    Double a theme's weight and it fails. Double it and recompute the digest's own hash to match, and it still fails, because the weight does not follow from the inputs. Change one like count in the corpus and the old digest fails against it. Pick each case in the panel.

    Source: src/chorus/receipt.py, verify

Walkthrough

Install it, run it once, then use the main feature. Each command below is real, and so is its output.

  1. Install

    Install from a checkout. Python 3.10 or newer; no service key and no model.

    $ git clone https://github.com/HarperZ9/chorus && cd chorus
    $ pip install -e .
  2. First run: digest a thread

    Score, weight and cluster the sample comments, then verify the digest.

    $ chorus run examples/discourse-sample.json --verify
      "contested": [ ... "term": "battery", "contested": 0.5386 ... ]
  3. Re-derive the digest

    In Python, verify a digest against the scored comments. An edited digest is rejected.

    >>> verify(edited, scored)
    False

Output from chorus at 29cefb2 on Windows with Python 3.12. The tamper cases were run through the Python API, chorus.receipt.verify.

What it does not do

Source: README.md at 29cefb2, "What you get" and "Release notes"; the digest's own method.coarseness field

Check what stuck

Answer each one in your head before you open it.

Why does c01 score 0 when it calls the battery incredible?

The lexicon does not hold the word incredible, so nothing in c01 scores.

What is c04's weight, and where do its two factors come from?

7.115: log(1 + 260 likes) times (1 + 0.5 x 0.5574), the size of its sentiment.

Clustering put battery praise and complaints in different themes. How does chorus still show the fight?

The contested lens measures every comment that mentions a term, so the battery reads contested at 0.5386.

A digest's weight is edited and its hash recomputed to match. Why does verify still reject it?

Verify re-derives the digest from the comments themselves, and the edited weight does not follow from them.