What it does for you
chorus reads a pile of comments and tells you what people are saying: the themes, ranked by how much the crowd engaged and how strongly it felt, the sharpest dissent, and the topics the crowd is split on. Every digest carries a receipt a stranger can re-run to get the same answer. It works on a corpus that gather captured, or on a plain JSON list.
- Themes, rankedEach theme has its size, a weight, a sentiment split, a controversy score and its strongest dissenting voice.
- The fight, namedA separate lens measures every comment that mentions a term, so a split topic is not hidden by clustering.
- A receipt you can re-run
--verifyre-derives the digest from the inputs and rejects one that does not follow. - Honest nullsA missing engagement signal is recorded as absent, and a thin corpus says so.
Source: README.md at 29cefb2 (version 0.3.0)
Watch
No concept film fits this tool closely yet. The walkthrough below covers it in text, with real commands and output.
Video walkthrough: coming with the next release.
How it works, one step at a time
Scroll, or use the step buttons. The panel follows the bundled sample, examples/discourse-sample.json: twelve comments under one phone review. Every number is output from chorus at commit 29cefb2, with no model.
- 01
Twelve comments under one review
Five comments talk about the battery, four about the display and three about the camera. Each carries its like count. The battery comments disagree: two love it and three hate it.
Source: examples/discourse-sample.json
- 02
Score each comment with a small lexicon
Sentiment comes from a lexicon of thirty words with rules for negation, intensifiers, capitals and punctuation. The same text always scores the same. c02 scores 0.79 on
amazingandbest; c03 scores -0.75 onterribleandworst.c01 says the battery is
incredible, a word the lexicon does not hold, so it scores 0. The digest states this coarseness about itself: English-only and literal.Source: src/chorus/sentiment.py,
score_text - 03
Weight by engagement, then by feeling
A comment's weight is log(1 + likes) multiplied by (1 + 0.5 times the size of its sentiment). A loud comment nobody engaged with stays small, and a neutral comment with many likes still counts.
c04 has 260 likes and a sentiment of -0.5574, so its weight is log(261) x 1.2787 = 7.115.
Source: src/chorus/synthesize.py,
item_weight - 04
Cluster into themes
Comments are clustered by hashed TF-IDF cosine over 512 dimensions, seeded most-engaged first. Five themes come out, ranked by weight. The display leads; the battery splits into a praise cluster and a complaint cluster; c04 stands alone and is labelled a singleton, so it is not presented as a broad crowd theme.
Source: src/chorus/synthesize.py,
cluster - 05
Name the contested topic
Clustering filed the battery praise and the battery complaints under different themes, which would hide the disagreement. The contested lens measures every comment that mentions a term. The battery appears in five comments, 60% negative and 20% positive: contested at 0.5386.
One-sided praise and neutral chatter are left out of this list. The display and the camera are not contested.
Source: src/chorus/synthesize.py,
contested_aspects - 06
A receipt that re-derives
The receipt hashes the inputs, the parameters and the digest body. Verify re-scores every comment from its text, re-clusters and re-weights, then compares. The unchanged digest verifies.
Double a theme's weight and it fails. Double it and recompute the digest's own hash to match, and it still fails, because the weight does not follow from the inputs. Change one like count in the corpus and the old digest fails against it. Pick each case in the panel.
Source: src/chorus/receipt.py,
verify
Walkthrough
Install it, run it once, then use the main feature. Each command below is real, and so is its output.
Install
Install from a checkout. Python 3.10 or newer; no service key and no model.
$ git clone https://github.com/HarperZ9/chorus && cd chorus $ pip install -e .First run: digest a thread
Score, weight and cluster the sample comments, then verify the digest.
$ chorus run examples/discourse-sample.json --verify "contested": [ ... "term": "battery", "contested": 0.5386 ... ]Re-derive the digest
In Python, verify a digest against the scored comments. An edited digest is rejected.
>>> verify(edited, scored) False
Output from chorus at 29cefb2 on Windows with Python 3.12. The tamper cases were run through the Python API, chorus.receipt.verify.
What it does not do
- The lexicon is English-only and literal: no sarcasm, irony or context. A word outside it, like incredible, scores 0.
- Clustering is lexical. Comments that say the same thing in different words can land in different themes.
- Sentiment is a weight and never a verdict. A digest reports what people said; whether they are right is outside it.
- A model overlay with
--modelis tagged with its provenance and never enters the re-checkable core. - The source-change gate reports whether sources changed. It does not decide whether a source claim is true.
Source: README.md at 29cefb2, "What you get" and "Release notes"; the digest's own method.coarseness field
Check what stuck
Answer each one in your head before you open it.
Why does c01 score 0 when it calls the battery incredible?
The lexicon does not hold the word incredible, so nothing in c01 scores.
What is c04's weight, and where do its two factors come from?
7.115: log(1 + 260 likes) times (1 + 0.5 x 0.5574), the size of its sentiment.
Clustering put battery praise and complaints in different themes. How does chorus still show the fight?
The contested lens measures every comment that mentions a term, so the battery reads contested at 0.5386.
A digest's weight is edited and its hash recomputed to match. Why does verify still reject it?
Verify re-derives the digest from the comments themselves, and the edited weight does not follow from them.