HarperZ9/gatherExplainer, built from commit bba8cceAll repository explainers

gather

Pull research out of hard sources, and keep a receipt for every piece.

What it does for you

gather collects research from the places other tools break on: arXiv, authenticated APIs, JavaScript pages, scanned PDFs, audio, video, feeds and your own documents. Every block it extracts is bound to where it came from by a hash. A value a model proposes is kept only if it can be found on the page that was fetched. What you store can be re-hashed later, and anything changed or missing is named.

Source: README.md at bba8cce (version 2.3.0)

Watch

Claiming got cheap. Checking did not. (2 min 57 s, narrated, captioned). Gather keeps a receipt for every source it pulls, so checking a citation stays cheap. Transcript, sources and recall questions.

Video walkthrough: coming with the next release.

How it works, one step at a time

Scroll, or use the step buttons. The panel follows a short local page and two notes files through gather's own CLI and Python API, with no network. Every hash, count and verdict is output from gather at commit bba8cce.

  1. 01

    Extract a page into blocks

    Start with a small HTML page about the aperiodic monotile: a heading and two paragraphs. gather extract article.html parses it and returns one record per block.

    Each block records its path in the document tree and a SHA-256 of its content. The whole record also carries hashes of the source and of the Markdown made from it.

    Source: src/gather/extract.py; README.md, "Quickstart"

  2. 02

    The Markdown it binds to

    gather markdown prints the same extraction as plain Markdown, ready for a model's context. The hashes in the previous step let anyone check later that this text came from that page.

    Source: README.md, "Features"

  3. 03

    A proposed value must be on the page

    Suppose a model extracts a record from the page. verify_record checks every value against the fetched content and marks each field grounded or not.

    Matching is on word boundaries, so a claimed 100 does not match inside 1000. A year the page never states is rejected too. Pick each record in the panel.

    Source: src/gather/schema_extract.py, verify_record

  4. 04

    Store by content hash

    Any fetch command takes --store DIR. Gathering two notes files writes each body under its own SHA-256, in objects/75/24e3e0... and objects/77/a65127..., and records them in a catalog.

    Run the same command again and nothing new is written: both bodies are already there under the same hashes.

    Source: README.md, "A durable local corpus"; src/gather/store.py, src/gather/corpus_cmd.py

  5. 05

    Re-check what you stored

    gather corpus verify re-hashes every stored body and compares it with its address. Untouched, both read MATCH and it exits 0.

    Change one word in a stored body and it reads CORRUPT. Delete a body and it reads MISSING. Either way the command exits 1. Pick each case in the panel.

    Source: README.md, "Quickstart"; src/gather/store.py, src/gather/corpus_cmd.py

  6. 06

    A scope filter, and a sealed digest

    --scope monotile keeps the items that mention the scope terms and counts what it dropped. The run ends with a digest: a seal over every item's receipt, verified on the spot.

    Source: README.md, "Filter ledgers"; src/gather/digest.py

  7. 07

    A tampered receipt breaks the digest

    The bundled demo parses one video into three items, metadata, transcript and a comment, each with a receipt that verifies. It seals them into a digest, then edits one receipt. The digest no longer verifies.

    Source: examples/demo.py

  8. 08

    A pilot run puts it together

    gather pilot takes a closed list of sources through capture, extraction, grounding and storage, then writes a redacted report and a hash-chained receipt. A source whose extracted values are not on the fetched page is recorded as an error and nothing from it is stored.

    A third party re-checks the bundle with no network. This step is described from the README and docs/PILOT.md; a pilot fetches live sources, so it was not run for this page.

    Source: README.md, "How a pilot run works"; docs/PILOT.md

Walkthrough

Install it, run it once, then use the main feature. Each command below is real, and so is its output.

  1. Install

    Install from a checkout to run the demo. Python 3.11 or newer; none of this needs the network.

    $ git clone https://github.com/HarperZ9/gather && cd gather
    $ pip install -e .
  2. First run: the demo

    The demo builds a sealed digest of three receipts, then tampers with one and shows the digest no longer verifies.

    $ python examples/demo.py
    witnessed digest: 3 receipts, seal 7da7dc456b11..., verified True
    
    after tampering one receipt, digest verifies: False  <- caught
  3. Extract a page

    Pull a saved page into structured blocks, each with its source position.

    $ gather extract article.html
    html[1]/body[1]/h1[1]  h1  sha256 b09abb7b64e1424e...
    html[1]/body[1]/p[1]   p   sha256 1a8f9554b2960300...
    html[1]/body[1]/p[2]   p   sha256 95355aaf2715a43d...
    content_sha256   4c907012d520ed44...
    markdown_sha256  94812767e4d6dd59...
    method           html-extract
  4. Store your notes and re-check them

    Store a folder by content hash, then verify the corpus later. A clean corpus exits 0.

    $ gather docs ./notes --store ./corpus
    $ gather corpus verify ./corpus

Output from gather at bba8cce, run from source. gather-engine 2.3.0 is the current PyPI release.

What it does not do

Source: README.md at bba8cce, "Current status" and "Features"; src/gather/schema_extract.py

Check what stuck

Answer each one in your head before you open it.

Why is a proposed copies value of 100 rejected when the page says 1000?

Grounding matches whole words, so 100 is not found inside 1000.

You run the same gather docs command twice with --store. What does the second run write?

Nothing new. The bodies are stored under their own hashes, so both are deduplicated.

What is the difference between CORRUPT and MISSING in corpus verify?

CORRUPT means the stored bytes no longer hash to their address. MISSING means the body is gone. Both make the command exit 1.

A pilot source's extracted fields are not on its fetched page. What is stored?

Nothing from that source. It is recorded as an error.