What it does for you
gather collects research from the places other tools break on: arXiv, authenticated APIs, JavaScript pages, scanned PDFs, audio, video, feeds and your own documents. Every block it extracts is bound to where it came from by a hash. A value a model proposes is kept only if it can be found on the page that was fetched. What you store can be re-hashed later, and anything changed or missing is named.
- Markdown with a receipt
gather extractturns a page into Markdown and binds every block to its node path and content hash. - No ungrounded fields
verify_recordrejects a proposed value that appears nowhere on the fetched page, matched on word boundaries. - A corpus you can re-checkBodies are stored by hash and deduplicated.
gather corpus verifyreports CORRUPT and MISSING and exits non-zero. - Grants before reachA run that would start a command, use the network or send a credential needs a grant you set at launch.
Source: README.md at bba8cce (version 2.3.0)
Watch
Video walkthrough: coming with the next release.
How it works, one step at a time
Scroll, or use the step buttons. The panel follows a short local page and two notes files through gather's own CLI and Python API, with no network. Every hash, count and verdict is output from gather at commit bba8cce.
- 01
Extract a page into blocks
Start with a small HTML page about the aperiodic monotile: a heading and two paragraphs.
gather extract article.htmlparses it and returns one record per block.Each block records its path in the document tree and a SHA-256 of its content. The whole record also carries hashes of the source and of the Markdown made from it.
Source: src/gather/extract.py; README.md, "Quickstart"
- 02
The Markdown it binds to
gather markdownprints the same extraction as plain Markdown, ready for a model's context. The hashes in the previous step let anyone check later that this text came from that page.Source: README.md, "Features"
- 03
A proposed value must be on the page
Suppose a model extracts a record from the page.
verify_recordchecks every value against the fetched content and marks each field grounded or not.Matching is on word boundaries, so a claimed
100does not match inside1000. A year the page never states is rejected too. Pick each record in the panel.Source: src/gather/schema_extract.py,
verify_record - 04
Store by content hash
Any fetch command takes
--store DIR. Gathering two notes files writes each body under its own SHA-256, inobjects/75/24e3e0...andobjects/77/a65127..., and records them in a catalog.Run the same command again and nothing new is written: both bodies are already there under the same hashes.
Source: README.md, "A durable local corpus"; src/gather/store.py, src/gather/corpus_cmd.py
- 05
Re-check what you stored
gather corpus verifyre-hashes every stored body and compares it with its address. Untouched, both read MATCH and it exits 0.Change one word in a stored body and it reads CORRUPT. Delete a body and it reads MISSING. Either way the command exits 1. Pick each case in the panel.
Source: README.md, "Quickstart"; src/gather/store.py, src/gather/corpus_cmd.py
- 06
A scope filter, and a sealed digest
--scope monotilekeeps the items that mention the scope terms and counts what it dropped. The run ends with a digest: a seal over every item's receipt, verified on the spot.Source: README.md, "Filter ledgers"; src/gather/digest.py
- 07
A tampered receipt breaks the digest
The bundled demo parses one video into three items, metadata, transcript and a comment, each with a receipt that verifies. It seals them into a digest, then edits one receipt. The digest no longer verifies.
Source: examples/demo.py
- 08
A pilot run puts it together
gather pilottakes a closed list of sources through capture, extraction, grounding and storage, then writes a redacted report and a hash-chained receipt. A source whose extracted values are not on the fetched page is recorded as an error and nothing from it is stored.A third party re-checks the bundle with no network. This step is described from the README and docs/PILOT.md; a pilot fetches live sources, so it was not run for this page.
Source: README.md, "How a pilot run works"; docs/PILOT.md
Walkthrough
Install it, run it once, then use the main feature. Each command below is real, and so is its output.
Install
Install from a checkout to run the demo. Python 3.11 or newer; none of this needs the network.
$ git clone https://github.com/HarperZ9/gather && cd gather $ pip install -e .First run: the demo
The demo builds a sealed digest of three receipts, then tampers with one and shows the digest no longer verifies.
$ python examples/demo.py witnessed digest: 3 receipts, seal 7da7dc456b11..., verified True after tampering one receipt, digest verifies: False <- caughtExtract a page
Pull a saved page into structured blocks, each with its source position.
$ gather extract article.html html[1]/body[1]/h1[1] h1 sha256 b09abb7b64e1424e... html[1]/body[1]/p[1] p sha256 1a8f9554b2960300... html[1]/body[1]/p[2] p sha256 95355aaf2715a43d... content_sha256 4c907012d520ed44... markdown_sha256 94812767e4d6dd59... method html-extractStore your notes and re-check them
Store a folder by content hash, then verify the corpus later. A clean corpus exits 0.
$ gather docs ./notes --store ./corpus $ gather corpus verify ./corpus
Output from gather at bba8cce, run from source. gather-engine 2.3.0 is the current PyPI release.
What it does not do
- A grounded value is one that appears on the fetched page. That shows the page says it. Whether the page is right is a separate check.
- Grounding is a word-boundary text match. A value written differently on the page, such as a number with a thousands separator, may read as ungrounded.
- Browser rendering, OCR, audio transcription and fast parsing are optional.
gather capsreports which ones your install has. - Corpus context selection is acquisition context. It makes no claim about truth, claim support or completeness.
- Video, Reddit and the scholarly APIs run under each service's terms and your own credentials.
Source: README.md at bba8cce, "Current status" and "Features"; src/gather/schema_extract.py
Check what stuck
Answer each one in your head before you open it.
Why is a proposed copies value of 100 rejected when the page says 1000?
Grounding matches whole words, so 100 is not found inside 1000.
You run the same gather docs command twice with --store. What does the second run write?
Nothing new. The bodies are stored under their own hashes, so both are deduplicated.
What is the difference between CORRUPT and MISSING in corpus verify?
CORRUPT means the stored bytes no longer hash to their address. MISSING means the body is gone. Both make the command exit 1.
A pilot source's extracted fields are not on its fetched page. What is stored?
Nothing from that source. It is recorded as an error.