It brings information in, from anywhere, and records how.
Network access, third-party tools, and credentials stay isolated behind source adapters. After intake, a run applies its scope filter, folds item receipts into one witnessed digest, and hands that digest to downstream indexing, refinement, or claim evaluation.
Every kind of intake is one small adapter behind a single shape: one string in, a list of receipted items out. The adapter can use the network, a tool, a credential, a headless browser, whatever the source demands; nothing else in gather imports any of that. Awkward access is an adapter problem, not a system problem, and the core stays pure standard library because of it.
The receipt is the rule. Each item records how it was obtained, so a transcript read from captions, text recognized from a scan, and a fact synthesized from fragments are all valid items, but they are not equal, and they are never confused.Reach anywhere. Then say exactly how.
gather ships a source adapter for each kind of intake, all behind the one shape. The cheap, local sources are pure standard library; the harder ones shell out to their own external tool, never a Python dependency. Each adapter records the method it used, so the accountability is in place before the harder reach is trusted. The method on every item is the load-bearing field: it keeps a direct read and a derived inference from ever being read as the same thing.
yt-dlp; pure parsing behind an impure shell, each item receipted
http-get, exactly that
pdftotext, and from a scanned image via tesseract; a best-effort machine reading, labelled ocr, never dressed up as clean text
transcribe so a listener knows it was machine-made
browser-extract, so you know JavaScript ran. This adapter has a narrower trust boundary than the static web route; see the limit below
The honesty matters most exactly where the reach is widest. The web adapter runs no JavaScript, so it is honest that a client-rendered page gives back only its shell. The browser adapter runs a real headless browser, and it is the most exposed adapter in the system: its host guard covers only the first navigation, and the rendered page then follows its own redirects and sub-requests unguarded. So do not point it at untrusted URLs in an environment where internal services are reachable. The browser-extract method is honest that JavaScript ran, but it cannot make the fetch safe. The threat model says so plainly, in the open, rather than hide it.
The receipt is the differentiator.
A provenance receipt rides on every item: the source, the reference, the method, the time, and a sha256 of the item’s own content. Re-hash the content and you can confirm it is what was obtained, unaltered. The method keeps the harder distinction on the record, yt-dlp, browser-extract, ocr, transcribe, synthesized, so a quote pulled straight from a source and an inference pieced together from fragments are both valid items, but they are never equally direct, and the receipt never lets one be read as the other.
A fact pieced together is not a fact pulled straight from the source. The receipt keeps them apart.
A derived item, one assembled or inferred from other items rather than fetched, is the sharp case, and the receipt is built for it. Its sha256 fingerprints the inference itself, not its sources, because it is a new statement that can only witness itself; a derived_from field records the content hash of each input, a re-checkable pointer back to the exact source content. The honesty is mechanical where it can be: the synthesized label is reachable only through one seam, and the bare builder refuses to stamp it and defaults to compiled, so a plain call can never forge a synthesis. With no model wired in, the default compiles the inputs verbatim and invents nothing. The digest seal folds in the method and the inputs alongside the hash, so relabelling an inference as a direct fetch, or quietly rewriting what it was built from, breaks the seal exactly as altering the content does.
Standard library at the core. The impurity is fenced at the edge.
The core runs on Python’s standard library with no third-party runtime dependency. An adapter may pull in whatever its source demands, a browser binary, an OCR engine, a transcription CLI, but only behind the Source shape and only at the impure edge. Clocks are injected everywhere a time is recorded, so given the same fetched items a run is deterministic and replayable, and verification stays trustworthy. The four stages below are the spine of a run, each one re-checkable from disk.
Source shape; network, tool, and credential use stays inside that adapter boundary
corpus verify re-hashes every stored body
A full gather run folds fetch, scope, optional synthesis, digest, and store into one re-checkable RunRecord with its own seal over what it did. Recall queries a stored corpus and re-verifies every body it returns, so a missing or corrupt item is reported, never handed back as if intact. The scope filter, the synthesizer, the store, and the external provenance verdict are all seams that default to a Null, so gather stands alone and a peer system plugs in without the core ever importing it.
Current state. How to run it.
Version 1.9.1, published as gather-engine, installs the gather command and the gather package. New corpus writes preserve exact UTF-8 source text while selection uses a readable LF-normalized view, so mixed-newline text can be quoted with its source identity intact. Selections stay pinned to the witnessed corpus digest used by --expect-digest. The core remains pure standard library; optional adapters call an external tool (yt-dlp, pdftotext, headless chromium, tesseract, or a Whisper-style CLI) only when that route is selected.
Version 1.8.1 added protection for retained corpus reads against directory replacement on unsupported Linux mounts. On WSL, keep corpora on native Linux storage or use native Windows Python: Windows-drive 9p/v9fs mounts are refused with UNSAFE_PATH. Python hosts can pass a retained same-process corpus descriptor; CLI and MCP continue to accept directory strings.
Install 1.9.1 or later. Every earlier release is inside the range of at least one of the four Gather advisories listed on the Security page. A tool found only inside the working folder now needs its absolute path in its GATHER_<TOOL> variable, and an MCP host that relied on the old gather.run or gather.pilot behavior adds a grant to its launch configuration.
Create sample.txt with this exact two-line body, then store it and select only the useful slice.
sample.txt body
Offline quickstart: Gather keeps acquired context separate from conclusions.
Decision fact: the release can select useful text from a stored corpus.
After the inspect command, replace ROW_REF and CORPUS_DIGEST with the values from the inspect JSON. 0:76 selects the first line of this fixture only.
$ pip install gather-engine==1.9.1 $ gather docs sample.txt --store corpus --json $ gather corpus context corpus --json --excerpt-chars 90 $ gather corpus context corpus --json --select ROW_REF:0:76 --expect-digest CORPUS_DIGEST
Sample selected text: Offline quickstart: Gather keeps acquired context separate from conclusions.
Release evidence and context limits
Version 1.9.1, released on 27 September 2026, fixes two more advisories, GHSA-r38f-cr69-jpp8 (high) and GHSA-j6j7-39vh-qrp4 (medium). A PATH entry that reaches the working folder, whether it names the folder, spells it another way or links to it, no longer lets a program planted there start in place of a tool. The folder of the Python that runs Gather stays on PATH, and on Windows so do the Windows, System32 and SysWOW64 folders. When Gather runs from a filesystem root, or from the home folder or a folder above it, the guard drops only the entry naming that folder itself. On Windows, the file sources and every MCP path argument refuse network shares and device paths. Every version before 1.9.1 is affected by at least one of the four advisories, so install 1.9.1 or later.
Version 1.9.0, released on 26 September 2026, fixes two high-severity security advisories, GHSA-pxvv-rg3f-4v5w and GHSA-r4f3-9xrf-72m5. Before gather.run or gather.pilot runs a command, reaches a network source or sends a credential, the MCP server now needs a grant given at launch. Every external tool now starts from an absolute path in a private folder. Every version before 1.9.0 is affected by both.
The 1.9.1 wheel installed from PyPI on 27 September 2026 and ran the quickstart above. The checks below were run on earlier releases.
The public 1.8.1 wheel passed isolated Windows CLI, MCP, retained-descriptor and tamper checks. On WSL x86_64, native reads passed and an actual 9p directory-replacement attempt was refused; descendant-directory and leaf controls were instrumented. The downloaded packages match the release commit: 64 Python modules and 133 tracked source archive files. Linux mount classification uses the opened descriptor and refuses unavailable mount information. Other POSIX platforms retain no-follow and type checks without a mount-detection claim.
The earlier released 1.7.1 wheel smoke run selected a newline-spanning DECISION/part fixture through both CLI and MCP. The payload used schema gather.readable-context/v1, returned verified: true, kept source SHA-256 b1736b2f3f466fdf55d900ccf4c28ba9a2170c60914ede77173873f447dba6d3 separate from readable view SHA-256 bfac793415f57cdb3aadb174993a58dc4767abe8005b78855daae6ab110b93cf, and produced selection digest 0450040a624224de25249cc42f1a68099898d8d8accf68d286e71c4771b7324b.
The publication receipt verifies the GitHub release, PyPI files, GitHub and PyPI download-back hashes, a clean Windows wheel install, CLI and MCP context selection, and tamper refusal. Those checks do not prove source truth, claim support, coverage completeness, downstream model use, or absence of sensitive material in selected text or URLs. Selection offsets belong to the readable view. For legacy corpus rows, reconstruction verifies source text but cannot prove historical raw object-byte integrity. Detecting removed or changed storage witnesses requires a previously pinned corpus digest.