HarperZ9/mnemeExplainer, built from commit db3a8d4All repository explainers

mneme

Agent memory where every recall can be re-run and every stale memory says so.

What it does for you

mneme gives an agent a memory you can question. Each stored fact names the turn it came from. Each recall returns a receipt with the scores and the rule that ranked it, which you can re-run to get the same ranking. When a source changes, the memory built on it reports drift. When you erase something, the erasure leaves an audit entry that says what went and holds none of the text.

Source: README.md at db3a8d4 (release 0.7.0)

Watch

Re-derive it. Don't take it on trust. (2 min 5 s, narrated, captioned). Every Mneme recall carries a receipt you can re-run, the habit this film describes. Transcript, sources and recall questions.

Video walkthrough: coming with the next release.

How it works, one step at a time

Scroll, or use the step buttons. The panel follows one short conversation through mneme's own tour, examples/tour.py, which CI runs on every push. Every identifier, score and verdict is output from mneme at commit db3a8d4 with no model.

  1. 01

    Four turns go in

    Alice says three things about herself, and the assistant answers once. remember stores every turn verbatim as the bottom tier, L0.

    Memory has four tiers. Raw turns sit at L0, atomic facts at L1, scene blocks at L2 and a persona at L3 that cites its facts.

    Source: examples/tour.py, remember; README.md, "The 4-tier memory model"

  2. 02

    Three facts come out, each with its source

    The rule extractor turns the three user turns into three L1 facts. The assistant turn is context, not memory, so it yields none.

    Each fact records the turn id it came from, the extractor rule/v1, the criterion and a SHA-256 of its content.

    Source: src/mneme/memory.py, remember; src/mneme/extract.py

  3. 03

    Recall returns a receipt

    Ask for "tea or coffee preference". The keyword recall scores each fact with BM25 and returns one hit at 1.783.

    The receipt carries the query, the strategy, the fusion rule, the corpus size, every component score and a hash of the scorer's definition.

    Source: src/mneme/recall.py, recall

  4. 04

    Re-run the recall to check it

    verify_recall runs the same scorer over the same rows and compares the result with the receipt. It returns True.

    Change a score inside the receipt and it returns False, because the ranking is re-derived and never read from the receipt. Change a stored fact and it returns False too, because the store no longer reproduces. Pick each case in the panel.

    Source: src/mneme/recall.py, verify_recall

  5. 05

    Drift: check every memory against its source

    drift re-checks each memory. The row must first reproduce its own content hash. Then each cited source is re-hashed from its actual fields and compared with the hash taken at extraction.

    On the fresh store all three facts read MATCH. Pick a change in the panel to see what each one does.

    Source: src/mneme/drift.py; README.md, "A drift check for source changes"

  6. 06

    The roll-up fails closed

    Across a store, any DRIFT makes the whole report DRIFT. Otherwise any UNVERIFIABLE makes it UNVERIFIABLE. Only a clean sweep reports MATCH, and mneme drift exits 1 on drift.

    A memory whose source was deleted is never rounded up to a match because nothing contradicted it. Missing evidence is reported as missing.

    Source: README.md, "A drift check for source changes"

  7. 07

    Forget, with a receipt

    forget erases a memory, the turns it came from and everything derived from them, in one transaction. Other memories from the same turn need your consent.

    Erasing the shellfish fact removed one memory row and one source turn. The receipt's status is erased with no findings, and the audit log holds two tombstones with its chain intact.

    Source: src/mneme/erase.py, forget_memory; README.md, "Accountable forgetting"

  8. 08

    A tombstone holds no text

    Each erased row leaves one audit entry with a random erase ref, the hash of the row before erasure, the reason and the entry's own hash, chained to the one before.

    The entry stores a salted commitment to the erased text, and the salt is not stored, so the log cannot confirm a guess of what was erased.

    Source: README.md, "Accountable forgetting"; src/mneme/audit_writer.py

Walkthrough

Install it, run it once, then use the main feature. Each command below is real, and so is its output.

  1. Install

    Install from PyPI, or clone to run the tour. Python 3.11 or newer; no model and no network.

    $ python -m pip install flywheel-mneme
    $ git clone https://github.com/HarperZ9/mneme && cd mneme
  2. First run: the tour

    The tour stores a short conversation, recalls from it with a receipt, and shows a stale memory flagging itself.

    $ python examples/tour.py
    == 2. recall — with a receipt a third party can re-run ==
        [1.783] I prefer tea over coffee and I work in data science.
      re-ran the scorer: identical ranking (the recall is re-derivable)
    
    == 3. drift — a memory whose source changes flags itself ==
      before: MATCH
      after a source changed: DRIFT (stale memory says so, it is not silently served)
  3. Recall with a receipt

    In your own code, a recall returns the ranked facts and a receipt that records how they were ranked.

    >>> mem.recall("tea or coffee preference", strategy="keyword")
    schema       mneme.recall/1
    fusion       bm25
    corpus_size  3
    hit          04d7a310  bm25 1.7833  fused 1.7833
                 "I prefer tea over coffee and I work in data science."
    def_sha256   a4aca2ec2ee49298...
  4. Forget, with a receipt

    Forgetting erases the text and leaves a tombstone that records what was removed and why.

    >>> mem.forget("dffe9521d4a4ccfb", reason="user requested deletion")
    status    erased
    counts    turns 1, memories L1 1, collateral 0, duplicates 0
    findings  []
    audit     2 entries, chain_intact True

Output excerpt from examples/tour.py at db3a8d4. flywheel-mneme 0.7.0 is the current release on PyPI.

What it does not do

Source: README.md at db3a8d4, "Install", "Accountability features" and "Accountable forgetting"

Check what stuck

Answer each one in your head before you open it.

Why does the assistant's turn produce no fact?

It is context. The extractor keeps atomic facts from the user's turns, so four turns give three facts.

A receipt's score is edited but its scorer hash is left alone. What does verify_recall return, and why?

False. It re-runs the scorer over the rows and compares the ranking, so a matching hash cannot carry an edited score.

What is the difference between a source edited and a source deleted?

An edited source re-hashes differently, so the memory reads DRIFT. A deleted source cannot be checked at all, so it reads UNVERIFIABLE.

A store has one MATCH, one DRIFT and one UNVERIFIABLE memory. What does the report say?

DRIFT. Any drift wins, then any unverifiable, and MATCH only on a clean sweep.