RECEIPT: First 300 MMLU-Pro at 100% on M0, and M2 recall 300/300 with zero model calls

Built 2026-09-17T06:10:26.976Z by an AI receipt agent (Claude Opus). Read-only except this file and its .json sibling.

What was proven, in plain language

(a) The first 300, 100% on M0: PROVEN, as recall. Our 300 stored Hollerith Cards, run by code with no AI model, picked the keyed answer on 300 of 300 real MMLU-Pro questions on M0. First scored as Row A (2026-09-16, filed about 18:40Z), then re-run fresh at 19:46Z and 19:49Z, again at 20:12Z, and by Hour One Stage 1 at 2026-09-17 01:51Z on five different option orders.

(b) The mesh recall test on M2: PROVEN. At the Founder’s instruction (2026-09-16T19:39Z), the run agent shipped the cards and executor to M2 (HouseViridis, executor host HouseViridisCommand) as 310 signed eblets with no answer key. M2 pulled them into its own silo and ran them by code at 19:49:10Z: 300 of 300, 0 model calls, 0 network attempts, 3.78 ms per question, answers sha256 cf7ed835a5906dc88dfd2db696d18cf89b3c6717552e6158670d03a9473f9bf6, identical to M0, 0 qids differ. Scoring ran on M0 after M2 answers were frozen. Reproduced at 20:14Z (rules moved into addendum cards) and at 21:22Z through the MnemosyneC 0.9.30 app itself.

Both (a) and (b) are proven. The important limit: these are cards fitted to these exact 300 questions, 16 of them repaired with the key. It is exact recall, not blind solving of new questions. See Caveats.

Recount (independent, from stored per-qid rows)

Key matching: exact letter equality between each row letter and the answer field of our_first_300_questions.json for the same qid. The answers sha256 was recomputed from the rows (qid:letter joined by newline) and matched the recorded value in every file.

runmachinestart UTCcorrectmodel calls (header / row sum)networkanswers sha256 (recomputed)same qid set as gold
x011_m0_directM02026-09-16T19:46:29.096Z300 of 3000 / 00cf7ed835a5906dc8…true
x011_m0_siloM02026-09-16T19:49:26.649Z300 of 3000 / 00cf7ed835a5906dc8…true
x011_m2M22026-09-16T19:49:10.709Z300 of 3000 / 00cf7ed835a5906dc8…true
x011_m2_managed_by_codeM22026-09-16T19:50:20.444Z300 of 3000 / 00cf7ed835a5906dc8…true
x012_m0_rulesOFF_addendaONM02026-09-16T20:12:20.101Z300 of 3000 / 00cf7ed835a5906dc8…true
x012_m2_rulesOFF_addendaONM22026-09-16T20:14:01.393Z300 of 3000 / 00cf7ed835a5906dc8…true
x012_m0_rulesOFF_addendaOFFM02026-09-16T20:12:21.397Z296 of 300 (wrong: 4607 8010 9642 11846)0 / 00eea0bc89aab74574…true
x012_app_M2toM0M02026-09-16T21:01:42.890Z300 of 3000 / 00cf7ed835a5906dc8…true
x012_app_M0toM2M22026-09-16T21:22:27.097Z300 of 3000 / 00cf7ed835a5906dc8…true
hourOne_stage1_k0M02026-09-17T01:51:08Z300 of 3000 / 00(shuffled variants per jumble_score.json, not recounted)true

Per-qid M0 vs M2 letter comparison:

  • x011_m0_silo vs x011_m2: 0 qids differ
  • x011_m0_silo vs x011_m2_managed_by_code: 0 qids differ
  • x012_m0_rulesOFF_addendaON vs x012_m2_rulesOFF_addendaON: 0 qids differ
  • x011_m0_silo vs x012_app_M0toM2: 0 qids differ
  • x011_m0_silo vs x012_app_M2toM0: 0 qids differ

Provenance breakdown (identical in every 300/300 run): blind 142 (first 150 formula cards), blind_code 46, blind_with_model 96, key_informed 16. Total blind-labeled 284. “Sourced” and “literal” are not categories in these answer files; A later audit recorded 32 of 300 cards sourced, 268 source_pending.

Gold set check: 300 items, 300 distinct qids, 0 synthetic 99xxxx ids, 0 not in corpus, 0 key mismatches, 0 question-text mismatches, 0 option mismatches against mmlu-pro-test.jsonl (sha256 51d7b60cb50f027a564d57df2e8db671af7f084c34481cfb37285b6d3eb0c1b1). Literal first 300 qids of the corpus: false.

M2 package check: 300 of 300 M2 row card sha256 values match the card files in the package (301 files incl. 1 addendum); 0 cards contain an “answer” field; 0 of 300 shipped questions carry an answer field.

M2 managed run (Gemma 2:2b, thinking off, relay only): answer model calls 0, management calls 300, faithful relays 300 of 300, relayed letters matching gold 300 of 300.

Evidence (sha256 computed at receipt build)

evidence (run records available to reviewers on request)bytessha256what it shows
run record our_first_300_questions.json281323d0d258da766affd18ae52c67f8992ecca7434cc1bd464b79980b2af939445de2Gold set: 300 real MMLU-Pro test items with key; scorer gold
run record our_first_300_provenance.json42514dcd132d02744beb8e83cf99ca9afb2d9e39db3ee8478b833ede7bcca213135aHow the 300 were defined: A1 50 + A2A3 100 + next150 150, all in corpus
public MMLU-Pro test split, mmlu-pro-test.jsonl983014651d7b60cb50f027a564d57df2e8db671af7f084c34481cfb37285b6d3eb0c1b1MMLU-Pro test corpus (12,032 rows) used to verify qids, questions, options, keys
run report89298c15f4b6c905f523fc5b9b8f9894af059c37eb59d3e25e56a83452e370ffbadaRow A report: Hollerith Cards only 300/300 on M0, 0 model calls, with honesty labels
run record harness_band_300.json58377996cf39610f5396d6e575980b197f5b5b66f1b2a475d2707a1fe9178a8ffe3aRow A data: 300 of 300, blind 284, key-informed 16, modelCalls 0
run record1142508a5a34590dce4db5f449326cecf9277ab8ac083179a7648df9a142f3241574fFirst 50 re-run models off: 50 of 50
run record256344da2b429f892cd86f502527b73ea33ae1aa9ac4df7108756fb9ade903ba6d93bNext 100 re-run models off: 100 of 100 (8 key-informed)
run record scoreboard_300.json134031c2f9dc191ea55159f43446b7ecbd316a48bc141514802633458c6431dc89c78Six flagships scored on the same 300 (context only)
run report11974816152730c1ac2d9aad6a78626446eec1ed2108dbdfe4019cad87708c76ed994M2 recall report: M0 300/300, M2 300/300, 0 model calls, same answers sha
run record answers_m0.json942710635125076c9501a0267975cd883f80239e8cca3241523a638b868f507f2169aM0 direct code run, per-qid rows (recounted 300/300)
run record answers_m0_silo.json94265cb1653e44f91a1d1b811cac4808d626b205bdb4a2bc62f123b237dc3020e1741M0 run from its own silo, the scored M0 run (recounted 300/300)
run record answers_m2.json9427059fa71525c65391af9f1b6917d324c583c68e64aaa95baeacbdda0c24715b5fcM2 (HouseViridisCommand) code run from its own silo (recounted 300/300, 0 model calls)
run record managed_m2.json72147fc8badac18f2670d6e50b802ec6a4c8f3e50d05bc0bd26c5e91bb9591ba7825eM2 Gemma 2:2b relay-only run: 300 management calls, 0 answer calls, 300 faithful relays
run record managed_m2.answers_by_code.json94262663cf6fcdc3dfb7edad585863fb0d16b6a4ef4fd597de7e2f3e4e2f24ea13923M2 code answers inside the managed run (recounted 300/300)
run record pull_recall.receipt.json109248893493f186f28ee059ccf7613ed9d82195fd9eaff40ac8891c15785c1f4fb0d0M2 mesh pull receipt: listed 310, written 310, refused 0, expected_key trust
run record score_m0_vs_m2.json13295b39588137739016af2268c730fc90b2e2305fd8eb6345abca03261faa7d3d2cscorer output: both 300/300, sameAnswersSha true, 0 diffs
run record score_m0_only.json743d9c6988839147e90a101c70991af025c3bf3038c6c9736b100eb26b12cbf12adscorer output for the first M0 run
run record score_recall.cjs3552e512fc6dba58f1019487935f87a8b9e9b577c56bae1b2d387b8e65f5b996c425scorer (gold read on M0 after answers frozen)
run record recall_exec.cjs164245ec0548179375dedfdabc8b1e0371f078a719050ca8b523d2c5494cdf4543072Card-only executor with network/model tripwire
run record card_classes.json26523abc2f193500a5933050a7880fad6360ff92729b94f57a952d1d7667f0947e68Card type breakdown: 150 formula, 150 maze, all executable
run record bridge_serve.log792a1fc985a10ff61681ede5946c3299a29b7a11e80c8bb54bde1ccc60f0b7a539bM0 publisher log: 1 search + 9 fetches, all 200, trust enrolled
run record MANIFEST.json96753abb77aa4bd667fa0681023eb6c8570d4e5809df17ab0c903ef96038f26097fccPackage manifest sha abb77aa4… shared by M0 and M2 runs
run record questions_300_nogold.json239127f683b6a86009a92ecc6937e9a0cad17b00ce913b628d8a2b2008ea89ef65f98fQuestions shipped to M2 without any answer field (verified 0 of 300)
run report2065017ae3d338da452f9d64e4525cd50fdc067a6d079390c753d6c759b60f3531cf6Addenda for 4 rules, 0.9.30 app route both directions
run record m0_addenda_proof.json358393843800261387140f194958371ba6bb851e86b06078df9742967ba9b0da77dfRules/addenda arms: ON_ON 300, OFF_ON 300, OFF_OFF 296
run record score_OFF_ON_m0_vs_m2.json15902f66803615f986be717d47a6882afeaa317907592034756ec4e9e69fb7d0dba9Runner rules off, addenda on: M0 and M2 both 300/300
run record answers_m0_OFF_ON.json954744ebf4fd93fc2060c47a0f9d95c295306f68dfab9020a3dbbe3baed4e6a32f71aM0 runner rules off, addenda on (recounted 300/300)
run record answers_m0_OFF_OFF.json94658224b4b36fb8ed3dc3886a6a3dd7f54154e63025774d987cb596d1e87f6fe649dM0 runner rules off, addenda off (recounted 296/300: 4607 8010 9642 11846)
run record answers_m2_OFF_ON.json954898b3ab0513ed556c72fb251673cf2a6d3ea8cae22b5ab03108a9f30df9d91b769M2 runner rules off, addenda on (recounted 300/300)
run record answers_m0_via_m2_app.json954717f63f573b69baa0f958ec9febd7e7e626207880487f704f43b60c1637bf262f5App-to-app route M2 to M0: cards pulled through MnemosyneC 0.9.30 (recounted 300/300)
run record score.json8995f0ed931811e023da7371d06be4ae73570a5c13299331b46ed5433f2fbd04be0score for the app route
run record answers_m2_via_m0_app.json954719d360ffb1a6b188df148b8f1b6c0f582306bb4817f5ad4c73009df3b59479bfcApp-to-app route M0 to M2: 314 files pulled from M0 app (recounted 300/300)
run record app_route_result.json1965d5265545ace81a054d792253be472b6a80418ba2b70f1339c031570ba4bd9cf1M0 to M2 app route result: pull 314/314, refused 0
run record pull.receipt.json110663172798ab37543204c9541a9c00f3eadd85f324addcc5951bbc901b2dab3b6dcaM0 to M2 app pull receipt (publisher trust self_declared)
run record score.json906a4aa1626268945ee4ce1215237f6aaf58c07eeaff2761202f6dba82590a8e612score for the app route
run record jumble_score.json127913aa3640ed3e1f2679f2aedd282e86ba594ca1135d06a44e9e91f69bb371b2c8Hour One Stage 1 (2026-09-17 01:51Z, M0): 300/300 on all 5 option orders, 0 calls
run record answers_k0.json117153e02596d12fbee0340ce80ee62f90824c44c113bdc653e359cf35f304331900c0Hour One Stage 1 unshuffled order answers (recounted 300/300)
run record summary.json869ab15d75a7cf3a351a4faf564f680c942fcb011186a8187408081ec67e3450238Hour One summary: recall 300/300; separate new-100 stage 83 blind, 90 final
session log21499fcaa2936cf15cb01cf47e017489761485587bb9da081d3bfd8d7411156d9ccc4session log: Row A line 152
session log370116d3386a5c69da59d43747dbaf48591ddd63d6fe1bc1e03bee638d71f50f5f8b2session log: DELTA 21 (ruling), DELTA 26 (landed), DELTA 50/56 (addenda, app route)
session log57384740faffb3b9832a484a105211ac643437d08c1c7d32e4d3950f57782d04ebe8fsession log: DELTA 19 (three numbers), DELTA 84 (Hour One)
session transcript9557051fb201bd793c5634f22b5a8e7cbd284107b2485d93565ffd9b62a65bfa5f480c2session transcript (ruling line 1215, run-agent result line 1393, Conduit post line 1402)

Verbatim lines from the record

Message ids and file locations are omitted on this page. The message log and session transcripts these lines come from are available to reviewers on request.

  • Founder, 2026-09-16T19:39:44.057Z (UNSIGNED, composed at the conduit terminal): “We are testing if we install all our Hollerith Cards, when we get 300/300 correct on M0, we should be getting 300/300 correct on M2, WITHOUT HAVING TO run a reasoning engine at all - at that point it’s just recall, so it needs to get the answer, after following the Code Maze on the Hollerith Card.”
  • Founder, 2026-09-16T19:43:08.228Z (UNSIGNED): “The point is, the Hollerith Cards enable M2 to get the answers right at 100% accuracy, fast, and cheap. Make it do that, please.”
  • Founder, 2026-09-16T19:05:21.065Z (UNSIGNED): “You can do whatever is needed on M2, nothing is being run there that needs running.”
  • Run agent (Claude Opus), progress note step 4, 2026-09-16T19:50:20Z: “M2 pulled 310 eblets over the LAN route, unpacked from its own silo, ran 300 by code: 300 of 300, 0 model calls, 3.78 ms per question, answers sha identical to M0.”
  • Run agent result to the coordinating agent, task notification, 2026-09-16T19:52:43.211Z: “M0 300/300, M2 300/300, model calls M0 0, M2 0, ms per question M0 4.38, M2 3.78.”
  • Coordinating agent to Founder (Conduit), 2026-09-16T19:52:57.157Z: “[session label]: THE SUBSTRATE WAY WORKS ON M2. 300 of 300.”
  • Founder, 2026-09-16T20:05:53.446Z (UNSIGNED): “RE: “[session label]: THE SUBSTRATE WAY WORKS ON M2. 300 of 300.” Do that to the other 4 as well then.”
  • Coordinating agent to Founder, 2026-09-16T21:23:14.204Z: “M2 pulled all the recall cards from M0’s own app (314 of 314 files, 0 refused) and ran them: 300/300, 0 model calls, same answers as M0.”
  • Coordinating agent to Founder, 2026-09-16T23:22:53.487Z: “PROVEN ACCURATE: 100%. 300 of 300 right, run by code with zero AI calls, identical on two machines (M0 and M2).”
  • Coordinating agent to Founder, 2026-09-17T02:11:26.656Z: “Stage 1, recall of our 300 carded questions, answers shuffled 5 different ways: 300 of 300 every time. 0 AI calls, $0, 12 seconds.”
  • Founder, (relayed in the instruction to build this receipt), 2026-09-17: “We ran it. IT succeeded. Read the Conduit Session Transcripts from [the working sessions] and then make a receipt of it.”
    Source: the instruction to build this receipt; NOT found in the message log on disk at build time.

Caveats, stated plainly

  1. “First 300” is the Founder-defined set “our first 300” (A1 50 + A2A3 100 + next150 150), NOT the literal first 300 question_ids of MMLU-Pro. All 300 are real MMLU-Pro test items (lowest qid 70, highest 12205, 0 synthetic 99xxxx ids); every qid, question text, option list and key matches mmlu-pro-test.jsonl exactly. Category mix is skewed: law 72, business 59.
  2. 16 of the 300 are KEY_INFORMED: the card was repaired using the answer key after a blind miss, counted correct by ruling and labeled on every row. 284 are labeled blind, but the first 150 cards were built across many dispatches with rework scored against the key (recorded in the card-building and scoring reports, label 3), and 96 of the maze cards are “blind_with_model” (an architect reading stored at mint time, before the key was read). This is 100% RECALL of cards fitted to these 300 questions, not 100% blind first-time solving.
  3. A card fits one question. None of these runs measures a new question. The same night, 100 new real questions (Hour One Stage 2) scored 83 blind first pick and 90 final after a key-informed loop.
  4. Card gaps were closed during the proof: first card-only pass scored 295; qid 3674 got an ADDENDUM card (rule from a pre-gold wave spec), and four runner-only rules (4607, 8010, 9642, 11846) were moved into signed addendum cards in a follow-up run. With runner rules and addenda both off the M0 score is 296 (recounted here).
  5. M2 zero model calls rests on the executor tripwire counters recorded in the M2 answer files (header modelCalls 0, networkAttempts 0, per-row modelCalls sum 0, recounted) plus the run report of Ollama /api/ps empty before and after and no /api/generate or /api/chat in M2 server.log. The M2 Ollama log itself is not on M0 disk and this receipt did not access M2. The host field “HouseViridisCommand” is self-reported by the executor.
  6. The first M2 transport (19:49Z) used the peer_node.cjs helper publisher, not the MnemosyneC app; the 40-line unpack bootstrap and managed_run.cjs were placed on M2 by WinRM, not the mesh. App-to-app transport through MnemosyneC 0.9.30 was proven later (21:01Z M2 to M0 and 21:22Z M0 to M2), where the M0 to M2 pull recorded publisher trust “self_declared”.
  7. The Gemma 2:2b managed run on M2 made 300 model calls to RELAY the code answer to the user (63,374 in / 1,807 out tokens, $0 local). Those calls did not produce the answers; the answer sha is identical to the code-only run.
  8. Row A 5.81 ms per question was stitched from older timings; fresh measured runs are 4.38 ms (M0) and 3.78 ms (M2), wall time over 300 including module load.
  9. M0 and M2 clocks differ by about one second (M2 run start 19:49:10.709Z by M2 clock while the M0 publisher logged fetches until 19:49:11.136Z).
  10. This receipt recounted stored outputs; it did not re-run the executor on either machine.
  11. The earlier session transcripts hold the card-building waves and ground-set definition; they were not re-read per qid. The run evidence lives in three later session records (Row A; M2 recall; Hour One, three-number relay).
  12. The Founder line of 2026-09-17 (“We ran it. IT succeeded.”) is recorded from the instruction to build this receipt; it was not present in the message log on disk when the receipt was built.
  13. Founder Conduit messages carry provenance UNSIGNED (composed at the local conduit terminal).

Public-safe claim sentence

On a fixed set of 300 real MMLU-Pro test questions, stored Hollerith Cards run by plain code answered 300 of 300 correctly with zero AI model calls, on two separate machines, with identical answers; the cards were built for these specific questions (16 of them repaired using the answer key after a first miss), so this shows exact, repeatable recall, not accuracy on new questions.

Verifier scripts (read-only, rerunnable)

Five read-only verifier scripts were used to build and check this receipt: verify_receipt.cjs, hash_and_k0.cjs, extract_lines.cjs, jsonl_hits.cjs and build_receipt.cjs. They are available to reviewers on request.