RECEIPT: First 300 MMLU-Pro at 100% on M0, and M2 recall 300/300 with zero model calls
Built 2026-09-17T06:10:26.976Z by an AI receipt agent (Claude Opus). Read-only except this file and its .json sibling.
What was proven, in plain language
(a) The first 300, 100% on M0: PROVEN, as recall. Our 300 stored Hollerith Cards, run by code with no AI model, picked the keyed answer on 300 of 300 real MMLU-Pro questions on M0. First scored as Row A (2026-09-16, filed about 18:40Z), then re-run fresh at 19:46Z and 19:49Z, again at 20:12Z, and by Hour One Stage 1 at 2026-09-17 01:51Z on five different option orders.
(b) The mesh recall test on M2: PROVEN. At the Founder’s instruction (2026-09-16T19:39Z), the run agent shipped the cards and executor to M2 (HouseViridis, executor host HouseViridisCommand) as 310 signed eblets with no answer key. M2 pulled them into its own silo and ran them by code at 19:49:10Z: 300 of 300, 0 model calls, 0 network attempts, 3.78 ms per question, answers sha256 cf7ed835a5906dc88dfd2db696d18cf89b3c6717552e6158670d03a9473f9bf6, identical to M0, 0 qids differ. Scoring ran on M0 after M2 answers were frozen. Reproduced at 20:14Z (rules moved into addendum cards) and at 21:22Z through the MnemosyneC 0.9.30 app itself.
Both (a) and (b) are proven. The important limit: these are cards fitted to these exact 300 questions, 16 of them repaired with the key. It is exact recall, not blind solving of new questions. See Caveats.
Recount (independent, from stored per-qid rows)
Key matching: exact letter equality between each row letter and the answer field of our_first_300_questions.json for the same qid. The answers sha256 was recomputed from the rows (qid:letter joined by newline) and matched the recorded value in every file.
| run | machine | start UTC | correct | model calls (header / row sum) | network | answers sha256 (recomputed) | same qid set as gold |
|---|---|---|---|---|---|---|---|
| x011_m0_direct | M0 | 2026-09-16T19:46:29.096Z | 300 of 300 | 0 / 0 | 0 | cf7ed835a5906dc8… | true |
| x011_m0_silo | M0 | 2026-09-16T19:49:26.649Z | 300 of 300 | 0 / 0 | 0 | cf7ed835a5906dc8… | true |
| x011_m2 | M2 | 2026-09-16T19:49:10.709Z | 300 of 300 | 0 / 0 | 0 | cf7ed835a5906dc8… | true |
| x011_m2_managed_by_code | M2 | 2026-09-16T19:50:20.444Z | 300 of 300 | 0 / 0 | 0 | cf7ed835a5906dc8… | true |
| x012_m0_rulesOFF_addendaON | M0 | 2026-09-16T20:12:20.101Z | 300 of 300 | 0 / 0 | 0 | cf7ed835a5906dc8… | true |
| x012_m2_rulesOFF_addendaON | M2 | 2026-09-16T20:14:01.393Z | 300 of 300 | 0 / 0 | 0 | cf7ed835a5906dc8… | true |
| x012_m0_rulesOFF_addendaOFF | M0 | 2026-09-16T20:12:21.397Z | 296 of 300 (wrong: 4607 8010 9642 11846) | 0 / 0 | 0 | eea0bc89aab74574… | true |
| x012_app_M2toM0 | M0 | 2026-09-16T21:01:42.890Z | 300 of 300 | 0 / 0 | 0 | cf7ed835a5906dc8… | true |
| x012_app_M0toM2 | M2 | 2026-09-16T21:22:27.097Z | 300 of 300 | 0 / 0 | 0 | cf7ed835a5906dc8… | true |
| hourOne_stage1_k0 | M0 | 2026-09-17T01:51:08Z | 300 of 300 | 0 / 0 | 0 | (shuffled variants per jumble_score.json, not recounted) | true |
Per-qid M0 vs M2 letter comparison:
- x011_m0_silo vs x011_m2: 0 qids differ
- x011_m0_silo vs x011_m2_managed_by_code: 0 qids differ
- x012_m0_rulesOFF_addendaON vs x012_m2_rulesOFF_addendaON: 0 qids differ
- x011_m0_silo vs x012_app_M0toM2: 0 qids differ
- x011_m0_silo vs x012_app_M2toM0: 0 qids differ
Provenance breakdown (identical in every 300/300 run): blind 142 (first 150 formula cards), blind_code 46, blind_with_model 96, key_informed 16. Total blind-labeled 284. “Sourced” and “literal” are not categories in these answer files; A later audit recorded 32 of 300 cards sourced, 268 source_pending.
Gold set check: 300 items, 300 distinct qids, 0 synthetic 99xxxx ids, 0 not in corpus, 0 key mismatches, 0 question-text mismatches, 0 option mismatches against mmlu-pro-test.jsonl (sha256 51d7b60cb50f027a564d57df2e8db671af7f084c34481cfb37285b6d3eb0c1b1). Literal first 300 qids of the corpus: false.
M2 package check: 300 of 300 M2 row card sha256 values match the card files in the package (301 files incl. 1 addendum); 0 cards contain an “answer” field; 0 of 300 shipped questions carry an answer field.
M2 managed run (Gemma 2:2b, thinking off, relay only): answer model calls 0, management calls 300, faithful relays 300 of 300, relayed letters matching gold 300 of 300.
Evidence (sha256 computed at receipt build)
| evidence (run records available to reviewers on request) | bytes | sha256 | what it shows |
|---|---|---|---|
run record our_first_300_questions.json | 281323 | d0d258da766affd18ae52c67f8992ecca7434cc1bd464b79980b2af939445de2 | Gold set: 300 real MMLU-Pro test items with key; scorer gold |
run record our_first_300_provenance.json | 4251 | 4dcd132d02744beb8e83cf99ca9afb2d9e39db3ee8478b833ede7bcca213135a | How the 300 were defined: A1 50 + A2A3 100 + next150 150, all in corpus |
public MMLU-Pro test split, mmlu-pro-test.jsonl | 9830146 | 51d7b60cb50f027a564d57df2e8db671af7f084c34481cfb37285b6d3eb0c1b1 | MMLU-Pro test corpus (12,032 rows) used to verify qids, questions, options, keys |
| run report | 8929 | 8c15f4b6c905f523fc5b9b8f9894af059c37eb59d3e25e56a83452e370ffbada | Row A report: Hollerith Cards only 300/300 on M0, 0 model calls, with honesty labels |
run record harness_band_300.json | 5837 | 7996cf39610f5396d6e575980b197f5b5b66f1b2a475d2707a1fe9178a8ffe3a | Row A data: 300 of 300, blind 284, key-informed 16, modelCalls 0 |
| run record | 11425 | 08a5a34590dce4db5f449326cecf9277ab8ac083179a7648df9a142f3241574f | First 50 re-run models off: 50 of 50 |
| run record | 25634 | 4da2b429f892cd86f502527b73ea33ae1aa9ac4df7108756fb9ade903ba6d93b | Next 100 re-run models off: 100 of 100 (8 key-informed) |
run record scoreboard_300.json | 13403 | 1c2f9dc191ea55159f43446b7ecbd316a48bc141514802633458c6431dc89c78 | Six flagships scored on the same 300 (context only) |
| run report | 11974 | 816152730c1ac2d9aad6a78626446eec1ed2108dbdfe4019cad87708c76ed994 | M2 recall report: M0 300/300, M2 300/300, 0 model calls, same answers sha |
run record answers_m0.json | 94271 | 0635125076c9501a0267975cd883f80239e8cca3241523a638b868f507f2169a | M0 direct code run, per-qid rows (recounted 300/300) |
run record answers_m0_silo.json | 94265 | cb1653e44f91a1d1b811cac4808d626b205bdb4a2bc62f123b237dc3020e1741 | M0 run from its own silo, the scored M0 run (recounted 300/300) |
run record answers_m2.json | 94270 | 59fa71525c65391af9f1b6917d324c583c68e64aaa95baeacbdda0c24715b5fc | M2 (HouseViridisCommand) code run from its own silo (recounted 300/300, 0 model calls) |
run record managed_m2.json | 72147 | fc8badac18f2670d6e50b802ec6a4c8f3e50d05bc0bd26c5e91bb9591ba7825e | M2 Gemma 2:2b relay-only run: 300 management calls, 0 answer calls, 300 faithful relays |
run record managed_m2.answers_by_code.json | 94262 | 663cf6fcdc3dfb7edad585863fb0d16b6a4ef4fd597de7e2f3e4e2f24ea13923 | M2 code answers inside the managed run (recounted 300/300) |
run record pull_recall.receipt.json | 109248 | 893493f186f28ee059ccf7613ed9d82195fd9eaff40ac8891c15785c1f4fb0d0 | M2 mesh pull receipt: listed 310, written 310, refused 0, expected_key trust |
run record score_m0_vs_m2.json | 1329 | 5b39588137739016af2268c730fc90b2e2305fd8eb6345abca03261faa7d3d2c | scorer output: both 300/300, sameAnswersSha true, 0 diffs |
run record score_m0_only.json | 743 | d9c6988839147e90a101c70991af025c3bf3038c6c9736b100eb26b12cbf12ad | scorer output for the first M0 run |
run record score_recall.cjs | 3552 | e512fc6dba58f1019487935f87a8b9e9b577c56bae1b2d387b8e65f5b996c425 | scorer (gold read on M0 after answers frozen) |
run record recall_exec.cjs | 16424 | 5ec0548179375dedfdabc8b1e0371f078a719050ca8b523d2c5494cdf4543072 | Card-only executor with network/model tripwire |
run record card_classes.json | 2652 | 3abc2f193500a5933050a7880fad6360ff92729b94f57a952d1d7667f0947e68 | Card type breakdown: 150 formula, 150 maze, all executable |
run record bridge_serve.log | 792 | a1fc985a10ff61681ede5946c3299a29b7a11e80c8bb54bde1ccc60f0b7a539b | M0 publisher log: 1 search + 9 fetches, all 200, trust enrolled |
run record MANIFEST.json | 96753 | abb77aa4bd667fa0681023eb6c8570d4e5809df17ab0c903ef96038f26097fcc | Package manifest sha abb77aa4… shared by M0 and M2 runs |
run record questions_300_nogold.json | 239127 | f683b6a86009a92ecc6937e9a0cad17b00ce913b628d8a2b2008ea89ef65f98f | Questions shipped to M2 without any answer field (verified 0 of 300) |
| run report | 20650 | 17ae3d338da452f9d64e4525cd50fdc067a6d079390c753d6c759b60f3531cf6 | Addenda for 4 rules, 0.9.30 app route both directions |
run record m0_addenda_proof.json | 3583 | 93843800261387140f194958371ba6bb851e86b06078df9742967ba9b0da77df | Rules/addenda arms: ON_ON 300, OFF_ON 300, OFF_OFF 296 |
run record score_OFF_ON_m0_vs_m2.json | 1590 | 2f66803615f986be717d47a6882afeaa317907592034756ec4e9e69fb7d0dba9 | Runner rules off, addenda on: M0 and M2 both 300/300 |
run record answers_m0_OFF_ON.json | 95474 | 4ebf4fd93fc2060c47a0f9d95c295306f68dfab9020a3dbbe3baed4e6a32f71a | M0 runner rules off, addenda on (recounted 300/300) |
run record answers_m0_OFF_OFF.json | 94658 | 224b4b36fb8ed3dc3886a6a3dd7f54154e63025774d987cb596d1e87f6fe649d | M0 runner rules off, addenda off (recounted 296/300: 4607 8010 9642 11846) |
run record answers_m2_OFF_ON.json | 95489 | 8b3ab0513ed556c72fb251673cf2a6d3ea8cae22b5ab03108a9f30df9d91b769 | M2 runner rules off, addenda on (recounted 300/300) |
run record answers_m0_via_m2_app.json | 95471 | 7f63f573b69baa0f958ec9febd7e7e626207880487f704f43b60c1637bf262f5 | App-to-app route M2 to M0: cards pulled through MnemosyneC 0.9.30 (recounted 300/300) |
run record score.json | 899 | 5f0ed931811e023da7371d06be4ae73570a5c13299331b46ed5433f2fbd04be0 | score for the app route |
run record answers_m2_via_m0_app.json | 95471 | 9d360ffb1a6b188df148b8f1b6c0f582306bb4817f5ad4c73009df3b59479bfc | App-to-app route M0 to M2: 314 files pulled from M0 app (recounted 300/300) |
run record app_route_result.json | 1965 | d5265545ace81a054d792253be472b6a80418ba2b70f1339c031570ba4bd9cf1 | M0 to M2 app route result: pull 314/314, refused 0 |
run record pull.receipt.json | 110663 | 172798ab37543204c9541a9c00f3eadd85f324addcc5951bbc901b2dab3b6dca | M0 to M2 app pull receipt (publisher trust self_declared) |
run record score.json | 906 | a4aa1626268945ee4ce1215237f6aaf58c07eeaff2761202f6dba82590a8e612 | score for the app route |
run record jumble_score.json | 1279 | 13aa3640ed3e1f2679f2aedd282e86ba594ca1135d06a44e9e91f69bb371b2c8 | Hour One Stage 1 (2026-09-17 01:51Z, M0): 300/300 on all 5 option orders, 0 calls |
run record answers_k0.json | 117153 | e02596d12fbee0340ce80ee62f90824c44c113bdc653e359cf35f304331900c0 | Hour One Stage 1 unshuffled order answers (recounted 300/300) |
run record summary.json | 869 | ab15d75a7cf3a351a4faf564f680c942fcb011186a8187408081ec67e3450238 | Hour One summary: recall 300/300; separate new-100 stage 83 blind, 90 final |
| session log | 21499 | fcaa2936cf15cb01cf47e017489761485587bb9da081d3bfd8d7411156d9ccc4 | session log: Row A line 152 |
| session log | 37011 | 6d3386a5c69da59d43747dbaf48591ddd63d6fe1bc1e03bee638d71f50f5f8b2 | session log: DELTA 21 (ruling), DELTA 26 (landed), DELTA 50/56 (addenda, app route) |
| session log | 57384 | 740faffb3b9832a484a105211ac643437d08c1c7d32e4d3950f57782d04ebe8f | session log: DELTA 19 (three numbers), DELTA 84 (Hour One) |
| session transcript | 9557051 | fb201bd793c5634f22b5a8e7cbd284107b2485d93565ffd9b62a65bfa5f480c2 | session transcript (ruling line 1215, run-agent result line 1393, Conduit post line 1402) |
Verbatim lines from the record
Message ids and file locations are omitted on this page. The message log and session transcripts these lines come from are available to reviewers on request.
- Founder, 2026-09-16T19:39:44.057Z (UNSIGNED, composed at the conduit terminal): “We are testing if we install all our Hollerith Cards, when we get 300/300 correct on M0, we should be getting 300/300 correct on M2, WITHOUT HAVING TO run a reasoning engine at all - at that point it’s just recall, so it needs to get the answer, after following the Code Maze on the Hollerith Card.”
- Founder, 2026-09-16T19:43:08.228Z (UNSIGNED): “The point is, the Hollerith Cards enable M2 to get the answers right at 100% accuracy, fast, and cheap. Make it do that, please.”
- Founder, 2026-09-16T19:05:21.065Z (UNSIGNED): “You can do whatever is needed on M2, nothing is being run there that needs running.”
- Run agent (Claude Opus), progress note step 4, 2026-09-16T19:50:20Z: “M2 pulled 310 eblets over the LAN route, unpacked from its own silo, ran 300 by code: 300 of 300, 0 model calls, 3.78 ms per question, answers sha identical to M0.”
- Run agent result to the coordinating agent, task notification, 2026-09-16T19:52:43.211Z: “M0 300/300, M2 300/300, model calls M0 0, M2 0, ms per question M0 4.38, M2 3.78.”
- Coordinating agent to Founder (Conduit), 2026-09-16T19:52:57.157Z: “[session label]: THE SUBSTRATE WAY WORKS ON M2. 300 of 300.”
- Founder, 2026-09-16T20:05:53.446Z (UNSIGNED): “RE: “[session label]: THE SUBSTRATE WAY WORKS ON M2. 300 of 300.” Do that to the other 4 as well then.”
- Coordinating agent to Founder, 2026-09-16T21:23:14.204Z: “M2 pulled all the recall cards from M0’s own app (314 of 314 files, 0 refused) and ran them: 300/300, 0 model calls, same answers as M0.”
- Coordinating agent to Founder, 2026-09-16T23:22:53.487Z: “PROVEN ACCURATE: 100%. 300 of 300 right, run by code with zero AI calls, identical on two machines (M0 and M2).”
- Coordinating agent to Founder, 2026-09-17T02:11:26.656Z: “Stage 1, recall of our 300 carded questions, answers shuffled 5 different ways: 300 of 300 every time. 0 AI calls, $0, 12 seconds.”
- Founder, (relayed in the instruction to build this receipt), 2026-09-17: “We ran it. IT succeeded. Read the Conduit Session Transcripts from [the working sessions] and then make a receipt of it.”
Source: the instruction to build this receipt; NOT found in the message log on disk at build time.
Caveats, stated plainly
- “First 300” is the Founder-defined set “our first 300” (A1 50 + A2A3 100 + next150 150), NOT the literal first 300 question_ids of MMLU-Pro. All 300 are real MMLU-Pro test items (lowest qid 70, highest 12205, 0 synthetic 99xxxx ids); every qid, question text, option list and key matches mmlu-pro-test.jsonl exactly. Category mix is skewed: law 72, business 59.
- 16 of the 300 are KEY_INFORMED: the card was repaired using the answer key after a blind miss, counted correct by ruling and labeled on every row. 284 are labeled blind, but the first 150 cards were built across many dispatches with rework scored against the key (recorded in the card-building and scoring reports, label 3), and 96 of the maze cards are “blind_with_model” (an architect reading stored at mint time, before the key was read). This is 100% RECALL of cards fitted to these 300 questions, not 100% blind first-time solving.
- A card fits one question. None of these runs measures a new question. The same night, 100 new real questions (Hour One Stage 2) scored 83 blind first pick and 90 final after a key-informed loop.
- Card gaps were closed during the proof: first card-only pass scored 295; qid 3674 got an ADDENDUM card (rule from a pre-gold wave spec), and four runner-only rules (4607, 8010, 9642, 11846) were moved into signed addendum cards in a follow-up run. With runner rules and addenda both off the M0 score is 296 (recounted here).
- M2 zero model calls rests on the executor tripwire counters recorded in the M2 answer files (header modelCalls 0, networkAttempts 0, per-row modelCalls sum 0, recounted) plus the run report of Ollama /api/ps empty before and after and no /api/generate or /api/chat in M2 server.log. The M2 Ollama log itself is not on M0 disk and this receipt did not access M2. The host field “HouseViridisCommand” is self-reported by the executor.
- The first M2 transport (19:49Z) used the peer_node.cjs helper publisher, not the MnemosyneC app; the 40-line unpack bootstrap and managed_run.cjs were placed on M2 by WinRM, not the mesh. App-to-app transport through MnemosyneC 0.9.30 was proven later (21:01Z M2 to M0 and 21:22Z M0 to M2), where the M0 to M2 pull recorded publisher trust “self_declared”.
- The Gemma 2:2b managed run on M2 made 300 model calls to RELAY the code answer to the user (63,374 in / 1,807 out tokens, $0 local). Those calls did not produce the answers; the answer sha is identical to the code-only run.
- Row A 5.81 ms per question was stitched from older timings; fresh measured runs are 4.38 ms (M0) and 3.78 ms (M2), wall time over 300 including module load.
- M0 and M2 clocks differ by about one second (M2 run start 19:49:10.709Z by M2 clock while the M0 publisher logged fetches until 19:49:11.136Z).
- This receipt recounted stored outputs; it did not re-run the executor on either machine.
- The earlier session transcripts hold the card-building waves and ground-set definition; they were not re-read per qid. The run evidence lives in three later session records (Row A; M2 recall; Hour One, three-number relay).
- The Founder line of 2026-09-17 (“We ran it. IT succeeded.”) is recorded from the instruction to build this receipt; it was not present in the message log on disk when the receipt was built.
- Founder Conduit messages carry provenance UNSIGNED (composed at the local conduit terminal).
Public-safe claim sentence
On a fixed set of 300 real MMLU-Pro test questions, stored Hollerith Cards run by plain code answered 300 of 300 correctly with zero AI model calls, on two separate machines, with identical answers; the cards were built for these specific questions (16 of them repaired using the answer key after a first miss), so this shows exact, repeatable recall, not accuracy on new questions.
Verifier scripts (read-only, rerunnable)
Five read-only verifier scripts were used to build and check this receipt: verify_receipt.cjs, hash_and_k0.cjs, extract_lines.cjs, jsonl_hits.cjs and build_receipt.cjs. They are available to reviewers on request.