RECEIPT: free open models on our first 300, and the SEALED PREDICTION for the next 700

Written 2026-09-17T06:45:49Z. Scores recomputed now from raw rows (latest HTTP 200 lettered row per qid).

Sealed prediction

  • SHA-256 4551cb9a377385b15c867f419cb20a0a88691da301cb580a7476458321694c1d, hashed 2026-09-16T20:39:38Z, file set read-only.
  • Re-hashed now: MATCHES, seal intact.
  • Covers: gemma4:12b rest of the 300; the next 700 (pre-chosen, seed 700223045, 0 overlap with the 300) for the six flagships, gemma4:12b, and the substrate way on M0 and M2.
  • Headline sealed numbers (from the sealing report): 700 flagships plain Anthropic 92.5, OpenAI 90.3, Google 90.1, xAI 90.1, DeepSeek 87.3, Mistral 64.2 (about $148); gemma4:12b thinking ON 73.2 (53.5 to 87.3); substrate way final 690 of 700 (674 to 697), blind 648, 0 model calls at replay, M2 answers identical to M0.
  • Disclosed inside: 80 of the 700 were already minted and scored the substrate way on 2026-09-15 (79 of 80).
  • Envelope stays closed until the runs it predicts are complete.

Free open models on the same 300 (no cost)

modelansweredcorrectaccuracystatus
llama3.1:8b (local, M0)3009531.7%complete
gemma4:12b thinking ON (local, M0)1288868.8%PARTIAL, run not finished
gemma4:12b thinking OFF (local, M0)392871.8%stopped by the Founder
gemma2:2b (local, M2)190 strict4423.2%separate run on M2

Paid six on the same 300: see six flagships on the first 300. Our cards by code: the first 300 by code.

Evidence files (sha256)

evidence (run records available to reviewers on request)bytessha256what
sealed prediction file191264551cb9a377385b15c867f419cb20a0a88691da301cb580a7476458321694c1dSEALED PREDICTION (read-only), gemma rest-of-300 and next 700, hashed 2026-09-16T20:39:38Z
sealed prediction hash receipt481683b3ba480893656dd0f189053cd2d8ae0298cb02a890369dd1b6db28a8132414hash receipt for the sealed prediction
run record PUBLIC_HASH_LINE.txt212a6356bbd288489718e4930677d81b620aa6ece7e9d36cbe885126eda26128a4epublic hash line
run record seal_hash_time.txt92f874359274391cc9bf2d8cb0f0e5b43e1913b6e07b54b29eaed790f27ab5ea45seal hash and time
run record prediction_calc.json972289bf1239b6578aa77493b6e53c1a4d85d4b75e3191d4670f60336089c2d1d190prediction math inputs/outputs
run record substrate_calc.json386888ebcd69478ba00e0277fed8e3ecc74a3f4ab7aa2f8025384349a355ae247dcfsubstrate-arm prediction math
run report69471c6a5a17310ba2264b57521457a7b09b4ec61da5d308f8707aee1e683ccf3ecbsealing report
run record next700_qids.json988078e6fc82f67fc1a9e50b70eea9db550cdb08907d99adbad8feaa662ab61090807the pre-chosen next 700 qids (seed 700223045)
run record scoreboard_nine_300.json10131573b2d6f5d95aa007275c4ab8f0c125dad7dc79f15ab65dc2d66f086db3ca4a8fnine-model scoreboard (19:10Z; open models partial at that time)
run record score_nine_300.cjs12385df29b0ef50dc67d21275ba355378ece654edcb0c606e1e152db4a6db9e2705f8nine-model scorer
run record open300_ollama_llama31_8b.jsonl178380685091db2f427659c403c5641f57a66e40159729aff179caaa3f06636fde52cd4raw rows: llama3.1:8b (local, M0)
run record open300_ollama_gemma4_12b.jsonl.partial6124445c75a60388e4fa5273e9300c063d2081d18dca9f2275076999f93ab064e2fa61craw rows: gemma4:12b thinking ON (local, M0)
run record plain300_gemma4_12b_thinkoff_M0.jsonl.partial24337367be94f370d105e3a648c36ff91f1af2de9dacfe8c823236938326bc3e64964fraw rows: gemma4:12b thinking OFF (local, M0)
run record open300_ollama_gemma4_12b.spend.json57206fe560ad421c1e3d98cb2627a7a80a4bd91fe41c3e0d7ed7c062697507785e9meter / audit / log
run record open300_ollama_llama31_8b.spend.json619942dcef9c65139067879af90d4988d9fdc9e77bd88b01c1e76264282dcf4f9ddmeter / audit / log
run record open300_ollama_gemma4_12b.DENOMINATOR_AUDIT.json361764722383b79be4ca005a95f24b4300e51d91ac78e2fc4706198afc6c7b02572meter / audit / log
run record open300_ollama_llama31_8b.DENOMINATOR_AUDIT.json3624ee1731bf0bf2ddaaa7d7dfcfc646ad496e3eccdd7a0e802e453030dc71c4a7dmeter / audit / log
run record open300_ollama_gemma4_12b.log593033eab4652a4baa059cc85acecb04aad1521d9a947b1f2b7747f8108d254f1ed79meter / audit / log
run record open300_ollama_gemma4_12b_THINKING_ON_PARTIAL.log9790b9a684459d7f7ea2a94f0227f2c06fb165671773f63b84a986c5ae741f9a0e9meter / audit / log
run record open300_ollama_llama31_8b.log259881f0babfb623d561ad907fe03582a9c8f23517f4053c5a2dca4a9b3502679ecc0meter / audit / log
run record validate_2q_ollama_llama33_70b.jsonl11466f9c32b4430ac8139e78849e64f2c899f412d958139aae64d28e528210f714239meter / audit / log
run record plain300_gemma4_12b_thinkoff_M0.spend.json581b75d13c822185b0018022994b61a91da43b29d2d7fc81fc064ecf9ebf0b9c4e5meter / audit / log
run record plain300_hosted_gemma4_26b_thinkon.spend.json585edf142de1bcf5a5effcd60f6bb17185a457dddc755a2e10bcd2cd240025311d3meter / audit / log
run record scoreboard_partial_gemma4_thinkon.json10061713b4c7d747377d1b54795d1cc08c50937afbd7cfc79e2ade691e942c7d67f2d3meter / audit / log
run report18655558ccfeb6e5692e2e95627534be1d8dc2d68d749343417aee78bc6493764e070gemma2:2b on M2 (190 strict, 23.2% plain, 38.4% model reads method text)