# Pre-registration v1 — "Judging Verified Truth" Timestamp of freeze: 2026-08-23 (before confirmatory statistics; SHA-256 committed below). ## Data (frozen snapshot) - Corpus: 71 papers + 607 public reviews fetched from api.recensorium.com/v1 on 2026-08-23 00:10–01:30 UTC (files: corpus_all.json, all_reviews.json, rev_table.json). - Ground-truth instrument: 18 mathematics-statistics papers making computationally checkable central claims. - 12 constant-weight-code papers: complete codeword witnesses re-checked mechanically (well-formedness, pairwise distance, claimed size, Schoenheim bound). Result at freeze time: 12/12 witnesses VALID, all central claims TRUE, 4 attaining the Schoenheim bound (= optimality established). - 4 exhaustion papers: Z_82/R(3,16) multiplier-invariant family independently re-enumerated (2047/2047 candidates contain a violation, 0 unresolved) — replicated. Remaining cells re-verified where compute completes before submission; otherwise excluded and reported as such. - No paper or review may be added/removed after freeze except exhaustion-cell verdicts governed by the rule above. ## Hypotheses (fixed now) - H1 (calibration level): mean rigour assigned to proven-TRUE papers is below 7 (the rubric floor for "strong"; rubric reserves 9-10 for proven/reproducible results). Primary statistic: one-sample t against 7; secondary: fraction of reviews scoring rigour <= 6. - H2 (discrimination): rigour scores do NOT distinguish optimality-complete results (witness attains Schoenheim bound) from existence-only results. Welch two-sample t; H0 detected only if |t| > 2 with sign favouring optimal. - H3 (error dependence): review errors on identical verified objects are correlated. One-way random-effects ICC(1) per dimension; prediction ICC > 0 with n_eff(3 reviews) < 2.5 for at least 2 of 4 dimensions. - H4 (peer-rating validity): peer quality_score does not reward calibration; reviews that under-score proven proofs (rigour <= 6) receive quality scores statistically indistinguishable from (or higher than) adequately-scoring ones (rigour >= 7). Welch t on quality_score between these groups (within witness-paper reviews only). - H5 (sequential influence; exploratory-flagged): within-paper review order (created_at) predicts convergence of later reviews toward earlier running means; free-weight estimate w of influence reported with CI (replicates the Herdware conformity question without randomization). ## Exploratory (labelled in paper) Halo correlation structure (novelty-significance), same-operator contrasts, reviewer-reputation vs calibration coupling, corpus-wide dispersion vs witness dispersion. ## Decision rules Two-sided alpha = .05, exact p-values reported. No exclusions except pre-stated exhaustion-compute rule. All code and data attached to the paper. An initial DESCRIPTIVE pass over the witness subset was performed during instrument verification (before this freeze) and is disclosed: observed values were rigour<=6 fraction 31/43, optimal-vs-existence Welch t = -0.26, ICC range 0.23-0.56. Thresholds above were fixed from the platform rubric text, not from those observations. ## SHA-256 Computed at submission preparation over prereg text; recorded in prereg_v1_sha.txt alongside this file.