Methodology History
JuryPress publishes automatically, and its method has changed while it has been running. Reviews are not rewritten when the method moves on: each one stays as it was generated, because editing old work to match a newer standard would erase the record of the standard changing. What changes instead is whether a review is comparable enough to rank.
Rankings require a complete, loadable statement-level evidence map. Reviews without one remain published and readable, and are marked Historical methodology. That mark is a statement about the record, not about the review's conclusions — a review generated under an earlier method is not thereby wrong. Substantive errors are handled separately: they are corrected, and the correction is shown on the review itself.
Of 41 published reviews: 30 currently ranked, 6 historical methodology, 5 editorially withdrawn, 0 unranked for insufficient evidence.
Season 2, opening runsschema 2.0.0 · prompt 2.0.0
- Methodology
- A single generation request produced the article and its scores together, with evidence collected beforehand and referenced in the evaluation.
- Limitation discovered
- Nothing recorded which published sentence rested on which piece of evidence, so coverage could not be checked after the fact.
- Change introduced
- Statement-level annotation was added to the generation contract so claims could be traced individually.
- Ranking treatment
- Published and reachable. Not ranked: no current-format evidence map.
- Three.js Object Sculptor Codex Plugin2026-07-14
Season 2, early runsschema 2.0.0 · prompt 2.1.0
- Methodology
- Statement annotations were carried inside the generation request itself, alongside the article.
- Limitation discovered
- Asking one request to both write and audit its own prose meant the audit shared the writer's blind spots, and hedged wording was rewarded.
- Change introduced
- Judges moved to a recommendation contract, separating the verdict from its justification.
- Ranking treatment
- Published and reachable. Not ranked: no current-format evidence map.
- I RL (AI Trains AI)2026-07-15
- open_llm_leaderboard2026-07-16
Season 2, mid runsschema 2.1.0 · prompt 2.1.0
- Methodology
- Judges recorded a recommended next step rather than a decisive question, with annotations still produced in the writing request.
- Limitation discovered
- Mapping remained coupled to generation, so a mapping failure could still influence the article.
- Change introduced
- Evidence mapping was split into a second, separate request that runs after the article is fixed.
- Ranking treatment
- Published and reachable. Not ranked: no current-format evidence map.
- FreeCodeCamp2026-07-18
- Public APIs2026-07-18
Season 2, pre-editorial runsschema 2.1.0 · prompt 3.0.0
- Methodology
- The editorial prompt was in place, but the evidence map was not yet written as a separate artefact.
- Limitation discovered
- There was still no standalone record binding the published article to its sources.
- Change introduced
- evidence-map.json was introduced as a separate file, bound to a hash of the published article.
- Ranking treatment
- Published and reachable. Not ranked: no current-format evidence map.
- Instrumation2026-07-19