Scoring Record
ParleyBot Intelligence · TruthSetsFree · Scoring Record
Scoring Record
Every daily edition publishes four predictions, each with a weight, a named falsifier and a hard close date no more than seven days out — three rather than four on a day a disclosed K5 shortfall applies. Each is graded 0–10 once its window closes, and the running record is recomputed and printed in the edition that grades it. This page reports that ledger. It does not keep a second one.
Lifetime
5.23
across 315 finalised predictions · 170 hits · 54.0 per cent
Form since the 16 August rebuild
5.62
146 predictions · 91 hits · 62.3 per cent · the twenty-second reading above five, on a tolerance of 5.61 to 5.64
As printed in the edition of 20 September 2026. A score of 6 or better counts as a hit. A mean below 5 means the calls read the world worse than a coin toss would have.
Four defects in the edition of 20 September
None of them is a factual error, and all four are visible on the page. The edition went through five drafts after a late correction round, and each defect is residue from that editing rather than from research.
The sourcing note carries the label Read in full and load-bearing twice, because an addition was inserted as a second list instead of merged into the existing one. A paragraph in the second section opens The second is that they now share a supply constraint, with no first: the sentence that enumerated it was replaced when two sections were merged and the enumerator was left dangling. The weaker-standard list still names a downed-aircraft claim and the Chinese foreign minister's remarks, both of which were cut from the body before publication, so the edition describes sourcing for material it does not contain. And the cut list states that any displacement figure with a stated method was cut for want of an anchor, in an edition whose lead is built on displacement figures with a stated method; that line survived from the draft written before the figures were found.
All four are corrected in the edition of 21 September rather than silently amended on the published page.
Form, by grading period
Bars run 0 to 10; the grey line marks 5. Hatched bars rest on fewer than ten predictions and should not be read as form. Blocks are grouped by the date a prediction was graded, closing every twenty. The block that closed 17 September stood at 5.85 across its full twenty; the block that opened on 18 September now carries twelve, enough to stop hatching it.
The most recent board, graded 20 September
Hit6Washington does not confirm the schedule. Continuity, priced 24.
Miss3A route is announced. Change, priced 28.
Hit6Outside the region: the authorisation is not renewed in public. Continuity, priced 24.
Hit6The Rome round gets no date. Continuity, priced 24.
Board of 20 September, mean 5.25 across four, three hits and one miss, all belonging to the panel of 13 September. The miss is the route call, and it was lost within hours of publication: the Salalah meeting it was built around was postponed late on 13 September and still has no date. It is graded three rather than lower because the exclusions were drawn tightly and the edition printed the defeating mechanism in its own against-case; no higher because a standalone of 72 was priced on the wrong side of an event already unravelling as it was written.
All three hits rest on nothing having been announced, and all three are capped at the floor. No American body confirmed the two-slot escort schedule; no American body announced a renewal of the cutter authorisation; no date was set for the Rome round. That shape is recorded rather than banked: on a week whose central argument is that nothing in this war gets agreed, a method that pays for predicting non-announcement returns sixes whether the argument is right or not. The caps are doing exactly the work they exist for, and a board of this composition should be discounted rather than read as confirmation.
The forty-prediction window, second reading not yet due
The window now running opened with the board graded 18 September and stands at twelve of forty. Change propositions run at 6.00 across five, continuity at 6.14 across seven. Twelve predictions is not a reading and none is claimed.
The direction has now reversed twice inside that dozen. The first reading, closed 17 September, had change ahead by 0.30. The board graded 19 September ran the other way, continuity 6.50 against change 5.50. The board graded 20 September widened it further, continuity 6.00 against a single change proposition at 3.00. A single change call per board is far too thin to carry a conclusion, which is precisely why the window runs to forty, and the standing caution still applies: change propositions are published only where this desk already believes them, so they are selected for changes visibly under way, while continuity carries no equivalent filter.
Open predictions and when they close
| Edition | Calls | Closes |
|---|---|---|
| 14 September | 4 | Monday 21 September |
| 15 September | 4 | Tuesday 22 September |
| 16 September | 4 | Wednesday 23 September |
| 17 September | 3 | Thursday 24 September |
| 18 September | 3 | Tue 22, Fri 25 & Sat 26 September* |
| 19 September | 4 | Thu 24, Fri 25 & Sat 26 September† |
| 20 September | 4 | Sat 26 & Sun 27 September‡ |
26 predictions open. *The three calls of 18 September carry three different close dates, not two as this page previously stated: the Gulf-meeting call closes Tuesday 22 September, tied to a fixed calendar event inside the seven-day cap; the EU-tariffs call closes Friday 25 September; and the call that the monitoring panel stays unreconstituted closes Saturday 26 September, set to the earlier of the two disputed mandate-expiry dates so it cannot grade a hit on a mandate that has already lapsed. †The four calls of 19 September likewise split: the pipeline call closes Thursday 24 September, the state-visit and Korea calls Friday 25 September, and the delegation call Saturday 26 September. ‡The four calls of 20 September split two ways: the American–Iranian meeting call closes Saturday 26 September, tied to the end of the high-level session, and the Red Sea, cutter-authorisation and ten-year calls close Sunday 27 September on the full seven days. The panel of 13 September has closed and been graded; the panel of 20 September takes its place. Long-horizon calls from special editions are held on a separate ledger and graded on their own dates; they are excluded from the figures above. One prediction from 18 August remains carried rather than graded, because the record that would settle it is contested and this desk does not grade a contested record on the convenient reading of it.
Five faults on the record, and now a sixth: the revision artefact
The dominant fault on this record through the first week of September was drafting. From 12 to 16 September the live fault was mispricing: propositions correctly written and simply priced wrong. The 17 September edition added a third and narrower kind: a statement about the record itself, rather than about the world, that turned out to be wrong. The 18 September edition added a fourth, one level further removed — a fault in the checking mechanism rather than in a proposition or in a statement about the record. The 19 September edition added a fifth, the inherited fact, set out below. All five are set out in full elsewhere on this page and none is retired.
The fifth, the inherited fact. The edition of 18 September placed a Korean naval unit near the Strait of Hormuz. It is in the Gulf of Aden, roughly fifteen hundred miles away. The 19 September edition repeated the error, because it took the fact from the prior edition rather than from the source, and a call was drafted against the wrong body of water before a direct read of the wire copy caught it. Nothing reached a reader in the second edition, but the first one published it and it stands corrected there now.
The sixth, added by the edition of 20 September, is the revision artefact, and it is the first fault on this list that has nothing to do with being wrong about the world. That edition was correct on every fact it printed and defective in four places anyway: a duplicated sourcing label, an enumerator with nothing to enumerate, and two sourcing lines describing a draft that no longer existed by the time it published. Every one was introduced by editing, in a revision round that was itself prompted by a correction.
Related, and recorded with it because the mechanism is the same: during that round a computed figure was written into three successive drafts as an unfilled placeholder. The script that should have substituted it matched only digits, so it found nothing to replace, printed the correct figure to the operator and wrote none of it to the file. The figure was reported as done three times on the strength of that printed output. It was the editor's check, not the desk's, that found the placeholder still sitting in the copy.
The general form: a fact is checked against the world, but an edit is only ever checked against the intention behind it. Research faults are caught by looking outward at a source. Revision faults can only be caught by reading the finished artefact as a stranger would, and that is precisely the reading a desk that has been through five drafts is least able to perform. The operational rules taken from it are two, and both are narrow. A verification step must inspect the file it claims to have changed rather than report what it computed. And the last pass before publication is a read of the whole edition for internal coherence — labels, enumerators, and whether the sourcing note still describes the edition that exists — with no reference to the sources at all.
The general form is worth stating because it is invisible by construction: a fact that has already appeared in this letter is the single least likely thing in any edition to be checked again. Every other claim arrives from outside and gets tested on the way in. An inherited one arrives carrying this desk's own imprimatur, and the second printing looks like corroboration of the first when it is only a copy of it. The rule taken from it is narrow and operational: a fact carried forward from a prior edition is re-verified against its original source before it is used to write a falsifier, because a call written on an inherited error can fire on a technicality that has nothing to do with the world.
The retired chart
The weekly bar chart this page replaced is preserved, unrevised, at the archived Weekly Accuracy Chart. The change was which page the record is kept on, not what the record says.
This page is regenerated from the most recent daily edition and, where more than one edition has passed since the last regeneration, from the anchors each of those editions itself published in between. It carries no population of its own, so it cannot drift out of step with the letter for long — this update closes a one-day gap, the cadence Part 11 has specified all along.
Seven known ways this scoring system can mislead, restated here so a reader can discount for them. First, most panels are built from change-propositions priced low, and most days most things do not change. Second, some hits rest on non-observation and are capped at the floor of a hit, lifted only on positive in-window evidence from the party who would have caused the thing. Third, a threshold occasionally gets set where routine events already clear it. Fourth, a proposition written in the direction the desk disbelieves, at a low standalone likelihood, can only lose and teaches nothing; propositions are now written in the direction the desk believes. Fifth, acting on a correction can switch off the measurement that produced it, which is why the per-panel change requirement was retired in favour of the rolling forty-prediction window. Sixth, an outcome can be decided by how a prediction was written rather than by what happened — five forms of this and the rules written against each are on record. Seventh, a proposition can be correctly drafted and simply mispriced, which is the live fault first identified in mid-September. None of the seven covers a fault in what this page says about its own arithmetic, a fault in the checking tool rather than the content, a fact inherited unchecked from a prior edition, or a defect introduced by revising the edition itself — all four are recorded separately above rather than folded into one of the seven.
The approach, the six coverage domains and our scoring record — graded daily and reviewed each month — are set out on the About page.
No financial advice is expressed or implied.
Last updated 20 September 2026 from the edition of 20 September 2026 · lifetime 5.23 across 315 finalised predictions, 170 hits · form since the rebuild 5.62 across 146 · 26 predictions open, closing between 21 and 27 September 2026 · one prediction from 18 August carried unresolved · special-edition and live-call items are held on separate ledgers and excluded.
Comments
Post a Comment
Comments will be displayed after moderation