Archived Weekly Accuracy Chart
Archived Weekly Forecast Accuracy
07 Jun
14 Jun
21 Jun
28 Jun
05 Jul
12 Jul
19 Jul
26 Jul
short week
02 Aug
short week
09 Aug
16 Aug
This rebuild: eight predictions finalised, from the editions of 8 and 9 August. The count moves from 153 to 161, the sum from 727.8 to 782.8, and the cumulative average from 4.76 to 4.86 — the largest single-rebuild move this chart has recorded. The hit rate goes from 44 to 47 per cent.
All eight were hits, and that is the problem with them. Every one took the same form: a proposition that something would change, priced below even, which then did not happen. The backlog did not clear and transits stayed in single digits. The blockade was not lifted. Tehran named no regional security vehicle and dropped none of its six conditions. Washington did not concede an administrative role at the strait — it moved the other way, claiming the waterway as territory. The reserve did not stop falling; it fell 6.1 million barrels to 298.7 million, below 300 million for the first time since 1983. Santo Domingo and Caracas produced no named step. Warsaw's appeal produced none either.
A run of eight consecutive hits should raise suspicion rather than satisfaction, and here is why. This desk's panels are built mostly from change-propositions priced low. Most of them resolve no, because on most days most things do not change. A forecaster who prices four unlikely events at twenty per cent every morning will accumulate hits indefinitely without demonstrating much. The rubric scores a correctly-priced non-event exactly as it scores a correctly-priced event, and the panel structure supplies far more of the former. Part of the jump from 4.76 to 4.86 is skill; part of it is the shape of the questions being asked. We flag it here rather than bank it.
The two lead calls both failed, and they failed the same way. The 8 August edition argued the fleet was waiting rather than frightened, and named its own test: transits should jump on an announcement. The 9 August edition argued Washington was preparing to concede the administration of the strait. Both were the desk's most confident reading of the week at 38 per cent, and both described movement toward a settlement that did not arrive. Set beside the misses graded yesterday, where we under-priced Washington restating a claim it had already made, a single pattern emerges: this desk over-predicts change and under-prices continuity. Eight low calls landing does not offset two lead calls missing, because the lead calls are where the analysis actually lives.
Two of the eight rest on non-observation. The Iranian security-framework call and the Iranian conditions call were graded no because we could not find evidence of the event, not because we found evidence of its absence. That is the same weakness recorded against a call graded yesterday, and it is now frequent enough to name as a category rather than a footnote.
One call was untestable as drafted. The backlog test was designed to distinguish a waiting fleet from a frightened one by watching transits after an announcement. No announcement came, so the discriminating event never occurred. The call resolves no on its literal wording and is graded accordingly, but the hypothesis underneath it remains untested and should not be treated as refuted.
W10 (3–9 August) can be drawn for the first time, at nineteen of roughly twenty-seven graded, with a whisker of 4.2 to 7.2. The graded nineteen average 6.00. The upward-bias warning printed at the last two rebuilds now has a full arc behind it: the first seven graded showed 5.71, eleven showed 5.36, and nineteen show 6.00. The bias was real but not monotonic, and we record the wobble rather than the tidy version.
The correction printed here this morning has itself been withdrawn. Earlier today this chart moved W8's label from six panels of seven to five, on the reasoning that a spread of 2.8 points across a twenty-eight-call week must mean eight ungraded calls, or two panels. That reasoning does not hold. The whisker's endpoints are printed to one decimal, so the true spread is 2.8 give or take a tenth, and at least three whole-number answers fit inside that tolerance: seven ungraded calls in a population of twenty-five, eight in twenty-eight, six in twenty-two. The method picks one of them only by assuming the population is twenty-eight. It cannot distinguish between them, and we should not have written as though it could.
The assumption behind it has now been falsified next door. W9's population is not twenty-eight either, for reasons set out below. Since W8 additionally carries live calls that sit on a separate ledger and never entered this population, its denominator is not established, and a ratio computed from an unestablished denominator cannot identify how many panels are missing. W8's label therefore states what is actually known — seven editions, five over-long calls moved out, panels graded not yet established — and the bar stays hatched until its panels are read one by one. Withdrawing a correction the same day it was published is embarrassing; publishing a confident figure derived from a method that cannot produce one would be worse.
That is a units inconsistency, not an arithmetic error, and it runs the length of this chart. W1 to W7 are labelled in predictions; W8 onward in panels. Nothing in the bars is affected — every bar is a mean of graded predictions either way — but a reader comparing the label under W4 with the label under W8 is comparing two different things, and until today so were we.
The archive confirms the week's shape and contradicts our first reading of it. W8 runs 20 to 26 July and contains seven daily editions, the first of which we had briefly mistaken for absent. The seven-panel, twenty-eight-call structure that made the whisker arithmetic come out exact is the real structure, and the audit above stands on it.
Five calls in W8 were written past the seven-day discipline, and they are moving. The 20 July edition set a China-facing blockade call at roughly two weeks. The 25 July edition set one on the patrons' split running to about 8 August, also fourteen days. The 26 July edition set three: an escalation call at roughly ten days, a diplomatic-track call at roughly two weeks, and an off-region call on a Senate floor vote at ten days — only its lead call, on the pause holding through the Tuesday summit, closed inside a week. All five move to the long-term ledger, on the same principle applied to the 15 August panel: graded exactly as published, on the page that suits their length.
The seven-day rule was already in force, which makes these breaches rather than legacy. We had wondered whether W8 predated the discipline. It did not. The editions of 14 to 17 July close at exactly seven days apiece — 14 July closing on the 21st, 15 July on the 22nd, 16 July on the 23rd, 17 July on the 24th — a clean seven-day cadence running straight into W8. The rule existed and was being kept; a minority of calls were written outside it anyway.
And the breaches are not confined to the hatched weeks. The 15 July edition — inside W7, a bar published solid and by our own rule unmovable — carried a fourth call resolving at the 29 July Federal Reserve decision, fourteen days out. So at least one closed week was completed with a call that ran to double the window, and we do not presently know how many others were. The solid bars are not being redrawn; that rule stands. But their labels currently imply a uniformity the archive does not support, and the honest position is that this problem is chart-wide rather than a late-July lapse.
The tell we named this morning has been falsified by the end of the day. We wrote that the marker of an over-long call is elastic phrasing rather than a date — that calls written “within roughly two weeks” drift, and calls written with a close date land on seven. The first half survives. The second does not. The 27 July panel's off-region call named an explicit date, 6 August, and still ran ten days from writing. A date disciplines the grading, because everyone knows when to look; it does not discipline the horizon, because a date can be any distance away. The fix is therefore the ceiling itself, not the drafting habit that made the ceiling easy to breach without noticing.
W9 is a six-day week, and nobody had noticed. This is the second pass's main finding and it settles the week's shape without needing a single grade. The run number and the day number advance independently: the day number follows the calendar, the run number counts numbered daily editions. Runs #79 to #86, spanning 19 to 26 July, sit at a constant offset of sixty-two from the day count. Run #90, on 31 July, sits at sixty-three. One calendar day in between produced no numbered daily edition, and the archive names it: 28 July carried only a special edition, filed before the Trump–Netanyahu meeting it was written to anticipate. The prev-bars of the editions on either side corroborate it, listing 28 July in the daily sequence explicitly marked as a special.
So W9 contains six daily panels and twenty-four calls, not seven and twenty-eight. Its label said six of seven panels graded. The denominator was wrong; the numerator may well have been right all along. If those six panels are in fact graded, W9 is complete, its whisker is an artifact of the missing seventh panel that never existed, and the week can close. Carrying the arithmetic across gives a closing mean near 5.4 — but that figure is derived from a low endpoint that was itself computed over twenty-eight, not summed from the grades themselves.
Two errors, one shape. Both of today's withdrawn findings came from inferring a population instead of counting one, and both would have been caught by asking how many editions a week actually contains before dividing by twenty-eight. That check is now a standing step at every rebuild: establish the panel count from the run numbers and the archive first, then compute.
W11 (10–16 August) is a six-day week by construction, and the reason is a drafting fault of ours. The four scenarios published on 15 August were written to close on 30 September — forty-six days out, against a house discipline of seven. Left in the daily panel they would have held this week's bar hostage until October and made the chart useless to a reader for six weeks. They have therefore been moved to the long-term ledger, where the 28 July base call and the specials calls already sit, and where a six-week horizon belongs. They are not reweighted, not withdrawn and not rescued: they will be graded exactly as published, on 30 September, on the ledger that suits their length. What changes is only which page they are counted on.
So this week's bar will rest on six days of daily panels rather than seven — the editions of 10, 11, 12, 13, 14 and 16 August, twenty calls in total. The 15 August edition contributes none. We print this rather than quietly drawing a seven-day bar from six days of evidence, because a week short by one panel is a week whose mean rests on less than it appears to.
The over-length problem is not confined to W8, and that is why W9 does not close either. The 27 July edition is W9's first panel. Three of its four calls run past seven days: the escalation branch at roughly ten, the diplomatic-track call at roughly two weeks, and the off-region Senate call to 6 August, ten days out. Only the lead call, on the pause surviving the Tuesday summit, closed inside the window — the following day, in fact. So the week whose ratio looked internally consistent, and which we declined to call audited on the strength of that consistency, turns out to carry at least three calls that belong on the long-term ledger before it can be summed at all. Declining to record it as audited was the right call for the wrong reason.
Grading is driven by the window, not by a fixed lag. A call is graded in the edition after its own window closes, and until then it is carried forward as open with its evidence updated. That is why a week's grades are scattered across several later editions rather than landing together, and it is the reason a week containing calls of four different lengths cannot be closed on any single day. It also means the number of editions between a call and its grade carries no information, which we had been half-assuming it did.
Two figures are now anchored to the published record rather than to this chart's own previous drawing. The 27 July edition states the ledger at 4.63 across 131 finalised predictions, 56 of them hits. The July review, published 1 August, independently states three of this chart's solid bars at 5.19, 4.87 and 5.61 for the weeks of 29 June, 6 July and 13 July, matching what is drawn here to the decimal, and confirms that the last two July weeks were still inside their scoring windows at that date. The solid bars check out against a source that is not this page.
One W8 grade is recovered and printed here. The 23 July lead call — that retaliation would reach Iranian bridges or power plants under the announced rule — scored 3 out of 10. It was right about the mechanism and wrong about the targets: the retaliation waves arrived on schedule and stayed on military sites.
The ledger is now anchored at three published points, not one. The editions of 27 July, 1 August and 4 August state the running record at 4.63 across 131 finalised predictions with 56 hits, 4.70 across 136 with 60, and 4.72 across 138 with 62 — each described as recomputed from the ledger and never estimated. Seven predictions finalised between the first anchor and the last, carrying a combined 44.8 points, an average close to 6.4. Two of the individual grades behind that are on the record: the 25 July call that Iran would hold the line scored 8, and the 2 and 3 August pause-breaking calls scored 6 apiece.
Those two sixes carry a warning this chart should repeat rather than bury. The letter graded them down itself, on the ground that vessels had already been struck before either call was written — so a proposition offered at around a quarter's confidence was close to true at the moment of writing. The bar had been set where routine events already cleared it. That is a fault in specification rather than a success in forecasting, and it belongs beside the non-observation category and the change-proposition warning as a third way this scoring system can flatter itself.
Two of the rules below were already in force, and this chart did not know it. The 4 August edition states plainly that from that day every call carries a hard calendar date and must not restate a question already open. The hard-date rule is therefore twelve days old, not new today, and the chart has spent this rebuild proposing a reform the letter had already adopted. The same edition shows the bounded-cycle and reserved-card branches already carried as a single standing call graded weekly, which places the consolidation at the start of August rather than on 9 August — that date was when the standing call was next due to be graded, not when the practice began. Both corrections point the same way: this page has been reasoning about the letter from its own previous drawings instead of from the letter.
W9's over-length problem is not three calls. It is most of the week. The 27 July panel carries three calls past seven days. The 29 July set runs to about 12 August. The 31 July panel's Gaza clause closes 14 August, fourteen days out. The 1 August panel sets three calls closing about 12 August, eleven days out. On the seven-day ceiling adopted below, the great majority of W9's twenty-four calls belong on the long-term ledger, which would leave a bar drawn from a handful of survivors and labelled as a week. Those calls have now been moved, and what remains is drawn below as a short week in its own colour.
Why the two teal bars are not comparable to the rest of this chart
W8 and W9 are drawn in a different colour because they are not the same kind of measurement as the bars beside them. Each was written as a seven-day and a six-day week — twenty-eight and twenty-four calls respectively. What survived to be graded on the daily ledger, once the calls written past a seven-day horizon were moved to the long-term ledger and the live calls were held on their own ledger, is three calls and two calls. The bars are honest about those three and two. They are not honest about the weeks, because three calls cannot describe a week and neither can two, and no amount of correct arithmetic fixes that.
W8 closes at 6.33 across three calls, two of them hits. The 23 July call that American retaliation would reach an Iranian bridge or power plant scored 3 — the retaliation arrived on schedule and deliberately left that target tier alone, so the mechanism read correctly and the target class did not. The 25 July call that Iran would hold the line with no ceasefire on the current terms scored 8, confirmed when the Omani proposal was rejected over who levies the fees. The off-region call from the same edition, that the Russia sanctions bill would reach a Senate floor vote, scored 8 when the combined Russia-and-Iran bill advanced 86 to 12 ahead of its window.
W9 closes at 7.00 across two calls, both hits. The 27 July call that the pause would survive the Netanyahu summit scored 7 — graded down for margin, since the meeting passed quiet exactly as called but the missile volley followed the same evening. The 29 July off-region call that the Federal Reserve would hold at 3.50 to 3.75 per cent and tilt hawkish on September scored 7, on a nine-to-three hold with all three dissenters demanding an immediate rise.
Every one of those five grades reconciles against the published ledger. The daily record moved from 4.63 across 131 finalised predictions on 27 July to 4.72 across 138 on 4 August. Seven entries entered in that span: the five above, plus the 2 August and 3 August pause-breaking calls at 6 apiece, which belong to W9's successor weeks. Those seven sum to 45 points against the 44.8 the two endpoints imply, and carry six hits against the six the endpoints imply. The itemisation and the arithmetic agree, which is the standard this page failed three times today before meeting it.
Two calls were moved out rather than graded, and they are named here. W8's 25 July call on the patrons' split resolving toward Beijing was written to a fourteen-day horizon, and W9's 31 July Gaza clause to a fourteen-day one. Both now sit on the long-term ledger. They are graded exactly as published, on the ledger that suits their length — never reweighted, never withdrawn. The same applies to the 27 July escalation branch, the 29 July bounded-war set, the 1 August panel's three long branches and the 2 August set, all written past the ceiling.
The real finding is what happened to the other forty-seven calls. W8 and W9 were written as fifty-two calls between them and produced five gradeable daily entries. A few went to the long-term ledger and a few to the live ledger, but most simply vanished into the next day's panel — restated or superseded rather than resolved, so they never became ledger entries at all. Across the eight editions from 27 July to 4 August the ledger grew by seven, against a nominal four calls a day. That is the mechanism behind the eight-item gap between the daily letter's population and this chart's, and it is why the twenty-eight-calls-a-week assumption that produced this morning's two withdrawn corrections was never true of any week on this chart.
So read the teal bars as what they are: the graded residue of two badly-drafted weeks, printed with their counts stated rather than suppressed. A week that produces two gradeable calls has not been measured. Publishing 7.00 beside 5.61 without this note would flatter the record, because two calls chosen by the accident of which windows happened to close are not a sample of anything. The alternative was to leave both weeks hatched indefinitely, and an honest small number with its denominator printed beside it is better than a permanent blank.
Five rules follow from this audit, and they take effect today.
One: a seven-day ceiling on daily-panel calls. Anything written with a longer horizon goes to the long-term ledger at the time of writing, alongside the 28 July base call and the specials calls. A daily bar should be readable within a week of the week ending.
Two: every daily call carries an explicit close date. Elastic phrasing is banned. This one is recorded rather than introduced — the letter adopted it on 4 August and this page is twelve days late in noticing. The recurring tell in every over-long call found by this audit is a phrase rather than a date — "within roughly two weeks", "within roughly ten days". Calls written with a date land on seven days; calls written with a phrase drift to ten or fourteen.
Three: every chart label declares its unit. W1 to W7 count individual predictions; W8 onward count daily panels of four. Nobody announced the change and for several rebuilds nobody noticed it, which is how a reader ends up comparing two different things placed side by side.
Four: a call still unresolved after three editions is graded unresolvable and struck from the ledger with a printed note. Holding a question open indefinitely because it has not obliged us with an answer is worse for the record than an honest void.
Five: retrospective moves are permitted for ledger placement only. Never for weight, wording or withdrawal. A printed call is never reweighted and never quietly disappears; the only thing that may change is which page it is counted on.
And one figure carried over unexamined. W10 is labelled nineteen of roughly twenty-seven graded. The run numbers show no skipped day between 3 and 9 August, so that week has seven panels and twenty-eight calls, and the approximate denominator is a soft number standing in for an exclusion set nobody has listed. It is the same defect as the eight-item gap described below, and it is named here rather than left to be discovered at the next rebuild.
Once a bar is published solid it never moves. Hatched and open bars are recomputed at each rebuild until their week closes. Next scheduled rebuild: 23 August 2026.
How to read this. Each bar is the mean finalised score of all daily predictions whose forecast window opened in that week. A prediction counts once, keyed to the edition that made it, and only after its window formally closed and the score entered the ledger. Early editions (June) finalised next-day; later editions carried predictions provisionally for up to a week before finalisation, so a call made in one week is grouped here in the week it was made, not the week it was scored. Six or better counts as a hit; a score under five means the call read the world worse than a coin toss would have, which is why the June figures are shown in full.
Why this differs from the figure printed in the daily letter. The daily editions report 4.89 across 169 finalised calls after today's grades; this chart reports 4.86 across 161. The two count different populations by design — the chart excludes provisional, faulted-premise, special and live-call items, which sit on separate ledgers. The gap has held steady at exactly eight items across three consecutive rebuilds, which is what a stable exclusion set should look like. It is nonetheless eight items whose exclusion is asserted rather than itemised here, and a reconciliation listing them individually is owed.
Comments
Post a Comment
Comments will be displayed after moderation