*This appendix makes the weighing of Chapter 16 recomputable. It lays the model bare, gives the four premises it computes through, and shows the outcomes. The weighing can be reproduced with a small Python script. Read this first: this model is not proof and delivers no composite p-value. It is an instrument of transparency — it shows how the ranking among AD 30, 31 and 33 moves with the assumptions one brings in. Whoever presents it as proof is building scaffolding.*
F.0 · What the model does and does not do
It does not combine independent probabilities into a single number — that would double-count the shared assumptions (calendar, textual reading). It does something more modest: it gives each year a merit score (1–3) per criterion, weights these, and sums them to an index of 0–100. The gain lies not in the final figure but in the sensitivity: shift the assumptions, and watch whether the ranking flips. Where it flips lies the hinge of the whole case.
F.1 · The eleven criteria and the score matrix
Eleven criteria in five families. Each year receives a merit score per criterion on the scale 0–3 (in practice the assigned scores run from 1 to 3). The score matrix below is the heart of the model. It gives the raw merit scores; the calendar symmetry (C1/C2 flex) counts as a weighting transformation in all four weightings, and the idiom reading (academic/sceptical) additionally overrides C9 — see F.2. [CALCULATED]
family
#
criterion
what it weighs
30
31
33
ground
Calendar/astr.
C1
calendar cleanness of the weekday
does the day of death fall on a halakhically clean erev Pesach (14 Nisan)?
3
2
2
30 robust Friday; 31 Adar II (system-normal, forced under rule A); 33 pre-equinox exception, P&D gives 2 May — 33 no longer "clean" (ch. 15)
Calendar/astr.
C2
astronomical robustness
how far does the new moon sit from the critical threshold (margin in days)?
does the year fit baptism-in-27/29 and the number of Passovers of the ministry?
2
3
2
31 baptism-27 + 4 Passovers natural, both halves of that ground weakened: the four-Passover reading is carried as support and not as compulsion (appendix E.3), the autumn-inclusive count behind the baptism of 27 as unattested and not refuted (ch. 9/ch. 14) — whether the cell should therefore drop to 2 is weighed in F.3, and why it has not been done stands there too; 30 baptism-27 + exactly-3 or co-reg-26; 33 baptism-29 standard (ch. 9/ch. 14)
Chron./real.
C4
46 years (John 2:20)
do the 46 years since the temple building rhyme with the first Passover of the ministry?
reads clean on Friday; Wednesday = "better buy", pays a price (ch. 13)
Reception
C11
majority/tradition
does the prevailing, ecclesiastical reading stand behind this year?
2
1
3
33 majority, 31 minority (ch. 6/ch. 14)
Within "Text" one criterion carries the most weight and the entire hinge: criterion 9 — the sign of Jonah. As a measure it scores AD 31 = 3, AD 30 = 1, AD 33 = 1 (the double construction points hard towards the Wednesday). As pliable idiom it becomes neutral (all 2), and the sharpest lead of AD 31 disappears. Two other criteria pull the other way: astronomy (C2: AD 30 = 3, AD 31 = 3, AD 33 = 1 — AD 33 fragile) and reception (C11: AD 33 = 3, AD 30 = 2, AD 31 = 1 — the consensus is comfortable with AD 33). These three contrasts explain almost the entire outcome.
The weights — from criterion to family. The eleven criterion weights sum to 100 and add up within each family to the family weight. Thus every family entry is traceable to its criteria. [CALCULATED]
family
criteria (weight)
family weight
Calendar / astronomy
C1 (12) + C2 (8)
20
Chronology / realism
C3 (9) + C4 (8) + C5 (7)
24
Exact fulfilment / convergence
C6 (8) + C7 (6) + C8 (6) — premise-dependent
20
Text
C9 (14) + C10 (8)
22
Reception
C11 (14)
14
total
100
F.2 · The computation rule and the four weightings
The computation rule. Each year receives one index of 0–100 per weighting. With weights w and cell scores s (per criterion c):
index(year) = 100 · ( Σ w·s ) / ( 3 · Σ w )
— where the sums run over the eleven criteria — rounded to one decimal. The factor 3 in the denominator is the maximum cell score, so that a year scoring the top mark on every criterion would score exactly 100; criteria with weight 0 drop out. This is exactly what `voorwerk/weging_herrekening_juli2026.py` (Python 3, no external packages) does — the four weightings below and their rankings roll straight out of it. [CALCULATED]
The four weightings. Each weighting is a transformation of the base weights (column "Σ" = the sum of the weights in that weighting), sometimes with overwritten score cells. This makes every outcome line recomputable from the matrix of F.1:
The symmetry correction C1/C2 "flex" — the calendar freedom of F.3, which sets C1 to 3/3/3 and C2 to 3/3/2 — counts in all four weightings, for it is a calendar judgement, not a judgement about the sign. The idiom reading (academic/sceptical) additionally overrides C9 to neutral (the sign is there no longer a measure). With these two tables and the computation rule the four outcomes are fixed:
30 > 33 > 31 (photo finish between 30 and 33, 0.7)
Sceptical
sign = idiom; convergence out; calendar freedom for all
80.1
75.5
76.6
30 > 33 > 31 (photo finish between 33 and 31, 1.1)
Worked check (default, AD 31), by way of illustration: with the calendar flex C1 for AD 31 is a 3 (not 2), so numerator = 12·3 + 8·3 + 9·3 + 8·3 + 7·3 + 8·3 + 6·3 + 6·3 + 14·3 + 8·2 + 14·1 = 264; denominator = 3·100 = 300; index = 100·264/300 = 88.0. The remaining cells compute the same way and reproduce the table to the decimal.
Read the table as a whole. Under the premise the book defends — the sign is a measure, God delivers to the day — AD 31 wins decisively: 93.3 against 54.8 for AD 33, and 88.0 against 61.7 in the balanced variant. That is not a neck-and-neck race; that is a win by a distance. Only when one rejects the core premise and reads the sign as a figure of speech does the picture flip to a cluster in which AD 30 leads and AD 31 and AD 33 lie within a few points of each other. The whole case thus turns on a single hinge — is the sign a measure? — and the model makes that visible instead of hiding it.
F.3 · The symmetry correction
An earlier version of this model was unfair on one point: it granted AD 33 the benefit of calendar freedom (Stern's "quite often a late Pesach") but not AD 31. That artificially depressed AD 31. The principle of the correction: whoever grants calendar freedom to one year grants it to both — both AD 31 (the Adar II intercalation) and AD 33 (the pre-equinox March moon) receive the normal calendar judgement (C1 → 3/3/3, C2 → 3/3/2). Neither year is depressed or inflated.
The first implementation (v3), however, applied that principle only in the academic and sceptical weightings — the two in which AD 31 loses anyway — and left it out of the exact-fulfilment and the default, the two in which AD 31 wins. There AD 33 was thus still depressed on calendar grounds that the principle itself forbids. That is a double standard, and it has been removed here: the calendar flex now counts in all four weightings. It is a judgement about the calendar, not about Matthew 12:40, and so must not hang on the sign premise.
By way of check, the four outcomes before and after this universal application:
weighting
before (v3) — AD 30 / 31 / 33
after (v4) — AD 30 / 31 / 33
Exact fulfilment
70.9 / 89.7 / 48.8
70.9 / 93.3 / 54.8
Default
73.7 / 84.0 / 55.0
73.7 / 88.0 / 61.7
Academic
79.1 / 72.3 / 78.4
79.1 / 72.3 / 78.4
Sceptical
80.1 / 75.5 / 76.6
80.1 / 75.5 / 76.6
The ranking changes in no weighting: AD 33 rises in the two measure-weightings (it was unfairly depressed there), AD 31 rises with it because its C1 is now a 3 as well, and the photo finish under the idiom reading stays unchanged — it already had the correction. AD 31 still wins the two measure-weightings by more than twenty points. Precisely because of this, AD 31's loss under the idiom reading is fair: AD 31 does not lose there because it is penalised, but because it has surrendered its sharpest trump (the sign as measure). [CALCULATED]
A second cell stands a point above its own ground. It has been computed through and it has not been lowered — this is the announcement of an open question, not the record of a correction. Criterion C3 gives AD 31 the top mark on the ground "baptism-27 + 4 Passovers natural". Both halves of that ground have been weakened, and weakened by this book's own registered tests: the four-Passover reading is carried as support and not as compulsion, on a test of which two of the three criteria for it were not met (appendix E.3), and the autumn-inclusive count that yields the baptism of 27 is carried as unattested, not refuted: not probable, but not impossible (ch. 9, ch. 14). A cell scoring 3 on a ground that no longer compels is a cell to be looked at.
What lowering it to a 2 would do, computed with the same script and the single cell changed:
weighting
current (C3 = 3) — 30 / 31 / 33
lowered (C3 = 2) — 30 / 31 / 33
ranking
Exact fulfilment
70.9 / 93.3 / 54.8
70.9 / 90.6 / 54.8
unchanged, 31 > 30 > 33
Default
73.7 / 88.0 / 61.7
73.7 / 85.0 / 61.7
unchanged, 31 > 30 > 33
Academic
79.1 / 72.3 / 78.4
79.1 / 69.1 / 78.4
unchanged, 30 > 33 > 31
Sceptical
80.1 / 75.5 / 76.6
80.1 / 72.0 / 76.6
unchanged, 30 > 33 > 31
It would cost AD 31 2.7, 3.0, 3.2 and 3.4 points on the unrounded indices, and it would flip no ranking in any of the four. What it would change is the size of the loss under the idiom reading: AD 31's distance behind AD 33 would grow from 6.1 to 9.3 points in the academic weighting and from 1.1 to 4.6 in the sceptical, so that the near-tie between AD 31 and AD 33 in the sceptical weighting would disappear. [CALCULATED — both columns roll out of `voorwerk/weging_herrekening_juli2026.py`, the second with that one cell set to 2]
And it has not been done, for a reason that is itself a symmetry test. C3 is not a cell of AD 31 alone. AD 30 and AD 33 both stand at 2 on that criterion — AD 30 on either exactly three Passovers or a co-regency from 26, AD 33 on the standard baptism of 29 — and whether those grounds hold their mark any more firmly than AD 31's has not been measured — nobody has asked.Whoever lowers his own cell without measuring whether the other two have earned the mark they hold is measuring with two standards in the other direction. Self-punishment is as asymmetrical as self-favour, and it is no more a measurement for running against one's own interest. The three cells are to be laid against each other first, on one and the same test; only then may any of them be adjusted — and it may then turn out that a different cell moves, or that two do.
Until that has been done, C3 for AD 31 stands at 3, with this section beside it, so that a reader who thinks the cell too high can read off at once what the alternative costs and can carry the whole weighing through with the lower figures.
The version history of this model, in one step:
step
what changed, and why
AD 31 under the default
v3 → v4
the calendar flex (C1/C2) counts in all four weightings, not only in the two AD 31 loses anyway — a double standard removed
84.0 → 88.0
That step raised AD 31 and AD 33 together, because a rule had been applied unevenly. It was not chosen for what it did to the outcome, and the open C3 question is not being settled by what it would do to the outcome either.
F.4 · What this appendix does not deliver
- The merit scores (1–3) are judgements, not measurements. They have been assigned as honestly as possible and checked for symmetry, but another assessor may set a cell one point differently — and one of them, C3 for AD 31, is openly under that question in F.3, where the figures with and without the lower mark stand side by side. The model is robust to such shifts in the default and exact weightings (AD 31 wins there by a clear margin), and sensitive in the academic/sceptical ones (where the differences are small) — exactly as it should be. - There is no composite p-value, and none is coming. The axes share assumptions; multiplying would be scaffolding (ch. 16). - The weights are a choice. The four premises cover the reasonable range of readers, but they are not the only ones possible. Whoever wants different weights adapts the script; the structure lies open. - The convergence family is premise-dependent and stands at zero in the academic/sceptical weightings — the firewall (ch. 2) in figures: reject the 457 line and the exact-fulfilment premise, and the measuring core (calendar, astronomy, chronology, text) remains standing without it.
Reproduction: `voorwerk/weging_herrekening_juli2026.py` (Python 3, no external packages) contains the score matrix of F.1, the four weight transformations of F.2 and the computation rule; it prints the four weightings and their rankings. The figures published above are reproduced one-to-one from it.
*This appendix makes the weighing of Chapter 16 recomputable. It lays the model bare, gives the four premises it computes through, and shows the outcomes. The weighing can be reproduced with a small Python script. Read this first: this model is not proof and delivers no composite p-value. It is an instrument of transparency — it shows how the ranking among AD 30, 31 and 33 moves with the assumptions one brings in. Whoever presents it as proof is building scaffolding.*
F.0 · What the model does and does not do
It does not combine independent probabilities into a single number — that would double-count the shared assumptions (calendar, textual reading). It does something more modest: it gives each year a merit score (1–3) per criterion, weights these, and sums them to an index of 0–100. The gain lies not in the final figure but in the sensitivity: shift the assumptions, and watch whether the ranking flips. Where it flips lies the hinge of the whole case.
F.1 · The eleven criteria and the score matrix
Eleven criteria in five families. Each year receives a merit score per criterion on the scale 0–3 (in practice the assigned scores run from 1 to 3). The score matrix below is the heart of the model. It gives the raw merit scores; the calendar symmetry (C1/C2 flex) counts as a weighting transformation in all four weightings, and the idiom reading (academic/sceptical) additionally overrides C9 — see F.2. [CALCULATED]
Within "Text" one criterion carries the most weight and the entire hinge: criterion 9 — the sign of Jonah. As a measure it scores AD 31 = 3, AD 30 = 1, AD 33 = 1 (the double construction points hard towards the Wednesday). As pliable idiom it becomes neutral (all 2), and the sharpest lead of AD 31 disappears. Two other criteria pull the other way: astronomy (C2: AD 30 = 3, AD 31 = 3, AD 33 = 1 — AD 33 fragile) and reception (C11: AD 33 = 3, AD 30 = 2, AD 31 = 1 — the consensus is comfortable with AD 33). These three contrasts explain almost the entire outcome.
The weights — from criterion to family. The eleven criterion weights sum to 100 and add up within each family to the family weight. Thus every family entry is traceable to its criteria. [CALCULATED]
F.2 · The computation rule and the four weightings
The computation rule. Each year receives one index of 0–100 per weighting. With weights w and cell scores s (per criterion c):
— where the sums run over the eleven criteria — rounded to one decimal. The factor 3 in the denominator is the maximum cell score, so that a year scoring the top mark on every criterion would score exactly 100; criteria with weight 0 drop out. This is exactly what `voorwerk/weging_herrekening_juli2026.py` (Python 3, no external packages) does — the four weightings below and their rankings roll straight out of it. [CALCULATED]
The four weightings. Each weighting is a transformation of the base weights (column "Σ" = the sum of the weights in that weighting), sometimes with overwritten score cells. This makes every outcome line recomputable from the matrix of F.1:
The symmetry correction C1/C2 "flex" — the calendar freedom of F.3, which sets C1 to 3/3/3 and C2 to 3/3/2 — counts in all four weightings, for it is a calendar judgement, not a judgement about the sign. The idiom reading (academic/sceptical) additionally overrides C9 to neutral (the sign is there no longer a measure). With these two tables and the computation rule the four outcomes are fixed:
Worked check (default, AD 31), by way of illustration: with the calendar flex C1 for AD 31 is a 3 (not 2), so numerator = 12·3 + 8·3 + 9·3 + 8·3 + 7·3 + 8·3 + 6·3 + 6·3 + 14·3 + 8·2 + 14·1 = 264; denominator = 3·100 = 300; index = 100·264/300 = 88.0. The remaining cells compute the same way and reproduce the table to the decimal.
Read the table as a whole. Under the premise the book defends — the sign is a measure, God delivers to the day — AD 31 wins decisively: 93.3 against 54.8 for AD 33, and 88.0 against 61.7 in the balanced variant. That is not a neck-and-neck race; that is a win by a distance. Only when one rejects the core premise and reads the sign as a figure of speech does the picture flip to a cluster in which AD 30 leads and AD 31 and AD 33 lie within a few points of each other. The whole case thus turns on a single hinge — is the sign a measure? — and the model makes that visible instead of hiding it.
F.3 · The symmetry correction
An earlier version of this model was unfair on one point: it granted AD 33 the benefit of calendar freedom (Stern's "quite often a late Pesach") but not AD 31. That artificially depressed AD 31. The principle of the correction: whoever grants calendar freedom to one year grants it to both — both AD 31 (the Adar II intercalation) and AD 33 (the pre-equinox March moon) receive the normal calendar judgement (C1 → 3/3/3, C2 → 3/3/2). Neither year is depressed or inflated.
The first implementation (v3), however, applied that principle only in the academic and sceptical weightings — the two in which AD 31 loses anyway — and left it out of the exact-fulfilment and the default, the two in which AD 31 wins. There AD 33 was thus still depressed on calendar grounds that the principle itself forbids. That is a double standard, and it has been removed here: the calendar flex now counts in all four weightings. It is a judgement about the calendar, not about Matthew 12:40, and so must not hang on the sign premise.
By way of check, the four outcomes before and after this universal application:
The ranking changes in no weighting: AD 33 rises in the two measure-weightings (it was unfairly depressed there), AD 31 rises with it because its C1 is now a 3 as well, and the photo finish under the idiom reading stays unchanged — it already had the correction. AD 31 still wins the two measure-weightings by more than twenty points. Precisely because of this, AD 31's loss under the idiom reading is fair: AD 31 does not lose there because it is penalised, but because it has surrendered its sharpest trump (the sign as measure). [CALCULATED]
A second cell stands a point above its own ground. It has been computed through and it has not been lowered — this is the announcement of an open question, not the record of a correction. Criterion C3 gives AD 31 the top mark on the ground "baptism-27 + 4 Passovers natural". Both halves of that ground have been weakened, and weakened by this book's own registered tests: the four-Passover reading is carried as support and not as compulsion, on a test of which two of the three criteria for it were not met (appendix E.3), and the autumn-inclusive count that yields the baptism of 27 is carried as unattested, not refuted: not probable, but not impossible (ch. 9, ch. 14). A cell scoring 3 on a ground that no longer compels is a cell to be looked at.
What lowering it to a 2 would do, computed with the same script and the single cell changed:
It would cost AD 31 2.7, 3.0, 3.2 and 3.4 points on the unrounded indices, and it would flip no ranking in any of the four. What it would change is the size of the loss under the idiom reading: AD 31's distance behind AD 33 would grow from 6.1 to 9.3 points in the academic weighting and from 1.1 to 4.6 in the sceptical, so that the near-tie between AD 31 and AD 33 in the sceptical weighting would disappear. [CALCULATED — both columns roll out of `voorwerk/weging_herrekening_juli2026.py`, the second with that one cell set to 2]
And it has not been done, for a reason that is itself a symmetry test. C3 is not a cell of AD 31 alone. AD 30 and AD 33 both stand at 2 on that criterion — AD 30 on either exactly three Passovers or a co-regency from 26, AD 33 on the standard baptism of 29 — and whether those grounds hold their mark any more firmly than AD 31's has not been measured — nobody has asked. Whoever lowers his own cell without measuring whether the other two have earned the mark they hold is measuring with two standards in the other direction. Self-punishment is as asymmetrical as self-favour, and it is no more a measurement for running against one's own interest. The three cells are to be laid against each other first, on one and the same test; only then may any of them be adjusted — and it may then turn out that a different cell moves, or that two do.
Until that has been done, C3 for AD 31 stands at 3, with this section beside it, so that a reader who thinks the cell too high can read off at once what the alternative costs and can carry the whole weighing through with the lower figures.
The version history of this model, in one step:
That step raised AD 31 and AD 33 together, because a rule had been applied unevenly. It was not chosen for what it did to the outcome, and the open C3 question is not being settled by what it would do to the outcome either.
F.4 · What this appendix does not deliver
- The merit scores (1–3) are judgements, not measurements. They have been assigned as honestly as possible and checked for symmetry, but another assessor may set a cell one point differently — and one of them, C3 for AD 31, is openly under that question in F.3, where the figures with and without the lower mark stand side by side. The model is robust to such shifts in the default and exact weightings (AD 31 wins there by a clear margin), and sensitive in the academic/sceptical ones (where the differences are small) — exactly as it should be. - There is no composite p-value, and none is coming. The axes share assumptions; multiplying would be scaffolding (ch. 16). - The weights are a choice. The four premises cover the reasonable range of readers, but they are not the only ones possible. Whoever wants different weights adapts the script; the structure lies open. - The convergence family is premise-dependent and stands at zero in the academic/sceptical weightings — the firewall (ch. 2) in figures: reject the 457 line and the exact-fulfilment premise, and the measuring core (calendar, astronomy, chronology, text) remains standing without it.
Reproduction: `voorwerk/weging_herrekening_juli2026.py` (Python 3, no external packages) contains the score matrix of F.1, the four weight transformations of F.2 and the computation rule; it prints the four weightings and their rankings. The figures published above are reproduced one-to-one from it.