How closeness scoring works
How closeness scoring works
Back to README · conventions
Three lenses. The site shows each tradition’s results through three lenses, side by side:
- Scripture closeness: agreement with the grammatical-historical reading of the biblical texts for each question (§3b).
- Fathers closeness (proximity-weighted): closeness to the patristic verdicts, weighted by how early the evidence is (§1–3a).
- Weighted average (reader’s lens):
lens(t) = w × scripture(t) + (1 − w) × fathers(t), with w = 0.5 by default and adjustable by the reader with the slider on the homepage. If a tradition has only one of the two scores, the lens shows that one.
The third lens is optional and not the verdict. The traditions themselves differ on how Scripture and tradition should be weighted (see Scripture and tradition), so the site does not choose a weighting; the default 50/50 is a neutral starting point, not a claim. The underlying verdicts and their reasoning, not any of the three numbers, are the research’s conclusions. The old unweighted patristic score is kept in the data as overall_unweighted for comparison only.
Traditions are read from the data, so the lenses and per-doctrine matrix show however many traditions the doctrine pages cover. All three lenses are tier-weighted (§3e).
The patristic scores summarise what the category files already say. They add no new judgement: each score is read mechanically off a category’s post-750 development table, using the rules below, and every score cites the row (with its anchor) it came from. The generated scores are in data/closeness.yaml, produced by scripts/build_data.py on every build, so new categories are scored automatically.
1. Rubric: verdict to a 0–4 closeness score
For each category question (Q1, B1, S1, …) and each of the nine traditions (Catholic, Orthodox, Oriental Orthodox, Lutheran, Anglican, Presbyterian, Reformed Baptist, General Baptist, Anabaptist), the tradition’s row in the post-750 table is split at commas and semicolons, and the question IDs it names are read (ranges such as B1–B4 count every question in the range).
| Score | Meaning | Rule |
|---|---|---|
| 4 | Matches the fathers | The question is listed under Closest to the fathers on and not under Departs or sharpens on. |
| 3 | Matches one patristic strand, or partly | Listed under Closest with “partly”, or the question’s verdict is itself split (wording “Split”, “vs”, or “East: … West”) so a tradition can only follow one strand. |
| 2 | Mixed | Listed under both columns, or under Departs with “partly”. |
| 1 | Departs or sharpens | Listed under Departs or sharpens on only. |
| 0 | Rejects | Listed under Departs with “rejected” or “no change” (denying the verdict outright). |
| not assessed | — | The file does not state this tradition’s position on this question. No score is guessed, and it is left out of the averages. |
Rule O (overrides). Where a row names a position in words rather than by question ID (for example “Nicaea II” or “Vigilantius’s side”), or the Overall verdict states it explicitly (“No tradition departs on O1”), the score is set by hand in scripts/closeness_overrides.json. Each override quotes its basis and cites its anchor, and is marked * on the scores page. Row notes that map to no question (e.g. “limited atonement”) are listed under each category on the scores page and are not scored.
2. Confidence weighting
Each question’s weight comes from the Confidence column of its verdict:
| Confidence | Weight w |
|---|---|
| High | 1.0 |
| Medium–high | 0.75 |
| Medium | 0.5 |
| Low–medium | 0.375 |
| Low | 0.25 |
When a cell gives two levels (e.g. “High / medium”, or “High; infants’ fate low consensus”), the lower one is used, so a partly uncertain verdict counts for less.
3. Formula
For tradition t, over the set Q(t) of questions where t has a score s:
closeness(t) = Σ_{q ∈ Q(t)} w_q · s_{q,t} ÷ Σ_{q ∈ Q(t)} w_q (0–4)
percent(t) = closeness(t) ÷ 4 × 100The same formula gives each category’s mini ranking (over that category’s questions) and the overall ranking (over all questions in all categories). Each ranking also shows how many questions a tradition’s score rests on; a score resting on fewer than half the questions is shown faded, since it is less reliable.
3a. Source-proximity tiers
| Tier | Span | Weight |
|---|---|---|
| Scripture | — | benchmark (separate axis) |
| T1 Apostolic Fathers | to c. 150 | 1.0 |
| T2 | 150–250 | 0.8 |
| T3 | 250–451 | 0.6 |
| T4 | 451–750 (Nicaea II, 787, included) | 0.4 |
The weights fall by 0.2 per tier. That is a deliberately simple, linear choice; it is not derived from anything. Proximity factor of a question = the mean tier weight of the dated key sources in its verdict row. The patristic weight is confidence weight × proximity factor. So a verdict resting only on 8th-century evidence counts 40% as much as an equally confident verdict resting on T1. Developmental questions (B2, R1, R4, I3: dates of attestation and history of disputes) get factor 1.0, because their late date is the finding.
3b. Scripture axis
For each non-developmental question, scripts/scripture_scores.json records the grammatical-historical reading (also printed in each reasoning file’s Step X.4), a Scripture confidence (weighted with the same scale as §2), and a 0–4 score per tradition with a one-line basis citing its confession. Rubric: 4 = the tradition’s position states what the texts state; 3 = consistent with the texts but goes beyond them, or underplays part of them; 2 = adds a claim the texts do not support, or strains a text; 1 = conflicts with the plain sense of a text; 0 = denies it. Questions that are not about Scripture (Q3, E3, O3) and developmental ones are left out. Only the canon all four traditions accept is scored. 2 Maccabees and the other deuterocanonical books are noted where relevant (L2), but they are not scored, and that choice disadvantages Catholic and Orthodox positions on L2.
3c. The weighted-average lens
An earlier version fixed a 60/40 blend as a headline score. That was removed. The blend now exists only as the reader-adjustable third lens described above, computed in the browser (with a 50/50 server-rendered default). It is never stored as a score in data/closeness.yaml.
3d. Judgement calls
The Scripture scores are judgements, not mechanical readings. The calls most likely to be disputed, in either direction, are listed here so readers can check them:
- B1 baptismal regeneration: the plain sense of John 3:5, Acts 2:38, Titus 3:5 and 1 Pet 3:21 is read as tying new birth to the washing.
- B3 mode: βαπτίζω and the burial imagery are read as favouring immersion.
- E1 real presence: the realist reading of 1 Cor 10–11 is judged the more natural one.
- E2 sacrifice: Hebrews’ ἐφάπαξ (“once for all”) is read against a repeated propitiatory offering.
- R2, R3 relics: miracles through relics are affirmed from 2 Kings 13 and Acts 19, while honorific prostration is read against from Acts 10 and Rev 19.
- I1, I2 images: the cherubim precedent is weighed against Deut 4 and Exod 20:5.
- S2, S3, S4 salvation: real choice and bondage held together; election by God’s purpose (Rom 9) preferred to foreseen faith; no parallel reprobation decree; preservation and warnings held together.
- Q1, Q2, Q6 justification: δικαιόω read as declarative; positive imputation treated as an inference, not an explicit text.
- L2, L3 afterlife: only the canon all four traditions share is scored, so 2 Maccabees is noted but not scored on L2, which affects the Catholic and Orthodox positions there.
- Where a tradition goes beyond the text in either direction, it is capped at 3.
Readers who dispute a reading can change one line in scripts/scripture_scores.json and rebuild.
3e. Tier weighting (“most correct”)
Since the tier migration (triage, structure), every ranking on the home page is tier-weighted and doctrine-balanced. First, each doctrine’s questions are averaged into one doctrine score, confidence-weighted as in §2 and §3. Then:
most_correct(t) = Σ_d w(tier_d) · score(t, d) / Σ_d w(tier_d) w = first-order 3, second-order 2, third-order 1- Doctrines a tradition isn’t scored on drop out of both sums. Doctrines not yet researched count for nothing.
- Each doctrine counts once, however many questions it has, so S19 (7 questions) does not outweigh S05 (1).
- The per-question variant, Σ w(tier)·w(q)·s / Σ w(tier)·w(q), is published as
most_correct_per_questionindata/closeness.yamlfor comparison. - The same rule applies to the fathers lens. The reader’s weighted-average lens blends the two tier-weighted scores.
- The untiered totals (
overall,overall_scripture) are still generated.scripts/parity_check.pychecks them, and every per-question score, against git tagpre-tier-migration. The migration changed none of them.
Tier sensitivity. Some traditions place a doctrine in a different tier (triage §3; each doctrine page lists them under “Where the tier is contested”). For each tradition, the score is recomputed with its own tiers, and the scores page shows whether its rank changes.
3f. In the faith
This uses first-order doctrines only, each scored against Scripture as a confidence-weighted doctrine average out of 4. The band follows the weakest-link rule:
| Band | Rule |
|---|---|
| holds all | every researched first-order doctrine scores at least 3/4 |
| holds with qualifications | at least one first-order doctrine scores below 3/4, and none at 1/4 or below |
| denies one or more | at least one first-order doctrine scores 1/4 or below |
- The band depends only on the lowest first-order doctrine.
- The percentage shown is the mean of all first-order doctrines. It is reported for information and never decides the band. That’s why a tradition with a lower mean can be “holds all” when every doctrine clears 3/4, while one with a higher mean is “with qualifications” because one doctrine falls below. For example, Anabaptist (93%, lowest 3.0) is “holds all”, and Orthodox (96%) is “with qualifications” because of F12 (2/4).
- The site names every doctrine that triggers a qualification.
- The band is marked provisional while any first-order doctrine is unresearched.
- A second figure uses each tradition’s own first-order list (scores page).
3g. Shaped by
This is built directly from scripts/emphasis.json. Each tradition gets an emphasis of 0–2 for each doctrine, from its own confessions, and each value carries a citation:
- 0: not a confessional point
- 1: stated
- 2: a defining or polemical point
shaped_by(t)[tier] = Σ emphasis(t, d) over doctrines d in that tier / Σ emphasis(t, d) over all doctrines
top 3 distinctives = the second- or third-order doctrines with the highest emphasis (ties: tier order, then doctrine number)- This doesn’t depend on Scripture or patristic scores. A doctrine counts here even if it has no Scripture score (for example S02, infant baptism, which is developmental).
- It describes how a tradition’s identity is weighted, not whether it is right.
- The earlier “distance from the median” rule was dropped.
Nine traditions
The old “Presbyterian” column is now Presbyterian (WCF 1646 with the Larger and Shorter Catechisms), and its scores are unchanged. Five traditions were added, all scored from classic formulations only: Anglican, Oriental Orthodox, Reformed Baptist, General Baptist and Anabaptist (sources in notes/confessions.md). Their fathers scores are set explicitly under rule O, in closeness_overrides.json, and each cites its tradition’s post-750 section. Their Scripture scores are in scripture_scores.json. Cells that were “not assessed” for Catholic, Orthodox, Lutheran and Presbyterian were filled where the category file supports a position; existing scores were not changed. Every Oriental Orthodox cell is to verify (thin sources).
4. Limits
- Scores inherit the provisional status of the verdicts. Change a verdict or a post-750 row, and the score changes on the next build.
- Every question counts as one unit (times its confidence); the rubric does not judge which questions matter most.
- “Not assessed” is common (see the scores page): a tradition with few assessed questions can rank high on little evidence. Check the question count next to each score.