Qurʾānic readings · numerical exploration

Where the twenty transmissions sit

A diffusion map, a map fitted to textual distances, and the exact comparison behind every pair. Explore the pattern without mistaking the picture for a genealogy.

Read the scale carefully. Diffusion coordinates measure similarity through the comparison network. They are not percentages of different Qurʾānic words. Use Metric MDS for an approximate percentage-point distance scale, and the exact distance bars or pair inspector for exact values.

Meccan transmitters: al-Bazzī and Qunbul

These are Ibn Kathīr’s two transmitters. Gold diamonds mark their true positions. This enlarged subset separates them visually from Ruways and Rawḥ (Basra); all twenty remain in the main map above.

Inspect an exact pair

These comparisons come from the full distance matrix. Choosing a pair highlights its locations on the map; clicking a matrix cell selects it here.

compared with

How each transmission compares with the other nineteen

All figures below use original textual dissimilarity, not map coordinates. Rank 1 means the lowest mean or lowest standard deviation. Equal profiles share ranks. The mean is equally weighted across the other nineteen transmission labels.

TransmissionMean %SD · ppMinimum %Maximum %Mean rankSD rank

What was compared

The source is the archived nQuran al-ʿashr al-ṣughrā inventory, associated with al-Shāṭibiyya and al-Durra. The archive reports all 114 surahs as processed. This is a source-defined, coded variant dataset, not a fresh collation of every transmission.

The default subset contains 1,093 fixed word-variant characters in 583 dependency blocks, occurring across 904 verses and 91 surahs, with no missing cells. Its denominator consists of selected documented differences. A result of 30% does not mean 30% of the Qurʾān differs.

The archive contains 33,602 source feature records: 29,349 are represented after the audited repairs (87.3%), while 4,253 are excluded. Some source features generate multiple analytic characters, producing 33,305 characters for the broad accepted-set scenarios. These are different counting units from the default 583 dependency blocks. This version retains the original dataset with three connecting-vowel components reassigned to their existing uṣūl block and one unresolved conditional pronoun-hāʾ component excluded; retained state assignments are unchanged.

Equal blocks gives each dependency block total weight 1, divided among its eligible characters. Equal characters / occurrences weights each eligible character equally, allowing recurring rules to count repeatedly. Neither weighting is historically neutral.

In the broad equal-block scenarios, fixed word variants still receive 93.3–94.0% of total weight. Their similarity to the default map is therefore partly built into the weighting, rather than independent validation.

Accepted sets uses Jaccard dissimilarity between the alternatives encoded for each transmission. These broad sets describe marginal compatibility and can combine options that do not form one coherent reading route. They are sensitivity comparisons.

What the maps can tell you

Diffusion: the kernel is exp(−D/ε), followed by density normalization with α, row normalization, and eigenvector coordinates λᵗψ. ε is the selected multiple of the median positive original dissimilarity. A two-axis view loses some full diffusion distance. Larger t emphasizes slower network patterns; it is not elapsed historical time.

Metric MDS: directly fits the original dissimilarities in two dimensions. Axes use percentage-point distance units and equal physical scale. The fit is approximate; the exact pair values remain authoritative.

Weighted PCA: uses a one-hot representation of fixed states. Full-space squared Euclidean distance equals weighted mismatch. The displayed coordinates have square-root mismatch units, not percentage-point units. Explained variance concerns this coded dataset.

Sidky comparison: his published consonantal-dotting study used PCA on 292 variant words and ten reader profiles, allowing transmitter alternatives. This twenty-transmission diffusion analysis is an extension using a different dataset, not a replication of his graph.

Geographic colors are traditional reader affiliations added after calculation. They do not enter the distances. Position, centrality, and clusters alone establish neither chronology, ancestry, fidelity, nor historical popularity.

Sources and reproducibility

nQuran source collection · Sidky, Consonantal Dotting and the Oral Quran (2023) · Author full text, especially Appendix A · Coifman & Lafon, Diffusion maps (2006)

This standalone file works offline and embeds all six scenarios and 36 diffusion parameter settings per scenario. The accompanying reproduction guide contains calculation scripts, numerical checks, coordinates, distance matrices, and the exact audit change log. Bootstrap intervals use 500 block-resampling replicates conditional on the default coded data and weighting; they are not historical confidence statements or estimates of annotation accuracy.