Qurʾānic readings · numerical exploration

Where the twenty transmissions sit

A diffusion map, a map fitted to textual distances, and the exact comparison behind every pair. Explore the pattern without mistaking the picture for a genealogy.

Read the scale carefully. Diffusion coordinates measure similarity through the comparison network. They are not percentages of different Qurʾānic words. Use Metric MDS for an approximate percentage-point distance scale, and the exact distance bars or pair inspector for exact values.

Meccan transmitters: al-Bazzī and Qunbul

These are Ibn Kathīr’s two transmitters. Gold diamonds mark their true positions. This enlarged subset separates them visually from Ruways and Rawḥ (Basra); all twenty remain in the main map above.

Inspect an exact pair

These comparisons come from the full distance matrix. Choosing a pair highlights its locations on the map; clicking a matrix cell selects it here.

compared with

How each transmission compares with the other nineteen

All figures below use original textual dissimilarity, not map coordinates. Rank 1 means the lowest mean or lowest standard deviation. Equal profiles share ranks. The mean is equally weighted across the other nineteen transmission labels.

TransmissionMean %SD · ppMinimum %Maximum %Mean rankSD rank

What was compared

The source is the archived nQuran al-ʿashr al-ṣughrā inventory, associated with al-Shāṭibiyya and al-Durra. The archive reports all 114 surahs as processed. This is a source-defined, coded variant dataset, not a fresh collation of every transmission.

The default subset contains 1,096 fixed word-variant characters in 586 dependency blocks, occurring across 907 verses and 91 surahs, with no missing cells. Its denominator consists of selected documented differences. A result of 30% does not mean 30% of the Qurʾān differs.

The archive contains 33,602 source feature blocks: 29,350 were semantically coded (87.3%), while 4,252 were excluded. Some source features generate multiple analytic characters, producing 33,306 characters for the broad accepted-set scenarios. These are different counting units from the default 586 dependency blocks.

Equal blocks gives each dependency block total weight 1, divided among its eligible characters. Equal characters / occurrences weights each eligible character equally, allowing recurring rules to count repeatedly. Neither weighting is historically neutral.

In the broad equal-block scenarios, fixed word variants still receive 93.3–94.1% of total weight. Their similarity to the default map is therefore partly built into the weighting, rather than independent validation.

Accepted sets uses Jaccard dissimilarity between the alternatives encoded for each transmission. These broad sets describe marginal compatibility and can combine options that do not form one coherent reading route. They are sensitivity comparisons.

What the maps can tell you

Diffusion: the kernel is exp(−D/ε), followed by density normalization with α, row normalization, and eigenvector coordinates λᵗψ. ε is the selected multiple of the median positive original dissimilarity. A two-axis view loses some full diffusion distance. Larger t emphasizes slower network patterns; it is not elapsed historical time.

Metric MDS: directly fits the original dissimilarities in two dimensions. Axes use percentage-point distance units and equal physical scale. The fit is approximate; the exact pair values remain authoritative.

Weighted PCA: uses a one-hot representation of fixed states. Full-space squared Euclidean distance equals weighted mismatch. The displayed coordinates have square-root mismatch units, not percentage-point units. Explained variance concerns this coded dataset.

Sidky comparison: his published consonantal-dotting study used PCA on 292 variant words and ten reader profiles, allowing transmitter alternatives. This twenty-transmission diffusion analysis is an extension using a different dataset, not a replication of his graph.

Geographic colors are traditional reader affiliations added after calculation. They do not enter the distances. Position, centrality, and clusters alone establish neither chronology, ancestry, fidelity, nor historical popularity.

Sources and reproducibility

nQuran source collection · Sidky, Consonantal Dotting and the Oral Quran (2023) · Author full text, especially Appendix A · Coifman & Lafon, Diffusion maps (2006)

This standalone file works offline and embeds all six scenarios and 36 diffusion parameter settings per scenario. The accompanying data and code bundle contains calculation scripts, numerical checks, coordinates, and distance matrices. Bootstrap intervals use 500 block-resampling replicates conditional on the default coded data and weighting; they are not historical confidence statements or estimates of annotation accuracy.