What was compared
The source is the archived nQuran al-ʿashr al-ṣughrā inventory, associated with al-Shāṭibiyya and al-Durra. The archive reports all 114 surahs as processed. This is a source-defined, coded variant dataset, not a fresh collation of every transmission.
The default subset contains 1,093 fixed word-variant characters in 583 dependency blocks, occurring across 904 verses and 91 surahs, with no missing cells. Its denominator consists of selected documented differences. A result of 30% does not mean 30% of the Qurʾān differs.
The archive contains 33,602 source feature records: 29,349 are represented after the audited repairs (87.3%), while 4,253 are excluded. Some source features generate multiple analytic characters, producing 33,305 characters for the broad accepted-set scenarios. These are different counting units from the default 583 dependency blocks. This version retains the original dataset with three connecting-vowel components reassigned to their existing uṣūl block and one unresolved conditional pronoun-hāʾ component excluded; retained state assignments are unchanged.
Equal blocks gives each dependency block total weight 1, divided among its eligible characters. Equal characters / occurrences weights each eligible character equally, allowing recurring rules to count repeatedly. Neither weighting is historically neutral.
In the broad equal-block scenarios, fixed word variants still receive 93.3–94.0% of total weight. Their similarity to the default map is therefore partly built into the weighting, rather than independent validation.
Accepted sets uses Jaccard dissimilarity between the alternatives encoded for each transmission. These broad sets describe marginal compatibility and can combine options that do not form one coherent reading route. They are sensitivity comparisons.