QIRAAT: MINIMAL AUDITED REPAIR OF THE ORIGINAL DATASET
Version: original-20260920-audited-repair-20261006-v1

This package starts from the original September 20 dataset used for the report.
It does not use the September 30 rebuilt dataset or recover excluded cases.

WHAT CHANGED
1. nq_t2_s004_f00707_c1, Quran 4:66, aw ukhruju
2. nq_t2_s017_f00695_c1, Quran 17:110, aw ud'u
3. nq_t2_s073_f00002_c1, Quran 73:3, aw inqus
   These three components retain every original transmission-state assignment.
   Their stratum becomes usul and their block becomes the EXISTING
   usul:connecting_two_sukuns. They no longer have separate farsh block weights.

4. nq_t2_s036_f00147_c1, Quran 36:35, 'amilathu aydihim
   This pronoun-ha sila component is excluded because its original short-ha
   default does not safely represent transmissions omitting that ha. No new
   inferred state is substituted. The original source record is retained in
   the baseline archive and the exact exclusion is in change_log.json.

Every other retained row and transmission-state code is unchanged. The script
asserts this row by row before any numerical computation. Original source
files are never overwritten. Changes concern analysis eligibility/grouping;
they do not assert that nQuran's source statements themselves changed.

COUNTS
Source-listed records: 33,602.
Records represented in repaired analysis: 29,349 (previously 29,350).
Excluded source records: 4,253 (previously 4,252).
Coded components: 33,305 (previously 33,306).
All-singleton components: 15,214; multistate components: 18,091.
Farsh components: 1,093; usul components: 32,212.
All blocks: 625 = 583 farsh + 42 usul.
All-singleton subset: 620 blocks.
Coverage: 114 surahs; 20 transmissions of the ten readings.
These are source-listed variant records/components, not all Quranic words and
not 33,305 independent observations.

METHODS PRESERVED
The six original scenarios retain their definitions and parameter choices.
Jaccard accepted-set dissimilarity is 1 - intersection_size / union_size.
Fixed components have exactly one state for every transmission.
Equal blocks assigns each retained block equal total weight; retained
components divide their block's weight equally. Equal-component alternatives
are also regenerated. No state choice is given an invented probability.

Diffusion: K = exp(-D/epsilon), epsilon = median positive off-diagonal D,
alpha = 1, t = 1. Reversible Markov normalization and all nonconstant modes
define full distances. Display axes are sign-aligned to the original axes;
no Procrustes rotation, axis mixing, stretching or geographical fitting occurs.

PCA: one-hot indicators weighted by sqrt(w/2), equally centered across the
20 transmissions, with no per-column standardization. Full squared PCA
distance equals fixed-feature weighted mismatch.

The original metric MDS settings, 36-parameter diffusion grid per scenario,
and 500-replicate fixed-farsh bootstrap are regenerated by analysis_engine.py.
The latter reproduces the original atlas sensitivity display; it is conditional
on analyzed blocks and is not a historical confidence interval.

Canonical results contain full mean/SD/CV for other19 and for other18 excluding
the sister. SD is population SD (ddof=0). Equal ranks use tolerance 1e-12.

FILES
results.json                     Canonical results for all report/figure tools.
matrix_corrected.json.gz         Full source-linked repaired matrix.
input_characters_corrected.json.gz Compact repaired numeric input.
counts_manifest.json             Counts, provenance and checksums.
change_log.json                  Exact four changes and reasons.
validation.json                  Independent computational checks.
report_statistics.csv            Sister-excluded full diffusion statistics.
full_diffusion_distances.csv      Repaired 20-by-20 full-distance matrix.
atlas_data/                      Original atlas schema, regenerated CSVs.
baseline/                        Immutable copies of original inputs/references.
analysis_engine.py               Audited original six-scenario implementation.
build_repaired_results.py         Portable repair/recomputation/check driver.

REPRODUCE
Python 3 with numpy, scipy and scikit-learn:
    python build_repaired_results.py
To produce only the matrices, changes and counts:
    python build_repaired_results.py --prepare-only
The code resolves paths relative to itself; no workspace-specific path is used.

VALIDATION
The corrected raw matrices are independently reconstructed from accepted sets
and weights. Full diffusion coordinates are recomputed independently and
verified against direct Markov transition-profile distances. PCA is checked
against an independent one-hot Gram-matrix computation and the squared-distance
identity. All checks must pass within 1e-9. The repaired Hafs mean/SD must match
the separately calculated audit scenario within 1e-9. Baseline hashes and
repaired compressed/uncompressed hashes are recorded.

INTERPRETATION
Distances and rankings are conditional on this source inventory, conservative
coding, block definitions and chosen distance metric. The equal-component
sensitivity remains relevant and must not be hidden. Diffusion distances are
not textual percentage differences. Spectral energy is global, so a 3D picture
does not retain that same percentage of every pair or point-origin distance.
The coordinate origin is the stationary-weighted mean, not an equidistant
centre. Region labels do not participate in computing positions. Historical
explanations such as Hajj require independent evidence.
