The core finding
At the innermost terminal positions — where post-mortem deamination peaks — standard MLM fine-tuning performs below random chance. Damage-Aware Masking is the only method consistently above random. p < 0.001 across all six terminal-zone widths.
scroll to expand
At the innermost 3 nucleotides from the fragment terminus — where post-mortem deamination probability peaks — standard MLM fine-tuning achieves 20.5% reconstruction accuracy. Random chance achieves 27.7%. DAM achieves 30.8% and is the only method consistently above random. All six terminal-zone widths: p < 0.001 (paired t-test, n = 626 windows).
Ancient DNA carries a forensic signature: cytosines deaminate to thymine preferentially at 5′ fragment ends; guanines oxidise at 3′ ends. This post-mortem damage (PMD) is non-random and position-dependent, empirically characterised by mapDamage2. Every DNA language model trained before this work applied uniform random masking — building no specialised ability at the exact positions where ancient sequence information is most degraded.
Standard MLM fine-tuning applies uniform 15% random masking across all positions — treating a site at the terminal end of a 44,800-year-old fragment the same as an interior position with near-zero damage probability.
DAM replaces this with a per-position, per-base-type masking probability matrix derived directly from the mapDamage2 misincorporation output. Cytosines at 5′ ends are masked proportionally to their empirical C→T deamination rate; guanines at 3′ ends follow the G→A profile. A/T tokens are never masked — biologically correct since deamination is C/G-specific.
The matrix is rescaled to maintain the same total 15% masking budget as standard MLM, making the comparison controlled: same budget, different spatial distribution.
The DAM advantage is localised to the terminal damage zone and decays inward, matching the shape of the PMD profile exactly. At background positions (authentic low-damage sites), MLM vs DAM shows p = 0.892 — confirming the effect is terminal-specific, not a generic artefact.
# DamageAwareDataCollator — core contribution # masking/collator_dam.py def compute_mask_probs(damage_profile, seq_len): # Per-position PMD probability matrix probs = torch.zeros(seq_len) for pos in range(seq_len): # 5′ C→T deamination (capped at position 69) p_ct = damage_profile['ct_5p'][min(pos, 69)] # 3′ G→A oxidation (from 3′ end) p_ga = damage_profile['ga_3p'][min(seq_len-1-pos, 69)] probs[pos] = max(p_ct, p_ga) # Rescale to 15% average — same budget as MLM mean_cg = probs[probs > 0].mean() probs *= (0.15 / mean_cg) return probs.clamp(0, 1)
Terminal zone reconstruction accuracy — DAM vs competing methods at T_END = 3 (innermost damage positions). All differences significant at p < 0.001.
Protein structural validation — ESMFold TM-scores for reconstructed CDS sequences. All > 0.95 confirms no fold-disrupting mutations introduced.
| Gene | Method | TM-Score | CαRMSD |
|---|---|---|---|
| TRPV3 | DAM | 0.9808 | 1.092 Å |
| KCNK9 | DAM | 0.9520 | 1.732 Å |
| HBB | DAM | 0.9714 | 0.840 Å |
| Biosecurity (615/626 safe) | AUC 0.934 | — | |
All 4 held-out genes (TRPA1, UCP1, ADRB3, FASN — never seen during training) show full generalisation. The 11 biosecurity flags are pre-existing in the Asian elephant reference genome, concentrated in TRPA1 ankyrin repeat domains conserved across kingdoms.