Portfolio

The core finding

30.8% DAM (proposed)
27.7% Random chance
20.5% MLM fine-tuning ⚠

At the innermost terminal positions — where post-mortem deamination peaks — standard MLM fine-tuning performs below random chance. Damage-Aware Masking is the only method consistently above random. p < 0.001 across all six terminal-zone widths.

scroll to expand

Python DNABERT-2 PyTorch HuggingFace ESMFold mapDamage2 W&B Manuscript Prep
arXiv:2607.27712 → GitHub
44,800 Years old — woolly mammoth aDNA
13% Lower validation loss vs standard MLM
626 Evaluation windows across 7 genes
0.934 Biosecurity classifier AUC
The core finding

Standard fine-tuning is worse than random at the positions that matter most

At the innermost 3 nucleotides from the fragment terminus — where post-mortem deamination probability peaks — standard MLM fine-tuning achieves 20.5% reconstruction accuracy. Random chance achieves 27.7%. DAM achieves 30.8% and is the only method consistently above random. All six terminal-zone widths: p < 0.001 (paired t-test, n = 626 windows).

Ancient DNA carries a forensic signature: cytosines deaminate to thymine preferentially at 5′ fragment ends; guanines oxidise at 3′ ends. This post-mortem damage (PMD) is non-random and position-dependent, empirically characterised by mapDamage2. Every DNA language model trained before this work applied uniform random masking — building no specialised ability at the exact positions where ancient sequence information is most degraded.

Pipeline

End-to-end reconstruction and validation

RAW aDNA SRA · ERP008929 2 specimens mapDamage2 PMD profiling C→T · G→A rates DAM COLLATOR PMD-weighted masking · 15% budget core contribution DNABERT-2 117M params fine-tuned ×17 ep ESMFold protein structure TM-score > 0.95 BIOSEC CNN 1D · AUC 0.934 10 virulence classes
Core implementation

Damage-Aware Masking — one principled change

Standard MLM fine-tuning applies uniform 15% random masking across all positions — treating a site at the terminal end of a 44,800-year-old fragment the same as an interior position with near-zero damage probability.

DAM replaces this with a per-position, per-base-type masking probability matrix derived directly from the mapDamage2 misincorporation output. Cytosines at 5′ ends are masked proportionally to their empirical C→T deamination rate; guanines at 3′ ends follow the G→A profile. A/T tokens are never masked — biologically correct since deamination is C/G-specific.

The matrix is rescaled to maintain the same total 15% masking budget as standard MLM, making the comparison controlled: same budget, different spatial distribution.

The DAM advantage is localised to the terminal damage zone and decays inward, matching the shape of the PMD profile exactly. At background positions (authentic low-damage sites), MLM vs DAM shows p = 0.892 — confirming the effect is terminal-specific, not a generic artefact.

# DamageAwareDataCollator — core contribution
# masking/collator_dam.py

def compute_mask_probs(damage_profile, seq_len):
    # Per-position PMD probability matrix
    probs = torch.zeros(seq_len)

    for pos in range(seq_len):
        # 5′ C→T deamination (capped at position 69)
        p_ct = damage_profile['ct_5p'][min(pos, 69)]
        # 3′ G→A oxidation (from 3′ end)
        p_ga = damage_profile['ga_3p'][min(seq_len-1-pos, 69)]
        probs[pos] = max(p_ct, p_ga)

    # Rescale to 15% average — same budget as MLM
    mean_cg = probs[probs > 0].mean()
    probs *= (0.15 / mean_cg)
    return probs.clamp(0, 1)
Results

Five-part validation

Terminal zone reconstruction accuracy — DAM vs competing methods at T_END = 3 (innermost damage positions). All differences significant at p < 0.001.

DAM (proposed)
30.8%
Random chance
27.7%
MLM fine-tuning
20.5% ⚠
Zero-shot DNABERT-2
15.5%
Val loss (DAM)
3.2736
Val loss (MLM)
3.7568

Protein structural validation — ESMFold TM-scores for reconstructed CDS sequences. All > 0.95 confirms no fold-disrupting mutations introduced.

GeneMethodTM-ScoreCαRMSD
TRPV3DAM0.98081.092 Å
KCNK9DAM0.95201.732 Å
HBBDAM0.97140.840 Å
Biosecurity (615/626 safe)AUC 0.934—

All 4 held-out genes (TRPA1, UCP1, ADRB3, FASN — never seen during training) show full generalisation. The 11 biosecurity flags are pre-existing in the Asian elephant reference genome, concentrated in TRPA1 ankyrin repeat domains conserved across kingdoms.

Full source code

Reproduce all five validation phases