Every result starts as a retail purchase and ends as a public file anyone can rerun. No vendor claims, no private demos. This page is the whole method.
a440 tests publicly observable guarantees. A capability counts only when an independent party can verify it through an accessible, reproducible path. A vendor with a stronger private tool has one move: expose a verifiable interface and we test it.
nota bene: a capability no outsider can check is, for our purposes, indistinguishable from no capability at all.
Every durability verdict on this site is one of these 38 marks. The suite is fixed, named, and cited per test: [SoK25] Wen et al., arXiv:2503.19176; [RAW25] Ozer et al., Interspeech 2025, arXiv:2505.19663; [AS24] San Roman et al., AudioSeal, ICML 2024. Attacks we added ourselves are labeled a440.
The music-specific detection literature anchors the framing: [AMR25] Cros Vila and Sturm, "The AI Music Arms Race", TISMIR 2025, DOI 10.5334/tismir.254 - a commercial AI-music detector is fooled by a 22.05 kHz resample alone; [AFC25] Afchar et al., arXiv:2501.10111 - the first published AI-music detector reports 99.8% accuracy that degrades under pitch shifts, re-encoding and unseen generators. Detector scores alone do not establish robustness, so a440 measures the end-to-end verification path.
The open-detector check runs one published detector against the corpus directly: [SNX25] Rahman et al., SONICS / SpecTTTra, ICLR 2025, arXiv:2408.14080. It flags 0 of 6 current un-attacked AI tracks. Details on the dive page.
Benchmark version, product and model version, test date - on every result. If a vendor fixes something, we test again and publish the next result. History is never rewritten; errata append. A vendor that disputes a result gets it rerun against the published corpus.