a440 News Dive in Corpus
a440 Benchmark v0.1 - Methodology

How a recording earns an a440 verdict.

Every result starts as a retail purchase and ends as a public file anyone can rerun. No vendor claims, no private demos. This page is the whole method.

The principle

a440 tests publicly observable guarantees. A capability counts only when an independent party can verify it through an accessible, reproducible path. A vendor with a stronger private tool has one move: expose a verifiable interface and we test it.

nota bene: a capability no outsider can check is, for our purposes, indistinguishable from no capability at all.

The evidence chain
01
Buy like a customer
Retail account, no special access. We test what any user can download.
02
Read the file
Declared provenance checked with the C2PA reference implementation.
03
Attack it 38 ways
The named battery: container, codecs, temporal, spectral, dynamics, spatial, edits, hostile chains.
04
Ask the vendor's detector
Every copy goes through the public verification path. Verdicts archived verbatim.
05
Publish everything
Corpus, hashes, commands, detector transcripts. Dispute us and we rerun and append.
The six layers
Provenance
Where did this recording come from?
Durability
Does the evidence survive normal use?
Interoperability
Can someone other than the vendor verify it?
Attribution
What existing work contributed to it?
Consent
Was that use authorized?
Accounting
Can the right people actually be paid?
The attack battery - 38 tests, 9 families

Every durability verdict on this site is one of these 38 marks. The suite is fixed, named, and cited per test: [SoK25] Wen et al., arXiv:2503.19176; [RAW25] Ozer et al., Interspeech 2025, arXiv:2505.19663; [AS24] San Roman et al., AudioSeal, ICML 2024. Attacks we added ourselves are labeled a440.

The music-specific detection literature anchors the framing: [AMR25] Cros Vila and Sturm, "The AI Music Arms Race", TISMIR 2025, DOI 10.5334/tismir.254 - a commercial AI-music detector is fooled by a 22.05 kHz resample alone; [AFC25] Afchar et al., arXiv:2501.10111 - the first published AI-music detector reports 99.8% accuracy that degrades under pitch shifts, re-encoding and unseen generators. Detector scores alone do not establish robustness, so a440 measures the end-to-end verification path.

The open-detector check runs one published detector against the corpus directly: [SNX25] Rahman et al., SONICS / SpecTTTra, ICLR 2025, arXiv:2408.14080. It flags 0 of 6 current un-attacked AI tracks. Details on the dive page.

container (2) codecs (5) temporal (4) spectral (3) dynamics (4) spatial (3) edits (6) hostile (2) literature (9)
Thirty-eight marks, thirty-eight attacks, nine families, drawn to scale. PASS criteria are published before the run: the C2PA channel passes if the manifest survives the c2pa reference SDK read, the watermark channel passes if the vendor public detector still finds the mark.
Verdicts - no mystery scoring
PASS
The layer still verifies after the attack.
FAIL
It does not. Reported with the count, never averaged away.
NOT TESTED
No public interface exists. Absence of a path is itself a finding.
N/A
Out of scope for this model. A future a440 score launches once the series can support a stable methodology.
Version everything

Benchmark version, product and model version, test date - on every result. If a vendor fixes something, we test again and publish the next result. History is never rewritten; errata append. A vendor that disputes a result gets it rerun against the published corpus.

See the method appliedSoundcheck 001 - Suno v6-mini and ElevenLabs Music v1 / v2 →