Method

How we measure key detection — the benchmark, the scoring, the numbers

"Accurate" is not a number. A key detection figure only means something once you know three things: which tracks it was measured on, how a near miss was scored, and what share of results the tool was willing to stand behind. This post answers all three for Harmoniq.

Every tool in this category says it is accurate. Almost none of them say what that means.

That is not usually dishonesty — it is that "accuracy" is genuinely ambiguous here, and stating it properly takes more than one number. So before we give ours, here is the frame you should be applying to anybody's figure, including ours.

Three questions that turn a claim into a measurement

Measured on which tracks? A detector tuned on a particular set of records will look excellent on that set. The number only transfers if the material is representative and, critically, if someone other than the toolmaker decided what the right answer was.

How was a near miss scored? Not every wrong answer is equally wrong. Call a track in A minor "E minor" and you have named its dominant — a fifth away, and in practice a blend that still works. Call it "C major" and you have named its relative major, which shares all the same notes. Call it "F# major" and you have produced something that will clash audibly. A scoring rule that treats those three the same is throwing away the information a DJ actually cares about.

What share of answers is the tool standing behind? A detector that gives an answer for everything, including the tracks that genuinely have no unambiguous key, is quietly padding its own score with guesses. One that flags its uncertain cases is doing something harder and more useful.

The reference set

There is a standard reference collection for key detection in electronic music, annotated by listeners rather than by another algorithm — which matters, because a set labelled by software inherits that software's mistakes and rewards anyone who makes the same ones. It is what published work in the field measures against.

We test on 205 unambiguously annotated tracks from it. "Unambiguously" is doing real work in that sentence: tracks whose annotation is contested or genuinely modal are excluded rather than counted as wins or losses, because scoring a track that has no single right answer tells you nothing about the detector.

The scoring rule

We report MIREX scoring alongside raw exact-match, because the two answer different questions.

MIREX is the weighting used in music information retrieval. An exact match counts fully. A perfect fifth counts half. A relative or parallel key counts less than that. It exists precisely because a fifth still mixes — the error is real, but it is not the same size as a random wrong key.

Raw exact match is the stricter number and the easier one to compare across tools. MIREX is the one that better predicts whether the tool will actually cost you a transition.

Our numbers

On that set, with the current engine:

Published independent measurements put the best commercial and academic engines on this benchmark in a comparable range on MIREX. Widely bundled DJ software sits lower, and tonal-profile methods — still the basis of most free tools — land lower still.

One honest caveat, which we also print on the site: annotation revisions of this collection differ between versions, so cross-tool figures should be read as indicative rather than exact. Anyone who tells you their benchmark comparison is precise to the decimal across tools measured at different times is overselling it.

Why two accuracy claims are often not comparable

You will see other tools publish accuracy figures, sometimes very good ones, measured on collections they assembled and verified themselves. That is a legitimate thing to do, and a large self-assembled set can be more representative of real DJ libraries than an academic benchmark is.

But it is not the same kind of evidence, and the two numbers cannot be lined up next to each other. A figure on the standard set is comparable to everything else published on the standard set. A figure on a private set is comparable to nothing — not because it is wrong, but because nobody else has measured on it.

When you are comparing tools, the question to ask is not "whose number is bigger". It is "are these two numbers measured on the same thing?" Usually they are not.

The number we care about more

Here is the part that actually changed how we built the product.

A detector's headline accuracy tells you how it does on average. It tells you nothing about the specific track in front of you right now — and that is the only track you care about at 1am.

So the engine reports how confident it is on every single result, on a calibrated scale. Our high-confidence threshold is 70, and on the benchmark set that threshold covered 80% of the tracks. The remaining fifth are the ones the engine marks rather than guesses at: minimal techno built on one looping chord that genuinely has no unambiguous key, modal deep house sitting between two answers. Those get flagged, with a second-best guess, for you to settle by ear.

A wrong key told confidently is worse than an honest "I'm not sure". A detector that is right 76.6% of the time and tells you which 76.6% is a fundamentally more useful instrument than one that is right 80% of the time and won't say.

What we do not claim

We do not claim to be the most accurate key detector available. We claim to have measured on the set the field publishes against, to have published the method alongside the number, and to tell you per track when we are unsure.

We also do not claim the confidence score is perfect. It is calibrated against the same annotated material, and calibration drifts as an engine changes — which is why the corrections users make in the app feed back into it.

If you want to check any of this, the most direct way is to run your own library through it and look at what gets flagged. You know your own tracks better than any benchmark does.

Try it on your own library Runs in your browser. Free to start — no card.

← All posts

© 2026 Harmoniq DJ Terms Privacy Refunds Contact