Measurement

Why key detection struggles with tech house — a 6,281-track measurement

6 min read

Same engine, same library, four genres — and a 16-point gap between them. We measured whether the gap comes from the engine hesitating or from the track having no key information to find. It is the second, and that changes how far you can trust a key tag.

Two tools give the same tech house track two different keys. Which one is lying?

Most of the time, neither. This post is the measured version of that answer.

What we measured, and what we did not

We have 6,281 tracks from a DJ friend's library, sorted into four folders: Deep house (1,833), House (1,698), Organic house (1,146) and Tech house (1,604). Every track carries a Mixed In Key tag. What we measure is how often Harmoniq agrees with those tags — not absolute accuracy. The tags are not perfect either: on the 635 tracks two libraries had in common, Mixed In Key contradicts itself on 45 (7.1%). Our full method is in the measurement post.

Two caveats up front:

We could not find a published genre-by-genre key detection measurement. That does not mean none exists — only that we found nothing to compare against.

The numbers

GenreAgreementCorrect key in top 2 candidatesIn top 3Agreement between a track's sections
Organic house90.9%97.4%99.1%84.2%
Deep house90.6%97.4%98.6%78.8%
House84.7%95.9%98.5%73.2%
Tech house75.0%86.3%90.8%68.7%

The first column is the headline: tech house sits 16 points behind organic house. The next two columns are the interesting part.

The real question: is the engine wrong, or is the track silent?

For every track the engine produces a ranked list of candidate keys. The "correct key in top 2" column asks: does the engine know the answer but fail to pick it?

In the other genres the answer is no: the correct key is almost always in the top two (96–97%). The engine is not struggling to choose. In tech house, the correct key is missing from the top two on 13.7% of tracks, and missing from the top three on 9.2%.

That separates two very different stories:

Tech house looks like the second story.

The evidence: the track contradicts itself

The engine splits a track into sections, decides a key for each one, and combines them. "Agreement between sections" is how many of them land on the final key.

In organic house, 84.2% of sections agree. In tech house, 68.7%. But the useful information appears when you split correct and incorrect tracks:

GenreSection agreement on tracks it got rightOn tracks it got wrong
Organic house86.3%63.2%
Deep house80.8%59.5%
House76.5%55.3%
Tech house74.4%51.4%

The same pattern in every genre: on the tracks the engine gets wrong, the track's own sections also agree with each other less. Wrong results cluster where the track itself gives conflicting key information. In tech house that conflict is both more common (68.7% on average) and deepest on the misses (51.4%).

The error types point the same way. A miss by a fifth — the kind you can still mix — is 13.9% of tech house tracks. But a completely unrelated key is 6.8% in tech house against 1.1–1.6% in the other three genres, and a half-step error is 2.2% against 0.0–0.2%. So in tech house the difference is not just more misses, but misses that land further away.

Why? (this part is a hypothesis, not a measurement)

We did not measure this, so we say so plainly. Tech house often runs on very little harmonic material: long drum sections, a single looping riff, a bass line that walks instead of sitting on the root. When a track carries little chord information, its sections can disagree, because each one holds a small, noisy piece of evidence.

We also tried finding the key's root from the bass line. The gain was +0.5 points, inside the measurement noise; in house and techno the bass often does not sit on the root. That fits the hypothesis, but it does not prove it.

What we tried, and what did not work

Once the gap looked like an evidence problem, we still tried two things. Both were measured, and both were closed:

1. Weighting sections by how certain each one is. The idea: listen harder to the section that speaks clearly. The result: the plain average was already the best. Overall agreement 85.1%, best weighted variant 85.0%; tech house went from 75.0% to 74.9%. Same on both libraries. 2. Taking confidence from a learned model. Instead of a hand-written formula, a confidence learned from nine features (section agreement, entropy and others). The result: how well the confidence scale matches reality got worse (deviation from 1.5 to 2.6 points), and the 80+ band's agreement dropped from 94.0% to 93.3%.

Measuring and stopping is also a result. With the signals we have, we cannot close this gap.

The right behaviour: say what you do not know

The engine sees this uncertainty. Tracks with a confidence score below 80 are flagged; in tech house 44.3% of tracks are flagged, in organic house 14.6% (deep house 22.5%, house 33.3%).

We also measured that the flag works. When the engine says 80 or above, agreement on tech house is 90.2%; between 60 and 79 it is 65.9%; under 60 it is 42.1%. So the badge tells the truth even inside one genre. It is not perfectly equal across genres, though: in deep house the 80+ band agrees 95.9%. The same badge carries a little less certainty in tech house.

On a flagged track the app now also says why. If fewer than half the sections land on the chosen key, for example: "Why: sections of this track disagree — only 2 of 8 point to this key. Listen to the part you will actually play." It also tells you when the track is consistent but two keys are close.

What it means at the decks

If you want to see what gets flagged in your own library, drop your tracks into Harmoniq; the analysis runs in your browser and nothing is uploaded. You know your own tracks better than any measurement will.

© 2026 Harmoniq DJ Terms Privacy Refunds Contact