← 제작기

100점짜리 오답

클래식 추천 앱은 곡만 고르지 않고 어느 연주의 음반인지까지 고릅니다. 그 음반 정보를 매번 모델에게 쓰게 하고 있었는데, 문제가 두 가지였습니다. 작곡 배경과 연주자 해설은 누가 언제 물어도 같은 내용인데 매번 새로 만들고 있었고, 더 나쁘게는 존재하지 않는 음반이 섞여 나왔습니다. 사용자가 검색해도 나오지 않는 음반을 추천하는 셈이었습니다.

검증된 카탈로그를 미리 만들어 앱에 넣기로 하고, 반드시 지킬 규칙 하나를 세웠습니다.

규칙: 연주자와 레이블은 모델이 절대 만들지 않습니다. MusicBrainz 에서만 가져옵니다.

모델은 작품을 제안하고, MusicBrainz 는 그것이 실제로 있는 녹음인지 확인합니다. 제안과 검증을 나누면 되겠다고 생각했습니다. 그 검증이 얼마나 여러 방향으로 틀릴 수 있는지는 나중에야 알았습니다.

점수 100 이 정답은 아니었다

MusicBrainz 검색은 후보마다 점수를 줍니다. 처음에는 그 점수를 믿었습니다. 감사 단계를 만들어 매칭된 제목을 우리 작품명과 직접 대조해 보고 나서야 무슨 일이 벌어지고 있었는지 알았습니다.

우리가 찾은 곡실제로 매칭된 것점수
하이든 교향곡 104번드보르자크 9번100
프로코피예프 1번베토벤 1번100
야나체크 크로이처베토벤 Op. 130100
사계 '봄' RV 269RV 315 '여름'100
호두까기인형 Op. 71백조의 호수 Op. 2096

점수가 낮아서 놓친 경우는 없었고, 모두 만점을 받고도 틀렸습니다. 그 점수가 같은 곡인지를 재지 않고 제목이 얼마나 비슷한지를 재기 때문입니다. "교향곡 1번"이라는 제목은 프로코피예프에게도 베토벤에게도 있습니다.

원칙: 검색 점수는 유사도이지 정답률이 아닙니다. 무엇을 재는 점수인지 확인하지 않고 임계값을 정하면, 그 임계값은 아무것도 막지 못합니다.

제목과 작품 번호와 곡 번호를 직접 대조하는 감사를 붙이자 뚜렷한 오매칭 37건이 나왔습니다.

검사를 어디에 걸 것인가

감사를 만들었으니 해결됐다고 생각했는데, 다시 돌려 보니 고쳐 둔 오매칭이 되살아나 있었습니다.

이유는 단순했습니다. 검사를 감사에만 걸어 두었는데, 매칭을 다시 하는 단계는 그 검사를 모르니 같은 후보를 다시 골라 완료 상태로 되돌려 놓았습니다. 감사는 나중에 표시만 할 뿐이고, 앞 단계가 계속 같은 결과를 다시 만들어 냅니다. 한 세션에서 이 일을 두 번 겪었습니다.

검사 내용은 맞았고, 거는 자리가 틀렸습니다. 판정을 후보를 거르는 필터로 끌어와서, 매칭 단계와 감사가 같은 함수를 쓰게 했습니다. 두 곳에 따로 두면 언젠가 어긋나고, 어긋나면 감사가 통과시킨 것을 매칭 단계가 저장합니다.

원칙: 뒤에서 거르는 검사는 앞 단계가 다시 만들면 소용이 없습니다. 판정은 만드는 자리에 겁니다.

이렇게 옮기는 것만으로 뚜렷한 오매칭이 37건에서 0건이 됐습니다.

검증기는 반대 방향으로도 틀린다

매칭에 실패해 보류로 빠진 곡이 67건 있었습니다. 손으로 확인하려고 열어 봤더니 대부분이 매칭 실패보다는 대조기의 결함으로 드러났습니다. 맞는 녹음이 검색 결과 안에 있었는데, 제가 만든 대조기가 떨어뜨리고 있었습니다.

  • 악센트를 없애지 않고 비교했습니다. Eternite 와 Eternité, Czardas 와 Czárdás 가 다른 글자로 남아서 같은 악장이 어긋나 보였습니다.
  • 괄호 속 별칭을 몰랐습니다. "Dance of the Knights (Montagues and Capulets)"는 같은 악장의 다른 이름입니다.
  • 우리 쪽이 더 자세히 적은 경우를 불일치로 봤습니다. 시드의 I. Allegro con brio 와 MusicBrainz 의 ...op. 25: 1. Allegro 는 겹치는 부분이 3분의 1 이라 탈락했는데, 실제로는 다른 악장이 아니라 상대가 줄여 적은 것이었습니다.

여기서 조심할 것이 있었습니다. 기준을 느슨하게 풀면 반대쪽이 뚫리는데, 원래 있던 약점이 바로 그랬습니다. 한 낱말짜리 I. Allegro 가 Symphony no. 9: IV. Allegro 와 겹침 1.0 으로 통과하고 있었습니다. 빠르기말은 아무 악장에나 붙을 수 있으니, 양쪽 앞머리에 번호가 있고 그 번호가 다르면 이름이 겹쳐도 거부하게 했습니다.

고친 뒤에는 회귀 시험 25쌍으로 결과를 고정했습니다. 통과해야 할 8쌍과 거부해야 할 17쌍입니다. 거부 쪽에는 같은 작품의 다른 악장과 빠르기말만 겹치는 경우를 일부러 넣었습니다.

원칙: 잘못 통과한 것만 세면 검증기는 점점 엄격해지고, 결국 정답까지 떨어뜨립니다. 잘못 거부한 것도 함께 셉니다.

가장 나빴던 일: 추정값이 확인된 값으로 둔갑했다

이 카탈로그에는 처음부터 알고 있던 제약이 하나 있었습니다. MusicBrainz 는 녹음 연도를 주지 못하고 발매 연도만 주는데, 재발매 모음집이 잡히면 수십 년이 어긋납니다.

실제 녹음MusicBrainz차이
굴드 골드베르크 19812012+31년
므라빈스키 비창 19602016+56년
리흐테르 라흐마니노프 19591995+36년

year_source 컬럼을 두고, 사람이 확인한 연도만 앱으로 내보내기로 했습니다. 막는 장치는 설계돼 있었는데 뚫려 있었고, 원인은 한 줄이었습니다.

year_src = "seed" if row["year"] else ...

row["year"] 는 시드 값이 아니라 데이터베이스의 현재 값입니다. 첫 매칭에서 MusicBrainz 발매 연도가 써 넣어지면, 다음 매칭에서는 "값이 있다"는 이유로 그 값이 사람이 확인한 연도로 둔갑합니다. 다시 돌릴 때마다 퍼졌습니다.

세어 봤습니다. 사람이 확인한 것으로 표시된 185건 가운데 실제로 근거가 있는 것은 57건이었습니다. 나머지 128건은 가짜였고, 그동안 앱에 녹음 연도로 표시되고 있었습니다.

더 나쁜 일이 남아 있었습니다. 그 연도를 "1981년 녹음"처럼 문장으로 쓴 해설이 이미 48곡 만들어져 있어서, 값만 되돌린다고 끝나지 않고 그 48곡은 해설을 다시 만들어야 했습니다. 잘못된 값은 지우면 되지만, 그 값을 근거로 쓴 문장은 따로 찾아가 고쳐야 합니다.

되돌린 뒤 확인해 보니 근거가 있던 57건은 시드 값과 정확히 일치했고, 어긋난 것은 한 건도 없었습니다. 막는 장치 자체는 옳았고, 판단 근거가 자기 자신을 보고 있었을 뿐이었습니다.

원칙: 출처를 값으로 판단하지 않습니다. 값은 이미 오염됐을 수 있습니다. 출처는 값과 따로 떨어진 자리에 적습니다.

2 의 차이

카탈로그 수가 926 이어야 하는데 928 로 나왔습니다. 그 2 가 단서였습니다.

배포용 뷰가 해설 생성 상태만 보고 매칭 상태는 보지 않고 있었습니다. 해설까지 만든 뒤에 오매칭으로 판정돼 보류로 내려간 곡은 그 조건에 걸리지 않아서 그대로 앱에 실렸습니다. 실제로 두 곡이 그렇게 실려 있었습니다. 텔레만 환상곡 2번 자리에 1번이, 워록 카프리올 모음곡의 "Pieds-en-l'air" 자리에 "Mattachins"가 들어가 있었습니다.

이것은 보류로 내리는 일이 생기기 전까지는 드러날 수 없는 결함이었습니다. 그전까지는 보류가 대부분 해설을 만들기 전에 정해졌기 때문입니다. 새로운 상태 변화를 만들면, 그 변화를 보지 않던 코드가 드러납니다.

전수 조사

표본에서 찾은 문제들이어서, 같은 종류가 더 있는지 악장이 지정된 755건을 모두 대조했습니다. 나온 것은 세 종류였습니다.

첫째는 우리가 틀린 2건입니다. 슈만 환상소곡집 Op. 12 의 "Warum?"과 드뷔시 녹턴의 "Sirènes"에서 시드가 적어 둔 악장 번호가 실제와 달랐습니다. 사실을 확인하고 고쳤습니다. 검증기를 의심하는 동안에도 입력이 틀렸을 가능성은 따로 남아 있습니다.

둘째는 꺼져 있던 검사 하나입니다. "같은 목록인데 번호가 다르면 다른 작품"이라는 검사가, K. 331 과 K. 300i 같은 대체 번호에서 오탐이 난다는 이유로 if False 로 막혀 있었습니다. 번호에 알파벳 접미사가 붙으면 건너뛰는 조건을 넣어 되살렸더니, 드보르자크 사이프러스 B. 152 자리에 들어가 있던 현악 4중주 B. 57 이 잡혔습니다.

셋째는 작품 번호가 우연히 같은 경우입니다. 쇼스타코비치 10번이 베토벤 8번에 매칭돼 있었는데, 둘 다 Op. 93 입니다. 번호 검사를 정당하게 통과한 것이어서, "제목 맨 앞에 다른 작곡가 이름이 나오면 다른 사람의 곡"이라는 검사를 더했습니다.

이 검사는 반드시 제목 '맨 앞'만 봐야 합니다. 제목 뒤쪽에 나오는 작곡가 이름은 헌정이나 주제 차용처럼 대개 정당한 이유가 있기 때문입니다. 브람스의 "헨델 주제에 의한 변주곡", 모차르트 K. 465 "하이든", 라벨의 "쿠프랭의 무덤"이 그렇습니다. 처음에는 "앞 15자 이내"로 두었다가 라벨이 걸려서 맨 앞 한 자리로 좁혔습니다.

남은 것

지금 앱에 들어가는 카탈로그는 음반 925종이고, 듣는 사람 기준으로는 691곡입니다. 한 작품의 여러 악장이 각각 한 행이어서 두 수가 다릅니다. 연도가 표시되는 것은 57종뿐입니다. 나머지 868종 가운데 849종은 MusicBrainz 발매 연도를 갖고 있지만 내보내지 않습니다. 틀린 연도를 보여 주느니 비워 두는 편을 택했습니다.

수가 늘기만 한 것도 아닙니다. 악센트를 없애지 않고 중복을 판정한 탓에 확장할 때 87행이 겹쳐 들어와 있었고(Erlkonig 와 Erlkönig, Traumerei 와 Träumerei), 이를 지우면서 1,051 에서 964 로 줄었습니다. 전수 조사에서도 한 곡이 더 빠졌습니다.

맞는 수가 큰 수보다 낫습니다.

정리하면

  • 검색 점수는 유사도이지 정답률이 아닙니다. 무엇을 재는 점수인지 모르고 임계값을 정하면 아무것도 막지 못합니다.
  • 검사는 내용만큼 거는 자리가 중요합니다. 뒤에서 거르면 앞 단계가 다시 만들어 되돌립니다. 만드는 자리와 검사하는 자리가 같은 함수를 봐야 합니다.
  • 검증기는 양쪽 방향으로 틀립니다. 잘못 통과한 것만 세면 점점 엄격해져 정답까지 떨어뜨립니다. 거부해야 할 쌍과 통과해야 할 쌍을 함께 회귀 시험으로 고정합니다.
  • 출처는 값과 따로 떨어진 자리에 적습니다. 자기 값을 근거로 출처를 판정하면, 다시 돌릴 때마다 오염된 값이 확인된 값으로 둔갑합니다.
  • 잘못된 값을 지웠다고 끝나지 않습니다. 그 값을 인용해 이미 만들어 둔 문장을 따로 찾아가야 합니다.
  • 두 숫자가 어긋나면 그 차이가 단서입니다. 926 과 928 사이의 2 가 결함 하나를 통째로 찾아 줬습니다.

같은 앱을 두 플랫폼으로 만든 이야기는 같은 앱을 iOS와 안드로이드로 두 번 만들면서 배운 것에, 재는 도구를 먼저 의심해야 했던 다른 사례는 흔들린 건 곡선이 아니라 재는 자였다에 있습니다. 없는 것을 지어내느니 비워 두는 편이 낫다는 같은 결의 판단은 지어낸 답보다 빈칸이 낫다에 적어 두었습니다. 이 카탈로그가 정상인데도 답이 틀린 뒷이야기는 쇼팽을 고르고 슈베르트를 설명했다에 있습니다.

A classical recommendation app does not only pick a piece — it picks which recording. And it was having the model write that too. Two problems. A composer's background and a performer's approach are the same for whoever asks, whenever, yet were being written fresh each time; and worse, albums that do not exist were coming out mixed in with the real ones. Recommending a record that returns nothing when the user searches for it.

So the plan was to build a vetted catalogue in advance and ship it inside the app, with one rule.

The rule: the model never invents performers or labels. Those come from MusicBrainz only.

The model proposes works; MusicBrainz confirms they are real recordings. Separate the proposing from the verifying and you are safe — that was the idea. How many directions the verifying could be wrong in, I learned later.

A perfect score on the wrong piece

MusicBrainz search returns a score per candidate, and at first I trusted it. Only after building an audit step that compares the matched title against our own work title did I see what had been happening.

What we searched forWhat it actually matchedScore
Haydn Symphony no. 104Dvořák no. 9100
Prokofiev no. 1Beethoven no. 1100
Janáček KreutzerBeethoven op. 130100
The Four Seasons, "Spring" RV 269RV 315, "Summer"100
The Nutcracker, op. 71Swan Lake, op. 2096

These were not missed because the score was low. They scored full marks and were wrong. That number answers "how similar is the title", not "is this the same piece". A title like Symphony No. 1 belongs to Prokofiev and to Beethoven alike.

Principle: a search score is similarity, not correctness. Set a threshold without checking what the number measures and the threshold blocks nothing.

With an audit comparing titles, opus numbers and ordinals directly, 37 outright mismatches surfaced.

Where you attach the check

Having built the audit, I thought that was that — until a later run showed the mismatches I had fixed were back.

The reason was simple. The check lived in the audit, and the step that redoes matching does not know about the audit. So it picked the same candidate again and marked it done. The audit only labels things afterwards while the earlier stage keeps re-creating them. I hit this twice in one session.

The check itself was right. Where I had attached it was wrong. I pulled the judgement into the candidate filter so that the resolver and the audit call the same function. Keep two copies and they drift, and when they drift the resolver stores what the audit would have rejected.

Principle: a check that filters at the end is void if the front keeps producing. Attach the judgement where the thing is made.

That move alone took outright mismatches from 37 to zero.

A verifier is wrong in both directions

Sixty-seven items had fallen into the review pile as failed matches. Opening them to check by hand, most turned out not to be failed matches at all but defects in my comparator. The right recording had been in the search results, and my own code had dropped it.

  • Accents were not folded. Eternite against Eternité, Czardas against Czárdás — the same movement looked different.
  • Parenthetical aliases were unknown to it. "Dance of the Knights (Montagues and Capulets)" is another name for the same movement.
  • Being more specific than the other side read as a mismatch. Our I. Allegro con brio against MusicBrainz's ...op. 25: 1. Allegro overlapped by a third and was dropped — but that is not a different movement, it is the other side writing it shorter.

Loosening it has its own hazard: the opposite direction opens up. There was already a weakness of exactly that kind — a one-word I. Allegro matched Symphony no. 9: IV. Allegro at full overlap, because a tempo marking attaches to any movement at all. So now, when both sides carry a leading number and those numbers disagree, the names no longer matter.

After fixing it I pinned the behaviour with 25 regression pairs: 8 that must pass and 17 that must be rejected. The rejection side deliberately includes other movements of the same work, and pairs sharing nothing but a tempo marking.

Principle: count only false accepts and a verifier grows steadily stricter until it rejects correct answers too. Count the false rejects with them.

The worst of it — provenance promoting itself

This catalogue had a known constraint from the start. MusicBrainz cannot give you a recording year. It gives release years. Land on a reissue compilation and you are off by decades.

Actual recordingMusicBrainzOff by
Gould, Goldberg, 19812012+31 years
Mravinsky, Pathétique, 19602016+56 years
Richter, Rachmaninov, 19591995+36 years

So there is a year_source column, and only curated years are exported to the app. The defence was designed in. The defence had a hole in it, and the hole was one line.

year_src = "seed" if row["year"] else ...

row["year"] is not the seed value — it is whatever is currently in the database. Once a first pass wrote a MusicBrainz release year there, the next pass saw a value present and promoted it to a curated year. Every re-run spread it further.

I counted. Of 185 rows marked as curated, 57 actually had a basis. The other 128 were fabricated, and had been displayed in the app as recording years.

There was worse underneath. Blurbs that narrated those years in prose — "recorded in 1981" — had already been generated for 48 pieces. Reverting the values did not finish the job; those 48 had to be regenerated. A wrong value can be deleted, but sentences written on top of it have to be hunted down separately.

After the revert, the 57 with a real basis matched their seed values exactly — zero discrepancies. The defence had been right all along; only its evidence was looking at itself.

Principle: never infer provenance from the value. The value may already be contaminated. Record provenance somewhere separate from the thing it vouches for.

A difference of two

The catalogue count should have been 926 and came out 928. Those two were the clue.

The export view was reading generation status but not resolution status. Anything judged a mismatch after it had already been generated, and demoted to review, failed to meet that condition and shipped anyway. Two pieces had done exactly that — a Telemann Fantasia no. 1 sitting where no. 2 belonged, and "Mattachins" standing in for "Pieds-en-l'air" in Warlock's Capriol Suite.

This defect could not surface until demotions started happening: until then, review was almost always decided before generation. Introduce a new state transition and you expose the code that never watched for it.

Checking all of them

Since these had come from samples, I compared all 755 works that specify a movement. Three kinds of thing came out.

Two where we were wrong. The movement numbers our seed recorded for "Warum?" in Schumann's Fantasiestücke op. 12 and "Sirènes" in Debussy's Nocturnes did not match reality. Checked and corrected. While you are busy suspecting the verifier, the input can still be the thing that is wrong.

One check that had been switched off. "Same catalogue, different number, therefore a different work" sat behind an if False, because alternate numberings like K. 331 and K. 300i produced false positives. A guard that skips numbers carrying a letter suffix brought it back — and it immediately caught a string quartet, B. 57, filed where Dvořák's Cypresses, B. 152, should have been.

Opus numbers that coincide. Shostakovich's Tenth was matched to Beethoven's Eighth. Both are op. 93, so it passed the number check legitimately. That called for another check: a different composer's name at the front of the title means a different composer's piece.

It has to be at the front. A composer's name later in a title is usually legitimate — Brahms's "Variations on a Theme by Handel", Mozart's K. 465 "Haydn", Ravel's "Le Tombeau de Couperin". I first allowed "within the first 15 characters", caught Ravel with it, and narrowed the window to position zero.

Where it landed

The catalogue now shipping in the app holds 925 recordings, which come to 691 distinct pieces — a work's movements each sit on a row of their own, so the two counts differ. A year is displayed for 57 of them; of the other 868, 849 do have a MusicBrainz release year, and it is not exported — better blank than wrong.

The count did not only grow, either. Because the duplicate check also failed to fold accents, expansion had pulled in 87 overlapping entries (Erlkonig beside Erlkönig, Traumerei beside Träumerei); removing them took 1,051 down to 964. The full survey dropped one more.

A correct number beats a big one.

In short

  • A search score is similarity, not correctness. A threshold set without knowing what the number measures blocks nothing.
  • Where a check is attached matters as much as what it checks. Filter at the end and the front will re-create what you removed; the making and the checking must consult the same function.
  • A verifier is wrong in both directions. Count only false accepts and it tightens until it drops correct answers. Pin both must-pass and must-reject pairs in regression.
  • Record provenance apart from the value. Judge provenance from the value itself and every re-run promotes its own contamination.
  • Deleting a wrong value is not the end of it. Prose already written on top of that value has to be chased down separately.
  • When two counts disagree, the difference is the clue. The two between 926 and 928 brought a whole defect with it.

Building the same app on two platforms is in What I Learned Building the Same App Twice, and another case of having to suspect the instrument first is in It Was the Instrument Shaking, Not the Curve. The same instinct — that having nothing beats inventing something — is in Empty-Handed Beats Invented. What went wrong later, with this catalogue intact, is in It Picked Chopin and Described Schubert.