← 제작기

쇼팽을 고르고 슈베르트를 설명했다

클래식 추천 앱에서 실제 기기 사례가 하나 올라왔습니다. 쇼팽의 '혁명' 연습곡을 골라 놓고 "이 작품은 슈베르트의 피아노 소나타 제23번입니다"라고 적어 놓은 것입니다. 온디바이스 엔진에서만 그랬고, Groq 와 Claude 는 멀쩡했습니다.

고치는 데 네 번이 걸렸습니다. 그중 세 번은 제가 넣은 안전장치가 스스로를 무력하게 만들거나, 너무 잘 들어서 다른 것을 망가뜨렸습니다.

어디가 원인이 아닌지부터 확인했다

먼저 데이터를 의심했는데 카탈로그는 정상이었습니다. 그 곡의 네 연주 모두 쇼팽에 대한 올바른 해설을 갖고 있었습니다. 이 카탈로그를 만들 때 걸러 냈던 문제와는 종류가 달랐습니다. 인덱스 매핑도 정상이어서, 프롬프트에 넣는 후보 배열과 답을 조립할 때 참조하는 배열은 같은 배열이었습니다.

화면에 나오는 작품 배경과 '이 음반인 이유'는 카탈로그에서 그대로 읽어 옵니다. 모델이 직접 쓰는 것은 '지금 이 곡인 이유' 한 줄뿐입니다. 그 한 줄에서 모델은 인덱스로는 쇼팽을 가리키고, 글로는 슈베르트를 설명했습니다.

1차: 후보 목록을 판정 기준으로 삼았다

판정에 쓸 어휘 목록을 따로 만들고 싶지는 않았습니다. 그런 목록은 카탈로그가 커지면 함께 낡기 때문입니다. 대신 후보 목록 자체를 기준으로 썼습니다. 목록에 있는 다른 작곡가 이름이 글에 나오는데 고른 곡의 작곡가는 나오지 않으면 버립니다. "브람스보다 직선적인 쇼팽"처럼 비교하느라 다른 작곡가를 언급하는 정상적인 글은 고른 작곡가도 함께 나오므로 살아남습니다.

같은 커밋에서 한 가지를 더 했습니다. 온디바이스 엔진의 후보를 40개에서 12개로 줄였습니다. 목록이 길면 인덱스와 글이 어긋나는 것으로 보였기 때문입니다.

실제 기기에서는 걸리지 않았습니다.

판정 기준이 후보 목록이었는데, 그 후보 목록을 같은 커밋에서 좁혔습니다. 슈베르트가 12개 안에 없으면 판정기는 그 이름을 모릅니다. 증상을 고치면서 그 증상을 보지 못하게 만든 셈입니다. 두 변경이 따로 보면 각각 말이 되는데, 함께 들어가니 하나가 다른 하나를 무력하게 만들었습니다.

2차: 판정 기준을 카탈로그 전체로 넓혔다

판정에 쓰는 이름을 후보 목록에서 떼어 내고, 카탈로그 전체에 있는 작곡가의 성으로 바꿨습니다. 지금 카탈로그로 세면 한글 기준 171개입니다. 후보를 몇 개로 줄이든 이 목록은 줄지 않습니다.

같은 자리에서 다른 문제도 하나 잡았습니다. 야나체크 현악 사중주에서 같은 문장을 계속 되풀이하는 답이 나와서, 문장 단위로 잘라 절반 이상이 같으면 버리도록 했습니다.

이번에도 실제 기기에서는 걸리지 않았습니다. 판정 로직만 놓고 짠 시험은 두 번 모두 통과했지만, 그 시험이 본 것은 "이런 글이 들어오면 버리는가"였지 "실제 기기에서 그런 글이 정말 줄어드는가"가 아니었습니다.

3차: 걸러 내는 대신 말할 자리를 없앴다

걸러 내는 방식을 두 번 시도해 두 번 다 실패했으니, 세 번째도 같은 방식일 이유가 없었습니다. 방향을 바꿨습니다. 이 모델이 쓰는 것은 '지금 이 곡인 이유' 한 줄이고, 그 줄은 듣는 사람의 상태에 관한 글이면 충분합니다. 작품에 관한 사실을 말할 필요가 처음부터 없었습니다.

온디바이스 엔진에만 적용하는 규칙을 넣어, 두세 문장 안에서 듣는 사람과 소리의 인상만 쓰게 했습니다. 작곡가, 작품명, 악장, 연도, 작품 번호, 나라, 역사적 사건은 쓰지 말라고 했습니다. 분량도 네댓 문장에서 줄였습니다. 길수록 헤맬 자리가 늘어나기 때문입니다.

금지해도 새어 나오는 경우에 대비해 잡아내는 검사를 하나 더 걸었습니다. 연도와 작품 번호는 형태가 뚜렷해서 문자열로 잡을 수 있습니다.

틀린 설명은 멈췄지만, 다섯 곡의 글이 똑같아졌다

한 번에 다섯 곡을 받으면 '지금 이 곡인 이유'가 다섯 개 모두 같았습니다.

원인은 제가 막은 방식에 있었습니다. 작품에 관한 모든 것을 금지하니 남은 재료가 듣는 사람의 상태뿐이었는데, 그 상태는 다섯 곡에 공통입니다. 쓸 재료를 빼앗아 놓고 다르게 쓰라고 시킨 셈이었습니다.

되돌리면서 재료를 다시 열되 출처를 가렸습니다. 후보 목록에는 에너지, 분위기 태그, 시대, 편성, 연주자처럼 곡마다 다른 정보가 이미 들어 있습니다. 이 값들은 모델의 기억이 아니라 카탈로그가 준 값이어서 틀릴 수가 없습니다. 그 값들을 소리에 대한 묘사로 바꿔 쓰되, 목록에 없는 것은 더하지 말라고 했습니다. 연도와 작품 번호는 여전히 금지입니다.

그래도 겹치면 버리는 검사를 하나 더 두었습니다. 낱말이 70% 이상 겹치면 같은 글로 봅니다.

버리는 검사를 넣으면 재시도 조건도 함께 봐야 한다

여기서 하나를 놓칠 뻔했습니다. 다섯 곡을 요청했는데 다섯 개가 같은 글로 오면, 중복을 버린 뒤 한 곡만 남습니다. 재시도 조건이 "비어 있지 않으면 통과"였으므로 한 곡짜리 답이 그대로 화면에 나가는데, 버리는 검사를 넣은 순간부터 그 조건은 더 이상 맞지 않게 된 것입니다.

조건을 "요청한 개수를 채우면 통과"로 바꿨습니다. 두 번까지 다시 물어보고, 더 많이 건진 쪽을 씁니다.

아직 모르는 것

다섯 곡의 글이 서로 달라진 것은 확인했지만, 요청한 개수를 늘 채우지는 못합니다.

후보가 모자라서는 아닙니다. 카탈로그에서 직접 재 보니 후보 120행에서 서로 다른 곡이 70개에서 107개 나옵니다. 모자라는 곡은 중복 판정이 버렸거나 재추천 간격 때문에 빠졌을 텐데, 둘 가운데 어느 쪽인지는 로그를 봐야 가려지고 아직 보지 않았습니다. 그럴듯한 쪽으로 덮지 않고 원인 미확정으로 적어 둡니다.

2026-09-21 — 원인이 가려졌습니다. 중복 판정이 버린 것이 원인이었고, 의도한 동작이라 고치지 않았습니다. 같은 판에서 다른 문제가 하나 더 나왔습니다. 분위기 구절이 다른 카드와 겹치고 있었는데, 추천 이유는 중복 검사가 보고 있었지만 분위기 구절은 아무도 보지 않았습니다. 위 문단은 지우지 않고 둡니다. 모른다고 적어 둔 것이 나중에 가려지는 편이, 그럴듯한 쪽으로 덮어 둔 것보다 낫습니다.

iOS 판과 안드로이드 판에 같은 검사가 같은 이름으로 들어가 있습니다. 한쪽에서만 고치면 같은 앱이라고 말할 수 없습니다.

남은 것

  • 판정 기준을 판정 대상에서 빌려 오면, 대상이 줄 때 기준도 함께 줄어듭니다. 후보 목록을 판정 어휘로 쓴 것은 낡지 않는 설계였는데, 같은 커밋에서 그 목록을 좁히자 판정기가 아무것도 보지 못하게 됐습니다. 지금은 카탈로그 전체에서 가져옵니다.
  • 틀린 답을 걸러 내는 것과 틀릴 자리를 없애는 것은 다릅니다. 걸러 내기로 두 번 실패한 뒤에야 방향을 바꿨는데, 더 일찍 바꿨어야 했습니다.
  • 재료를 빼앗고 다양성을 요구할 수는 없습니다. 금지가 잘 들으면 남는 것은 모두에게 공통인 한 가지뿐입니다. 틀릴 수 없는 재료가 어디에 있는지부터 찾는 편이 낫습니다.
  • 버리는 검사를 넣으면 "비어 있지 않으면 통과"는 틀린 조건이 됩니다. 걸러 내기와 재시도 조건은 함께 봐야 합니다.
  • 판정 로직만 보는 시험은 판정 로직만 보증합니다. 세 번의 수정이 모두 시험을 통과하고도 실제 기기에서 실패했습니다. 같은 스택이 세 번 되돌아온 적도 있습니다.

A report came back from a real device using the classical recommendation app. It had picked Chopin's "Revolutionary" étude and written "this work is Schubert's Piano Sonata No. 23" underneath it. Only the on-device engine did this; Groq and Claude were fine.

It took four attempts to fix. Three times the guard I added either disabled itself or worked so well that it broke something else.

Ruling out the places it wasn't

The data was the first suspect, and the catalogue turned out to be fine: all four recordings of that piece carried correct notes about Chopin. What the audit caught while building this catalogue was not this kind of error. The index mapping was fine too. The candidate array written into the prompt and the array read when assembling the answer are the same array.

The work background and the "why this recording" text on screen are read straight from the catalogue. The only line the model writes itself is "why this piece right now." In that one line it pointed at Chopin by index and described Schubert in prose.

First attempt — the candidate list as the test

I did not want a separate vocabulary list for detection, because a list like that goes stale as the catalogue grows. The candidate list itself became the test: if the text names another composer from the list but never names the composer of the chosen piece, throw it away. Ordinary comparative writing survives, since "more direct than Brahms, this Chopin" names the chosen composer too.

The same commit did one more thing. It cut the on-device candidate count from 40 to 12, because the mismatch looked like something that happened over long lists.

It didn't catch anything on the device.

The test's vocabulary was the candidate list, and the same commit narrowed that list. If Schubert isn't among the twelve, the detector has never heard of him. I had fixed the symptom and blinded myself to it in one move. Each change made sense on its own; together, one cancelled the other.

Second attempt — widening the test to the whole catalogue

The names used for detection came off the candidate list and became every composer surname in the catalogue, which measures 171 in Korean against the current catalogue. However far the candidate list is trimmed, this one does not shrink with it.

The same pass caught something else. An answer about a Janáček quartet kept repeating the same sentence, so a check now splits the text into sentences and rejects it when half or more are identical.

This didn't catch anything on the device either. Both fixes passed the tests written against the judging logic, but those tests asked "does this text get rejected," not "does such text actually become rarer on a device."

Third attempt — removing the place to be wrong

Two rounds of filtering had failed, so there was no reason for the third to be more filtering. What this model writes is one line about why this piece fits right now, and a line about the listener's state is enough for that. There was never a reason for it to assert facts about the work.

An on-device-only rule went in: two or three sentences, the listener and the impression of the sound, nothing else. Composer, work title, movement, year, opus number, country and historical events are not to be written. The length came down from four or five sentences, since more room is more room to wander.

One more check catches what leaks through the ban anyway. Years and opus numbers have a distinct shape, so a string match finds them.

The lying stopped, and five pieces came back with one text

Asking for five recommendations at once produced five identical explanations.

The cause was how I had blocked it. Forbidding everything about the work left the listener's state as the only material, and that state is common to all five pieces. I had taken the material away and then asked for variety.

Rolling back meant reopening the material while sorting it by origin. The candidate list already holds something different for every piece: energy, mood tags, era, instrumentation, performer. These come from the catalogue rather than the model's memory, so they cannot be wrong. The rule now says to turn those values into a description of the sound and to add nothing that isn't on the list. Years and opus numbers stay banned.

A further check throws away what still overlaps: 70% shared words counts as the same text.

Adding a reject step means revisiting the retry condition

This part nearly slipped past me. When five pieces come back with the same text, discarding the duplicates leaves one. The retry condition was "pass if not empty," so a one-item answer went to the screen unchallenged. The moment a reject step exists, that condition is a lie.

It now reads "pass if it fills the requested count," asks up to twice, and keeps whichever attempt salvaged more.

What I still don't know

The five explanations do differ now. The requested count, though, is not always filled.

It isn't candidate scarcity. Measured against the catalogue directly, 120 candidate rows yield somewhere between 70 and 107 distinct pieces. Whether the shortfall comes from the similarity check discarding answers or from the re-recommendation interval is something only the logs can separate, and I haven't read them. Rather than cover it with the more plausible of the two, it goes down as cause undetermined.

2026-09-21 — separated. The similarity check discarding answers was the cause, and since that is the intended behaviour nothing was changed. The same release turned up one more: mood phrases were repeating across sibling cards, because the reasoning had a duplicate check watching it and the mood had none. The paragraph above stays as written — an unknown that later gets resolved beats one papered over with the likelier answer.

Both editions, iOS and Android, carry the same checks under the same names. Fixing one side only would make "the same app" untrue.

What's left

  • A test that borrows its vocabulary from the subject shrinks when the subject does. Using the candidate list for detection was a design that would never go stale, and narrowing that list in the same commit shut the detector's eyes. It now reads from the whole catalogue.
  • Filtering wrong answers and removing the chance to be wrong are different moves. It took two failed rounds of filtering before I switched. That should have come sooner.
  • You cannot take the material away and still ask for variety. When a ban works, what remains can be the single thing everyone shares. Better to find first which material cannot be wrong.
  • Once a reject step exists, "pass if not empty" is a lie. Filtering and the retry condition have to be read together.
  • A test aimed at the judging logic guarantees the judging logic. Three fixes passed their tests and failed on the device. One stack came back three times, too.