← 제작기

앱이 제 실수를 내 탓으로 채점했다

말하기 시험 앱을 iOS 로 옮기면서, 안드로이드판이 지키던 제약 하나를 더 잘 지킬 기회가 생겼습니다. 키 없이도 돌아야 한다는 제약입니다. 안드로이드에서는 무료 API 키 하나로 전 과정이 끝나야 한다는 뜻이었는데, iOS 에는 온디바이스 모델과 시스템 음성 인식기가 있으니 키가 아예 없어도 되겠다고 생각했습니다.

붙여 보니 됐습니다. 실제 기기에서 키 없이 5문항을 끝까지 돌렸습니다. 질문 생성, 음성 재생, 자동 녹음, 받아 적기, 채점, 저장까지 네트워크 요청이 한 건도 없었으니, 여기서 멈췄다면 이 글의 제목은 "키가 필요 없어졌다"였을 것입니다.

받아 적은 글을 읽어 봤다

끝까지 돌았다는 것은 연결에 빠진 데가 없다는 뜻일 뿐, 결과를 쓸 만하다는 뜻까지 담고 있지는 않습니다. 저장된 받아쓰기 결과를 열어 봤습니다.

  • 제가 사는 도시 이름 → "Swan"
  • trip(여행) → "tree"
  • photo(사진) → "potter"
  • 쇼핑몰 이름으로 보이는 것 → "stock page one"
  • "인테리어가 우드톤"으로 보이는 것 → "story teller is very woody"

틀린 곳은 고유명사와 생활 어휘에 몰려 있었습니다. 이것만으로는 온디바이스 인식기가 나쁜 것인지 제 발음이 나쁜 것인지 알 수 없습니다. 비원어민이 말한 29초짜리 자기소개였기 때문입니다.

조건을 맞춰 비교했다

가려낼 방법은 같은 녹음 파일을 두 인식기에 넣어 보는 것 하나뿐입니다. 다른 회차의 다른 답변을 비교하면 발음 차이가 섞여서 아무 말도 할 수 없게 됩니다.

실제로 말한 것온디바이스Groq Whisper
제 영문 이름전혀 다른 이름 ✗정확 ✓
Legend of Zelda"Legend of Jareda" ✗정확 ✓
I live in ○○ in Korea"I'm really through one in Korea" ✗도시명까지 정확 ✓
friends"friend" ✗정확 ✓
always live"already" ✗정확 ✓

단어 일치도는 76% 였고, 둘이 갈린 다섯 곳에서 모두 Whisper 가 맞았습니다. 한 곳이라도 반대로 갈렸다면 "서로 다르다"에 그쳤겠지만, 다섯 대 영이면 방향이 있는 것입니다.

보기 싫은 정도로 끝나는 문제가 아니었다

세 번째 줄을 다시 보겠습니다. "I live in ○○ in Korea"가 "I'm really through one in Korea"가 됐습니다. 질문은 어디 사는지를 물었고, 저는 답했는데, 인식기가 그 답을 지웠습니다.

앞선 회차에서 같은 문항이 질문 일치율 0% 로 가장 낮은 등급을 받았습니다. 채점기는 받아 적은 글만 보므로, 지워진 답을 "답하지 않았다"고 읽은 것입니다.

정리하면, 앱이 자기의 인식 오류를 사용자의 잘못으로 채점했습니다.

이 글의 핵심은 온디바이스 인식의 대가가 정확도 몇 퍼센트보다 책임이 누구에게 가는가에 있었다는 것입니다. 사용자는 자기가 무엇을 말했는지 압니다. 앱이 "당신은 어디 사는지 말하지 않았다"고 하면, 사용자는 앱보다 자기 영어를 의심합니다.

이것을 증명할 수 있었던 것은 녹음을 지우지 않았기 때문입니다. 원본 저장소가 채점 뒤에도 답변 오디오를 남겨 두기로 한 결정이 여기서 제 몫을 했습니다. 오디오가 없었다면 "받아쓰기가 이상하다"는 짐작만 남고, 실제로 무엇을 말했는지는 끝내 밝히지 못했을 것입니다.

채점기도 재 봤다

인식이 문제라면 채점은 괜찮을 수도 있습니다. 답변 5개를 같은 받아쓰기 결과로 두 채점기에 넣었습니다. 같은 녹음 대신 같은 받아쓰기 결과를 넣은 것은, 녹음을 넣으면 인식기 차이가 섞여서 채점이 다른지를 가려낼 수 없기 때문입니다.

각각 3회씩 돌렸습니다. 한 번씩만 비교했다면 차이가 나와도 엔진 차이인지 그날그날의 변덕인지 몰랐을 것입니다.

Groq온디바이스
중앙 등급NHIM2
자체 흔들림(3회 폭)1.0단계1.4단계

엔진 사이의 차이는 2.8단계였고, 각자의 흔들림은 1.0 단계와 1.4 단계였습니다. 차이가 흔들림의 두 배가 넘으니 우연으로 보기 어렵습니다. 3회씩 매긴 것은 이 문장 하나를 쓸 수 있게 하려는 것이었습니다.

어느 쪽이 맞는지 정답은 없지만, 온디바이스 쪽에는 스스로 모순되는 결과가 있었습니다.

  • 어떤 답은 질문 일치율이 100% 인데 발화 구조 점수가 17 이었습니다. "완전히 답했다"와 "단어를 늘어놓은 수준이다"가 동시에 참일 수는 없습니다.
  • 같은 답변을 세 번 채점한 결과가 IM1, IM2, AL 로 갈려 4단계나 벌어졌습니다. 문법이 깨진 76단어짜리 답변에 최상위권 등급이 나온 것입니다. 같은 답에 Groq 는 세 번 모두 같은 등급을 줬습니다.

더 중요한 것은 두 오류가 겹친다는 점이었습니다. 이 받아쓰기 결과들은 이미 온디바이스 인식기가 망가뜨린 것이었는데, Groq 는 그 손상을 감점했고 온디바이스 채점기는 알아채지 못했습니다. 그 결과는 이렇습니다.

  • 인식과 채점을 모두 온디바이스로 하면 점수는 그럴듯한데 이유가 틀립니다.
  • 온디바이스로 인식하고 Groq 로 채점하면 앱의 오류를 사용자가 뒤집어씁니다.

어느 조합이든 먼저 고칠 것은 인식이었습니다.

원칙을 버렸다

온디바이스 경로를 모두 걷어냈습니다. 지금 이 앱은 형제 앱과 구성이 같습니다. 키가 필요하고, 받아 적기는 언제나 Whisper 가 합니다.

버린 것이 기능보다 무거운 원칙이었다는 점에서 이 결정은 가볍지 않았습니다. "키 없이 돈다"는 편의 기능이 아니라 설계 제약이었고, 저는 그것을 지키려고 코드를 썼습니다. 그 제약을 지킨 결과가 사용자에게 틀린 성적을 주는 앱이라면 지킬 이유가 없습니다.

코드는 지우지 않고 따로 보관했습니다. 되살리려면 대가를 치러야 합니다. 온디바이스 기능이 요구하던 두 프레임워크가 빠지자 배포 대상을 iOS 26 에서 17 로 낮출 수 있었고, 덕분에 지원하는 기기가 훨씬 넓어졌습니다. 되돌리려면 이것을 다시 내줘야 합니다. 알고 한 맞교환입니다.

뜻밖의 이득이 이쪽에서 왔다는 점이 흥미롭습니다. 최신 기능을 쓰려고 최신 운영체제를 요구하고 있었는데, 그 기능이 쓸 만하지 않다는 것을 재고 나니 요구할 이유도 함께 사라졌습니다.

이 결론의 한계

표본은 답변 5개, 한 세션, 한 사람입니다. 그 한 사람이 비원어민이어서 인식기에는 어려운 조건이기도 했습니다. 방향은 뚜렷하지만 "온디바이스가 나쁘다"로 일반화하지는 않고, 이 화자의 이 용도에는 대신 쓸 수 없었다는 것까지만 제가 재서 아는 사실로 둡니다.

해 보지 않은 것도 남겨 둡니다. 설정에 영문 이름을 넣어 두면 인식기에 힌트로 줄 수 있어서, 위의 다섯 곳 가운데 이름 하나는 고칠 수 있습니다. 지명과 작품명은 질문에 없는 단어라 힌트로도 잡지 못합니다.


정리하면

  • 끝까지 돌려 보는 것은 연결 점검일 뿐 품질 측정이 아닙니다. "끝까지 돌았다"와 "결과를 쓸 만하다"는 다른 문장입니다.
  • 두 후보를 비교할 때는 같은 입력을 넣습니다. 다른 입력으로 비교하면 무엇이 달랐는지 말할 수 없습니다.
  • 흔들리는 대상을 판정하려면 여러 번 재서 자체 흔들림의 폭부터 구합니다. 차이가 그 폭보다 커야 차이입니다.
  • 정확도 손실이 사용자의 책임으로 바뀌는 지점이 있는지 봅니다. 그 지점을 넘으면 더는 품질만의 문제가 아닙니다.
  • 원칙도 실측 앞에서는 물러서야 합니다. 지키려던 제약이 사용자에게 손해로 돌아가면, 고칠 것은 결과보다 제약입니다.

같은 두 앱의 엔진 선택은 채점기가 소리를 듣는가에, 재는 도구를 먼저 의심해야 했던 이야기는 흔들린 건 곡선이 아니라 재는 자였다에 적어 두었습니다. 없는 것을 지어내느니 비워 두는 편이 낫다는 같은 결의 판단은 지어낸 답보다 빈칸이 낫다에 있습니다.

Porting the speaking-test app to iOS looked like a chance to honour one of the Android version's constraints even better: it has to run without a key. On Android that meant one free API key had to cover the whole flow. On iOS there is an on-device model and a system transcriber — so maybe it could need no key at all.

I wired it up, and it worked. Five questions end to end on a real device with zero keys: question generation, playback, automatic recording, transcription, scoring, storage — zero network requests. Had I stopped there, this post would be titled "It no longer needs a key."

Then I read the transcripts

Getting through the flow means the plumbing does not leak. It does not mean the output is usable. So I opened the stored transcripts.

  • the city I live in → "Swan"
  • trip → "tree"
  • photo → "potter"
  • what I take to be a shopping mall's name → "stock page one"
  • what I take to be "the interior is wood-toned" → "story teller is very woody"

Concentrated in proper nouns and everyday vocabulary. But on its own this cannot tell you whether the recogniser is weak or my pronunciation is. It is a 29-second self-introduction by a non-native speaker.

So I ran a controlled comparison

There is only one way to separate those: put the same recording through both recognisers. Comparing different answers from different sessions mixes in pronunciation differences and leaves you unable to claim anything.

What I actually saidOn-deviceGroq Whisper
my own name in Englishan entirely different name ✗correct ✓
Legend of Zelda"Legend of Jareda" ✗correct ✓
I live in ○○ in Korea"I'm really through one in Korea" ✗city name and all ✓
friends"friend" ✗correct ✓
always live"already" ✗correct ✓

76% word agreement — and all five disagreements went Whisper's way. Had even one gone the other direction, the finding would have been "they differ". Five to nothing is a direction.

This was not a cosmetic problem

Look at the third row again. "I live in ○○ in Korea" became "I'm really through one in Korea." The question asked where I live. I answered. The recogniser erased the answer.

And in an earlier session that same question scored 0% coverage and the lowest grade. The grader only ever sees the transcript, so it read the erased answer as "did not answer".

Put plainly: the app billed its own recognition error to the user as a grade.

That is the heart of it. The price of on-device recognition was not a few percent of accuracy; it was the direction of blame. A user knows what they said. When the app says "you never mentioned where you live", the thing they doubt is not the app — it is their own English.

I could prove this only because the recordings had not been deleted. The original repo's decision to keep answer audio after scoring paid for itself here. Without the audio there would be a suspicion that "the transcript looks wrong", and no way to ever establish what was actually said.

I measured the grader too

If recognition is the problem, scoring might still be fine. So five answers went to both graders as the same transcript — the same text, not the same recording, because feeding audio would mix the recognisers back in and make "do they score differently?" unanswerable.

And three times each. Compared once apiece, a difference tells you nothing about whether it is the engine or the day.

GroqOn-device
median gradeNHIM2
own spread over 3 runs1.0 step1.4 steps

The engines sat 2.8 steps apart while wobbling 1.0 and 1.4 on their own. More than double their own noise is hard to call coincidence. Grading three times each existed so that this one sentence could be written.

Which one is right has no answer key. But the on-device side contradicted itself.

  • One answer scored 100% question coverage and 17 for delivery structure. "Answered it completely" and "barely strings words together" cannot both be true.
  • The same answer over three runs came back IM1, IM2, AL — four steps apart, awarding a near-top grade to a 76-word answer with broken grammar. Groq returned the same grade all three times for that answer.

More important still: the two errors compound. These transcripts were already damaged by the on-device recogniser. Groq marked that damage down; the on-device grader never noticed. So:

  • on-device throughout gives plausible-looking grades for the wrong reasons, and
  • on-device recognition with Groq scoring makes the user carry the app's error.

Either way, recognition was the thing to fix first.

So I gave up the principle

The on-device path came out entirely. The app now matches its sibling: it needs a key, and transcription always goes through Whisper.

What made this decision weigh something is that what I gave up was not a feature but a principle. "Runs without a key" was a design constraint, not a convenience, and I had written code to honour it. But if honouring it produces an app that hands people the wrong score, there is nothing left to honour.

The code was set aside rather than deleted — though restoring it has a price. With the two frameworks that demanded it gone, the deployment target dropped from iOS 26 to 17, which widened device support considerably. Going back means handing that in again. It is a trade made with open eyes.

The interesting part is which direction the unexpected benefit came from. The app was demanding the newest OS in order to use its newest features — and once those features measured out as not good enough, the reason to demand it left with them.

What this conclusion cannot carry

The sample is five answers, one session, one speaker — and that speaker is a non-native one, which is a hard condition for any recogniser. The direction is clear, but it does not generalise to "on-device is bad". What I measured, and all I know, is that for this speaker and this use it was not a substitute.

What I have not tried is also worth recording. An English name stored in settings can be handed to the recogniser as a hint, which would fix exactly one of the five rows above. Place names and game titles never appear in the question, so no hint reaches them.


In short

  • An end-to-end run is a plumbing check, not a quality measurement. "It completed" and "the output is usable" are different sentences.
  • Compare two candidates on the same input. Compare different inputs and you cannot say what differed.
  • To judge something that wobbles, measure it repeatedly and get its own spread first. A difference counts only when it exceeds that spread.
  • Check whether the accuracy you lose turns into the user's fault. Past that point it stops being a quality issue.
  • Principles lose to measurements too. When a constraint you were protecting arrives as a loss for the user, the thing to fix is the constraint, not the result.

The engine choice across these two apps is in Whether the Grader Listens, and the case for suspecting the instrument before the subject is in It Was the Instrument Shaking, Not the Curve. The same instinct — that having nothing beats inventing something — is in Empty-Handed Beats Invented.