채점기가 소리를 듣는가
말하기 시험을 준비하는 앱을 두 개 만들었습니다. 먼저 영어 시험(OPIc)용 오픽 리허설을 만들어 Play 비공개 테스트까지 올린 뒤, 그 앱을 통째로 복사해 중국어 시험(TSC)용 TSC 리허설을 시작했습니다. 빌드 설정, 디자인 시스템, 저장 계층이 그대로 넘어왔으니 복사 자체는 쉬웠습니다. 형제 저장소는 모듈을 공유하지 않고, 파일을 복사한 뒤 각자 갈라져 나갑니다.
두 앱을 근본적으로 가른 것은 질문 하나였습니다. 채점기가 글을 읽는가, 소리를 듣는가. 사소해 보이는 이 한 가지가 엔진 선택과 탭 구성, 그리고 무엇에 점수를 매길 수 있는지까지 새로 정하게 만들었습니다.
오픽: 채점기는 글만 읽는다
오픽 앱에서 질문은 화면에 글로 나오지 않습니다. 음성(TTS)으로만 나오고, 재생이 끝나면 자동으로 녹음으로 넘어갑니다. 읽으면서 답하는 것은 말하기와 다른 능력이어서 실제 시험의 제약을 그대로 두었습니다.
답변은 녹음된 뒤 Groq Whisper 로 받아 적히고, 채점기는 그 받아 적은 텍스트만 봅니다. 여기서 한 가지가 저절로 정해집니다. 발음은 채점할 수 없습니다. 받아 적은 글에는 발음 정보가 없으니, 발음 점수를 만들면 지어낸 숫자가 됩니다.
채점 프롬프트에 이 원칙을 어기지 못하게 하는 문단을 넣고, 점수 필드의 주석에 이유를 적어 두었습니다. 발음 점수를 되살리려면 Azure 발음 평가를 붙이거나 오디오를 그대로 받는 엔진으로 바꿔야 합니다. 주석 한 줄을 고쳐서 될 일이 아닙니다.
엔진 쪽에도 같은 성질의 제약이 있습니다. Claude 는 오디오를 입력으로 받지 못합니다(Messages API 에 오디오 콘텐츠 블록이 없습니다). 엔진을 Claude 로 골라도 받아 적기는 늘 Groq Whisper 가 합니다. 엔진을 가르는 분기는 화면마다 두지 않고 한 곳(TestEngine)에만 둡니다.
나머지는 정직함을 지키기 위한 장치들입니다.
- 무료로 쓸 수 있는 경로는 꼭 지켜야 하는 제약입니다. Groq 키 하나로 질문 생성, 받아 적기, 채점, 저장이 모두 끝나야 하고, 어떤 코드 경로도 유료 키를 요구하면 안 됩니다. 심사하는 사람도 사용자도 키 하나로 전 과정을 돌 수 있어야 합니다.
- 세션 등급은 문항 등급의 평균이 아니라 중앙값입니다. 받아 적기가 깨진 답 하나가 세션 전체를 끌어내리지 않게 하려는 것입니다.
- 등급은 추정치입니다. 실제 채점자의 기준표 가중치와 등급 경계는 공개되지 않아서, "공식 성적이 아님"이라는 문구를 설정 안내, 결과 위쪽, 결과 아래쪽 세 곳에 넣었고 빼지 않습니다.
- 녹음은 캐시에만 쓰고 채점이 끝나면 지웁니다.
.gitignore도*.m4a와*.wav를 막습니다. 녹음이 커밋되면 사용자의 목소리가 새는 것이어서 미리 넣어 두었습니다.
TSC: 소리를 듣게 만들었다
TSC 에서는 발음이 공식 4대 평가 영역(문법, 어휘, 발음, 유창성) 가운데 하나이고, 중국어는 성조가 있는 언어입니다. 받아 적은 글만 보는 오픽의 구조를 그대로 가져오면 시험의 한 축이 통째로 빠집니다.
이 앱은 오디오를 그대로 받아 인식과 채점을 함께 하는 Gemini 를 처음부터 유일한 엔진으로 정했습니다. 엔진을 고르는 화면이 없는 것은 기능이 빠진 것이 아니라 결정한 것입니다. 텍스트만 보는 경로를 하나라도 남기면 그 경로에서는 네 영역 가운데 하나를 채점하지 못하는데, 그런 앱은 설정만 다른 같은 앱이라고 볼 수 없습니다. AIEngine 열거형에서 고를 수 있는 엔진이 Gemini 하나뿐인 이유입니다.
무엇을 베끼지 않았나
복사해서 만든 앱에서 가장 조심한 것은 옛 동작을 무심코 물려받는 일이었습니다.
설문 탭을 없앴습니다. OPIc 에는 배경 설문이 있어서 오픽 앱은 이것을 재현하느라 탭 하나를 씁니다. TSC 에는 사전 설문이 없습니다(공개된 응시자 안내를 다시 확인했습니다). 시험이 응시자에게 아무것도 묻지 않는데 앱이 '설문' 탭을 두면, 시험에 대해 틀린 것을 가르치게 됩니다. 대신 다양성은 직전 회차들이 쓴 화제(최근 14개)를 "이것은 피하라"는 목록으로 프롬프트에 넘겨 확보합니다. 사용자에게 묻지 않고 앱이 이미 만들어 둔 것에서 가져온다는 점이 중요합니다.
공개된 시험 구성은 표 그대로 넣었습니다. 7개 파트 26문항과 파트별 준비 시간, 답변 시간을 공식 사이트에서 직접 확인해 TscPart.kt 한 곳에 표 그대로 옮겼습니다. 여기 숫자가 틀렸다면 공개된 표와 다르게 틀린 것이니, 호출하는 쪽에서 우회하지 말고 이 파일을 고칩니다.
예상 시험 시간은 계산해서 얻은 값입니다. 이 앱은 23분으로 표시하는데, 공개된 문항별 준비 시간과 답변 시간을 더한 1,011초에 문항마다 재생 시간 12초를 더한 값입니다. 공식 시험은 약 30분입니다. 그 30분에 맞추려고 숫자를 키우지는 않았습니다. 차이는 파트 사이의 안내 방송처럼 이 앱이 재현하지 않는 부분에서 나오는 것으로 보이지만, 확인하지는 못했습니다. 맞추는 것과 재는 것은 다릅니다.
남의 저작물은 그대로 넣지 않습니다. 레벨 설명 원문은 저작권이 있는 자료여서, 읽고 레벨 사이의 차이만 뽑아낸 뒤 전부 새로 썼습니다. 기출 문항과 공식 설문 항목표도 복제하지 않았고, 질문은 모두 AI 가 만듭니다.
아직 못 한 일도 적어 둡니다. 오픽 앱은 이름에 "오픽/OPIC"만 단독으로 쓸 수 없습니다. 이미 등록된 다른 사람의 상표와 같아서, 반드시 다른 말과 결합한 이름이어야 합니다. 그렇게 해도 스토어에 신고가 들어올 수 있고, 그런 분쟁에서 가장 강한 반박은 본인 명의의 상표 등록인데 아직 하지 않았습니다. 열린 과제로 남겨 두었습니다.
정리하면
- 형제 앱은 코드보다 결정을 물려받고, 결정 하나(글인가 소리인가)가 엔진과 탭과 채점할 수 있는 범위를 새로 정합니다.
- 채점기가 관측하지 못하는 것(발음, 공식 등급 경계)에는 숫자를 만들지 않습니다. 되살리려면 코드를 바꿔야 하고, 주석을 바꿀 일이 아닙니다.
- 시험이 응시자에게 묻지 않는 것은 앱도 묻지 않습니다.
- 확인한 것과 확인하지 못한 것(23분과 30분의 차이, 상표 등록)을 나눠 적습니다.
I built two apps for rehearsing spoken-language exams. First OPIC Rehearsal for the English test (OPIc), which I took as far as Google Play closed testing, and then I copied it wholesale to start TSC Rehearsal for the Chinese one (TSC). The copy itself was easy — the build config, the design system, and the storage layer all came straight over. Sibling repos here don't share a module; a file is copied and then the two diverge.
But one question separated the two apps at the root. Does the grader read a transcript, or does it listen to the audio? That single, small-looking axis redrew the engine choice, the tabs, and what could be scored at all.
OPIc — the grader only reads text
In the OPIc app, a question is never printed on screen. It is spoken (via TTS) only, and when playback ends it switches to recording on its own. Reading while you answer is a different skill from speaking, so I kept the real test's constraint intact.
The answer is recorded, transcribed by Groq Whisper, and the grader sees only that transcript. One thing follows by force: pronunciation cannot be scored. A transcript carries no pronunciation signal, so producing a pronunciation number would be inventing one.
So the scoring prompt has a paragraph that forbids breaking this, and the score field's comment records why. Reversing it means adding Azure pronunciation assessment or moving to an audio-native engine — not editing a line of comment.
The engine side has a constraint of the same kind. Claude cannot take audio input (the Messages API has no audio content block). So even with Claude selected as the engine, transcription always goes through Groq Whisper. The engine branch lives in one place (TestEngine), not on any screen.
The rest are devices for staying honest.
- A free path is a hard constraint — one Groq key must run question generation, transcription, scoring, and storage end to end, and no code path may demand a paid key. A reviewer, and a user, must be able to run the whole thing on a single key.
- A session's level is the median of the per-question levels, not the mean — so one answer with a broken transcript doesn't drag the whole session down.
- The level is an estimate. Real raters' rubric weights and cutoffs are private, so "not an official score" sits in three places in the UI — the settings note, the top of the result, and the bottom. It stays.
- Recordings live only in cache and are deleted after scoring.
.gitignorealso blocks*.m4aand*.wav— a committed recording is a user's voice leaking, so I put that in pre-emptively.
TSC — so this one was made to listen
In TSC, pronunciation is one of the four official scoring areas (grammar, vocabulary, pronunciation, fluency), and Chinese is tonal. Carry over OPIc's transcript-only structure and a whole axis of the exam goes missing.
So this app made Gemini — which takes audio natively and does recognition and scoring together — the only engine from the start. Having no engine-picker screen is a decision, not a missing feature. Leave even one text-only path and that path can't produce one of the four areas — that's a different app, not a setting. That's why the AIEngine enum has exactly one selectable engine, Gemini.
What I refused to carry over
In a copied app, the thing to guard against most is inheriting old behaviour without noticing.
I removed the survey tab. OPIc has a background survey, so the OPIc app spends a tab reproducing it. But TSC has no pre-test survey (I re-checked the published candidate guide). If the exam asks the candidate nothing and the app keeps a “Survey” tab, it teaches the user something false about the test. Variety instead comes from feeding the last 14 rounds' topics back into the prompt as "avoid these" — the point being to draw from what the app already produced, not to ask the user.
The published test structure is baked into a table. Seven parts, 26 questions, and each part's prep and answer seconds were checked directly on the official site and put verbatim into one place, TscPart.kt. If a number here is wrong, it's wrong about the public table — so fix that file rather than working around it at the call site.
The estimated test length is a measured value. The app shows 23 minutes — the sum of the published per-question prep and answer times (1,011s) plus 12s of playback each. The official test runs about 30 minutes. I did not inflate the number to hit that 30. The gap looks like things this app doesn't reproduce, such as the between-part announcements, but I could not confirm that. Fitting a number is not the same as measuring it.
Someone else's text doesn't go in as-is. The level-descriptor source is copyrighted material, so I read it, extracted only the differences between levels, and rewrote all of it. Past question banks and the official survey category tables aren't copied either; every question is AI-generated.
And I write down what I haven't done. The OPIc app can't use "OPIC" alone in its name — it's identical to an already-registered third-party mark, so it must be a combined term. Registering it can still draw a store complaint, and the strongest rebuttal in that dispute is a trademark registered in my own name, which I don't have yet. I've left it as an open task.
In short
- A sibling app inherits decisions, not code — and one decision (read or listen) redraws the engine, the tabs, and what can be scored.
- Don't invent a number for what the grader cannot observe (pronunciation, the official cutoffs). Reversing it means changing code, not a comment.
- What the exam does not ask the candidate, the app does not ask either.
- Record what was verified apart from what was not (the 23-minute gap, the trademark).