← 제작기

고쳤는데 스택이 그대로였다

번역 판본을 찾아 주는 앱을 앱스토어에 올렸더니, 심사 기기에서 검색을 누르는 순간 앱이 죽는다는 보고가 왔습니다. 크래시 로그의 스택은 이랬습니다.

SIGABRT
__cxa_pure_virtual
  → swift::AsyncTask::completeFuture

제가 가진 기기에서는 한 번도 나지 않았습니다. 재현할 수 없는 크래시를 로그만 보고 고쳐야 하는 상황이었습니다.

첫 번째 진단

__cxa_pure_virtual 은 순수 가상 함수를 호출했다는 뜻입니다. 구현이 없는 자리를 부른 것인데, Swift 에서 그런 자리가 생기는 곳은 프로토콜의 witness table 입니다. 마침 이 앱은 세 엔진(Claude, Groq, 온디바이스)을 any TranslationEngine 존재 타입으로 들고 다녔고, 그 엔진들은 actor 였습니다.

설명이 맞아떨어졌습니다. actor 가 구현한 async 메서드를 존재 타입을 거쳐 호출할 때 런타임이 witness table 의 빈 자리를 부른다고 봤습니다. 존재 타입을 걷어 내고, 구체 타입을 감싸는 래퍼를 만들어 직접 호출하게 한 뒤 빌드 2 를 올렸습니다.

같은 스택으로 다시 죽었습니다.

두 번째 진단

그렇다면 존재 타입보다는 actor 가 async throws 프로토콜을 준수한다는 사실 자체가 문제일 것이라고 봤습니다. 준수하기만 해도 witness table 에 그 항목이 생기기 때문입니다. 이번에는 프로토콜 선언을 통째로 지웠습니다. 래퍼가 이미 유일한 입구여서 동작은 그대로였습니다. 빌드 3 이었습니다.

또 같은 스택이었습니다.

여기서 멈췄어야 했다

세 번의 크래시 로그가 완전히 같았습니다. 같은 시그널, 같은 프레임, 같은 순서였습니다. 고칠 때마다 무언가는 달라져야 하는데 아무것도 달라지지 않았습니다.

원칙: 고친 뒤에도 증상이 조금도 움직이지 않았다면, 덜 고친 것이 아니라 원인을 잘못 짚은 것입니다.

저는 두 번째 진단을 하면서 이 신호를 그냥 지나쳤습니다. 이유는 분명합니다. 두 진단 모두 그럴듯했기 때문입니다. 존재 타입을 거친 호출도, 프로토콜 준수도 실제로 witness table 을 만드는 일이고, 스택에 찍힌 이름과도 맞아떨어집니다. 설명이 되니 믿었고, 믿었으니 같은 방향으로 한 걸음 더 팠습니다.

첫 수정이 실패했을 때 물었어야 할 질문은 "그럼 witness table 의 어느 부분이 문제일까"보다 "witness table 이 원인이 맞기는 한가"였습니다. 앞의 질문은 이미 답을 정해 놓고 묻는 것입니다.

범인은 제 코드가 아니었다

방향을 바꾸고 나서야 보였습니다. 크래시가 나는 조건은 기기의 상태였습니다. Apple Intelligence 가 켜져 있지 않거나 모델 자산을 아직 내려받지 않은 기기(notOptedIn / assetIsNotReady)에서, 온디바이스 모델 프레임워크가 태스크가 끝나는 시점에 스스로 죽고 있었습니다.

제 기기에서 나지 않았던 이유도 이것으로 설명됩니다. 제 기기는 그 기능이 켜져 있었고, 심사 기기는 그렇지 않았습니다. 재현이 안 됐던 것은, 재현에 필요한 조건이 기능이 꺼진 상태였기 때문이고 저는 그 상태를 만들어 보려 한 적이 없었습니다.

내 버그가 아니어도 비켜 가는 것은 내 일이다

원인을 알았다고 고칠 수 있는 것은 아니었습니다. 다른 회사의 프레임워크이기 때문입니다. 그렇다고 "제 버그가 아닙니다"로는 심사를 통과할 수 없습니다.

고칠 수 있는 것은 버그 자체가 아니라 버그에 닿는 경로였습니다.

  • 온디바이스를 고른 채 저장돼 있던 설정값은 다른 엔진으로 바꿔 줍니다. 이미 그렇게 설정해 둔 사람이 업데이트하자마자 앱이 죽으면 안 되기 때문입니다.
  • 그래도 그 경로로 들어오면 크래시 대신 이유가 적힌 오류를 냅니다.
  • 설정 화면에서 그 선택지를 숨겨 새로 고를 수 없게 했습니다.

기능을 빼다가 버그를 또 만들었다

이 문제는 코드 리뷰에서 잡았습니다. 예전에 온디바이스를 골라 두었던 사용자가 업데이트하면, 설정 화면의 엔진 타일이 아무것도 선택되지 않은 상태로 뜨고 키 입력란도 사라졌습니다. 저장된 값이 목록에 없는 값이 됐기 때문입니다.

기능을 빼는 일은 코드를 지우는 것으로 끝나지 않았습니다. 이미 그 기능을 고른 사람들을 어디로 보낼지 정하는 일이었습니다. 그 사람들은 제가 무엇을 뺐는지 모르고, 화면이 이상해진 것만 봅니다.

지우지 않은 것

흥미로운 점은 빌드 2 에서 만든 래퍼가 그대로 남아 있다는 것입니다. 틀린 진단에서 나온 수정이었지만 결과물 자체는 나쁘지 않았습니다. 엔진을 가르는 분기가 한 곳으로 모였기 때문입니다. 온디바이스 엔진 코드도 지우지 않고 두었습니다. 프레임워크가 고쳐지면 다시 열 자리이고, 그때 필요한 것은 빈 파일보다 왜 막았는지에 대한 기록입니다.

덧붙임: 경로를 막는 것으로도 부족했다

이 글을 올린 뒤 네 번째로 같은 스택이 났습니다. 위의 세 가지를 다 했는데도 그랬습니다.

코드가 그 경로로 가지 않는 것만으로는 부족했습니다. 온디바이스용으로 표시해 둔 타입이 바이너리 안에 들어 있기만 해도, 그 기능이 설정되지 않은 기기에서는 그 타입을 등록하는 과정이 Swift 비동기 런타임의 vtable 을 오염시켰습니다. 그러면 그다음 첫 async 작업이 끝나는 순간 앱이 죽습니다. __cxa_pure_virtual → swift::AsyncTask::completeFuture, 앞의 세 번과 완전히 같은 자리입니다.

import 와 그 타입 정의 전체를 컴파일에서 빼는 블록으로 감쌌습니다. 구현은 블록 안에 그대로 두었고, 되살리는 순서(먼저 대체 코드를 지우고 그다음 블록을 연다)를 파일 맨 위에 적었습니다. 그러지 않으면 블록만 지웠을 때 같은 이름이 두 개가 되어 컴파일이 깨집니다.

이번에는 원인은 맞았는데 처방이 부족했습니다. 앞의 세 번은 "증상이 움직이지 않았으니 원인을 잘못 짚었다"였는데, 네 번째는 원인을 제대로 짚고도 같은 스택이 다시 났습니다. 같은 신호가 두 가지를 뜻할 수 있다는 것입니다. 원인을 잘못 짚었거나, 처방이 원인에 닿지 못했거나입니다. 이번에는 "부르지 않는다"와 "들어 있지 않다" 사이의 거리였습니다.


정리하면

  • 증상이 움직이지 않으면 원인에 대한 이해가 틀린 것입니다. 같은 스택이 두 번 나왔다면 "덜 고쳤다"는 뜻이 아닙니다.
  • 그럴듯한 설명이 가장 위험합니다. 맞아떨어지는 이야기를 찾으면 그 방향으로만 더 파게 됩니다.
  • 수정이 실패한 뒤에는 "어디가 문제인가"보다 "여기가 원인이 맞기는 한가"를 묻습니다.
  • 재현이 안 되는 것처럼 보여도, 재현 조건이 "무언가 꺼진 상태"일 수 있습니다. 다 갖춰진 환경만 시험하면 그 조건은 끝내 만들어지지 않습니다.
  • 남의 버그라도 비켜 가는 것은 제 일입니다. 고칠 수 없으면 그 버그에 닿는 경로를 막습니다.
  • 같은 신호가 두 가지를 뜻할 수 있습니다. 원인을 잘못 짚었거나, 처방이 원인에 닿지 못했거나입니다. 네 번째 크래시는 뒤쪽이었고, "부르지 않는다"와 "들어 있지 않다"는 서로 다릅니다.
  • 기능을 빼는 일은 이미 그 기능을 고른 사람을 옮기는 일입니다. 코드만 지우면 그 사람들의 화면이 깨집니다.

같은 앱에서 결론이 뒤집혔던 다른 이야기는 범인은 서명이 아니었다에, 재는 도구를 먼저 의심해야 했던 사례는 흔들린 건 곡선이 아니라 재는 자였다에 있습니다. 온디바이스 모델을 실측 끝에 걷어 낸 다른 앱의 이야기는 앱이 제 실수를 내 탓으로 채점했다에 적어 두었습니다.

I shipped an app that finds Korean translations of a book, and the review device reported it dying the moment you hit search. The crash log's stack read:

SIGABRT
__cxa_pure_virtual
  → swift::AsyncTask::completeFuture

It never once happened on a device I owned. A crash I could not reproduce, to be fixed from the log alone.

The first diagnosis

__cxa_pure_virtual is a pure virtual call — something reached a slot with no implementation, and in Swift such slots live in a protocol's witness table. As it happened, this app carried its three engines (Claude, Groq, on-device) around as an any TranslationEngine existential, and those engines were actors.

It explained itself neatly: dispatching an actor's async method through an existential makes the runtime call an empty slot in the witness table. So the existential came out — a wrapper holding concrete types, dispatching directly — and build 2 went up.

It crashed again, on the same stack.

The second diagnosis

Then it isn't the existential; it must be the fact that an actor conforms to an async throws protocol at all, since conforming alone creates that entry. So the protocol declaration was deleted outright. The wrapper was already the only entry point, so nothing behaved differently. Build 3.

The same stack again.

This is where I should have stopped

Three crash logs, identical. Same signal, same frames, same order. Every fix should move something, and nothing had moved.

Principle: if the symptom didn't budge after a fix, you didn't under-fix it — your model of the cause is wrong.

And yet I walked past that signal into the second diagnosis. The reason is plain: both explanations were plausible. Existential dispatch and protocol conformance genuinely do produce witness tables, and the names line up with the stack. It explained things, so I believed it, and believing it I dug one more step in the same direction.

The question to ask after the first failed fix was not "which part of the witness table is at fault" but "is it the witness table at all". The first question has already decided its answer.

The culprit was not my code

Only after changing direction did it show. The condition for the crash was a device state — Apple Intelligence not turned on, or the model assets not yet downloaded (notOptedIn / assetIsNotReady). On such a device the on-device model framework dies of its own accord at task completion.

That also explains why my devices never showed it. Mine had the feature switched on; the review device did not. It was not that the bug wouldn't reproduce — reproducing it required something to be switched off, and I had never set out to arrange that.

Not my bug, still my job to step aside

Knowing the cause did not mean I could fix it — it is someone else's framework. And "this isn't my bug" does not get an app through review.

So what got fixed was not the bug but the route to it.

  • A stored setting still pointing at the on-device engine is moved to another one — someone who had chosen it must not die on first launch after updating.
  • If that path is entered anyway, it now throws an error that says why, rather than crashing.
  • The option is hidden in settings, so nobody new can choose it.

Removing the feature created its own bug

A code review caught this one. A user who had previously selected the on-device engine, on updating, found the engine tiles with nothing selected and the key field gone — their stored value was no longer among the options.

Removing a feature was not a matter of deleting code. It was deciding where to send the people who had already chosen it. They don't know what was removed; they only see that the screen has gone strange.

What wasn't deleted

The amusing part is that the wrapper from build 2 is still there. It came out of a wrong diagnosis, yet the thing itself was no loss — engine branching ended up in one place. The on-device engine's code stayed too. If the framework is fixed, that door reopens, and what will be needed then is why it was closed, not an empty file.

Postscript — closing the route was not enough

After this was published, the same stack came back a fourth time — with all three of the above in place.

It was not enough for the code to never take that route. Merely having the on-device-annotated types present in the binary was enough: on a device where the feature is not configured, registering those types corrupted the Swift async runtime's vtable, and the next async task to complete died — __cxa_pure_virtual → swift::AsyncTask::completeFuture, the same frames as the first three.

So the import and every one of those type definitions went inside a block excluded from compilation. The implementation stays inside it, and the order for bringing it back — delete the stub first, then open the block — is written at the top of the file, because deleting only the block leaves two declarations of the same name and the build breaks.

This time the cause was right and the remedy fell short. The first three rounds said "the symptom didn't move, so the cause is wrong". The fourth had the cause right and still produced the same stack. Which means the same signal can mean two things — the cause is wrong, or the remedy never reached it. Here the gap was between "we don't call it" and "it isn't in there".


In short

  • If the symptom doesn't move, the model of the cause is wrong. The same stack twice does not mean "not fixed enough".
  • A plausible explanation is the dangerous kind. Find a story that fits and you will only dig further along it.
  • After a failed fix, ask "is this the right place at all", not "which part of this place".
  • A bug that won't reproduce may need something switched off. Test only well-provisioned environments and that condition never arises.
  • Someone else's bug is still yours to avoid. If you cannot fix it, close the route to it.
  • One signal can mean two things. The cause is wrong, or the remedy never reached it — the fourth crash was the latter. "We don't call it" is not "it isn't in there".
  • Removing a feature means relocating whoever chose it. Delete only the code and their screen breaks.

An earlier conclusion overturned on a different app is in The Culprit Wasn't the Signature, and the case for suspecting the instrument first is in It Was the Instrument Shaking, Not the Curve. Another app where an on-device model was measured and dropped is in The App Graded Its Own Mistake as Mine.