혼자, 노트북 한 대로 일합니다. 그동안 앱 다섯 개를 스토어에 출시했고, 펌웨어부터 직접 짠 책상 위 기기를 만들었고, 장비 1,443대 규모의 라인 트윈을 세웠고, 매일·매주 사람 손 없이 돌아가는 자동화를 붙였습니다. 제가 빨라서 가능했던 일은 아닙니다.
판단은 한 사람이, 실행은 여럿이
일을 여러 AI 에이전트에게 역할별로 나눠 맡기면 실행을 병렬로 진행할 수 있습니다. 결정은 제가 혼자 내리니 순서대로 할 수밖에 없지만, 실행까지 그럴 필요는 없습니다. 한 작업이 진행되는 동안 다른 작업을 시작할 수 있어서, 동시에 진행할 수 있는 일의 수가 사람 수에 묶이지 않습니다.
대가도 따릅니다. 에이전트는 매번 처음부터 시작하고 이전 맥락을 기억하지 못하기 때문에, 그 역할이 알아야 할 것을 계약서처럼 문서로 먼저 정해 둡니다. 기억에 기대는 대신 문서로 고정하는 것입니다. 병렬로 진행하는 만큼, 각 작업이 무엇을 알고 있는지는 제가 관리해야 합니다.
구분해 둘 것이 하나 있습니다. 라인 트윈 위에서 교대마다 판단하고 서로 협상하는 에이전트 14개는 제가 만든 결과물입니다. 여기서 말하는 AI 에이전트는 그것을 만드는 동안 함께 일한 쪽입니다.
병렬로 맡길 수 없는 일이 하나 있습니다. 아래에 적은 되돌린 결정은 모두 사람이 내린 판단입니다. 모델은 스스로 틀렸다고 먼저 말하지 않습니다. 무엇을 잴지, 결과를 믿을지, 되돌릴지를 정하는 일은 나눠 맡길 수 없습니다. 그 과정을 제작기에 기록합니다.
재 보고 되돌린 결정 여섯 가지
만든 것보다 되돌린 결정을 더 많이 기록합니다. 그중 여섯 가지를 꼽으면 다음과 같습니다.
- 모델 대신 기준을 고쳤습니다. 장비 1,443대 규모의 라인 트윈이 기준 실행 로그와 3% 어긋난다고 나왔는데, 그 기준값의 추이를 그려 보니 4년이 지나도 수렴하지 않는 누적 평균이었습니다. 같은 로그의 재공과 산출로 Little's law 를 적용해 기준을 다시 잡자 오차가 +0.5%로 줄었습니다. 모델은 한 줄도 고치지 않았습니다. 4년째 자라고 있는 자로 쟀다 →
- 모델의 대표 결론을 철회했습니다. 우선순위는 대기시간을 줄이지 못하고 재분배할 뿐이라는 결론을 문서 첫 줄에 적어 뒀는데, 그 정리는 일감이 있으면 설비를 놀리지 않는다는 전제 위에 서 있었습니다. 실제 라인에는 재공도 설비도 멀쩡한데 품질 기준에 막혀 진행이 멈추는 구간이 있어 그 전제가 성립하지 않았습니다. 찾아낸 편향이 예외 없이 같은 방향이어서, 답을 점추정 대신 상한으로 제시하기로 했습니다. 설비는 노는데 재공은 쌓였다 →
- 12배 저렴한 모델로 바꾼 결정을 되돌렸습니다. 성조만 다른 최소대립쌍 10쌍을 합성해 실제 채점 경로에 넣어 봤더니 10쌍 중 하나도 잡아내지 못했습니다(0/10). 그 모델은 들은 소리를 옮기는 대신 말했어야 할 문장을 병음까지 적고 있었습니다. 알려진 오답을 넣어 보기 전에는 구별할 방법이 없었습니다.
- API 키 없이 동작해야 한다는 제약을 버렸습니다. 온디바이스 전사 경로를 붙여 API 키 없이 끝까지 동작하게 만들었는데, 같은 녹음으로 비교해 보니 두 전사 결과가 갈린 다섯 곳 모두 다른 쪽이 맞았습니다. 동작한다는 것과 쓸 만하다는 것은 다른 이야기였습니다.
- 목록의 수를 늘리는 대신 줄였습니다. 추천이 존재하지 않는 음반을 지어내지 않도록 검증된 카탈로그를 만들고 악장 번호를 전수 조사했습니다. 다른 곡에 잘못 연결된 항목이 나왔고 곡 수가 하나 줄었습니다. 정확한 수가 큰 수보다 낫습니다.
- 범인으로 지목한 곳은 결백했습니다. 책상용 기기의 목록 화면이 굼떠서 통신 속도를 의심했습니다. 재 보니 400kHz 에서 0.38ms, 100kHz 에서 1.04ms 로 세 배 가까이 차이가 나 범인처럼 보였습니다. 터치 쪽 속도만 올리고 다시 물었더니 체감은 같다고 했습니다. 0.6ms 는 손끝으로 느낄 수 있는 차이가 아니었고, 원인은 화면을 다시 그리는 쪽에 있었습니다. 예제대로 구웠으면 카드가 지워졌다 →
여섯 가지 모두, 재 보기 전에는 반대쪽이 맞아 보였습니다. 잰 수치와 그 수치가 바꾼 결정을 모두 모아 두었습니다 →
무엇을 만드는가
시작점은 저마다 다릅니다. 이론에서 출발한 것도 있고, 직접 쓰다가 아쉬웠던 점에서 출발한 것도 있습니다. 아래 여섯 가지가 하나의 파이프라인으로 이어져 있지는 않습니다.
- 앱은 iOS(SwiftUI)와 Android(Kotlin·Compose)로 만듭니다. 같은 앱을 두 플랫폼에서 만들면서, 플랫폼이 지원하지 않는 기능을 억지로 흉내 내는 대신 무엇을 포기했는지 설계 문서에 남기는 쪽을 택했습니다. 다섯 개가 스토어에 출시돼 있고, 스토어를 거치지 않고 자체 업데이트로 배포하는 안드로이드 문자 알림 앱이 하나 있습니다. 나머지는 비공개 테스트 중이거나 개발 중입니다.
- 윈도우 프로그램으로는 사무실 동료들이 쓰는 메모 프로그램을 만들었습니다. 바탕화면에 붙여 쓰는 메모이고, 서버 없이 PC 끼리 직접 공유해 양쪽에서 모두 고칠 수 있습니다. 스토어를 거치지 않으니 설치와 업데이트도 직접 만들었습니다. 관리자 권한 없이 설치되고, 새 판이 나오면 프로그램이 먼저 알린 뒤 버튼 한 번으로 업데이트합니다. 메모는 직접 지우기 전까지 사라지지 않도록 휴지통과 매일 자동 백업을 두었습니다.
- 기기는 펌웨어와 그것을 관리하는 앱을 함께 만듭니다. 책상에 두는 화면 기기로, 보드에서 도는 펌웨어는 C 로 짰고 사진·일정·화면 구성과 새 펌웨어를 보내는 쪽은 아이폰 앱으로 만들었습니다. 펌웨어는 무선으로 교체하고, 새 펌웨어가 정상적으로 뜨지 않으면 이전 버전으로 되돌아갑니다. 물리적인 기기는 되돌리는 비용이 크다는 것을 여기서 배웠습니다. 예제를 그대로 구웠다면 SD 카드가 지워질 뻔했습니다.
- 자동화는 매일·매주 사람 손 없이 돌아가는 작업들입니다. 실행할 때마다 처음부터 시작하기 때문에, 맥락을 기억에 맡기는 대신 문서로 고정해 둡니다.
- 시뮬레이션·최적화 모델은 문제를 대기행렬 이론으로 정식화하고, 이산사건 시뮬레이션으로 검증한 뒤, 실제로 쓸 수 있는 도구로 옮깁니다. 가장 큰 것은 장비 1,443대 규모의 라인 트윈이고, 그 위에서 교대마다 판단하고 서로 협상하는 에이전트 14개가 움직입니다. 기준 실행 로그 대비 최대 오차는 2.1%입니다.
- 교육 자료는 제가 익힌 것을 다른 사람에게 전하는 일입니다. 네 단계로 나눈 매뉴얼 127장을 썼고, 장마다 발표자 노트에 해설과 실습 정답을 넣었습니다. 쓸 줄 아는 것과 다른 사람이 따라 할 수 있게 적는 것은 다른 일이었습니다.
어떻게 일하는가
- 재는 자부터 의심합니다. 검증 결과가 표본마다 뒤집힌 적이 있습니다. 모델은 그대로였고 반복 실행 횟수가 모자랐습니다. 흔들린 것은 곡선이 아니라 재는 자였습니다. 그 뒤로는 판정에 95% 신뢰구간 반폭을 함께 적습니다. 다음번에는 자가 흔들린 것이 아니라 자라고 있었습니다. 기준으로 삼은 값이 4년이 지나도 수렴하지 않는 누적 평균이었습니다. 두 번 다 대상은 멀쩡했습니다.
- 상한값을 목표로 삼지 않습니다. 정적 용량 모델은 "가장 잘 풀리면 이만큼"을 재는 도구인데, 이를 운영 목표로 삼으면 설비를 늘려야 할 것처럼 보입니다. 실제로 필요했던 것은 투입 규칙의 상수 하나였고, 설비 투자 없이 처리량 3.5%를 얻을 수 있었습니다. 배치는 늘 꽉 찬다는 거짓말 →
- 자동 검사는 물어본 것만 답합니다. 4페이지 × 화면 폭 11가지 × 2개 언어, 88개 조합이 모두 통과했는데 두 번 다 틀린 곳은 눈으로 보고서야 찾았습니다. 넘침 여부만 확인하게 했으니 넘침만 확인한 것입니다. 화면이 바뀌는 수정은 지금도 반드시 눈으로 한 번 봅니다.
- 기본값도 한쪽으로 기울어 있습니다. 모르는 값을 채워 둔 기본값 다섯 개를 하나씩 바꿔 가며 다시 재 보니 넷이 같은 방향으로 기울어 있었고, 다섯을 한꺼번에 적용하자 답이 3.4배가 됐습니다. 답만 보고할 것이 아니라 그 답이 기대고 있는 가정까지 보고해야 했습니다. 중립인 줄 알았던 기본값 →
- 지어낸 답보다 빈칸이 낫습니다. 내놓을 것이 없을 때 그럴듯한 것을 만들어 내는 대신 비워 둡니다. 서로 다른 세 프로젝트에서 같은 결론에 이르렀습니다. 같은 이유로, 공개 데이터로 만든 지표는 리포트에 그 한계가 매번 함께 출력되게 했습니다. 읽는 사람이 따로 문서를 찾아볼 거라고 기대하지 않기 때문입니다. 공개 데이터는 열려 있지 깨끗하지 않다 →
지향하는 바
만드는 범위가 넓어져 왔습니다. 스토어에 출시하는 모바일 앱에서 시작해, 동료들이 쓰는 윈도우 프로그램과 펌웨어부터 직접 짠 책상 위 기기까지 왔습니다.
제가 지향하는 바는 현실에서 보고, 판단하고, 움직이는 AI입니다. 흔히 피지컬 AI 라고 부르는 분야입니다. 그 절반은 이미 만들어 두었습니다. 라인 트윈 위에서는 에이전트 14개가 교대마다 판단하지만 시뮬레이션 안의 일이고, 책상 위 기기는 현실을 측정하지만 판단은 하지 않습니다. 이 둘을 잇는 것이 다음 목표입니다.
물리적인 기기는 되돌리는 비용이 크다는 것을 기기를 만들며 배웠습니다. 움직이는 기계라면 그 비용은 더 커집니다. 이곳에 기록해 온, 재 보고 되돌리는 방식이 그 분야에서 더 쓸모 있으리라 생각합니다. 아직 직접 해 본 일은 아니어서, 해 보게 되면 여기에 기록하겠습니다.
왜 업계를 밝히지 않는가
여기 있는 글에는 산업명도 소속도 없습니다. 일부러 뺐습니다.
이름과 이메일을 걸고 쓰기 때문입니다. 업계를 밝히면 소속까지 짐작할 수 있게 되는데, 그건 제가 감당하기로 한 범위를 넘습니다.
대신 뺀 것보다 넣은 것이 훨씬 많습니다. 문제의 구조, 규모, 알고리즘과 이론, 검증 방법, 실측값, 그리고 제약과 그 이유는 모두 적었습니다. 어떤 판단이 옳았는지 따지는 데 필요한 것은 그쪽이고, 산업명은 없어도 됩니다.
빠진 것은 산업명뿐이라 여기 있는 수치는 직접 검증해 보실 수 있습니다. 틀린 곳이 있으면 그 수치를 들어 지적해 주세요.
저장소도 같은 이유로 비공개여서 직접 재현해 보실 수는 없습니다. 이 글들의 분명한 한계이고, 짐작에 맡기는 대신 이렇게 밝혀 둡니다.
One person, one laptop. In that time five apps went onto the stores, a desk device was built from the firmware up, a twin of a 1,443-machine line was built, and automations that run daily and weekly went into service. Not because I am fast.
One place decides, many places carry out
Handing work to several AI agents, each with its own role, makes execution run in parallel. Deciding is serial because there is one person doing it; executing has no such reason to be. One branch can be running while another is started, so the number of branches opened in a day is not bound to the number of people.
It has a price. An agent starts from an empty desk every time and remembers no context. So what a role needs to know is handed to it as a document, like a contract — pinned down rather than remembered. The more branches run at once, the more deliberately I have to control what each one knows.
One distinction is worth keeping. The fourteen agents deciding and negotiating each shift on top of the line twin are something I built, not someone I worked with. The AI agents here are the ones that were beside me while building.
And one seat cannot be parallelised. The reversals below were all human judgements — a model does not volunteer that it was wrong. Deciding what to measure, whether to believe the result, and whether to go back cannot be handed off. The build log is where I write down what happened at that seam.
Six decisions I measured, then reversed
More gets written down here about what I reversed than about what I made. Six of them, to start with.
- I fixed the baseline, not the model. A twin of a 1,443-machine line came out 3% off its reference run log. Plotting the trajectory of that baseline showed a running mean still unconverged after four years. Re-deriving it from the same log's WIP and output via Little's law brought the error to +0.5% — without changing a line of the model. I Measured With a Ruler That Was Still Growing →
- I withdrew the model's headline conclusion. Its documentation opened with the claim that priority only redistributes waiting time rather than reducing it, and the theorem behind that claim assumes a machine is never left idle while work is waiting. On the real line there are places where the WIP and the machine are both fine and a quality gate stops the work anyway, so that assumption does not hold. Every bias I found pointed the same way, so the model's answer is now read as an upper bound instead of a point estimate. The Machines Idled While the Queue Grew →
- I went back on switching to a model that was 12× cheaper. I synthesized ten minimal pairs differing only in tone and fed them through the app's actual grading path: 0 detected out of 10. The model was transcribing — pinyin and all — not the sound it heard but the sentence that should have been spoken. Short of feeding it known-wrong answers, there was no way to tell.
- I dropped the constraint that it must run without an API key. I wired an on-device transcription path and got the whole flow working with zero keys. Compared on the same recording, the five places the two transcripts disagreed went 5 : 0 the other way. Working end to end and being good enough are not the same claim.
- I made a list shorter instead of longer. To stop recommendations from inventing albums that do not exist, I built a vetted catalogue and then audited every movement number. Recordings filed under the wrong work came out, and the count of pieces went down by one. A correct number beats a big one.
- The suspect I pointed at was innocent. A list on the desk device felt sluggish, so I looked at the bus speed. Timing it gave 0.38ms at 400kHz against 1.04ms at 100kHz — nearly three times, which reads like a culprit. I raised only the touch controller and asked again: it felt the same. 0.6ms is not a quantity a fingertip can feel, and the sluggishness was the redraw. Flashing the Example Would Have Wiped the Card →
In all six, the opposite looked right until I measured. Every number I measured, and what each one changed →
What I build
The starting points differ. Some begin in theory; some begin at a place where something I was using annoyed me. The six below are not one pipeline in sequence.
- Apps — iOS (SwiftUI) and Android (Kotlin, Compose). Building the same app twice, I chose to write down what I gave up rather than fake what a platform cannot do. Five are on the stores, and an Android SMS alert app ships through its own updater instead of a store. The rest are in closed testing or still being built.
- Windows programs — a memo program that colleagues at the office use. Notes stick to the desktop and are shared directly between PCs with no server, with either side able to edit. With no store in between, I built the installer and the updater too: it installs without administrator rights, announces a new version itself and updates with one button. A note never disappears until someone deletes it; there is a bin and a daily automatic backup.
- Devices — firmware and the app that manages it, built together. A panel for a desk: the firmware on the board is written in C, and an iPhone app sends it photos, calendar events, screen layouts and new firmware. Firmware is replaced over the air, and if a new image fails to come up the board rolls back to the previous one. A physical device taught me how expensive undoing is — flashing the vendor's example would have wiped the SD card.
- Automations — things that run daily and weekly with nobody watching. Work that starts from an empty desk every time has to pin its context down in a document rather than remember it.
- Simulation and optimization models — formulate the problem in queueing theory, verify it with discrete-event simulation, then carry the result into a tool someone can actually use. The largest is a twin of a 1,443-machine line, with fourteen agents deciding and negotiating with each other every shift running on top of it. It validates to within 2.1% of the reference run log.
- Training material — the side that hands what I built to someone else. A 127-slide manual in four stages, every slide carrying speaker notes with the explanation and the exercise answers. Knowing how to use something and writing it down so another person can follow turned out to be different work.
How I work
- Suspect the instrument first. A verification result kept flipping between samples. The model had not changed — there were too few replications. What was shaking was not the curve but the instrument. Every verdict since carries the half-width of a 95% confidence interval. The next time the ruler was not wobbling but growing — the figure I had taken as a baseline was a running mean still unconverged after four years. Both times the subject was fine.
- A ceiling is not a target. A static capacity model measures "this much if everything goes well". Read as an operating target it makes buying equipment look necessary — when what was actually needed was one constant in the release rule, and a 3.5% throughput gain was sitting there for no capital spend. The Lie That Batches Are Always Full →
- An automated check answers only what it was asked. Four pages × eleven widths × two languages: all 88 combinations passed, and both times I found the broken layout by looking at it. I had asked about overflow, so overflow was all it answered. Anything that changes the screen still gets looked at.
- Defaults are not neutral. I reverted five "reasonable" defaults one at a time and re-ran; four of them leaned the same way, and stacking all five made the answer 3.4× larger. What needed reporting was not the answer but the assumptions it stood on. The Defaults Were Not Neutral →
- Empty-handed beats invented. When there is nothing to show, leave it blank rather than manufacture something plausible. Three separate places arrived at the same decision. For the same reason, the metrics built from public data make the report print its own limits every run — a reader is not expected to go looking for the documentation. Public Data Is Open, Not Clean →
Where I am heading
What I build has been widening. It started with mobile apps on the stores and has reached a Windows program colleagues use and a desk device written from the firmware up.
What I am aiming for is AI that sees, decides and moves in the real world — what is usually called physical AI. Half of it is already built. On the line twin, fourteen agents make decisions every shift, but inside a simulation; the desk device measures the real world, but decides nothing. Joining the two is the next goal.
Building the device taught me that undoing is expensive once there is hardware. With a machine that moves, it costs more still. I expect the habit recorded on this site, measuring and then being willing to reverse, to matter more there. I have not done that work yet; when I do, it will be written up here.
Why no industry is named
Nothing here names an industry or an employer. That is deliberate.
I write under my own name and address. Naming the industry would make the employer inferable, and that is past what I agreed to carry.
What is left out is far less than what is put in. The structure of the problem, the scale, the algorithms and the theory, how it was validated, the measured numbers, and the constraints with their reasons — all of it is written down. Judging whether a decision was sound needs those; it does not need the industry.
So the numbers here can be checked. If something is wrong, argue with the number.
The repositories are private for the same reason, which means you cannot reproduce any of this. That is a real limit of these write-ups, and I write it down rather than leave it to be assumed.