A4 한 장입니다. ⌘P → PDF로 저장 하면 메일에 붙일 수 있습니다. 지금 보이는 언어로 인쇄됩니다. 이 안내문과 위 메뉴는 종이에 나오지 않습니다.
One A4 page. ⌘P → Save as PDF to attach it to an email. It prints in whichever language is showing. This note and the menu above stay off the paper.
한 사람 · 노트북 한 대 · 여러 AI 에이전트
이론에서 정식화하고, 시뮬레이션으로 검증하고, 손에 쥐여 준다
앱과 기기, 자동화, 시뮬레이션·최적화 모델을 만듭니다. 판단은 제가 하고 실행은 여러 AI 에이전트에게 나눠 맡겨서, 한 사람이 여러 일을 동시에 진행합니다. 만든 것보다 되돌린 결정을 더 많이 기록하며, 현실에서 보고 판단하고 움직이는 AI 를 지향합니다.
대표 작업
- 대규모 라인 트윈 + 에이전트 조직 — 장비 1,443대 · 스텝 926~4,013의 재진입 라인을 이산사건 시뮬레이션으로. 그 위에 교대마다 판단하고 서로 협상하는 에이전트 14개. 기준 실행 로그 대비 최대 오차 2.1%
- Classical Mood (iOS · Android) — 몸 상태·시간·날씨로 클래식 곡과 특정 연주 음반까지 고르는 앱. 양쪽 스토어 배포. 추천을 모델 생성이 아니라 검증된 카탈로그 조회로 바꿔, 없는 음반을 지어내던 문제를 없앰
- 우선처리 물량 산출 모델 — 대기행렬 이론으로 정식화하고 시뮬레이션으로 교차검증. 외부 패키지 의존 0, 검증 94건
- 사무실 메모 프로그램 (Windows) — 동료들이 쓰는 바탕화면 메모. 서버 없이 PC 끼리 직접 공유하고 양쪽에서 고침. 스토어 없이 설치 파일과 자체 업데이트로 배포하며, 관리자 권한 없이 설치되고 새 판을 프로그램이 먼저 알림
- 데스크 대시보드 (ESP32-S3 · iOS) — 책상에 두는 800×480 화면 기기. 펌웨어는 C, 관리 앱은 SwiftUI 로 직접 만듦. 화면 구성·일정·사진·새 펌웨어를 앱에서 블루투스로 보내고, 새 펌웨어가 뜨지 않으면 이전 판으로 되돌아감
- 업계 재고 사이클 판독기 — 공개 공시 데이터만으로 국면 판정. 누계 차분 복원을 연간 원본과 대조해 최대 오차 0.0007%
재 보고 되돌린 결정
- 모델이 아니라 기준을 고쳤다 — 기준값의 궤적을 그려 보니 4년이 지나도 수렴하지 않는 누적 평균이었다. 같은 로그의 재공과 산출로 Little's law 환산해 기준을 다시 잡자 오차가 +3.0% → +0.5%. 모델은 한 줄도 안 고쳤다
- 12배 저렴한 모델로 바꾼 결정을 되돌렸다 — 성조만 다른 최소대립쌍 10쌍을 실제 채점 경로에 넣었더니 탐지 0/10. 그 모델은 들은 소리가 아니라 말했어야 할 문장을 옮겨 적고 있었다
- 목록의 수를 늘리는 대신 줄였다 — 악장 번호를 전수 조사하니 다른 곡에 걸려 있던 것들이 나왔다. 정확한 수가 큰 수보다 낫다
어떻게 일하는가
- 재는 자부터 의심한다. 판정은 점추정이 아니라 95% 신뢰구간 하한으로 한다 — 반복이 모자라면 노이즈가 버그처럼 보이고, 그걸 쫓으면 멀쩡한 코드를 고친다
- 자동 검사는 물어본 것만 답한다. 4페이지 × 화면 폭 11가지 × 2개 언어, 88개 조합이 전부 통과했는데, 틀린 것은 두 번 다 눈으로 보고서야 찾았다
- 한계를 결과보다 앞에 적는다. 뒤에 적으면 안 읽힌다. 리포트가 자기 한계를 매번 같이 출력하게 만든다
- 지어낸 답보다 빈칸이 낫다. 내놓을 것이 없을 때 그럴듯한 것을 만드는 대신 비워 둔다
One person · One laptop · Many AI agents
From theory, through simulation, into someone's hands
I build apps, devices, automations, and simulation and optimization models. I make the decisions and hand the execution to several AI agents, so one person can run many pieces of work at once. I write down more about decisions I reversed than about what I made, and I am aiming for AI that sees, decides and moves in the real world.
Selected work
- Large-line twin + agent organization — a re-entrant line of 1,443 machines and 926–4,013 process steps as a discrete-event simulation, with 14 agents deciding and negotiating every shift on top of it. Within 2.1% of the reference run log
- Classical Mood (iOS · Android) — picks a classical piece, down to the recording, from body data, time and weather. On both stores. Picks now come from a vetted catalogue instead of model generation, which ended invented albums
- Priority-work sizing model — formulated in queueing theory, cross-checked by simulation. Zero third-party dependencies, 94 tests
- Office memo program (Windows) — desktop notes colleagues use, shared PC to PC with no server, self-updating
- Desk dashboard (ESP32-S3 · iOS) — an 800×480 desk panel with C firmware and a SwiftUI app; updates over Bluetooth, rolls back on failure
- Inventory cycle reader — reads the cycle phase from public filings alone. Reconstructed quarters checked against annual originals to within 0.0007%
Decisions I measured, then reversed
- I fixed the baseline, not the model — plotting its trajectory showed a running mean still unconverged after four years. Re-deriving it from the same log's WIP and output via Little's law moved the error from +3.0% to +0.5%, without changing a line of the model
- I went back on a model that was 12× cheaper — ten tone-minimal pairs through the real scoring path: 0/10 detected. It was transcribing the sentence that should have been spoken, not the sound it heard
- I made a list shorter instead of longer — auditing movement numbers surfaced recordings filed under the wrong work. A correct number beats a big one
How I work
- Suspect the instrument first. Verdicts use the lower bound of a 95% confidence interval, not a point estimate — too few replications make noise look like a bug, and chasing it means editing code that was fine
- An automated check answers only what it was asked. All 88 combinations of four pages × eleven widths × two languages passed, and both times the broken layout was found by looking at it
- Limits go before the results. Put them after and they do not get read. Reports print their own limits every run
- Empty-handed beats invented. When there is nothing to show, leave it blank rather than manufacture something plausible