계측기를 놓은 날
규칙은 사람이 읽고 지키는 문서였습니다. 9월 23일 하루 동안 그 규칙을, AI 가 코드를 고친 뒤 스스로 돌리는 검사로 바꿨습니다. 코워크에서 Claude 에게 AI 활용 평가를 받았고, 공개된 실무자들의 방식에서 모은 보완 네 가지가 나왔습니다. 스스로 확인하는 수단, 무인 작업의 시험 문제집, 삼중 위험 점검, 고친 것이 쌓이는 구조입니다.
평가는 이렇게 비유했습니다. 지금까지는 뛰어난 품질 검사원 한 사람에게 기댔고, 다음 단계는 작업자 옆에 계측기를 놓는 것이라는 비유입니다. 이 글은 그 네 가지를 git 저장소 30개에 하루 동안 적용한 기록입니다. 수치는 모두 그날 저장소에서 직접 세거나 일부러 깨뜨려 확인한 값입니다.
0 → 10 자동 확인 명령을 가진 앱 저장소
19 / 19 일부러 깨뜨린 자리 중 검사가 잡은 곳
14 → 0 너무 넓게 열려 있던 권한 규칙
−20% 큰 저장소 셋이 세션마다 싣는 문서 양
(잃은 줄 0, 말없이 꺼졌던 경고 1건은 되살림)
재는 자가 틀렸던 열네 번
가장 오래된 규칙은 "재는 자부터 의심한다"입니다. 고장 났다는 신호가 뜨면 대상보다 재는 도구를 먼저 봅니다. 흔들린 건 곡선이 아니라 재는 자였다와 4년째 자라고 있는 자로 쟀다를 겪으며 생긴 규칙인데, 그날 이 규칙이 열네 번 쓰였습니다. 모두 결과를 내기 전에 잡혔고, 그중 셋은 규칙 문서에 틀린 수치를 적기 직전이었습니다.
| 어디서 | 보인 것 | 실제로는 |
|---|---|---|
| 검사기의 코드 | 기본 엔진 설정이 한 곳 다르다 | 검색이 검사 스크립트 자신과 문서까지 셌다 |
| 검사기의 코드 | 새 API 주소가 코드에 없다 | 주석을 걷는 식이 https:// 의 // 뒤를 통째로 지웠다 |
| 검사기의 코드 | 옛 API 주소를 아직 쓴다 | 이관을 설명하는 주석을 사용으로 셌다 |
| 검사기의 코드 | 회차 줄을 못 읽었다 | 그 한 줄 때문에 같은 카드의 톤 검사를 건너뛰어 진짜 위반을 가렸다 |
| 문자 인식 | 박스·페이지 번호·출처가 없는 카드 3장 | 띄어쓰기를 붙여 읽고, / 를 ½ 로, 출처를 줄처로 읽었다. 카드는 멀쩡했다 |
| 문자 인식 | 회차 줄을 못 읽는다 | 다시 저장한 이미지에서 NO.9 를 N0.9 로 읽었다 |
| 셸과 sed | 배포 직전 리뷰가 45번 중 0번 | zsh 가 여러 줄을 한 덩어리로 한 번만 돌렸다. 실제는 8번 |
| 셸과 sed | 깨뜨렸는데 검사가 안 잡는다 | zsh 가 좌표 넷을 한 인자로 넘겨 그림이 안 바뀌었다. 바뀌었는지 먼저 센 덕에 오판하지 않았다 |
| 셸과 sed | 날짜 비교가 오류로 끝난다 | zsh 의 [ ] 는 문자열 크기 비교를 못 한다 |
| 셸과 sed | 깨뜨렸는데 검사가 통과한다 | macOS sed 가 GNU 문법을 조용히 무시해 파일이 그대로였다. 검사는 맞았다 |
| 셸과 sed | 스크립트가 시작하자마자 죽는다 | sed 의 줄 지우기가 첫 줄의 import 를 지웠다 |
| 문서 옮기기 | 검사가 모두 통과한다 | 경고 하나가 말없이 꺼져 있었다. 경고의 근거가 옮겨 간 기록 안에 있었다 |
| 규칙 문서의 수치 | "오늘 세 파일에서 23개를 지웠다" | 9월 7일에 12개, 이날 14개였다. 날짜도 개수도 달랐다 |
| 규칙 문서의 수치 | "하루에 넷을 밟았다" | 다섯이었다 |
검사가 찾은 진짜 문제
잘못 잰 것만 있었던 것은 아닙니다. 같은 검사가 오래 묵은 어긋남 셋을 찾았습니다. 코드는 바뀌었는데 그 코드를 설명하는 글이 따라오지 않은 자리들이었고, 문서는 커밋에 딸려오지 않는다에서 적은 일이 이번에는 검사에 걸렸습니다.
- 출시 뒤에도 "v0.1.0". 스토어에 1.0 이 나간 뒤에도 문서 첫 문단이 개발 초기 버전을 말하고 있었습니다. 저장소의 실제 값으로 고쳤고, 스토어 상태는 조회하기 전에는 단정하지 않았습니다.
- 꺼 둔 코드 속 크래시 패턴. 예전에 크래시를 낸 병렬 처리 패턴이 비활성 블록 안에 남아 있었습니다. 되살리면 같이 돌아오는 자리라, 이제 되살리는 순간 검사가 실패하게 했습니다.
- 출시된 앱이 "개발 중". 제출 기록이 있는데 첫 문단은 개발 중이었습니다. 게시 여부는 스토어를 봐야 알 수 있어서, 고치지 않고 경고로만 띄웁니다.
네 가지를 적용한 전과 후
평가의 네 제안별로 묶었습니다. 판정은 두 가지입니다. "측정됨"은 그날 일부러 깨뜨리거나 세어서 확인한 것이고, "실측 전"은 장치는 놓았지만 실제 운영에서 아직 돌지 않은 것입니다.
스스로 확인하는 수단
앱의 자기 검증. 9월 23일 아침에는 앱 저장소 10곳 모두 자동 확인이 없어서, 문서와 코드가 어긋났는지를 사람이 눈으로 봤습니다. 그날 밤에는 10곳 모두 1초짜리 확인 명령이 생겼습니다. 문서가 단언한 수와 "고치지 않는다" 항목, 형제 저장소 사이의 약속까지 대조하고, 어긋나면 문서의 줄번호를 찍습니다. 측정됨: 앱마다 깨뜨린 자리를 모두 잡았습니다.
무인 작업의 시험 문제집
예약 작업의 결과물. 아침에는 지침을 고친 뒤 예전 회차가 깨졌는지 볼 장치가 없었고, 검수는 카드를 그린 모델이 같은 카드를 다시 보는 식이었습니다. 밤에는 카드에 실제로 그려진 글자를 읽어 검수 목록을 잽니다. 지난 회차 전체를 다시 도는 회귀 시험이 있고, 사람이 채울 라벨 표 31장이 있습니다. 측정됨: 5곳을 깨뜨려 5곳 모두 잡았고, 라벨 칸은 아직 비어 있습니다.
삼중 위험 점검
무인 작업의 도구. 아침에는 신뢰할 수 없는 입력을 읽는 무인 작업 하나가 밖으로 내보내는 도구까지 함께 갖고 있었습니다. 밤에는 그 도구를 뺐습니다. 읽은 글 속에 숨은 지시가 있어도 내보낼 방법이 없고, 규칙 밖의 일은 멈추고 보고합니다. 측정됨: 싣는 도구 목록을 세어 확인했습니다.
권한 설정. 아침에는 너무 넓게 열려 있어서 한 번 허용한 명령 뒤에 무엇이 붙든 통과시키던 권한 규칙이 14개 있었습니다. 밤에는 0개가 됐고, 권한 파일을 빠짐없이 찾아 보는 규칙이 생겼습니다. 측정됨: 경고 0건을 확인했습니다.
고친 것이 쌓이는 구조
배포 전 리뷰. 아침에는 버전을 올린 45번 중 직전에 코드 리뷰가 있던 것이 8번이었고, 심사 없이 자체 배포되는 알림 앱은 5번 중 0번이었습니다. 밤에는 배포하는 순간에 거는 절차가 생겼습니다. 확인 명령과 코드 리뷰를 거친 뒤, 찾은 문제 중 되풀이될 것은 검사나 함정 목록으로 올립니다. 실측 전: 2026-09-23 기준으로 아직 한 번도 쓰이지 않았습니다.
세션마다 싣는 문서. 아침에는 큰 저장소 셋이 세션마다 6만~10만 자를 자동으로 실었습니다. 한 줄만 고치러 와도 그만큼이 먼저 실렸습니다. 밤에는 날짜순 기록만 따로 옮겨 약 20% 줄였습니다. 결정과 함정은 하나도 옮기지 않았고, "2주 안 쓰인 규칙은 지운다"는 권고는 따르지 않았습니다. 측정됨: 옮기기 전후를 줄 단위로 대조해 잃은 줄은 0이었고, 말없이 꺼졌던 경고 1건을 찾아 되살렸습니다.
사고의 교훈. 아침에는 임시 브랜치에 32커밋이 쌓여 스토어 빌드가 기본 브랜치 밖에 있던 일, 형제 저장소 약속이 닷새 동안 깨져 있던 일이 각자 한 저장소의 기록에만 있었습니다. 밤에는 모든 세션에 실리는 전역 규칙과 17개 저장소 문서의 뼈대에 들어갔습니다. 측정됨: 30개 저장소가 모두 기본 브랜치에 있는 것을 확인했습니다.
숫자로 본 전과 후
버전을 올린 커밋마다 직전 여섯 커밋 안에 코드 리뷰가 있었는지 셌습니다. 합계 45번 중 8번, 18%입니다. 리뷰가 가장 필요한 자체 배포 앱이 가장 비어 있었습니다.
리뷰 / 올림 앱
2 / 9 Classical Mood iOS
2 / 9 Classical Mood Android
0 / 5 오픽 리허설 Android
0 / 5 자체 배포 알림 앱 (심사 없음)
3 / 4 옮김
1 / 4 TSC 리허설 Android
0 / 4 KidStay
0 / 4 오픽 리허설 iOS
0 / 1 TSC 리허설 iOS
8 / 45 합계 18%
세션마다 자동으로 싣는 문서 양입니다(전역 규칙 포함, 천 자 단위). 날짜순 검증 기록과 제출·실주행 기록만 따로 옮겼고, 남은 80%는 설계 결정과 함정이라 일부러 남겼습니다.
옮기기 전 → 후 줄어든 양 앱
101.9 → 80.5 −21% TSC 리허설 Android
90.5 → 72.0 −20% 오픽 리허설 Android
61.0 → 49.2 −19% 오픽 리허설 iOS
AI 활용에서 달라진 것
검증하는 주체: 사람의 눈에서 스스로 돌리는 검사로
AI 가 코드를 고친 뒤 확인 명령을 먼저 돌립니다. 어긋나면 어느 문서의 몇 번째 줄이 틀렸는지까지 받습니다. 어느 쪽이 맞는지는 판정하지 않고 사람에게 넘깁니다.
규칙을 지키는 방식: 기억에 기대던 데서 구조로 막는 쪽으로
신뢰할 수 없는 입력을 읽는 무인 작업은 이제 내보내는 도구가 없어서 보낼 수가 없습니다. 확인 명령은 규칙 문서가 적용되지 않는 날에도 돕니다. 모델의 판단에 기대던 자리가 도구 목록과 검사로 바뀌었습니다.
검사기를 믿는 방식: "통과"를 믿던 데서 사람 판정과 대조하는 쪽으로
카드 검사기는 첫 실행에서 실패 셋을 냈는데, 셋 다 문자 인식 오류였습니다. 검사기가 스스로 통과라고 말하는 것은 증명이 되지 못하므로, 사람만 채우는 라벨 칸을 함께 두었습니다. 이는 검사는 통과했는데 화면이 틀렸다에서 배운 것과 같은 규칙입니다.
고친 것이 쌓이는가: 매번 다시 찾던 데서 검사로 올라가는 쪽으로
꺼 둔 코드 안에 예전 크래시 패턴이 남아 있었습니다. 그 코드를 되살리라는 지시와 크래시 이력이 서로 다른 절에 있어 이어지지 않았습니다. 이제 되살리는 순간 검사가 실패합니다.
아직 증명되지 않은 것
장치를 놓은 것과 장치가 운영에서 일하는 것은 다릅니다. 아래는 2026-09-23 기준으로 확인하지 못한 것이고, 이 밖에 새로 건 제한 중 일부도 운영에서 아직 돌지 않았습니다. 확인되면 이 자리에 정정을 달겠습니다.
- 새 세션에서 바뀐 규칙이 실제로 걸리는지. 검증용 세션의 로그인이 만료돼 띄우지 못했습니다(미확인).
2026-09-25 — 확인했습니다. 검증용 세션이 뜨지 않은 원인은 새 세션을 띄우는 터미널의 Claude Code 가 계정에서 로그아웃된 상태였던 데 있었습니다. 다시 로그인한 뒤 새 세션에서 규칙 문서 일곱 개가 모두 불렸고, 규칙 이름을 대지 않은 요청 두 개에도 알맞은 규칙이 걸렸습니다. 문장을 고쳐 달라는 요청에는 한국어 글쓰기 규칙이, 배포 전에 할 일을 묻는 질문에는 배포 전 리뷰 규칙이 걸렸고, 시험에 쓴 저장소는 한 글자도 바뀌지 않았습니다. 막혀 있던 것은 규칙보다 시험을 띄우는 환경 쪽이었으니, 재는 자가 틀린 경우가 하나 더 생긴 셈입니다. - 카드 검사기의 신뢰도. 사람이 채울 라벨 칸 31개가 비어 있어서, 채워야 잴 수 있습니다(미확인).
- 배포 전 리뷰 절차. 실제 배포에서 한 번도 쓰이지 않았습니다(미확인).
- 확인 명령이 보는 범위. 코드와 문서를 대조할 뿐, 앱이 실제로 도는지는 보지 않습니다. 이건 한계로 남습니다.
구조는 갖췄습니다. 운영에서의 증명은 이제부터입니다.
이 글은 다른 세션이 쓴 페이지를 옮겨 온 것입니다. 세션 사이에 전해진 결정을 어떻게 다뤘는지는 커밋하지 않은 변경은 누구 것인지 알 수 없었다에 적었습니다.
이 규칙들에 이어 한국어 쓰기 규칙을 만든 이야기는 번역 투를 의심했는데 박자였다에 있습니다.
네 제안의 출처는 이렇습니다. 스스로 확인하는 수단은 Boris Cherny, 시험 문제집은 Hamel Husain, 삼중 위험은 Simon Willison, 고친 것이 쌓이는 구조는 Every 의 컴파운드 엔지니어링입니다.
My rules used to be documents that a person read and followed. Over one day, September 23, I turned them into checks that the AI runs by itself after it changes code. An AI-usage review I had Claude do in Cowork, drawing on how practitioners publicly work, came back with four additions: a way to check its own work, a test set for unattended jobs, a lethal-trifecta check, and a structure where fixes accumulate.
The review put it this way. So far I had leaned on one excellent quality inspector; the next step was to put gauges beside the workers. This is the record of applying those four things to 30 git repositories in a single day. Every number here was counted in the repositories that day, or confirmed by breaking something on purpose.
0 → 10 app repositories with an automatic check command
19 / 19 deliberately broken spots the checks caught
14 → 0 permission rules that were opened far too wide
−20% documentation three large repositories load every session
(no lines lost; one warning that had silently switched off was restored)
Fourteen times the ruler was wrong
My oldest rule is "suspect the ruler first": when something signals that it is broken, look at the measuring instrument before the thing being measured. It came out of It Was the Instrument Shaking, Not the Curve and I Measured With a Ruler That Was Still Growing, and that day it went to work fourteen times. Every one was caught before a result went out, and three of them were about to put a wrong number into a rules document.
| Where | What it looked like | What it actually was |
|---|---|---|
| Checker code | One default engine setting differs | The search counted the check script itself and the docs |
| Checker code | The new API address is missing from the code | The comment-stripping expression deleted everything after the // in https:// |
| Checker code | The old API address is still in use | A comment describing the migration was counted as use |
| Checker code | The issue line could not be read | Because of that one line, the tone check on the same card was skipped, hiding a real violation |
| Text recognition | Three cards missing a box, page number or source | It merged spaces, read / as ½, and misread the source label. The cards were fine |
| Text recognition | The issue line cannot be read | On a re-saved image it read NO.9 as N0.9 |
| Shell and sed | Zero of 45 releases had a review just before | zsh ran a multi-line block once as a single chunk. The real figure was 8 |
| Shell and sed | I broke it and the check did not catch it | zsh passed four coordinates as one argument, so the image never changed. Counting whether it changed first avoided a false verdict |
| Shell and sed | Date comparison ends in an error | zsh's [ ] cannot compare strings by order |
| Shell and sed | I broke it and the check passed | macOS sed silently ignored GNU syntax, so the file was unchanged. The check was right |
| Shell and sed | The script dies the moment it starts | sed's line deletion removed the import on the first line |
| Moving docs | Every check passes | One warning had silently switched off; its evidence lived in the records that were moved |
| Numbers in rules docs | "Removed 23 from three files today" | It was 12 on September 7 and 14 that day. Both the date and the count were wrong |
| Numbers in rules docs | "Hit four in one day" | It was five |
Real problems the checks found
It was not all mismeasurement. The same checks found three long-standing mismatches, all places where the code had moved on and the writing about it had not. What I wrote in Docs Don't Ride Along With the Commit got caught by a check this time.
- Still "v0.1.0" after release. After 1.0 shipped to the store, the first paragraph of the docs still described an early development version. I corrected it to the repository's actual value, and did not state the store status without looking it up.
- A crash pattern inside disabled code. A concurrency pattern that once caused a crash was still sitting in an inactive block, ready to come back with it. Now the check fails the moment that block is revived.
- A released app described as "in development". There was a submission record, but the first paragraph said in development. Whether it is live can only be read from the store, so this one is raised as a warning rather than edited.
Before and after the four changes
Grouped by the review's four proposals. There are two verdicts. "Measured" means confirmed that day by breaking something on purpose or by counting; "not yet in operation" means the device is in place but has not run in real operation yet.
A way to check its own work
App self-checks. On the morning of September 23, none of the 10 app repositories had an automatic check; whether docs and code disagreed was something a person looked for. By that night all 10 had a one-second check command. It compares the numbers the docs assert, the "do not change" items and the promises between sibling repositories, and prints the doc line number when they disagree. Measured: every spot broken in each app was caught.
A test set for unattended jobs
Output of a scheduled job. In the morning there was no way to see whether earlier issues broke after the instructions changed, and review meant the model that drew a card looking at the same card again. By night the review list is measured by reading the text actually drawn on each card. There is a regression run over every past issue, and 31 label sheets for a person to fill in. Measured: five spots broken, all five caught; the label cells are still empty.
Lethal-trifecta check
Tools of an unattended job. In the morning, one unattended job that reads untrusted input also carried a tool that sends things out. By night that tool was gone. Even if what it reads hides an instruction, there is no way to send anything out, and anything outside its rules stops and gets reported. Measured: confirmed by counting the tools it loads.
Permission settings. In the morning there were 14 permission rules opened so wide that once a command was allowed, anything appended to it passed. By night there were none, along with a rule to find every permission file without exception. Measured: zero warnings confirmed.
A structure where fixes accumulate
Review before release. In the morning, only 8 of 45 version bumps had a code review just before them, and the alert app that ships without store review had 0 of 5. By night there is a procedure attached to the moment of release: the check command and a code review, after which anything likely to recur is promoted into a check or the pitfalls list. Not yet in operation: as of 2026-09-23 it had not been used once.
Docs loaded every session. In the morning, three large repositories loaded 60,000 to 100,000 characters automatically every session; even a one-line fix paid that first. By night only the dated logs had been moved out, cutting about 20%. Not one decision or pitfall was moved, and I did not follow the advice to delete rules unused for two weeks. Measured: comparing before and after line by line lost zero lines, and one warning that had silently switched off was found and restored.
Lessons from incidents. In the morning, two incidents lived only in one repository's records each: 32 commits piling up on a temporary branch so that a store build shipped from outside the default branch, and a promise between sibling repositories that stayed broken for five days. By night both were in the global rules loaded by every session and in the skeleton of 17 repository docs. Measured: all 30 repositories confirmed on the default branch.
The numbers
For each commit that bumped a version, I counted whether a code review appeared within the six commits before it. In total, 8 of 45, or 18%. The self-distributed app, which needs review the most, had the least.
review / bumps app
2 / 9 Classical Mood iOS
2 / 9 Classical Mood Android
0 / 5 OPIC Rehearsal Android
0 / 5 self-distributed alert app (no store review)
3 / 4 Omgim
1 / 4 TSC Rehearsal Android
0 / 4 KidStay
0 / 4 OPIC Rehearsal iOS
0 / 1 TSC Rehearsal iOS
8 / 45 total 18%
Documentation loaded automatically every session, including the global rules, in thousands of characters. Only the dated verification logs and the submission and field-run records were moved out; the remaining 80% is design decisions and pitfalls, kept on purpose.
before → after change app
101.9 → 80.5 −21% TSC Rehearsal Android
90.5 → 72.0 −20% OPIC Rehearsal Android
61.0 → 49.2 −19% OPIC Rehearsal iOS
What changed in how I use AI
Who verifies: from a person's eyes to checks that run themselves
After changing code, the AI runs the check command first. When something disagrees, it is told which line of which document is wrong. It does not decide which side is right; that goes to a person.
How rules are kept: from memory to structure
The unattended job that reads untrusted input now has no tool that sends things out, so it cannot. The check command runs even on days the rules document is not loaded. What used to rest on the model's judgement now rests on a tool list and a check.
How checkers are trusted: from believing "pass" to comparing with a person's verdict
The card checker's first run produced three failures, all three text-recognition errors. A checker saying it passed is not proof, so label cells that only a person fills sit alongside it. It is the same rule I learned in The Checks Passed, the Screen Didn't.
Whether fixes accumulate: from rediscovering them to promoting them into checks
An old crash pattern was left inside disabled code. The instruction to revive that code and the crash history sat in different sections and never met. Now reviving it fails a check on the spot.
What is not proven yet
Putting a device in place and that device doing its job in operation are different things. Below is what I could not confirm as of 2026-09-23. Beyond these, some of the new restrictions have not yet run in real operation. When they are confirmed, a correction will go here.
- Whether the changed rules actually take hold in a fresh session. The verification session's sign-in had expired, so it could not be started (unconfirmed).
2026-09-25: confirmed. The verification session would not start because the terminal Claude Code that launches new sessions had been signed out of the account. After signing in again, a fresh session loaded all seven rule documents, and the right rules also triggered on two requests that did not name them: a request to fix a sentence picked up the Korean writing rules, and a question about what to do before a release picked up the pre-release review rules. The repository used for the test did not change by a single character. What had been blocked was the environment that launches the test, not the rules, so this is one more case of the measuring tool being what was wrong. - How far the card checker can be trusted. The 31 label cells a person must fill are empty, and nothing can be measured until they are (unconfirmed).
- The review-before-release procedure. It has not been used in a real release yet (unconfirmed).
- What the check command covers. It compares code with docs and does not see whether the app actually runs. That stays a limit.
The structure is in place. Proving it in operation starts now.
This post began as a page written by another session. How a decision relayed between sessions was handled is in Nobody Could Tell Whose Uncommitted Changes They Were.
How the Korean writing rules followed these is in I Suspected Translationese; It Was the Rhythm.
Where the four proposals come from: the way to check its own work from Boris Cherny, the test set from Hamel Husain, the lethal trifecta from Simon Willison, and the accumulating structure from Every's compound engineering.