거짓 날짜도 날짜처럼 생겼다
어제 실측 탭을 만들면서, 다음에 이 파일을 여는 사람이 줄을 추가하기 직전에 읽도록 데이터 파일에 규칙을 적어 두었습니다.
새 줄을 넣기 전에: 그 수를 내가 재서 본 것인지,
아니면 어딘가에서 읽고 옮긴 것인지 먼저 확인합니다.
후자면 여기 자리가 없습니다.
오늘, 그 페이지의 날짜 다섯 개를 제가 지어냈다는 것을 발견했습니다.
왜 통과했나
이 페이지의 다른 칸들은 스스로를 지킵니다. 방법 칸이 비면 줄이 그려지지 않고, 가리키는 글을 찾지 못하면 출처 줄이 통째로 사라집니다. 어제 글에 그렇게 만든 이유까지 적어 두었습니다.
그런 장치가 날짜에만 없었고, 이번 일의 핵심이 여기에 있습니다.
날짜는 참이든 거짓이든 똑같이 날짜처럼 생겼습니다. 2026-08-14 는 제가 어디에서도 읽지 않고 그럴듯해서 적은 값인데, 화면에서는 실제 실행 기록에서 옮긴 다른 열두 개와 조금도 달라 보이지 않습니다. 숫자에는 어떻게 얻었는지를 물으면서 날짜에는 묻지 않았습니다.
수는 방법을 요구하면 막을 수 있습니다. 방법을 적지 못하면 수를 넣을 수 없기 때문입니다. 날짜는 방법이 필요 없어 보인다는 점이 함정이었습니다. 언제 쟀는지는 사실을 묻는 질문인데도 서식을 채우는 일처럼 느껴지고, 그렇게 열일곱 칸 가운데 다섯 칸이 그냥 채워졌습니다.
다시 잴 수 없는 것을 어떻게 잡나
발견한 뒤에 문제가 하나 더 있었습니다. 지나간 날은 다시 잴 수 없습니다. 성조 실측 자체는 다시 돌릴 수 있어도, 그것을 8월 14일에 했는지 8월 7일에 했는지는 이제 어디에서도 나오지 않습니다.
다시 잴 수 없으니 재는 대신 맞대어 보기로 했습니다. 기록 두 개를 나란히 놓고 서로 모순되는 자리를 찾는 방식입니다.
측정은 그 측정을 쓴 글보다 나중일 수 없다.
당연한 말이지만, 이 한 줄로 열일곱 건을 훑으니 여섯 쌍이 있을 수 없는 조합이었습니다.
comment-density 측정 08-14 글 08-07 불가능
what-actually-ships 측정 08-14 글 08-07 불가능
layout-combinations 측정 08-16 글 08-09 불가능
lot-model-replications 측정 08-17 글 08-16 불가능
lot-model-factors 측정 08-17 글 08-16 불가능
여섯 번째는 반대쪽이 틀린 경우였습니다. 「고쳤는데 스택이 그대로였다」가 8월 23일 자로 적혀 있었는데, 그 글이 말하는 빌드 2, 3, 4 는 각각 08-20, 08-22, 08-24 입니다. 08-23 은 그 이야기의 어느 사건과도 맞지 않아서, 이번에는 글 쪽 날짜를 고쳤습니다.
서로를 반증할 수 있는 두 기록을 나란히 두면, 어느 쪽도 혼자서는 드러내지 못하는 거짓이 드러납니다. 날짜 하나만 놓고 보면 언제까지나 그럴듯합니다.
검사를 넣었더니 멀쩡한 줄에서도 경고가 울렸다
손으로 한 번 훑고 끝낼 일이 아니어서 이 대조를 코드로 옮겼습니다. 첫 실행에서 경고가 둘 떴는데, 그중 하나는 진짜 오류가 아니었습니다.
카탈로그 항목은 8월 25일에 잰 것인데, 가리키는 글은 8월 21일 자입니다. 순서가 뒤집힌 것처럼 보이지만, 그날 그 글에 새 수치를 반영했습니다. 글이 먼저 있었고, 나중에 잰 수가 그 글에 실린 것입니다.
여기서 고르기 쉬운 방법은 규칙을 느슨하게 만드는 것입니다. "며칠까지는 봐준다" 같은 조건을 넣으면 경고는 사라지지만, 이 검사가 잡으려던 것도 함께 사라집니다. 8월 14일과 8월 7일은 일주일 차이여서, 그 정도 여유를 주면 제가 지어낸 다섯 건 가운데 몇 건은 그냥 통과합니다.
규칙을 건드리는 대신 예외를 두는 방식을 바꿨습니다. 예외를 두려면 글을 고친 날을 함께 적어야 합니다(postAmended). 검사를 끄는 스위치를 만든 것이 아니라 근거를 가져오게 한 것이어서, 왜 예외인지 모른 채 통과하는 줄은 남지 않습니다.
울리기만 하는 검사는 둘째 주부터 아무도 보지 않습니다. 검사를 넣는 일의 절반은 거짓 경보를 없애는 일이었습니다.
같은 실수가 몇 분 뒤에 또
고친 것을 배포하고 실제 사이트에 반영됐는지 확인했습니다. 방법은 간단했습니다. 배포된 파일에 2026-08-07 이 있으면 새 버전이라고 본 것입니다.
첫 시도에 "배포됨"이 떴는데, 그것은 옛 파일이었습니다.
2026-08-07 은 제가 고친 날짜이면서, 원래 어떤 글의 날짜이기도 합니다. 그 문자열은 고치기 전 파일에도 있었습니다. 확인 스크립트가 물은 것은 새 버전인지가 아니라 이 여덟 글자가 어딘가에 있는지였고, 답은 처음부터 "예"였습니다.
바이트 단위로 다시 비교하니 배포된 사이트는 아직 옛 버전이었습니다. 데이터에서 한 번, 그것을 확인하는 자리에서 또 한 번, 같은 날 두 번이나 모양이 맞는 것을 증거로 착각했습니다.
기계에 맡길 수 있는 것과 없는 것
어제 이 페이지에 관해 쓰면서 "그 압박을 막는 규칙은 코드 대신 사람에게 거는 것이어서, 눈에 보이는 자리에 적어야 합니다"로 맺었습니다. 하루 만에 그 절반이 틀렸다는 것을 알았습니다.
"측정은 글보다 늦을 수 없다"는 기계가 확인할 수 있습니다. 두 기록이 같은 파일에 있고, 둘의 관계가 한 줄로 적히기 때문입니다. 이제 그 페이지를 열 때마다 이 검사가 돕니다.
지금 열일곱 줄 가운데 일곱 줄에는 연결된 글이 없습니다. 저장소 기록에만 있고 그 저장소는 공개돼 있지 않으니, 그 일곱 줄의 날짜는 맞대어 볼 상대가 없습니다. 제가 지어낸 다섯 가운데 셋이 바로 그런 줄이었습니다.
이 글을 쓴 뒤 코드 리뷰에서, 그 일곱 줄에도 적용되는 약한 검사 두 가지를 더 넣었습니다. 날짜가 YYYY-MM-DD 형식인지, 그리고 오늘보다 뒤가 아닌지입니다. 앞의 검사는 대조가 조용히 뒤집히는 것을 막아 줍니다. 2026-8-7 은 화면에서는 거의 맞아 보이지만, 문자열로 비교하면 2026-12-01 보다 큽니다. 그래도 두 검사 모두 그 날짜에 근거가 있는지는 보지 못합니다. 제가 지어낸 다섯은 형식도 맞았고 미래도 아니었으니, 이 검사들이 있었어도 모두 통과했을 것입니다.
코드 주석에 검사의 한계도 함께 적었습니다. 검사가 있다는 사실이 "여기는 다 확인됐다"로 읽히면, 검사가 보지 않는 자리가 오히려 더 위험해집니다.
덧붙임: 하루 뒤에 반대 방향으로 또 틀렸다
이 글을 올린 다음 날, 같은 실수를 반대 방향으로 했습니다. 배포됐는지 확인하려고 받은 파일을 /tmp/x 에 쓰고 비교했는데, /tmp/x 는 이미 디렉터리였습니다. 파일은 하나도 써지지 않았고, 비교는 없는 파일을 상대로 돌았고, 확인 스크립트는 열 번 내리 "아직"이라고 답했습니다. 배포는 진작 끝나 있었습니다.
어제는 거짓 양성("배포됨"인데 옛 버전)이었고, 오늘은 거짓 음성("아직"인데 이미 배포됨)이었습니다. 방향은 반대이지만 모양은 같습니다.
두 확인 스크립트 모두 자기 전제를 한 번도 확인하지 않았습니다. 첫 번째는 "이 문자열이 있으면 새 버전"이라면서 바뀌기 전에도 그 문자열이 있었는지를 보지 않았고, 두 번째는 "파일이 다르면 아직"이라면서 파일이 써지기는 했는지를 보지 않았습니다. "예"와 "아니오"를 말하는 도구는 자기가 둘을 구별할 수 있다는 것부터 보여야 합니다. 그렇지 않으면 한쪽으로만 답하는 고장 난 계기와 겉보기에 똑같습니다.
이번에는 규칙을 적는 대신 도구를 만들었습니다(tools/check-deploy.py). 바이트 단위로 비교하고, 재지 못했으면 "다름"이라고 답하지 않습니다. 받기에 실패했거나 응답이 200 이 아니거나 본문이 비었으면 ? 로 보고하고 다른 종료 코드를 냅니다. "다름"과 "재지 못함"은 다른 사건이기 때문입니다.
처음 돌린 날 그 도구가 제 몫을 했습니다. 네 파일 모두 ? HTTP 403 이 떴는데, 파이썬 기본 User-Agent 가 막힌 것이었습니다. 옛 방식이었다면 "배포 안 됨"으로 읽었을 자리입니다. 고치고 나니 이번에는 글 한 편이 ? HTTP 308 이었습니다. .html 주소가 확장자 없는 주소로 넘어가는데, 파이썬 기본 처리기가 308 을 따라가지 않았습니다. 제 전역 규칙에 "308 을 따라가지 않아 빈 응답을 쟀다"고 적혀 있는 바로 그 함정이 새 도구에서 그대로 재현됐습니다. 지금은 리다이렉트를 따라가되, 최종 주소가 다르면 어디를 쟀는지를 결과 옆에 적습니다.
어제 이 글은 "규칙의 절반은 기계가 볼 수 있고, 나머지 일곱 줄은 여전히 제가 지켜야 합니다"로 끝났습니다. 제가 지키겠다던 쪽이 하루를 가지 못했습니다.
정리하면
- 형식이 맞는 값은 검사를 통과합니다. 날짜는 참이든 거짓이든 날짜처럼 생겨서, 근거를 묻지 않으면 아무 저항 없이 채워집니다.
- 어떤 칸은 "사실"보다 "서식"으로 느껴집니다. 수에는 방법을 물으면서 날짜에는 묻지 않은 이유가 그것입니다.
- 다시 잴 수 없으면 맞대어 봅니다. 서로를 반증할 수 있는 두 기록을 나란히 두면, 어느 쪽도 혼자서는 드러내지 못하는 거짓이 드러납니다.
- 거짓 경보를 남겨 두면 검사는 버려집니다. 규칙을 느슨하게 하지 말고, 예외가 근거를 가져오게 합니다.
- 확인하는 도구도 틀릴 수 있습니다. "이 문자열이 있으면 새 버전"은 그 문자열이 옛 버전에도 있으면 언제나 "예"라고 답합니다.
- "예"와 "아니오"를 말하는 도구는 자기가 둘을 구별할 수 있다는 것을 먼저 보여야 합니다. 재지 못했을 때 "아니오"라고 답하면, 한쪽으로만 답하는 고장 난 계기와 겉이 같아집니다.
- 검사가 있다는 사실이 안전하다는 증거는 아닙니다. 무엇을 보지 않는지도 함께 적습니다.
이 페이지를 만든 이야기는 빈칸을 인쇄했다에, 자동 검사가 물어본 것에만 답했던 사례는 검사는 통과했는데 화면이 틀렸다에 있습니다. 재는 도구 자체가 틀렸던 이야기는 흔들린 건 곡선이 아니라 재는 자였다에 적어 두었습니다.
Yesterday, while building the Measured tab, I wrote a rule into the data file.
Before adding a row: check whether you measured this
number yourself, or read it somewhere and copied it.
If the latter, it does not belong here.
And today I found that five of that page's dates were ones I had made up.
Why it passed
The page's other fields defend themselves. An empty method field means the line isn't drawn, and a write-up that can't be found makes the whole source line disappear. I wrote up the reasoning for that yesterday.
The date field had no gatekeeper at all. And here is the point —
A date looks like a date whether it is true or false. 2026-08-14 is a value I read nowhere and wrote because it seemed about right, and on screen it is indistinguishable from the twelve taken out of real run records. I asked the numbers how they were obtained. I never asked the dates.
Demanding a method defends a number: no method, no number. But a date doesn't seem to need one. "When was this measured" feels less like a claim of fact than like filling in a form. So five of seventeen fields simply got filled.
Catching what can't be re-measured
Then a second problem: a past day cannot be measured again. The tone probe can be re-run, but whether it happened on 14 August or 7 August is now recoverable from nowhere.
So instead of measuring, I collided.
A measurement cannot postdate the write-up that describes it.
Obvious — and running that one line across seventeen entries made six pairs impossible.
comment-density measured 08-14 post 08-07 impossible
what-actually-ships measured 08-14 post 08-07 impossible
layout-combinations measured 08-16 post 08-09 impossible
lot-model-replications measured 08-17 post 08-16 impossible
lot-model-factors measured 08-17 post 08-16 impossible
The sixth was wrong on the other side: The Stack Did Not Move was dated 23 August, while the builds it describes — two, three and four — landed on 08-20, 08-22 and 08-24. The 23rd matches no event in that story. There, the post's date was the one to fix.
Put two records that can refute each other side by side, and a falsehood neither could expose alone becomes visible. A date on its own stays plausible forever.
The check fired on a healthy row
Since this shouldn't depend on my remembering to look, the collision went into code. On its first run two warnings appeared, and one of them was not a real error.
The catalogue entry was measured on 25 August and points at a post dated the 21st. It looks reversed, but the new figures were folded into that post that day. The post existed first, and a later measurement went into it.
The easy move here is to loosen the rule — allow a few days' slack and the warning disappears. So does what the check was for. The gap between 14 August and 7 August is a week; any slack generous enough to quiet this would let some of my five straight through.
So the exception changed instead of the rule. Taking one now requires recording the day the post was amended (postAmended). Not a switch that turns the check off — a place where the exception has to bring its evidence. No row passes without saying why it may.
A check that only ever cries wolf goes unread from its second week. Half the work of adding one was removing its false alarm.
The same mistake, minutes later
I deployed the fix and checked that it had gone live. The method: if the live file contains 2026-08-07, it's the new version.
The first attempt said "deployed". It was the old file.
2026-08-07 is one of my corrected dates — and also, already, the date of a post. The string was in the file before the change too. What the probe asked was not "is this the new version" but "do these eight characters occur anywhere", and the answer had always been yes.
A byte-level comparison showed the live copy was still the old one. Once in the data, once in the thing checking the data — twice in a day, I mistook well-formedness for evidence.
What a machine can hold, and what it can't
Writing about this page yesterday, I closed with "the rule binds a person, not the code, so it has to be written where it will be seen". A day later, half of that turned out to be wrong.
"A measurement cannot postdate its write-up" is something a machine can see. Both records live in one file and the relation fits on a line. The check now runs every time the page opens.
But seven of the seventeen rows have no write-up. They exist only in repositories that aren't public, so their dates have nothing to be checked against. Three of the five I invented were exactly that kind.
After writing this, a code review added two weaker checks that do reach those seven: is the date YYYY-MM-DD, and is it no later than today. The first keeps the comparison above from silently inverting (2026-8-7 looks nearly right on screen and sorts above 2026-12-01 as a string). Neither, though, can see whether a date has any basis. My five were well-formed and safely in the past — they would have passed these too.
So the code comment states the check's blind spot alongside it. If the existence of a check reads as "everything here is verified", the places it cannot see get more dangerous, not less.
Postscript — the day after, in the other direction
The day after publishing this, I made the same mistake with the sign reversed. Checking whether a deploy had landed, I wrote each downloaded file to /tmp/x and compared — and /tmp/x was already a directory. Nothing was written, the comparison ran against files that did not exist, and the probe answered "not yet" ten times in a row. The deploy had finished long before.
Yesterday a false positive ("deployed" while the old file was live); today a false negative ("not yet" while it already was). Opposite signs, one shape.
Neither probe ever checked its own premise. The first said "this string means the new version" without asking whether the string was there before. The second said "files differ, so not yet" without asking whether a file had been written at all. A tool that answers yes or no has to first show it can tell the two apart. Without that it is indistinguishable from a broken instrument stuck on one answer.
So this time I built the tool instead of writing the rule down (tools/check-deploy.py). It compares bytes, and when it could not measure it does not say "different" — a failed fetch, a non-200, an empty body all report as ? with a distinct exit code. "Different" and "couldn't tell" are different events.
On its first run the tool earned its place. All four files came back ? HTTP 403 — Python's default User-Agent is blocked, and the old style would have read that as "not deployed". With that fixed, one post came back ? HTTP 308: .html redirects to an extensionless path and Python's default handler does not follow 308. The exact trap my own global notes describe — "curl didn't follow the 308 and measured an empty response" — reproduced in a brand-new tool. It now follows the redirect and prints where it actually measured when the final URL differs.
This post ended yesterday with "the check can see ten; the other seven are still mine to keep". My half did not last a day.
In short
- A well-formed value passes. A date looks like a date either way, so unless provenance is demanded it gets filled with no resistance at all.
- Some fields feel like formatting rather than fact. That is why the numbers were asked for a method and the dates never were.
- When you can't re-measure, collide. Two records able to refute each other expose what neither could alone.
- Leave a false alarm in and the check dies. Don't loosen the rule; make the exception bring evidence.
- The thing doing the checking is also a value. "This string means the new version" always says yes if the old version had it too.
- A tool that answers yes or no must first show it can tell them apart. Answering "no" when it could not measure makes it indistinguishable from an instrument stuck on one answer.
- A check existing is not proof of safety. Write down what it doesn't look at.
How this page was built is in Printing the Blank; a case of automated checks answering only what they were asked is in The Checks Passed and the Screen Was Wrong. The instrument itself being wrong is in It Was the Instrument Shaking, Not the Curve.