설정을 적지 않은 수는 한 달 만에 되짚을 수 없었다
8월 26일 실측 탭에 수 하나를 적었습니다. 소규모 라인 트윈의 처리량이 한 번 돌렸을 때는 정적 용량 상한의 101% 로 나와 버그처럼 보였지만, 반복 30회로 다시 재니 99.6%(9.485 ± 0.105, 상한 9.52)였다는 기록입니다. 배치는 늘 꽉 찬다는 거짓말에도 같은 표를 실었습니다.
한 달 뒤인 9월 23일, 그 트윈의 저장소 문서를 새로 쓰면서 사이트가 인용한 수를 원본과 하나씩 맞춰 봤습니다. 이 저장소를 인용하는 실측은 세 건이었고, 그중 두 건은 수를 확인할 파일이 없었습니다.
화면에만 찍히던 수
배치 정합 결과와 디스패칭 룰의 처리량은 저장소의 CSV 와 그대로 맞았습니다. 등급별 공기와 용량 대조 결과는 사정이 달랐습니다. 실험 코드가 두 표를 화면에만 찍고 파일로 남기지 않았기 때문에, 사이트가 공개해 둔 수 가운데 둘은 근거 파일 없이 서 있었던 셈입니다.
실험 코드를 고쳐 두 표를 CSV 로 남기게 하고 다시 돌렸더니, 등급별 공기는 사이트의 수와 정확히 맞았습니다. 우선순위 룰에서 급한 물량 1,674분, 일반 물량 2,469분입니다. 한 달 전에 만든 결과 파일과 그림은 바이트까지 똑같이 다시 나왔습니다. 난수 씨앗을 고정해 둔 덕분입니다.
하나는 조금 넘쳐 있었습니다. 사이트는 "FIFO 는 모든 등급이 2,288분으로 같다"고 적었는데, 등급이 셋이고 가장 급한 등급은 2,301분이었습니다. 글의 표처럼 두 등급으로 좁혀 고쳤습니다.
세 수가 한 설정에서 나오지 않았다
용량 대조는 재현되지 않았습니다. 처음 떠오른 설정인 재공 18 로 반복 수를 바꿔 가며 돌렸습니다.
반복 처리량/일 반폭 상한 대비
1 9.489 0.000 99.6%
5 9.458 0.294 99.3%
10 9.520 0.145 100.0%
20 9.484 0.074 99.6%
30 9.480 0.054 99.5%
사이트의 9.485 와 99.6% 는 반복 20회 값과 맞았는데, 사이트는 30회라고 적었습니다. 반폭 0.105 는 어느 반복 수에서도 나오지 않았고, 이 설정의 1회는 101% 가 아닌 99.6% 였습니다. 다른 설정도 찾아봤습니다. 재공 19·20·22 는 1회 96~97%, 재공 24 는 102%, 투입 간격을 고정하는 방식은 85~99% 였습니다. 재공 21 의 1회가 101.7% 로 가장 가까웠지만, 그 설정의 30회는 9.522 로 사이트의 9.485 와 달랐습니다.
한 설정에서 세 수가 함께 나오는 곳은 끝내 찾지 못했습니다. 8월에 어느 설정으로 돌렸는지 어디에도 적어 두지 않았기 때문입니다. 수는 적었는데 그 수를 낸 조건은 적지 않았고, 한 달이 지나자 그 수를 되짚을 길이 없었습니다.
다시 재니 하나가 더 보였다
그 실측이 말하려던 것은 "한 번만 돌리면 상한을 넘어 보인다"입니다. 재공 18 은 1회도 상한 아래라 그 현상을 보여 주지 못합니다. 그 현상이 드러나도록 라인이 충분히 포화되는 재공 21 로 설정을 고정하고 다시 쟀습니다.
반복 처리량/일 반폭 구간 하한 상한 대비
1 9.689 0.000 9.689 101.7%
5 9.720 0.120 9.600 102.1%
10 9.562 0.134 9.428 100.4%
20 9.533 0.075 9.458 100.1%
30 9.522 0.054 9.468 100.0%
(상한 9.524)
5회까지는 신뢰구간 하한마저 상한 9.524 를 넘었습니다. 8월에 세운 규칙은 "판정은 점추정 대신 신뢰구간 하한으로 한다"였는데, 반복이 5회뿐이면 그 규칙으로 판정해도 버그라는 결론이 나옵니다. 하한이 상한 아래로 내려온 것은 10회부터였습니다. 규칙을 바꾼 것만으로는 충분하지 않았고, 규칙이 믿을 만해지는 반복 수까지 함께 정해야 했습니다.
사이트의 실측 줄과 글의 표는 이 결과로 바꿨습니다. 글에는 옛 수와 바꾼 이유를 정정 문구로 남겼고, 이 표는 이제 저장소에 CSV 로 남아 있습니다.
남의 일화가 제 것처럼 실려 있었다
같은 날 대규모 라인 트윈의 저장소 문서도 새로 썼는데, 그 README 에 똑같은 일화가 실려 있었습니다. "처리량이 정적 용량의 101% 로 나와 버그를 의심했으나 반복 30회로 재면 99.6% 였다." 그런데 그 저장소에는 시뮬레이션 처리량을 정적 용량과 대조하는 코드 자체가 없었습니다. 작은 트윈에서 겪은 일이 큰 트윈의 문서로 넘어가면서 출처가 지워진 것입니다. 그 문단은 수를 빼고, 형제 저장소의 일이라고 밝히는 정정 표시로 바꿨습니다.
같은 README 안에서 두 절이 서로 반대를 말하는 곳도 있었습니다. 한 절은 LLM 을 실제로 호출해 비용까지 쟀다고 적었고, 다른 절은 실제 호출로 검증하지 못했다고 적었습니다. 실행 기록을 확인하니 앞쪽이 맞았습니다. 문서는 커밋에 딸려오지 않는다에서 적은 일이 이번에는 한 문서 안에서 일어났습니다.
옮겼는지는 값을 바꿔 봐야 안다
그 저장소는 검증 수치를 한 파일에 모아 두고 문서 생성기가 거기서 읽도록 만들어 둔 상태였습니다. 그런데 발표자료와 보고서 생성기 다섯 곳에 "재공 오차 21.7% → 1.4%"가 손으로 박혀 있었습니다. 다섯 곳을 모두 그 파일에서 읽도록 옮긴 뒤, 고치기 전과 후의 코드로 문서 10개를 만들어 글자를 비교했습니다. 모두 같았습니다.
같다는 결과는 새 코드가 실제로 쓰였다는 증거가 되지 못합니다. 값이 같으니 글자가 같은 것은 당연하기 때문입니다. 확인하려고 그 파일의 두 값을 잠깐 −77.7 과 9.9 로 바꿔 다시 만들었고, 다섯 곳이 모두 가짜 값으로 바뀌는 것을 확인한 뒤 값을 되돌렸습니다. 그 저장소 README 는 함정 표 끝에 "변경 후 결과가 완전히 같으면 그건 안정성이 아니라 신호다"라고 적어 두었는데, 그 문장이 그대로 들어맞는 자리였습니다.
이 맥에서만 돌던 것
문서 생성에 쓰는 패키지 셋이 의존성 목록에 빠져 있었습니다. 이 맥의 가상환경에는 이미 깔려 있어서 아무도 몰랐습니다. 빈 가상환경을 새로 만들어 옛 목록만 깔아 보니, 문서 생성 모듈 셋이 각각 패키지가 없다며 import 단계에서 멈췄습니다. 목록을 고친 뒤 같은 방법으로 다시 확인했습니다. 되짚을 수 있어야 한다는 말은 수에만 해당하지 않았습니다. 그 수를 만드는 환경도 다른 곳에서 다시 세울 수 있어야 했습니다.
정리하면
- 밖에 싣는 수는 그 수를 낸 설정까지 파일로 남깁니다. 화면에만 찍힌 수는 한 달이면 근거를 잃습니다.
- 판정 규칙도 표본이 모자라면 속습니다. 신뢰구간 하한으로 판정하기로 해 놓고도, 5회에서는 하한마저 상한을 넘었습니다.
- 문서에 옮겨 적은 일화는 출처와 함께 옮깁니다. 출처가 빠지면 다른 저장소의 일이 제 일처럼 읽힙니다.
- 제대로 옮겼는지는 같은 결과로 알 수 없습니다. 값을 바꿔 보고 따라 바뀌는지 확인합니다.
- 재현은 제 기계 밖에서 확인합니다. 빈 환경에서 다시 세워 보기 전에는 의존성 목록을 믿지 않습니다.
재는 자를 먼저 의심한다는 규칙은 흔들린 건 곡선이 아니라 재는 자였다에서 생겼습니다. 이번에는 자가 틀린 것도 흔들린 것도 아니었고, 눈금을 어디에 댔는지를 적어 두지 않은 것이 문제였습니다.
On August 26 I put a number on the measurements page. Run once, the small line twin's throughput came out at 101% of its static capacity ceiling and looked like a bug; re-measured over 30 replications, it was 99.6% (9.485 ± 0.105 against a ceiling of 9.52). The same table went into The Lie That Batches Are Always Full.
A month later, on September 23, I was writing new documentation for that twin's repository and checked every number the site quotes from it against the original. Three measurements on the site cite that repository, and for two of them there was no file to check the numbers against.
Numbers that only reached the screen
The batch-alignment results and the dispatching throughput matched the repository's CSV files exactly. The per-class cycle times and the capacity comparison were another matter. The experiment code printed both tables to the screen and never saved them, so two of the numbers the site publishes had been standing without a source file.
I changed the experiment code to save both tables as CSV and ran it again. The per-class cycle times matched the site exactly: 1,674 minutes for urgent work and 2,469 for ordinary work under the priority rule. The month-old result file and chart came out identical to the byte, because the random seed is fixed.
One claim was slightly too broad. The site said "FIFO gives every class 2,288 minutes", but there are three classes and the most urgent one took 2,301. I narrowed it to the two classes the post's table shows.
No single setting gave all three numbers
The capacity comparison did not reproduce. I started with the setting that came to mind first, 18 lots in the line, and varied the number of replications.
reps throughput/day half-width vs ceiling
1 9.489 0.000 99.6%
5 9.458 0.294 99.3%
10 9.520 0.145 100.0%
20 9.484 0.074 99.6%
30 9.480 0.054 99.5%
The site's 9.485 and 99.6% matched the 20-replication row, yet the site said 30. A half-width of 0.105 appeared at no replication count, and a single run at this setting gave 99.6%, not 101%. I tried other settings too: 19, 20 and 22 lots gave 96 to 97% on a single run, 24 lots gave 102%, and releasing at a fixed interval gave 85 to 99%. A single run at 21 lots came closest at 101.7%, but that setting's 30-replication mean was 9.522, not the site's 9.485.
I never found a setting that produced all three numbers together, because nowhere had I written down which setting the August run used. The number was recorded; the conditions that produced it were not. A month later there was no way back to it.
Measuring again showed one more thing
What that measurement meant to show was that a single run can look as if it exceeds the ceiling. At 18 lots even a single run stays below it, so it cannot show that. I fixed the setting at 21 lots, where the line is saturated, and measured again.
reps throughput/day half-width lower bound vs ceiling
1 9.689 0.000 9.689 101.7%
5 9.720 0.120 9.600 102.1%
10 9.562 0.134 9.428 100.4%
20 9.533 0.075 9.458 100.1%
30 9.522 0.054 9.468 100.0%
(ceiling 9.524)
Up to 5 replications, even the lower bound of the confidence interval sat above the 9.524 ceiling. The rule I set in August was to judge by the lower bound rather than the point estimate, and with only 5 replications that rule still concludes there is a bug. The lower bound dropped below the ceiling only from 10 replications on. Changing the rule was not enough; the number of replications at which the rule becomes trustworthy had to be fixed along with it.
The measurement line and the post's table now carry these results. The post keeps the old numbers and the reason for the change in a correction note, and this table now lives in the repository as a CSV.
Someone else's story, told as its own
The same day I also wrote new documentation for the large line twin, and its README carried the very same story: throughput came out at 101% of static capacity, a bug was suspected, and 30 replications gave 99.6%. That repository has no code that compares simulated throughput with static capacity at all. Something that happened in the small twin had travelled into the large twin's documentation and lost its source on the way. I took the numbers out of that paragraph and marked it as a correction saying it belongs to the sibling repository.
The same README also had two sections contradicting each other. One said the LLM had been called for real, with costs measured; another said real calls had never been verified. The run records showed the first one was right. What I wrote in Docs Don't Ride Along With the Commit happened this time inside a single document.
Whether it moved is known only by changing the value
That repository keeps its validation figures in one file, and the document generators read from it. Yet the slide and report generators had "WIP error 21.7% → 1.4%" typed in by hand in five places. I moved all five to read from that file, then built ten documents with the old code and the new code and compared the text. It was identical.
Identical output is no evidence that the new code was used; with the same values, the text is bound to be the same. So I changed the two values in that file to −77.7 and 9.9 for a moment and rebuilt. All five places switched to the fake values, and then I put the values back. That repository's README ends its pitfalls table with "an identical result after a change is not stability but a signal", and this was exactly where it applied.
What only ran on this Mac
Three packages used for document generation were missing from the dependency list. This Mac's virtual environment already had them, so nobody noticed. I built a fresh virtual environment with only the old list installed, and the three document modules each stopped at import, missing a package. After fixing the list I checked the same way again. Being able to retrace a number turned out to cover more than the number: the environment that produces it has to be rebuildable somewhere else too.
In short
- A number that goes out gets saved to a file together with the settings that produced it. A number that only reached the screen loses its footing within a month.
- A verdict rule can be fooled too when the sample is small. Even judging by the lower bound, at 5 replications the lower bound was above the ceiling.
- A story copied into documentation travels with its source. Without it, another repository's experience reads as your own.
- Identical results do not show that something moved correctly. Change the value and see whether the output follows.
- Check reproduction outside your own machine. Do not trust a dependency list until it has been rebuilt in an empty environment.
The rule of suspecting the ruler first came from It Was the Instrument Shaking, Not the Curve. This time the ruler was neither wrong nor shaking; what was missing was a note of where it had been laid.