4년째 자라고 있는 자로 쟀다
장비 1,443대 규모의 생산 라인을 이산사건 시뮬레이션으로 세웠습니다. 만든 것이 맞는지 확인하려면 비교할 대상이 있어야 하는데, 다행히 원자료에 원저자가 상용 엔진으로 돌린 4년 치 실행 로그가 들어 있었습니다.
그 로그의 평균 공기를 기준으로 잡고 제 모델이 얼마나 어긋나는지 쟀습니다. 한쪽 데이터셋에서 +3.0% 가 나왔습니다. 3% 면 못 쓸 정도는 아니지만 설명은 되어야 하는 수여서, 모델을 뜯어보기 시작했습니다.
틀린 것은 모델이 아니었습니다.
자가 아직 자라고 있었다
그 값 하나만 보고 있다가, 문득 시간에 따라 어떻게 움직이는지가 궁금해졌습니다. 로그에 일별 스냅샷이 다 들어 있어서 변화 궤적을 그려 볼 수 있었습니다.
시점 30일 90일 365일 730일 1,460일
평균 공기 15.10 29.60 37.10 38.40 39.00
1,459일 가운데 값이 오른 날이 191일, 내린 날이 10일이었고, 나머지는 반올림하면 같았습니다. 꾸준히 오르기만 하는 값입니다. 4년이 지난 마지막 날에도 아직 오르고 있었습니다.
정상상태의 평균이라면 이렇게 움직이지 않습니다. 이 값은 첫날부터 쌓아 온 누적 평균이었습니다.
왜 낮게 나오는지도 분명합니다. 이런 모델은 빈 라인 대신 이미 작업이 들어차 있는 상태에서 출발합니다. 출발 시점에 라인 끝부분에 있던 작업들은 남은 공정이 몇 개 없어서 금방 빠져나가고, 그 짧은 공기들이 평균에 섞입니다. 4년이 지나도 그 영향이 다 사라지지 않을 만큼입니다.
제가 기준으로 삼은 값은 정상상태보다 낮게 치우친 수였습니다. 그 값을 기준으로 재면, 정상상태를 제대로 재현한 모델이 "너무 느리다"고 나옵니다.
기준을 다시 세운다
같은 로그에는 재공(라인 안에 들어 있는 작업 수)과 산출(하루에 나오는 작업 수)도 있습니다. 이 둘은 누적 평균과 달리 그 시점의 값이고, 공기까지 셋을 잇는 관계가 있습니다.
재공 = 산출 × 공기
리틀의 법칙(Little's law)입니다. 정상상태에서 성립하고 도착 분포나 서비스 분포를 가정하지 않아서, 로그의 재공과 산출로 공기를 거꾸로 계산했습니다.
| 로그가 적어 둔 값 | 거꾸로 계산한 값 | 차이 | |
|---|---|---|---|
| 데이터셋 A | 38.84일 | 39.94일 | +2.8% |
| 데이터셋 B | 36.90일 | 37.72일 | +2.2% |
그러자 제 모델의 오차가 이렇게 바뀌었습니다.
| 옛 기준 | 새 기준 | |
|---|---|---|
| 데이터셋 B 의 공기 오차 | +3.0% | +0.5% |
모델은 한 줄도 고치지 않았습니다.
이것을 잡지 못했다면 저는 +3.0% 를 보고했을 것이고, 그 2.5%p 를 없애려고 있지도 않은 버그를 찾아 모델을 뜯어고쳤을 것입니다. 고치다 보면 언젠가는 맞아떨어졌을 텐데, 틀린 기준에 맞아떨어졌을 것입니다. 맞는 모델을 틀린 기준에 맞춰 망가뜨리는 것으로 끝났을 것입니다.
이 사이트에 재는 자가 흔들린 이야기를 적은 적이 있습니다. 그때는 자가 흔들렸고, 이번에는 자가 자라고 있었습니다. 두 번 다 재는 대상은 멀쩡했습니다.
그다음에는 한쪽만 20% 어긋났다
기준을 고치고 나니 데이터셋 A 는 1% 안에 들어왔는데 B 만 크게 어긋났습니다. 처리율은 1% 이내로 맞는데 재공이 −21.7%, 공기가 −20.4% 로, 실제보다 덜 막히는 라인이었습니다.
한쪽만 어긋난다는 것은 코드가 통째로 틀렸다기보다 B 쪽의 특성을 놓쳤다는 뜻입니다. 가설 세 개를 세우고 하나씩 계측했습니다.
| 가설 | 어떻게 쟀나 | 결과 |
|---|---|---|
| 대기 자리 정원을 넘기는 것을 무시했나 | 동시 사용량 계측 | 최대 148/400 — 무관 |
| 묶음을 너무 쉽게 만드는가 | 부분 묶음·대기 시간 계측 | 부분 묶음 0건 — 무관 |
| 셋업 정책이 지나치게 효율적인가 | 두 정책을 대조 실행 | −16.5% → −16.0% — 무관 |
셋 다 기각됐습니다. 여기서 짐작을 멈추고, 원자료 압축 파일 안에 함께 들어 있던 16쪽짜리 문서를 열었습니다.
다섯 가지가 나왔습니다.
- 기준 디스패칭 룰이 데이터셋마다 다릅니다. 저는 둘 다 같은 룰로 돌리고 있었습니다.
- 임계 공정은 연속된 여러 층을 같은 장비로 처리해야 합니다. 데이터에 그 컬럼이 있었는데 통째로 무시하고 있었습니다.
- 장비 용량 2 는 독립된 두 자리가 아니라 2단 파이프라인입니다. 저는 처리 능력을 2배로 부풀려 보고 있었습니다.
- 납기가 원자료에 들어 있습니다. 저는 "투입 시각 + 이론 공기 × 계수"로 지어내서 쓰고 있었습니다.
- 스텝을 옮길 때마다 이송 시간이 붙습니다. 저는 이것이 있다는 사실 자체를 몰랐습니다.
넷은 데이터 컬럼을 눈으로 보고도 뜻을 몰라 무시한 것이고, 하나는 아예 있는지도 몰랐습니다.
마지막 것이 컸습니다. 작업 하나가 거치는 스텝 전환이 300~500번이어서 1.5~2.6일에 해당하는데, 이 수가 딱 들어맞는 자리가 있었습니다. 그 시점에 제 모델은 작업량 지표(총 공정 처리 건수)가 오차 0.0% 로 완벽했는데 공기만 2.31일 모자랐습니다. 할 일은 다 했는데 이동하며 기다린 시간만 빠져 있었던 것이고, 이송 시간의 가중 평균 기여가 2.07일이었습니다.
가설 셋을 세워 하나씩 계측하고 모두 기각하는 데 든 시간보다 16쪽을 읽는 데 든 시간이 훨씬 짧았습니다.
해석이 갈리는 자리에서 무엇을 근거로 고르는가
문서 내용을 다 구현한 뒤에도 한 지점이 남았습니다. 2단 파이프라인에서는 앞 작업이 첫 칸을 비운 뒤에야 다음 작업을 실을 수 있는데, 그 싣는 시간이 간격에 더해지는지를 문서가 말하지 않습니다. 해석이 셋으로 갈렸습니다.
세 해석을 모두 구현해서 두 데이터셋에 각각 적용했습니다.
| 해석 | 재공 오차 | 공기 오차 |
|---|---|---|
| A — 더하지 않는다 (문서의 글자 그대로) | −5.3% | −5.7% |
| B — 싣는 시간만 더한다 | −1.5% | −1.7% |
| C — 싣고 내리는 시간을 모두 더한다 | +0.4% | −0.4% |
B 를 골랐습니다. 무엇을 근거로 골랐는지가 중요합니다.
셋 가운데 하나는 어차피 가장 좋게 나오기 마련이어서, B 가 한쪽 오차를 5.5% 에서 3.1% 로 줄였다는 것은 근거가 되지 못합니다. 근거는 다른 쪽이 나빠지지 않았다는 것입니다. 다른 쪽은 1.2% 에서 1.4% 로 거의 움직이지 않았습니다. 한쪽 수치에 억지로 맞춘 것이었다면 다른 쪽이 나빠졌어야 합니다.
두 데이터셋을 처음부터 동시에 맞추기로 한 이유가 이것입니다. 하나만 맞추면 맞은 것인지 맞춘 것인지 구별할 방법이 없습니다.
세 해석 모두 설정 하나로 바꿔 돌릴 수 있게 남겨 두었습니다. 제가 고른 것을 다른 사람이 다시 재 볼 수 없다면, 그것은 근거라기보다 주장입니다.
덧붙임: 변하지 않은 것도 신호다
같은 프로젝트에서 "결과가 똑같다"가 두 번이나 버그의 신호였습니다.
한 번은 운영 정책을 공격적인 쪽과 보수적인 쪽으로 바꿔 가며 A/B 를 돌렸는데, 달성률과 산출과 공기와 재공이 소수점 마지막 자리까지 같게 나왔습니다. 승인 건수는 달라지는데 결과가 같다는 것은 승인된 조치가 아무 일도 하지 않는다는 뜻입니다. 실제로 매니저의 주력 수단이 코드 분기 하나 때문에 한 번도 실행되지 않고 있었습니다.
또 한 번은 기능 하나를 구현했는데 검증 수치가 비트 단위까지 이전과 같았습니다. 코드를 넣는 치환의 패턴이 실제 소스와 달라서 아무것도 바꾸지 못했는데, 문자열 치환은 대상을 찾지 못해도 예외를 내지 않고 조용히 지나갑니다.
자동 검사가 물어본 것에만 답한 이야기와 같은 종류입니다. 그때는 검사가 묻지 않아서 보지 못했고, 이번에는 아무것도 변하지 않은 것을 안정성으로 읽을 뻔했습니다. 변경 후 결과가 완전히 같다면, 그것은 안정성보다 문제가 있다는 신호입니다.
남은 것
지금 이 트윈의 최대 오차는 원저자 로그와 비교해 데이터셋 A 가 2.1%, B 가 1.4% 입니다. 남은 차이의 원인은 밝히지 못했습니다. 상용 엔진 내부의 대기열 처리와 난수 구현의 세부는 바깥에서 재현할 방법이 없습니다. 여기까지가 제가 아는 것이고, 그 너머는 모른다고 적어 둡니다.
남는 것
- 검증할 때 먼저 의심할 것은 대상보다 기준입니다. 기준이 틀리면 맞은 것도 틀렸다고 나오고, 그 차이를 없애려다 맞는 것을 망가뜨립니다.
- 값 하나보다 궤적을 봅니다. 39.00 하나만 보면 평범한 수인데, 다섯 시점을 늘어놓으면 그 값이 무엇인지 드러납니다.
- 데이터만 받고 딸린 문서를 읽지 않으면 조용히 틀립니다. 가설을 세워 기각하는 것보다 함께 온 16쪽을 읽는 편이 훨씬 빨랐습니다.
- 하나만 맞추면 맞은 것인지 맞춘 것인지 모릅니다. 해석을 고르는 근거는 좋아진 수치보다 나빠지지 않은 다른 쪽입니다.
- 변경 후 결과가 완전히 같다면 그것은 신호입니다.
재는 자를 먼저 의심하는 규칙이 하루에 열네 번 쓰인 날의 기록은 계측기를 놓은 날에 있습니다.
I built a production line of 1,443 machines as a discrete-event simulation. To know whether it was right I needed something to compare it against, and the source data happened to include the original author's run log from a commercial engine. Four years of it.
I took the average cycle time from that log as my baseline and measured how far my model drifted. One dataset came out at +3.0%. Three percent is not unusable, but it is the kind of number that has to be explained, so I started pulling the model apart.
The model was not the thing that was wrong.
The ruler was still growing
I had been staring at the single figure when it occurred to me to ask how it moved over time. The log carried daily snapshots, so the trajectory was there to plot.
Day 30 90 365 730 1,460
Avg cycle 15.10 29.60 37.10 38.40 39.00
Across 1,459 days the figure rose on 191 and fell on 10, with the rest identical after rounding. It is monotonic. And on the last day of the fourth year it is still climbing.
A steady-state average does not behave like that. This was a running mean from day one.
Why it sits low is equally clear. A model like this starts from a line that is already loaded, not an empty one. The work sitting near the end of the route at time zero has only a few steps left, leaves almost immediately, and those short cycle times mix into the average — deeply enough that four years does not wash them out.
So the value I had chosen as my baseline was biased below the steady state. Measured against it, a model that reproduces steady state correctly reads as too slow.
Rebuilding the baseline
The same log also carries WIP (how much work is inside the line) and output (how much leaves per day). Those are point-in-time values, not running means. And there is a relationship joining all three.
WIP = output × cycle time
Little's law. It holds in steady state and assumes nothing about arrival or service distributions. So I derived cycle time from the log's own WIP and output.
| As logged | Derived | Difference | |
|---|---|---|---|
| Dataset A | 38.84 days | 39.94 days | +2.8% |
| Dataset B | 36.90 days | 37.72 days | +2.2% |
And my model's error moved like this.
| Old baseline | New baseline | |
|---|---|---|
| Dataset B cycle-time error | +3.0% | +0.5% |
Not one line of the model changed.
Had I not caught this, I would have reported +3.0% and then gone hunting for a bug that did not exist, shaking the model to close a 2.5-point gap. And shake it long enough and it would eventually have landed — on the wrong baseline. I would have finished by fitting a correct model to an incorrect target.
I have written once before about an instrument that was shaking. That time the ruler wobbled; this time it was growing. Both times the thing being measured was fine.
Then one dataset was off by 20%
With the baseline fixed, dataset A came inside 1% while B stayed badly off. Throughput matched to within 1%, but WIP was −21.7% and cycle time −20.4% — a line less congested than the real one.
Only one side being wrong means the code is not broadly broken; something specific to B is missing. I formed three hypotheses and measured each.
| Hypothesis | How it was measured | Result |
|---|---|---|
| Ignoring a queue-capacity limit | Concurrent occupancy | Peaked at 148/400 — unrelated |
| Batch formation too permissive | Partial batches and waits | Zero partial batches — unrelated |
| Setup policy too efficient | Ran both policies | −16.5% → −16.0% — unrelated |
All three rejected. At that point I stopped guessing and opened the 16-page document that had shipped inside the same source archive.
It produced five findings.
- The baseline dispatching rule differs by dataset. I was running both with the same rule.
- Critical layers must be processed on the same machine across a run of consecutive steps. The column was in the data and I had ignored it entirely.
- "Machine capacity 2" is not two independent slots but a two-stage pipeline. I had been overstating capacity by a factor of two.
- Due dates are in the source data. I had been inventing them from release time and a theoretical-cycle multiplier.
- Every step transition carries a transport time. I did not know this existed.
Four were columns I had looked at without knowing what they meant. One I did not know about at all.
The last one was large. A lot goes through 300 to 500 step transitions, which comes to 1.5 to 2.6 days. And there was a place where that fit exactly. At the time, my model's workload metric — total processing operations — was accurate to 0.0% while cycle time fell short by 2.31 days. All the work was being done; only the waiting was missing, and transport contributed a weighted average of 2.07 days.
Forming three hypotheses, measuring each, and rejecting all of them cost far more time than reading sixteen pages.
What justifies a choice when readings diverge
Even after implementing everything the document described, one point remained. In a two-stage pipeline the next lot cannot be loaded until the previous one clears the first slot, and the document does not say whether that loading time is added to the interval. Three readings were possible.
I implemented all three and ran each against both datasets.
| Reading | WIP error | CT error |
|---|---|---|
| A — add nothing (the literal text) | −5.3% | −5.7% |
| B — add loading only | −1.5% | −1.7% |
| C — add loading and unloading | +0.4% | −0.4% |
I chose B. But the reason matters.
That B cut one dataset's error from 5.5% to 3.1% is not the reason. One of any three will come out best regardless. The reason is that the other side did not get worse — it moved from 1.2% to 1.4%, essentially nothing. Overfitting to one dataset would have damaged the other.
This is why both datasets were matched simultaneously from the start. Match one alone and there is no way to distinguish a fit from a fitting.
All three remain switchable by a single setting. If nobody else can re-measure what I chose, it is not evidence — it is an assertion.
A note — an unchanged result is also a signal
In the same project, "the result is identical" was a bug signal both times it appeared.
Once, running an A/B across aggressive and conservative operating policies, attainment, output, cycle time and WIP came out identical to the last decimal. The number of approvals differed while the outcome did not, which means the approved action did nothing. The manager's main lever had never once executed, because of a single stale branch in the code.
Another time I implemented a feature and the validation figures were bit-for-bit unchanged. The pattern in the string replacement that inserted the code did not match the actual source, so nothing was replaced — and a string replacement that finds nothing raises no exception. It passes quietly.
Same seam as the automated check that answered only what it was asked. That time the check did not ask; this time I nearly read "nothing changed" as stability. An unchanged result after a change is not stability — it is a signal.
What is left
As it stands, the twin's largest error against the reference log is 2.1% on dataset A and 1.4% on B. The cause of the remaining gap is unexplained. The queueing and random-number internals of a commercial engine cannot be reproduced from outside. This is as far as I know, and past that I write down that I do not know.
What stays
- In validation, suspect the baseline before the subject. A wrong baseline reports a correct thing as wrong, and closing that gap breaks what was right.
- Look at the trajectory, not the value. 39.00 on its own is an ordinary number; five points in a row tell you what it is.
- Taking the data without reading the documentation fails quietly. Reading the sixteen pages that came with it was far cheaper than forming and rejecting hypotheses.
- Match one and you cannot tell a fit from a fitting. The reason for choosing a reading is not the number that improved but the one that did not degrade.
- An identical result after a change is a signal.
The day the rule of suspecting the ruler first went to work fourteen times is written up in The Day the Gauges Went In.