← 제작기

흔들린 건 곡선이 아니라 재는 자였다

앞 글에서 우선처리 물량의 적정 비중을 계산하는 모델 이야기를 했습니다. 그 모델은 같은 질문을 닫힌 수식으로 계산한 뒤, 같은 조건을 시뮬레이션으로 돌려 다시 재는 두 가지 방법으로 풉니다. 둘이 맞으면 믿고, 어긋나면 어느 쪽이 틀렸는지부터 찾습니다.

이 구조에는 전제가 하나 있습니다. 두 방법이 서로 독립이어야 합니다. 시뮬레이션이 수식의 가정을 그대로 되풀이하면 검증이 되지 못하고 같은 말을 두 번 하는 셈입니다. 같은 말을 두 번 해 놓고 "일치한다"고 적는 것과 다르지 않습니다.

독립은 저절로 생기지 않는다

수식이 입력으로 받는 값을 시뮬레이션은 하나도 쓰지 않게 만들었습니다.

수식이 받는 입력          시뮬레이션에서는
─────────────────────────────────────────────
셋업 간 작업 수 N_s   →   안 씀. 품목 전환에서 창발
묶음 충진율 φ        →   안 씀. 강제 착수 규칙에서 창발
대기 근사식          →   안 씀. 실제 큐를 시간축에서 돌림
가동률 A            →   안 씀. 설비별 고장·수리를 실제로 발생
우선순위 실효도 η     →   안 씀. 배차 규칙의 결과로 창발

오른쪽 칸은 모두 "안 씀"입니다. 시뮬레이션에는 우선순위가 얼마나 잘 듣는지 알려 주는 숫자가 아예 없습니다. 작업을 어떤 순서로 집어 들지에 관한 규칙만 있고, 실효도는 그 규칙을 돌린 결과로 나옵니다.

이렇게 된 것은 일부러 그렇게 짰기 때문입니다. 시뮬레이터를 만드는 내내 "이 값은 수식에서 그냥 가져다 쓰면 편할 텐데" 하는 유혹이 계속 있었습니다. 가져다 쓰는 순간 두 답이 잘 맞기 시작합니다. 같은 값을 넣었으니 당연합니다. 잘 맞는 것 자체는 목적이 될 수 없습니다. 안 맞을 수도 있는 상태에서 맞아야 합니다.

검증 결과가 표본마다 뒤집혔다

독립까지 확보했는데 이상한 일이 생겼습니다. 같은 모델, 같은 설정인데 판정이 왔다 갔다 했습니다.

복제 5회  → 최적점 16%,  비용손실 21.97%   FAIL
복제 16회 → 최적점 12%,  비용손실  0.00%   PASS

모델은 그대로였고, 바뀐 것은 표본이었습니다. 이런 상태에서는 PASS 가 떠도 모델을 평가한 것이 못 되고, 표본 크기를 평가한 것이 됩니다. 운 좋게 통과한 것과 정말로 통과한 것을 구분할 수 없습니다.

원인 1: 난수 흐름이 하나였다

시뮬레이터가 도착 시각, 물량 구성, 공정 시간, 고장까지 모든 난수를 한 흐름에서 뽑아 쓰고 있었습니다.

이 상태에서 우선처리 비중을 바꾸면 등급을 추첨하는 횟수가 달라집니다. 그러면 그 뒤에 뽑히는 난수가 통째로 한 칸씩 밀립니다. 비중만 바꿨는데 공정 시간도, 고장 시점도, 도착 간격도 모두 다른 값이 나옵니다.

비중을 바꿔 가며 그린 곡선은 사실 매번 다른 세계를 그린 것이었습니다. 곡선이 오르내리는 데 진짜 신호와 난수가 밀린 효과가 섞여 있었고, 표본을 바꾸면 순서가 뒤집혔습니다.

고친 방법은 난수 흐름을 용도별로 나누는 것입니다. 도착, 물량 구성, 공정 시간, 고장에 각각 다른 흐름을 줍니다. 그러면 비중을 바꿔도 공정 시간과 고장은 같은 값을 유지하고 물량 구성만 달라집니다.

순위상관 0.57 → 0.98

같은 시뮬레이터, 같은 복제 수인데 곡선이 안정됐습니다. 흔들리고 있던 것은 곡선이 아니라 재는 자였습니다.

원인 2: 잡음의 폭을 판정에 넣지 않았다

두 번째 원인은 더 근본적입니다. 두 방법의 최적점이 같은지 볼 때, 비용 차이가 잡음보다 큰지는 보지 않고 있었습니다.

복제별 비용의 95% 신뢰구간 반폭을 계산해 판정에 넣었습니다. 비용 차이가 그 폭 안에 들면 "일치"나 "불일치" 대신 "통계적으로 구분되지 않음"으로 판정합니다.

검증 결과가 표본 크기에 따라 뒤집힌다면 아직 검증이 아닙니다. 잡음을 통제하거나, 적어도 그 크기를 잰 다음에 판정해야 합니다.

그렇게 만든 검증이 잡아낸 것

이 과정에서 오류 아홉 개를 찾았고, 모두 검증을 거치고서야 드러났습니다. 묶음 타임아웃의 부동소수점 비교 때문에 시뮬레이터가 무한 루프에 빠진 것처럼 뻔한 것도 있었지만, 무서운 것은 결과가 멀쩡해 보이는 쪽이었습니다.

가장 인상 깊었던 것은 소요 시간의 분산 식에서 항 하나가 빠져 있던 일입니다.

  • 소요 시간 평균은 3% 이내로 정확했습니다. 어디를 봐도 맞는 모델처럼 보였습니다.
  • 분산이 틀리니 납기 준수율이 틀렸습니다. 0.913 으로 예측했는데 실제는 0.809 였습니다.
  • 지연 페널티가 크다 보니, 그 차이가 최적해를 12% 에서 0% 로 뒤집었습니다.

모델은 오류 없이 그럴듯한 숫자를 냈는데, 그 숫자로 계산한 결론은 "우선처리를 아예 쓰지 말라"였습니다. 평균만 맞추는 모델은 위험합니다. 평균은 눈에 보이고 분산은 보이지 않는데, 결정을 뒤집는 쪽은 분산이었습니다.

비슷한 일이 하나 더 있었습니다. 강건 최적화에서 실행 불가능 벌점이 위험도 계산을 지배해 답이 늘 0 으로 나온 적이 있습니다. 원인은 원래 있던 위험과 우선처리가 새로 만든 위험을 나누지 않은 데 있었습니다. 비중이 0 일 때도 실행 가능 확률이 85% 인 라인이었는데, 나머지 15% 를 우선처리 탓으로 돌리면 무엇을 넣든 답은 0 이 됩니다. 원래 있던 위험을 빼지 않으면 모든 수단이 책임을 뒤집어씁니다.

맞다는 것을 어떻게 보여 줄까

여기까지는 "내가 이 결과를 믿을 수 있는가"의 문제였습니다. 남은 문제는 남에게 어떻게 보여 줄까입니다. 결과가 이상하다고 느낀 사람이 가장 먼저 열어 볼 곳이 있어야 합니다.

그런 곳이 없었습니다. 수식은 설명 문서 안에 수식 표기법으로만 들어 있어서, 렌더러 없이 열면 원문 기호가 그대로 보였습니다. 게다가 다품종 구조를 넣은 뒤에 생긴 수식들은 문서에 반영되지 않은 상태였습니다. 코드는 앞서 나갔는데 문서는 그대로였습니다. 이 사이트에 이미 같은 이야기로 글을 하나 써 두었는데, 다른 저장소에서 똑같은 일을 되풀이한 것입니다.

수식만 따로 문서 한 장으로 뺐습니다. 기호 정의부터 최적화 문제까지 식 28개입니다. 만들면서 세 가지를 정했습니다.

  • 각 식이 어느 코드에 대응하는지 옆에 적었습니다. 식과 구현은 언젠가 갈라지기 마련인데, 대응 관계가 적혀 있으면 갈라진 것을 알아볼 수는 있습니다.
  • 흔한 함정 네 가지를 식 옆에 붙였습니다. 공칭 시간과 유효 시간을 섞어 쓰는 것, 묶음 설비의 서비스 시간과 용량 시간을 혼동하는 것, 기준 크기에 정격을 넣는 것, 분산의 항 하나를 빠뜨리는 것입니다. 마지막은 제가 실제로 저지른 실수라 위에 적었습니다.
  • 외부 라이브러리 없이 CSS 만으로 조판했습니다. 수식 렌더링 라이브러리를 불러오면 인터넷이 없는 곳에서는 빈 페이지가 됩니다. 결과를 못 믿겠다는 사람이 열어 보는 문서가 그 사람의 자리에서 열리지 않으면 없는 것과 같습니다.

특히 세 번째가 중요했습니다. 문서를 어디서 여는지는 문서의 내용만큼 중요합니다.


정리하면

  • 두 방법으로 같은 답을 내는 검증은 두 방법이 독립일 때만 검증입니다. 독립은 저절로 생기지 않으니, 한쪽이 쓰는 값을 다른 쪽이 쓰지 않도록 일부러 짜야 합니다.
  • 목적은 잘 맞는 것이 아니라, 안 맞을 수도 있는 상태에서 맞는 것입니다.
  • 판정이 표본마다 뒤집히면, 그 판정은 모델을 평가한 것이 못 되고 표본 크기를 평가한 것입니다.
  • 난수를 한 흐름으로 쓰면 조건 하나만 바꿔도 그 뒤가 전부 밀립니다. 용도별로 나누면 비교하려던 것만 달라집니다. 순위상관은 0.57 에서 0.98 이 됐습니다.
  • "같다"와 "다르다" 사이에 "구분되지 않는다"는 판정이 필요합니다. 잡음의 폭을 판정 조건에 넣습니다.
  • 평균만 맞추는 모델은 위험합니다. 평균은 3% 이내로 맞았는데 분산 항 하나가 빠져 최적해가 12% 에서 0% 로 뒤집혔습니다.
  • 원래 있던 위험을 빼지 않으면, 검토하는 수단이 모든 위험을 뒤집어씁니다.
  • 결과를 못 믿겠다는 사람이 먼저 볼 곳을 만들어 둡니다. 그 문서가 그 사람의 자리에서 열리는지까지 챙기는 것이 문서의 일입니다.

여기서 생긴 "재는 자부터 의심한다"는 규칙이 한 달 뒤 하루 동안 열네 번 쓰인 기록은 계측기를 놓은 날에 있습니다.

The previous post was about a model that computes the right share of priority work on a production line. That model answers the same question two ways — once in closed form, once by simulating the same conditions. If they agree, I believe it; if not, the first job is finding which one is wrong.

This rests on one premise: the two methods must be independent. If the simulation simply replays the formulas' assumptions, it is not verification but tautology — saying the same thing twice and recording "they agree."

Independence does not happen by itself

So the simulation uses none of the values the formulas take as input.

formula input               in the simulation
──────────────────────────────────────────────────────
jobs between setups N_s  →  unused. emerges from product changeovers
batch fill φ             →  unused. emerges from forced-start rules
queueing approximations  →  unused. real queues run on a time axis
availability A           →  unused. per-machine failures actually occur
priority effectiveness η →  unused. emerges from the dispatch rule

Every entry on the right is "unused." The simulation contains no number telling it how well a priority marking works. It has only a rule for which job to pick up next, and effectiveness comes out as a result of running that rule.

This was deliberate, not automatic. While writing the simulator there was a constant pull toward "I could just reuse that value from the analytic side." The moment you do, the two answers start agreeing — because you fed them the same number. The goal is not agreement; it is agreement that was free to fail.

Then the verdict started flipping between samples

With independence secured, something odd appeared. Same model, same settings, and the verdict moved.

5 replications  → optimum 16%,  cost loss 21.97%   FAIL
16 replications → optimum 12%,  cost loss  0.00%   PASS

The model had not changed; the sample had. In that state a PASS does not evaluate the model, it evaluates the sample size — and you cannot tell a lucky pass from a real one.

Cause 1 — one random stream

The simulator drew every random number from a single stream: arrivals, work mix, process times, failures.

Now change the priority share. The number of class draws changes. Every random number after that point shifts by one. You altered one dial, and process times, failure moments, and arrival gaps all became different realisations.

The curve traced across shares was, in fact, tracing a different world each time. Real signal and stream displacement were mixed into its ups and downs, and changing the sample reordered them.

The fix is one stream per purpose — arrivals, work mix, process times, failures. Change the share now and process times and failures keep the same realisations; only the mix moves.

rank correlation 0.57 → 0.98

Same simulator, same replication count, stable curve. The curve had not been shaking. The instrument had.

Cause 2 — noise was not part of the verdict

The second is more fundamental. When asking "do the two methods find the same optimum," nothing checked whether the cost difference was larger than the noise.

So the half-width of the 95% confidence interval across replications now enters the verdict. If the cost difference falls inside it, the result is neither "agree" nor "disagree" but "statistically indistinguishable."

If a verification flips with sample size, it is not yet a verification. Control the noise, or at minimum quantify it, before judging.

What that verification caught

Nine real errors surfaced, all of them only through verification. Some were blunt — a floating-point comparison in a batch timeout that sent the simulator into an infinite loop. The frightening ones were those whose output looked fine.

The most striking: a term was missing from the variance of lead time.

  • The mean lead time was accurate to within 3%. By every visible measure, a correct model.
  • But with the variance wrong, the on-time rate was wrong — predicted 0.913, actual 0.809.
  • Lateness penalties being large, that gap flipped the optimum from 12% to 0%.

Nothing errored. It produced plausible numbers, and the conclusion drawn from them was "do not use priority work at all." A model that only matches the mean is dangerous. The mean is what you look at; the variance is what flipped the decision.

One more of the same family: in the robust optimisation, an infeasibility penalty dominated the risk measure and every answer came back 0. The cause was failing to separate pre-existing risk from risk the mechanism introduced. The line was only 85% feasible at zero priority share; charge that 15% to priority work and no amount of it can ever look good. Without subtracting the baseline, every mechanism is guilty.

Showing that it is right

All of that is "can I trust it." What remains is how anyone else can. Someone who finds a result suspicious needs a first place to look.

There wasn't one. The equations existed only as inline notation inside an explanatory document, and without a renderer you saw the raw source. Worse, the equations added for multi-product structure had never made it into the document at all. The code had moved on; the prose had not. I have already written a post on this site about exactly that, and then repeated it in another repository.

So the equations moved into a document of their own — 28 of them, from symbol definitions to the optimisation problem. Three decisions while building it:

  • Each equation names the code it corresponds to. Formula and implementation drifting apart is a matter of time, but with the mapping written down, the drift is at least visible.
  • Four common traps sit beside the equations. Mixing nominal and effective time; confusing a batch machine's service time with its capacity time; using the rated size as the reference; dropping a term from the variance. The last is the one I actually committed, above.
  • Typeset in CSS alone, with no external library. Pull in a maths renderer and the page is blank wherever there is no internet. A document meant for the sceptic is worthless if it will not open at the sceptic's desk.

The third especially. Where a document opens matters as much as what it says.


In short

  • Answering the same question two ways only verifies anything if the two are independent — and independence must be built by deliberately making one side not use what the other consumes.
  • The goal is not agreement, but agreement that was free to fail.
  • If the verdict flips with the sample, you evaluated the sample size, not the model.
  • One random stream means changing a single condition displaces everything after it. One stream per purpose leaves only the thing you meant to compare. Rank correlation 0.57 → 0.98.
  • Between "same" and "different" you need "indistinguishable." Put the noise width into the verdict.
  • A model that matches only the mean is dangerous. Mean accurate to 3%, one variance term missing, and the optimum went from 12% to 0%.
  • Without subtracting pre-existing risk, the mechanism under review is blamed for all of it.
  • Build the first place a sceptic will look — and making sure it opens at their desk is part of the document's job.

"Suspect the ruler first" came out of this post; a month later it went to work fourteen times in one day. That record is in The Day the Gauges Went In.