설비는 노는데 재공은 쌓였다
우선처리 물량을 전체의 몇 % 로 둘지 계산하는 모델을 만들어 두고 있었습니다. 닫힌 식으로 한 번 풀고, 같은 조건을 이산사건 시뮬레이션으로 다시 돌려 둘이 맞는지 확인하는 구조입니다. 두 방법의 답은 서로 맞았습니다.
이번에는 그 모델의 답을 실제로 운영되는 라인에 대어 봤습니다. 결론부터 적으면, 틀린 것은 입력값보다 모델의 전제였고, 그중 하나는 이 모델의 대표 결론이 기대고 있는 전제였습니다.
합격 기준부터 바꿔야 했다
모델끼리 맞춰 볼 때의 합격 기준은 비용 곡선의 최소점이 일치하는가였습니다. 이 모델이 내놓는 것은 시간 예측보다 "몇 % 로 할 것인가"라는 결정이어서, 예측이 얼마나 정확한지보다 두 방법이 같은 결정을 내리는지를 봤습니다.
실제 라인에는 그 기준을 쓸 수 없었습니다. 최소점은 관측되지 않습니다. 다른 비중으로 운영해 본 적이 없으니 "그렇게 했다면"을 비교할 자료가 없고, 결정을 다루는 대상은 이렇게 실제로 택한 쪽의 결과만 관측됩니다.
판정할 수 있는 범위부터 좁혔습니다. 현재 운영 조건을 그대로 넣었을 때 출력이 실측과 맞는지는 볼 수 있고, 값을 조금 움직였을 때의 방향과 크기도 과거에 바꿔 본 이력이 있으면 볼 수 있습니다. 최소점이 맞는지는 볼 수 없습니다.
이 기준을 미리 정해 두지 않으면, 격차를 본 뒤에 "이 정도면 맞은 셈"인지 "틀린 것"인지를 정하게 됩니다. 그러면 어느 쪽으로든 정할 수 있습니다.
낙관적인 쪽으로 틀렸다
모델은 사이클 타임을 실제보다 짧게 예측했고, 실측은 훨씬 길었습니다.
틀린 방향이 중요합니다. 사이클 타임을 짧게 보면 우선처리를 늘렸을 때의 피해도 작게 보고, 따라서 권고값이 실제 최적보다 높게 나옵니다. 낙관적으로 틀린 모델은 그냥 부정확한 데 그치지 않고 위험한 쪽으로 부정확합니다.
이 모델에는 같은 양상의 이력이 이미 있었습니다. 평균은 몇 % 이내로 맞는데 분산 항이 빠져서 납기 관련 지표를 부풀려 예측했고, 그 결과 최적해가 두 자릿수에서 0 으로 뒤집힌 일입니다. 그때 붙인 보정 계수의 기본값을 이론값으로 두지 않은 이유가 그것입니다. 이번 결과는 그 보정으로도 부족하다는 뜻이었습니다.
참값은 숫자가 아니라 모양이었다
여기에 이르는 길은 곧지 않았습니다. 처음에는 대기를 늘릴 만한 항을 찾았습니다. 우선처리가 배치 구성을 깨뜨리는지, 흐름 도중에 등급이 바뀌는지, 고장이 서로 얽혀 있는지, 재진입 구간에 상관관계가 있는지를 봤습니다. 가장 유력하게 꼽았던 배치 깨짐은 이미 모델에 들어 있었고 영향까지 계산돼 있었습니다. 그 지점에서 방향을 바꿨습니다. 후보들이 하나같이 흐름이 보존되고 부하가 정상상태라는 틀 안에서 항을 더하는 것이었기 때문입니다. 항을 더해도 설명되지 않으면 의심할 것은 틀입니다. 항을 찾는 일을 멈추고 전제를 확인했고, 아래는 거기서 나온 결과입니다.
원인을 찾느라 두 번 돌렸는데, 두 결과가 참값을 사이에 둔 범위의 양 끝이었습니다.
완주 물량만 → 전 구간 부하 과소 → 시간 과소예측
투입 전부 → 하류 부하 과대 → 가동률 1 초과, 계산 불가
이 라인은 흐름이 보존되지 않습니다. 들어간 것이 모두 나오지 않습니다. 재작업이 있고, 평가하느라 소모되는 것이 있고, 중간에 빠지는 것이 있습니다. 그러면 앞 공정과 뒤 공정의 부하가 달라지고, 공정을 따라 그린 부하의 모양이 직사각형 대신 사다리꼴이 됩니다.
모델은 앞 공정과 뒤 공정을 똑같이 봤기 때문에, 어떤 단일 투입률을 넣어도 맞지 않습니다. 낮게 넣으면 모든 구간이 가볍고, 맞게 넣으면 뒤 공정이 무거워져 발산합니다. 참값이 숫자 하나로 정해지지 않고 모양으로 주어져서, 입력 숫자를 바꾸는 것으로는 고칠 수 없었습니다.
이것은 실수로 빠뜨린 것도 아니었습니다. 모델의 입력 데이터 명세가 끝까지 가지 못하는 물량을 빼도록 지시하고 있었으니 설계였습니다. 그 물량도 설비를 실제로 쓴다는 사실만 반영되지 않았을 뿐입니다.
확증처럼 보이는 숫자를 그대로 받아들이지 않았다
발산했을 때의 초과 부하가 사다리꼴 가설의 증거처럼 보였는데, 계산해 보니 설명이 되지 않았습니다.
감쇠율이 r 이면 직사각형 대비 초과 부하의 상한은 1/(1−r) 입니다. 끝까지 가는 물량으로 직사각형을 깔면 라인 초입에서 그 값이 나오는데, 초입에서는 나중에 빠져나갈 물량까지 모두 들고 있기 때문입니다. 뒤 공정으로 갈수록 1 에 가까워지므로 구간 평균은 상한보다 작아야 하는데, 관측된 초과 부하는 그 상한을 넘어 있었습니다.
사다리꼴을 모두 반영해도 그 숫자는 나오지 않습니다. 가설을 지지하는 것처럼 보이던 값이 실은 다른 오류가 남아 있다는 신호였습니다. 자기 가설에 유리한 숫자가 나왔을 때 계산을 한 번 더 해 보는 습관이 이 작업에서 가장 쓸모 있었습니다.
설비는 노는데 재공은 쌓였다
이 모델의 대표 결론은 이것이었습니다.
우선순위는 대기시간을 줄이지 못하고 재분배할 뿐이다.
대기행렬 이론의 보존 법칙에서 나오는 결론이고, 문서에도 첫 번째로 적어 두었습니다. 그 법칙에는 작업 보존이라는 전제가 붙습니다. 일감이 있으면 설비를 놀리지 않는다는 전제입니다.
실제 라인에는 그렇지 않은 상태가 있었습니다. 재공도 있고 설비도 정상인데, 품질 기준을 맞추지 못해 진행이 막히는 경우입니다. 설비는 노는데 재공은 쌓이고, 대기열이 비어 있지 않은데 설비가 쉬고 있으니 작업 보존이 성립하지 않습니다.
그러면 보존 법칙이 보장하던 관계가 무너지고, 대표 결론은 근거를 잃습니다. 파라미터가 틀린 것이 아니라, 정리를 적용할 조건이 성립하지 않는 것입니다.
여기서 실무적으로 더 큰 것이 나왔습니다.
모델이 풀던 문제 총 대기가 고정이니 그 고통을 누구에게 배분할까
실제로 있는 레버 버려지는 용량을 되찾아 고통의 총량을 줄일 수 있다
재분배와 총량 감축은 다릅니다. 뒤의 것이 더 큰 수단인데, 모델은 그런 수단이 있다는 것조차 모릅니다. 이 파라미터를 두고 논쟁하기 전에, 진행이 막혀 설비가 쉬는 시간이 가용 시간의 몇 % 인지부터 재야 했습니다.
덤으로 알게 된 것이 하나 더 있습니다. 설비 상태를 분류하는 표준은 설비가 돌 수 있느냐를 기준으로 시간을 나눕니다. 이 상태는 설비는 돌 수 있는데 물량을 넣지 못하는 경우라 표준에 해당하는 축이 없고, 어느 칸에 넣을지가 사업장마다 다릅니다. 어느 칸에 넣었느냐에 따라 모델에 넣을 용량이 정반대가 됩니다. 가동 가능 시간에 들어 있다면 모델이 그 시간을 여유로 착각하니 용량을 깎아 넣어야 하고, 비가동으로 잡혀 있다면 이미 빠져 있으니 또 반영하면 두 번 세는 셈입니다. 같은 현상인데 분류 방식이 답을 뒤집습니다.
병목은 부하의 모양으로 예측되지 않는다
부하의 모양이 사다리꼴이면 앞 공정이 무거우니 병목도 앞쪽으로 옮겨 간다고 추론했습니다. 그 추론은 틀렸습니다.
숨은 가정이 있었습니다. 용량이 공정 경로를 따라 대체로 고르게 배치돼 있다는 가정입니다. 실제로는 그렇지 않습니다. 용량은 부하의 모양을 알고 나서 배치된 것입니다. 앞 공정이 무겁다는 것을 여러 해 동안 알았으니 거기에 투자했고, 용량 배치는 이미 그 모양에 맞춰져 있었습니다.
용량이 이렇게 부하에 맞춰 정해진다면, 병목을 결정하는 것은 부하의 모양보다 투자하고 남은 차이입니다. 실제로 병목이 되는 조건들은 대개 모델 밖에 있었습니다. 설비가 비싸거나, 설치 면적이 커서 기존 설비를 옮기거나 철거해야 들일 수 있거나, 공정 자체가 어려워 진행이 막히는 경우입니다.
모델 문서에는 이미 "병목 후보만 정확히 넣으면 충분합니다"라고 적혀 있었습니다. 설계는 처음부터 그랬는데, 중간에 제 추론이 틀렸습니다.
상수 하나로 메울 수 없는 것
또 하나의 전제는 정상상태였습니다. 해석 모델은 평균 부하 하나를 받는데, 이 라인의 부하는 몇 달 단위로 오르내립니다.
사이클 타임은 가동률에 대해 아래로 볼록한 함수이고, 가동률이 1 에 가까워지면 급격히 커집니다. 그러면 다음 부등식이 늘 성립합니다.
평균 부하로 계산한 시간 < 실제 평균 시간
격차는 부하의 변동 폭이 클수록, 평균 가동률이 1 에 가까울수록 커집니다. 정상상태 모델에 평균 부하를 넣으면 구조적으로 짧게 예측하고, 파라미터를 아무리 잘 맞춰도 이 차이는 없어지지 않습니다.
보정 계수가 근본적인 해결책이 될 수 없는 이유가 여기서 설명됩니다. 상수 하나로는 국면에 따라 달라지는 항을 흡수할 수 없어서, 한 국면에 맞추면 다른 국면에서 틀립니다.
증거는 제 문서 안에 있었습니다. 그 계수의 실측 적합값이 값 하나 대신 범위로 적혀 있었습니다. 상황에 따라 움직이는 값이라는 뜻이었는데, 저는 그것을 상수로 쓰고 있었습니다.
그렇다면 이 모델을 어떻게 쓰나
찾아낸 편향은 예외 없이 모두 같은 방향이었고, 반대쪽으로 상쇄하는 항은 하나도 없었습니다. 흐름이 보존되지 않는 것, 부하의 모양, 정상상태가 아닌 것, 진행이 막혀 쉬는 시간이 빠진 것, 뒤 공정의 배치 충진율이 떨어지는 것이 모두 권고값을 올리는 쪽입니다.
그러면 답이 하나 나옵니다.
쓰지 말 것 "적정값은 X 다"
쓸 것 "적정값은 X 를 넘지 않는다"
한쪽으로만 치우친 편향이 있는 추정은, 한쪽 경계로는 여전히 쓸 수 있습니다. 상한을 알면 그 아래에서 고를 수 있으니, 모델을 버릴 필요 없이 점 추정으로만 쓰지 않으면 됩니다.
쓸 수 있는 답을 잃은 것은 없고, 답을 쓰는 방법이 바뀌었습니다.
판정하지 못한 것
합격인지 불합격인지는 판정하지 못했습니다. 허용 오차를 실측 전에 정해 두지 않았기 때문입니다. 이것은 나중에 되살릴 수 없는 유일한 항목이었습니다. 예측 자체는 모델이 결정론적이라 같은 입력으로 다시 돌리면 되살아나지만, 판정 기준은 결과를 보고 나서 만들면 이미 오염됩니다. 이번에 얻은 것은 판정 대신 "벌어진 폭"입니다.
발산했을 때의 초과 부하 가운데 감쇠만으로 설명되지 않는 부분의 원인은 아직 확정하지 못했습니다. 계수의 단위와 실물 단위가 맞지 않는 회계상의 문제가 유력하다고 보지만 확인하지 못했습니다.
진행이 막혀 설비가 쉬는 시간의 실제 비율은 재지 못했고, 제품과 설비와 공정마다 달라서 값 하나로 정할 수도 없습니다. 값을 정하는 대신, 그 비율을 범위 안에서 바꿔 가며 답이 뒤집히는지만 보기로 했습니다. 뒤집히지 않는다면 그 불확실성은 이 결정과 무관하고, 그 판정 자체가 결과가 됩니다.
정리하면
- 결정을 다루는 대상은 실제로 택한 쪽의 결과만 관측됩니다. 최적점은 볼 수 없으니, 합격 기준을 최적점 일치에서 현재 운영 조건의 재현으로 바꿔야 했습니다.
- 판정 기준만은 나중에 되살릴 수 없습니다. 예측은 모델이 결정론적이면 다시 돌려 되살릴 수 있지만, 허용 오차를 나중에 정하면 이미 오염됩니다.
- 자기 가설을 지지하는 듯한 숫자가 나오면 계산을 한 번 더 합니다. 상한을 넘은 값은 확증이기는커녕 다른 오류의 신호였습니다.
- 정리를 인용할 때는 전제도 함께 확인합니다. 작업 보존이 깨진 라인에서는 우선순위에 관한 보존 관계가 보장되지 않습니다.
- 재분배와 총량 감축은 다릅니다. 모델이 앞의 것만 다룰 수 있다면, 더 큰 수단은 모델 밖에 있을 수 있습니다.
- 부하가 무거운 곳이 병목이라는 추론은 용량이 부하와 상관없이 주어질 때만 성립합니다. 이미 부하를 알고 용량을 배치한 조직에서 병목은 용량을 늘리지 못한 이유에서 생기고, 그 이유는 대개 모델 밖에 있습니다.
- 보정 계수의 적합값이 값 하나 대신 범위로 적혀 있다면, 상수로 쓸 수 없다는 신호입니다.
- 편향이 모두 한 방향이면 한쪽 경계로 쓸 수 있습니다. 점 추정으로만 쓰지 않으면 됩니다.
한 바퀴를 돌고 얻은 것은 권고값보다 모델을 어디까지 쓸 수 있는지에 대한 답이었습니다. 처음 물은 질문의 답은 아니지만, 그 질문에 답하려면 먼저 알아야 했던 것들입니다. 모르는 채로 답을 냈다면 그 답은 틀린 채로 쓰였을 것입니다.
이 모델의 보수적인 기본값이 왜 중립이 아닌지는 중립인 줄 알았던 기본값에, 같은 모델을 재다가 재는 자를 의심하게 된 이야기는 흔들린 건 곡선이 아니라 재는 자였다에 있습니다. 기준으로 삼은 수가 틀려서 모델이 어긋나 보였던 이야기는 4년째 자라고 있는 자로 쟀다에 적어 두었습니다.
I had a model that works out what share of a line should run as priority work. It solves the problem in closed form, then runs the same conditions through a discrete-event simulation to check that answer. The two methods agreed with each other.
This time I took the model's answer to a line that is actually running. To put the conclusion first: what was wrong was not an input value. It was the model's premises — and one of them was what the model's headline conclusion stands on.
The pass criterion had to change first
When checking one method against the other, the criterion was whether the cost curve's minimum lands in the same place. The output is not a time prediction but a decision — what share to run — so decision equivalence mattered more than predictive accuracy.
Against a real line that criterion is unavailable. The minimum is not observable. The line has never run at another value, so there is no counterfactual. A decision-shaped target is only ever observed on one arm.
So the first job was narrowing what could be judged at all. Whether the model reproduces today's operating point can be checked. Whether it gets the direction and size right for a small move can be checked, if there is a history of such moves. Whether the minimum is correct cannot.
Settle this in advance or you end up deciding, after seeing the gap, whether it counts as close enough — and then it can be decided either way.
Wrong in the optimistic direction
The model underpredicted cycle time. The measurement was far longer.
The direction matters. Underestimate cycle time and you also underestimate the damage from raising priority work, which means the recommendation comes out higher than the real optimum. A model that is wrong optimistically is not merely inaccurate; it is inaccurate in the dangerous direction.
This model already had one episode with the same signature. The mean was right to within a few percent, but a variance term was missing, so a due-date metric was overpredicted and the optimum flipped from double digits to zero. That is why the correction coefficient added afterwards does not default to its theoretical value. This result meant even that correction was not enough.
The true answer was a shape, not a number
Getting here was not a straight line. The first search was for terms that would add waiting — priority work breaking up batch formation, grades reassigned mid-flow, failures that are not independent of one another, correlation across re-entrant visits. But the leading candidate, batch disruption, was already in the model, with its effect quantified. That is where the direction changed. Every candidate was an addition inside a frame that assumed conserved flow and a steady-state load. When adding terms will not close the gap, the frame is what to doubt. So the search for terms stopped and the premises got checked instead — and what follows came out of that.
Chasing the cause took two runs, and the two results were the ends of a bracket.
only what finishes → load too low → time underpredicted
everything entering → downstream too heavy → utilization past 1
This line does not conserve flow. What goes in does not all come out. There is rework, there is material consumed for evaluation, and there is work that leaves partway through. So upstream and downstream carry different loads, and the load profile is not a rectangle but a taper.
The model treated upstream and downstream as equal. So no single release rate fits — set it low and the whole route is too light; set it right and the downstream end diverges. The true answer was a shape, so changing an input number could not fix it.
And this was not an accidental omission. The model's input data specification instructs you to exclude the work that does not finish. It was by design. What went unrepresented was that this work still occupies machines.
I did not accept a number that looked like confirmation
The excess load at the point of divergence looked like evidence for the taper. Then the arithmetic refused to agree.
If the attrition rate is r, the excess load over a rectangle has an upper bound of 1/(1−r). Against a rectangle drawn from the completing volume, that value occurs at the head of the line — where everything that will later drop out is still being carried — and it falls toward 1 downstream, so the route average is below the bound. The observed excess was above that bound.
No amount of taper produces that figure. So a number that appeared to support the hypothesis was in fact a signal that another error remained. Doing the arithmetic once more precisely when the number favours your own hypothesis was the most valuable habit in this whole exercise.
The machines idled while the queue grew
This model's headline conclusion was this.
Priority cannot reduce waiting time; it only redistributes it.
It follows from a conservation law in queueing theory, and I wrote it down as finding number one. But that law comes with a precondition — work conservation. A machine is never left idle while there is work for it.
The real line has a state that violates it. Work is waiting, the machine is healthy, and the process is blocked because a quality criterion is not met. The machine idles while work piles up. The queue is not empty and the server is idle. That is not work-conserving.
Then the invariant is not guaranteed, and the headline conclusion loses its footing. No parameter is wrong; the theorem's conditions of application do not hold.
And something larger came out of that, practically.
what the model solved total waiting is fixed —
so whom do we hand the pain to
the lever that exists reclaim wasted capacity —
the total itself shrinks
Redistribution and reduction are not the same thing. The second is the bigger lever, and the model does not know it exists. Before arguing about this parameter, the thing to measure was what share of available time that blocked idle accounts for.
One more thing came with it. The equipment-state standard divides time by whether the machine is able to run. This state is machine able, work unable to enter — an axis the standard does not have. So which bucket it lands in varies, and where it lands flips the model's capacity input: counted inside available time, the model mistakes it for slack and capacity must be reduced by that share; counted as downtime, it is already excluded and adding it again double-counts. Same phenomenon, and the classification decides the answer.
A bottleneck is not predicted by the load profile
If the load profile tapers, upstream is heavier, so the bottleneck moves upstream — that was the reasoning. It was wrong.
It carried a hidden assumption: that capacity is roughly uniform along the route. It is not. Capacity was placed after the load profile was known. Years of experience had established that upstream is heavy, so that is where the investment went. It was already optimized against that shape.
When capacity is endogenous, the load profile does not determine the bottleneck. The residual after capacity investment does. And the conditions that actually create a bottleneck were mostly outside the model — the machine is expensive, or its footprint is large enough that installing it means moving or removing something else, or the process itself is hard enough to block.
And the model's own documentation already said so — "only the bottleneck candidates need to be accurate." The design had it right from the start; the reasoning went wrong in between.
What no single constant can absorb
Another premise was stationarity. The analytic model takes one average load. This line's load rises and falls over spans of months.
Cycle time is convex in utilization and explodes as utilization approaches 1. Which means
time computed from average load < actual average time
holds always. The gap widens the more the load varies and the closer average utilization sits to 1. Feed an average load to a steady-state model and it underpredicts structurally — no amount of parameter fitting removes it.
This is also why a correction coefficient is not a real fix. One constant cannot absorb a term that changes with the regime. Fit it to one regime and it is wrong in another.
The evidence was inside my own documentation. The coefficient's fitted value was written as a range, not a number. That meant it moves with the situation — and I was using it as a constant.
So how does one use this model
Every bias I found pointed the same way, without exception. Nothing offset anything. Non-conserved flow, the load profile's shape, the stationarity violation, the unrepresented blocked idle, the degraded downstream batch fill — all of them push the recommendation up.
Which gives one answer.
do not write "the right value is X"
write "the right value is no higher than X"
A bias that runs one way is still valid as a one-sided estimate. Knowing the ceiling lets you choose below it. The model does not need to be thrown away; it needs to stop being used as a point estimate.
It is not that no value came out. How the value is used changed.
What went unjudged
There is no pass or fail. The tolerance was not set before the measurement was taken, and that turned out to be the one item that cannot be restored after the fact. The prediction itself survives — the model is deterministic, so the same inputs regenerate it — but a criterion invented later is already contaminated. So the output of this round is not a verdict; it is a gap.
The part of the excess load that attrition alone cannot explain is undetermined. An accounting mismatch between the counting unit and the physical unit is my leading suspicion, but it is unverified.
The actual share of blocked idle is unmeasured. It differs by product, by machine, by process, and cannot be pinned to one number. So rather than fixing a value, the plan is to sweep it as a sensitivity axis and see only whether the answer flips. If it does not, that uncertainty is irrelevant to this decision — and that verdict is itself a result.
In short
- A decision-shaped target is observed on one arm only. The optimum is invisible, so the pass criterion had to move from matching the optimum to reproducing the current operating point.
- Only the criterion cannot be restored after the fact. A deterministic model's prediction comes back on re-run; a tolerance decided later is already compromised.
- When a number favours your own hypothesis, do the arithmetic once more. A value above the bound was not confirmation but a signal of another error.
- Cite a theorem and check its preconditions too. On a line where work conservation breaks, the priority invariant is not guaranteed.
- Redistribution and reduction are different. If a model can only express the first, the bigger lever may sit outside it.
- "The heaviest load is the bottleneck" holds only when capacity is exogenous. Where an organization already placed capacity knowing the load, the bottleneck comes from why capacity could not be added — and that reason usually lives outside the model.
- If a correction coefficient's fitted value is written as a range rather than a number, that is a sign it cannot be used as a constant.
- When every bias runs one way, use the result as a one-sided estimate. Just do not use it as a point.
What one round produced was not a recommendation but a domain of validity. That is not the answer to the question I started with, but it is what had to be known before answering it. Answer it without knowing, and the answer gets used while wrong.
Why this model's conservative defaults are not neutral is in The Defaults Were Not Neutral, and measuring the same model until the instrument itself fell under suspicion is in It Was the Instrument Shaking, Not the Curve. A case where the baseline was wrong and made the model look off is in I Measured With a Ruler That Was Still Growing.