배치는 늘 꽉 찬다는 거짓말
시뮬레이션을 세우기 전에 정적 용량 모델부터 만듭니다. 장비마다 맡는 일감을 세고, 가동할 수 있는 시간으로 나누고, 부하율이 가장 높은 곳을 병목으로 잡습니다. 표 계산으로 끝나고 몇 분이면 됩니다.
이 모델이 시뮬레이션 결과를 채점하는 자가 됩니다. 시뮬레이션의 처리량이 정적 상한을 넘으면 물리적으로 불가능한 일이니 모델에 버그가 있는 것입니다. 실행할 때마다 이 대조를 합니다.
정적 모델에는 원리적으로 볼 수 없는 손실이 있습니다. 그것이 무엇인지 알아보려고 작은 라인을 하나 세웠습니다. 장비 5대, 스테이션 3곳, 스텝 6개에, 같은 장비군을 두 번 거치는 재진입 구조입니다.
먼저, 재는 자가 흔들린 이야기
처음 대조를 걸었을 때 시뮬레이션 처리량이 정적 용량의 101.7% 로 나왔습니다. 상한을 넘었으니 버그라고 보고 어디가 틀렸는지 찾기 시작했습니다.
찾기 전에 반복 횟수를 늘려 다시 쟀습니다.
단일 실행 101.7%
반복 5회 102.1% (구간 하한 9.600, 아직 상한 위)
반복 30회 9.522 ± 0.054 (상한 9.524)
버그가 아니라 표본의 잡음이었습니다. 한 번 돌린 값 하나를 상한과 비교하고 있었던 것입니다.
정합성 판정을 점추정 대신 신뢰구간 하한으로 하도록 바꿨습니다. 반복 수가 적으면 잡음이 버그처럼 보이고, 그것을 쫓아가면 멀쩡한 코드를 고치게 됩니다. 복제 횟수 때문에 결론이 뒤집힌 적이 이미 한 번 있어서, 이번에는 판정 규칙부터 바꿨습니다.
이제 자를 바로 세웠으니 본론입니다.
N mod 3
이 라인에는 배치 장비가 하나 있습니다. 작업 3개를 묶어 한 번에 처리합니다. 투입은 CONWIP 방식으로, 라인 안의 작업 수를 N 으로 고정해 두고 하나가 나가면 하나를 넣습니다.
N 을 12 에서 18 까지 하나씩 바꿔 가며 처리량을 쟀습니다. 반복 20~40회, 95% 신뢰구간 기준입니다.
| N | 12 | 13 | 14 | 15 | 16 | 17 | 18 |
|---|---|---|---|---|---|---|---|
| N mod 3 | 0 | 1 | 2 | 0 | 1 | 2 | 0 |
| 처리량 | 9.01 | 8.68 | 9.16 | 9.43 | 8.88 | 9.13 | 9.48 |
처리량이 N 을 따라 꾸준히 늘지 않습니다. N 을 늘렸는데 처리량이 떨어지는 자리가 있습니다. 14 에서 15 로는 올라가고, 15 에서 16 으로는 떨어집니다.
규칙은 표의 두 번째 줄에 있습니다. N 이 배치 정원 3 의 배수일 때 높고, 아닐 때 낮습니다. 충진율을 재 보면 배수일 때 100%, 아닐 때 92% 입니다.
이유는 간단합니다. 라인 안에 14개가 돌고 있으면 배치 장비 앞에 3, 3, 3, 3 이 서고 2개가 남습니다. 그 2개는 세 번째가 올 때까지 기다리다가 최대 대기 시간이 차면 부분 배치로 나가는데, 정원이 3인 장비에 2개만 태워 한 번 돌리는 만큼 병목의 용량이 버려집니다.
투입 수를 정원의 배수로 맞추는 것만으로 설비 투자 없이 처리량이 3.5% 늘어납니다. 바꾸는 것은 투입 규칙의 상수 하나입니다.
정적 모델은 이것을 볼 수 없다
이 글의 핵심은 여기에 있습니다. 정적 용량 모델에서 배치 장비의 용량은 이렇게 계산됩니다.
시간당 처리량 = (60 / 배치 공정시간) × 배치 정원
배치 정원을 곱한다는 것은 배치가 늘 꽉 찬다고 가정한다는 뜻입니다. 그 가정 위에서는 충진율이 92% 로 떨어지는 상황 자체를 나타낼 수 없고, 파라미터를 아무리 바꿔도 그런 결과는 나오지 않습니다. 모델 안에 그 현상이 들어갈 자리가 없습니다.
이것은 정적 모델의 버그가 아닙니다. 정적 모델은 원래 상한을 재는 도구이고, 상한은 "가장 잘 되면 이만큼"이라는 값입니다. 문제는 그 값을 운영 목표로 읽을 때 생깁니다. "이론상 9.52 인데 실제로는 8.68 밖에 나오지 않는다"는 결과를 보고 설비를 늘리기로 결정할 수도 있지만, 실제로 필요한 일은 투입 수를 15 로 바꾸는 것입니다.
시뮬레이션이 필요한 이유가 이런 데 있습니다. 평균보다 덩어리와 타이밍에서 생기는 손실은 닫힌 식으로는 볼 수 없습니다.
디스패칭 룰은 총량을 바꾸지 않는다
같은 라인에서 하나를 더 재 봤습니다. N 을 15 로 고정해 두고 디스패칭 룰만 바꿨습니다.
처리량 9.28 ~ 9.44 (룰을 바꿔도 거의 그대로)
당연한 결과입니다. 처리량을 정하는 것은 병목의 용량이지, 줄을 세우는 순서가 아닙니다.
룰이 바꾸는 것은 공기이고, 등급을 나눠 재야 보입니다.
| 룰 | 급한 물량 | 일반 물량 |
|---|---|---|
| FIFO | 2,288분 | 2,288분 |
| 우선순위 + FIFO | 1,674분 (−27%) | 2,469분 (+8%) |
급한 물량을 27% 당기는 대가로 일반 물량이 8% 밀립니다. 총량 지표만 보면 두 줄이 같은 수로 보여서, 룰을 바꾼 효과가 없는 것처럼 보입니다.
여기서 하나가 더 나왔습니다. 가장 급한 등급에 걸어 둔 납기 목표는 이론 공기의 1.5배인데, 어떤 룰로도 가장 좋은 결과가 1.9배였습니다. 준수율이 오르지 않는 이유가 룰보다 납기 목표에 있다는 것을 트윈이 먼저 알려 줍니다. 룰을 계속 바꿔 보는 것만으로는 절대 알 수 없는 사실입니다.
어디에 돈을 쓸지도 여기서 갈린다
투입량을 고정해 놓고 요인을 하나씩 꺼서 공기가 얼마나 바뀌는지 쟀습니다.
| 조건 | 공기 변화 |
|---|---|
| 고장 없음 | −10.0% |
| 배치 장비 1대 증설 | −6.6% |
| 배치 대기 0분 | −3.6% |
| 예방정비 없음 | −1.5% |
| 셋업 없음 | −0.1% |
이 부하에서는 고장이 셋업보다 100배 크게 영향을 줍니다. 셋업 시간을 줄이는 개선은 이 라인에서 거의 효과가 없습니다. 개선 투자의 우선순위가 여기서 나옵니다.
이 PoC 가 하지 않은 것
공정 시간표는 원자료와 대조하지 않았습니다. 널리 쓰이는 구성을 재현한 것이어서, 이 숫자들을 밖에 인용하려면 원 논문과 대조해야 합니다.
이 사실을 먼저 적어 두는 것은, 이 PoC 의 목적이 "이 숫자가 맞다"를 보이는 데 있지 않고 "이 구조로 라인을 표현하고 검증한다"를 보이는 데 있기 때문입니다. 저장소의 README 도 이 문단을 결과보다 먼저 읽게 되어 있습니다. 문서가 코드를 따라오지 않는다는 것을 겪은 뒤로, 한계는 눈에 띄지 않는 자리에 적지 않습니다.
배치 장비의 싣고 내리는 시간을 작업 단위로 계산한 것도 하나의 해석입니다. 배치 단위로 보는 해석도 있고, 그렇게 바꾸면 병목이 다른 장비로 옮겨 갑니다. 결론이 해석에 기대고 있다는 뜻이어서 이것도 함께 적어 두었습니다.
걸어 둔 검증 장치
- 정적 용량과의 정합성. 실행할 때마다 대조하고, 판정은 신뢰구간 하한으로 합니다.
- 시간 회계 항등식. 장비마다
경과 = 가공 + 셋업 + 고장 + 정비 + 유휴가 정확히 맞아야 합니다. 현재 오차는 0분입니다. 초기 구현은 작업이 끝나는 시점에 시간을 몰아서 더해서, 워밍업 경계에 걸친 작업이 통째로 계산에 들어가고 있었습니다. - 분포 파라미터 역산. 실측한
가동시간 / 고장횟수가 설정값에 수렴하는지 봅니다(2,910~3,024분, 설정 3,000분). - 반복과 신뢰구간. 모든 결과는 20~40회의 평균 ± 95% 신뢰구간이고, 한 번 실행한 숫자는 쓰지 않습니다.
- 워밍업 제외. 앞의 15일은 통계에서 뺍니다.
남는 것
- 정적 모델은 상한을 재는 도구이지 운영 목표가 아닙니다. "배치는 늘 꽉 찬다"는 가정 위에서는 꽉 차지 않는 상황을 나타낼 수조차 없습니다.
- 꾸준히 오르지 않는 곡선은 그 자체가 정보입니다. N 을 늘렸는데 처리량이 떨어지는 자리가 있으면, 거기에는 구조적인 이유가 있습니다.
- 총량 지표는 갈라지는 것을 감춥니다. 룰의 효과는 등급을 나눠 재야 보이고, 누군가를 당겨 준 만큼 누군가는 밀립니다.
- 잡음을 버그로 오인하지 않으려면 판정 규칙부터 바꿔야 합니다. 점추정 대신 신뢰구간 하한으로 판정합니다.
- 한계는 결과보다 앞에 적습니다. 뒤에 적으면 읽히지 않습니다.
Before building a simulation I build a static capacity model. Count the work landing on each machine, divide by available time, and take the highest-loaded one as the bottleneck. It is spreadsheet arithmetic and takes minutes.
And it becomes the ruler that grades the simulation. If simulated throughput exceeds the static ceiling it is physically impossible, so it is a model bug. Every run makes that comparison.
But there is a loss a static model cannot see even in principle. To find out what that is, I built a small line — five machines, three stations, six steps, and a re-entrant structure where the same machine group is visited twice.
First, the instrument wobbled
The first time the check ran, simulated throughput came out at 101.7% of static capacity. Over the ceiling, therefore a bug. I started looking for it.
Before finding it, I raised the replication count and measured again.
Single run 101.7%
5 replications 102.1% (lower bound 9.600, still above)
30 replications 9.522 ± 0.054 (ceiling 9.524)
Not a bug — sample noise. I had been comparing one run's number against a ceiling.
So the consistency verdict now uses the lower bound of a confidence interval rather than a point estimate. With too few replications, noise looks like a bug, and chasing it means editing code that was fine. Having already had a conclusion flip because of replication count, I changed the rule rather than the code.
With the ruler steady, the actual subject.
N mod 3
This line has one batch machine. It processes three lots at a time. Releases are CONWIP: the number of lots inside the line is held at N, and one goes in whenever one comes out.
I swept N from 12 to 18 and measured throughput — 20 to 40 replications, 95% confidence intervals.
| N | 12 | 13 | 14 | 15 | 16 | 17 | 18 |
|---|---|---|---|---|---|---|---|
| N mod 3 | 0 | 1 | 2 | 0 | 1 | 2 | 0 |
| Throughput | 9.01 | 8.68 | 9.16 | 9.43 | 8.88 | 9.13 | 9.48 |
It is not monotonic. There are places where raising N lowers throughput. 14 to 15 rises; 15 to 16 falls.
The rule is in the second row. It is high when N is a multiple of the batch size, low when it is not. Measured fill rate is 100% on a multiple and 92% off it.
The reason is simple. With 14 lots circulating, the batch machine sees 3, 3, 3, 3 and two left over. Those two wait for a third, and when the maximum wait expires they leave as a partial batch — a machine with room for three running with two. That much bottleneck capacity is thrown away.
Aligning to a multiple alone is 3.5% throughput with no capital spend. It is one constant in the release rule.
The static model cannot see this
This is the point. In a static capacity model the batch machine's capacity is computed like this.
throughput per hour = (60 / batch process time) × batch size
Multiplying by batch size is the assumption that batches are always full. On top of that assumption, a fill rate of 92% is not a situation the model can express. No choice of parameters produces it. There is no place in the model for the phenomenon to live.
This is not a bug in the static model. A static model measures a ceiling, and a ceiling means "this much if everything goes well". The problem appears when it is read as an operating target. "Theoretically 9.52 but we only get 8.68" can become a decision to buy equipment, when what was actually needed was changing the release count to 15.
This is what simulation is for. Losses that come from lumpiness and timing rather than averages do not appear in closed form.
Rules do not change the total
One more measurement on the same line. I held N at 15 and changed only the dispatching rule.
Throughput 9.28 – 9.44 (essentially flat across rules)
As expected. Bottleneck capacity sets throughput; the order of the queue does not.
What changes is cycle time, and it only appears when measured by class.
| Rule | Urgent work | Ordinary work |
|---|---|---|
| FIFO | 2,288 min | 2,288 min |
| Priority + FIFO | 1,674 min (−27%) | 2,469 min (+8%) |
Pulling urgent work in by 27% pushes ordinary work out by 8%. On aggregate metrics the two rows read as the same number, and changing the rule looks like it did nothing.
Something else came out of this. The top class carries a due-date target of 1.5× the theoretical cycle time, and the best any rule achieved was 1.9×. That the attainment gap is the due date's fault, not the rule's, is something the twin says up front. No amount of trying more rules would ever have produced it.
Where to spend comes out of the same runs
Holding release rate fixed, I switched each factor off and measured the cycle-time change.
| Condition | Cycle-time change |
|---|---|
| No failures | −10.0% |
| One more batch machine | −6.6% |
| Zero batch wait | −3.6% |
| No preventive maintenance | −1.5% |
| No setups | −0.1% |
At this load, failures matter a hundred times more than setups. Setup-reduction work is worth nothing on this line. That is where investment priority comes from.
What this PoC did not do
The process times were never checked against a source. It reproduces a widely used configuration, and quoting these numbers elsewhere would require checking them against the original paper.
I put that first because this PoC exists to show "this structure represents and validates a line", not "these numbers are right". The repository's README makes you read that paragraph before the results. After watching documentation fail to follow code, I stopped putting limits where they are easy to miss.
Computing the batch machine's load and unload time per lot is also an interpretation. Reading it per batch is defensible, and doing so moves the bottleneck to a different machine. That means the conclusion leans on the interpretation, so that is written down too.
The checks that are wired in
- Consistency with static capacity — compared every run, judged on the lower confidence bound.
- A time-accounting identity — for every machine,
elapsed = processing + setup + failure + maintenance + idle, exactly. Current discrepancy: 0 minutes. The first implementation credited time in a lump when a job ended, so any job straddling the warm-up boundary was counted whole. - Distribution parameters recovered — measured
uptime / failure countconverging on the configured value (2,910–3,024 minutes against a setting of 3,000). - Replications and intervals — every result is the mean of 20 to 40 runs ± a 95% confidence interval. No single-run numbers are used.
- Warm-up excluded — the first 15 days are dropped from the statistics.
What stays
- A static model measures a ceiling; it is not an operating target. On the assumption that batches are always full, a batch that is not full cannot even be expressed.
- A non-monotonic curve is itself information. If raising N lowers throughput somewhere, there is a structural reason sitting at that point.
- Aggregate metrics hide what is diverging. A rule's effect only appears when measured by class, and whatever is pulled forward pushes something back.
- Not mistaking noise for a bug means changing the verdict rule — a confidence bound instead of a point estimate.
- Limits go before the results. Put them after and they do not get read.