← 제작기

중립인 줄 알았던 기본값

생산 라인에서 우선처리 물량을 전체의 몇 % 로 두어야 적절한지 계산하는 모델을 만들었습니다. 대기행렬 이론으로 풀고, 같은 조건을 시뮬레이션으로 돌려 답을 확인하는 구조입니다.

모델에 들어가는 값은 대부분 공정 순서나 설비 대수처럼 이미 있는 자료에서 나오고, 대여섯 개만 실측하기 전에는 알 수 없습니다. 우선순위를 매겼을 때 실제로 얼마나 먼저 처리되는지, 묶음 설비가 평균 몇 개를 채워 도는지, 대기 시간 분포의 꼬리가 얼마나 두꺼운지 같은 값들입니다.

이런 칸에는 기본값을 넣게 되는데, 기본값은 대개 1.0 이나 정격값이나 이론값처럼 생겼습니다. 중립적으로 보이고, 어느 쪽 편도 들지 않는 숫자 같습니다.

정말 중립인지 재 봤습니다

재는 방법은 단순합니다. 보정을 마친 값으로 푼 답을 참값으로 두고, 파라미터를 하나씩 순진한 기본값으로 되돌리면서 답이 어디로 움직이는지 봅니다.

참값이 11.6% 인 라인에서 이렇게 나왔습니다.

분산 꼬리계수 = 1.0     → 21.0%   (+9.4%p)
우선순위 실효도 = 1.0    → 19.0%   (+7.3%p)
묶음 실현크기 = 정격      → 17.4%   (+5.8%p)
셋업 간 작업 수 = 8      → 16.2%   (+4.6%p)
도착 변동성 = 1.0       →  5.8%   (−5.8%p)   ← 유일하게 안전
────────────────────────────────────────
다섯 개 전부            → 40.0%   (+28.4%p, 3.4배)

다섯 가운데 넷이 같은 방향으로 밀었고, 핵심은 마지막 줄입니다. 하나씩 되돌렸을 때의 차이는 5~9%p 였는데, 전부 함께 되돌리니 서로 상쇄되지 않고 28.4%p 가 됐습니다. 11.6% 가 상한인 40% 까지 밀려 올라간 것입니다.

기본값은 중립이 아닙니다. 방향이 있고, 그 방향은 잴 수 있습니다.

왜 하필 같은 방향이었나

처음에는 우연이라고 생각했지만 그렇지 않았습니다. 이유는 하나로 모입니다.

순진한 기본값은 모두 "이상적인 경우"를 뜻합니다. 실효도 1.0 은 우선순위를 매기면 100% 그대로 통한다는 뜻이고, 묶음 정격은 언제나 꽉 채워서 돈다는 뜻이고, 꼬리 계수 1.0 은 분포가 이론상 가장 얌전하다는 뜻입니다.

이 모델이 재는 것은 우선처리라는 수단이 얼마나 잘 듣는가이고, 이상적인 조건은 언제나 그 수단을 좋아 보이게 만듭니다. 차이들이 서로 독립이 아니었던 이유가 여기에 있습니다. 파라미터는 달라도 같은 낙관을 공유하고 있었습니다.

편향이 서로 독립이면 상쇄되지만, 같은 뿌리에서 나오면 겹칩니다. 겹치는지 확인하는 방법은 하나씩이 아니라 전부 함께 되돌려 보는 것뿐입니다. 저는 개별 표를 먼저 만들고 마지막 줄을 나중에 더했는데, 그 줄이 없었다면 "많아야 9%p 정도 틀리겠구나"로 끝났을 것입니다. 실제로는 그 3.4배였습니다.

기본값을 옮겼습니다

템플릿의 기본값을 "모르면 우선처리에 불리한 쪽"으로 다시 잡았습니다. 실효도는 묶음 설비 0.30, 그 밖의 설비 0.90 이고, 묶음 실현 크기는 정격의 80%, 꼬리 계수는 2.5 입니다. 모두 실측 범위의 보수적인 끝입니다.

이 값들로 같은 라인을 풀면 7.8% 가 나옵니다. 참값 11.6% 보다 3.8%p 낮습니다. 여전히 틀렸지만, 덜 권하는 쪽으로 틀립니다.

틀릴 방향을 고르는 기준

어차피 틀릴 테니 아무 쪽이나 괜찮다는 말은 아닙니다. 이 문제는 손실이 비대칭입니다.

  • 적게 넣어서 틀렸을 때는 급한 몇 건이 늦습니다. 다음 주에 비중을 올리면 회복됩니다.
  • 많이 넣어서 틀렸을 때는 셋업이 잦아지고 묶음이 깨지면서 일반 물량의 소요 시간이 무너집니다. 비중을 되돌려도 쌓인 재공이 빠지기까지 몇 주가 걸립니다.

회복에 걸리는 시간이 다르면 두 오류의 크기도 같지 않습니다. 기본값은 되돌리기 쉬운 쪽으로 기울여야 합니다. 모델의 정확도보다 운영에 관한 문제여서, 코드와 문서 양쪽에 함께 적었습니다.

경고를 두 종류로 나눴습니다

입력을 점검하는 명령이 있습니다. 원래는 보정되지 않은 칸을 모두 같은 경고로 띄웠는데, 지금은 방향에 따라 나눕니다.

  • [위험]: 지금 이 값이 답을 지나치게 권하는 쪽으로 밉니다
  • [임시값]: 보정은 안 됐지만 안전한 쪽입니다

둘 다 아직 실측하지 않았다는 점은 같지만, 읽는 사람이 해야 할 일은 다릅니다. 앞의 것은 재고 나서 결론을 쓰라는 뜻이고, 뒤의 것은 그대로 두어도 된다는 뜻입니다. 다른 상태를 같은 경고로 묶으면 급한 것과 급하지 않은 것이 같은 색으로 보입니다.

모르는 칸은 비워 두라고 적었습니다

가장 마음에 드는 것은 마지막 변화입니다. 사용 안내 첫 장에 모르는 칸은 비워 두세요라고 적어 두었습니다.

비워 두면 읽는 시점에 보수적인 값이 자동으로 들어갑니다. 반대로 감으로 채우면 대개 순진한 값이 들어가고, 앞에서 본 것처럼 답이 한 방향으로 밀립니다. 모른다고 말하는 편이 아는 척하는 편보다 안전한 구조를 만들어 두면, 사용자에게 성실함을 요구하지 않아도 됩니다.

보정이 끝나기 전까지 나온 값은 답으로 읽지 말고 하한으로 읽으라고 적었습니다. 숫자 하나를 내놓으면서 그 숫자를 어떻게 읽어야 하는지 함께 알려 주는 일은, 정확도를 높이는 일과는 다른 종류의 일입니다.


정리하면

  • 모르는 값을 채우는 기본값에는 방향이 있습니다. 1.0 이나 정격값이 중립으로 보이는 것은 숫자가 깔끔해서일 뿐입니다.
  • 그 방향은 잴 수 있습니다. 보정된 답을 참값으로 두고 하나씩 되돌려 답이 어디로 가는지 보면 됩니다.
  • 반드시 전부 함께도 되돌려 봐야 합니다. 편향이 같은 뿌리에서 나오면 상쇄되지 않고 겹칩니다. 하나씩은 많아야 9%p 였던 차이가 합쳐서 28%p 가 됐습니다.
  • 순진한 기본값이 한 방향으로 몰리는 데는 이유가 있습니다. 대개 "이상적인 경우"를 뜻하고, 이상적인 조건은 검토하는 수단을 좋아 보이게 만듭니다.
  • 틀릴 방향은 손실을 회복하는 데 걸리는 시간으로 고릅니다. 되돌리는 데 몇 주가 걸리는 쪽으로 기울이지 않습니다.
  • "아직 보정 안 됨"은 한 종류가 아닙니다. 위험한 미보정과 안전한 미보정을 나눠서 경고합니다.
  • 모른다고 말하는 편이 안전해지도록 만들면, 사용자에게 성실함을 부탁하지 않아도 됩니다.

I built a model that works out what share of a production line should run as priority work. It solves the question with queueing theory, then runs the same conditions through a simulation to check the answer.

Most of its inputs come from records that already exist — process sequences, machine counts. But five or six of them cannot be known before measurement: how much a priority marking actually advances a job, how full a batch machine really runs, how heavy the tail of the waiting-time distribution is.

Those cells get defaults. And defaults tend to look like this — 1.0, the rated value, the theoretical value. They look neutral. They look like numbers that are not taking a side.

So I measured them

The method is simple. Take the answer produced with fully calibrated values as the truth, then revert one parameter at a time to its naive default and watch where the answer moves.

On a line whose true answer is 11.6%:

tail coefficient = 1.0     → 21.0%   (+9.4pp)
priority effectiveness = 1.0 → 19.0%   (+7.3pp)
batch fill = rated          → 17.4%   (+5.8pp)
jobs between setups = 8     → 16.2%   (+4.6pp)
arrival variability = 1.0   →  5.8%   (−5.8pp)   ← the only safe one
──────────────────────────────────────────────
all five together           → 40.0%   (+28.4pp, 3.4×)

Four of five pushed the same way. And the last line is the point. Individual deviations were 5–9pp; together they did not cancel, they compounded into 28.4pp. 11.6% became 40% — pinned against the ceiling.

Defaults are not neutral. They have a direction, and the direction is measurable.

Why the same direction

I first assumed coincidence. It wasn't. The reasons converge on one.

Every naive default means "the ideal case." Effectiveness 1.0 means a priority marking lands perfectly. Rated batch size means every batch runs full. A tail coefficient of 1.0 means the distribution is as well-behaved as theory allows.

But what this model measures is how well the priority mechanism works — and ideal conditions always make a mechanism look good. So the deviations were not independent. Different parameters, sharing one optimism.

Independent biases cancel. Biases from a common root compound. The only way to find out which you have is to revert them all together, not just one at a time. I built the per-parameter table first and added the last row later; without it I would have concluded "off by at most 9pp." It was 3.4×.

Moving the defaults

So the template defaults were reset to "when unknown, assume it hurts priority work": effectiveness 0.30 for batch machines and 0.90 elsewhere, batch fill at 80% of rated, tail coefficient 2.5. All at the conservative end of the measured ranges.

With those, the same line solves to 7.8% — 3.8pp below the true 11.6%. Still wrong, but wrong toward recommending less.

Choosing which way to be wrong

Not "wrong is wrong, pick either." The losses here are asymmetric.

  • Wrong by too little — a few urgent jobs run late. Raise the share next week and it recovers.
  • Wrong by too much — setups multiply, batches break, and ordinary work's lead time collapses. Dialling the share back does not drain the accumulated backlog for weeks.

When recovery times differ, the two errors are not the same size. So defaults should lean toward the one you can undo. That is an operations question rather than an accuracy question, which is why it went into the documentation and not only the code.

Splitting the warnings in two

There is a command that checks an input file. It used to raise one kind of warning for every uncalibrated cell. Now it separates them by direction.

  • [risk] — this value currently pushes the answer toward over-recommending
  • [placeholder] — uncalibrated, but on the safe side

Both mean "not measured yet." But they ask different things of the reader: measure this before you write a conclusion, versus leave it alone. Collapse both into one warning and the urgent and the harmless arrive in the same colour.

Telling people to leave cells blank

My favourite part is last. The first page of the guide says: leave what you don't know blank.

Blank cells get the conservative value at read time. Filled-in-by-instinct cells usually get a naive one, and naive ones lean the way we just saw. Build it so that admitting ignorance is the safer path and you no longer have to ask users to be conscientious.

And until calibration is done, the guide says to read the output not as the answer but as a lower bound. Handing over a number together with instructions for how to read it is a different kind of work from making the number more accurate.


In short

  • Defaults for unknown values carry a direction. 1.0 and rated values look neutral only because they are tidy.
  • The direction is measurable — hold the calibrated answer as truth and revert one parameter at a time.
  • Always revert them all together as well. Biases from a common root compound instead of cancelling. At most 9pp each became 28pp combined.
  • Naive defaults cluster in one direction for a reason: they usually encode "the ideal case," and ideal conditions flatter whatever mechanism is under review.
  • Choose which way to be wrong by recovery time. Do not lean toward the error that takes weeks to undo.
  • "Not yet calibrated" is not one state. Separate the dangerous kind from the safe kind.
  • Make admitting ignorance the safe path and you stop having to ask users for diligence.