공개 데이터는 열려 있지 깨끗하지 않다
업계 재고 사이클이 지금 어느 국면에 있는지를 공개된 공시 데이터만으로 판정하는 도구를 만들었습니다. 비공개 데이터는 한 줄도 쓰지 않습니다. 쓰면 남에게 보여 줄 수 없고, 보여 줄 수 없으면 검증도 받을 수 없기 때문입니다.
계산 자체는 어렵지 않습니다. 재고를 원가 기준으로 며칠 치 들고 있는지 재고, 재고 증가율에서 매출 증가율을 빼서 "쌓이는 속도가 팔리는 속도보다 빠른가"를 봅니다. 그 값의 수준과 방향으로 네 국면 가운데 하나를 고릅니다. 몇 줄이면 됩니다.
파이프라인 코드의 대부분은 계산에 쓰이지 않습니다. 아래에 적는 함정들을 피하려고 있는 코드이고, 모두 원본 응답을 열어 보고 확인한 것이지 일반론이 아닙니다.
① 회계 4분기가 없다
연차 보고서는 연간 값만 태그를 달고 4분기 값을 따로 주지 않습니다. 1·2·3분기는 있는데 4분기가 없습니다.
현금흐름 쪽은 더합니다. 처음부터 연초 이후의 누계로만 공시돼서, 90일짜리 구간은 1분기 하나뿐이고 나머지는 180일, 270일, 365일입니다.
시작일이 같은 누계들을 종료일 순서로 세우고 이웃한 값끼리 뺍니다. 270일 누계에서 180일 누계를 빼면 3분기가 나오는 식입니다.
② 기업이 도중에 태그를 바꾼다
같은 기업이 어느 해에 매출 태그를 바꿉니다. 긴 이름에서 짧은 이름으로 옮기는 식인데, 옮긴 시점부터 옛 태그에는 아무 값도 들어오지 않습니다.
"값이 잡히는 첫 태그를 쓴다"로 짜면, 옛 태그를 잡은 뒤로 그 뒤 4년 치가 통째로 사라집니다. 화면에는 오류가 아니라 "데이터 없음"으로 나옵니다.
이 문제는 태그를 하나 고르지 않고 여러 태그의 값을 합쳐서 풀었습니다.
③ 회계 기준마다 이름이 다르다
어느 회계 기준을 쓰느냐에 따라 같은 "재고"에 다른 이름의 태그가 붙습니다. 한쪽에서는 복수형 한 단어이고, 다른 쪽에서는 "순"이 붙습니다. 이름 하나로 찾으면 절반이 빕니다.
④ 분기 손익이 아예 없는 기업이 있다
어떤 기업은 특정 회계연도의 분기 매출원가에 아예 태그를 달지 않고 연간 값만 냅니다. 이런 값은 앞의 누계 빼기로도 되살릴 수 없습니다.
대신 매출 − 매출총이익 = 매출원가 라는 항등식으로 채웁니다. 나머지 두 값은 있기 때문입니다.
되살린 값을 어떻게 믿을 것인가
①의 누계 빼기는 이 파이프라인의 첫 단계입니다. 여기가 틀리면 그 위에 쌓은 모든 지표가 조용히 틀리고, 국면 판정이 이상해도 어디서 틀렸는지 찾을 수 없습니다.
되살린 네 분기의 합을 같은 출처의 연간 원본 값과 대조했습니다.
12개사 × 36개 연도 최대 오차 0.0007%
연간 값은 되살리는 데 쓰지 않은 독립된 수라서, 이 대조는 자기 자신과 비교하는 것이 아닙니다.
검증하는 쪽이 스스로를 근거로 삼고 있었던 적이 한 번 있어서, 대조할 상대가 정말 독립인지부터 확인합니다.
두 출처가 서로를 메운다
한 출처에는 분기 값을 아예 내지 않는 기업들이 있습니다. 제출 양식이 달라서 연간 값만 냅니다. 최신 재무 태그가 1년 반 전에 멈춰 있는 기업도 있습니다.
그 자리를 다른 출처의 월 단위 매출 공시가 채웁니다. 이쪽은 법에 따라 매달 정해진 날까지 공시하게 되어 있어서, 분기 공시를 기다릴 필요가 없습니다.
어느 쪽도 혼자서는 업계 전체의 단면을 만들지 못합니다. 둘이 함께 있으면 할 수 있는 일이 하나 더 생기는데, 바로 교차 확인입니다.
어느 달 월매출이 +720% 로 튄 종목이 나왔습니다. 한 출처만 보면 이것이 신호인지, 공시 오류인지, 비교 시점의 값이 낮았던 탓인지 가릴 방법이 없습니다. 같은 시기 다른 출처의 분기 매출을 보니 +346% 였고, 독립된 두 출처가 같은 방향을 가리켰습니다.
이 확인이 없으면 한쪽의 공시 오류가 그대로 결론이 됩니다.
한계는 보고서가 매번 함께 출력한다
이 도구는 실행할 때마다 자기 한계도 함께 출력합니다.
- 재고 갭에는 상한이 없습니다. 증가율의 차이라서 매출이 급변한 기업은 갭의 크기를 다른 기업과 비교하는 데 쓸 수 없습니다(한 종목은 −347% 가 나옵니다). 방향을 판정하는 데만 씁니다.
- 연간 값만 있는 기업은 시점이 뒤처집니다. 보고서가 종목마다 이 사실을 밝힙니다.
- 회계연도를 역년 분기로 맞췄습니다. 결산월이 제각각이라 맞추지 않으면 업계 전체를 한 단면으로 볼 수 없는데, 맞추는 과정에서 시점이 최대 한 달까지 어긋납니다.
이런 한계를 각주 대신 본문에 넣은 이유는 지어낸 답보다 빈칸이 낫다와 같습니다. 수치는 그럴듯하게 생겼고, 한계를 적지 않으면 읽는 사람은 그 수치를 있는 그대로 믿습니다.
원본 응답은 받은 그대로 남겨 둡니다. 지표가 이상해 보일 때, 계산이 틀렸는지 원본이 원래 그런지를 열어서 가릴 수 있어야 하기 때문입니다.
덧붙임: 그림은 좌표가 맞아도 눈으로 봐야 한다
결과를 카드 이미지로 뽑는데, 여기서 걸린 문제도 모두 직접 열어 보고 잡았습니다.
- 이상치 하나가 차트를 망칩니다. +720% 짜리 막대 하나 때문에 나머지 열 개가 모두 3px 가 됐습니다. 위쪽 이상치를 뺀 축척을 쓰고, 넘치는 막대는 끊긴 표시로 나타냅니다.
- 4분면 그래프는 두 축의 축척을 따로 잡아야 합니다. 갭은 수백 % 까지 벌어지는데 갭의 분기 변화는 몇 %p 여서, 같은 축척이면 모든 점이 가로 한 줄에 붙습니다.
- 라벨끼리 겹치지 않게 할 때는 점과 축 문구까지 함께 피해야 합니다. 라벨끼리만 피하게 했더니 라벨이 자기 점 위에 올라앉았습니다.
- 글꼴에 빼기 기호와 화살표 글리프가 없어서 두부 문자(⊠)로 나왔습니다. 안전한 문자로 바꿔 주는 단계를 따로 두었습니다.
자동 검사는 통과했는데 화면이 틀렸던 일과 같은 종류입니다. 좌표는 모두 맞았습니다.
남는 것
- 공개 데이터는 열려 있을 뿐 깨끗하지는 않습니다. 파이프라인 코드의 대부분은 계산보다 함정을 피하는 데 쓰입니다.
- 첫 단계일수록 독립된 값과 대조합니다. 되살리는 로직이 틀리면 그 위의 모든 지표가 조용히 틀리고, 틀린 자리를 찾을 수 없습니다.
- 한 출처의 값이 튀면 다른 출처에 물어봅니다. 하나만 보고는 신호인지 오류인지 가릴 수 없습니다.
- "값이 잡히는 첫 것"으로 짜면 데이터가 조용히 사라집니다. 오류 대신 "없음"으로 나오는 실패가 가장 늦게 발견됩니다.
- 한계는 보고서가 직접 말하게 합니다. 읽는 사람이 문서를 찾아볼 것이라고 기대하지 않습니다.
I built a tool that reads which phase of the inventory cycle an industry is in, from public filings alone. Not one line of non-public data — because non-public data cannot be shown to anyone, and what cannot be shown cannot be checked.
The arithmetic is not hard. Measure how many days of cost sit in inventory, subtract revenue growth from inventory growth to see whether stock is piling up faster than it sells, and place the level and direction of that figure into one of four phases. A few lines of code.
But most of the pipeline is not arithmetic. It exists to avoid the traps below, every one of which was confirmed by opening the raw response — none of it is general advice.
1. There is no fiscal fourth quarter
Annual reports tag the year and do not break out Q4. Q1, Q2 and Q3 are there; Q4 is not.
Cash-flow items are worse. They are disclosed only as year-to-date figures, so exactly one 90-day window exists — the first quarter. The rest are 180, 270 and 365 days.
So cumulative figures sharing a start date are ordered by end date and differenced pairwise. Subtract the 180-day from the 270-day and the third quarter falls out.
2. Companies switch tags mid-history
A company changes its revenue tag in some year — moving from a long name to a shorter one — and from that point the old tag returns nothing.
Written as "use the first tag that returns a value", latching onto the old tag makes the following four years vanish entirely. And on screen it appears not as an error but as "no data".
So the tags are not chosen between. They are merged.
3. Standards diverge
Depending on which accounting standard a filer uses, the same "inventory" is tagged under a different name — one plural word in one, the same word with "net" appended in the other. Search by name and half the set comes back empty.
4. Some filers have no quarterly income statement at all
For certain fiscal years, some companies never tag quarterly cost of revenue. Annual only. Differencing cannot recover it.
Instead it is filled through the identity
revenue − gross profit = cost of revenue. The other two are there.
How the reconstruction earns trust
The differencing in (1) is the pipeline's first stage. If it is wrong, every metric stacked above it is quietly wrong, and a strange phase call becomes untraceable.
So the four reconstructed quarters were summed and compared against the annual originals from the same source.
12 companies × 36 fiscal years largest error 0.0007%
The annual figure plays no part in the reconstruction, so this is not a comparison with itself.
Having once had a verifier that used itself as its own evidence, the first question is now whether the thing being compared against is genuinely independent.
Two sources covering each other
One source has a whole class of filers with no quarterly figures at all. A different filing form means annual only. Some have a latest financial tag standing still for a year and a half.
That hole is filled by the other source's monthly revenue disclosures, which are required by law by a fixed day each month — no waiting for a quarter.
Neither source alone produces a cross-section of the industry. And having two allows one more thing: a cross-check.
One month a name's revenue jumped +720%. Looking at that alone, there is no way to tell a signal from a filing error from a base effect. The other source's quarterly revenue for the same period was +346%. Two independent sources pointing the same way.
Without that check, a filing error on one side becomes the conclusion.
The report states its own limits every run
The tool prints its limitations alongside its results, every time.
- The inventory gap has no upper bound. It is a difference of growth rates, so for a company whose revenue moved sharply the magnitude cannot be compared across companies (one comes out at −347%). It is used for direction only.
- Annual-only filers lag. The report names them individually.
- Fiscal years are folded onto calendar quarters. Year-ends vary, and without folding there is no industry cross-section — but the folding introduces up to a month of timing error.
These sit in the body rather than in a footnote for the same reason as empty-handed beating invented. A number looks credible on its own, and a reader who is not told the limits will believe it exactly as printed.
Raw responses are also kept as received. When a figure looks wrong, "is the calculation wrong or is the source like that" has to be answerable by opening the source.
A note — a drawing has to be looked at, not just computed
The results go out as card images, and everything that went wrong there was caught by opening the file.
- One outlier kills a chart. A single +720% bar flattened the other ten to 3px. The scale now excludes top outliers and clips overflowing bars with a break mark.
- A four-quadrant plot needs separate scales per axis. The gap spans hundreds of percent while its quarterly change is a few points, so one shared scale collapses every point onto a single horizontal line.
- Label collision avoidance has to include the points and the axis text. Avoiding only other labels let a label sit on top of its own point.
- The font has no glyph for the minus sign or the arrow. Both render as tofu. A substitution layer maps them to safe characters.
The same species as the checks that passed while the screen was wrong. Every coordinate was correct.
What stays
- Public data is open, not clean. Most of a pipeline's code goes into avoiding traps rather than into arithmetic.
- The earlier the stage, the more it needs an independent comparison. A wrong reconstruction makes every metric above it quietly wrong and impossible to locate.
- When one source spikes, ask the other. One source alone cannot separate a signal from an error.
- "The first one that returns a value" fails silently. A failure that surfaces as "no data" rather than as an error is the last one anybody finds.
- Make the report say its own limits. Do not expect the reader to go looking for the documentation.