← 제작기

세지 않기로 한 것은 세어지지 않는다

사이트 아래쪽에 방문자 수를 붙였습니다. 오늘 몇 명, 누적 몇 명이 나옵니다. 붙여 놓고 처음 든 질문은 기능에 관한 것이 아니었습니다.

그럼 어제까지 온 사람들은?

알 수 없습니다. 버그도 아니고 빠뜨린 것도 아닙니다. 그렇게 하기로 한 결과입니다.

없는 것은 만들어 낼 수 없다

이 사이트는 접속 로그를 남기지 않고, 원본 IP 를 저장한 적도 없으니, 어제까지 다녀간 사람은 되살릴 방법이 없습니다. 어딘가에 저장된 것을 꺼내 오는 문제가 아니라, 처음부터 기록되지 않은 것입니다.

그렇다고 아무것도 없지는 않았습니다. 두 군데에 흔적이 남아 있었습니다.

  • 호스팅 쪽 집계. 대시보드에는 지난 트래픽이 남아 있지만, 그것이 세는 단위가 지금 카운터와 같은지 확인해야 합니다. 단위가 다르면 두 숫자를 더한 값은 아무것도 뜻하지 않습니다.
  • 게시판이 도배를 막으려고 저장한 해시. 같은 방식에 같은 솔트를 써서 지금의 방문 기록과 그대로 비교할 수 있지만, 이 해시는 글을 쓴 사람만 담고 있습니다. 읽기만 한 대다수는 없습니다.

둘 다 이전 방문자라고 부르기에는 다른 것을 세고 있었습니다. 오늘부터 0 에서 시작하기로 했습니다.

0 은 아무도 오지 않았다는 뜻이 아닙니다. 그때 세지 않기로 했다는 뜻입니다.

이 글에서 하고 싶은 말은 결국 이것입니다. 무엇을 기록할지는 그 숫자가 필요해지기 훨씬 전에 정해집니다. 그때는 결정처럼 느껴지지 않고 그저 만들지 않은 것으로 보이다가, 나중에 그 숫자가 필요해졌을 때에야 그것이 결정이었다는 것을 알게 됩니다.

숫자는 이름만큼만 정직하다

가장 쉬운 방법은 페이지가 열릴 때마다 1 씩 더하는 것입니다. 이 사이트는 탭이 넷이고 글이 여럿이어서, 한 사람이 둘러보면 대여섯이 찍힙니다.

그 수를 방문자라고 적으면 네 배쯤 부푼 숫자를 사실처럼 내놓는 셈입니다. 속이려는 의도가 없었어도 결과는 같습니다.

코드를 쓰기 전에 단위부터 정했습니다. 그날 다녀간 사람입니다. 같은 사람이 탭을 옮겨 다녀도 그날은 1 이고, 다음 날 다시 오면 그때 또 1 이며, 누적은 그 합입니다.

그 정의는 표의 구조로 정해 두었습니다. (날짜, 해시) 한 쌍이 기본 키여서 같은 날 같은 사람은 두 번 들어갈 수 없습니다. 정의를 주석으로 적어 두면 다음 사람이 보지 못하고 지나칠 수 있지만, 스키마로 적어 두면 어기고 싶어도 어길 수 없습니다.

정확해지려면 사생활을 내줘야 한다

이 숫자는 실제보다 적게 나옵니다. 이동통신망이나 회사망에서는 여러 사람이 IP 하나를 나눠 쓰는데, 그 사람들은 모두 한 명으로 셉니다.

고치는 방법은 압니다. 브라우저에 식별자를 하나 심으면 되지만, 그것이 바로 하지 않기로 한 일입니다.

지금 저장하는 것은 (날짜, 되돌릴 수 없는 해시) 한 쌍뿐입니다. 원본 IP 는 솔트를 섞어 되돌릴 수 없게 만들고, 그 사람이 어느 글을 봤는지는 처음부터 기록하지 않습니다. 나중에 누가 그 정보를 요구하더라도 내줄 것이 없습니다.

원칙: 어느 쪽으로 얼마나 틀렸는지 말할 수 있으면 틀린 숫자도 거짓말은 아닙니다. 그런 숫자가, 어떻게 틀렸는지 말할 수 없는 정확해 보이는 숫자보다 낫습니다.

세지 못한 것과 0 은 다르다

서버에 닿지 못하면 카운터는 0 을 적지 않고 자리를 비웁니다. 0 은 "아무도 오지 않았다"는 주장이지만, 실제로 일어난 일은 "세지 못했다"입니다. 화면에서는 둘이 똑같이 생겼어도 읽는 사람에게는 다른 이야기입니다.

같은 판단을 다른 앱에서도 한 적이 있어서 여기서는 다시 따지지 않았습니다. 그때의 이유는 지어낸 답보다 빈칸이 낫다에 있습니다.

아끼려는 것을 쓰면서 아낄 수는 없다

남용이 걱정됐습니다. 누가 이 주소를 계속 호출하면 데이터베이스 작업이 그만큼 늘어납니다. 마침 이 사이트의 게시판에는 이미 속도 제한이 있으니 그대로 가져오면 될 것 같았습니다.

게시판의 방식은 "이 사람이 최근 한 시간에 몇 번 썼나"를 데이터베이스에서 세는 것인데, 방문 카운터에서 아끼려던 것이 바로 그 데이터베이스 작업이었습니다. 막을지 말지 판단하려고 조회를 하나 더 붙이면, 호출하는 쪽의 한 번이 오히려 더 비싸집니다.

같은 도구도 자리를 옮기면 반대로 작동합니다. 무엇을 아끼려는지 먼저 말하고, 그 도구가 바로 그것을 쓰는지 봐야 합니다.

속도 제한 대신 평소 비용을 줄였습니다. 브라우저가 그날 한 번만 기록하고 그 뒤로는 읽기만 하도록 나눴고, 누적은 매번 전부 세지 않고 한 줄에 담아 두게 했습니다. 진짜 남용을 막는 일은 이 코드보다 아래 단계, 곧 요청이 코드에 닿기 전에 할 일이고, 실제로 그런 호출을 당하면 그때 걸 생각입니다.

영리한 방법보다 스스로 복구되는 방법

마지막에 문제가 하나 남았습니다. 방문 기록을 넣는 일과 누적을 1 올리는 일이 서로 다른 트랜잭션이어서, 그사이에 서버가 죽으면 기록은 남고 누적만 올라가지 않습니다. 그러면 그 1 은 영원히 모자랍니다.

한 트랜잭션으로 묶는 방법이 있었습니다. 데이터베이스가 "방금 문장이 몇 줄을 바꿨나"를 알려 주는 함수를 쓰면 됩니다. 그 함수가 묶음 안에서 바로 앞 문장을 가리킨다는 보장은 끝내 확인하지 못했고, 어긋나도 오류 없이 조용히 계속 틀립니다.

결국 둘 중 하나를 골라야 했습니다.

  • 드물게 1 이 모자랍니다(그리고 모자란 줄도 모릅니다)
  • 확인하지 못한 보장 위에 올려 두고, 어긋나면 조용히 계속 틀립니다

둘 다 마음에 들지 않아서 세 번째를 택했습니다. 묶지 않고, 그날 첫 방문자가 올 때 전체를 한 번 다시 셉니다. 하루에 한 번이라 비용은 거의 없고, 그사이에 생긴 어긋남은 하루 안에 저절로 바로잡힙니다.

원칙: 확인하지 못한 보장 위에 정확함을 올려 두지 않습니다. 대신 틀려도 스스로 바로잡히는 쪽을 고릅니다.

정리하면

  • 지금 세지 않는 것은 나중에도 셀 수 없습니다. 그 결정은 숫자가 필요해지기 전에, 결정처럼 보이지 않는 모습으로 내려집니다.
  • 숫자는 그 이름만큼만 정직합니다. 페이지뷰를 방문자라고 부르는 순간, 의도와 상관없이 부풀린 수를 내놓게 됩니다.
  • 정의는 주석 대신 스키마에 적습니다. 주석은 지나칠 수 있지만 기본 키는 어길 수 없습니다.
  • 어느 쪽으로 얼마나 틀렸는지 말할 수 있으면 거짓말이 아닙니다. 정확해지는 대가가 사생활이라면 틀린 채로 두는 편이 낫습니다.
  • 같은 도구도 자리를 옮기면 반대로 작동합니다. 아끼려는 것을 쓰면서 아낄 수는 없습니다.
  • 영리한 방법보다 스스로 바로잡히는 방법을 고릅니다. 조용히 계속 틀리는 것이 가끔 모자란 것보다 나쁩니다.

이 사이트에 게시판을 붙이면서 무엇을 브라우저에 맡길 수 있는지 따진 이야기는 화면은 권한이 없다에, 재는 도구를 먼저 의심해야 했던 이야기는 흔들린 건 곡선이 아니라 재는 자였다에 있습니다.

I added a visitor count to the bottom of this site — how many today, how many in total. The first question that came up once it was live was not about the feature.

So what about everyone who came before today?

There is no way to know. And that is not a bug, or an oversight. It is the result of a decision.

You cannot make what was never kept

This site keeps no access logs. It has never stored a raw IP address. So the people who came yesterday are unrecoverable — not buried somewhere waiting to be dug out, but never written down in the first place.

It wasn't quite nothing. Two traces existed.

  • The host's own analytics. Past traffic is sitting in the dashboard. But I would have to confirm it counts the same unit as this counter does, and if it doesn't, adding the two numbers produces a figure that is nothing at all.
  • Hashes the board stored for rate limiting. Same method, same salt, so they compare directly against visit records. Except they only capture people who posted. Everyone who only read is absent.

Both were counting something other than "past visitors". So the counter starts at zero, today.

Zero does not mean nobody came. It means I decided not to count back then.

That is the whole of this piece. The decision about what to record is made long before the number is wanted — and at the time it does not feel like a decision at all. It looks like something you simply didn't build. You only learn it was a decision when you need the number.

A number is only as honest as its label

The easiest implementation adds one every time a page opens. But this site has four tabs and a pile of posts. One person looking around registers five or six.

Label that "visitors" and you are presenting a number inflated several times over as if it were fact. Not intending to lie doesn't change the result.

So the unit was settled before any code: a person who came that day. Move between tabs all you like and the day still counts one; come back tomorrow and that is one more. The total is the sum of those.

And that definition is pinned into the shape of the table — the pair (day, hash) is the primary key, so the same person cannot land in the same day twice. Write a definition in a comment and the next person walks past it. Write it in the schema and it cannot be broken.

Accuracy would have to be bought with privacy

This number runs below reality. Mobile carriers and office networks put many people behind a single IP, and all of them count as one.

I know the fix: plant an identifier in the browser. That is precisely the thing I decided not to do.

What gets stored is one pair: (day, an irreversible hash). The raw address is salted beyond recovery, and which posts someone read is never written down at all. So if anyone later asks for it, there is nothing to hand over.

Principle: a wrong number is not a lie if you can say which way it is wrong and roughly how much. That beats an accurate number you cannot account for.

"Couldn't count" is not "zero"

When the server cannot be reached, the counter leaves the space empty rather than printing zero. Zero is a claim that nobody came; what actually happened is that nothing could be counted. On screen the two look identical, but to a reader they are different statements.

I made the same call in another app, so I won't re-argue it here — the reasoning is in Empty-Handed Beats Invented.

You cannot protect a resource by spending it

Abuse was a worry: someone hammering this endpoint means database work in proportion. The board on this site already has rate limiting, so it looked like a matter of copying it over.

The board's approach is to count, in the database, how many times this person has posted in the last hour. But database work is exactly what the visit counter was trying to conserve. Adding a query to decide whether to conserve makes each hit by the attacker more expensive, not less.

The same tool inverts when it changes position. Name what you are conserving first, then check whether the tool spends it.

So instead of rate limiting I cut the ordinary cost: the browser writes once a day and only reads after that, and the running total is held in a single row rather than recounted every time. Stopping real abuse belongs to the layer below this code — before a request ever reaches it — and that is something to switch on if it ever actually happens.

Something that heals, over something clever

One thing remained. Recording the visit and incrementing the total are two separate transactions, so if the server dies between them the record survives while the total does not move. That one is then missing forever.

There was a way to fuse them into one: use the database function that reports how many rows the last statement changed. But I could not confirm the guarantee that it refers to the immediately preceding statement inside a batch. And if that assumption is wrong, what happens? No error. It is quietly wrong, forever.

The choice came down to this:

  • rarely short by one (and never knowing it)
  • built on an unverified guarantee, and if it slips, quietly wrong forever

Disliking both, I took a third. Don't fuse them — recount everything once, when the day's first visitor arrives. Once a day costs effectively nothing, and any drift that opened up in between disappears within a day, on its own.

Principle: don't rest correctness on a guarantee you haven't verified. Choose the design that heals instead.

In short

  • What you don't count now cannot be counted later. That decision gets made before the number is wanted, in a form that doesn't look like a decision.
  • A number is only as honest as its label. Call page views visitors and you publish an inflated figure regardless of intent.
  • Put the definition in the schema, not a comment. A comment can be walked past; a primary key cannot be broken.
  • If you can say which way it is wrong, it isn't a lie. When accuracy has to be bought with privacy, leave it wrong.
  • The same tool inverts when it changes position. You cannot conserve a resource by spending it.
  • Prefer what heals to what is clever. Quietly wrong forever is worse than occasionally short.

What can and cannot be trusted to the browser, worked out while adding the board to this site, is in The Screen Has No Authority; the case for suspecting the instrument first is in It Was the Instrument Shaking, Not the Curve.