How a post becomes a solid

← the game · the atlas

A sentence embedding is a point in 768 dimensions. Nobody can read one. The usual responses are to project it to 2D and lose almost everything, or to draw something decorative beside it and call that an explanation.

This page does neither. It uses the vector as the spectrum of a surface — in the one basis where "the spectrum of a surface" is not a metaphor.

The map

Real spherical harmonics Ylm are the Fourier basis on the sphere. Any well-behaved radial function has exactly one expansion in them, so a coefficient vector is a shape and a shape is a coefficient vector:

r(θ,φ) = R · ( 1 + amp · Σl Σm ĉlm · Ylm(θ,φ) )

The whole design is the choice of which number goes in which slot. Dimensions are ranked by their variance across the corpus, so the ranking describes the corpus rather than any one post:

bandslotswhat it carries
l = 01 Deliberately no dimension at all. It carries ‖z‖ — how far this post sits from the average post. Ordinary posts are small pebbles; strange ones are big.
l = 1–424 The 24 loudest dimensions, one slot each. These are the lobes: the silhouette you can read across a room, and the part you learn first.
l = 5–1096 The remaining ~740 quiet dimensions, carried in by a seeded sparse random projection. Not discarded and not invented: the Fourier transform of everything too small to earn a lobe of its own. This is the grain, and it is what makes two posts with the same gross silhouette still tell apart.

Colour is a second, independent linear readout — the corpus's top three principal directions mapped into OKLCH. It is redundant with the shape on purpose: two channels saying the same thing is how anyone learns to read one of them.

What this buys, provably

Two posts that mean the same thing cannot look different

The map from embedding to coefficients is linear, and the harmonics are orthonormal. Parseval's theorem then gives, exactly:

‖ surfaceA − surfaceB ‖L²(S²)  =  ‖ cA − cB ‖2

The distance between two shapes, integrated over the sphere, is the distance between the two posts' vectors. That is not a design goal we approximated — it is an identity, and embed-geometry.selftest.mjs §2 checks it against a Gauss–Legendre quadrature to 1 part in 108.

Measured, on a corpus with realistic embedding statistics

The selftest fits the map on a synthetic corpus built to look like real sentence embeddings — topic clusters, the dominant "common direction" every real model has, per-post noise, L2-normalised — and then asks the question the game asks:

questionmeasureresult
Given a post, is a same-topic post the more similar-looking one? AUC over same- vs different-topic pairs> 0.95
Same question via the full L² surface distance AUC> 0.78
…with the size channel held constant AUC> 0.95
Does shape similarity track raw embedding cosine? Pearson r> 0.55
Does it preserve fine ordering among unrelated posts? Spearman ρ> 0.25 only

Where it loses information — three places, named

1. Fine ordering among unrelated posts is not preserved

Spearman ρ ≈ 0.3 against raw cosine, versus AUC > 0.95 for the same/different-topic split. The band gains deliberately magnify the loud dimensions, so the map keeps the big structure very well and scrambles the fine ranking of things that were never similar anyway. Reading "these two are both about football" off the geometry is sound. Reading "this one is the 4th most similar rather than the 9th" is not.

2. Shape distance mixes meaning with strangeness

The L² surface distance carries size as well as form, and size comes from ‖z‖ — how unusual the post is, which has nothing to do with its topic. Two posts about the same thing at different strangeness are genuinely different-sized objects, which is why the AUC above drops from 0.95 to 0.78 and recovers to 0.95 when size is held constant. The game therefore judges ripeness on direction only (cosine), never on L² distance.

3. The quiet 740 dimensions are compressed, not shown

96 high-band slots cannot hold 740 numbers. The projection is a Johnson–Lindenstrauss embedding, which preserves pairwise distances in expectation — the selftest pins the median distortion under 35%. So the grain is a real reading of the quiet subspace, at a stated resolution. It is not a per-dimension display and never could be.

Two decisions that were measured, not chosen

How hard to whiten

Each dimension is divided by stdiα · globalScale1−α. Textbook whitening (α = 1) sounds correct and is quietly destructive here: it flattens every dimension to unit variance, which annihilates the variance ranking the whole map is built on and amplifies dead dimensions to full authority. α = 0 lets one runaway dimension swamp the silhouette. Sweeping α:

αsame-topic separation (AUC)noise gain on dead dimensions
0.000.9971.0×
0.500.9981.6×
1.000.9992.3×

Separation is a near-tie across the whole range, so it decides nothing. The tiebreak is the second column, and α = 0.5 ships: essentially the best separation, half the noise amplification of full whitening, and a real square-root trace of the variance ranking left in the magnitudes.

Why text-only posts

Not squeamishness — the premise. A post carrying an image has already told you what it is about through a channel the embedding never saw, so its shape would be judged against information the player can see and the geometry cannot. Replies are dropped for needing a parent to make sense of, reposts because the reposter said nothing, and non-English posts because bge-base-en-v1.5 would produce a vector, and a shape, and no meaning.

The game is the measurement

A pretty visualisation can always be defended by saying it feels meaningful. Hunt mode refuses that defence. It shows one anchor solid, rains posts down, and marks a post ripe if its shape is within τ of the anchor's — where τ is a quantile of the round's similarities, so a round is never accidentally impossible or trivial.

At the end, of the posts you chose to slice, were more of them ripe than if you had sliced at random? Under the null hypothesis "you cannot read the shapes", your hits are Binomial(slices, base rate), and the page runs a one-sided exact test. If p ≥ 0.05 it tells you plainly that you were indistinguishable from guessing.

The test is calibrated rather than decorative: rounds.selftest.mjs §5 runs six hundred simulated guessers and asserts the false-positive rate lands at the nominal α, and runs perfect readers and asserts they are recognised over 95% of the time. The page is allowed to be wrong about you, which is the only thing that makes it worth being right.

Honest provenance

Every page states, on screen, what its shapes were computed from. Two things can fail independently and both are announced rather than hidden:

Source: zest/embed-geometry.js (the map), zest/rounds.js (the rules and the statistics), with embed-geometry.selftest.mjs and rounds.selftest.mjs alongside them. Run both before changing either.