· articles

Your Glasses Aren't Green, They're Grue

General disclaimer about the handwaving nature of these articles

While this blog features mostly longform content, the issues discussed and thinkers presented deserve their own 3-parts university course (and usually they have one).

Everything you read here is a handwaving summary, lacking to the point of liability.

Hopefully it will still be an interesting read.

At the Gates of the Emerald City, Part 2: Your Glasses Aren’t Green, They’re Grue

Let’s recap. We’ve spent part 1 on Hume’s old riddle of induction, exploring some back and forth until finally conceding that we have no rational justification for our inductive inferences and projections.

In this part, we’re going to take it for granted that induction is permissible, then see what we can build on top of that and how far it can take us. Spoiler: it won’t be far enough. Not in general, and nowhere close once we consider the challenges that GenAI introduces.

Lots of philosophical and practical nuance ahead. Let’s dive in.

Prologue

There’s a nice little detail in Baum’s The Wonderful Wizard of Oz that didn’t make it into the movie, but will provide us with a good metaphorical hook to begin with.

When Dorothy and her gang reach the gates of the Emerald City, they don’t just walk in. The Guardian of the Gates gives them green glasses they must wear before entering. These are locked at the back of the head by a PG-13 version of a Saw-like mechanism, so nobody can take them off. It’s for their own protection, you see, as the glory of the emerald city would otherwise blind them.

So they walk into a city and everything gleams with emerald, absolutely dazzling. However, as we eventually work out, it’s all a con, the city isn’t green at all, only the glasses are. Everybody’s wearing them and can’t take them off, so everyone agrees the city is emerald. 100% consensus out of a worthless observation, because the instrument is producing its own conclusion.

Part 1 left us with a sample we have no rational justification to project from; but now we’re over that whole thing, and just take the license to inductively infer as an axiom. This part will show the instrument we’re sampling with has a green tint, skewing our conclusions.

But don’t worry, things will get much worse. Turns out that’s actually the optimistic part, because unlike an actual tint, this skew isn’t the kind of thing we can calibrate away. The metaphor is a bit heavy handed, sure, but it will allow us to cover some surprisingly practical territory.

Our glasses were always green

If the metaphor is about a tinted set of glasses, what does it translate to, philosophically and practically?

The glasses represent the theory-ladenness of observation and fact. Theory-ladenness means there never are raw facts, only facts as interpreted by some theory, minimalist as it may be. You’d expect this claim to come from relativists, but in fact it’s extremely mainstream and widely accepted across philosophical views. We’ve met it in an extreme form with Kuhn’s paradigms, but one of the sharpest early statements around this actually belongs to Popper.

As the anecdote goes, Popper once opened a lecture by telling students “to observe”, then to write down their observations. The obvious question was quick to follow: observe what, exactly? Popper made explicit the fact that observation is always selective, and as such it needs a chosen object, or a target to be directed at (or a context, a point of view, a problem, etc.). There’s no such thing as a bare instruction “to observe” without some theory already dictating what’s worth paying attention to. Put differently, without a preexisting idea of what a signal looks like, everything is just noise.

Imagine Tycho Brahe and Kepler on a hill at dawn. Tycho thinks the earth is fixed, Kepler thinks it moves, and the physical image hitting their retinas is for all intents and purposes identical. But Tycho watches the sun slowly climb the stationary horizon, while Kepler watches the horizon roll away from a stationary sun. They don’t share an observation and then disagree about it; the disagreement is already baked into the seeing.

We don’t need the strong version of this, where theory reaches alllllll the way into your eyeballs (yuck). Let’s take for granted that the retina is honest (or that the log is accurate and the screenshot hasn’t been retouched, etc.). None of that is evidence yet. It only becomes evidence when it gets described, and that presupposes a vocabulary, which means classification, which means somebody has already decided on how things should be grouped. A theory doesn’t need to reach into your eyes in order to reach into your data.

In fact, part 1 already had a demonstration of a theory shaping data into evidence: the cliché of the urn with the balls. Split the red category into red and dark red, and suddenly your starting distribution changes while the urn in question remains exactly the same. Nothing happened to the actual balls (eyeballs or coloured balls in the urn). Something happened to the vocabulary, hence the categories, hence the theory, hence the way reality is structured into evidence.

Testers work through something similar in the sense that there is no such thing as a raw failure. A failure isn’t the behaviour itself, but the behaviour as it violates the expectations laid out by an oracle, and from this POV an oracle is a theory. The same observed output can be a defect, or a feature or a complete non-event; this depends entirely on which theory you brought to the screen. Or worse, it depends on the Jira taxonomy someone settled on a decade ago when they set up the template. Theory-ladenness isn’t philosophers trying to scam you into buying more philosophy. It’s as concrete as that Jira dropdown you hate.

And of course, it’s theories all the way down. The logs show you what somebody once had a theory about, which is why you can usually only see failures whose shape was preconceived. Your equivalence classes are a claim about which inputs are the same kind of thing, and that claim is drawn on a whitebox presumption that you bring to the system under test, not discover within it. And your severity taxonomy decides ahead of time what counts as a bad outcome, which means it also decides what quietly won’t be recorded as one.

This isn’t intended as a harsh criticism, BTW. You can’t test without a set of categories any more than you can observe without one. The green tint isn’t a defect in the lens, it’s what makes the lens a lens. But it does mean the green in “all green ✅” is (partly? mostly?) a property of the eyewear. Similarly, it means the post-incident question of “why didn’t the tests catch it” is less important than asking “what did this apparatus keeps us from seeing”. This is a more grounded POV than Land’s anti-rationalist defence system, but the practical consequences are the same: the test suite is obstructing some sight as the price of enabling any.

It also means consensus is cheap. If everyone on the team trained on the same syllabus, works from the same spec and inherited the same defect taxonomy, then their agreement isn’t independent evidence about the system. It’s evidence that the supply of company-issued glasses was well stocked. 3 people signing off on a release while wearing the same tinted glasses is a single observation, not 3.

The machine wears its own alien tinted glasses

Machine learning / AI shows up in this story in dual roles. It can be the product we’re testing, and it can be the tester itself (generating the cases, reviewing the diffs and judging the outputs). In this part we’ll focus on the first role; but before we can see what’s so alien about the machine’s glasses, let’s be clear about what’s so familiar about ours.

The family we know and love…

Say I give you 3 points on a graph and ask you to draw the curve passing through them. Turns out the “the” in the ask is misleading. There isn’t a curve that connects them, but an infinite multitude of curves. Whatever you drew, there’s also a wild sinusoid through those same 3 points, and a 10th degree polynomial that threads all 3 and then rockets off between them. Every one of those curves agrees where you have evidence, and disagrees everywhere else.

So no finite set of points can determine a single curve. In philosophy this is known as the underdetermination of theory by data, and it applies to test suites as it applies to graphs: for any given test suite, there are many systems and system states that pass it.

Some of them are the system you aimed for; others have real defects sitting in areas the suite never observed; others have defects that cancel out other defects in exactly the scenarios you test for. A green run doesn’t indicate a single specific implementation (the system), but an entire family of them (everything consistent with the evidence the test suite collected).

Usually we don’t really care about this nuance (read: nonsense), because those family members are very closely related. Humans only draw a few families of curves (yes, even humans who think themselves uniquely creative [yes, even you]). A human wrote the code aiming at an intended behaviour, so the ways it can deviate are constrained by human error patterns. Even when Dave (damn it Dave) put in that magic number spaghetti code to fix that client escalation, the code base still behaves like every other messed up codebase you’ve ever tested. It isn’t alien (one might say it’s famili-ar).

So when the test suite can’t distinguish if we’re dealing with system A or system B, we just say “ship it”, since we can implicitly rely on A and B behaving closely enough. You don’t really worry about your SaaS service exploding if the user uses the toilets. The service probably has defects, yes, and you might get surprised, for sure, but not jaw-droppingly astonished. Your mind won’t be blown by the system any more than the toilets will.

…is dead now

Machine learning components and systems built by GenAI nullify all of these implicit assumptions, and undermine decades of lessons learned and clever optimisations. ML and GenAI models draw their own curves between the points, and these are alien curves. There are many reasons these models end up alien, or alien-esque. A couple of big ones are randomness and simplicity.

Which rule you get is decided by noise: Imagine taking a standard training pipeline and running it repeatedly, just changing the random seed. Turns out you get models that are statistically indistinguishable on the relevant benchmark test set (same accuracy, same everything the benchmark reports) but that diverge substantially the moment you stress them on even the slightest variations. Read: a process indistinguishable to humans run twice, produces two rules that agree inside our test set and break apart outside it. Our human way of grouping instances is no longer a reliable indicator of them being “close” in any meaningful way (well, at least not meaningful to humans).

When the evidence can’t choose, simplicity chooses: If infinitely many curves fit the data points, the data can’t break the tie, so something else does. In modern model training this turns out to be the simplicity bias. Given a complex rule that captures the true pattern, and a simple rule that merely correlates with the answer in the training data, the produced model will bias towards the simple rule. This has been repeatedly documented in case after case, even when the complex one is more predictive and more robust. Read: the model ends up tracking whatever correlates with the pattern humans care about in the training data, rather than tracking that pattern itself. And since that correlation is an accident of the training data, extending the rule beyond it produces an extension alien to human judgement.

The duck on the lawn

This can be demonstrated with the infamous Waterbirds dataset, built as part of an attempt to catch (and mitigate) exactly these types of alien behaviours. Researchers took photos of birds, pasted them onto landscapes, and asked models to classify whether they contain a waterbird or a landbird. The pattern we’d want the model to track is bird morphology. But morphology is complex and hard to learn, and the training data was rigged (though with a realistic enough pattern). 95% of the waterbirds sit against water, and 95% of the landbirds against land. Which means there was a simpler feature that correlates with the correct classification. Blue and wet => waterbird. Green and leafy => landbird.

Now let’s transpose this experiment onto our own professional lives. Suppose someone who for sure wasn’t me or you (because we’d know better and would never do such a thing), built their test set the same way they built their training set. When this completely made up person ran their benchmark, simplicity bias would do exactly what you’d expect.

The model learns to track the background of the image. It scores amazingly well on this imaginary person’s test set, because that set came from the same distribution and carries the same background correlation. Dashboard’s all green boss, no worries. Then the model sees a duck standing on the lawn and it says landbird with total confidence, because it was never tracking the bird, only looking at the grass.

The model learned an alien rule that is perfectly correct on every single example you evaluated it on, and wrong the moment the correlation it silently depended on stops holding.

There’s no defect in the fit. Nor is there a defect in the test procedure as conventionally understood: you held out data properly, you didn’t peek, there was no tuning of the test set, the accuracy number is honest. And yet the thing you thought you were doing all along (training to classify birds) is not the thing you actually verified (classifying backgrounds, on a distribution where backgrounds just happen to correlate with birds).

Simplicity bias explains why the model reached for a simple correlation, but the training noise explains which alien rule it happened to end up with. Retrain the exact same pipeline, change nothing but the random seed, and you might get a model that tracked lighting, or JPEG artefacts. Same accuracy on the same benchmark, different rule underneath, and different projections in real world settings.

”Tinted” was the optimistic diagnosis

Everything above says the glasses are tinted. That’s bad, but it’s the friendly kind of bad, because a consistent error is a solvable engineering problem. You take an object whose true colour you know, compare it against what the instrument reports, compute the offset, and subtract it from every future reading.

So let’s calibrate the Waterbirds model. Take your known-truth cases, compare, compute the offset. The offset is zero. Of course it’s zero, as that was the whole point.

The model classified essentially every image in your test set correctly. Its readings and the truth matched exceptionally well across your entire evaluation sample set. Your calibration procedure reports a wonderfully accurate instrument, and it’s right: on all available evidence, the instrument is very accurate. Until you take it out for a spin in the real world and encounter a duck on grass.

The machine is following a rule that no human given the original task would even consider a candidate. We wouldn’t group backgrounds into a category worth having in our taxonomy at all; to us it’s unclassified noise. So to us, there’s no offset to subtract, because there is no offset in the things we’re tracking.

Connecting us to part 1, this is another way of turning Hume’s riddle on its head. The problem isn’t that we can’t justify drawing a curve through our data points. It’s that there are infinitely many, equally fitting curves, each extending and projecting into unobserved cases in wildly different ways. It’s the infinitely-many-straight-rules problem, back with a vengeance: no longer an abstract curiosity about what converges in the limit, but a practical question about which rule our instrument is actually tracking right now.

Let’s explore this further through the work of one of my personal favourite philosophers of science, Nelson Goodman, and his new riddle of induction (dam-dam-dam!).

According to Goodman, the problem isn’t that our glasses were green, it’s that they’re grue.

Grue

Goodman’s main body of work attempts to abolish any metaphysical, non-actual references from theories of science, epistemology and ontology. We’ll for sure visit those at a later time, but for now we’ll focus on one of his most encapsulated pieces: the new riddle of induction. Let’s walk through it.

Goodman defines a predicate (an adjective or an attribute) called grue. Something is grue if it’s observed before 2030 and is green, or not observed before then and is blue. Similarly he defines bleen as being observed before 2030 and blue or not observed before then and green (just one of the many grue-some terms these examples can be constructed on).

Now let’s think of our evidence regarding emeralds. Every emerald ever observed has been green. So, projecting inductively, it’s not unreasonable to claim that all emeralds are green, including the ones we dig up next year, and the year after that, and in 2031.

However…

Every emerald ever observed has also been grue. Think about it, each was observed before 2030 and found green, which satisfies grue’s first clause. So our evidence supports “all emeralds are grue” exactly as well as it supports “all emeralds are green”. Again these are different curves sharing all data points but diverging outside them; only this time it’s set up to be directly translatable to inductive hypotheses.

All observations equally support both hypotheses, but the hypotheses disagree about the future. “All emeralds are green” predicts the next emerald is green, and more importantly, it predicts the first emerald of 2031 to be green as well. “All emeralds are grue” predicts it’ll be blue, because after 2030, being grue means being blue.

What Goodman managed to construct is a clean scenario where we can prove the same set of observations equally support two hypotheses or rules, that produce a pure contradiction when projected into the future, and there’s no referee in the data itself. This is the new riddle of induction. Goodman actually thought Hume’s problem dissolves, and then showed that what it dissolves into is worse: if induction is permissible at all, our observations support infinitely many contradictory projections.

An objection that immediately comes to mind is that grue is a gerrymandered predicate, artificially engineered, while green is a natural predicate, and that only the latter is worthy of being inductively projected. That feels instinctively right, for sure, but turns out to be VERY hard to cash out. Any attempt to gesture to the notion of something being “natural” is immediately undermined: if no observation is theory-free, then “natural” can’t mean “read directly from reality”, because nothing is read directly from reality. Whatever makes green the sensible predicate over grue has to come from somewhere other than the emeralds.

You may want to say grue is defined in terms of a time and a reference to green, so it’s parasitic. But Goodman notes a speaker whose vocabulary is grue would in turn define green as the parasitic term with “grue-if-observed-before-2030-or-bleen-if-not”. From inside that language, green is the artificial one. It’s like in the classical urn example: there is nothing more natural about 3 colours than about 4. All options are determined by human language, even (read: especially) the ones we consider “natural”.

Goodman’s term for the property we’re chasing here is projectibility: the status of being legitimately extendable from observed to unobserved cases. We can agree green has it and grue doesn’t, but the problem is that there’s no formula for telling them apart that doesn’t smuggle in a prior judgement about which predicates are “natural”. We’ll just end up arguing in a circle.

Mainstream philosophy has mostly accepted Goodman’s problem as real. Meaning that unlike deduction, the validity of inductive claims relies on something outside the context of the claim itself. Different philosophers have proposed their own criteria for appealing to something beyond the claim: Quine’s natural kinds, restrictions to “qualitative” predicates, appeals to lawlikeness, and Goodman’s own answer, entrenchment (we’ll touch on what that means in a short while).

Most of these do real, informative work, and are very interesting to explore (and in future posts we’ll for sure do that). For today’s discussion, the one relevant thing that they all share is being rooted in human life: as linguistic communities, as products of specific evolutionary lineage, as sharing a common history, etc. None of these solutions is transferable to machines (overstepping for dramatic effect etc.).

Why this is our problem, precisely

Let’s map this back onto our Waterbirds case. The model learned a grue-predicate: waterbird-if-background-is-water. Contrast this with the intended green-predicate a human would learn: waterbird-if-morphology-matches-waterbird. On our highly correlated training distribution those two rules overwhelmingly agree (emeralds before 2030), so they’re confirmed identically by every scrap of evidence you currently have, and diverge only outside it. That is the essence of grue. But wait, there’s more:

The green and grue rules are not meaningfully distinguishable by ANY amount of additional evidence drawn from the same distribution. Run 10 times as many test images or run a million, as long as they come from a distribution where the backgrounds always line up with the birds, every single one confirms both rules equally. This is hopelessly different from ordinary sampling error, where more data narrows the error.

So scale doesn’t fix this, and neither do bigger evals or higher coverage. What fixes it is evidence from outside that distribution. Sounds easy enough, but notice why that real waterbirds test set exists in the first place. Someone has already suspected the backgrounds were doing the work, so they built the test to verify that. That sharp initial observation is like winning the lottery. Cool story bro (and it is indeed cool), now let’s turn to the millions of losing tickets.

Humans don’t structure the world through grue rules, so we don’t register the bulk of these corrupting distributions as something to be noticed, let alone tracked, and for sure not fixed. It’s not that we’re choosing to ignore the problem, but that our glasses obscure the very taxonomy that would allow us to see it at all.

Just to make clear, this is not some bored philosopher inventing problems with the hope of tenure. We already see this playing out empirically; for example, take GSM-Symbolic. This benchmark took an original maths benchmark (GSM8K) and turned each of its problems into a template, regenerating it with different names and numbers while keeping the logical structure identical.

If a model passed the original benchmark by learning arithmetic rules, changing “Sophie” to “Liam” and 5 to 7 should be irrelevant. It was anything but irrelevant: performance varied noticeably across regenerations of logically identical problems. Similar issues were discovered in other studies by attaching irrelevant clauses about cats or other meaningless details in the question. As we noticed when we explored GenAI knowledge, if changing these details changes the math answer, then GenAI wasn’t following the rules of math to begin with.

So controlling for what’s being learned is hopeless from inside the distribution. But wait, there’s even more. Let’s put a pin in the issues around learning a rule; turns out controlling for what rule is being followed in runtime is also impossible.

Wittgenstein, as usual, makes things more intense

Goodman shows us 2 rules agreeing on all evidence and diverging beyond it. Wittgenstein (as read by Kripke), makes a claim that’s somehow worse: from any finite body of behaviour (notice, not observations to construct a rule from, but realtime behaviour), there’s no attainable knowledge about which rule is being followed in actuality.

Kripke’s example, riffing on Wittgenstein’s is deliberately a Mickey-mouse one focusing on simple addition. You’ve added numbers your whole life, right? When asked for 68 + 57, you’ll produce 125. For sure, that’s trivial, because you’re simply applying the rule “plus”, the rule you’ve always followed.

Have you though? What about quus? It agrees with plus for every pair of numbers you have ever actually added in your life, but answers 5 for any pair of numbers bigger than any you’ve previously encountered. Imagine 68 + 57 is the first sum you’ve done with numbers this big (because you’re a mathematician, and they ironically almost never deal with actual numbers). Plus answers 125, quus answers 5, and every calculation you have ever performed in your life is equally consistent with your having meant quus all along. So how can you tell what you should answer?

Wittgenstein’s own formulation (though he ultimately dissolved this tension): no specific course of action can be determined by a rule, because any course of action can be brought into accord with the rule under some interpretation. The full version of this will have you not know, internally, personally, what rule you’re actually following. I get how this sounds insane when separated from the entirety of Wittgenstein’s language analysis context, but luckily a weak version will suffice.

The weak version will have you not being able to prove, behaviourally to an external inquirer, that you’ve followed plus rule and not quus (as any finite set of past behaviours can fit both rules). This version still allows you to have an internal private “screen” on which you project your inner thoughts, and that screen will say (privately, in a manner you could never share with the external inquirer), “plus”. This diverges from Wittgenstein but is more accessible (Wittgenstein passionately argued this private screen doesn’t do this kind of work).

Cool. Now ask, does an AI model have an inner screen? I would say it most certainly doesn’t, but we don’t need to settle the question for our purposes. Even if it does, we have no access to it. All we ever get is finite runtime behaviour as seen by us as external inquirers, permanently. i.e., we can never tell what rule actually is being followed in runtime.

Let’s ground this. Like the finite data the model was trained on, its behaviour we observe at runtime is also finite, so which rule is being followed at runtime is undetermined as well. And when you evaluate on held-out data from the same distribution, you’re just asking for more sums inside the range you already fitted. You aren’t probing the plus/quus divergence at all, and as we said, you aren’t even aware there’s a boundary to probe around.

Folding all of this to underdetermination in the broad sense: our test suite identifies a family of systems, and we used to reassure ourselves that its members were close cousins, because a human wrote the code. Goodman and Wittgenstein gesture towards what that family looks like when a human didn’t. It contains plus, it contains quus, and much stranger rules than either. The data loves them all equally, so they are equally justified and look identical in every benchmark you throw at them. You can’t tell what rule was learned, and you can’t tell what rule is being followed at runtime.

Read: you can’t tell if your test suite is green or grue.

The tiebreaker is a form of life, and the model doesn’t share ours

So underdetermination is real and inescapable, and the tiebreaker has to come from outside the evidence and the immediate context. How can we proceed? Well if you can’t beat them, join them; or in our case, let’s double-down on theory ladenness. Not for the theory in question, but for the overall governing theory we employ when we inhabit the world.

Wittgenstein’s way for doubling down is that following a rule isn’t a private, personal mental act, but a shared practice, embedded in what he calls a form of life. Meaning the whole inherited scaffolding of trained dispositions, agreed reactions and communal custom, against which “125 is the correct continuation” is simply what we humans do. On this reading of Wittgenstein, there’s no publicly available fact deeper than the shared practice that makes plus “right” and quus “wrong”. It’s right because we, the community of calculators, are trained into agreeing on it, and that agreement is foundational.

Goodman’s answer has similar shape. Why is green projectible and grue not, given that no formula separates them? Because green is entrenched. Entrenched means having a track record of successful projection in the history of a language community. Green has it plenty and grue has none. Notice this isn’t a claim about ontologically “real” truth, but of a shared communal history. Projectibility isn’t a logical property a predicate holds by itself, rather it’s a status earned within a community’s accumulated practice of using it and having it work.

You might think this sounds somewhat similar to Kuhn’s all encompassing paradigm, and you’d be correct. The emphases are different, but the consequences are similar in spirit. What’s defined to be natural is indeed defined as such. Defined by a particular sub-community of humans sharing a history, language and form of life. There is no “objective”, transcendent rational foundation for induction. There’s only human-rationality. And sub communities of different rationalities at that.

This is already quite deflating, but it turns into a practical nightmare when coupled with software practices in the age of GenAI. An LLM does not share our form of life, it’s only trained on the textual exhaust of it. Its tiebreaker, the thing filling in for our entrenchment and our agreement-in-practice, is simplicity bias over the training data distribution. It doesn’t merely fail to prefer green over grue, rather its inductive bias prefers grue, because from within its internal language grue was cheaper to encode.

That is a form of life, in a loose sense, sure. It just isn’t ours.

What this means

We shouldn’t be surprised that a model’s projections are foreign to ours. The shortcuts it takes aren’t stupid, they’re perfectly rational within its own form of life. They’re just shortcuts no human would ever take, or even consider (or even imagine). The model isn’t so much a bad reasoner as a different kind of reasoner, running an alien inductive practice that happens to emit tokens in our language. Locke said the madman reasons correctly from mistaken premises; well, that would make a madman more trustworthy and legible than the model.

To distill this into one practical lesson to be learned:

You cannot treat a model’s reasoning and projection the way you treat a human’s. With a human you lean on an enormous shared substrate, a common form of life, a common sense of what’s natural, a common set of shortcuts you’d both never take, to fill the gaps that finite behaviour leaves open. That substrate is exactly what you don’t share with the model, so the gaps left by underdetermination can’t be filled by trust. They can only be filled by checking. This means testing isn’t quality assurance sprinkled on top of an otherwise trustworthy reasoner. For an alien reasoner, testing is the entire thing, and it needs to be adversarial to an unprecedented, almost absurd degree.

We’ll see how this lesson is cashed out in part 3. Meanwhile, especially since this has been a long one, let’s sprinkle a few bite-size takeaways:

<Obligatory LinkedIn So what do I actually do with this on Monday?>

Want to know if the rule that produced your answer was green or grue? Generate variance around boundaries that seem crazy and irrelevant. Change a name you’ve made up for an arbitrary example, add a dash to the entities, swap the positions of the numbers, paraphrase the sentence, drop the apostrophe. If the way the output changed no longer meets the expected logic, you may have detected a grue.

Report the worst group of a benchmark, not the average score. The average hides precisely the slice where the shortcut fails, and any assumptions about smooth, projectable distribution into the real world are meaningless when you’re dealing with an alien curve. More often than not, a model at 85% average and 40% worst-group turns out to be a 40% model in the real world.

Treat every public benchmark number as contaminated until proven otherwise, and possibly after if the proof was part of the PR hype. Always assume the benchmark failed at being sufficiently adversarial due to lack of imagination.

Accept that your instruments and your object may be wearing glasses from the same box, and so that their agreement is worth less than it feels or registers on paper.

<Obligatory LinkedIn So what do I actually do with this on Monday? />

Epilogue

So, let’s recap how exactly this trainwreck unfolded.

We conceded part 1 entirely and assumed induction works anyway. It didn’t help us at all. Nobody observes without glasses, and they tint what you see. This was true for your oracle, equivalence classes and your taxonomy waaaay before any model showed up. GenAI makes the tint much worse, because a family of rules stops being close cousins in any way meaningful to humans, and the tiebreaker deciding which one the model learned or followed is a random seed and a bias for simplicity.

This tint can’t be calibrated away, because we’re dealing with a rule that agrees with the truth across all available evidence. More evidence from the same distribution doesn’t help, because it confirms all these rules equally. Worse yet, we’re blind to the distribution that’s actually relevant. Behaviour tracking doesn’t fix the rule either, because there’s no way to tell what rule the model actually follows at runtime.

What breaks the tie is a form of life, a community’s entrenchment; not the transcendent view from nowhere. The model brings its own tiebreaker, but it’s an alien entrenchment, which more often than not is tuned in the wrong direction relative to the thing we care about.

This has dire consequences. We cannot verify a learned rule via agreement within the distribution it was fitted to. That’s asking for more sums in the range you already computed, and it can’t separate plus from quus, or green from grue. Every in-distribution benchmark, however large or prestigious, is structurally blind to distinctions its distribution never separates, and we only ever separate the ones we thought of.

And our usual trust-heuristics can’t bridge that gap. Not because they’re worthless, but because the inference from fluent, human-shaped explanation to human-shaped reasoning requires a shared form of life, and the model doesn’t share ours.

So, what now? Is it all hopeless gloom and doom?

Well, Wittgenstein told us where to look, and Popper how to look for it. You can’t read the rule off the weights and you can’t confirm it from inside the training distribution, but you can attack the boundary, and you can dial the severity to 11. You’ll need to extend yourself into the absurd and the irrelevant, because from inside an alien language, those might be rational and justified.

Goodman told us the other half: projectibility is earned, through a community’s repeated practice of projecting and surviving contact with the world. Nobody can hand your model that status. But a community can confer it, from outside, through a standing practice of adversarial checking.

Think of it this way: if testing is like buying lottery tickets, then the Waterbirds authors won the draw by having a brilliant insight and suspecting the correct thing in advance. Unfortunately, deciding “today I’m going to have a brilliant one-in-a-million insight” won’t get you far. What you can do is stop buying one ticket at a time. Every absurd variation you generate is another ticket; no particular one needs to be inspired, you just need enough of them, with different enough variations. You won’t confirm the rule is green, but you might just catch it being grue.

So next time: we stop trying to take the glasses off, because we can’t, and start composing enough differently-tinted views into a technicolor vision, that a grue-shaped rule has (almost) nowhere left to hide from.

Part 3 is all about imagining those absurd unimaginables and putting them into practice, including the model in its second role, as the thing doing the testing rather than the thing being tested.

Still a hoot. Marginally less depressing.

Comments

(must be logged on to comment)