· articles

The City Isn't Made of Emerald, Just a Few Stones Are

General disclaimer about the handwaving nature of these articles

While this blog features mostly longform content, the issues discussed and thinkers presented deserve their own 3-parts university course (and usually they have one).

Everything you read here is a handwaving summary, lacking to the point of liability.

Hopefully it will still be an interesting read.

At the Gates of the Emerald City, Part 1: The City Isn’t Made of Emerald, Just a Few Stones Are

I’ve been entertaining the idea of opening a blog on philosophically inclined testing for many (many) years. What finally drove me to actually open the thing was my thoughts on relevant themes and practices for testing in the new age of GenAI development and tooling.

Specifically, I was interested in exploring the connection between testing and induction (trivial), through the prism of the old and new riddles of induction (maybe less trivial). I wanted to see if the different emphasis between the old and new riddles of induction can inform us on the differences between traditional software testing challenges and the new challenges introduced in the wake of GenAI.

This series attempts a first pass through these ideas. This piece will be the 1st part of 3, in which we’ll lay the epistemological foundation of testing and its practical implications. Nothing that will shock you, for sure, but a needed stepping stone for part 2, where we’ll explore what happens to this foundation when the thing we’re testing (or the thing doing the testing) stops sharing our human notion of a rule.

From there we’ll proceed to part 3, where we’ll attempt to sketch a practical guide for the future, drawing on all the theoretical tools we’ve developed through parts 1 and 2.

Let’s start with a rather simple observation:

Prologue: What you actually learn when the dashboard goes green

The pipeline finished, the dashboard goes green. 120 passing tests, the bot posts “all green ✅” in the channel, and the release goes out. Cool. Let’s reflect on that for a moment and ask: what did we just learn?

Not metaphorically and in some moral philosophy meaning of life notation, but precisely. We executed a finite set of specific inputs against a specific build in a specific environment, and observed that each produced an output matching a stored expectation. Those are the technical raw facts of the actual observation. However, that wasn’t the claim. The claim being made in the channel is enormously larger. That the system works; that it’s fit to ship. Not that 120 stones are green, but that the whole city is emerald.

There’s a huge leap between the observation and the claim being made. This leap is part of all facets of human life, science, philosophy, and yes, also testing: induction. Like deduction it’s a form of rational inference. Unlike deduction, which can be demonstrated to always be truth-preserving (i.e. a valid deduction will always lead from true premises to a true conclusion), induction isn’t guaranteed to do so. Still, it’s considered a justifiable inference move in some cases.

There’s induction in the broad sense, which pretty much means any inference that isn’t deductive, and there’s induction in the narrow sense, which means projecting from a sample to the next instance (or all instances) of that population. All observed x-s have been y-s, so in some cases we feel it’s justified to conclude the next x will also be y.

This series examines the inductive leap through several modern lenses, which will hopefully lend themselves to interesting thoughts and insights. Let’s start with why the leap is unavoidable.

You can never test the city, only a few stones

You know how the joke goes: A tester walks into a bar, orders a beer; orders 2 beers; orders 65535 beers; orders -1 beer; orders “three” beers. All is well. A client walks into a bar, asks to use the toilets. The bar explodes.

So, why didn’t we test for the toilet? Well, besides the crushing imagination destroying pressure of being an employee in a late stage capitalist world? Take something trivial, say a function that adds two integers. With just 2 simple arguments you’ll get somewhere around 10¹⁹ input pairs. Test a billion per second and you’ll still be testing 500 years from now. By the time we depart the Mickey Mouse examples and get to something like a real input form, working off a stateful session over 6 browser versions, you’ll find yourself engaged till the heat death of the universe 🤷‍♂️. Great job security, though.

The combinatorics don’t explode so much as stop being meaningful as a number. And that’s before you count different ways to order the operations, the timing, the back button, the token that expires halfway through, etc. Of course the toilets didn’t make it to the test suite. All of this to say that this is the oldest known fact in the testing world: complete testing is impossible, even in principle, for any non-trivial system, forever, whatever your budget actually is.

For sure this does not come as a surprise, and that’s fine. We will however explore this triviality in a little more detail than usual. Specifically, we’re making explicit the fact that every act of testing is an act of sampling. You choose a tiny fraction of a tiny subdomain of a tiny subset of the possible inputs, states and paths, exercise them and observe. And then, because the whole point was to say something about the system rather than about your subset of chosen inputs, you project from the sample to the whole.

You examined a few stones and found them to be green. You concluded the city is emerald. We are, all of us testers, in the projection business. The test cases are the sample. “Ship it” is the projection.

Naturally, not all subsets and not all inductive projections are cut from the same cloth. Even before we do some philosophical deep dives, practically, professionally, no one is naive enough to glance at a few random green stones and declare the city is emerald. Of course not. We’ve come up with pretty nifty techniques to make our projections more reliable and our sampling punch waaaaay above its weight. To name a few:

  • Equivalence partitioning: split the input domain into classes you have reason to believe behave similarly, sample a single representative from each. Save orders of magnitude in time and resources.
  • Boundary value analysis: defects cluster at edges, so oversample the edges. Zero, one, max, max+1, empty, null, off-by-one, malformed. Toilet.
  • Risk-based prioritisation: weight the sample by consequence × likelihood, and spend the finite budget where being wrong is expensive.
  • And a whole lot of other techniques, enough to fill the shelves of all public libraries in the world (this mainly speaks to the dismal state of public libraries, but I digress).

At its best, a testing team’s projection is some version of statistical acceptance sampling. Devise acceptable quality limits, sampling plans, and sampling statistics informed by detection probability. Sample the batch / release, estimate the defect rate, and you can now state your confidence with confidence.

So, nobody really says “the tests passed therefore the system is perfect”. When the watercooler talk gets distilled onto the slide deck, it ends up saying something closer to: given this coverage, this risk analysis and the project’s history, the residual probability of a severe defect in production is acceptable to a stated tolerance.

That’s still an inductive projection from a small sample to a broader population, yes. But when done professionally and responsibly, it seems reasonable and justifiable.

Well, Hume would beg to differ.

Hume’s radical deconstruction of induction

The modern analysis of induction starts with David Hume in the 18th century. Hume pretty much obliterates the rational standing of induction. He doesn’t merely say that it doesn’t always work (that’s a given as it’s not deduction), or that it’s sometimes unjustified. Take the best possible case of induction you can think of, and Hume claims you have NO rational justification whatsoever for it. That’s pretty radical.

Possible justifications for induction

Let’s say all observed copper conducts electricity, and we want to use induction to claim that all copper conducts electricity. Hume breaks this into two possible paths:

Justify through Demonstrative Reasoning: This means through relations of ideas / deduction. Hume tries to see if there’s a sort of a priori justification for the claim that unobserved instances will resemble observed ones. Usually this is done through pointing at some logical absurdity, contradiction or metaphysical necessity (though Hume hated the latter) that would force induction as a consequence. However, there is none to be found. The world could have easily been one where the future doesn’t resemble the past (and Hume’s entire point is that maybe it is).

Justify through Probable or Moral Reasoning: Having eliminated deduction, Hume considers whether the foundation can come from empirical experience; reasoning from past observations of cause and effect to predict future occurrences. Basically, this would claim that since induction has worked exceptionally well in the past, we can count on it working pretty well in the future (and specifically for using it in this next instance for the copper thingy). Surprisingly, this isn’t as immediately circular as it seems at first glance (there’s some nuance around induction in the broad and narrow sense and which is used where in the full argument); but at the end of the day, this does derail into circular reasoning pretty fast.

Hume identifies that all available justifications of induction have to rely on some version of uniformity assumption (the future will resemble the past, laws of nature exist, etc.); and that assumption is in itself unjustified without induction already being established. The bridge is held up by the bridge.

The circularity can be made vivid by imagining a devout counter-inductionist. If the elevator pitch for induction is “more of the same”, for counter-induction it’s something like “time for a change”. You say to your counter-inductionist friend: hey, remember all those times your counter-induction failed miserably? You said it was time for boiling tea to be delightful to gulp down, and burned your tongue? He shrugs and replies: You’re right! It failed miserably in the past, so I think next time it will do great! I’m justified in continuing being a devout counter-inductionist!

You might say counter-induction seems like a stupid, circular, unjustified belief system. And you’ll be right, of course. The problem is, according to Hume, our induction based rationality is just as circular and unjustified. He doesn’t think we should stop using induction (or that we can, for that matter), just that we’re not rationally justified when doing so.

In our everyday professional lives

Translating the above to our day-to-day professional context, our projections from a passing test suite to “the system works” rest on a uniformity-esque assumption about the system: the code paths, inputs and internal states, constraints and processing steps, the coding style in the development team etc. We have decades of aggregate experience about how these creatures are created, maintained and behave, where the pitfalls are, what the person in charge would’ve overlooked, and what the technological safeguards would’ve let slip through.

All that experience may offer some form of mitigation, or promise a better ROI, but they have no bearing on whether or not this uniformity assumption is justified, rationally. Remember, scientists projecting from past samples get to depend on the broad, hard-won regularities of nature, and Hume’s argument dismisses those just as hard.

Our situation is much worse. Yes, nature has some weird discontinuities, but the discontinuities in software are denser, deliberate and adversarially placed. Because of course Dave made this codebase non-uniform on purpose last Thursday, attempting to fix some customer escalation, and told nobody (and for sure there’s always a Dave). The area you’re implicitly filling between your tested points has been engineered to be full of sharp edges, with little smooth surface between them.

Reactions to Hume

So, Hume is kind of a downer and rains not only on the parade of scientific inquiry and enlightenment flavoured optimism, but on the very foundations of human rationality themselves. Unsurprisingly, this ruffled a few feathers (not to mention waking Kant from his dogmatic slumber [yes yes, that was about causation, we know] etc.). A lot of very smart people have offered a lot of very smart ways to fill the hole Hume punched through human rationality.

Let’s go through some of the main reactions to Hume, and see how they fare, and what they might mean in our day-to-day translation of the riddle of induction:

Just say yes

This is known as the linguistic response to Hume (AKA Strawson’s dissolution response). A handwaving summary of it is: performing inductive inferences under favourable conditions is just an integral part of what we mean by human rationality. So when Hume asks us to rationally justify induction, he’s actually asking if it’s rational to be rational, and we should not be ashamed to answer him with a resounding YES.

A more serious detailing of this objection can be made by imagining what we’d say if Hume required us to rationally justify deduction, without allowing us the resources of deduction itself. Of course it’s not something we could ever do, but neither is it a realistic demand; failing to meet this demand doesn’t reflect on us or on deduction, it just means that deduction is foundational to reason (philosophers would say it cannot be verified by a more basic framework). This does seem to answer Hume.

The gaping hole in the linguistic response is that induction and deduction do differ significantly. While it’s true we can’t offer a more fundamental justification for neither of them, deduction still has a forward looking justification of the sorts that induction doesn’t and could never have.

We can build truth-tables and prove to our own satisfaction that deduction will live up to its promise, meaning that a valid deduction process guarantees that true inputs will always lead to true outputs. This is a solid justification to rationally rely on deduction. Unfortunately, induction can demonstrate no such thing to no such degree.

So to be clear, this isn’t a case of philosophers doubting things for doubt’s sake. Induction seems much shakier than deduction, and the linguistic objection to Hume ultimately fails.

Math can sort things out

Hume’s attack seems fitting for the implicit ways us humans infer in our day to day life, and maybe even for the ways philosophy of science builds theories out of distinct observations. Fair enough. But what about the specific fields where inductive inference is done through extremely robust, detailed, mathematically proven mechanisms? Can’t those provide a robust enough foundation to justify (some cases of) induction from?

Hey People Who Actually Know Math

Forgive the mishmash train wreck of the next 2 paragraphs. Just gesturing in an extremely handwaving manner to make a general point 🤷‍♂️

For example, if you ask someone who professionally does political polling for a living, or works for a government statistical office, or even is just a principal data engineer, they might harken to combination of the central limit theorem and the law of large numbers as the foundational justification for simple, well controlled induction.

Basically they’ll say you can get as close as you’d like to whichever confidence level you choose for a projection, by randomly sampling a large enough amount of a population. Notice, sampling absolute amount, not relative proportion. This means that even when our population size is billions, we could randomly sample just a few thousand representatives of that population, and get > 95% confidence level in projecting their observed attribute to the entire thing.

It would be really nice to have induction justified by something as robust and logically verified as math, but unfortunately this reaction to Hume is also broken. There are some nitpicky technical pitfalls in how the math actually maps to in the actual world (the math allows you to aggregate repeated frequentist statistics, and doing that in the real world might sneak in some inductive presumptions); but more relevant to our discussion is the requirement for a random sample.

The math says random sampling of a population will give you these confidence levels. Random here means roughly that every member of the population has an equal probability to end up in the sample we examine (there’s some mathematical nuance I’m skipping); this opposed to biased sampling, which just means it’s not random. As an easy example, if you’re doing political polling by surveying people in the street, you’ll get a biased sample, because parents of young kids usually stay at home to take care of them so they’ll be underrepresented in your sample (these Mickey-mouse examples sometimes do actually happen in the real world, e.g. Roosevelt 1936 won the election, despite a HUGE 2.4 million sample projecting he’d lose due to bias).

Well, is our sampling of things we inductively infer about random? Well, the situation is almost, definitely, completely… 180° from being a random sample. Our situation is so horribly biased it’s actually dumbfounding.

Take the conductivity of copper for example: you may think we have a geographically varied, temporally spaced sampling of copper; after all, it was mined all over the globe throughout a lot of human history. Compared to the sampling that drives most of our day-to-day inferences and projections, it seems amazingly more random.

But now think of how our sampling is positioned compared to the actual population we project on. The universe is full of copper, and has been for billions of years. Compared to the actual population, our sample is incredibly biased - it’s from a teeny-tiny single spot in a teeny-tiny sliver of time. It’s like doing a political poll in a single room of a single house, and claiming you’ve randomly sampled the country. To stress, the issue isn’t the sample size (our math mitigates that), it’s its bias.

Transposed to our professional lives, our test suite isn’t a random sample of the application it covers. As we’ve mentioned, it was designed and built on decades of technical and psychological insights and lessons learned about where risks reside and where defects accumulate. And doing that was a reasonable thing to do, for sure. But it isn’t random sampling, so it’s downstream from taking inductive projection as a given, not upstream where it can justify it. BTW, missing the toilet scenario is another case of biased thinking that generates a biased sample (of imagined scenarios).

Another math-based counter to Hume invokes Bayesian analysis (actually raised by Price shortly after Hume published his work). It also fails, and in interesting ways (it’s all about the priors baby), but we’ll save that for a more dedicated exploration of Bayesianism. Sufficient to say that math can’t properly counter Hume and provide a rational justification for induction.

If anything works, induction is guaranteed to work

This is also known as the pragmatic vindication of induction. And it sets its sights lower by making a critical distinction: it doesn’t prove there is an answer to our question (maybe the conductivity of copper fluctuates wildly, never settling on a final figure). Rather, it claims that if there is a valid pattern to project our sample onto, induction is the only methodology that’s guaranteed to eventually converge on that pattern.

Maybe you could get there quicker by some other means (maybe reading tealeaves, maybe by flipping coins), but none of the other methods give you the guarantee that you will (again, assuming there is an actual pattern to converge on). So given we do need to infer things about the world in order to live in it, the only pragmatically justifiable bet is on induction.

This approach was mainly proposed by Reichenbach, and technically it invokes what’s called “the straight rule”, meaning we project the observed frequency from our sample to the entire population in a straightforward manner and correct as more data comes in (think of it as perpetually assuming that our current sample is random).

It won’t prove our projection is correct, but Reichenbach does show that if there’s an answer to converge on, induction eventually will (and shows counter-induction doesn’t achieve that). Other methods might get you to the correct answer more quickly, but you won’t get the guarantee, and that’s where the justification lies. It’s a genuinely good argument, but still has two major caveats:

The 1st one hangs on the “eventually” the argument sneaks in. As the saying goes, in the long run we’re all dead, that tells me nothing about tomorrow. Convergence in the limit promises nothing about the next observation, and as humans, the next observation is the one we actually face. Knowing that when the universe recollapses into a black hole induction will have been vindicated is nice, for sure. Nice, but useless.

The 2nd caveat is quite surprising. Turns out the same convergence argument works in infinitely many ways. Meaning that yes, the inductive straight rule is unique compared to reading tea leaves, but there are infinitely many straight rules to choose from. “Predict what the straight rule predicts, but inflate by 20% until 2036” also converges in the limit. Asymptotic convergence (read: after an extremely long time) can’t distinguish the sensible method from the deranged one until LONG after both have stopped mattering.

Notice that this problem almost turns Hume’s riddle on its head. Hume says we can’t justify even a single instance of induction. The straight rule answer to Hume seems at first to provide a light at the end of the tunnel. Then infinitely many rules show this light to actually be an incoming train, and what we’re left with is a trainwreck, not a usable answer.

All these inductive rules to choose from, all with equal justification and no rational way to choose between them. Reichenbach tried to solve this issue, as did many others, but they all failed. This kind of flipping of Hume’s riddle will come back with a vengeance in the 2nd part of the series as Goodman’s new riddle of induction. Until then, we’ll leave it at that.

How does the pragmatic vindication fare in our day to day professional context? Our version is the track record defence, specifically that our process has history. Our escape rate is low, change failure rate is trending nicely, this suite has carried us through 10 prior releases. Our method has been converging, so let’s keep running with it.

The caveats in this context are similar to the broader inductive context. The 10 prior releases are a nice series to examine, for sure, but a converging series tells you nothing about whether the next member of it will follow the trend, or will be an outlier. And usually we have shamefully short series to draw on.

Worse, it doesn’t even tell you what the trend it’s converging on is. The test suite may be a massive misguided error waiting to be uncovered by the user who uses the toilet. Similarly, defect criticality and testing resource allocation grow out of our current taxonomy, and that may shift dramatically when new user groups are onboarded, an underlying tech stack changes, or new organisational KPIs are introduced.

Our track record is subject to definitional and language changes that might paint it as a significantly different creature altogether. Just think of what happens to your P1 defect distribution when a new priority / severity taxonomy is introduced, and how it can retroactively invalidate your projections and inference.

At the end of the day

None of them (or the many others who attempted a clever response) managed to end the argument, and IMHO every objection I’ve heard collapses to the linguistic objection, which only ends up vindicating Hume. So looks like we’re stuck with having to infer inductively, but not having any rational justification for doing so.

Or not, if we embrace Popper’s point of view. He doesn’t rescue induction, but brilliantly changes the question and frame of thought.

Popper, and why green was never what we looked for

Hume tells you the leap has no foundation, which is unsettling but not actionable, as we seem to have no choice but to still infer inductively. Human psychology and the practicality of everyday life forced us to look at the green stones and infer about the emerald city. Popper can be worked into a surprising practical advice: don’t look for green stones, but red ones.

We’ve briefly visited Popper’s approach in our intro to Kuhn. Let’s recap. Popper accepted Hume’s verdict and drew a radical conclusion from it: stop trying to confirm theories, because you can’t. No finite number of observations establishes a universal claim. But there’s an asymmetry, as a counterexample can deductively refute one (greatly oversimplified, just to make the general point). You propose a bold theory, you try your hardest to break it, and if you fail the theory is corroborated, which means it “has survived our attempts so far” and emphatically not “has been shown true”.

So a naive reading of Popper would have us not counting green stones, but instead looking for red ones (and yellow, and white, etc.). Failing to find them, we can say we’ve failed to kill the theory that the city is emerald, and that’s seemingly the best we could ever say.

Adopting this naive interpretation of Popper as an actual answer to Hume is a non-starter. It is actionable, and can lead to a more well-rounded point of view, but at the end of the day, without induction it holds no practical value. For example, it’s hard to say what value corroborated theories have over theories that haven’t been tested at all. And think, if you really can’t rely on past success as a guide to the future, would you use your banking app?

So Popper isn’t a real counter to Hume (not surprising, as Popper begins by agreeing with Hume). Still worth taking a moment to reiterate the insights this point of view offers us, practically.

A Popperian point of view reframes the dashboard. A green suite does not say the system works. It says our attempts to refute the claim that this system works have failed. Put that way, the crucial question stops being how many tests you ran and becomes: How severe were the attempts?

Severity is Popper’s own term. A test that a theory was always going to pass tells you almost nothing. A test the theory would have very likely failed, had it been false, tells you a great deal. 1000 tests poking gently at the happy path are a weak corroboration. 10 tests that each represent a serious, well-aimed attempt at murder are a strong one.

This approach somewhat turns coverage metrics on their head, as coverage inherently measures how much of your code you executed, not how severely you tried to break it. You can drive branch coverage to 95% with assertions so loose that no realistic bug could fail them (in fact, you can do 0 assertions and still get the coverage 🤷‍♂️). So coverage counts the stones you’ve seen, but it’s silent on whether you did that in different lighting conditions (OK, this metaphor is getting increasingly hard to work with).

<Obligatory LinkedIn So what do I actually do with this on Monday?>

Since this is a 3 part series where the takeaways are waiting in the last part, let’s left shift a few of them here:

Write the leap explicitly. Next to “ship it”, add the sentence that makes it valid, i.e. the uniformity assumption: assuming the untested payment providers behave like the tested one (similar to the risk planning thingy). A surprising number of those statements will be flagged as false the moment somebody reads them out loud.

Ask what each test would have caught, not what it covers. If the answer is “nothing realistic”, it’s a green stone you’ve now looked at twice. Or, calling back to the 1st article on the blog, it’s just kicking up dust, not shining light.

Stop quoting coverage as if it were severity. It measures what you executed, not how hard you tried.

Develop a nihilistic, hedonistic lifestyle devoid of all pretence of meaning or content, because the future is uncertain.

Well, maybe best to wait and see about that last one for the time being.

<Obligatory LinkedIn So what do I actually do with this on Monday? />

Epilogue

So, at the end of part 1, let’s survey the landscape we’ve uncovered. We can’t test everything, so we sample. Acting on our sample means projecting, and projection is induction. Hume shows that we have no rational justification for this; not in the sloppy cases, not in the best of cases, not ever.

The linguistic response ends up conceding the point and rebranding it as a virtue. The mathematical response is great for random samples, which ours will never be. The pragmatic vindication buys a guarantee that pays off at the end (for a useless definition of “end”); but then hands the identical guarantee to infinitely many rules that disagree where it practically matters.

Then Popper, who doesn’t rescue induction and never claimed to, but does change what we’re looking for. Green stones were never the evidence, the absence of red ones is, and that absence is worth something only in proportion to how hard we looked for them.

Our test suite isn’t a random sample of the application, it’s a deliberately biased one, built out of decades of hard-won lessons about where defects cluster. That bias is a feature (it’s most of the craft), but the inductive leap is still not rationally justified. Our track record is a short series, and it’s expressed in a taxonomy that can be redrawn next quarter, retconning our projections.

The city is not made of emerald, only a few stones are green. Everything else is a projection resting on a uniformity assumption that isn’t a robust assumption to begin with. We’re on much shakier ground than the one physicists get to stand on, and Hume wasn’t impressed with theirs either.

If you think that’s depressing, in part 2 we’ll see how much worse things can get when GenAI disrupts and breaks our hard earned lessons. Then, even when we concede to Hume and just take induction as a given to get some results, it still won’t be enough. But do not despair, as part 3 might pave a new way forward.

So next time: We enter the Emerald City and discover it’s not the stones that were green, but the glasses we were looking through all along. Then, shockingly, we’ll learn the glasses aren’t even green, but grue.

Stick around, it’ll be a hoot. A depressing hoot, but still.

Comments

(must be logged on to comment)