Kuhn! In the Software Industry - Part 2: For and Against Method
In the first part of the series we focused on Kuhn and his revolutionary way of doing philosophy of science. Unlike normative philosophy of science (i.e. what science should look like, conceptually and rationally), Kuhn saw himself as more of a historian of science (we might consider him an historian / sociologist), giving an actual account of what science looks like in the real world, then working backwards from that to the concepts and patterns that describe it in a meaningful way.
This second part adds some much needed nuance and detail to Kuhn’s foundations. We’ll start by making explicit the immediate criticisms raised against his work. That gives us the context to examine two philosophers whose reactions to Kuhn form two sides of a larger, more balanced whole.
So, what’s wrong with Kuhn?
Quite a lot, actually (his taste in ties, for starters). Kuhn’s point of view was revolutionary, but some of the details of his work were… shall we say, contested.
The historical account is a stretch at times. Going through Kuhn’s own case studies makes his distinct phases - normal science, anomalies, crisis, revolution, new normal - seem more imposed than discovered. At the very least, the transitions are messier, slower and not as total as the model offers. Old and new paradigms often coexist for years (decades, even), and practitioners across them communicate quite effectively, without a hint of incommensurability1 Kuhn himself did pull incommensurability back in his later work, from a global failure of communication to a local, taxonomic affair. .
The vocabulary is loose, which raises internal consistency concerns. As we mentioned in part 1, paradigm is badly defined, and the definition drifts across Kuhn’s later works. Worse, Kuhn deploys several separate criteria for the same terms, without demonstrating that they stand or fall together.
The N is small and biased. Kuhn built his model on a handful of episodes, mostly from physics and astronomy. These are the fields with the strongest theoretical unification and the most dramatic upheavals. He does reach further (Lavoisier and oxygen for example), but stray into geology or most of biology, and the well defined crisis-and-revolution pattern becomes harder to find. It’s not that there are no candidates; it’s that they’re far messier than the ones Kuhn cites.
These and other issues are ultimately fine, especially for such a groundbreaking work as the one Kuhn produced. It isn’t perfect and it doesn’t need to be. But for our purposes, there’s a more practical issue: the picture Kuhn paints is extremely fatalistic. Paradigms harden; anomalies accumulate; crisis arrives; the winner is settled by persuasion, funding and funerals. For all its grounding in the practical, Kuhn’s work remains close to useless as forward looking advice. For that, we need the reactions to Kuhn, and the formation of post-Kuhnian thought.
Reactions to Kuhn
Kuhn’s work was enormously influential inside and outside philosophy of science. The Structure of Scientific Revolutions made its way onto syllabi ranging from business management to sociology and anthropology, and sent ripples that changed entire fields of thought.
Within philosophy of science, Kuhn left the field in shock. All the normative justifications for science’s supposed epistemic merits fall flat if in actuality no one follows them. If science is just another sub-society with arbitrary rituals, non-rational power struggles and crisis fatalism, why do we grant it the prestige and authority it seems to warrant?
Well, one might say science earned its place through its remarkable track record. Science seems to be the best epistemic endeavour in human history, which raises a central question in Kuhn’s wake: how? How can science be so successful if there’s nothing epistemically special about it? Kuhn’s own answer centres on science striking an unusually effective balance between the dogma of normal science and the Wild West of crisis science, the “essential tension” between tradition and innovation.
Whether you find his answer convincing or not, every post-Kuhn reaction now carries the burden of answering how science manages to succeed. Generally, the reactions to Kuhn’s initial work clustered into the following groups:
Acceptance through division of labour. Accept the core tenets. Yes, social factors and power dynamics affect the practice of science; and since science is done by humans, it will forever be biased and never entirely rational. But that’s a matter for historians and sociologists. Philosophy of science reflects on the parts of science that can be rationally reconstructed. On this view Kuhn marked an important and valuable boundary. Inside it rational epistemology proceeds as before, and does have an effect to the proceedings of science as a whole.
Rejection. Reject Kuhn’s description of what science is and how it’s carried out, usually by attacking the details (which, as noted, had some real weak points); i.e., things aren’t as bad as Kuhn said. A second form of rejection insisted that what makes science science is still normative. If most of what happens in labs is different and can’t be reconstructed to match, then that isn’t really science but merely science-adjacent2 Popper himself noted that much of the actual work scientists do doesn’t fall into the narrow scope of falsifying empirical claims (his normative criterion). For example, coming up with theories has nothing to do with falsification, but there’s nothing to falsify without that part. . Either way, science stays epistemically special in virtue of its normative and rational analysis.
Doubling down. Accept Kuhn’s analysis as a step in the right direction, then radicalise it. This usually meant not merely weakening the social standing of science but demolishing it. The reductio ad absurdum of this category lives in the Strong Program in the sociology of knowledge, which degenerates into nullifying the very notion and value of truth3 IMHO at least. Full disclosure: I’m extremely biased against this school of thought. I hate it 🤷♂️. .
Fine-tuning and detailing. Kuhn’s historical account can be made to fit a range of close theoretical frameworks. Many post-Kuhnians accepted Kuhn’s point of view, but developed it into more nuanced and detailed mechanisms. They usually tried to describe mechanisms that, while still matching the historical and sociological record, preserved enough epistemic merit to justify science’s special standing.
An example of this kind of fine-tuning is where we’ll delve into post-Kuhnian thought, through the work of one Imre Lakatos.
Lakatos’s solution: telling honest modifications from ad-hoc excuses
Lakatos’s focus was the tension between two opposite extremes. Popper’s normative criterion, stating that an empirical claim is scientific if and only if it can be falsified, would have scientists drop their theories at the first observed anomaly4 This is a straw man dramatisation of his view scoped for this article. We will visit Popper’s view in a less caricatured manner in a later dedicated piece. . Kuhn’s normal science would have them suppress counter evidence until the complete break of crisis science, which had no normative, rational, evidence-based mechanism at all. One extreme isn’t practical: good theories are very hard to come by and shouldn’t be easily dropped. The other is epistemically reprehensible.
Lakatos wasn’t the only one who found Popper’s view, while extremely valuable, hard to operate in practice, and agreed with Kuhn that it had almost nothing to do with what was actually going on in labs. Lakatos and others tried to keep the value of falsification while mitigating its implications. After all, almost every theory has anomalies, we need a license to explore them before dropping the theory they inhabit. One common mitigation makes use of a theory’s auxiliary assumptions. A naive reading of falsification says a failing test kills a theory. But a failing test never falsifies just one thing. It falsifies the conjunction: the theory, the implicit claims it leans on, the instruments, the background conditions, and the competence of whoever ran the test, etc.
Kuhn had his paradigms weaponising that mitigation into a wholesale suppression of refuting signals, but that doesn’t make the mitigation itself unwarranted. Various thinkers cashed out the concept differently, theoretically and practically, trying to strike a better balance between epistemological merit and real-world usability. Quine’s web of belief is one popular example, sitting in the normative-theoretical corner. Lakatos’s model is another, occupying a more sociological-structural vantage point.
Naive, bite-sized falsification won’t do. Both Quine and Lakatos answered by changing the unit being judged. Quine argued that our entire web of belief faces the tribunal of experience together. Lakatos’s move was from the ground up: he doesn’t judge a theory against its evidential support at all5 As usual, that’s a hand-waving dramatisation. , he assesses scientific research programs over time.
Research programs are the loose social, practical and theoretical structures that the practice of science is actually made of. A research program has a hard core and a protective belt. The hard core is the set of commitments held constant by decision, not by evidence. Opening them for debate usually means the end of the research program in its current form. Around the core sits the protective belt of auxiliary hypotheses and assumptions. When results go badly, researchers modify the belt, not the core.
So far this sounds like a more nuanced cashing-out of Kuhn’s paradigms and normal science. It is. Lakatos gives the descriptive analysis first, then builds the redeeming normative part on top of it. He judges a research program as healthy or sick according to how the belt gets modified. Some modifications are honest, valid, epistemically acceptable moves (progressive); others are ad-hoc excuses (degenerating). The belt modification allows scientists to hold on to their theories (the core) without collapsing into blanket dogmatism, while still calling out the ones who abuse the process.
Lakatos, like Kuhn, argues from historical case studies. His most famous one is Uranus in the 1840s. The planet’s orbit didn’t match the predictions astronomers derived from Newtonian mechanics. A naive falsificationist would see this as refuting Newton. Instead, astronomers patched the belt. Urbain Le Verrier and John Couch Adams independently suggested an unseen planet and computed where it would have to be; in 1846 Johann Galle pointed a telescope at the coordinates and found Neptune within a degree of the prediction. The patch made a novel prediction that was risky (independently testable, and it could have failed), and it came true. That’s the hallmark of a progressive research program.
So far, so good. The next chapter of the story shows what a degenerating research program looks like. 13 years later, Le Verrier faced another anomaly: Mercury’s perihelion advanced more than Newtonian mechanics could account for. Reasonably, he tried the same move, divining an unseen planet he named Vulcan, with a prediction of where to look for it. Everyone looked, for decades. There was no Vulcan (the anomaly was finally accounted for in 1915, by general relativity).
What makes this case degenerate isn’t the failure. Vulcan started as a perfectly valid progressive patch: risky, novel, testable. It happened to be wrong, but that is exactly what a good conjecture is allowed to be. Degeneration came from what happened next: suggesting successively smaller and dimmer Vulcans, then a diffuse asteroid belt, then a nudge to the exponent in the inverse-square law; each one cut to fit only the anomalies already discovered. That’s the mark of a degenerating program. It patches the belt endlessly, but only to accommodate what has already gone wrong; explaining retroactively, never anticipating. From the outside, on any given Tuesday, the two activities look identical. Epistemically Lakatos saw them as opposites.
So Lakatos injects normative judgment and extra rules into normal science, heading off the fall into Kuhnian crisis science. He despised the rule-less Kuhnian description, calling it a matter of mob psychology6 The psychoanalysts in the room might connect Lakatos’s rejection of free-for-all crisis science to his family’s experiences in WW2 and the Holocaust. His mother and grandmother died in Auschwitz. The analysis is irrelevant to the theory, but not completely unfounded. . His solution is doubling down on Method with a capital M, strict and heavily rule-governed. This also yields a detailed, value-laden explanation of science’s track record. Science succeeds because of its strict method, and that method is epistemically good. Case closed.
Time to pause the history lesson and see what parts of Lakatos transpose onto our software testing team from part 1. Hopefully, astronomers hunting for missing planets and QA teams hunting for defects could be shown to be playing a similar epistemological game.
Your project methodology is a Lakatos research program
Take heavyweight planned testing: waterfall proper, or the meticulously groomed agile variety with its risk register, entry and exit criteria, etc (just make sure it really is agile, and not the yeah, for sure we’re doing agile, we have no specs, no documents, no rituals, no retrospectives, no point system and no cadence meetings “agile”).
The methodology has a hard core: quality comes from systematic, pre-planned verification against specified requirements; coverage of the specification is a reasonable proxy for coverage of risk; testing is a measurable, plannable, estimable activity. Drop those, and you’re no longer doing planned testing altogether.
It also has a protective belt: the case templates, the traceability matrix, the severity taxonomy, the entry and exit gates, the tooling, the sign-off ritual, the reporting cadence. This is where the patching happens, and it means every post-incident action item in the history of software has been a belt modification. That’s fine; the mechanism itself is value-neutral. A patch can be an honest, justifiable modification that makes predictions (progressive), or an ad-hoc excuse that only fits the already available data (degenerating).
Let’s differentiate the two in light of the Neptune case. Does our testing approach make novel predictions? Not “does it generate numbers”, but does it tell you, in advance, where the defects are in a way that could turn out to be wrong?
A good model says: the defects will be in the reconciliation batch, in the edge cases around the fiscal year boundary. Then you point the telescope where it told you to, and there they are. That’s the recall half; you also point it somewhere else, find little, and get your model’s precision as well. The model has earned its keep, or maybe it failed, but either way that’s a verifiable, falsifiable claim7 Just make sure this prediction isn’t derivable from last quarter’s defect list; if it is, it’s retrodiction wearing a prediction’s clothes. .
Mere coverage percentages, on the other hand, are not predictions, just descriptions of an artefact you built yourself. “x cases covering y% of requirements” forbids nothing and risks nothing. And if the belt-patches are limited to changing the coverage, always responding to new defects by adding exactly the cases that would have caught them (and only those), you aren’t predicting; you’re merely retrodicting. That’s a hallmark of a degenerating testing program.
Sometimes the degeneration is even easier to spot, because it’s a phrase repeated wholesale at every retrospective. Ask what the standard explanation is when a defect escapes:
- The process is great, it just wasn’t followed properly
- The requirements were incomplete
- We didn’t have time to do it properly
Any of these can be true in a given instance. Used every time, they’re unfalsifiable, non-predictive belt-patches, and what they guarantee is that your methodology can never fail, it can only be failed. Vulcan is out there somewhere; just give us one more sprint and we’ll find it.
Remember however, that exploratory context-driven testing is also a testing program, with its own hard core: testing is a skilled human craft that can’t be reduced to procedure; the value of any practice depends on context; the tester is the instrument. Fine, good, defensible, if that’s your cup of tea. So when that program suffers an escaped defect, what’s the ad-hoc belt patch? Usually a more-or-less rude version of “we need a better tester”.
Also unfalsifiable. Also degenerate. If every miss is explained by individual skill and every catch is explained by individual skill, you’ve stopped making predictions and started doing character assessment. An evergreen defence of exploratory testing that no outcome can falsify is in exactly as much trouble as a Gantt chart that survives any possible feedback from reality.
Both programs can be progressive, of course. Usually neither is, because people are usually the worst (not you and I, naturally, we are outstanding, honest professionals). Whichever you choose, Lakatos would just want you to be honest about one central question: is my methodology currently predicting anything, or is it only explaining? If it’s the latter, it’s time for a change.
Feyerabend would also have a word
Paul Feyerabend was Lakatos’s friend, colleague and favourite antagonist. He thought the entire project of pinpointing “the correct Method” was misconceived. The search for a single, universally applicable standard, licensing this and forbidding that across the whole of science, was itself a mistake. Feyerabend’s reading of the historical record was that every episode of scientific advance we now celebrate involved somebody breaking whatever rule prevailed at the time. Any methodology strong enough to be enforceable would have strangled the science it was meant to protect.
Feyerabend is hard to read. He deliberately juggles words and meanings, and makes the medium into the message (Yes, I’m aware that’s not really what the expression means). He’s so intent on not being reduced to a single point that it’s very hard to ascribe any specific “message” to him at all. He’s ironic half the time, contradicts himself the other half, then comes back for another round with everything in reverse. No wonder he preferred to be thought of as an entertainer rather than a philosopher. Still, there’s a deep argument in there, with enough content to sink our teeth into.
Lakatos and Feyerabend planned to have it out formally in a joint book, For and Against Method, with Lakatos writing for and Feyerabend against. Feyerabend wrote his half, but Lakatos tragically died in 1974, aged 51, and the case for method was never written up as a book. Feyerabend’s half was published as Against Method and became quite influential, though diminished by the absence of its counterpart. What survives of Lakatos’s side is his LSE lecture course and his and Feyerabend’s correspondence, published together in 1999 under the original title.
Lakatos yields plenty of actionable insight in a modern setting. Feyerabend is by far the more fun, but we’ll have to take him more seriously than the usual TL;DR summary will have us do. So, to stress again, even though he often described himself as an epistemological anarchist, Feyerabend does not claim scientists should work with no method, rhyme or reason, or that all approaches are equally good. His “anything goes” approach is sarcasm. If you forced him to crown a single universal method, then sure, “anything goes” is the only one that could possibly fit. That, though, is a reflection on the absurdity of the demand, not on his actual view.
His defiance makes him fun to read, but that’s not a positive argument and it doesn’t entail anything actionable. For that, we’ll walk through the substantive parts of his argument that do transfer to testing, using our mileage under Lakatos as context along the way.
The historical argument against method
Feyerabend’s main case study was Galileo. Galileo asked people to trust his telescope at a time when no accepted theory of optics justified it; he met the (good) objections to a moving Earth with ad-hoc moves invented on the spot; and he was an extremely effective propagandist. He advanced science by violating the methodological standards of his period. He also happened to be correct (This doesn’t mean rule-breaking is a net positive; most rule breakers achieve nothing). Any methodology strict enough to have stopped him would have deprived science of a critical advance, and that covers more or less every methodology proposed since. As we noted, this reflects more on the absurdity of the demand for a monolithic methodology than on any specific answer Feyerabend is willing to give.
In testing, think about your project’s last 5 defects that actually mattered. Not the ones that got logged, but the ones that changed a decision. Where did they come from? Probably from someone poking at something they weren’t assigned. A developer building a demo. A support engineer with an unreasonable customer and a hunch. A tester who side-quested off the plan because something smelled wrong. Maybe a contractor who got bored.
Now open your process documentation and find the section describing those routes. It isn’t there. Of course it isn’t there, it can’t be, because those routes aren’t procedurally reproducible and the document format only accepts things that are. Your documentation describes a testing method that systematically excludes the part that finds the bugs. Well, the bugs that matter.
Counterinduction
This was Feyerabend’s most constructive proposal. Deliberately develop working theories that are inconsistent with well-confirmed, commonly shared facts and beliefs. Not for contrarianism’s sake, but because a framework’s problems are often invisible from inside it. Some evidence only becomes visible as evidence once you’re holding an incompatible alternative; which means you sometimes have to build the alternative first and let the supporting evidence arrive later. Something like a gestalt shift, or if you’re feeling fancy, Kierkegaard’s leap into faith (my metaphors are getting fuzzier by the minute, apologies).
In testing, this means sometimes testing against the oracle rather than only with it. Treat the spec as an hypothesis rather than as ground truth. Assume the requirement itself is wrong and go looking for the user it would harm. Ask what would have to be true about the world for this agreed, signed-off and approved behaviour to be a disaster, then go and check whether the world is like that. This shouldn’t consume the bulk of your time, but it should have some time set aside for it.
Or put another way: treat the expected result column as the place where defects go to hide. That column’s whole point is to convert an open question into a binary, aggressively scoped assertion. This happens to be at exactly the moment when the open question was the valuable artefact.
Proliferation
This is counterinduction at the systemic level. Proliferation asks us to maintain multiple incompatible approaches, or whole theories, at once. Not as a transitional brainstorm to be forced into consensus, but permanently, as policy. No single framework can see its own blind spots, and you can only locate your unexamined assumptions from somewhere outside them, so make sure every position has someone on the outside looking in. In science, Feyerabend explicitly championed a plethora of partially contradicting, partially overlapping theories coexisting and competing.
In testing, this means treating standardisation as a coverage risk. Two testers with genuinely different mental models of the system will find different defects. That isn’t inefficiency to be optimised away; it’s a mechanism to be utilised. Pushing to make the team test “consistently” will get you beautifully comparable metrics, yes. It will also get you a team that has quietly stopped triangulating on anything.
Rational Reconstruction: rewriting the retrospective
In Kuhn’s original account, the victorious paradigm rewrites the textbooks after the crisis, and that’s presented as a brute fact about power dynamics. Lakatos has something similar, but paints it in epistemically positive terms. He called it rational reconstruction, and it means cheerfully putting a spin on the history of science to justify a program’s point of view. On Lakatos’s terms this isn’t cheating, because it’s supposed to include explicitly how the spin was produced, and to build an honestly persuasive case out of the factual record. You write the history as a tight narrative in which the research program’s method did well, and relegate the messy actual history to the footnotes. Think of it as a more honest and virtuous version of the near indoctrination Kuhn ascribed to the ruling paradigm.
Feyerabend rejected this outright as self-serving fiction. You write the history so the methodology comes out looking necessary, then cite the history as evidence for the methodology. It’s nice that Lakatos tries to keep everyone honest about the manoeuvre, and nicer still that he thinks the honesty makes it constructive. Feyerabend argued that at the end of the day it remains an exercise in wilful and harmful self-delusion.
In testing, you’ve probably lived through a version of this. The release was saved by a weekend in which the plan was abandoned and 5 people poked the system with their bare hands. The retrospective, however, credits the process improvements adopted afterwards. The entire war room is reduced to a resourcing note. That DevOps junior with the log-scraping script doesn’t even make the footnotes, because there’s no category they fit into. No wonder the resulting action items are always a heavier plan for next time, justified by a reconstructed history in which the plan is the part that worked. The process sets itself up for repeat failure.
It’s also why our crises never resolve the way Kuhn’s revolutions do. Kuhn’s revolutions are largely irreversible; nobody goes back to Aristotelian physics8 Newtonian mechanics is the awkward case for this, since it never really left. It still fits the overwhelming majority of human-sized calculations perfectly well. This itself is a nice illustration of how partial and untidy these “replacements” are. . Ours oscillate: plan, crisis, exploration, ship, new plan. What’s the new plan? Well, it’s the old plan plus a section about the specific defects we just found, retitled “lessons learned” and slightly longer. The internally reconstructed history goes in the text; the actual history ends up in the footnotes, or nowhere at all. Lakatos demanded an honest lampshading of the bias. Feyerabend correctly called it out as a demand that will never be met.
The case for method, as there’s always a method
If the last few sections read as an attack on planned testing, let’s set the record straight. Feyerabend is not a permission slip for vibes-based testing, and the arguments for tightly managed testing are considerably stronger than the people making them usually manage to articulate.
The original Kuhn takeaway still holds. Normal testing is where the bulk of the work actually gets done. Normal science isn’t second-rate science; it’s productive and effective. For all his self-proclaimed anarchism, Feyerabend has nothing specific against normal science as such. Testing-wise, the despised regression suite is the machinery that converts open questions into puzzles, and puzzles are the only thing an organisation can schedule, staff and finish on a quarterly plan. Just try building a delivery plan out of “our best tester will follow her nose”.
You can’t build an onboarding plan out of it either. Method is how testing survives the departure of the person who understood the system. Implicit knowledge and personal genius doesn’t transfer; it walks out with two weeks’ notice. Artefacts transfer badly, incompletely, with distortion, yes. All true, but that’s still better than nothing. Over time, unmethodical testing degenerates in an entirely predictable direction: comfortable paths, familiar features, the parts that are pleasant to use, and a debrief note reading “looks fine to me”. The on-the-job training the next generation absorbs will be the most mediocre, middle-of-the-pack version of all that, if anything at all.
Not to mention that there are entire fields where the artefacts and processes just are part of the product. Regulated and safety-critical work needs the evidence trail not as a byproduct of testing but as a deliverable in its own right. What’s being purchased is a defensible demonstration that specified behaviours were verified by a repeatable process. That’s not going anywhere any time soon, even if people shipping standalone webapps say otherwise.
A final argument for method is surprisingly in the very spirit of what mattered most to Feyerabend. Much of his attack on scientific method was rooted in anti-authoritarian concerns: his real target was science operating as a church, with methodologists as its clergy. He thought this had tangible political implications that undermined democratic resource allocation (a rabbit hole for another day). But in a software project, a rule governed method is often the only thing standing between the team and an overly confident manager or self-appointed expert. “Trust my judgment, I’m the senior” is an appeal to authority of exactly the kind Feyerabend warned about. At its best, a rule governed method gives the junior tester a voice and a veto they would otherwise never have.
Notice we’re using several different notions of rules throughout the discussion. There are rules about what gets recorded and who has standing: every release needs a documented sign-off, anyone may file a blocker, a “no” requires a written reason and findings get logged regardless of where they came from. Then there are rules regulating how and where to look: use this template, work off that checklist, cover the specs in this order, never go off-plan. The first kind of rules is what gives the junior tester a veto, what survives the departure of the person who understood the system, etc. The second kind is what manufactures correlated blind spots. Most arguments for method rely on the former type of rules; most arguments against method target the latter type of rules. Heavyweight processes tend to bundle the two and then defend the bundle, which is how you end up mandating the blind spots in the name of protecting the juniors.
Questions instead of practices
If there’s one universal to take from the for-and-against-method debate, it’s the Feyerabendian insistence on context. The answer is not 70% planned and 30% exploratory, nor the reverse. The whole notion of “an answer” is as false as the supposed dichotomy between planned and exploratory. We’ve accumulated enough mileage to see that for and against method were never really opposed. Feyerabend wasn’t against methods as such (note the plural). Lakatos wasn’t defending a single rulebook for the sake of there being a rulebook. Their letters come across far more like two people circling a shared problem than a debate between unyielding sides, which is presumably why a joint book seemed like a good idea in the first place.
What’s left standing, once the dichotomies fall, is the meta-observation that the choice of approach is itself a judgment call that no approach can determine, and it has to be remade constantly as the context moves.
You’re shocked, I’m sure, and not at all reminded that the testing world already has a school of thought saying exactly this. Context driven testing, as championed by Kaner, Bach, Pettichord and Marick, opens its 7 principles with the claim that the value of any practice depends on its context, and follows with “There are good practices in context, but there are no best practices”. No one had to read Feyerabend to reach that conclusion. So rather than dwell on the obvious, let’s extract some non-trivial action items as questions to be explored.
Where did the last 5 defects that mattered actually come from, and does your method contain that route? If not, your process isn’t a description of value-producing testing. It’s a description of a parallel activity that you also happen to fund.
Does your risk model make a prediction that could turn out to be wrong? If it only produces percentages and completion states, it’s describing paperwork, not the system. Derive your Neptune, point your test suite where the model says your defects should be, and check. Then change the model according to what you find.
Are your escaped-defect rate and your test suite pass rate on the same chart? If not, put them there, even if only in your own internal report. If they diverge, no amount of additional test cases in the existing shape will save you. Your anomalies are accumulating, and crisis testing won’t be far behind.
What’s your standard explanation when something escapes? If it’s the same one every time, and nothing could contradict it, you’re in a degenerating test program and the belt has stopped touching anything. That may be an inescapable reality, but it should at least be made explicit. Lakatos was fine with degenerating programs as long as they were honest about it.
Is anyone on the team permitted to test differently from everyone else? If not, you’ve optimised for consistent blind spots. Note this is a question about how and where people look, not about what gets recorded. Standardise the paper trail as much as you like, those are the rules that protect people. Standardising the attention is the type of rule to avoid (well, mitigate for).
Name what you’d lose if you drop the plan. If you can’t name anything, it’s theatre and you should say so. If you can name a great deal (traceability, audit, handover, a regulator) don’t drop it, and stop apologising for having it. There’s always a method, and it’s usually fine for it to be a boring cliche. Cliches became cliches because they usually fit the bill.
Epilogue, with the obligatory part about GenAI
We started this dive into post-Kuhnian thought by noting that Kuhn’s work was revolutionary and eye-opening, but that his fatalism left us empty-handed in practical terms. At the end of the journey, the most valuable tool for circumventing that fatalism turns out to be human, context-sensitive judgment. The capacity to notice that our model has gone stale; to smell that something is wrong before you can say why; to break the rule that needed breaking, even when nobody wrote that rule down explicitly. Awesome possum.
Meanwhile, the industry is busy delegating testing to AI systems that are exceptionally good at generating thick protective belts. A thousand plausible test cases before lunch, a traceability matrix on request, a slick coverage report (look mom, I vibe-coded this myself, etc). The standard complaint at this point is that AI systems inherently have no judgment: they will never wander off the plan because something smelled off, because to them nothing smells like anything. They cannot be surprised, which means they cannot be usefully wrong, which means counterinduction isn’t available to them at all. All true, for sure, but by now we’ve accumulated the tools for a slightly more detailed analysis.
The Feyerabendian worry is worse than a simple lack of judgment. Proliferation, intended or not, used to be our insurance policy. Two testers who model the system differently fail differently, which is why a team of humans had a chance to triangulate beyond any one tester’s blind spots. But an AI generated test suite doesn’t have a human mental model; it has a distribution. And it is largely the same distribution as the test suite your competitor generated, from the same handful of models, with the same training data, the same idea of what an edge case looks like, and the same silence about the test cases nobody ever wrote a blog post about. Blind spots used to be local. Every team had its own peculiar ones, which meant that somewhere out there, somebody’s weird process had a chance of covering your gap, if only by luck and idiosyncrasy. Now everything correlates. The review process correlates too, incidentally, since that’s increasingly done by the models similar enough to the one that generated the thing. We all get to discover the same class of defects in the same quarter, and there is no outside left to look in from.
How to counter this fatalism? Some would see this as a golden opportunity, and call for a smart division of labour: hand over the puzzle-solving of normal testing to AI, keep the open questions, the value judgments and the belt adjustments for ourselves. Being somewhat of a cynical fatalist myself, I expect we’ll all converge on an elaborate, token-wasting fast route to the same green dashboard nobody believes. Actually, calling that an expectation is generous - it’s already been happening for a while.
Anyway.
Your test suite is green. Go and use the product for an hour. You might even be surprised.
OMG, you actually read the entire thing? You beast. Good on you.
Forget using the product for an hour, you deserve the rest of the day off.
Comments
(must be logged on to comment)