Testing Saints
In this piece we’ll transpose one of my favourite moral philosophy papers onto the software testing profession, and present a rational, well structured, thought out argument licensing you to hate your co workers. It’s great, I know!
The paper in question is Susan Wolf’s “Moral Saints”. It doesn’t attack theories of morality, but a hypothetical person that would follow any one of them to the maximal possible extent. Wolf argues (I think successfully), that such a person would be severely lacking in most things we value, and would basically come off as a horrible person you’d never want to be, or to have in your life.
We’ll explore the details and structure of her argument, then build the relevant mapping to transpose it onto the software testing world. There, we’ll use it to examine what would that entail for a hypothetical testing professional that follows a methodology to its maximal possible extent. If you you’ve been in the industry for long enough, maybe that won’t even be a hypothetical person, as you’ve surely had the misfortune of working with someone similar.
Moral license to despise the kind of people you already hate in your gut? What could be better?
Let’s go.
Prologue
Susan Wolf is an American moral philosopher who spent a career on meaning in life, free will, and moral responsibility. In 1982 she published what came to be one of her most famous works, “Moral Saints”, which I personally absolutely love. It’s a rather short piece as these things go, not too formal, somewhat soft language, and absolutely brutal. She’s great.
The paper teases its provocative target right off the bat, saying that whether or not moral saints exist, Wolf is glad she isn’t one of them and nor are the people she loves. She’s not here to argue technicalities about what’s morally possible or not, but to claim that moral perfection is just not the kind of thing you should aim for. Seems paradoxical and in direct conflict with the very notion of morality, but by the end of the paper she actually manages to deliver on that.
The whole thing works through analysing what a hypothetical moral saint would look like (in the moral theory flavour of your choice). Meaning, what would a person that attempts to maximise their moral worth actually be; what would this person do, practically, and what exactly would that entail. We’ll borrow her structure and analysis for our own use in the second half, but for now, let’s see what she makes of it.
Have I mentioned she’s great? Because she really is. If you have the time, I’d actually recommend you go read her original paper (you can probably find the full text in a simple google search 🤷♂️). It’s a delight.
Wolf’s saints
Two internal ways one can be a saint
Wolf splits her sainthood analysis into two: the internal psychological motivation, and the external functional behaviour. Internally, she calls out two possibilities:
The Loving Saint is all in, through and through. Not only a true believer in the cause, but a wilful enthusiast. Their own happiness genuinely consists in the happiness of others, so they sacrifice nothing and suffer no strain. In a nutshell, Wolf sees this saint as having no independent self in them to do the loving; their desires have been so thoroughly aligned with the moral demand that there’s nothing left of the person that could have wanted otherwise.
The Rational Saint has their own internal interests, which they suppress for the cause. They would rather be reading, or playing tennis, or, god forbid, sleeping. They wanted to and didn’t, because there was something better to do with the hour, and morally speaking, that’s an evergreen statement; there is always something better to do with the hour. They white-knuckle it everyday, keeping their internal self honest, but spending their entire life overruling it. Wolf notes this looks less like virtue and more like a form of self-alienation with good PR.
Wolf’s description already paints these as bad in general, but she goes through the trouble of detailing a couple of mechanisms for why, specifically.
The first is crowding out. See, there are a lot of things a saint simply can’t afford to have in their life: gourmet cooking, fine wines, Victorian novels, the oboe, a really good tennis backhand. These aren’t vices, and usually we’d think of them as a-moral (in the sense of not having moral content; Wolf says nonmoral). Just the ordinary excellences of a rich human life. However, they’re expensive, or time-consuming, or self-indulgent, or simply trivial when weighed against a world containing preventable suffering.
The saint can’t cultivate all of these, because compared to the dogma of their moral theory, the justification to do so never arrives. Whatever your dominant value is, if it’s genuinely dominant, then everything else must earn its place against it, which would always be beyond what the a-moral virtues can meet. That may be acceptable from the internal point of view of the moral system in question; but from a general, common sense, human and humane point of view, life without any of these a-moral virtues isn’t admirable, it’s impoverished.
The second is character. The saint has to be very (very [very {very}]1 very ) nice, and this rules out A LOT of the spectrum of what we’d consider to be human excellence. Any kind of mocking humour is out, because it requires taking pleasure in someone’s absurdity. Cynical wit is out as well. So is a certain kind of aesthetic ruthlessness. The saint isn’t merely dull as a side effect, but rather the dullness is entailed by sainthood. It’s like trying to crack witty jokes with the collective from Pluribus (a hive mind of relentlessly contented people). They’re just too darn nice to make it work.
Critiquing morality from the outside
At this point one might stop and ask what we (i.e. Wolf) are even trying to achieve here? If a hypothetical saint maximises a moral theory implementation, any critique we mount against them will tautologically fail. They are defined as a saint, which literally begs the question and sets the inevitable conclusion.
Wolf makes her point of view clear, and explicitly states her critique is from the outside. Of course there can be no moral reason to be less than a moral saint. From inside morality, the saint is unimpeachable. That’s a given. The objection comes from a broader human point of view.
Wolf invokes what she calls the point of view of individual perfection: a standpoint from which we judge lives as flourishing or stunted, rich or thin, well-lived or wasted. If we grant such a point of view exists (or at least entertain the possibility because it feels intuitively valid), then indeed the saint’s life looks bad. Not wicked, for sure, just bad in the way a life spent entirely indoors is bad. Lacking. Starving, even.
The very fact we can entertain the point of view hints at a deeper truth. If morality is unimpeachable from inside, but produces a life we can judge as lacking from the outside, then morality is just ONE value among others, rather than the sovereign arbiter of all value. Armed with this understanding, Wolf proceeds to examine two major flavours of moral theory from this outside point of view:
Utilitarianism crashes and burns almost immediately, as it was always vulnerable to this kind of critique. Maximise aggregate welfare, and every hour you spend on your own projects is an hour stolen from some better use of your time. There’s no threshold above which you’ve done enough, and no room for what philosophers call supererogation: acts that are good and praiseworthy but not actually required of you. What we’d label as going beyond the call of duty.
Utilitarianism can’t have that category, because if an act would produce more good then it was already obligatory; and if it wouldn’t, it isn’t praiseworthy in the first place. The whole middle ground where most of human decency actually exists just disappears. The theory is total by construction, and the saint it produces is total too. Easy peasy lemon squeezy.
Kantianism looks like it might get away. Kant has room built in: imperfect duties come with latitude, you’re allowed (even required) to develop your own talents, and there’s no demand to maximise anything. So instead of a head-on approach, Wolf’s attack is based on what people have always found unintuitive in Kant; what we might call a wrongly furnished life. The motives are all backwards, in a manner that almost comes off as alien.
You cultivate your talents because duty requires the cultivation of talents. Yes Sir. Right away Sir. Wouldn’t want to fail my duties, Sir! The oboe is permitted, yes, but you’re playing it under a prescription, in an allotted manner. Self perfection, as required, twice a day; not because you love the sound of it or have fun doing so, god forbid. Remember, Kant insists that moral worth attaches only to action done from duty rather than from inclination. Imagine being friends with someone, because it’s your duty to be a good friend, and that this is the be all and end all of the entire thing (for a saint, at least).
So Kantianism is also a no-go. What now? Wolf has no inclination to produce a replacement theory, or to land on some “balance is important” deepity. That would just be another theory, with better PR. Wolf offers a structural diagnosis: moral theories have been asked to do a job they can’t do, i.e. serve as complete guides to life. A reasonable person gives morality serious weight without granting it sovereignty, and there’s no formula for that. This is not a gap to be filled, but a fact of human life.
That’s a quick, limited rundown of her paper, but I do recommend you make some time to read it yourself. You’re no saint, so you should have an hour to spare.
Next up, we’re going to copy Wolf’s argument structure wholesale.
Bridging moral theory to software testing practices
Take a testing methodology; imagine a hypothetical practitioner who satisfies it perfectly: every judgement it prescribes, made correctly; every artefact it requires, produced; every value it holds, held with full weight and no competitor. Then we ask Wolf’s question: not “does this work?”, but is this a practitioner you’d want to be, or want to work with?
Throughout this argument, we’ll deploy Wolf’s analytical tools: working from an external point of view; the loving/rational split; what does crowding out look like; and what character “defects” are entailed from this whole framework. Granted, morality pretends to be much more universal than testing methodologies (I hope you’re not using ISTQB standards when talking to your significant other), so “maximising” will mean something thinner for the testing saint2 If we really want to stretch things, we can make the case that there are people for whom a methodology isn’t an instrument, but an identity. “I’m context driven”, “We’re a proud TMMi Level 4 shop”, “I’m Agile” etc. People do change jobs over this, and some lose friends over it (well, “friends”). So Wolf’s sovereignty thingy may still apply, it’s just that here it’s sovereignty over the professional self. A stretch, but still. . Still, I think we’ll gain some nice insights from the process.
With that in mind, let’s begin.
Our saints
The Kantian saint of the traceability matrix
Our Kantian saint practises planned, documented, technique-driven testing to the highest degree. All your favourite (or not so favourite) syllabus techniques, the test basis, entry and exit criteria, the whole shebang. This approach maps nicely to a Kantian morality through a simple prism: what’s considered right is determined by the process it was derived by, rather than the outcome.
A test case is correct because it falls naturally out of a procedure. Equivalence partitioning gave you the classes; boundary value analysis gave you the test data; the pairwise table scoped the condition combinations to focus on; etc. The case is justified by the process that produced it. Whether it finds anything is a whole other question, of secondary importance.
After all, a case that found nothing has done its job, which is exactly the structure of a Kantian duty discharged. Another nice parallel is that just like Kant’s categorical imperative, the justification is a priori to the test execution run: you can defend the case before you execute it, from the test basis and the procedure alone.
OK, all well and good. We all like the hierarchical, ordered, procedurally produced, traceable test suite. Nothing wrong with that. Except that there is quite a lot wrong with that, once you consider it’s being carried out by a maximising testing saint.
The whole thing reeks of putting duty over other motives that seem equally good (or better). Let’s say our saint notices something odd on a screen, follow it for 2 minutes, and stumble across the best defect anyone will find this quarter. The methodology has nowhere to put it. It’s not traceable to a requirement; it doesn’t increment coverage. It goes in Jira with “found through ad hoc usage” in the origin field, which is an apology if there ever was one (notice, usage, not even testing).
The methodology has looked at the single most valuable event of the week and declined to recognise it as a form of testing, because it wasn’t derived from a procedure. Kant devalues the action done from love; our saint devalues the defect found from curiosity.
But wait, you say, the official dogma syllabi do make room for unscripted techniques. Room, yes. But like in Kant, the room is wrongly furnished. Exploratory testing appears as a taxonomy entry, with a definition, a place in the hierarchy, a recommended proportion of effort, a scheduled slot in the plan and a deliverable at the end of it. Prescribed like your test manager ordered, twice a day for the duration of SIT.
You are permitted to follow your nose, in the following manner, for the following duration, producing the following artefacts. Sorry, did I say permitted? I meant required. The love of the thing and the natural curiosity that usually drives it has been replaced by its authorised, slightly alien facsimile. Our saint’s exploratory testing session resembles a joke reconstructed by AI. The material’s all there, sure, but the delivery’s all wrong. Would you hire an incurious person to do your exploratory testing? Actually, would you hire an incurious person to do anything at all?
This works nicely with Wolf’s crowded-out excellences: reading the source code for curiosity’s sake when the technique is specified as black-box; learning the business domain deeply enough that it pays off a year from now; building a throwaway tool because it was fun; teaching a junior something you can’t write down; and caring whether the product is good rather than whether the testing was correct. None of these are forbidden. They just never generate a justification, so they don’t make it on the Gantt (or even on the extra 10 minutes you got back from a team meeting cut short).
How about our two inner versions of the saint? Maybe a loving testing saint can fare better than a rational one, or vice versa?
Well, the Loving version of this saint genuinely enjoys the matrix (no pun intended [well, a little bit]). They find the derivation satisfying, feel the click of a decision table closing, and have no independent curiosity about the product whatsoever. The product is where the requirements go. Ask them “wait, does this even make sense for our users?” and they’ll stare at you confused; it’s the spec, what other meaning of “sense” is there?
Think that paints a sad picture? The version the Rational saint paints is even worse. They are curious and do think about things, and actually they did see something in the logs on Tuesday. They had to white-knuckle it and let it go because it wasn’t in scope. They were right, though; that was from a hidden reconciliation error. At the end of the day, both they and the system were left worse for it on Monday.
The Utilitarian saint of the risk register
This somewhat writes itself, as risk-based test management is explicitly a maximising calculus. Probability times impact. Effort allocated in proportion to risk-weighted expected loss. Stop when the marginal cost of further testing exceeds the marginal expected cost of the defects it would have caught. This might as well be marketed as a Benthamite plugin for Jira.
So Wolf’s demandingness analysis just transposes verbatim. As the utilitarian saint can never justify an hour on themselves (because there’s always a better use for an hour), so our testing saint can never justify an hour on a low-risk area, because that is a straight up stolen hour from a higher-risk one.
That means slack, even as a source of serendipity, is never defensible. Not the afternoon someone spent on a component that intuitively looked funny, not the tester who wanted to learn the payments domain properly, not the exploratory session in the boring corner of the product.
Every one of those can only be justified retrospectively, by having paid off; but that future is blocked by the methodology’s justification structure beforehand. The saint has abolished supererogation. There’s no such thing as testing above and beyond, because if the extra hour was worth spending it was already mandatory, and if it wasn’t, it’s immediately waste. All that’s left is this dichotomy.
The risk utilitarianism saint also maximises the systematic day to day distortions that plague the industry as a whole. Think of Goodhart’s and Campbell’s laws (when a measure becomes a target, it becomes corrupt and no longer a good measure). Our saint maximises the corrosive effects of both. Take any testing KPI or measurement, DDP, DRE, defect density - once it becomes the objective, our saint has to maximise their effect, and that usually means finding cheap shallow defects fast (as finding deep meaningful defects usually take more time and resources). And of course the drift in quality is invisible from inside because all the numbers are improving.
Another way to look at this be made using Williams’ One thought too many analysis (that Wolf cites). Imagine a man who, seeing his wife and a stranger drowning, stops to calculate and indeed conclude that saving his wife is permissible before jumping in to help her. The end result was OK, but there was something alien and broken in the process; namely, he had one thought too many.
How does this translate to our test manager saint? Imagine a tester walks in and says something about the reconciliation job is badly wrong, they can’t yet say what. Then our saintly manager thinks, and his first move is to check whether allocating 3 hours to that area will be optimal given the current risk register. He may well arrive at yes, for sure. He has still done something alien to the profession and corrosive, and everybody in the room felt it.
So, that’s from an external behaviour analysis. Internally, how are things looking for our saint? The Loving version has fully internalised this and experiences no friction; the register is his conscience and it never troubles them. You can replace them with a good enough Excel formula. The Rational version knows the model is a fiction, knows the probabilities are made up in a workshop, knows “impact” was almost a coin toss, and defers to it anyway. Meaning not only they’re a person who overrules their own judgement for a living, but they know it.
Just to be clear, we all have bosses and we all work in imperfect systems under imperfect risk models, and we all override our own judgement for a living. We don’t however, do any of these in a maximising, total capacity that overrides all the nuance and situations of our professional lives.
This total intensity is what makes the saint a horrible person to be, or to be with.
The Virtuous saint of the exploratory session
Let’s extend our mapping even further than what Wolf provides. She wrote her paper in 1982, when the revival of virtue ethics was still gaining ground (basically modernising Aristotelian thought about virtue and ethics and, well, virtue ethics). Wolf isn’t nostalgic for Aristotle and doesn’t want modern morality rolled back to it; but either way, let’s use her structural elements to do the work ourselves.
One of the pillars of virtue ethics is the fact that morals are inherently too complex and too intertwined with the human way of life to be captured, codified and written down. Well, exploratory testing in its developed form is anti-codificationist on purpose and proudly so. Testing is a performance; maybe an art, and therefore not an artefact. There are no best practices, only practices in context. The heuristics are explicitly heuristical; all of them are just aids to a judgement they can never fully replace.
That is virtue ethics. Right action is what the practically wise person would do in these circumstances; and “practically wise” cannot be cashed out into rules without destroying the very thing being described. One of Aristotle’s central concepts is the mean as the prudent person would determine it. Swap “prudent person” for “skilled tester” and you have the RST position on why there’s no such thing as a test case.
Surprisingly (or not, if you’ve met your share of devout exploratory testers), the saint of this methodology is the inversion of Wolf’s. Not too nice and too dull, but rather too interesting (and sometimes too rude). Notice the dominant value here is the tester’s own engaged, skilled consciousness in contact with the product; so naturally this saint refuses anything that reduces its role.
Automation is checking, not testing, so it’s beneath the practice. The tedious 400-configuration compatibility sweep can’t be justified as an exercise of judgement, because it isn’t one. Writing things down in a form an auditor can read is a translation that can’t fully capture or recognise what actually happened. Every session is a fresh perceptual encounter with a unique context, which, taken to the saint extreme, means no two runs are commensurable. The answer to “did this regress?” becomes an unworkable story rather than an actionable comparison.
Let’s follow it through to its absurd extreme. Uncodifiable means non-delegable, as the practice can’t be handed to someone with less skill, so it doesn’t scale past the people who have it. It also means unauditable, which is disqualifying outright in some areas where the evidence trail is the deliverable. And it means unfalsifiable, because when the saint says the product feels wrong and can’t say why, there’s no route for anyone to disagree (well, not one the saint will accept, anyway).
The Loving version finds every corner of the product fascinating, and therefore has no stopping heuristic that isn’t the clock. The most insignificant cosmetic UI issue gets 3 hours, just like the critical DB lock, because they’re both very interesting indeed. The Rational version deliberately refuses to accumulate routine, declines to let anything become habit, insists on meeting each situation fresh, denies themselves the efficiency they’ve earned3 BTW, interestingly, this makes the saint betray the origins of the theory, as Aristotelian virtue is a hexis - a stable disposition built by habit. Routine isn’t the enemy of practical wisdom, it’s the material practical wisdom is made from. An exploratory saint who refuses to let anything harden into habit isn’t a distilled Aristotelian; they’re a Romantic who became a caricature of one. Which is a nice summary of Wolf’s whole point about saints. .
Epilogue
Let’s borrow Wolf’s final conclusion as well, because she’s great (did I mention how great she is?). Put our testing saints side by side and generalise what actually went wrong. The shared failure isn’t zealotry, over-documentation or under-documentation, or any specific detail you can tease out. It’s that in each case the methodology has been promoted from an account of how to do part of the job to an account of what the job is.
That was Wolf’s whole point (well, it was part of her point, but for our scoped discussion I’ll focus on that). The problem isn’t with sainthood per se; it’s with a conception of morality that has expanded to fill the whole space of practical reason.
By taking the practice into its extreme, our testing saints demonstrated that Testing with a capital ‘T’ is susceptible to the same. To them a testing school of thought doesn’t have an instrumental, functional use, but they act as if it defines the profession as a whole. That’s a claim to sovereignty, which is what Wolf’s argument actually attacks.
Remove the totality and the all consuming sovereignty, and of course all saints make good and valid points. The Kantian scripted saint is right that without derivation you cannot distinguish skill from luck, and that in a regulated domain the trail is not idle bureaucracy but an end in itself.
The risk saint is right that attention is finite and that refusing to prioritise is itself a choice, usually a worse one when not made explicitly. And the exploratory saint is right that no document has ever noticed anything, and that every interesting defect in history was found by somebody paying attention. As in irreducibly human attention, driven by undefinable common sense.
Just to drive home the point, this doesn’t mean that synthesis is the answer. Wolf didn’t offer a balanced moral theory as a way out, because that’s just another system claiming sovereignty. Similarly, prescribing 30% exploratory, 20% risk based and 60% process based (I’m bad at maths) will just spawn a different breed of saint. One who has a documented rationale for exactly how much of each school to apply, treats that as the be all and end all, and somehow produces an even worse outcome.
What’s left is the thing Wolf actually says. The plurality is irreducible. Methodology is one value among several, and it gets serious weight without getting sovereignty. The judgement that arbitrates between competing values and practices is not something any of them can supply, because each of them thinks itself to be the judgement. There’s no formula.
One practical thing to look out for from all of this: not which methodology someone holds, but what happens when their methodology and the product disagree. When the process says stop and the tester’s nose says keep going, or when the nose says trust me and the auditor says show me. A practitioner who overwhelmingly resolves these cases in favour of the methodology is on their way to sainthood.
Trust me, you don’t want to be that person, or even worse, hire them.
Though now you have the moral license to hate them after the fact, which is nice 🤷♂️.

Comments
(must be logged on to comment)