I took the Gallup CliftonStrengths assessment recently. Thirty four themes, ranked one to thirty four, the whole thing.
Now the sensible person opens the report, reads the top five, feels quietly wonderful about themselves, and gets on with their day. I did not do the sensible thing. I opened a chat instead and said: before I tell you anything, you rank all thirty four for me. In order. With reasoning for each.
Partly this was curiosity. Mostly it was that I wanted to catch the thing out.
It got the number one exactly right. It also got the number thirty four exactly right, which is the one that actually matters, and we will come to why.
I am not publishing either of them. You will have to live with this, and by the end I think you will agree the names are the least interesting part of the story.
The top was a coin flip dressed up as insight
Let us take the win away from the machine first, because half of it was not earned.
The top of anybody’s list is over-determined. If you have spent fifteen years in enterprise architecture, collecting certifications the way some people collect stamps, writing a blog nobody asked for, and turning up to conferences voluntarily, then a small cluster of themes is obviously going to be sitting in your top five. Anyone who knows you could name that cluster. The machine picked the right one out of the cluster, sure, but it was choosing between four or five plausible candidates that all point the same direction.
That is a coin flip with good odds. Impressive at a party. Not evidence of anything.
The bottom, though. The bottom was arithmetic, and the arithmetic is the interesting bit.
Ipsative is a fancy word for “you cannot have everything”
Here is the thing about CliftonStrengths that most people never bother to learn, and it changes how you read your own report.
It is not scored against a population. You are not being told you are more curious than 84 percent of New Zealanders. The instrument is ipsative, which means you are being scored against yourself. It works by making you choose. Two statements, pick the one that sounds more like you, again and again and again, and the ranking falls out of who beat whom.
So every theme spends the whole test in a series of fights. And this is where the structure gives the game away, because the themes are not evenly matched fighters.
Some themes are defined by doing something. Taking charge. Starting the thing. Standing on a principle. When those show up in a pairing they sound decisive, because they are describing an action, and actions are vivid.
Other themes are defined by not doing something. By holding back. By seeking the least friction available. And when a theme like that meets a theme like the other kind in a forced choice, it loses. Not because the person genuinely lacks it, but because “I would rather find common ground than argue” reads as smaller than “I take charge” when the two are sitting side by side on a screen and you have four seconds to choose.
Do the multiplication across the whole instrument and the withholding themes drift towards the floor for most action-oriented professionals. It is a property of the format as much as the person. Which means the bottom of the list is, in a small way, predictable before you have met the human at all.
It was not reading my soul, it was reading my sentences
The other half of the prediction was less mystical still, and I have only myself to blame.
I had told it, in writing, weeks earlier, that I do not want a consensus builder. I want disagreement, pushback, a position defended rather than softened the moment I lean on it. I had written that down as a standing instruction because I meant it.
And the theme that landed at thirty four is more or less the definition of the thing I had explicitly asked it not to be.
So the model was not peering into my depths. It was reading a near verbatim negation of a theme definition that I had handed over myself, and then doing the pairing arithmetic on top. That is what inference is. Pattern completion over the evidence available. The feeling of magic comes entirely from us forgetting how much evidence we gave away, and we give away an enormous amount, casually, in the way we phrase a request at nine in the morning.
I find this deflating and also quite reassuring, in roughly equal measure.
But is this not just the horoscope problem
Fair question, and I would ask it too.
Personality instruments are famous for the Barnum effect. Tell somebody they are independent minded but also value close relationships, and they will nod along, because it is true of everybody and nobody at once. Vague, flattering, unfalsifiable. Very good business model.
Two things save this particular exercise from that. First, the prediction was made in advance and written down, which makes it falsifiable, and a fair chunk of the middle of the ranking was in fact wrong. It was not a clean sweep. Second, and more importantly, the hit was at the bottom. Barnum statements are always flattering, because that is the whole trick. Nobody ever built a booming trade telling people the thing they are worst at. You will not find “at heart, you avoid difficult conversations” printed on a mug in an airport.
The floor is where a personality instrument stops being entertainment.
Nobody came first
My seven year old asked me the other day what I was doing, and I said a test told me I am first at one thing and last at another thing, out of thirty four.
He thought about this for a second and asked, quite reasonably, who came first.
And I opened my mouth to explain and realised he had gone straight past me to the actual point. It is not a race. There is no field. Nobody came first, because everybody comes first at something and last at something else, and the ranking only ever compares you to yourself.
Seven years old, no report, no certification, no forty page PDF. He worked out the definition of an ipsative instrument in about four seconds and then asked if he could have a biscuit.
The floor is the useful end
Here is what I actually took from it, and it is nothing to do with celebrating my strengths.
Your top five tells you what you will do anyway, under pressure, without deciding to. It is descriptive. Nice to have on a slide.
Your bottom five tells you where willpower will not save you, and that is operationally different, because it means you need a rule instead of an intention. If the thing at the bottom of your list is the instinct to seek common ground, then in a forum you are chairing, “I will try to be more inclusive of other views” is a wish, not a plan. It fails the first time someone says something you think is wrong.
What works is mechanical. Ask the question, then count to five before you speak. Make somebody else summarise the room before you do. Small, dumb, external rules that do not depend on you feeling generous in the moment.
I do not need a machine to tell me that. But I did need something to point at the floor rather than the ceiling, and the report, plus a model willing to argue with me about it, did exactly that.
Which brings me to the last bit
The genuinely unsettling thing is not that a language model guessed my results. It is how little it needed. A few weeks of ordinary conversation, one grumpy instruction about how I like to be argued with, and the structural properties of a survey. From that, the top and the bottom of a thirty four item ranking.
We are all far more legible than we think. Every request we type is a small deposit.
But legible is not the same as known, and I keep coming back to that. The machine worked out what I do when the pairings get hard. It has no idea why, and neither did the report, and honestly the why is the only part I would have paid for. That part is still mine, still being worked out, and apparently still capable of being explained better by a seven year old holding a biscuit than by forty pages of Gallup.
Good. Long may that continue.
Written for KiwiGPT.co.nz · Generated, Published and Tinkered with AI by a Kiwi