Written by Emil O. W. Kirkegaard.
A reader once asked me the following:
I read your article about the intelligence quotient being very arbitrary and not being able to divide or multiply with it, what about raw scores on the WAIS for example? Can someone claim for example that he has 60% better memory than someone based on raw points?
This question is actually very deep and important to psychology. But let’s backtrack a little first.
In the Middle Ages, humans had the concepts of warm and cold, but no way to measure temperature precisely. During the Renaissance, various simple thermometers were invented which could capture some amount of temperature variation—in a nonsensical unit, such as the level of mercury in a glass tube. Although the level of mercury doesn’t really have much to do with temperature, it does tell you whether it is relatively warmer in one place than another, and by how much.
Later, scientists developed the theory of thermodynamics. We now understand that what we feel as temperature is really the movement of invisible particles. The more they move, the higher the temperature. If they don’t move at all, that’s called absolute zero. A mercury thermometer cannot measure this, nor can any other household thermometer. But scientists eventually devised methods for measuring the full range of temperatures.
Importantly, 0 Kelvin really is zero temperature (unlike 0 degrees Fahrenheit or 0 degrees Celsius), so you can do any mathematical operation with temperature data in Kelvin. The data has a ratio scale (a true zero and true intervals). By contrast, concepts like multiplication don’t make sense when a scale has no true zero. Suppose it’s 10 °C outside and 20 °C inside. Is it twice as warm inside? Some might say yes. But what if it’s –10 °C outside and 10 inside, is it –1 times warmer inside? Clearly, the math doesn’t work this way.
The Flynn effect and mental chronometry
Psychological measures are not like the Kelvin scale at all. In 1987, James Flynn famously reviewed the evidence on what we now call the Flynn effect. He noted the following:
The Ravens test measures a correlate of intelligence that ranks people sensibly for both 1952 and 1982, but whose causal link is too weak to rank generations over time. This poses an important question: If a test cannot rank generations because of the cultural distance they travel over a few years, can it rank races or groups separated by a similar cultural distance? The problem is not that the Ravens measures a correlate rather than intelligence itself, rather it is their weak causal link.
The problem is more general than just the Raven’s test, or even just intelligence tests. Nearly all psychological tests are afflicted. Let’s clarify a few things.
The Raven’s Standard Progressive Matrices test (RPM) has 60 questions or “items” as they are called in psychometrics. So when a person takes the test, they get a raw score from 0 to 60. That raw score is ratio scale. However, it doesn’t map onto intelligence itself as a ratio scale. 0 items correct on the Raven’s test does not mean zero intelligence. Some people will fail even the first item, but they have some level of intelligence, so 0 on the Raven’s is not really zero intelligence, but rather some low, indefinite amount.
This is why it doesn’t make sense to talk about how much smarter one person is than another in percentages. If you get all 60 items right on the RPM test, and I only get 30, you aren’t twice as smart as me, since a score of 0 doesn’t correspond to zero intelligence. If we both took the Colored Raven’s test (which is designed for children), we’d probably both score close to the maximum, making the percentage difference disappear. The test is easy for adults so even moderately smart people hit the ceiling. In the same way, a poor thermometer will not be able to establish who has the highest fever if it cannot measure over 40 °C.
There are some other options available. Digit span is sometimes mentioned as a ratio scale of sorts. The subject is asked to read or listen to some numbers, after which he must repeat them back to the test administrator. The administrator keeps adding numbers, making the task more and more difficult.
Here again, one can obtain a score of 0, and this is somewhat less arbitrary than the same score on Raven’s. But still, a digit span score of 0 does not really correspond to zero intelligence either. It could reflect some problem with the subject, say, he was deaf, spoke another language, or was too young to understand the task. Of course, we can stipulate that the testing must take place under certain conditions in order to be valid—just as a mercury thermometer doesn’t work under all conditions. Yet even then, a score of 0 doesn’t correspond to zero intelligence in the same way that 0 Kelvin corresponds to zero temperature.
The problem here is that we have no mathematical theory of intelligence itself. And this is true not just for intelligence but for every psychological trait. No one knows how these traits can be defined in the way that temperature is defined in physics and chemistry.
Another attempt at constructing a ratio scale measure of intelligence is summarized in Arthur Jensen’s final book Clocking the Mind:
Mental chronometry is the measurement of cognitive speed. It is the actual time taken to process information of different types and degrees of complexity. The basic measurement is an individual’s response time (RT) to a visual or auditory stimulus that calls for a particular response, choice, or decision. The elementary cognitive tasks used in chronometric research are typically very simple, seldom eliciting RTs greater than one or two seconds in the normal population.
Chronometry offers us data that is truly ratio scale because we measure time. One problem, however, is that zero time corresponds to maximum intelligence, not zero intelligence! It’s an inverted scale of measurement. If a brain can process incoming information at infinite speed, it can respond with no delay, thus giving a reaction time of 0. A very slow brain would take roughly forever to process information—like an old computer trying to load a modern game—and it would never finish.
How far can we get with chronometry in terms of measuring intelligence? Quite far actually. The basic apparatus looks something like this:
This is the eight-button variant of the Jensen box. You hold your finger on the middle button (the start or “home”). At some point, a light is shown at one of the other buttons, and you press it as fast as you can. The measurement can be varied based on how many buttons you need to choose between. It turns out that this is quite important. The more buttons and lights there are to keep track of, the longer the reaction time.
It might seem very simple and perhaps not important to measure, but it turns out that the Jensen box works well as an intelligence test when you combine a few variants of the task.
In this study, eight variants of reaction time were measured, and three groups were compared: university students (Un), gifted children (G) and non-gifted children (NG). You can seen that the gifted children perform about as well as the university students. Perhaps in some sense, then, they are about equally intelligent—despite the age difference.
If one combines several different reaction time measures, they can predict school grades just about as well as a standard battery of intelligence tests. Here’s a complicated looking model from a paper by Dasen Lou and colleagues:

They studied 532 primary school children, who were given the full Wechsler Intelligence Scale for Children (WISC), as well as some reaction time-like tests from a battery called CAT. The researchers extracted a g factor from the reaction time tests (CAT G) and one from the traditional intelligence tests (g) to see how well they would correlate.

The answer is .761. Not bad at all. And if they had used more different reaction time tests, they probably could have obtained a higher value in the 0.8 to 0.9 range. Near unity between traditionally measured g and g from reaction time-like tests. Indifference of the indicator.
How do we get a mathematical theory?
Chronometry shows that one can get ratio scale data for intelligence, and that this data can be used to measure intelligence almost as well as with traditional testing. But still, a g factor from a bunch of reaction time tests is not a ratio scale measure of intelligence either. The factor analysis or structural equation modelling that is necessary to compute g transforms the data into z-scores, and therefore degrades it from ratio to interval scale.
If we actually had a real, working, mathematical theory of intelligence, we could plug the numbers into some kind of model that would provide an estimate of intelligence on a scale that would be the same for humans of any age, dogs, cats, fish etc. But no one has the slightest idea how to construct such a model.
Again, this criticism is not specific to intelligence research. In fact, I am not aware of personality psychologists discussing the issue at all. Whoever heard of a mathematical theory of extraversion? If extraversion is a stable trait that differs between people, and members of at least some other species, it must be quantifiable on a ratio scale—at least in theory.
The most we can say is that extraversion must relate to some property of the brain’s neural network, such as its tendency to seek out external stimuli. (Hans Eysenck tried to get at some of this with his personality constructs. He was too early for the neuroscience revolution and died without making any serious progress.)
Intelligence, on the other hand, will presumably be related to whatever properties of the neural network make it more capable of processing complicated information. But this is merely restating our usual verbal definitions of intelligence with “neural network” added in. Not exactly a major leap forward.
Our lack of a ratio scale measure of intelligence is why we get all the Flynn effect problems. As Flynn noted in the quotation above, Raven’s tests work well as a measure of the output of intelligence. They can gauge differences in intelligence among people born at the same time, but not so much differences across cohorts. It’s as if we have a thermometer, but every 10 years manufacturers have to subtract an arbitrary value from the reading to keep them consistent—because of some cosmic force that changes how the reading relates to temperature.
Various intelligence tests show remarkably different cohort changes, and exactly why they do is something of a mystery. Here’s data from testing of children on the WISC, taken from a paper by James Flynn and Lawrence Weiss:

The numbers in brackets can be ignored, as these are estimates of gains, not measured gains. Putting those aside, there are some stark differences. Similarities gained about 24 IQ points over 54 years, but Information (i.e., general knowledge) and Arithmetic gained only 2 IQ points! How can these tests measure the same thing?
We know they measure the same thing fairly well within a cohort of same-aged people, but we also know that they don’t measure the same thing across cohorts.
Like Similarities, Raven’s shows huge gains over time. Yet these two tests don’t seem to have much in common. Raven’s is a non-verbal test involving figures and abstract patterns, whereas Similarities is explicitly verbal and involves finding a relationship between two or more things. What’s the connection?
Arithmetic is about applying basic rules of mathematics, and somehow it didn’t increase even though there’s been a huge rise in the amount of time children spend learning mathematics. Indeed, any school-based explanation will fail because neither Arithmetic nor Vocabulary increased much despite these being the explicit targets of schooling. The mysteries of the Flynn effect illustrate why we need a proper mathematical theory of intelligence.
What’s even stranger is that, while every subtest of the WISC shows at least some increase over time, simple reaction time may have actually declined. Indeed, this has been reported in both Britain and Sweden.
Imagine you were trying to measure the heights of men in different cohorts of a population. But instead of using a ruler, I only gave you pictures of the men’s shadows. When you looked at the data across cohorts, the shadows appeared to get longer. Did the men really get taller? Or were the pictures simply taken at different times of day? This is akin to what is going on with intelligence, since we lack any direct measures.
Conclusion
Almost all scales in psychology lack a true zero. Such scales therefore can’t be used for calculations involving multiplication or percentages. Most work done on ratio scales relates to intelligence, and although there are some interesting ideas about simple reaction times, there is no mathematical theory of intelligence on the horizon.
In my view, any mathematical theory of intelligence will have to invoke properties of the brain’s neural network, which is inherently complicated. Perhaps devising a mathematical theory of intelligence is beyond human intelligence.
Our lack of a ratio scale measure of intelligence lies behind the various confusions and paradoxes regarding changes in scores across cohorts. One can try to resolve these using advanced statistical methods (e.g., testing for measurement invariance), but such methods can only go so far. Other areas of psychology don’t seem to even think about the issue. As usual, intelligence research is the frontrunner.
This essay is adapted from one originally published here.
Emil O. W. Kirkegaard is a social geneticist. You can follow his work on Twitter and Substack.
Become a free or paid subscriber:
Like and comment below.







Very interesting article...
"Perhaps devising a mathematical theory of intelligence is beyond human intelligence."
...and a very interesting conjecture.
The question is whether it's human intelligence now or a future enhanced intelligence.