The Mismeasure of Samuel Morton and Psychometrics and the Myth of Multiple Intelligences

Stephen Jay Gould, The Mismeasure of Man (London: Penguin, 1981)

Stephen Jay Gould’s ‘The Mismeasure of Man’ is an almost entirely scientifically worthless work which has nevertheless proved enormously influential, in scientific discourse, social science discourse, political discourse and among the wider educated public.

Gould’s central argument is that objective, disinterested scientific research on sensitive and politically charged subjects, such as intelligence differences, and differences in brain size, as between groups and individuals, is almost impossible, because such research invariably reflects and is corrupted by the social and political biases of scientists themselves, and of the wider culture in which they live and work and were raised.[1]

Ironically, the closest he comes to actually proving this thesis is through his own work, which manifestly reflects the egalitarian, Marxist, radical leftist political milieu in which Gould was, by his own account, raised, as well as the racially egalitarian ideology prevailing in the wider American society in which his work, and this work in particular, was assiduously championed and promoted in academia and the media.[2]

‘Debunking Scientific Fossils’

Celebrated educational psychologist Arthur Jensen, himself a main target of Gould’s invective, and one of the few such targets still alive and hence able to defend himself at the time Gould published his book, wrote a devastating review of Gould’s book in the journal Contemporary Education Review soon after the book was published, which I strongly recommend reading in conjunction with Gould’s work as a necessary corrective to the latter’s errors and misrepresentations.

Jensen titled his review, ‘The Debunking of Scientific Fossils and Straw Persons’ (Jensen 1982).

This reference to ‘scientific fossils’ had, one presumes, a double meaning. First, it was a reference to Gould’s background as a palaeontologist.

Second, and more importantly, it was a reference to the fact that most of the research that Gould purports to debunk was conducted decades, even centuries, before Gould’s own work was written.[3]

It therefore almost totally irrelevant to assessing the merits of contemporary research in psychometrics, intelligence research and behaviour genetics, research which, save for his attack on factor analysis and the g factor (see below), Gould almost entirely ignores.

As Jensen puts it:

“Instead of taking on the real issues of contemporary research in these fields, paleontologist Gould tilts at a museum collection of scientific fossils” (Jensen 1982: 124).

In focussing almost exclusively on nineteenth and early twentieth century research in craniology and psychometrics rather than the latest cutting-edge research, Gould’s implication seems to be that, since such early research was, at least in Gould’s own telling, invariably crude, methodologically flawed and beset with statistical errors and fallacious thinking, the same must necessarily be true of all modern research in the same fields.

Yet this is a fallacious assumption, as must be obvious to anyone with a remote familiarity with the history of science, or the nature of scientific progress and the scientific method.

In short, science is a cumulative self-correcting process, both building on and supplanting what has gone before but also often invalidating it, through a method involving the testing and falsification of hypotheses.

As Arthur Jensen himself puts it in his review:

“Should we ridicule the early astronomers for claiming that the earth is the center of the universe, or the early anatomists for claiming that the heart is the seat of emotion?… Gould’s exclusive critical focus on forebears (and the worst examples, at that) is much like trying to condemn the modern automobile by merely pointing out the faults of the Model T” (Jensen 1982: 125).

In short, if we today see further than previous generations, it is, not just through ‘standing on the shoulders of giants’, but also through trampling on the corpses of midgets.

Defaming the Dead

However, to say this is to impliedly concede that Gould’s critique of nineteenth and early twentieth century work in craniology, and early work in intelligence testing, is largely valid. This, however, is questionable.

Indeed, one suspects that Gould chose to debunk early research in psychometrics and anthropology, rather than the latest cutting-edge research, not only because, since science is a cumulative, self-correcting process, the former is almost inevitably weaker and hence more amenable to legitimate criticism, but also because few of the scientists whom he critiqued were around to defend themselves and point out Gould’s own errors and misrepresentations in his characterization of their work.

Therefore, most readers would simply take Gould’s assertions about nineteenth century craniologists and early-twentieth century pioneers in intelligence testing on trust.[4]

Yet, interestingly, in the years since Gould’s work was published, a few bold intrepid researchers with too much time on their hands have indeed taken the trouble to track down and reread some of this literature, just as Gould himself claimed to have done.

Unfortunately, however, their conclusions regarding the content and merits of this research have often diverged dramatically from those of Gould himself.

Indeed, they have found that work conducted over a hundred years before that of Gould himself, in an age before the Modern Synthesis, before the discovery of DNA, and before even Mendel’s work on inheritance and Darwin’s on evolution, when science itself was still very much in its infancy, and conducted by the very scientific pioneers that Gould himself chose to posthumously defame, was often methodologically and mathematically superior to that of Gould himself, and moreover conducted with greater scientific integrity and disinterest.

The Mismeasure of Samuel Morton

In writing this, I allude in particular to the Gould’s critique and purported reanalysis of the data of Samuel Morton, a pioneering and once celebrated early craniologist and anthropometrist, who is regarded as both the founder and leading exponent of what was once called ‘the American School of Anthropology’, yet who is today largely forgotten, and chiefly remembered, to the extent that he is remembered at all, primarily on account of Gould’s critique.

Morton, famous for his collection of human skulls, the largest such collection in the world at that time he lived, which he had accumulated from all around the world, sought to measure and compare the cranial capacity of the various skulls by a technique that involved filling the various skulls with material such as seed or, later, lead shot, then moving the seed or shot to a different container in order to measure its volume.

Gould, however, purports to show that, despite the seeming objectivity of his methodology, Morton nevertheless systematically underreported the average cranial capacity of non-white skulls, while systematically overrepresenting that of white Caucasian, and especially white Nordic, skulls, on account of his own supposed “unconscious finagling” (p55).

Thus, Gould gratuitously speculates:

“Plausible scenarios are easy to construct. Morton, measuring by seed, picks up a threateningly large black skull, fills it lightly and gives it a few desultory shakes. Next, he takes a distressingly small Caucasian skull, shakes hard, and pushes mightily at the foramen magnum with his thumb. It is easily done, without conscious motivation; expectation is a powerful guide to action” (p65).

“I detect no sign of fraud or conscious manipulation. Morton made no attempt to cover his tracks and I must presume that he was unaware he had left them. He explained all his procedure s and published all his raw data. All I can discern is an a priori conviction about racial ranking so powerful that it directed his tabulations along preestablished lines” (p69).

Unfortunately, however, two subsequent reanalyses of Gould’s own reanalysis found little if any evidence for any such finagling, conscious or unconscious, at least on the part of Morton (Michael 1988; Lewis et al 2011).

On the contrary, the authors of the most recent and rigorous of these reanalyses found that virtually all the errors were on the part of Gould, not Morton, and, whereas Morton’s errors seem to have been random in their direction, and hence not obviously reflective of any biases on the part of Morton, Gould’s errors all had the effect of minimizing the reported racial differences.

The authors conclude:

“Ironically, Gould’s own analysis of Morton is likely the stronger example of a bias influencing results” (Lewis et al 2011).

In contrast, whatever biases or preconceived notions Morton may have had regarding either the intellectual ability or cranial capacity of different races, which is itself a matter of some dispute, Morton evidently did not allow them to influence either his measurements or how he reported his results.[5]

Thus, the authors conclude:

“Rather than illustrating the ubiquity of bias, [Morton’s research] instead shows the ability of science to escape the bounds and blinders of cultural contexts” (Lewis et al 2011).

Of course, it might be suggested that, just as I criticise Gould himself for critiquing the work of dead scientists no longer able to defend their own work, the same could be said also of Lewis et al (and perhaps of myself), since Gould too had died a few years before their paper was published.

However, an earlier analysis published in the journal Current Anthropology reported similar conclusions well over a decade before Gould’s death, yet the latter, to my knowledge, never deigned to acknowledge, let alone respond to this critique (Michael 1988).

Moreover, lest anyone dismiss this analysis on the same grounds that Gould dismissed the earlier work of Morton himself, namely by suggesting that the findings reflected only confirmation bias, or the racial prejudice of the author and researcher himself, it ought to be noted that the author of this earlier study was certainly no racist, nor even a race realist, emphatically rejecting the very biological validity of the race concept, let alone the notion of race differences in brain size.[6]

Yet he nevertheless concludes:

“Morton’s research was conducted with integrity” (Michael 1988: 353)

The contrast with Gould’s own work is, of course, striking.

Brain Size and Intelligence

The odd notion that brain size is unrelated to differences in intelligence enjoys surprisingly wide currency. I suspect this idea originates with Gould’s book, and the latter is indeed, my experience, sometimes cited as authority for this claim.

In fact, however, Gould, in his defence, makes no such claim.

Gould is, after all, a palaeontologist, and no remotely competent palaeontologist would deny that the massive increase in hominid brain size that occurred during the course of human evolution is related to our superior intelligence as compared to other species.

After all, brain tissue is hugely metabolically expensive and larger brains would hence never have evolved unless they conferred some compensating advantage, presumably intellectual.

As well as being related to intelligence differences as between different species, brain size is also correlated with differences in IQ as between individuals (Pietschnig et al 2015; Rushton and Ankney 2009).

Indeed, by so strenuously attempting to refute and discredit the research of figures like Morton, whose research demonstrated an association between brain volume and racial ancestry, Gould might even be taken as implicitly conceding that brain volume is indeed likely related to intelligence. Otherwise, he would have had no need to so strenuously and mendaciously attempt to dispute in their findings.[7]

Instead, rather than outright denying that there exists any association between brain-size and IQ, Gould instead acknowledges that brain size may indeed correlate with IQ, but argues that, rather than being causal, this is a consequence of the superior material conditions of children from economically privileged backgrounds, which enables them to both better achieve their full intellectual potential and also to grow bigger brains.

Thus, in a passage from the ‘Postscript’ to his chapter on nineteenth century craniology (‘Measuring Heads’) included in his original 1981 edition of ‘The Mismeasure of Man’ that is oddly omitted from later “Revised and Expanded” editions, Gould argues:

“The primary environmental hypothesis for correlations of head size with social class holds that they are artifacts of a causal correlation between brain size and status. Large bodies tend to carry large heads and proper nutrition and freedom from poverty fosters better growth in childhood” (p111).

Thus, he explains:

“Brain size would then be an imperfect measure of height and IQ might correlate with it… for the primary environmental reason that poverty and poor nutrition can lead both to reduced stature and poor IQ scores” (p108-9).[8]

This seems superficially plausible.

Severe malnutrition during childhood is indeed associated with reduced IQ and reduced stature, and presumably reduced brain size too (Soliman et al 2021; Kirolos et al 2022).

Indeed, malnutrition and disease were indeed surely a factor in explaining reduced stature among the poor in previous centuries, and surely remain a factor in explaining the reduced stature, brain size and measured IQs of many people in the developing world today.

However, this explanation cannot explain the association between brains-size and IQ in the west today, since the sort of extreme poverty and undernutrition that results in reduced stature and brain growth is now all but unknown in the west.

Indeed, in the west today, obesity is a far greater health problem in the west than is undernutrition, even among the poor.

Indeed, it is a problem especially among the poor. Thus, at least in the developed world, obesity is actually more common among those of lower socio-economic status than it is among the wealthy (Zhoi 2021; Kim & von dem Knesebeck 2018).

Moreover, while height does indeed correlate with measured IQ, the correlation of brain size with IQ is stronger than that observed for height.[9]

So, reduced brain-size cannot be solely a result of malnutrition or poverty unless these factors, for some reason, reduce brain size to a greater extent than they do overall stature.

Moreover, even within a single family, the sibling with the larger brain tends to have the higher IQ, which suggests that differences between families, including in socioeconomic status and nutrition, cannot be the explanation for this association (Jensen & Johnson 1994; Lee et al 2019).

IQ Testing

Whereas the first part of Gould’s book deals with the nineteenth century science of craniology, the second half ostensibly focuses on intelligence testing.

Again, however, he focuses largely on early research, focussing on the work of early pioneers in intelligence testing, such as Alfred Binet, Lewis Terman, Cyril Burt, Goddard, and Yerkes, and the methods used to test the ability of US military recruits during World War I. This work was mostly conducted in the early twentieth century, when IQ testing was crude and still very much in its infancy.

Thus, Gould’s critiques of the supposed methodological and statistical errors of these early pioneers may or may not be valid, but they are, like his critique of eighteenth and nineteenth century craniology and anthropology, almost entirely irrelevant to an assessment of the modern science of intelligence testing.

To the extent that Gould does critique of modern intelligence research, he focuses almost entirely on two related issues – first, the statistical method of factor analysis as employed by psychometricians, and, second, the concept of a general factor of intelligence (‘g’) as revealed by this method of statistical analysis.

The g Factor

The so-called ‘g-factor’ is a concept invoked to capture the positive correlation between performance on virtually all different types of test item on intelligence tests, even those measuring very different forms of intellectual ability.

In short, people who score high in one form of intellectual ability (e.g. verbal ability) also tend to score high in other forms of intelligence (e.g. spatial ability), while those scoring lower on one type of test item likewise tend to score lower on other types as well. The ‘g-factor’ refers to this general factor of intelligence, which explains most of the variation in IQ test performance as between both individuals and groups.

In itself, the g factor is, as Gould himself repeatedly emphasises, only a statistical construct. Yet Gould contends that, by treating ‘g’, or general intelligence, as a real thing, psychometricians commit what he refers to as the ‘fallacy of reification’.

This all sounds very impressive, but what exactly does Gould mean by the ‘fallacy of reification’?

After all, if performance on different types of test item, even those with very different content and testing very different types of ability, do indeed consistently correlate, then there must surely be some sort of explanation for this.

Moreover, while ‘g’ itself is indeed a mere statistical construct, it nevertheless itself correlates with measures of real physical and structural differences in the brains of different individuals, including with differences in cortical thickness, neural efficiency and glucose metabolism, functional connectivity, nerve conduction speed, white matter integrity, regional grey matter volume – and, of course, as already discussed, brain size.

For all his obfuscations, Gould’s critique of the g factor ultimately seems to come down to two points.

First, he reiterates the famous admonishment that ‘correlation does not imply causation’.

Therefore, Gould argues, even if performance in different spheres of intelligence do strongly correlate with one another, this does not necessarily mean that we can posit a general factor of intelligence that causes this correlation.

However, while it is indeed an axiom of scientific research that ‘correlation does not imply causation’, nevertheless, if the performance of different individuals on different types of intelligence test item do indeed consistently correlate with one another, then, again, there must be some sort of explanation for this.

Thus, while it may be true that ‘correlation does not imply causation’ in the sense of either A causing B, or B causing A, or indeed even A and B both being caused by C, nevertheless, if two variables (A and B) do indeed consistently correlate and the finding is robust and well-replicated, then this does indeed imply some sort of causal nexus, howsoever distant and indirect.[10]

Indeed, Gould himself comes close to admitting this in his second point, namely that, even if the g-factor is a real thing, it is not necessarily a heritable thing, but could instead reflect environmental factors.

Thus, he rightly recognises, “we cannot infer the cause from the correlation” and, seemingly admitting the existence of g, at least as a statistical construct, contends that, rather than “reflect[ing] an inherited level of mental acuity”, it could instead “record… environmental advantages and deficits (some people do well on most tests because they are well schooled, grew up with enough to eat, books in the home, and loving parents)” (p282).

In other words, contradicting his first point, Gould now seems to be conceding that correlation does indeed imply causation, at least in this case, but that the cause is to be found, not in genes, but in environmental factors.

“If the simple existence of g can be theoretically interpreted in either a purely hereditarian or purely environmentalist way, then its mere presence—even its reasonable strength—cannot justly lead to any reification at all” (p282).

Gould is right that, even if the different components of intelligence do intercorrelate (and they do), this does not necessarily mean that the reason for this correlation is to be found in heredity.

However, a great deal of evidence (e.g. from twin studies and adoption studies) does indeed suggest that heredity plays a large role in explaining differences between individuals in g.

Yet, besides dismissing the data on identical twins reported by Cyril Burt as fraudulent and hence worthless, Gould says almost nothing about his large body of research – research which has, incidentally, entirely corroborated the supposedly fraudulent and worthless findings of the much-maligned Burt.[11]

The g Factor and Heritability

In arguing that, if the g factor does exist, then it could just as easily be caused by environmental factors as by heredity, Gould gets one thing right – namely, he recognises that the question of whether individual variation in IQ scores is usefully captured by a single g factor is an entirely separate one from that of whether differences in intelligence, howsoever they are best represented in terms of their underlying factor analytic structure, are substantially heritable.

Yet, elsewhere, he seems to confuse these two questions, namely heritability and factor analytic structure, seeing the question of the existence of ‘g’ as somehow, and for some reason, inextricably connected to that of heritability, with theories of the heritability of intelligence necessarily depending on the unitary structure of intelligence.

Thus, he writes:

“I will at least say this for Arthur Jensen. He recognizes that his hereditarian theory of IQ depends upon the validity of g” (p295).

It may indeed be true that Jensen’s own specific “hereditarian theory of IQ depends upon the validity of g”, since Jensen did indeed devote much of his distinguished career to studying the nature of the g factor.

However, it is not true to say that hereditarian theories of IQ in general depend on the existence of a g factor.

On the contrary, if, as Gould has claimed, the g factor could be caused by environmental factors, then is it not also true that the multiple intelligences posited by fringe scholars like Howard Gardner could themselves be substantially heritable?[12]

In short, there is no reason to anticipate that multiple independent factors of intelligence, were they shown to exist, would be any less heritable than a general factor of intelligence.

Thus, Gould comes close to contradicting himself thrice over: first, he denies the existence of the g factor, contending that correlation does not prove causation; next, he admits that maybe the g factor does exist, but contends that the existence of the g factor says nothing about whether it is caused by genetic differences or environmental factors; then, contradicting himself yet again, he says that hereditarian theories of IQ necessarily depend on the existence of the g factor.

Group Differences and the g Factor

With respect to the vexed question of innate group differences, Gould again sees the g factor as central, writing:

“A reified Spearman’s g is still the only promising justification for hereditarian theories of mean differences in IQ among human groups” (p350).

Yet, here, he is again confused, for different “human groups” could conceivably differ in their innate intellectual ability, even if intellectual ability is itself not a unitary construct, but rather captured by several independent factors of intelligence (e.g. spatial and verbal ability).

For example, one group might score higher than another in verbal ability, while the other group scores higher in spatial ability, and these differences in ability could be innate in origin.[13]

Indeed, there is a particular irony here, because, in fact, the more different independent factors of intelligence that we recognise, then the less likely it becomes, on purely statistical grounds, that different races would nevertheless, purely by chance, have evolved to perform exactly equally in every single one of these different domains.

Thus, it is perfectly possible, if somewhat unlikely, that two races, say blacks and whites, that evolved on separate continents, in sufficient reproductive isolation from one another (and, presumably, subject to sufficiently divergent selection pressures) as to have evolved the obvious, and not so obvious, physiological differences that distinguish these two races, might nevertheless, purely by chance, have evolved exactly equal levels of innate cognitive ability – if cognitive ability is indeed usefully characterized as a single unitary factor, namely ‘g’.

However, the more different domains of intelligence that we recognise, each supposedly statistically independent of each of the others, then the correspondingly less likely it becomes that any two races would be equal on every single one of these independent factors.

If, like Howard Gardner, we posit as many as seven, or, in later expositions of his theory, as many as eight or even nine different ‘intelligences’, then it becomes statistically extremely unlikely that two different races, would, purely by chance, have evolved to be exactly equal on every single one of these different allegedly independent factors of ability.

When, on the other hand, one broadens one’s gaze to all the different races and populations in the world, the notion that all these different races and populations might have evolved exactly equal levels innate ability in all seven, or all eight, or all nine, of these supposedly independent spheres of ability is all but unthinkable on statistical grounds alone.[14]

Thus, far from ‘g’ representing, as contended by Gould, “the only promising justification for hereditarian theories of mean differences in IQ among human groups”, in fact a unitary g factor represents the best case scenario under which it is possible, but still unlikely, that different races might conceivably have evolved almost exactly equal ability.

Even then, however, it would remain extremely unlikely that this would apply to every single race or population on the planet.

It is therefore a supreme, yet rarely acknowledged, irony that Gould, a famously fervent opponent of the notion that different races differ in intellectual ability, whose abhorrence of this notion is such that it seems to have formed the main impetus for his authoring the very book I am currently reviewing, should nevertheless, in this very same book, have also so vehemently opposed the existence of the g factor that is so firmly established by modern psychometric research, since the positing of such a factor is surely the most likely scenario under which complete equality of the races in intellectual ability, an unlikely scenario under the best of conditions, could most plausibly be supposed to have evolved.


[1] Thus, Gould writes:

“I criticize the myth that science itself is an objective enterprise, done properly only when scientists can shuck the constraints of their culture and view the world as it really is” (p21).

Here, he skirts dangerously close to the anti-scientific epistemologically nihilistic cultural relativism of postmodernism, though, as a working scientist, necessarily ultimately rejects this view.
Gould is, of course, correct that scientific research does not occur in a social, cultural or political vacuum.
Scientists are not only well aware of the social, political and cultural values and dogmas prevailing in the wider society in which they live and work, they also typically share these values themselves, having, like non-scientists, internalized these values themselves during the process of socialization. It is therefore unsurprising that many biologists who should know better nevertheless parrot the biologically untenable leftist dogmas of sex denial, race denial and biology denial.
Moreover, even those scientists able to see that these dogmas are biologically untenable, not to mention empirically unsupported, are well aware of the fate that has befallen other scientists courageous or foolish enough to challenge the conventional wisdom with regard to these subjects (e.g. Helmuth Nyborg, Bo Winegard, Noah Carl, Bryan Pesta, even Lawrence Sommers and James Watson), and therefore know only too well to either parrot conventional dogmas themselves or else keep their mouths firmly shut on these matters if they want to keep their jobs.

[2] In referring to “egalitarian, Marxist, radical leftist political milieu in which Gould was… raised”, I refer to Gould’s own admission, “I learnt my Marxism at my daddy’s knee”. I make feel the need to clarify this since Gould often accused critics who suggested his Marxist upbringing and leftist views may have influenced his scientific views of “red-baiting” (e.g. here). Yet, as we have seen, the idea that a person’s political and social views influence their scientific claims is among his central claims in ‘The Mismeasure of Man’. Are leftists like Gould himself, then, in Gould’s view, somehow exempt from this sort of bias? Or is it just distasteful to accuse them of it? Gould’s resort to the charge of is “red-baiting” particularly ironic, not to mention hypocritical, given his own rather distasteful attempt, in ‘The Mismeasure of Man’, to smear hereditarianism by reference to its supposed association with Nazism, segregation, slavery and eugenics.

[3] Although Gould’s ‘The Mismeasure of Man’ was first published in 1981, nevertheless, according to Jensen’s count:

“Of all the book’s references, a full 27 percent precede 1900. Another 44 percent fall between 1900 and 1950 (60 percent of those are before 1925); and only 29 percent are more recent than 1950” (Jensen 1982: p124)

[4] Even modern scientists working in the fields of psychometrics, biological anthropology and neuroscience would be unlikely to be familiar with such research, since modern scientists are, quite naturally, concerned with the latest state-of-the-art research in their field, not the work of early pioneers that has long previously been built upon and supplanted. As Jensen, himself a leading researcher in the field of intelligence research at the time Gould’s book was authored, put it when reviewing Gould’s book:

“Frankly, I feel little inclination to comb the many archaic references on which most of Gould’s debunking depends, especially because they are no longer of any concern to modern researchers in these fields” (Jensen 1982: p125)

Only a researcher working in the field of the history of science who happened to have studied this particular area of research during the relevant time period might be expected to wade in, but, even then, s/he could hardly be expected to be familiar with much of the methodological minutiae upon which Gould focusses.

[5] Whether Morton did indeed have any preconceived biases regarding the data he collected is itself disputed. Thus, Lewis et al (2011) claim, “It is doubtful that Morton equated cranial capacity and intelligence”, which would suggest that Morton would not have had any motivation, conscious or otherwise, to fudge his findings in the first place. On this view, Moron was primarily interested in studying and measuring human morphological variation, rather than proving the superiority or inferiority of any particular race. Thus, Michael infers:

“He was trying to understand human racial variation, and not, as Gould claims, trying to prove Caucasian racial or intellectual superiority” (Michael 1988: p353).

[6] Thus, Michael himself concludes in 1988 paper by writing:

“Although Gould is mistaken in many of his assumptions and his work, he is correct in asserting that [Morton’s data] is unsound. He fails, however, to mention the overriding reason for rejecting them, namely Morton’s acceptance of the concept of race” (Michael 1988: p353).

Indeed, some fifteen years later, Michael even took to the internet to, among other things, disavow the use that some racialists such as Rushton had been making of his research.

[7] Interestingly, even Gould’s flawed purported reanalysis of Gould’s data fails to eliminate entirely the population differences reported by Morton and others. Thus, even Gould’s “Corrected values for Morton’s Final Tabulation” report the largest cranial capacity for “Mongoloids” and “Modern Caucasians” and the smallest among “Africans” (p66: Table 2.5). (No figures are given in respect of other groups such as Eskimos and Australian Aboriginals. Morton apparently lacked any Eskimo skulls and had very few Australian Aboriginal skulls.) Thus, Gould is reduced merely to asserting that:

“My correction of Morton’s conventional ranking reveals no significant differences among races for Morton’s own data [emphasis added]” (p67).

[8] In writing this review, and searching for the relevant passages to quote in both a hard copy of the 1981 edition and a pdf version of “Revised and Expanded” 1996 edition, I was surprised to discover that, as mentioned above, the whole section of text from which both this quotation and the quotation immediately preceding it were taken (namely, what was, in the original edition, the final section of Gould’s third chapter on “Measuring Heads: Paul Broca and the Heyday of Craniology”, and which has the subheading “Postscript” and runs from p108-112) is omitted from the “Revised and Extended” edition. As a result, there is no systematic treatment of Gould’s position on the relationship between brain size and intelligence among humans in this more recent edition. In his own review of the “Revised and Expanded edition”, Philippe Rushton infers:

“Gould realized that repeating this section verbatim, given the weight of the new evidence, would destroy his entire thesis. Rather than revise his arguments in light of the truth, Gould chose to repeat them without change and to withhold any evidence to the contrary. Both Gould and his publisher owe it to their readers to explain why this supposedly ‘new’ edition studiously avoids any mention of all the new evidence” (Rushton 1997).

[9] The correlation between height and IQ is around r = 0.2, a modest but robust positive correlation. The correlation between external head measurements (e.g. head circumference) is similar, but when brain size is measured directly (e.g. by MRI scans) the correlation is higher, at around r = 0.3, which would be described as a moderate or modest-to-moderate correlation.

[10] Indeed, despite its popularity as a slogan, the very phrase ‘correlation does not imply causation’ itself strikes me as problematic, or at least potentially misleading. It would, in my opinion, be more accurate to say:

‘Correlation does not always necessarily imply a direct causal relationship in and of itself, and moreover of itself implies nothing regarding either the direction of causation nor the nature and directness of the causal nexus.’

Perhaps, however, this is somewhat less catchy and easy to remember and quote and hence less likely to catch on.

[11] Burt is, of course, another researcher whom Gould and others have expended much ink in posthumously defaming when he was safely deceased and hence no longer around to defend himself or his research. Actually, however, although his research has been written up in textbooks as a classic case of scientific fraud, the case against Burt has always struck me as surprisingly weak.
Thus, the first, and, still, to this day, the main evidence cited in support of the contention that Burt was responsible for fabricating his data was the observation of leftist psychologist and political activist Leon Kamin of an odd apparent anomaly in the data collected by Burt regarding the correlation in IQs among identical twins, Thus, according to Gould, Kamin noticed that, in his later publications:

“Burt had increased his sample of twins from fewer than twenty to more than fifty in a series of publications, [but] the average correlation between pairs for IQ remained unchanged to the third decimal place—a statistical situation so unlikely that it matches our vernacular definition of impossible” (p265).

Actually, Gould’s phraseology here in sloppy and statistically imprecise. It was not the “average correlation between pairs for IQ” that remained unchanged, but the correlation coefficient between the IQ scores of identical twins.
Yet the idea that this statistical anomaly was indeed necessarily anomalous, let alone decisive proof of fabrication, seems to me to rest on an elementary statistical error.
Thus, it is indeed true, as alleged by Gould and Kamin, that it is very unlikely that correlation coefficient would have remained unchanged to three decimal places despite an increase in sample size “from fewer than twenty to more than fifty”.
However, it is no less unlikely that the average correlation would have changed by, say, exactly 0.003, or precisely 0.014.
In short, after increasing your sample size by more than 150%, any specific average correlation that you might predict beforehand would equally unlikely. An average correlation exactly the same as that found prior to increasing your sample size is, in principle, no more improbable than a correlation, say, exactly 0.002 higher, or 0.002 lower than that found formerly.
(It must also, of course, be borne in mind that one would expect the correlation to be very similar, since other studies of identical twins have also reported a very similar correlation coefficient, and moreover Burt’s later samples were appaently only expansions of his earlier samples and therefore the data overlapped.)
Indeed, given the amount of scientific research that is conducted each year, statistically unlikely results are statistically very likely to occur quite often.
Indeed, one might even argue that the identical correlation coefficients reported by Burt actually argue against deliberate fabrication, since, if Burt were indeed fabricating his findings, he would surely never have reported such an apparently suspicious result.
Interestingly, Burt is among the few intelligence researchers whom Gould does not accuse of racism. After all, Burt lived and worked in Britain in the early twentieth century, a time when the population of Britain was almost exclusively white. Indeed, he spent a large part of his career working for London County Council in a London that was, in stark contrast to the London of today, overwhelmingly white. As a result, neither he nor his employer evinced much interest in racial disparities in achievement.
Yet, rather than accusing Burt of racial prejudice, Gould instead accuses him of class prejudice, and blames Burt for his role in the introduction of the 11+ examination system in the UK, whereby pupils were allocated to different types of school, emphasising different forms of education, based on their performance in standardized tests at age 11, whose “major effect… in terms of human lives and hopes, surely lay with its primary numerical result—80 percent branded as unfit for higher education by reason of low innate intellectual ability” (p295).
He is dismissive of Burt’s claim that “his primary goal in supporting 11+ was… to provide access to higher education for disadvantaged children whose innate talents might otherwise not be recognized”, since, according to Gould, “Burt himself did not believe that many people of high intelligence lay hidden in the lower classes” (p295).
Burt did indeed argue that intellectually gifted children were relatively less frequently to be found among poor families than among the families of the wealthy and eminent. This is indeed inevitable, as Burt himself observed, and Richard Herrnstein was later to famously reiterate, if one accepts both that intellectual ability is associated with upward social mobility, and that intellectual ability is passed down in families, whether biologically or indeed culturally.
However, from this, it does not follow, as Gould asserts, that Burt therefore “did not believe that many people of high intelligence lay hidden in the lower classes” (p295).
On the contrary, although there may be a lesser relative proportion of gifted children among the working classes than among the upper- and middle-classes, there could still be a great number in absolute terms, especially since, at this time, working class people constituted the great majority of the population of the UK. Indeed given the greater numbers of working classes people, gifted children from the working-classes could even have outnumbered those from among the upper- and upper-middle classes.
As for the 11+ examination system (introduced, incidentally, by perhaps the most radical and reforming left-wing government in British history), this system surely not uncoincidentally coincided with higher proportions of children from working-class backgrounds ascending to professional occupations than has been observed either before the system’s introduction or after its abolition.

[12] Along with his rejection of the g factor, Gould explicitly champions Gardner’s theory of multiple intelligences, approvingly citing Gardner’s critique of the g factor and, in his ‘Introduction to the Revised and Expanded Edition’ crediting the latter’s the “theory of multiple intelligences [as] the major challenge to Jensen in the last generation, to Herrnstein and Murray today, and to the entire tradition of rankable, unitary intelligence marking the mismeasure of man” (p22).

[13] Indeed, to some extent, this is indeed true. Thus, although most variation in intelligence, as between groups no less than as between individuals, is captured by the g factor, nevertheless some races do indeed score relatively higher in certain specific abilities, for example, East Asians in spatial-visual and mathematical ability; Ashkenazi Jews in verbal ability; blacks in rote memory; and Australian Aboriginals in spatial memory. However, most variation in intelligence, including that between groups, is captured by the g factor. In other words, groups that score higher in one form of intelligence generally also score higher in other forms of intelligence too. Groups, no less than individuals, differ in g.

[14] We can illustrate this point by looking at physiological differences between groups. Thus, it is perfectly possible that two races might have evolved to be exactly equal in one specific physiological measure. For example, the reported average heights for black and white American adult males is virtually identical. Despite their very different evolutionary histories, black and white males resident in America seem then to be almost exactly equal for this trait.
However, the same is obviously not true of other racial groups resident in the USA, for example, East Asians, South Asians or Hispanics, most of whom are considerably shorter on average than either white or black Americans. Moreover, even blacks and whites differ in respect of other bodily measures and dimensions, even those other bodily measurements that correlate with height.
Thus, to take one example I have discussed earlier in this piece, black American men, despite being of the same average height as white males, have smaller head and brain size, on average. They also have somewhat longer legs, on average, than white Americans. This is despite the fact that both these measures (brain size and leg length) are themselves positively correlated with height, whereas the different forms of intelligence posited by Gardner and Sternberg were originally conceived as being largely, if not wholly, independent of one another (though Gardner, faced with the overwhelming evidence in favour of a g factor, later revised this view).
In short, the more different measures employed, and the more different racial groups in respect of whom the measurements are conducted, the less likely it is that every single group will be exactly equal in respect of every single one of the measurements. Moreover, this is especially unlikely to be the case where the measurements in question are uncorrelated with one another, as was alleged to be the case in respect of the multiple intelligences posited by Gardner and Sternberg.

___________________________

References

Jensen (1982). The Debunking of Scientific Fossils and Straw Persons Contemporary Education Review. 1(2):121–35.

Jensen & Johnson (1994) Race and sex differences in head size and IQ, Intelligence 18(3): 309-333

Kim & von dem Knesebeck (2018) Income and obesity: what is the direction of the relationship? A systematic review and meta-analysis BMJ Open 8(1):e019862

Kirolos et al (2022) Neurodevelopmental, cognitive, behavioural and mental health impairments following childhood malnutrition: a systematic review BMJ Global Health 7(7): e009330

Lee et al (2019) The causal influence of brain size on human intelligence: Evidence from within-family phenotypic associations and GWAS modeling, Intelligence 75: 48-58.

Lewis et al (2011) The Mismeasure of Science: Stephen Jay Gould versus Samuel George Morton on Skulls and Bias. PLoS Biol 9(6): e1001071.

Michael (1988) A new look at Morton’s craniological research 29(2):349-354.

Pietschnig et al (2015) Meta-analysis of associations between human brain volume and intelligence differences: How strong are they and what do they mean? Neuroscience & Biobehavioral Reviews 57: 411-432.

Soliman et al (2021) Early and Long-term Consequences of Nutritional Stunting: From Childhood to Adulthood Acta bio-medica: Atenei Parmensis 92(1):2021168

Rushton (1996) Race, intelligence and the brain: The errors and omissions of the ‘revised’ edition of S. J. Gould’s ‘The Mismeasure of Man’ (1996), Personality & Individual Differences, 23(1): 169-180.

Rushton & Ankney (2009) Whole Brain Size and General Mental Ability: A Review. International Journal of Neuroscience, 119(5):692-732.

Zhou (2021) The shifting income-obesity relationship: Conditioning effects from economic development and globalization SSM – Population Health 15: 100849