Plus ça change
The method is simple: construct pairs of applications that are equivalent in every respect except the applicant's name, send them to real vacancies, and count the callbacks. In 2019 researchers at Nuffield College's Centre for Social Investigation reported the results of nearly 3,200 such applications to real British jobs across a range of occupations — software engineering, marketing, chef, shop assistant. Applicants with minority ethnic names had to send around sixty per cent more applications than white British applicants with identical qualifications to receive the same number of positive responses. The penalty appeared across occupations and did not disappear when the applications made the candidate's British education and citizenship explicit. The researchers' comparison with earlier British field experiments, going back to work done in the late 1960s and 1970s, found no clear evidence of decline over half a century.
Fifty years. Whatever else has changed in British working life, that has not.
from A Guide to Occupational Psychology: People, Work and Organisations in the Modern UK (2026)
GPT-5.6 Sol:
The core claim is substantially right, but the strongest version of it needs slightly tighter wording.
What makes correspondence experiments unusually powerful is that they get close to the counterfactual we actually care about: what would have happened to the same applicant had the employer perceived them as belonging to a different ethnic group? Qualifications, experience and other job-relevant information are held constant while ethnic origin is experimentally signalled, principally through the name. Random assignment means ordinary observational explanations—different educational attainment, work histories, job-search behaviour, geography, and so on—cannot explain the callback difference. That gives this sort of experiment much stronger causal leverage than simply observing that different ethnic groups have different employment rates.
There is one methodological correction to your description. The 2016–17 British GEMM experiment reported by Nuffield in 2019 did not simply send a matched white and minority application to each vacancy. Applicants of different ethnicities were randomly assigned across real vacancies, with one fictitious application sent to each post. The underlying logic is still experimental—equivalent applicants, randomized ethnic background—but “construct pairs … send them to real vacancies” describes some correspondence studies better than this particular one.
The headline number is accurately stated. White British applicants received a positive response about 24% of the time, compared with 15% for minority-ethnic applicants overall. Since 24/15 = 1.6, the latter group had to make about 60% more applications to obtain the same expected number of positive responses. One useful clarification is that this does not mean their callback probability was “60% lower”: 15% is 37.5% below 24%. “Sixty per cent more applications” is the correct formulation.
The British-background point matters considerably. The applications indicated that minority candidates were British-born or had arrived by age six and had received their education and training in Britain. So a straightforward explanation of the result as employers reacting rationally to uncertainty about English proficiency, unfamiliar foreign qualifications, immigration status or recent migration becomes very difficult to sustain. Subsequent cross-national research similarly finds that ethnoracial minority status predicts hiring penalties independently of birthplace, citizenship and education, although foreign qualifications and lack of citizenship can create additional penalties of their own.
But it is worth being exact about what has been experimentally manipulated. A name is not a laboratory switch labelled race. It can simultaneously communicate perceived ethnicity, religion, ancestry and cultural background. “Mohammed,” for example, may cause an employer to infer something different from “James” along several dimensions. The experiment therefore demonstrates discrimination on the basis of perceived ethnic or ethnocultural background as signalled by the application. It cannot tell you whether an individual employer's underlying mechanism was racial animus, religious prejudice, statistical stereotyping, anticipated customer preferences, assumptions about “cultural fit”, or something else.
That distinction does not make the finding less serious. In the ordinary labour-market sense, discrimination is precisely unequal treatment of otherwise comparable candidates because of group membership or its signal. Nor does the experiment need to establish conscious racism. An employer who sincerely believes themselves unbiased but systematically rejects “foreign-sounding” names is producing the phenomenon the experiment is designed to measure.
There is also substantial heterogeneity concealed by the 60% average. Applicants of western European or US origin were treated considerably more like white British applicants, while Pakistani, Nigerian and Middle Eastern/North African identities incurred larger penalties. The reported application penalty was roughly 70% for Pakistani-origin candidates, 80% for Nigerian-origin candidates and 90% for Middle Eastern/North African candidates. So the result should not be translated into the proposition that every ethnic minority experiences one uniform 60% penalty. “Minority ethnic” is an analytical aggregate here, not a homogeneous social category.
The historical claim is the most striking part, but also the place where I would tighten your prose most. Heath and Di Stasio's 2019 meta-analysis assembled British field experiments going back to Daniel's work in 1967, Jowell and Prescott-Clarke's 1969 correspondence study, and subsequent studies. Their statistical conclusion was that there was no significant diminution over time in discrimination risks for Caribbean applicants or for South Asians taken as a whole. That finding is entirely consistent with the Nuffield summary saying discrimination remained at levels seen in the late 1960s and 1970s.
Yet “unchanged for fifty years” is rhetorically a little stronger than the statistical evidence. “No statistically detectable decline across the available field experiments” is not quite the same proposition as “the true magnitude was literally constant for fifty years.” The early studies were small, their occupational coverage differed, methods evolved, and Britain's ethnic composition and labour market changed enormously. A meta-analysis can make those studies more comparable, but cannot turn them into a single continuous measuring instrument that has been operating unchanged since 1967.
There is another boundary worth preserving. These experiments investigate entry into employment, at the initial employer-response stage. They do not establish that discrimination has been unchanged for fifty years in pay, promotion, dismissal, occupational mobility or workplace treatment. Indeed, different dimensions of ethnic labour-market inequality can move differently. The experiment is exceptionally good evidence about one mechanism, but precisely because it is clean evidence we should resist making it bear claims outside its scope.
So I would modify the final flourish:
“Fifty years. British working life has changed almost beyond recognition. Yet across comparable field experiments, there is no clear evidence that the ethnic penalty at the first stage of hiring has diminished.”
That is actually, to my mind, more powerful than “Whatever else has changed … that has not.” The latter gives a critic an easy objection: obviously we cannot know that nothing about discrimination has changed. The narrower formulation makes a claim that is harder to escape. The extraordinary fact is not that every aspect of British racial inequality has remained frozen since the late 1960s. It is that after half a century of anti-discrimination law, enormous educational change, altered social attitudes, professionalised HR practices and a radically different labour market, experiments designed to isolate differential treatment at the point of application still failed to detect the improvement one might reasonably have expected.
That is the genuinely unsettling result.