By

Britain’s Statistics Have a Missing Variable

 

Britain’s Statistics Have a Missing Variable

Why Sex Data Matters in Public Policy

 

 

Britain’s Statistics Have a Missing Variable

A government-commissioned review has warned that Britain’s public bodies do not always collect sex data clearly enough to answer basic questions about health, justice, education and work. The dispute is no longer abstract: a census measure has been downgraded, the Supreme Court has clarified the meaning of sex in equality law, and a City of London policy now applies that ruling to a fiercely contested public service.

The conflict is often presented as a choice between recognising transgender people and preserving women’s rights. That is a false choice for statisticians. They need to know which variable they are measuring before they can discover whether it matters. A question about sex is not a question about gender identity; a record that combines both cannot later be separated by clever analysis. And a category that is omitted from the evidence cannot be restored by a forceful argument about what the evidence ought to show.

In 2025, sociologist Alice Sullivan led an independent review for the Department for Science, Innovation and Technology. It examined surveys, public records, research and institutional practice across the UK. Its central recommendation was blunt: record sex and gender identity as separate variables where relevant, and make clear which one a question is meant to capture.1 The recommendation is not a licence to collect intimate information indiscriminately. It is a demand to stop asking one question, recording another answer and pretending the resulting number remains dependable.

That case deserves a hearing. So does the evidence against the broadest claims made in its name. Confused questions, poor testing and weak communication are not, by themselves, proof that an institution has been captured by an ideology. On the 2021 Census, the official regulator found serious weaknesses in the gender-identity statistic but also said it had found no evidence that the Office for National Statistics was biased in the way critics alleged. The distinction between an error and a conspiracy is not a courtesy to officials. It is what keeps criticism useful.

The question that decides what can be known

Statistics do not begin with a spreadsheet. They begin with a decision about what a question means. “What is your sex?” and “What is your gender identity?” may appear on the same form, but they seek different information. The first can be used to compare outcomes between females and males. The second can help describe the experiences of people whose identity differs from their sex. A form that asks only “gender” may leave respondents unsure which one is wanted; an analyst may then be unable to tell what the answers represent.

This is not a semantic quibble. Researchers cannot recover a distinction that their instrument never recorded. If a hospital database stores a single field that changes when a patient’s marker changes, it may lose the ability to compare treatment or screening outcomes by sex over time. If a survey asks a hybrid question, it may gather a mixture of answers about birth registration, legal documents, presentation or identity. A large sample does not cure a muddled target. It merely produces a large number of answers to a question whose meaning is uncertain.

Sullivan’s review argues that the loss can cut in more than one direction. It says people with diverse gender identities are also poorly served when sex and identity are conflated, because researchers cannot isolate the outcomes of distinct groups. The point is not that sex explains every difference, or that identity is irrelevant. It is that a study cannot test the influence of both factors if it has collapsed them into one field. The review recommends that data owners specify the purpose of each question, collect sex by default in many research and administrative contexts, and add a separate identity measure where the use justifies it.

There are real reasons not to ask. A question may be unnecessary for the service, intrusive in context, risky if identifiable records are shared, or likely to deter responses. Researchers should minimise data and tell respondents why information is collected. But those are arguments for a proportionate design, not for assuming that missing information is harmless. The correct question is not “Should every form ask everything?” It is “What decision is this data meant to support, and what could no longer be tested if we remove it?”

The census correction is a warning, not a verdict on motive

The 2021 England and Wales Census offers the clearest case in which a contested variable produced an official correction. The census asked whether a person’s gender identity was the same as their sex registered at birth. After publication, concerns arose about the estimate and its breakdowns. New evidence from Scotland’s census indicated that the England and Wales question could have been misunderstood, particularly by some people whose first language was not English. The Office for National Statistics then asked that the figures lose their accredited status and be relabelled “official statistics in development”.

The regulator’s report was critical. It found that the question had not worked as intended and that ONS had failed to communicate uncertainty clearly. The regulator also said the national estimate of roughly one in 200 people aged 16 and over broadly triangulated with other evidence, while warning that some local breakdowns were more vulnerable to bias. It described the census question as novel and the field as one where measurement methods were still developing. Crucially, the report said it found no evidence of the kind of interest-group capture alleged by some critics; it did criticise ONS for defensiveness and for insufficiently engaging with emerging evidence.2

That finding cuts both ways. It validates the practical concern that question design and interpretation can make data unreliable. It also rebuts the claim that this episode, on its own, proves a deliberate plan to conceal biological sex. The ONS had defended the estimate; it later asked the regulator to change the statistics’ status. A correction does not erase the cost of weak design, but it is evidence that a quality-control mechanism can operate.

The sequence matters. First, an unfamiliar question was used at national scale. Then users challenged the output, further evidence exposed limitations, and the regulator required a clearer warning. That is a serious failure in the production and communication of statistics. It is not identical to suppressing a result because it is politically unwelcome. The public deserves an account of both: how the question passed testing and approval, and what safeguards will stop another measure from being treated as settled before its meaning is secure.

“Capture” is a claim that needs a chain of evidence

“Capture” is a powerful word because it converts a dispute about judgment into an accusation about control. In its classic economic use, George Stigler’s theory of regulatory capture described an industry acquiring influence over regulation so that the rules are designed and operated primarily for its benefit. The concept is not simply a synonym for a bad decision, a consultation with campaigners or a policy that critics dislike. It implies a durable channel through which organised interests redirect an institution away from its public purpose.3

Applying that idea to statistical agencies requires care. An advocacy group can raise a legitimate concern; a public body can take its advice and still act independently. Capture becomes plausible where a record shows who influenced a decision, how access or pressure changed the process, what alternatives were excluded, and whether the resulting practice persisted despite contrary evidence. Those are testable questions. “The institution uses language I reject” is not yet an answer to them.

Sullivan used the term in a 2021 case study of how ONS handled guidance for the census sex question. Her account described a draft guide that appeared to permit respondents to answer the sex item by reference to subjective identity, and a later legal challenge that forced a last-minute change before census day. She argued that the agency gave undue weight to lobbyists and neglected quantitative social scientists. That criticism is substantial, and should be judged against the contemporaneous correspondence and the legal record rather than dismissed as a culture-war complaint.

There is a further discipline for those making the charge. A case study can show how one decision was made; it does not establish that every institution, survey or official is governed by the same mechanism. Nor does proof of influence prove every downstream number is false. To turn a useful concept into a slogan is to make it unfalsifiable: a correction can be presented as admission of wrongdoing, while the absence of evidence becomes proof that influence was hidden. A serious account should say which policy changed, who pressed for it, what the evidence was, and what record would change the author’s mind.

A cancelled seminar shows how debate can shrink

The most concrete example in Sullivan’s account is not a national data series but a research-methods seminar. She says she had been invited to speak at an event organised by the National Centre for Social Research and City University about measuring sex and gender identity. After a complaint from an internal LGBT staff group and a discussion about whether she should be included, the event was cancelled. Sullivan says emails recovered through a subject access request documented the effort to keep her off the platform, and that NatCen’s chief executive later acknowledged that avoiding her participation was a factor.9

The episode matters because it concerns a professional question—how a survey should define and measure its variables—not merely a disagreement over political language. If an event is cancelled because organisers fear a speaker’s presence will upset colleagues, a discussion about the quality of a public measure has been displaced by a judgment about who may be heard. That can discourage researchers who see the subject as professionally hazardous, even if no formal rule tells them to stay silent.

But the evidence has a boundary. Sullivan’s article is a first-person case study, not an independent investigation into every participant’s motives. She reports emails, meetings and what the organisation told her; readers should distinguish those records from her interpretation that the episode illustrates policy capture. It establishes a troubling decision about one event and gives a reason to ask how it was made. It does not, without more, prove that all researchers who disagree with her are silenced or that each subsequent ONS choice followed the same pressure.

Institutions need not host every proposed event. They do owe staff and the public intelligible reasons when a professional discussion is stopped, especially when the topic is the quality of official data. A sound process would identify the concern, consider proportionate alternatives, keep a record of who made the decision, and explain why a methods debate could not proceed. The test is not whether everyone feels comfortable. It is whether the organisation can defend the decision without treating discomfort as a substitute for evidence.

Postmodern theory is not a licence to abolish facts

In the interview source, Sullivan links disputes over sex data to a “postmodern” view that naming a category or asking a question creates the reality it records. That description compresses a varied intellectual tradition into a single political accusation. Lyotard’s famous definition of the postmodern as “incredulity towards metanarratives” refers to distrust of sweeping stories that claim to explain and legitimise knowledge. It does not amount to a claim that bodies are unreal, facts do not exist or a survey creates biological sex.4

There is a more serious point beneath the shorthand. Categories used by states and institutions do not merely describe the world; they can shape how people are treated. A form can decide who is eligible for a service. A legal status can affect rights. A medical record can guide screening, medication and follow-up. Feminist and social theorists have long studied how classifications organise power and make some experiences visible while obscuring others. That insight can support better measurement: ask who created the category, for what purpose, and what happens when it is applied.

It does not follow that measurement creates every material fact it records. The census may influence how a population understands itself, but its question does not create the body of each respondent. A legal document can alter a person’s status in law without changing every physical characteristic relevant to medical treatment. Conversely, an observed biological fact does not settle the social or legal treatment a person should receive. These propositions can coexist. Confusing them produces a debate in which one side treats every category as oppression and the other treats every measurement choice as politically neutral.

The useful challenge is to ask advocates for a causal account. If collecting a particular variable causes harm, what kind of harm, through what mechanism, and can it be reduced by privacy protections or different wording? If the concern is that a question forces people into a binary, can an identity measure be collected separately? If the concern is discriminatory use, can access be controlled and decisions audited? “Categories have effects” is a reason to design them responsibly. It is not a reason to assume that silence improves knowledge or fairness.

The law settled one definition; it did not settle every service

The UK Supreme Court’s 2025 ruling in For Women Scotland Ltd v Scottish Ministers clarified that “sex”, “man” and “woman” in the Equality Act 2010 refer to biological sex. The court was deciding a question of statutory interpretation: whether a Gender Recognition Certificate changed a person’s status as a woman for the purposes of that Act. The judgment gave public bodies a clearer legal reference point than they had before. It did not declare that every service must adopt one access rule regardless of its purpose, circumstances or other legal duties.5

That distinction is visible at Hampstead Heath. In July 2026 the City of London Corporation agreed that Kenwood Ladies’ Pond would continue to admit biological and trans women, while Highgate Men’s Pond would admit biological and trans men; the mixed pond remains open to everyone. The City said 86 per cent of more than 38,000 consultation respondents favoured keeping the existing arrangements. It also said its decision followed legal advice and that a judicial-review challenge was continuing.6

Those facts are not proof that the policy is lawful, or that it is unlawful. The authority’s policy page is an account from the decision-maker, not a court ruling. The consultation itself was not a representative referendum: a large number of submissions demonstrates strong engagement, but does not show how the whole population—or even all regular users—would answer a neutral survey. The eventual legal challenge will test the policy against the statute and the facts of that service.

The debate at the ponds is also a reminder that a general legal definition and a practical access policy are different layers. A law can establish what a protected characteristic means; a provider still has to decide how an exception or service rule applies, and justify it. This is where accurate data can help rather than replace judgment. Records of use, complaints, safety incidents and user experience could inform whether a policy meets its stated purpose. The evidence must be collected in a way that does not presume in advance which group’s experience matters less.

Health records cannot answer questions they were never built to answer

Health is where the measurement dispute is least amenable to slogans. Sex can be relevant to pregnancy, reproductive care, disease risks, screening and treatment response. Gender identity can be relevant to a patient’s experience, communication, discrimination and care needs. Neither field is a complete substitute for the other. The appropriate measure depends on the clinical question, and sensitive information should be available only to staff who need it.

The Sullivan review warns that a record system that overwrites sex with a changed marker may prevent longitudinal research or leave clinicians without information relevant to care. It recommends keeping the underlying variable and recording forms of address separately. Critics of such an approach worry that retaining sensitive data could expose a person’s history or invite disrespectful treatment. That is not an imaginary risk. The answer must include clear access controls, accurate notes about how a variable was obtained, secure records and a reason for every item collected.

Referral patterns among young people show why sex can matter in analysis without explaining the individual. The Cass Review described the shift in the UK’s specialist referrals: the earlier pattern was more male-heavy, while referrals later rose sharply among adolescents registered female at birth. A peer-reviewed systematic review by researchers at the University of York found a twofold to threefold rise in referrals across studies and a growing ratio of birth-registered females to males over time. It also warned that evidence about the characteristics and pathways of these young people remained limited. The figures describe who reached specialist services; they do not explain why, establish a single cause or tell us what care any one patient needs.7

That qualification is essential. An aggregate shift can generate hypotheses about clinical practice, social change, referral patterns or unmet needs. It cannot adjudicate among them by itself. If datasets fail to record sex, age, referral source, co-occurring conditions and outcomes consistently, researchers may be unable to examine what changed. The conclusion is not that sex should dominate a patient’s care. It is that a service needs the right data to test whether its model works for different patients, and the humility to say when the answer is not known.

Public opinion is evidence, not a substitute for method

Arguments over sex and gender often arrive in the language of “what most people think”. The claim can be true and still be methodologically weak. A poll depends on the exact wording, the sample, the order of questions and whether respondents understand the terms. “Should trans people be treated with respect?” is not the same question as “Should this single-sex service admit a particular group under these conditions?” Nor does an answer to one automatically predict an answer to the other.

The Hampstead Heath consultation illustrates the distinction. Thirty-eight thousand responses are too many to dismiss as a handful of activists. Yet self-selection means the number cannot be treated as a probability sample of Britain, London or all pond users. The City itself says the consultation gathered views to inform a decision rather than determine the law. A public body can take those responses seriously while declining to claim that they establish a population-wide mandate.

Sullivan and Helen Joyce discuss preference falsification: people may hide views when they fear social or professional penalties. That mechanism can make public expression a poor guide to private belief. But it cannot serve as a universal solvent for inconvenient polling. If someone says a majority is silent, the claim needs independent evidence: confidential surveys, consistent responses to neutrally worded questions, or records of sanctions for dissent. An assumption that quiet people secretly agree with one side is not better evidence than an assumption that loud campaigners speak for everyone.

There is a practical way out of the manufactured binary. Ask separate questions about dignity, legal definitions, sex-based data, service access and the effects of a specific policy. Publish the wording and method. Show disagreement rather than compressing it into a single “support” score. Public opinion should shape democratic decisions; it should not be used to override statistical standards, legal protections or evidence about a specific service. A poll can tell policymakers what people said. It cannot relieve them of explaining what they asked.

The price of dissent is visible in the Ngole case

A test case of how contested beliefs can carry professional consequences. Felix Ngole, a Christian social worker, won a Court of Appeal case in 2019 after the University of Sheffield had removed him from its social-work course over public Facebook posts about homosexuality. He completed his master’s degree in 2021. In 2022, Touchstone Leeds offered him a mental-health support-worker role, then withdrew the offer after managers found reports about his views online. The charity said it needed confidence that he could support LGBTQI+ service users, comply with its policies and carry out the job safely.

The dispute is not reducible to “a Christian was fired for his beliefs” or “a charity protected vulnerable clients”. The Employment Tribunal found that the initial withdrawal of the conditional offer was direct discrimination because of his religious belief, but it rejected parts of his claim concerning a second interview and the final refusal to reinstate the job. In February 2026 the Employment Appeal Tribunal held that the tribunal had not properly analysed whether concern about service users discovering his beliefs was separable from concern about how he would do the work. It sent those parts back for reconsideration, while confirming that the employer could seek reassurance about how he would support service users and meet the role’s requirements.8

The case does not prove that “progressive ideology” controls public employment. It does show why assertions about a chilling climate should be handled through evidence rather than rhetoric. The employer’s stated safeguarding concern is real; so is the legal risk of penalising someone merely because others may object to a protected belief they discover online. The appeal tribunal required those reasons to be separated and assessed, not bundled together under a vague appeal to inclusion.

That is the standard institutions should bring to public data as well. A belief can be unpopular without being a disqualification. A professional can be questioned about conduct that affects a role without being required to pretend that a contested moral or religious belief has disappeared. The line is difficult, but the tribunal’s demand for a reason-by-reason analysis is more useful than declaring either side beyond scrutiny. Written reasoning, fair process and a record of what decision-makers actually relied upon are not bureaucratic luxuries. They are how a dispute can be tested after the people involved have left the room.

Counting inequality does not create it

Sullivan’s income analogy exposes a recurring error. If asking about a category is said to create inequality, should government stop asking what people earn and declare poverty ended? Of course not. Income differences exist whether or not a survey records them. Refusing to measure deprivation may make an uncomfortable chart disappear, but it does not pay a bill, improve a wage or redistribute a fortune.

The analogy has limits. Income is not sex; the privacy risks and purposes of each question differ. It would be a mistake to use the comparison to justify collecting every characteristic on every form. But it clarifies a basic point: measurement can reveal a disparity without causing it. A category may affect how people are treated, and a badly designed question may reinforce stereotypes, but the argument that measurement itself creates the underlying material difference needs evidence.

Better statistics do not automatically produce fair policy. A government can count poverty accurately and still fail to reduce it. A health service can know who is at risk and still provide poor treatment. Data can be misused, and the publication of small-cell statistics can identify individuals. These are arguments for careful collection, restricted access and transparent limits. They are not reasons to confuse ignorance with equality.

The opposite error is to assume that every difference between females and males is innate, fixed or policy-relevant. Data identify patterns; they do not explain them. Researchers must distinguish association from cause, disclose uncertainty and avoid treating averages as descriptions of each person. If the answer is that no meaningful difference appears in a particular setting, accurate sex data can establish that too. A category kept available for analysis is not a conclusion imposed in advance.

Repair the method before declaring victory

The practical programme begins with definitions. Every survey and administrative system should state whether it records biological sex, legal sex, gender identity, a form of address or another construct. The label on the field should match what the respondent is being asked. If multiple concepts are necessary, separate them. If a variable is not relevant, do not collect it. If it is relevant but sensitive, state the purpose, explain who will see it and make refusal possible where law and the study design allow.

Second, test questions on the people expected to answer them, not only with a small group of advocates or specialists. Cognitive testing should reveal how respondents interpret a question, including those with different language skills and levels of familiarity with the subject. Publish the wording, instructions, response options, coding rules and any changes made after testing. Report non-response and uncertainty plainly. When a question fails, say so before users build policy on its output.

Third, create a route for statisticians, clinicians and public servants to challenge a method without being forced to prove bad faith by the people who designed it. An allegation of bias should be investigated; it should not become an automatic verdict. Equally, officials should not dismiss criticism as ideological merely because it comes from an organised group. Keep minutes, disclose conflicts, distinguish evidence from advocacy and provide a review process independent of the programme being criticised. A frank correction is less damaging to trust than a prolonged defence of a weak measure.

Finally, evaluate a policy against its stated purpose. For a health service, test whether patients receive suitable care and whether records permit safe follow-up. For a single-sex facility, identify the legal basis, the interests the rule protects and the effects of the alternatives. For a census measure, ask whether respondents understood the question and whether the resulting figures can bear the level of detail users want. Distinct data will not settle political disagreement. It will make it harder for either side to claim certainty where the evidence has not been collected.

The public argument has spent years treating data collection as a referendum on identity. It is better understood as a promise made by institutions: ask a question clearly, record the answer faithfully, protect the person who supplied it, and permit the result to challenge the policy that follows. When the question is built to blur what it claims to measure, that promise cannot be kept.

  1. Alice Sullivan, Review of data, statistics and research on sex and gender: executive summary, UK Department for Science, Innovation and Technology, 19 March 2025. https://www.gov.uk/government/publications/independent-review-of-data-statistics-and-research-on-sex-and-gender/review-of-data-statistics-and-research-on-sex-and-gender-executive-summary ↩
  2. Office for Statistics Regulation, Review of statistics on gender identity based on data collected as part of the 2021 England and Wales Census: Final report, 12 September 2024. https://osr.statisticsauthority.gov.uk/publication/review-of-statistics-on-gender-identity-based-on-data-collected-as-part-of-the-2021-england-and-wales-census-final-report ↩
  3. Sam Peltzman, “Stigler’s Theory of Economic Regulation After Fifty Years”, University of Chicago Law School, 2021. https://chicagounbound.uchicago.edu/law_and_economics/944 ↩
  4. Peter Gratton, “Jean François Lyotard”, Stanford Encyclopedia of Philosophy, first published 21 September 2018. https://plato.stanford.edu/entries/lyotard/ ↩
  5. UK Supreme Court, For Women Scotland Ltd v The Scottish Ministers, [2025] UKSC 16, judgment given 16 April 2025. https://supremecourt.uk/cases/uksc-2024-0042 ↩
  6. City of London Corporation, “Hampstead Heath’s Bathing Ponds Access Policy”, updated 31 July 2026. https://www.cityoflondon.gov.uk/things-to-do/green-spaces/hampstead-heath/activities-at-hampstead-heath/swimming-at-hampstead-heath/hampstead-heath-bathing-ponds-access-policy ↩
  7. Jo Taylor et al., “Characteristics of children and adolescents referred to specialist gender services: a systematic review”, Archives of Disease in Childhood, 2024. https://pubmed.ncbi.nlm.nih.gov/38594046/ ↩
  8. Employment Appeal Tribunal, Mr F Ngole v Touchstone Leeds, [2026] EAT 29, judgment 16 February 2026. https://www.gov.uk/employment-appeal-tribunal-decisions/mr-f-ngole-v-touchstone-leeds-2026-eat-29 ↩
  9. Alice Sullivan, “Sex and the Office for National Statistics: A Case Study in Policy Capture”, The Political Quarterly, 2021. https://discovery.ucl.ac.uk/id/eprint/10130795/9/Sullivan_Sex%20and%20the%20Office%20for%20National%20Statistics-%20A%20Case%20Study%20in%20Policy%20Capture_AOP.pdf ↩

This post contains affiliate links. If you purchase through these links, I may earn a commission at no extra cost to you.

 

Leave a Reply

Discover more from Thoughts on Technology

Subscribe now to keep reading and get access to the full archive.

Continue reading