Last September, the government unveiled new league tables ranking every NHS provider. When I raised concerns that they could drive gaming rather than genuine improvement, and that they wouldn’t help the public make informed choices, the then Secretary of State Wes Streeting dismissed my concerns as “elitist nonsense”. It was suggested that the Nuffield Trust wanted to gatekeep important information from the public.
My response then, as now, is that the Nuffield Trust has always believed that good, accessible, meaningful and actionable data is essential. The question was never whether or not good, insightful data should be published to help the public’s understanding, it was whether these tables would deliver that.
Nine months and three quarterly publications later, I don’t think they do. In fact, I think they are actively obscuring the picture rather than illuminating it.
How are they meant to work?
The tables rank 205 NHS trusts – acute, mental health, community, and ambulance services – across approximately 30 indicators covering waiting times, cancer access, urgent and emergency care, and financial balance. Scores run from 1 (best) to 4 (worst), producing an average delivery score that places each trust into one of four segments. Segment 1 represents the strongest performers, and segment 4 those with the broadest range of challenges. There are three separate tables covering acute trusts (134), non-acute and mental health trusts (61), and ambulance services (10). What is new about the league tables, as distinct from the national oversight framework, is that they were presented as patient facing and of direct benefit to patient choice.
The then Secretary of State’s explicit hope was that trust leaders would be shamed into taking actions they weren’t currently taking. The history of shame-based performance management in the NHS is mixed at best. Shame can produce movement – organisations scramble to escape the bottom of the rankings and, yes, to escape the shame. But if the league table process is seen to be flawed, does that really drive the right behaviours, or the right culture?
League tables can mean doing things that improve a league position without improving the care that patients actually receive. The academic literature on gaming in public sector performance regimes documents this pattern extensively, and the structural features of these tables – ranking organisations relative to each other rather than against absolute standards, and updating quarterly – create exactly the conditions in which gaming behaviour tends to emerge.
When the tables don’t add up
My core concern last year was the significance given to how healthy (or unhealthy) a trust’s finances are in the rankings, and nothing in nine months since has allayed it. Financial performance matters – NHS organisations need to live within their means and value for money is a legitimate public interest. But while financial balance is a reasonable indicator of efficiency and performance (with its own caveats), it’s a poor proxy for the quality of care a patient will receive, and is irrelevant for patient choice. Treating it as a central quality signal is problematic for several reasons.
First, with the majority of NHS trusts currently in deficit, financial position reflects systemic funding and cost pressures as much as it reflects local management performance. It is hardly fair to penalise a trust whose local population is sicker or older, and therefore more costly to care for, than a neighbouring area where the people are healthier.
Second, and more fundamentally, there is no evidence that financial deficit is central to how the public thinks about health care. The British Social Attitudes survey suggests that, when people are asked about NHS priorities, they focus overwhelmingly on access and capacity: easier access to GP appointments, shorter waits in A&E, shorter waits for planned operations, and more NHS staff.
The public are also capable of holding two views at once: most think the government spends too little on the NHS, while many also doubt that the NHS uses its existing money efficiently. But that is very different from treating an individual provider’s balance sheet as a meaningful signal of the care that a patient will receive. Presenting financial performance as if it were a central measure of quality risks confusing patients rather than informing them.
This problem is most apparent in the mechanism within the league tables that applies to automatically cap any trust in financial deficit (the “financial override”), which means it cannot be ranked any higher than segment 3, even if the quality of care is exemplary.
Gloucestershire Hospitals fell 21 places in the acute trust table in Q3 – not because patient care deteriorated, but because of the financial override. Sheffield Teaching Hospitals fell 60 places in Q4 – the largest single-quarter drop recorded across all quarters of the framework – on a similar basis. These are financial signals being relabelled as indicators of quality in a public-facing ranking, which is a problem if they are interpreted as reflecting the quality of care that patients will receive.
There is a related structural problem that the framework has not adequately addressed: the majority of indicators, based on waiting times and infection control, are not adjusted for the complexity or deprivation of the population that a trust serves. A trust serving a highly deprived, elderly or clinically complex catchment faces structurally harder targets on many metrics relative to a trust serving a healthier, more affluent population.
Richard Lilford and colleagues have identified this as a basic validity requirement for any performance comparison – comparative measures of institutional performance are only meaningful if you account for the differences in the populations being served. That adjustment is absent here, which means the tables systematically disadvantage trusts serving the populations with the greatest need. That is not a minor point – it is a fundamental problem with what the tables are actually measuring.
There is also a deeper statistical problem that runs through the entire framework. Roughly one in five trusts changed segment every quarter – 38 in Q2, 42 in Q3. On the face of it, that looks like a system in constant flux, with substantial numbers of organisations meaningfully improving or deteriorating. But the academic literature would predict that a large proportion of this churn is statistical noise rather than genuine quality change.
The NHS framework does attempt to address this: it applies 95% and 99.8% confidence thresholds and flags changes that cross those boundaries as significant. But in Q3, for instance, only two trusts met the 99.8% threshold and only 15 cleared 95% – yet 42 changed segments.
NHS England themselves acknowledge this, cautioning that segment changes alone should not be taken as evidence of meaningful performance change. Which raises an obvious question: if the regulatory body publishing the tables urges that caution, what signal are they intended to send? And how many people have the financial literacy to understand this? Producing tables with many pages of footnotes that most people don’t have the skills or knowledge to interpret certainly doesn’t help the public’s understanding.
The importance of real transparency
None of this critique should be taken to mean that I believe transparency is wrong, or that comparison is not valuable. It clearly is. Nor does it mean the NHS should be exempt from public accountability – it absolutely should not be. The Nuffield Trust has argued consistently for better, richer, more meaningful data about NHS performance to be in the public domain, and we will continue to do so.
But data that confuses more than it illuminates – which mislabels financial distress as poor care, and statistical noise as performance change, and which fails to adjust for the populations that trusts are serving – is not transparency. It is the appearance of transparency. It creates the impression of rigorous public accountability but does nothing to aid the public in navigating the complexity of evaluating the services they are being directed towards.
Management tool masquerading as public information
We all need clear information that is presented in a timely fashion, about a service we are using, and which helps us to make the informed choices we need to make or simply to understand the performance of the service we have been referred to.
Nine months ago, I said these tables were unlikely to help the public navigate care or drive the right improvements in the service. Three sets of data, and some very dramatic rank movements that have more to do with accounting than with clinical quality, have not changed that view.
At the end of the day, the tables are a management tool masquerading as public information. The question to ask of them is not whether they work, but who are they for.
As a performance management instrument – internal and read by people who understand their limits – they are understandable. Driving trust performance and monitoring financial sustainability are legitimate aims. But that is not how the league tables are presented. Offered to the public as a guide to the quality of their care, they fail. The framework’s own caveats explain why: segments turn on tiny movements near arbitrary boundaries, financial deficits cap operational scores regardless of clinical quality, and there is not yet enough data to separate noise from real change.
What patients need is clear, relevant, easy-to-understand information about the services they are using and, where it applies, the ability to make a timely, informed choice. The league tables do not provide that, and no amount of refinement will make them do so. As a public-facing product, they should go.