Pages

Friday, 30 April 2021

Should GPs turn off their online systems to control demand?

There has been a huge amount of comment in recent discussions about GPs being overwhelmed by the current level of demand. Many have proposed turning off their online systems to hold back the wave. This is problematic for many reasons but has highlighted many deep misunderstandings about total triage and online tools.


The right answer isn't that they should turn off online access, but to test whether or not doing so makes any difference. Or, better, they should design effective triage processes to ensure their limited capacity is used to treat the patients that most need their help without driving all the doctors into burnout.



How did GPs control demand before online systems?

In the world before online total triage solutions were common and before the pandemic changed everything, GPs were used to limiting access as a way of controlling their workload. 


The majority ran systems where the only way to get an appointment was to phone the practice between 8am and 9am. Then, even if you got through to reception, you would often be lucky to find an appointment slot free, especially if you wanted an urgent one. Some GPs offered open access clinics for urgent cases, but those require the patient to turn up in person in a narrow time slot and then wait a potentially extended period of time in a crowded waiting room.


There are problems with the traditional approach. One is that many patients can't get through on the phone because the lines are too busy. We don't know how many requests are lost to the system that way because almost nobody even tries to measure them. Another is that turning up in person for a long, unpredictable wait is often not possible for the patient so this discourages them from going to the clinic at all. Yet another is that the allocation of requests to same day slots is, in effect, unrelated to the patients' need as it depends more on the luck of the draw during the early morning phone lottery than it does on the urgency or severity of their condition.


On the upside, these approaches do allow GPs to limit the demand they see (though partially by making it entirely invisible).



Does turning off online tools work?

The big change from moving significant volumes onto online triage systems is that the capacity of the phone lines is irrelevant: online demand can't be blocked by switchboard capacity. It is far easier to make an online request than it is to make 30 attempts to get connected to reception by phone in the early morning rush. 


So GPs see more of the actual demand. Especially if the online systems are open 24 hours a day and at weekends. Many commentators have speculated that this greatly increases demand. They argue that making it easy for patients to contact the GP must increase demand (and some economists say this is exactly what theory would predict). But few bring data to the argument. 


Not looking at the data risks two serious errors. 


The first is to blame online tools for the increase in demand. The current debate has been triggered because many GPs are unusually busy right now. But many of those GPs have been online for a year and the excess demand has only materialised in the last month or two. The only thing being online has done is to make visible actual demand that would otherwise be lost to the morning phone lottery. It isn't the use of online tools that have caused the demand leap. If demand is going up now, the cause is something else.


The second error is the assumption that limiting the time online systems are open makes any difference to the overall weekly demand. This seems to feel right to many GPs used to limiting demand by limiting appointment slots or rejecting requests that can't get through in the morning phone rush. But we have good data on how patterns of online demand change when systems are turned off. Some smart GPs have run tests of how much demand they get with 24hr systems compared to systems only open in the working day. 


Everyone who has properly tested closing their online system has found that the major effect is not to lower demand but to redistribute it to the times when the system is open. Some, but certainly not all, have found that there is a small reduction in total weekly demand. And, for some, that might be a good outcome (though see what this implies below). 


Absolutely nobody–and this is my most important message here–should be shutting down online access without having done a proper trial to see what the result is. Closing down your online system in the vain hope that it will reduce demand and make your life easier is futile.


In short, limiting access to online systems might help, but no GP should do it unless they know it will, and that requires a controlled experiment collecting good data.


What does limiting access achieve for patients?

The rarely asked question about all schemes to control GP demand is what are the implications for patients? This question is obviously relevant to any debate on online triage tools but it is also relevant to every GP who uses the phone lottery to limit incoming requests or who denies the possibility of appointment bookings when the current slots are all filled.


There are two levels of problem with the traditional "phone the GP between 8am and 9am to get an appointment" model. 


The first is that many patients never even get through to the switchboard. We don't know how many because few GPs have set up their phone systems to monitor that statistic. This has the useful benefit of limiting the demand that reaches the practice. It also has the "benefit" that the unmet need never appears on any report.


The second level, where the patient gets through on the phone but can't get an appointment because all the slots are full, has a similar effect, though the numbers are more likely to be visible somewhere.


If your only concern is to manage the workload, then these methods are fine. The same–apparent–logic applies to limiting the opening hours of an online system (though, as I said above, it is a great deal less clear whether managing access online actually works as effectively as a busy switchboard).


But, if your concern is to actually help the patients who most need your attention, then all attempts to limit incoming requests are very bad. The morning phone lottery at best limits access randomly and, at worst, limits it to the patients with the fastest redial fingers who are, needless to say, not the ones with the greatest medical needs. Limiting appointment slots by a "first come, first served" process has a similar effect. 


The big difference with online triage tools is not that they create more demand; it is that the practice knows what the real demand is. There are no busy switchboards with online demand to arbitrarily reject the request because the patient can't get through. And closing the online system out of hours has a far smaller effect on incoming demand than a busy switchboard (patients quickly learn when the online system is open and make their requests then).


Bizarrely, one of the most common objections to using online systems at all is that they might bias access to particular groups. But compared to the traditional system that randomly rejects patients with slow redial fingers, this is a pretty weak objection (also, for well designed online tools that patients choose to use, the phone lines are far less congested and access for those who don't want to go online gets better). 


There are real tradeoffs to consider when GPs reject online tools. I think they don't increase the underlying level of demand (much of my data says so, but more providers need to release their data to independent researchers to show this is true in general). But they do make that demand more visible. If the response is to reject online tools because they don't arbitrarily limit visible demand, then practices must recognise the cost to patients of the traditional process. The arbitrary selection of requests, unrelated to need. The frustration to patients of the morning phone lottery. And the safety implications of random request prioritisation. GPs sometimes complain that they are overwhelmed by too many trivial, unimportant, non-urgent requests, but it is hard to see how the traditional process solves that when it rejects requests at random not by triage based on need.


And this brings us to a possible solution: effective triage.


If demand is overwhelming then triage is a better solution than limiting access

GPs and primary care researchers sometimes have a strong antipathy to triage (online or otherwise). Helen Salisbury (a GP and researcher) made this argument in a recent BMJ column called Triage is for disasters, not everyday general practice. She argued that, while the idea might be appropriate for front-line military hospitals where there just isn't the capacity to save every wounded soldier, it was completely inappropriate for GPs. But her argument contains several problems.


In a front line military hospital a wide range of injured soldiers arrive. Some have injuries that won't kill them soon and can be treated later; some might die whatever the medics do; some will survive if they are treated immediately. Triage is the sorting process that identifies which is which and allows the medics to focus their limited capacity where it will have the biggest benefit. GPs don't have decisions of the same magnitude, but they do have to see patients with a wide range of needs and should focus their attention where it can do the most good.


Salisbury argues that GPs should see everybody (and should probably do so face to face f-to-f). The implication is that GPs have the capacity to do this and that a one-size-fits-all f-to-f is the best response for every patient. But the whole point of GPs unhappiness with online tools is that they don't have the capacity. We don't seem to live in the world she describes.


She also makes an argument–and one that many researchers seem to agree with–that GP triage actually takes more time than just seeing everyone. This suggests some very deep misunderstandings about what triage is supposed to do or, possibly, how GPs need to do it. The military analogy should make this clear. If sorting the wounded took more time than treating them, triage would be killing more soldiers than not triaging. But triage was invented to save lives by sorting quickly so medical effort could do the most good with the limited resources available. If GPs are taking more time overall with a triage process than without one, then, whatever they are doing, it isn't triage


I don't want to pretend that a good triage process is easy. It isn't, especially when GPs are trained to be one-club golfers who only ever handle patients using short f-to-f meetings. But a disciplined process, especially an online one where the patient request can be assessed rapidly and asynchronously, can be fast and effective. The point is to sort the patients by urgency and need. And to use a flexible range of responses many of which are faster than bringing the patient in for an f-to-f. Even better, if a notable proportion of the requests are the trivial sort that GPs complain overwhelm their capacity, they can be dealt with instantly by a simple, pre-written, short message. 


One of the huge benefits of an effective triage process is that it can do a better job of managing demand than a process that involves limiting access. The urgent serious cases can rise to the top of the queue and get an appropriate amount of GP time; the trivial ones can be dealt with by short standard responses; the middling concerns can be dealt with quickly by message, saving the patient and GP the time taken by an f-to-f appointment. 


If the practice finds it is overwhelmed by demand, it can respond by altering the thresholds for each type of response. Why this is widely considered a worse way to handle demand than a system that arbitrarily denies patients–some of whom might have a serious or urgent need–the ability to make a request at all is something of a mystery.


Some experts object to this because many patients won't get an f-to-f response. They argue that patients mostly want an f-to-f response and will be unhappy if they don't get one. Or that f-to-f is vastly safer than alternatives. Martin Marshall recently made this argument in The Guardian and other places. There are two overwhelming objections to this argument. First, GPs don't have the capacity to do f-to-f appointments for all their current demand (that was, after all, how the whole debate about them being overwhelmed by online demand started). And randomly limiting patient access to even make a request can hardly be safer than triaging requests. Secondly, his claim about what patients want is just wrong. I know this because I have the data. AskmyGP is an online triage tool that has handles most of the incoming requests for several hundred GP practices (70% online, 30% by phone). For every request the patient is asked what sort of response they prefer (f-to-f, message, phone call, video call…). Before covid took off only about one in three requested an f-to-f.


In short, an effective triage process can direct limited GP time to those who need it and can also adjust to prevent that demand causing GP burnout.


Conclusion

There are several key themes here:

  • Online tools don't create extra demand, though they may make changes in demand more visible

  • Limiting access to online tools might sometimes work, but you need to look at the evidence for your practice rather than making knee-jerk decisions

  • Traditional methods of "demand management" have serious and often ignored issues

  • Online tools alongside efficient triage processes offer a better, safer route to manage GP capacity and avoid burnout

Monday, 22 March 2021

What the military can teach the NHS about how to get the right things done in the most effective way

Everyone has a plan until they get punched in the mouth

Mike Tyson, boxer


...no plan of operations extends with any certainty beyond the first contact with the main hostile force


Helmuth von Moltke the Elder, 19th century Chief of the Prussian General Staff


The NHS swings between bouts of centralised power in the hands of ministers and decentralised decision making by its regions and units in repeated attempts to achieve its goals. But it never seems to achieve the goals desired. I suspect its leaders (both politicians and bureaucrats) have never learned an important lesson that many military forces learned the hard way.



The current NHS white paper proposes a great deal of new centralisation of power in the NHS, a major reversal of Blair-era reforms which tried to decentralise decision-making. My best guess as to why this is happening is that ministers feel the NHS is not achieving what they want and the only effective way to do so is to take back power. This is wrong. Though it is far from obvious that decentralisation worked either and the real problem is a fundamentally misguided view of how to get things done in a big organisation that has persisted in the NHS for decades.


Other big organisations have taken a different approach. Many militaries learned the hard way how to achieve their goals and how they did so provides useful lessons for the NHS.


The fundamental idea that governs the NHS is that people should "follow the plan" and do as they are told. They should achieve things by following the how of the plan and are frequently judged on whether they follow the right process to achieve the goal. Metrics are often if not usually process metrics not outcome metrics. This philosophy of management (or culture) has persisted over time despite major changes in the structure of NHS management which has swung from centralised to decentralised and back over the last few decades.


This follow-the-plan philosophy is how many militaries used to exercise command and control in their troops. This was how the British military worked at the start of WW2 until the outcomes forced a major rethink. 


A good illustration of the problem is the campaign in North Africa. Until the decisive defeat of the Afrika Corps at the second battle of El Alamein, the British forces were repeatedly defeated by a vastly inferior German and Italian force (the Allied forces almost always outnumbered the Axis forces in men, tanks, guns equipment and had far fewer persistent shortages of supplies). According to the insightful analysis of Stephen Bungay (in his books Alamein and The Art of Action) this was not–as is often reported–because their leaders were better (though Rommel was an excellent tactician) but because the command and control doctrine of the German army was far superior.


A simple way to understand why is encapsulated succinctly by the Mike Tyson quote and at more length by Moltke: rigid plans are derailed by the messiness of reality which discombobulates the plan as soon as the enemy doesn't act as expected. The British army doctrine at the start of the war demanded that soldiers follow the plan, in detail. Junior officers were disciplined if they deviated from it. Following the orders in detail was the goal. The Germans didn't do that. They knew what Mike Tyson and Moltke knew: the details of the plan fall apart as soon as you meet the enemy. Instead, the Germans issued objectives for all levels of command about the key goals rather than detailed plans about how to achieve the goal. All levels of officers were given wide discretion about the how. They were expected to innovate when they saw the reality on the ground as long as they pursued the key goal they had been set. This led to persistently better decision making in actual battles leading to their inferior forces often winning key engagements with the rigid, hidebound Allied forces (at least until the allies had overwhelming superiority).


German command doctrine (now called Mission Command and widely adopted as a central philosophy by the modern US and British armies) was that a detailed plan is derailed by real events (like getting punched in the face or when the enemy doesn't act as expected in the plan) so tactical innovation in response to reality is far more likely to achieve the goal than blindly following the details of out-of-date orders cooked up by commanders a long way from the action.


But this is pretty much how the NHS works. Goals are not set in terms of outcomes but in terms of following the plan. Do things the way the centre tells you to. And suffer punishment if you deviate from the process, which we will measure. And they won't measure the outcomes because that is far too hard to measure in the short term. The centraising plan will make this worse, not least because it is the equivalent of a frustrated Churchill demanding direct control of the deployment of individual tanks in the desert because his generals were not winning their battles. The NHS has rarely tolerated local innovation that delivers better outcomes and often seems to regard innovation as "rocking the boat" or "deviating from the process".


This is a critical failure. Unless the NHS radically changes its management philosophy and culture from one where following the plan is rewarded to one where local leaders are rewarded for achieving better outcomes by innovating, it will never get better at achieving better outcomes however the overall structure is changed. 


The NHS plan should focus on the management culture not the structure. It should define the broad goals (outcomes) desired by the government or the leadership and leave the front line leaders to find better ways to achieve those goals. And it should radically cut the process ("do things this way") metrics and use more achieve this outcome metrics. And don't forget culture change: culture may reappear and control the results even if the structure is new. A change in management structures will always be trumped by a persistent management culture.


It cannot be said that the military have always remembered this lesson. An analysis of the army's poor performance in the second Gulf War pointed out that Mission Command is not very useful when you have no idea what your mission is (see the introduction to The Good Operation which also has other useful conclusions about public policy that don't seem to have influenced the NHS plan). Perhaps this too is a problem affecting the NHS. Which suggests that the NHS might have been better analysing why it frequently fails to achieve better outcomes before writing a plan that reinforces the very cultural problems that make innovation and flexibility so difficult.



Friday, 8 January 2021

The next Chernobyl will be in the west

The brilliant HBO/Sky drama Chernobyl tells the story of the world's worst nuclear accident. The accident was avoidable. It happened because previous dangerous events in reactors of similar design (including one of the other reactors at Chernobyl) were suppressed by the Orwellian Soviet government because they didn't want to admit their technology designs were flawed. 


The first words spoken in the series are:


What is the cost of lies?


It is not that we will mistake them for the truth.


The real danger is that if we hear enough lies, then we will no longer recognise the truth at all.


The deaths, the huge cost, the vast radioactive contamination across Europe, the damage to Soviet credibility were all avoidable if people only had access to the truth. 


And the key facts were suppressed because an Orwellian government embarrassed to admit flaws in its technology deliberately withheld it.


Plenty of other disasters occured in Soviet times because the government sought to hide the truth from the people who needed it.


Suppressing the truth is bad.


I'm worried that the United States and some other western countries are now facing a similar truth deficit that will lead to many similar disasters. But they have chosen a totally different way to suppress truth.


Let me explain why.


Orwell vs Huxley

Neil Postman's book, Amusing Ourselves to Death,  was published in 1985 and argued–before the existence of Facebook, Twitter, social media or the World Wide Web–that the media of the day had adopted the values of entertainment and had elevated keeping the audience's attention over telling them the truth. Needless to say, he thought this was dangerous in a society where knowing the truth was important when decisions have to be made.


He pointed out that many thinkers in the west had recognised the importance of truth in public discourse but were worried about the danger of Orwelian truth suppression, which they saw in communist and other totalitarian regimes. He thought they had missed the alternative model where truth is suppressed by swamping it with "soma", the model of Huxley's alternative dystopia, Brave New World. As Postman wrote:


What Orwell feared were those who would ban books. What Huxley feared was that there would be no reason to ban a book, for there would be no one who wanted to read one. Orwell feared those who would deprive us of information. Huxley feared those who would give us so much that we would be reduced to passivity and egoism. Orwell feared that the truth would be concealed from us. Huxley feared the truth would be drowned in a sea of irrelevance. Orwell feared we would become a captive culture. Huxley feared we would become a trivial culture, preoccupied with some equivalent of the feelies, the orgy porgy, and the centrifugal bumblepuppy. As Huxley remarked in Brave New World Revisited, the civil libertarians and rationalists who are ever on the alert to oppose tyranny "failed to take into account man's almost infinite appetite for distractions." In 1984, Huxley added, people are controlled by inflicting pain. In Brave New World, they are controlled by inflicting pleasure. In short, Orwell feared that what we hate will ruin us. Huxley feared that what we love will ruin us.


This, remember, was well before the invention of the sort of social media that provides the equivalent of an intravenous feed of soma to the masses. Or, where an American President's Twitter feed might reasonably be described as centrifugal bumblepuppy. 


Bullshit versus lies

The second key source relevant to the current crisis of truth is Henry Frankfurt, the author of the excoriating analysis On Bullshit. Though the book version was only published in 2005, the original argument was made in 1986 not long after Amusing Ourselves to Death.


In a core section he argues this (my highlighting):


Someone who lies and someone who tells the truth are playing on opposite sides, so to speak, in the same game. Each responds to the facts as he understands them, although the response of the one is guided by the authority of the truth, while the response of the other defies that authority and refuses to meet its demands. The bullshitter ignores these demands altogether. He does not reject the authority of the truth, as the liar does, and oppose himself to it. He pays no attention to it at all. By virtue of this, bullshit is a greater enemy of the truth than lies are.


And later adds:


Thus the production of bullshit is stimulated whenever a person’s obligations or opportunities to speak about some topic are more excessive than his knowledge of the facts that are relevant to that topic.


My reason for referencing Frankfurt is that bullshit seems to have become the main ingredient of the proliferating amount of soma in the current world. We are being swamped by an avalanche of distracting and entertaining bullshit on media that did not even exist when Postman and Frankfurt first made their arguments about the state of the world.


And the specific nature of bullshit is its power to castrate the truth.


So, to butcher the words from Chernobyl: "The real danger is that if we hear enough bullshit, then we will no longer recognise the truth at all."


Not an Orwellian Chernobyl but a Huxleyan one

America has no notable problem with free speech (though neither do many other western countries, even those, like the UK, who have no constitutional protection for it). Being able to say what you want is not the problem.


The truth will not be suppressed in the USA by Orwellian means. But it can be extinguished by too much bullshit if people lose their ability to recognise objective facts. This appears to be the state of a disturbing portion of the current US population. Trump has a lot to answer for.


Here is a trivial example to illustrate the point. I was engaged in some twitter banter and pointed out that Trump started his presidency with the lie that his inauguration crowd was bigger than Obama's. A Trump supporter claimed the photos were faked and provided a link to a CNN story where a photographer admitted some manipulation. But I read the link–supposedly refuting my point about the first big lie of the presidency–and what it actually claimed was that the official government photographers had cropped the photos to make Trump's crowds look bigger. Trump's claims are so powerful that believers will quote evidence that clearly refutes the claim in a belief that it validates the claim. Truly the ability to distinguish true from false has departed much discourse.


Whether we care about crowd size (or the size or body parts for that matter–hands, I'm talking about hands you filthy minded cynics) doesn't matter much. Who cares? But the same disregard for truth applied to public health is somewhat more worrying. Or to trust in democracy itself. All of which is a notable problem in the dying days of the Trump presidency.


Trump has elevated bullshit to a whole new and stunning level. In his early days he damned the mainstream media as "fake-news" when they refused to pander to his ego. He continued throughout his presidency to assert his own ego-driven world view against more objective analysis. His conversation with the Ukranian president was "perfect". Wearing masks as a public health measure was "a choice not a recommendation" and "I'm not going to be doing it". Miracle cures for covid were on the horizon (HCQ, injected bleach or "light"). The US response to the virus was "perfect" and no better job could have been done. He asserted his own intelligence agency's advice was wrong on Russian involvement in the 2016 election or in the major 2020 hack of many government and private systems, preferring Putin's judgement.


His supporters loved his bullshit. 


Neutral public health advice was undermined during the pandemic. Instead of rational scientific assessments of options, he politicised everything. Belief in miracle cures like HCQ was no longer a matter of calm assessment of objective medical trials but a signifier of political allegiance (and many of his supporters then claimed the opposition had politicised it). Mask wearing became a signal of government oppression not a sensible public health precaution. Social distancing became a sign of effeminate fear and mass political rallies and White House events without it became a signifier of masculine strength not a reckless way to spread the virus (even after the party announcing the Coney Barrett supreme court nomination probably became a super spreader event infecting Trump himself and a disturbing proportion of the senior staff present at it). 


His supporters still loved his bullshit.


Since the 2020 election his bullshit has got even more detached from reality. He claims he won by millions of votes (despite the certified counts showing him more than 7m behind). He claims he easily won the electoral college (despite the official count giving a result he described as a "landslide" when he won by the same margin in 2016). He continues to assert he won in 2020 despite losing every legal challenge everywhere. His legal team have made arguments so outrageous it is simply impossible to fathom how anyone took them seriously. In the Texas argument to the Supreme Court, for example, the one in a trillion chance of Biden winning in swing states assumed–when the statistical verbiage was stripped away–that voters voted as they did in 2016 and that batches of vote counts were perfect random samples from the voting population (a problem confounded when the statistician claimed there was no evidence this was not true in a response to legal challenges despite it being a clear and known consequence of counting votes by precinct and county where populations vary greatly). He claimed that voting machines had been hacked and had switched votes, despite verification by hand recounts of paper ballots that cannot be manipulated electronically. He sacked his own appointee responsible for scrutinising cyber-security for saying the election was secure. He lost his own–previously obsequiously loyal attorney general–for saying there was no evidence of fraud on a scale that could tip the result.


His supporters continued to trust him rather than, well, the objectively checkable facts. Worse, some intelligent political leaders decided to back his ludicrous claims (presumably out of the political calculation that pandering to his support would be good for them) despite knowing the claims were as strong as a wet sheet of toilet paper.


Fuck truth. 


Truth matters and Chernobyl is what you get when you ignore it

This brings me back to the point of all this. 


Trump, many of his voters, and far too many of his political backers seem to have no regard for the truth. They are not liars but–to use Frankfurt's definition–people who simply don't care about whether truth even exists. To use my modified quote from the Chernobyl TV series:


The real danger is that if we hear enough bullshit, then we will no longer recognise the truth at all.


Looking at the behaviour of many Trump-supporting voters and far too many political leaders it looks like the tsunami of bullshit has done its work and many no longer have any idea that there is such a thing as objective truth. As I write this a Trump-incited rally has invaded and disrupted the US Congress delaying the process of certifying the next president. And rapid opinion surveys suggest that nearly half of republicans think this is acceptable. Their conviction seems to be that a conspiracy robbed them of the true result (and, presumably that conspiracy now includes the Senate majority leader Mitch McConnell and the Vice President).


The USA has achieved the same situation as the pre-Chernobyl Soviet Union not by suppressing truth but by swamping it with bullshit. They have used soma rather than the "jackboot in the face" to do it. President Trump's tweets are the new centrifugal bumblepuppy. We are in the world of Huxley not Orwell. 


The result will be the same. There will be a catastrophe caused because too many people can not perceive the drops of truth in the ocean of bullshit. It took the Soviet system half a century of truth suppression to have a disaster so public they couldn't hide it (though they had plenty of others beforehand some of which killed millions). How long will it take for the USA's alternative approach to have the same result? And how many other countries in the west will go down the same path?




Tuesday, 22 December 2020

The final NHSE proposals for measuring A&E performance are still a shonky mess

It helps to understand how A&E works as a system before proposing how to change the public measurements used to manage it. There is little evidence that the latest NHSE proposals do. While the goals are good, the details are contradictory and inconsistent. And a great deal of the evidence used to justify the changes is either wrong or laughably inconsistent.


I've said much of this about previous versions of the proposals so, now I feel compelled to say it again, there might be swearing. And apologies in advance for the length of this post, I thought a thorough review was worth it.


A new consultation has just started on the long awaited proposals from NHS England on how to change how we measure A&E performance (it is part of this larger consultation: Transformation of urgent and emergency care: models of care and measurement).


I've had plenty to say about how metrics can be used to improve performance (see https://policyskeptic.blogspot.com/2020/02/nhs-performance-can-be-improved-by.html) and my specific reaction to the original proposal about replacing the 4hr target: https://policyskeptic.blogspot.com/2019/03/publishing-wider-metrics-about.html).


To be fair, there are some improvements in this version of the proposals but they don't fix the basic problems with the evidence and logic for the changes. Let's look at why that is.


The rationale and content of the new proposals


The rationale given for the new proposals is (my highlighting):

The intention is to enable a new national focus on measuring what is both important to the public, but also clinically meaningful. ... The CRS has concluded that these indicators are critical to understanding, and driving improvements in urgent and emergency care, and proposes a system-wide bundle of new measures that, taken collectively, offer a holistic view of performance through urgent and emergency care patient pathways. This bundle ... will enable both a provider and system-wide lens to assess and understand performance. The review findings show how these metrics will enable systems to focus on addressing what matters to patients and the clinicians delivering their care. 


There appear to be 10 key new metrics and 6 or 7 of those are relevant to A&E:

  • % of ambulance handovers inside 15mins

  • % initial assessments done inside 15mins

  • Mean time in A&E non-admitted patients

  • Mean time in A&E admitted patients

  • Clinically ready to proceed [I know what this is meant to measure but have no idea what the specific measurement is despite reading the explanation.]

  • Patients spending more than 12hr in A&E [presumably a count]

  • Critical time standards [again, though, the specifics are poorly explained]


This is a mix of good ideas, bad ideas and muddled definitions driven by good intentions. But, perhaps the worst idea is that they propose to completely stop using the 4hr metric. 

So, what is so wrong with these proposals? Let's have a look at some of the problems.


They obfuscate performance instead of clarifying it and make improvement harder to deliver.


The big problem here is the use of averages as performance metrics (the average wait time for different types of patient). These are good in just one context: retrospectively monitoring changes in performance over time. They are operationally meaningless for supporting staff to spot problems in real time, unlike the 4hr metric which is useful both in real time and in retrospective analysis of changing trends. 


And an average time metric obfuscates performance problems by failing to distinguish between situations where excessively long waits can be balanced by many short waits and a situation where the typical wait is reasonable. To illustrate consider the following:

  1. 100 patients get discharged in 1hr but another 35 wait 12hr: mean wait 3.8hr

  2. 135 patients get discharged in 3.8hr: mean wait 3.8hr

Very long waits are very bad. But the mean time metric totally fails to distinguish the two situations (the 4hr target *would* clearly show 1 as worse than 2). 


So those metrics fail to help operationally and fail to clarify performance problems. 


One possible excuse that has been used is that they do distinguish the following scenarios:

  1. 100 patients leave in in 3hr but another 35 take 4.1hr: mean wait 3.3hr

  2. 100 patients leave in 3hr but another 35 take 12hr: mean wait 5.3hr

These two scenarios have the same performance on the 4hr metric but 2 is clearly worse on the mean time metric. This is a good point but hardly a compelling reason to use only a mean time target as it clearly has exactly the same sort of flaws as the 4hr target when used alone.


And the proposals for new metrics have always been adamant that the 4hr metric must go: these are not supplements to the existing target, but substitutes for it.


They are much harder to understand than the simple current metric


Does anyone seriously think that the mean time measured separately for different types of patient will be easier for the public to understand? A count of the patients who have waited too long and a clear time limit that defines that wait is about the simplest metric that can be communicated to the public. What will they make of the new metrics? I don't have any idea and neither does anyone else as nobody has asked them. How will a member of the public judge their personal wait time against an average metric? "I'm sorry you waited for 8hr Ms Fisher, but our average wait is well within the target so there can't be a problem" is not a conversation I want to staff to be having with patients in the future.


Later the consultation also proposes to use some sort of weighted bundle of several very different metrics as a sort of simplified overall score. Again, this contradicts the goal of something that is easy to understand and act on. So not only are the individual metrics less clear but a new jumbled score will be introduced that will be hopelessly incomprehensible to anyone but its creators.


The report introduces a pathetically weak justification for a change based on a Healthwatch survey that claims the public don't understand the current 4hr target. I'm fairly sure the survey didn't ask the public to compare options for alternatives and I'm very sure that any future survey will find a much lower level of comprehension of targets based on mean performance. Any lack of understanding the 4hr target looks like a problem with NHS communication not an inherent problem with the specification of the metric, though that is not how the report wants you to interpret the HealthWatch result. I struggle to see how A&E staff will either comprehend, communicate or use the new targets, never mind the public.


The argument that the system has changes and therefore we need to abolish the 4hr target are weak


The claim is that the system has changed a lot since the 4hr target was introduced and, therefore "this single standard is no longer on its own driving the right improvement". The specific changes given as examples are:

  1. the introduction of specialist centres for stroke care, for example the reconfiguration of London and Manchester stroke services 

  2. the development of urgent treatment centres 

  3. the introduction of NHS 111 

  4. the creation of trauma centres, heart attack centres and acute stroke units 

  5. increased access and use of tests in emergency departments 

  6. the introduction of new standards for ambulance services 

  7. the increasing use of Same Day Emergency Care (SDEC) to avoid unnecessary overnight admissions to hospital. 


1 & 4 are irrelevant. Volumes in these specialist centres are not large compared to A&E attendance, they use their own clinical standards already (so don't need new ones) and they have little overall impact on major A&E metrics. 


2 could, in principle, be relevant if UTCs diverted patients who would otherwise attend a type 1 A&E. The remaining mix of patients in A&E would then skew towards those with more serious conditions. But there is no evidence that isolated UTCs divert anyone from A&E and those managed alongside A&Es are, basically, a form of patient streaming at the front door and do not imply a different performance metric is needed for a combined unit. 


3 is a WTF? I have no idea why the existence of NHS 111 makes any difference to how we should measure A&E performance unless you accept the insane techno optimist idea that booked appointments in A&E would make demand "more predictable" which should be a ludicrous idea to any who has ever done analysis of A&E data. A&E demand is very, very predictable: for any given department the weekly attendance at a given time of year, the busiest day of the week and the busiest time of the day are predictable to a few %. Bookable appointments offer no conceivable gain for performance (though patients might like them for other reasons).


5 offers no good rationale for different targets either (it might justify additional reporting of data already collected, though). The report argues that for sicker patients having more tests we need to have a different waiting time metric. But this is based on the flawed idea that the bottleneck in the patient flow depends on the time taken to do tests or to treat the patient. It has never been the bottleneck in any department where I have seen the detailed analysis. The sicker patients wait longer because the flow out of A&E is blocked (usually because free beds are hard to find). The report does promise to measure the "clinically ready to proceed" time which might highlight that problem. This is a good idea, but less radical than it sounds. The metric is–at least in principle–already accessible from the patient-level data already collected. This is why I'm fairly certain that waiting for treatment is never the bottleneck in the overall A&E wait: i've done the analysis on multiple steps in the patient journey. So reporting this metric routinely would be good especially if it disabuses policy makers about the dominant cause of long waits. But, yet again, this does not justify the abolition of the 4hr target in any way.


6 No, No, NO! The only reason why new A&E targets could be impacted by new ambulance targets is if there were gaming of the handover from ambulance to A&E. That was ruled out more than a decade ago by starting the A&E clock 15mins after ambulance arrival whether the patient had been transferred or not. 


7 There is good justification for more SDEC. But it is pure obfuscation to claim this demands new ways to measure A&E performance. Not least because it has been happening for a long time and some hospitals have already dealt with it in ways that are entirely compatible with the 4hr target metric. Sure, we should report something about the use of SDEC and that is a good additional metric for assessing departments. But it doesn't support the abolition of the 4hr target. Some departments have specialist units for SDEC and account for that activity by stopping the A&E clock when patients move into the unit while measuring the % of same-day discharges from it (this being the goal of such units). In this context there is no impact requiring a change to the 4hr metric and the additional metric could be added to public reports without any extra data collection effort.


On the positive side there are some good ideas about additional metrics that might support a holistic view. But none that suggest 4hr should die. My suspicion about the desperate desire to get rid of the 4hr metric is that this was driven by political pressure not from compelling clinical arguments. It is just too embarrassing for the government to be seen to keep failing to meet it.


The idea of adding extra metrics to give a more holistic view of performance is not new. It was first proposed in 2010 but never taken seriously by the incoming government (though many of those metrics can be found in obscure corners of the government web). Most of the possible metrics don't even require a lot of work. The big ones are, in principle, already available to any analyst who can access HES data or local hospital activity data. Indeed any competent A&E analyst should already be looking at them to support local improvement. For reasons known only to the god of bureaucracy that isn't how we collect the public data. Instead hospitals have to report the data via an entirely different route meaning that, to report the new targets they have to do extra work (a sane system would make sure the universally collected activity dataset was the basis for both the public and private reporting which would mean that new metrics required just a few extra equations to be written to automatically generate new reports for public consumption).


I should praise the best idea that the latest report has introduced. That is the measurement of 12hr waits. The report correctly recognises that there is never, ever the slightest justification for patients to wait >12hr in an A&E. The metric is already countable from HES data and should have been a public metric a decade ago. It would be the single most effective metric to highlight and correct the worst behaviours the report claims to want to fix. But the system has strongly resisted publishing the numbers until now. I was once fired from an analyst job for leaking the actual numbers (which I clearly should not have done: key lesson don't tweet when drunk). But they have been easily available and publication was strongly resisted precisely because they provided an extra insight into how bad A&E performance was. In my cynical moments I suspect that they are only willing to publish 12hr waits now because nobody will notice how bad they are in the blizzard of other new metrics being released.


So much for the key arguments that supposedly compel abolition of the 4hr metric and its replacement by other metrics.


The rationale is often based on a flawed analysis of what is currently happening and how the system works


But there is another problem with the proposals. While claiming the NHS needs metrics that drive improvement across the system of emergency care, the report betrays a lack of understanding of how the parts of the system actually interact. It is also confused about what the current trends are.


Different public targets for different types of patient is a flawed idea

Firstly the idea of setting separate public targets for different groups of patients is deeply flawed. The justification is that some seriously ill people need very rapid treatment. So specific targets are required for some clinical groups (actually this idea is not a bad one but every A&E should already be applying such targets internally). The proposal, however, is to have different public waiting targets for patients needing only simple treatments and those needing faster or more complex care. 


The public say to Healthwatch that the most sick should be prioritised and the report naively accepts this as a justification for allowing longer waits for "minors". This, and other comments in the report, betray a naive understanding of how A&E departments work and how queues work. The goal–getting fast treatment for some patient groups where urgent action is required–is a good one. The method–achieving this by slowing down treatment for minors–is as dumb as a box of spanners. One common issue in pre-4hr target A&Es was that this was exactly how they handled the queue of patients. Crudely speaking, those patients with minor cuts and bruises would be repeatedly bumped to the back of the queue when a heart attack appeared in the single queue. The consequence was that departments would fill up with minors who waited for a long time and whose treatment was frequently interrupted dramatically lowering the productivity of the staff. And. eventually, the department would be too full to allow fast treatment of the heart attacks and the minors. This is basic queueing theory. Unfortunately it appears that the experts who wrote the report don't know any queuing theory and didn't consult anyone who does. I can excuse the public for not getting this as it isn't something usually taught in schools or even widely in universities. But when the report uses the public's flawed understanding to justify its conclusions I can only assume either unforgivable ignorance or mendacity.


The problem with some patients needing faster treatment was solved in the early days of the 4hr target by recognising that long single queues were part of the problem and made waits worse for all patients. The solution was to split the queues and reserve enough capacity to guarantee speedy treatment for the majors while achieving efficient treatment for the minors at the same time. An efficient process for the minors keeps the overall queue small and leaves enough space for the department to cope with the lower volume of major injuries and problems that need guaranteed fast treatment. Faster treatment for minors actually enables faster treatment for majors: it is not a trade-off. Rapid assessment at the front door enabled departments to stream patients to the right queue. This was one of the key insights that helped A&Es meet the original 4hr target. 


The proposal to relax the target times for minors is entirely counterproductive and ignores those insights about how the system works. It might provide some justification for additional metrics that measure time for specific clinical groups, but provides no justification at all for abolishing the 4hr metric.


It is worth noting that individual A&E departments should have many more detailed metrics to track their performance internally. This has always been true even when the 4hr target was the dominant public target. But the specific changes to the public targets are likely to entrench mistaken views about how the overall system works and encourage bad trade-offs. 


Some arguments for change are based on a flawed view of system trends

Another naive argument is that attendance in A&E has grown unpredictably quickly and this is a major cause of crowding and slow flow. But when the document quotes numbers for A&E growth and volume, it quotes the numbers that include UTCs and their ilk. This is either naive or dishonest. UTCs, WICs and other type 3 departments were rare two decades ago but are common now. Most of the fast volume growth exists because many new units of this type were opened over that time period. The volume growth in major A&Es (type 1 units) is much lower and has been relatively stable year to year for three decades. And all those new units opening has had no discernible impact on that growth rate (though part of the original rationale was to divert minors away from major A&Es). While UTCs might fulfill a useful and popular need for patients, they have little interaction with major A&Es and should not be grouped together in public reporting (reporting a system performance number including the UTCs has the political advantage of diluting the clear signal of how badly the type 1 A&Es are doing in headlines but a diluted signal makes identifying the need for improvement harder). 


If there is a clear problem with volume, it is caused by a failure to match core hospital and A&E funding to the highly predictable long term rate of growth not because type 1 A&Es are overwhelmed by unpredictable demand.


Unpredictability of A&E demand isn't the problem

This last point and the other observation that the pattern of demand by hour, day and season is very predictable also highlights the absurdity of some of the proposed solutions to A&E crowding. One proposal is to have NHS 111 send booked appointments to A&E to make unpredictable demand easier to handle. But this only works if unpredictability is the problem, which it is not. The idea reeks of techno-utopianism where magic new technology somehow solves a big problem. 


It is also worth noting that there is another proposed use of direct booking by NHS 111 and that is for GP appointments (this is not part of this consultation). This is also a bad idea as booked slots are the enemy of GP responsiveness and insisting that GPs offer them reduces the speed of response and the overall flexibility of a good GP service. Both booked A&E appointments and GP appointments look like an example of "cool technology will fix all problems" techno-optimistic naivety.


So what?

The report claims its new proposals for targets are needed because we need to have clearer targets that are easier to understand and are more focussed on driving improvement. What it actually proposes are harder to understand targets that muddy the signals required to drive improvement. 


The report claims that the 4hr metric must go. But the arguments used to justify this are so weak that cynics will be compelled to conclude that the real motivation is political embarrassment. If we keep publishing performance using the 4hr standard it will be very obvious that the system is failing: a mixed bunch of new obscure targets will obfuscate that failure and reduce demand for change and investment in the resources that might improve it. And the proposals admit the need for new metrics to provide multiple perspectives on performance, but reject the simpler idea of supplementing the current target with additional metrics (for example, why has england never published a 12hr wait metric before?)


Some who take a top-down perspective on NHS economics might argue that improvement is impossible in the current climate where the NHS is focussed on priorities other than increasing bed numbers. I disagree. Yes, this isn't a big part of the NHS plan. But the effective use of the right data can be a powerful lever for improvement even when resources are tight. But replacing a good target with a multitude of bad ones will not drive improvement. The NHS would do far better to encourage more investment in analysis of the rich data it already has than in inventing new public metrics that won't help.