Pages

Showing posts with label modelling. Show all posts
Showing posts with label modelling. Show all posts

Thursday, 7 April 2022

We are having the wrong debate about modelling the NHS workforce

It sounds intuitive that having a good model of the NHS workforce would be useful for solving many observable problems in the current NHS. But there are reasons to assume any such model would not be as useful as expected and could even be harmful. More importantly, the starting point suggested for the model is wrong in multiple ways that almost certainly guarantees the model would fail to solve the real problems.


When I first heard that a coalition of Labour MPs and the ex-NHS SoS Jeremy Hunt were proposing an amendment to the Health Bill to insist on regular publication of a workforce model for the NHS, I thought the idea was a good one. Better, transparent information about the staffing needs of the future NHS sound like a useful idea to test against government policy. 


But, when I reflected on some of the issues I have seen in the past when Strategic Health Authorities did workforce plans I started to have doubts. Then, after some further conversations and cogitation, those doubts grew. Considered alongside my analysis of what the biggest challenges are for the NHS (in short, the front line workforces is far from the biggest problem), my skepticism strengthened. 


So, I'm going to argue that, while a good workforce plan might help, the one we are likely to get is likely to be somewhere between useless and harmful. It is starting in the wrong place, has unclear goals that will likely make it far less useful than expected and has some risk of making things worse. 


That's a big claim. Let me walk through the potential issues I see step by step. 


Where the workforce model might go wrong


It doesn't start by considering productivity


The starting assumption in the debate is the almost universal belief that the problem is a shortage of front line staff. Commentators observe busy A&Es, overwhelmed GPs risking staff burnout, hospitals where waiting lists are growing not falling, and leap to the conclusion that the only way to address these is more medical staff.


But the problem framed this way distracts from any analysis that concludes anything other than "more staff" can influence the amount of work done. Concluding that only more staff matters rules out many known interventions that should be part of the debate, especially those that improve productivity..


Here is a simple example. NHS A&Es are currently catastrophically crowded and patients are getting treatment so slow it is killing them. But we know from analysis that has existed since the 4hr target was introduced that the number of A&E doctors has very little influence on the speed (for a fun review of the evidence on this and how little appetite the system has to listen to it read this BMJ piece and the replies). The dominant cause of long waits in the last decade has been slow access to free beds for admitted patients. The problem isn't even inside the A&E, so adding more A&E staff won't fix it. To be fair, the workload needs in A&E do increase sharply with the length of the queue. But they don't help to make the queue shorter. So any workforce model that ignores the external factors causing the queue will end up recommending far more A&E staff than needed if the external bottleneck causing the queue is ever solved.


Another example shows that not thinking more widely about the mix of staff required is a major problem in planning. A Royal College of Surgeons blog reported in 2017 that the productivity of surgeons had declined sharply as their numbers rose because numbers of support staff and nurses had not risen. It reports "Between 2010 and 2016, consultant numbers rose by 22%, compared to just 1% for nurses and 2% for all staff."  It Also pointed out that "... consultants in hospitals that invested more in infrastructure and building … are more productive". (it is worth reading the original analysis by the Health Foundation as well for much more detail). So, is a workforce model is built to forecast the number of surgeons required, but ignores the number of nurses and support staff or the capital required to create a productive theatre, it will vastly overestimate the number needed.


The point is that productivity depends on the mix of staff and other factors like equipment. Driving higher output across the NHS requires a good understanding of where the bottlenecks to productivity are so the right mix of capital and people can be deployed. Adding more of the most visible front line staff is often not that effective. But it is what a "more resources to the front line" workforce model is likely to achieve.


It is unclear what decisions a long term workforce model is intended to support


The only point of any model is to support better decisions. If you are vague about what decisions, then the model is likely to be unhelpful and even misleading. 


So what decisions could a long term workforce model support? So far the discussion has tended to focus on the need for a model to inform the NHS about its workforce need in 5, 10 or 15 years. 


What decisions could such a model influence? Not many. If we know we need more nurses in 5 years time the NHS might just have enough time to increase the number of training places to increase the numbers qualifying in 5 years time. But it is unclear whether increasing the number of doctors in training right now would lead to higher available numbers in a decade's time. 


If the NHS is short of particular skills right now, it is unclear how a long term model can possibly help. What the system needs most urgently is some idea of what the options are today


If the model focuses–as much of the discussion about it has–on front line staffing need, then it also misses critical groups that contribute to the productivity of the front line (see the section on productivity for why this is important). The question that desperately needs an answer is what different mix of staff, equipment, buildings and new clinical processes would give the biggest increase in the number and quality of treatments delivered. Many of those questions are not workforce questions at all and even a workforce-only model needs a good understanding of how different staff interact to make the front line more productive. That understanding will only come from a significant piece of careful analysis that doesn't seem to exist. 


What the NHS needs right now is that analysis. It needs to know where the bottlenecks to higher activity are. It needs to know what mix of capital spending, front line staff, support staff and managers would lead to the largest improvement. Without that a workforce plan will be about as useful as a one-legged trapeze artist with an itchy bum.


A workforce model designed to tell the system how many staff it needs to put into training now will get the wrong answer because it ignores all the interactions and other factors that matter and, even if that wasn't true, could not influence any decision that will have an effect for 5-10 years at best.


Major factors that influence the workforce today are likely not part of the model


The NHS has a workforce problem right now. And many of the factors causing problems are not relevant to the long term supply of qualified staff. Or anything else likely to appear in the proposed workforce model.


Right now the biggest factors influencing staffing gaps are recruitment problems and high turnover (plus illness, if temporary pandemic-specific problems count). These problems don't just affect the front line staff groups but are common in the other groups where a lack of staff has a lot of leverage over front line productivity.


There are many causes of recruitment problems and high turnover. In some groups NHS pay is inadequate compared to other jobs the same staff can do. This is a big issue for nurses but a huge problem for support staff like data scientists. It is also a bigger problem in some geographies like London where the cost of living is much higher and there are more alternative well-paid jobs. But the NHS finds it almost impossible to flex salaries to retain the people it needs both because the pay scales are national and because some groups are vastly undervalued in AfC grading compared to the market.


Working conditions are also a huge factor for recruitment and turnover. If the space is badly adapted to the work being done (~14% of buildings predate the NHS!) then the environment will be poor. Badly maintained buildings add to this (the maintenance backlog is about £10bn). Old, shonky equipment is slower and harder to use than modern equipment. IT systems are often slow and not seamlessly integrated so staff waste time waiting to log on or logging in to a dozen separate systems to complete a clinical task. Front line staff end up spending too much time doing tasks that should be done by support staff or managers (where staffing levels have been cut to "put more staff on the front line") instead of caring for patients. 


Too much of the people management in the NHS is bad. Staff are treated badly and insensitively by managers but also by senior doctors and nurses (the Ockenden report didn't just blame "staff shortages", it clearly blamed senior staff of all professions for ignoring clear signals about problems and even suppressing whistleblowers). 


Very few, if any, of the factors that discourage recruitment and drive high turnover are part of any proposed workforce model.


So the model won't tell the NHS whether a big increase in capital spending, creating better buildings, equipment and IT systems, would yield rapid gains in a better working environment. Nor will it conclude that recruiting more support staff to enable the front line to focus on treating, rather than admin paperwork, would improve their job satisfaction. And nobody in NHSE would allow the model to conclude that salary flexibility might yield immediate benefits in both lower turnover and higher recruitment rates.


That means that the model is likely to have nothing to say about the major factors that could impact the workforce any time in the next 5 years. What was the point of it again?


A long term workforce model risks fossilising current mistakes and practices


More than a decade ago I was part of a team auditing some workforce models for SHAs (when they still existed). One of the problems the team spotted was that complex models with very large amounts of detail tended to be very hard to audit properly and often contained errors in their code. That's bad when you rely on their outputs. 


But that complexity also had a side effect that is, though not an error, worse: they fossilised current assumptions about the mix of the workforce. In particular, they made assumptions about the need for very small specialist subgroups of staff (humorously like the number of orthopaedic surgeons specialising in only left hands). The problem is that practices often change faster than the model. So, if the model spits out the demand for some small highly specialised group in a decade's time, it may have been overtaken by major changes in the way that specialty works. Once upon a time, for example, most cataract operations were done under general anaesthetic. Then it became obvious that local anaesthesia was faster and safer and the mix of activity changed rapidly. Any model built before that change was obvious would forecast a completely incorrect mix of staff or number of staff.


When you build complicated models there is always a big risk that the assumptions in the model persist long after the change as many models are even harder to update than clinical practice. Or the modellers just don't notice the changes and the NHS keeps relying on their model as the users of the model don't understand it well enough to understand the assumptions it makes.


In another case I studied a model built for NICE on staffing in A&E departments (see my commentary). The original report gained a lot of credibility when NHSE allegedly suppressed it. But it was leaked alongside the full documentation on a simulation model built by external consultants that had been a major evidence source for their recommendations. I read the documentation. I wept. The assumptions about how an A&E worked had almost no relationship to reality and ignored very clear, well-known, data about actual performance. It looked like it had been built by someone who had never visited a real A&E or mapped a real world operational process. I suspect that most people who read the report didn't understand the model or that it was a major part of the evidence behind the recommendations. But bad models make bad recommendations. Worse, complex models make those mistakes harder to spot.


Even when a model works it may fail to influence the right decisions


The debate about the need for an NHS workforce model seems to assume that models have a magical ability to change the decisions people make. Decision scientists know this isn't true. 


Many analyses and models completely fail to influence actual decisions even when they are reliable and the data behind them is correct. For example, the NHS has reported on the hospital maintenance backlog for years, including an estimate of the need for urgent action to limit the immediate risk to patients. Yet decision makers have repeatedly chosen to spend far too little on capital (the budget has been about half that of peer health systems for most of the last two decades). And the high risk maintenance budget backlog grows every year. Maybe bad decision making is very resistant to modelling or analytical data.


Models work best not when they give a highly specific and precise answer but when they help decision makers to understand the core issues behind the decisions they have to make. A complex and detailed model of the NHS front line workforce is unlikely to achieve this. Not least because, if its focus is just the front line, it will fail to help decision makers to understand the tradeoffs involved in between different decisions they could make today.


What if, for example, a small increment in the number of managers greatly improved the productivity and quality of the work done in hospitals? How would that choice interact with the future needs for front line staff? We already know that more managers do have significant effects (see this summary from the NHS Confederation) but the idea that a workforce model should think about them is entirely absent from the current discussion on workforce modelling.


Also missing in the discussion on workforce modelling is any hint of how capital spending on better buildings, equipment or IT could contribute to productivity. But decision makers have to make tradeoffs today about how to split the budget among front line staff, support staff (including managers) and capital spending. And for most of the last decade that choice has skewed towards the front line leaving the NHS with a chronic deficit of support staff, managers and adequate modern buildings and IT (for some data see my analysis here). Those choices have led to declining front line productivity, a much worse working environment and, arguably, contributed to recruitment problems and higher staff turnover. A model focussed just on the long term needs for front line workforce numbers will encourage continued neglect of those other factors which directly impacts the immediate workforce.




Conclusion: a long term workforce model is a distraction not a solution


It seems obvious that the NHS has a serious shortage of front line staff. But that observation is very deceptive. It is a symptom of widespread problems of productivity and blocked flow of patients. As with many medical conditions, there is a strong temptation to treat the symptom and assume that this cures the disease. In NHS language, focussing on the front line workforce assumes that investment in the front line workforce cures the problem. But, like trying to cure headaches caused by a brain tumour with stronger painkillers, treating the symptom doesn't solve the underlying problem.


A model that assumes that frontline overload is cured purely by adding more staff distracts attention from all the other factors causing overload of front line staff. So attention will be distracted from inadequate buildings, obsolete equipment, slow and badly designed IT, admin overload caused by a lack of support staff and managers, and poorly designed clinical pathways. And "more staff" doesn't fix problems where the issue is having the wrong mix of staff. 


Some of those problems might be fixed, in principle. We could, perhaps, build a model that takes into account the staff mix, not just the overall number of staff. It could even highlight the areas where extra staff would most improve overall productivity (eg, to fix A&E overload, invest in staff who can improve the flow through beds). Unfortunately the first step would be to develop an analysis of how the whole system fits together and therefore identify which incremental investments would most improve the productivity of the system. There is no such analysis, though there are plenty of hints that workforce isn't the biggest problem. 


The workforce model as currently discussed seems likely to further distract the NHS from other major problems. A focus on the front line risks distracting from even bigger staff shortages behind the front line. And from a long term neglect of capital spending (leading to major issues with buildings, equipment, and IT). Furthermore there is a big risk that a long term workforce model could ignore the short term decisions that might help the immediate problems with workforce and could encourage the NHS to build in false assumptions that fossilise bad current choices.


That's a lot of risks for an unclear outcome. Whatever the intuitive attraction of a transparent model of workforce needs, it is far from obvious that the NHS would get anything useful.


We need a more informed debate about what holds back NHS productivity, not a model focussed on the front line workforce.



Friday, 8 April 2016

The debate on BREXIT illustrates how little evidence determines major public policy decisions

People don't decide their position on BREXIT by looking at the numbers: they choose their position and then seek numbers to justify it. Most of the numbers quoted on either side can't be trusted. This is corrupting public discourse by damaging the credibility of numbers that do tell a clear story.


The ferocity of debate in politics is often inversely proportional to the amount of actual hard evidence available on the topic. Whether the UK should stay in the EU is a typical example. Many decide they don't like the loss of freedom required when you are a member of the club and call for exit; others, like me, decide that the compromises required and the bureaucratic overhang of membership are worth it because the gains from collaboration are worth it. Some choices are more atavistic: many presume (probably incorrectly) that uncontrolled immigration is caused by EU membership (and they also believe the populist myth that immigrants are the cause of many other problems in society). Neither side reaches their conclusion because they have done some calculations: the conclusions come from deep emotional value choices not statistics. But the debate pretends otherwise and dredges up volumes of statistics to confirm the emotionally reached decision.


The numbers thrown around in the debate are proxies for that emotional choice. Most people can't admit that they didn't reason their way to their choice and they seek what look like the rational arguments that got them there. But this process suffers from all of the cognitive biases that so beset much thinking (especially confirmation bias where people are more likely to believe numbers supporting the position they already hold). People seek the numbers that confirm their side of the debate.

Nothing illustrates this better than the meme of European bureaucracy. Even EU supporters often complain about bureaucracy and many even quoted this tweet as an example:

EU cabbage bureaucracy tweet.jpg

Luckily for rational thinkers the BBC's More or Less programme (see this BBC article) decided to test the assertion. Their conclusion: there are no EU regulations specifically about cabbage sales and the 26,911 number originated in a complaint in the post-war USA about government regulations (and was probably mythical even then). Amazingly the same number of words has been repeated by many other objectors to bureaucracy without ever being validated in even the most cursory way against real documents.

The meme of EU bureaucracy is so strong that even pro-EU campaigners didn't think to challenge the basic facts in the assertion.

My main point, though, is about the harder economic numbers bandied about by the two sides and the extent to which they corrupt public discourse when we are debating significant issues. Both sides in the debate, for example, agree that leaving the EU would be disruptive. I agree: that is one of the few things that is clear. But there is a lot of disagreement about how disruptive leaving would be and even more on the economic benefits of staying or leaving. The leave campaign tend to assume that new trade treaties would be easy to negotiate quickly (not that there is any evidence to support this) and they assume that the benefits of an independent UK would be large when freed from the dead hand of EU regulation. The stay camp assume that renegotiation is a long and costly process and that many current jobs in the UK that depend on EU trade would be lost.

Both sides quote specific estimates for the economic benefit of their position. This is where the problem comes. Although specific numbers are quoted, the estimates have no credibility. John Kay, the economist and FT columnist addressed some of the problems with these kinds of estimates in a column in 2011. His concerns, though originally made about the economic rationale for big government projects, apply equally to any analysis of the case for staying or leaving:

...Because so many inputs to the analysis are invented, they can be chosen with a view to the desired result...

...The only information exercises such as these convey is the limits of the imagination of the people who have undertaken them….

...Yet the mistaken belief persists that these procedures provide an objective basis for decision making...

...We do great damage by claiming to know things that are not known, by asserting certainty in the face of uncertainty and ambiguity, and by attaching a veneer of rationality to decisions that have in fact been made on other, rarely articulated, grounds. The paradoxical result is all too obvious. The public sector and large bureaucratic organisations appear as paragons of good decision making process and exemplars of bad decisions.

I would add an extra concern. The specific nature of the forecasts made by each side damage the credibility of almost any analysis intended to illuminate a major public decision. This is bad because, for example, when we have to decide how much to spend to mitigate global warming, the numbers we are shown will have no credibility and we will, most likely, make poor decisions as a result. Sometimes the numbers do point one way on a major decision. If we pollute public discourse with spurious, over-precise analysis, we will undermine our ability to make good decisions when the evidence is clear.

When deciding whether to stay in the EU or to invest in major infrastructure projects we can never have precise estimates of the costs and benefits. We should not deny the uncertainty by pretending that we do. We can't decide on whether to stay in the EU by purely objective analysis: the uncertainties are too large. We should admit that our choice is based on values not numbers. I believe that international cooperation is a better way to conduct our affairs than standing independent and alone, despite the cost in bureaucracy. I don't pretend I can prove it with economic models.

But, if you care about honesty in public debate, there is something you can do whatever position you hold on the EU. Donate to Full Fact's campaign to fact-check the debate. They don't care how you vote but they do care whether the debate is conducted honestly.

Thursday, 4 February 2016

The NICE guidance on safe staffing in A&E adds nothing to our understanding of how to run a safe A&E department



The controversial NICE guidance on staffing in A&E isn't worth arguing about. The evidence base is almost nonexistent and disagrees with better, older analysis. The NHS should ignore its recommendations and focus on gathering better evidence.


When the HSJ prompted the release of the NICE guidance about safe staffing in A&E I thought we might see something interesting. Then I read it and changed my mind. If anything the analysis sets back our understanding of how to run a safe A&E department. In fact it stands as a case study in how not to do useful analysis of an important operational issue for the NHS. Here is why I reached that conclusion.


What NICE did and didn't do

The NICE guidance is based on three sources of evidence: expert judgement; a literature survey; and an economic modelling study. The documents describing these sources of evidence are now available either from the NICE website (the economic model) or the HSJ.

What NICE didn't do is to gather systematic evidence from actual A&E departments in the UK either about staffing or performance (the model used limited evidence from a handful of departments and supplemented this with some average performance evidence from SITREPS and HES data).


What's wrong with the evidence

The NICE review itself sums up some of what is wrong with the literature evidence. Two problems stand out from their own summary: almost none of the evidence relates to the UK and there is very little high-quality evidence to start with.

In addition to this their evidence specifically excluded evidence relating to certain important practices that are common in English A&E departments. A critical example is the exclusion of evidence relating to Emergency Nurse Practitioners (ENPs) and related specialists. This seems to have been a choice so that the recommendations could be focussed on the general level of nurse staffing.

NICE commissioned a simulation model to help clarify some of the relationships that were simply missing from the actual literature evidence. This simulation forms the only significant basis for the actual recommendations (the literature evidence is simply too flimsy and contradictory to support any solid recommendations).

The trouble is that the simulation model is itself deeply flawed. So flawed that it is hard to take its recommendations seriously. It makes assumptions that were known to be naive a decade ago, some that directly contradict common practices in most actual A&E departments and produces results which disagree with actual observations about both staffing and performance. These flaws deserve a whole section to themselves.

Simulation modelling is just a way to hide the link between bad assumptions and your recommendations

[Actually I don't really mean that. Simulation modelling is an effective tool in the right hands and when the right assumptions are made. When a system is well understood but its performance is not it can provide valuable insight into how to improve. But the NICE model shows a failure to understand how A&E works and, therefore, cannot say anything useful about performance.]

The model used by NICE embeds false assumptions about how A&E operates. This is a critical failure in such an important model but NICE didn't pay much for it so perhaps that's all we can expect. Whatever the reason the assumptions are sufficiently bad that the output of the model can tell us nothing useful about staffing in a real A&E department.

Here are three examples of where the model makes really unrealistic assumptions and fails to represent the reality in A&E.

The model assumes a single process for treating patients.
This means that the model assumes that patients with single minor conditions are treated in the same way and by the same people as patients with complex or multiple problems or injuries. Real A&Es don't do that. One of the major innovations that led to much faster A&E treatment times in the early 2000s was the introduction of streaming for different types of patient. The idea of "see and treat" for minors was a major innovation that recognised that many patients don't need multiple investigations or multi-skilled teams to treat them. So many A&Es designed much simpler processes which cut out multiple stages of assessment, investigation or treatment. The process is staffed by people fully qualified to both assess and treat minor injuries or conditions. Patients get assessed and, if they don't have anything complex wrong with them, they get treated immediately, often by a specially qualified nurse. This is fast and efficient. It reduces the number of staff required (by eliminating unnecessary steps in the process for the majority of patients) and speeds the treatment. It leaves far fewer patients waiting around and clogging up the waiting room, thereby reducing crowding (which is good for staff and other patients).

By ignoring this major innovation, the NICE model becomes a hypothetical model of how an A&E department might operate if nobody ever had any good ideas about how to organise one effectively. By modelling something which mostly doesn't exist the model tells us nothing useful about staffing or performance in real A&E departments in England.

The model assumes the patients with more severe injuries go to the front of the queue
This sounds reasonable. But coupled with the previous assumption that all patients get treated in the same process it turns out to be both unrealistic and bad for patients.

It is unrealistic because it generates the output where the majors get treated faster than the minors (that what it assumes should be the process so, of course, that is the output the model generates). This is the opposite of what the data actually shows. In reality the majors--especially the ones who need to be admitted--have the longest waiting times. In well functioning departments the majority of minors are treated in less than 90 minutes, but it isn't uncommon for patients requiring admission to have an average waiting time of 4hrs or more. Moreover there is good evidence about why they wait and it isn't, mostly, caused inside the A&E department at all but by the failure of most hospitals to manage their beds effectively (see this Monitor report on the causes of A&E delays). The model doesn't consider these delays at all.

Streaming of patients into separate processes was developed because a single process is bad for all patients when there is a mix of different patient needs. A single process is wasteful; it creates unnecessary delays for minors; and it uses more staff time for no benefit at all to the majority of patients. Streaming minors into a separate efficient process frees up staff time for the more complex needs of majors and allows the separate process of treating them to operate quickly without interfering with the process for treating minors. Having two processes achieves the result of rapid initial treatment for majors without having to bump the minors to the back of the queue.

By ignoring streaming and modelling a different treatment process that no longer exists, the model fails to address anything useful in real-world A&E staffing or performance.

The model assumes that staff are all much the same
The focus of the model is to understand whether nurse staffing affects performance so it assumes that there are few differences among nursing grade staff and ignores issues with doctor staffing. Again the assumptions made ignore the reality of how A&E departments work.

There are two things that are well known by A&E experts that relate staffing to performance. One is that, when you stream minors to a "see and treat" process, you can use experienced nurses to deliver a lot of the treatment. These specialist staff (called advanced or Emergency Nurse Practitioners--ANPs or ENPs) are dedicated to the stream dealing with minors and allow fast treatment to be delivered efficiently for patients who don't have complex problems. Both the NICE model and the evidence review explicitly exclude anything relating to these specialists. The other staffing issue is that senior medics "on the shop floor" improve performance everywhere, probably because they can make fast confident decisions for edge-case patients where more junior staff would dither or make poor judgement calls. This is also ignored in the NICE evidence and model.

In summary: modelling the wrong thing won't provide any useful insights
I could go on but I won't. The key point here is that if you develop a model that isn't based on the real world you won't get any useful insights about the real world. A Lego model of the Empire State Building won't tell you about the structural integrity of the real Empire State Building. If the engineers used a Lego model for this purpose you would be well advised to stay out of New York.

So NICE have created a model that is uninformed by real world observations about how A&E actually operates; it ignores observations about real A&E departments are staffed; it doesn't have any inputs about how they actually perform; it ignores observations about where problems exist and models a process that doesn't consider the biggest problem (finding beds). Why does anyone think its conclusions are useful?

What NICE should have done

Given the admitted lack of evidence about real A&E departments in England what NICE should have done is to look for useful evidence rather than waste time on summarising poor quality analysis of irrelevant systems in other countries. There are more than 150 major A&Es in England and their performance is measured both in public SITREP data and in less public but more detailed HES data. Most of these departments should have some idea of their staffing profiles and rosters. Putting those two sets of observations together would allow a rich set of "experiments" to be done by comparing the departments to each other. It might take more effort (and actual statistical skill as opposed to modeling or literature review skill). But the results would tell us about the system we actually have.

NICE did none of this.

What is worse, the exercise has been done before and nobody at NICE, it seems, noticed.

When the Audit Commission existed and still did some work on hospital performance they had a programme called the Acute Hospital Portfolio Review. When the programme reviewed A&E it looked at staffing and performance on a range of clinical metrics including speed but also including quality of care. In other words, they did exactly what NICE didn't. The last of their reports that I know of is preserved here (pdf download).

The reports reached some startling and unexpected conclusions about A&E staffing which were credible because they were based on extensive real evidence on actual English A&E departments not on models or academic speculation. Here are two conclusions (with my emphasis):
Common sense would suggest that a large part of the improvements in times spent in A&E departments since 2000 has been due to the increases in staff. However, when comparisons are made at individual department level, there is no association between relative increases in staff and improvement in times spent in A&E.

for comparability, staffing levels need to be expressed as a ratio between actual staff numbers and the numbers of annual attendances (a reasonable measure of the size of a department). When expressed in this way, there is no relationship between times spent in A&E and staffing levels. Tightly staffed departments perform as well as generously staffed departments. This is consistent with the findings in the 2000 review.
Staffing in A&E has improved significantly since the last of these reports was written.

I suspect that the detailed data behind these conclusions has been lost with the abolition of the Audit Commision and its successor. But I know that the evidence was comprehensive and solid over several successive periods of data collection.

If you want to produce guidance about staffing that disagrees with their surprising conclusions then you need to generate some better evidence. Nothing in the NICE recommendations does that.

We also have other recent analysis that shows the biggest problem in A&E performance is nothing to do with A&E staffing but is about coordinating the A&E demand for beds with the flow through the beds in the rest of the hospital. No amount of extra staffing in A&E will help that. So not only is the evidence behind the NICE staffing recommendations as weak as wet toilet roll, it completely fails to address the biggest actual problem in our A&Es.

The controversy over the non-publication of the work has given it a credibility it doesn't deserve. The right response would have been to publish it and ignore it as it has nothing credible to say.