A paper landed in the American Economic Review this month that is going to get quoted badly by almost everyone who quotes it. Physician groups will pull three numbers out of it. Nursing organizations will pull a different three. Both sets are in there, which is exactly why the paper is worth an hour of your attention instead of a headline.
David Chan and Yiqun Chen looked at 1.1 million emergency department visits across 44 Veterans Health Administration sites, involving 156 nurse practitioners and 1,348 physicians (1)(2). What makes it different from the usual scope-of-practice study is the design. VA provider schedules are set months in advance. Patients show up when they show up. That mismatch means the question of whether you got an NP or a physician on a given night was close to a coin flip rather than a reflection of how sick you looked, and it lets the authors claim causation instead of the correlation that plagues most of this literature.
What the numbers say
Patients seen by NPs had emergency stays 11 percent longer and cost about 7 percent more, roughly 66 dollars per visit (1)(3). Thirty-day preventable hospitalizations ran 20 percent higher. The AMA ran the arithmetic forward and estimated that routing a quarter of VA emergency patients to NPs adds about 129 million dollars a year net, after accounting for the salary difference between the two groups (3).
Thirty-day mortality showed no statistically significant difference (2).
Hold onto both of those. People are going to publish articles this fall that mention one and not the other.
The number everyone is going to skip
Here is the finding I think actually matters, and it is the one I expect to see least in the press coverage. It needs a slow walk, because it is the part that gets garbled every time.
Everything in the section above compares two averages. Add up all 156 NPs and take the mean. Do the same for the 1,348 physicians. Compare the two. Physicians win that comparison, and I am not waving that away.
Now throw the other profession out and look at physicians alone. We are not all the same. Some of us order a great deal of testing and some order very little. Look at length of stay instead, or at thirty-day bouncebacks, and the same wide scatter turns up. The NP group has an equally wide scatter inside it.
What Chan and Chen found is that the spread inside each profession is larger than the distance between the two averages. The gap between a low-resource physician and a high-resource physician is wider than the gap between the typical physician and the typical NP.
Put those two facts together and the distributions overlap most of the way. The strongest NPs sit well inside the physician range. The physicians at the expensive end sit well inside the NP range.
So the authors ran the obvious test. Pull one NP at random. Pull one physician at random. Compare what each of them actually did, and repeat that many times over. The NP is the better performer in 38 out of every 100 draws (1)(2).
That is not a rounding error, and the number is worth calibrating against the two ends it could have landed on. If the professions were genuinely separate tiers, with the weakest physician still ahead of the strongest NP, you would expect something near zero. If they were interchangeable you would expect 50. Thirty-eight sits far closer to interchangeable than to separate.
Two things it does not say. It does not say NPs are 38 percent as good. It does not say that 38 percent of NPs outperform physicians across the board. It describes random one-to-one matchups and nothing wider than that.
And physicians still take 62 of those 100 draws. The average difference did not evaporate. Both of those are true at the same time, and holding both at once is the whole trick with this paper.
The gap also moved around depending on what walked in the door. For the least complicated cases, the extra cost attached to NP care fell by roughly 80 percent compared with the average case (2). It narrowed further as NPs accumulated years, and narrowed again as they accumulated reps with a specific condition. Chan framed the takeaway as a question of “which patients they should see” rather than whether NPs should practice at all, and I think that is the honest reading of his own data.
The mechanism looks like uncertainty, not carelessness. NPs ordered more diagnostic testing and more specialist consults. For sepsis, stroke, and heart failure they were substantially more likely to admit. They wrote fewer opioid prescriptions and more antibiotic prescriptions (2). Read that list as a set and a pattern shows up: when the picture was ambiguous, the threshold to spend a resource dropped. Anyone who has been six months out of residency recognizes that behavior in themselves, because we all did it, and the thing that fixed it was not a different diploma but two thousand more patients.
An aside, because the word keeps getting misused. “Productivity” here is an economics term. It means outputs relative to inputs consumed, not effort expended and not how hard someone works. Nothing in this paper says NPs work less hard. It says a given clinical result cost more to produce.
Where I think it is weakest
One health system. One care setting. One hundred fifty-six NPs, against a national workforce of more than 461,000 licensed NPs (4)(5). The AANP’s objection that you cannot generalize from that sample to every emergency department in the country is fair, and I would make the same objection if the finding had gone the other way.
The VA population is also not America. Older, more male, more comorbidity, and enrolled in an integrated system with a shared record. The VA also grants full practice authority, which means these NPs were working without the collaborative arrangement most of my colleagues in private systems actually have. What the paper cannot see is the physician who glanced at a chart, said one sentence in passing, and quietly changed a plan. That interaction leaves no data trail and it happens constantly.
None of that makes the effect estimates wrong. It makes them local. An 11 percent length-of-stay difference in a VA emergency department is a measurement of that department, and treating it as a national verdict on a profession is a category error.
What this means for a virtual urgent care visit
Most of my work is telemedicine, so this is the setting I thought about first. Virtual urgent care now runs largely on nurse practitioners and physician assistants, and the mechanism Chan and Chen identified does not carry over to a video call cleanly. I cannot order a CBC in the middle of an encounter. I cannot walk anyone down the hall for imaging. When uncertainty rises the levers available are prescribe empirically or send the patient somewhere with hands, which happen to be the same two the study found NPs pulling more often (2).
There turned out to be more to say about that than belongs inside a piece about an economics paper. What the payment and rating structures do to the decision. Why recognizing the sick patient matters more in this setting than almost any other. What happens when the platform cannot order a test at all. I put all of it in its own article: Spotting the Sick One: The Only Decision That Really Matters in Virtual Urgent Care.
Virtual weight management is a different problem, and a more forgiving one.
Weight management is a different animal, and I think the gap mostly closes
Obesity medicine breaks almost every condition that produced the ED result. There is no undifferentiated chest pain arriving at 2 a.m. The diagnosis is usually made before the visit starts. Care is longitudinal, protocol-heavy, and forgiving of a decision revisited in four weeks. The study’s own results predict a smaller gap here, because it found the difference shrinking by about 80 percent on the least complex cases and shrinking again with condition-specific experience (2). An NP who has titrated a thousand patients through semaglutide dose escalation has more relevant pattern recognition than a physician who has titrated forty.
That said, the complexity in this field is real. It just shows up in different places than people expect. Sorting expected GLP-1 nausea from something needing imaging. Recognizing that a patient on 30 units of basal insulin and a sulfonylurea will need those doses coming down as the weight comes off, before the hypoglycemia arrives rather than after. Pancreatitis history. Family history of medullary thyroid carcinoma. Restrictive eating patterns that look like excellent adherence on a video call and are not. Secondary causes, and the long list of psychiatric medications that drive weight gain and never get revisited.
So the risk in virtual weight management is not the credential on the screen. It is the eight-minute refill visit, whoever is running it. A physician doing rushed protocol care and an NP doing thorough protocol care are not close, and I would put my patients with the second one.
The tell is what got asked. Did anyone go back through the medication list once the weight started coming off, or was the box checked and the refill sent? Did anyone ask what the patient is actually eating on the days the nausea is bad, which is the question that separates a tolerable side effect from six weeks of accidental starvation. Nothing on that list has a degree attached to it. It has time attached to it, and time is a scheduling decision made by somebody in an office who has never met the patient.
For patients
You are allowed to ask who you are seeing and what their background is, and no reasonable clinician will be offended. What you should not do is treat the letters after the name as the whole answer. This study says a randomly chosen NP outperforms a randomly chosen physician 38 times out of 100. Experience with your specific problem is the more useful question. If you are starting a GLP-1, ask how many patients they have managed on it. If you are calling a virtual urgent care with chest pain or a severe headache, understand that any competent clinician in that setting is going to send you somewhere with a CT scanner, and that is the correct answer rather than a failure of the visit.
For clinicians
Two things I would take into practice from this paper. First, the within-profession spread being wider than the between-profession spread should change how we think about quality improvement. We spend enormous political energy on scope-of-practice fights and almost none on identifying and coaching the outliers inside our own group, and the data says the second one has more room in it.
Second, the actionable finding for anyone designing care is the complexity gradient rather than the average effect. Straightforward cases should route broadly. Diagnostic ambiguity and high-acuity complaints should route to whoever has the most reps with them, assigned on individual performance data rather than on license class. That is a solvable engineering problem and almost nobody is solving it.
The Bottom Line
This is a serious paper with a genuinely strong design, and its central finding is not the one being headlined. Physicians came out ahead on average in a VA emergency department. NPs came out ahead in 38 percent of head-to-head matchups, the difference nearly disappeared on straightforward cases, and thirty-day mortality was a wash. The spread inside each profession was wider than the gap between them, which is the part worth carrying around.
For virtual weight management I expect the gap to be small, and I care far more about how much time the visit gets and how deep the protocol runs than about which degree is on the screen.
Match the case to the clinician. That is the finding.
Scott Rennie, D.O.
Board Certified in Obesity Medicine and Family Medicine
This blog is for educational purposes only and does not constitute individual medical advice. Always consult your own physician before making changes to your health, medications, or treatment plan.
Sources
- Chan DC Jr, Chen Y. The Productivity of Professions: Evidence from the Emergency Department. American Economic Review, August 2026. https://www.aeaweb.org/articles?id=10.1257/aer.20241007
- Berkeley Research. New study upends traditional thinking about doctors versus nurse practitioners. August 22, 2026. https://vcresearch.berkeley.edu/news/new-study-upends-traditional-thinking-about-doctors-versus-nurse-practitioners
- American Medical Association. Nurse practitioners’ care linked to 11% longer stays in the ED. https://www.ama-assn.org/practice-management/scope-practice/nurse-practitioners-care-linked-11-longer-stays-ed
- Clinician.com. Organizations Take Issue with Data Regarding Nurse Practitioner Care in the ED. https://www.clinician.com/articles/organizations-take-issue-with-data-regarding-nurse-practitioner-care-in-the-ed
- American Association of Nurse Practitioners. Nurse Practitioners in Primary Care (2025 NP count). https://www.aanp.org/advocacy/advocacy-resource/position-statements/nurse-practitioners-in-primary-care
- Chan DC Jr, Chen Y. The Productivity of Professions: Evidence from the Emergency Department. NBER Working Paper No. 30608, issued October 2022, revised August 2026. https://www.nber.org/papers/w30608