Spine surgery can answer its critics only with its own data

Advertisement

On Sept. 24, The Economist published a leader titled “Back and Shoulder Surgery Is Often Worse than Useless.” Its subtitle was blunter: “Millions of operations should be scrapped.” It placed lumbar decompression among six common procedures it described as “no better than drugs, physiotherapy or just Father Time.”

Many of us will want to respond by defending our specialty. I would urge our field to resist that instinct. The editorial overstates its case, and I will explain where. But it raises a question we must answer one patient at a time: Were our recommendations appropriate? Would our peers agree? Would the data agree? We cannot settle that question by counting operations or recalling our best results. Until we can answer it with our own data, we will keep losing this argument, and we will deserve to.

I say this as a neurosurgeon who operates on the spine every week and who has spent much of the past decade studying unnecessary spine surgery. Those two commitments come from the same conviction. The right operation, in the right patient, at the right time, gives people their lives back. The wrong one takes something from them that we cannot return.

What the editorial gets right

Surgical procedures spread without the proof we require of drugs. In a 2020 review in Pain, Harris and colleagues found that among randomized trials of common surgical procedures for chronic musculoskeletal pain, only about 1% compared the operation with not operating. We have studied how to perform these operations far more carefully than whether to perform them.

When operations have been tested against placebo, several have failed. In the CSAW trial, arthroscopic subacromial decompression for shoulder impingement did no better than an arthroscopy in which nothing was removed.

Our own field has also adopted riskier operations ahead of the evidence. In a Medicare analysis of surgery for lumbar stenosis, complex fusion carried more than twice the rate of life-threatening complications of decompression alone (5.6% versus 2.3%) and more than three times the hospital charges. It was being used without trial evidence that it improved outcomes for stenosis.

In its 2025 analysis of Medicare claims, the Lown Institute classified more than 200,000 spinal fusions, laminectomies and vertebroplasties as low-value, at a cost of more than $1.9 billion over three years. These are counts derived from claims-based criteria, such as fusion for low back pain without a recorded radicular, traumatic or structural diagnosis. They do not confirm that every one of those operations was unnecessary. But they describe a pattern we should not dismiss.

Where the editorial goes too far

The Economist’s central claim draws on a 2021 umbrella review in The BMJ of 10 common orthopedic procedures. That review found randomized evidence of superiority over non-operative care for two: carpal tunnel release and total knee replacement. Hip replacement had never been compared with nonoperative care in a randomized trial. The editorial nonetheless counts hip replacement among the proven procedures despite the absence of a randomized comparison, then treats a lack of demonstrated average benefit for six others as proof that they are useless.

The review’s authors were more careful than the editorial. They wrote that these procedures “may be effective overall or in certain subgroups.” “Not shown to be superior on average” and “useless” are different claims. The first describes the state of the evidence. The second is a verdict on every patient who might receive the operation.

Lumbar stenosis shows why the distinction matters. Delitto and colleagues randomly assigned patients who were already surgical candidates to decompression or structured physical therapy. At two years, physical function was similar. Fifty-seven percent of the therapy group eventually had surgery, and analyses designed to account for that crossover still found no significant difference. The honest reading is important: a structured course of therapy is a reasonable first strategy for many surgical candidates, with surgery available if it fails. It does not follow that decompression is useless for the patient who has completed that therapy and is still losing the ability to walk.

The editorial commits the error it charges us with: certainty beyond what the evidence supports. Our answer should not be an equal and opposite certainty. It should be more disciplined than our critics.

The line that matters is the indication

“Lumbar spine surgery” is not a single intervention, and the evidence does not treat it as one. Sorted by indication, the literature is clearer than either the editorial or our own marketing suggests.

IndicationWhat the best evidence showsWhat it means for practice
Lumbar disc herniation with persistent sciaticaSurgery relieves leg pain faster than non-operative care. In a trial of patients treated after several weeks of sciatica, outcomes converged by one year (Peul et al.). In SPORT, heavy crossover limited the randomized comparison, and a smaller advantage persisted in as-treated analyses. For sciatica lasting 4 to 12 months, a later trial found a more sustained advantage with surgery (Bailey et al.).Early on, surgery is a legitimate choice for faster relief, and so is waiting; the patient’s goal should decide. Longer-lasting sciatica strengthens the case for surgery.
Lumbar stenosis with neurogenic claudicationAs-treated analyses of SPORT favor surgery. Delitto et al. found similar function at two years with structured therapy, despite 57% crossover to surgery.Structured therapy is a reasonable first step for many. Surgery for those who do not respond or are losing function.
Degenerative spondylolisthesis with stenosisMixed results on adding fusion (Försth et al.; Ghogawala et al.). Decompression alone was non-inferior in NORDSTEN-DS (Austevoll et al.).Consider fusion only for a specific, documented reason, such as instability, not out of habit.
Degenerative cervical myelopathyA prospective multicenter surgical cohort showed improved function and neurological scores after decompression across mild, moderate and severe disease (Fehlings et al.). A 2017 clinical practice guideline recommends surgery for moderate and severe disease.In moderate to severe or progressive myelopathy, delay can cost function that may not return, and timely surgery is the standard. Mild disease may be managed with supervised rehabilitation and close follow-up, or with surgery.
Painful osteoporotic compression fractureVertebroplasty was no better than sham in two 2009 blinded trials (Buchbinder et al.; Kallmes et al.). A later sham-controlled trial limited to very recent, severely painful fractures found a benefit (VAPOUR).Should not be routine. A possible benefit in selected patients with very recent, severe fractures remains contested.
Chronic axial low back pain without instability, deformity or nerve compressionFusion showed a modest advantage over usual care (Fritzell et al.) but little or no meaningful difference against structured rehabilitation (Brox et al.; Fairbank et al.).Should be rare, reviewed by a team, and preceded by comprehensive non-operative care.

The line is not between surgery and no surgery. A structural finding that matches the patient’s symptoms strengthens an indication, but it does not by itself establish benefit. We should defend operations, without apology, when the anatomy, symptoms, examination, clinical course and the patient’s goal all support them. Operations aimed at pain alone, without that support, should be held to a standard they often cannot meet. And because degenerative findings on MRI are common in people with no pain at all, the imaging report alone can never be the indication.

Overuse is concentrated, and review changes decisions

Two findings suggest this is a solvable problem rather than a verdict on our specialty.

First, overuse is concentrated. In the Lown analysis, the average hospital’s overuse rate for fusion and laminectomy was 13%, but rates varied widely, from 1.4% at one California hospital to 57.2% at one in Pennsylvania. For spinal fusions meeting Lown’s overuse criteria, the 10% of physicians with the most such procedures accounted for 60% of the total. That ranking reflects the number of procedures, not each surgeon’s personal rate, and claims data cannot see every clinical detail. Still, a pattern this concentrated can be measured, reported and addressed.

Second, structured review changes decisions. In a prospective Brazilian study of patients who had already been recommended for spine surgery, only 66 of the 425 who completed a full second opinion, or 15.5%, received the same surgical recommendation. In a later analysis of that program, 737 recommended fusions fell to 65 after review. Because the originally proposed operations were not performed, the comparison with their expected results was modeled, and those modeled outcomes suggested comparable rates of meaningful improvement.

In a 2017 study my colleagues and I published in Spine, a multidisciplinary conference recommended non-operative care for 58 of 100 patients who had been told elsewhere that they needed lumbar fusion, and revised the surgical plan for 28% of those who went on to surgery. That study measured changed recommendations, not outcomes, and it does not prove that all 58 fusions would have been unnecessary. What it showed is how much a surgical recommendation can depend on who makes it and how closely it is examined. The conference was, in effect, asking two questions of every case: Would our peers agree? Would the data agree? Review can also change the surgical plan for patients who still undergo an operation.

Why the data are missing

None of this requires bad surgeons. It requires only ordinary surgeons working inside a system that rewards action and rarely measures results. An MRI finding becomes the diagnosis, and the diagnosis becomes the plan. Patients arrive after weeks of referrals expecting an operation, and in my experience “failed conservative care” is too often two or three therapy sessions. Many expect surgery to restore the back they had at 20, and too often we do not correct them. Productivity models pay more for fusion than for decompression, and no more than an office visit for a sound decision not to operate. And without following comparable patients who do not have surgery, we can too easily credit an operation for every improvement that follows it.

Four answers every spine program should be able to produce

Here is a simple way to find out whether your program can answer the questions above:

Take 100 consecutive patients evaluated for elective spine surgery at least a year ago, including those advised against it. For each, ask from the record: Was the recommendation appropriate? Would our peers agree? Would the data agree? Then find out what happened to each patient, and count how many you could not follow.

Most of us will say we already know which patients need surgery. But we remember the patients who return, and we remember the striking recoveries. Consecutive cases force us to look at the uncertain decisions too, and the number lost to follow-up tells us how much confidence to place in the results. A disagreement between a surgeon and reviewers calls for investigation, not an automatic verdict.

Doing this routinely requires four answers. Could your program produce them today? Every spine program, of any size, should be able to report them and review them at least twice a year.

  1. Outcomes by indication. Collect patient-reported outcomes at baseline, three months and one year, and report them by indication rather than procedure code, with complications and reoperations. Report the share of eligible patients who actually supplied one-year outcomes, and the share who reached a meaningful improvement, not only the average change.
  2. A reviewable rationale for every elective decision. For each recommendation, to operate or not, record the indication, whether imaging and examination agree, the patient’s goal and the non-operative care already tried. Send a sample to colleagues who were not involved: would our peers agree, and would the data agree?
  3. Variation in fusion use. Track fusion rates by surgeon, adjusted for case mix, including fusion as a share of lumbar decompressions.
  4. Outcomes for patients treated without surgery. Follow them with the same instruments at the same intervals. That comparison is essential for testing our selection decisions, though it cannot replace trials that determine what surgery itself adds.

None of this requires a research grant. Outcome collection can be built into the electronic health record at scale. It is an infrastructure decision, and it is overdue.

What spine surgeons should do

Program data matter only if they change what happens in the exam room. The Economist calls for shared decision-making, and on this we agree completely. But a patient cannot share in a decision without the information that decision requires. For every elective spine operation we recommend, or decline to recommend, we should:

  1. Work through five questions with the patient, and answer them honestly. What happens if I wait? What exactly are we treating, and how certain are we? What are my reasonable options? What am I trying to achieve? For me, are the expected benefits worth the risks and burdens? Name the goal behind the recommendation, whether the fastest relief, the most durable result or the lowest immediate risk. And describe waiting accurately: most sciatica from a disc herniation improves without surgery; many patients with tolerable, moderate lumbar stenosis remain stable for years; moderate to severe or progressive cervical myelopathy is the exception, where delay can cost function that does not return.
  2. Invite the question that tests all the others: What would make you change your recommendation? And ask one of ourselves: would I advise the same for my own mother, if she could see our outcomes for patients like her, including those we advised not to have surgery?
  3. Let the answers change our practice. That may mean narrowing a fusion indication, sending patients to physiatry or physical therapy first, recalibrating what we tell patients to expect, stopping an operation for a group it does not help, or operating sooner where delay costs function.
What health system leaders should do

The Economist’s remedy is for governments to run more trials and for payers then to refuse coverage for procedures that do not benefit patients. Where trials are decisive, that is reasonable. Applied bluntly to averages, it would deny care to the patients who clearly benefit along with those who do not. Health systems can act more precisely, and most of the tools are already within our control:

  1. Fund one-year outcome follow-up, including for patients who do not have surgery, and the infrastructure to produce the four answers above.
  2. Build independent review into the pathway. Make physiatry or physical therapy the front door for non-urgent referrals, route elective fusion, multilevel and revision cases through a multidisciplinary indications conference, and offer a structured second opinion before elective fusion, so that each of those decisions faces the question: would our peers agree?
  3. Give surgeons private, case-mix-adjusted feedback on their own utilization. If a small share of physicians accounts for much of the overuse, this is a tractable problem, and peer comparison has changed physician behavior in other areas of care.
  4. Stop measuring spine programs by volume alone. A decision not to operate, followed by a good outcome, is program value, not lost revenue. Compensation and program scorecards should reflect that.
  5. Contribute to the evidence. Join registries and pragmatic trials that compare operations with structured non-operative care, because only trials can tell us what surgery itself adds.

These are budget decisions, not research projects. They protect patients. They also protect the programs that adopt them. The scrutiny The Economist has brought to a general audience is already arriving from payers and regulators. Programs that can show their indications and outcomes will be able to answer it. Programs that cannot will have the answer written for them.

What spine societies should do

Individual programs cannot make their answers comparable on their own. Our societies have already built much of the foundation: the American Association of Neurological Surgeons and the American Academy of Orthopaedic Surgeons jointly run the American Spine Registry, and the North American Spine Society built a diagnosis-based registry designed to capture surgical and non-surgical care. The next steps are within their reach:

  1. Align existing registries around a common minimum dataset: the indication, the patient’s goal, the alternatives tried, the reason surgery was recommended or declined, and patient-reported outcomes at one year, for patients who do not have surgery as well as those who do, reported with the share of patients actually followed.
  2. Publish shared definitions of common indications and of an adequate trial of non-operative care, so that “failed conservative care” means the same thing everywhere.
  3. Support confidential peer review and benchmarking, so that any program can ask whether its peers would agree with its decisions and see how its results compare.

That does not ask societies to defend every operation or to condemn their members. It asks them to set the standard by which decisions are judged.

The alternative is already taking shape. If we cannot define and measure an appropriate decision ourselves, others will define it for us from claims and coverage criteria. Since January, Medicare’s WISeR model has asked providers in six states to choose between prior authorization and post-service, pre-payment review for selected Original Medicare services, including a selected cervical fusion procedure. A reviewer sees only what our records contain. If our records do not capture the patient’s goal, the examination and the therapy that preceded a recommendation, no one can defend the decision, including us.

The question for every program

I hold my own program and my own practice to the same test. Recently, I cared for a woman in her 70s who could no longer walk to the end of her driveway. After a few minutes on her feet, her legs went heavy and numb. Her spinal canal was severely narrowed. She had completed months of physical therapy and was getting worse. We operated, and within weeks she was walking her neighborhood again. Had she read The Economist first, she might have refused the operation that gave her life back.

I can tell you why I recommended her operation. I can tell you she walked her neighborhood again. But that is one patient’s good result, and confidence is exactly what The Economist has asked the public to stop accepting from surgeons. If I ask the public to trust my judgment, my program must be able to show that our peers would agree with that decision, that the data would agree, and how patients like her, and those we advised not to have surgery, actually fared. Those data cannot, on their own, prove what her operation added; that is the work of trials. But a program that can answer those questions will not need to defend spine surgery in general. Its data will show how it decides, and where the results fall short, they will tell us what to change.

Dr. Yanamadala is vice chair of neurosurgery and system medical director for Neurosurgery Quality, Research and Innovation at Hartford (Conn.) HealthCare. The views expressed are his own.

At the Becker’s 32nd Annual Meeting: The Business and Operations of ASCs, taking place October 29-31 in Chicago, ASC leaders, surgeons and healthcare executives will explore strategies to drive growth, enhance operational performance, navigate reimbursement challenges and prepare for the future of ambulatory surgery. Apply for complimentary registration now.

Advertisement

Next Up in Orthopedic

Advertisement