The ethics of digital twins in spine care
Imagine that you are considering spine surgery. Your surgeon reviews your MRI, symptoms, medical history, prior treatments and goals. Now imagine that a computer combines those details with laboratory results, movement data and information from a wearable device to create a virtual model of you. As your condition changes, the model changes too.
This is the promise of a healthcare “digital twin.” Unlike a traditional risk calculator, which produces a prediction based on information collected at a single point in time, a digital twin is intended to evolve. In spine care, it could help clinicians compare treatments, anticipate complications and detect when recovery is moving off course. That possibility is compelling because patients with similar imaging findings can experience very different levels of pain, disability and treatment response.
We argue that a digital twin should inform a clinical conversation but not fully guide clinical decision-making. In our recent scoping review of digital twins and multimodal artificial intelligence in spine care, we found that the field remains in a concept stage because most models do not always capture data dynamically. Most importantly, no spine-specific digital twin has been shown to improve clinical decisions or patient outcomes so far.
Applying digital twin technology to spinal surgery also creates several ethical risks. The first risk is promoting a false sense of certainty. A model may tell a patient that the risk of complication is 17%. While that number looks precise, it may be untrue. A model developed at one hospital may perform differently when the same model is implemented at another. The model may also be less accurate for patients who were poorly represented in the data used to train it.
In our related review of spine surgery risk prediction, artificial intelligence often produced only modest and inconsistent improvements over simpler statistical methods. Testing in new hospitals, determining whether predicted risks match what occurs, and showing that the tool improves care remain essential.
The second risk is that the model may capture the medical record more accurately than it captures the patient. Factors such as pain, mental health, employment, caregiving responsibilities, transportation and access to rehabilitation can all shape patient outcomes after surgery. These factors are often missing or poorly measured. A system trained on incomplete data can also reproduce inequities. Continuous monitoring may also favor patients with reliable technology, internet access and the ability to use wearable devices. A model cannot be considered personalized if it systematically overlooks important parts of a person’s life.
The third risk relates to assigning responsibilities for when a digital twin model is wrong. The surgeon may have followed an AI-generated recommendation, the hospital may have deployed it, and a company may have designed or updated it. Who has what responsibility?
Responsible progress in digital twin technology means more testing across different hospitals and patient populations before it guides routine care. Clinicians and patients should understand what the technology can and cannot do. People should know what information is being collected to generate digital twins, how it will be used and who is responsible for reviewing and acting on the model’s recommendations.
Most importantly, technology must be aligned with the patient’s goals and the clinician’s judgment. A prediction cannot decide how much pain a patient is willing to live with, which risks are acceptable or what outcome matters most in that person’s life.
A digital twin may eventually become a valuable map of a patient’s changing condition. The goal should be to make shared decision-making more informed, responsive and genuinely personal.
By Samer G. Salman and Rohan A. Phadke, second year medical students at Baylor College of Medicine
