Close Menu
  • AI
  • Content Creation
  • Tech
  • Robotics
AI-trends.todayAI-trends.today
  • AI
  • Content Creation
  • Tech
  • Robotics
Trending
  • AWS Strands Brokers Staff Releases Strands Harness: An Open-Supply Agent Harness With 28% Decrease Token Price at Comparable Accuracy
  • AI, Tariffs and Rare Minerals: Expectations for Trump’s Upcoming Meeting with Xi Jinping
  • Alibaba Qwen Releases Qwen-Picture-2.1: A 7B Open-Weight Mannequin for Picture Era and Enhancing
  • Collaboration Should Sit On the Coronary heart of Manufacturing’s Multi-Agentic AI Strategy. Right here’s How. – Unite.AI
  • Obtained an Android Cellphone? Google Thinks You’ll In all probability Desire a Googlebook Laptop computer
  • Tips on how to Use Substack: Classes from Creators
  • UN AI Panel Invokes Precautionary Precept on Loss-of-Management Threat – Unite.AI
  • US and China Talk about Alerting Every Different to AI Nationwide Safety Threats
AI-trends.todayAI-trends.today
Home»Robotics»Decades-Old Anonymized Medical Data May Cause AI Misdiagnoses Now – Unite.AI

Decades-Old Anonymized Medical Data May Cause AI Misdiagnoses Now – Unite.AI

Robotics By Gavin Wallace20/09/20268 Mins Read
Facebook Twitter LinkedIn Email
Lumo AI Review: This AI Chatbot Keeps My Data Private
Lumo AI Review: This AI Chatbot Keeps My Data Private
Share
Facebook Twitter LinkedIn Email

An anonymous patient record that was included in a medical dataset decades ago could cause AI trained on it to identify the person. real patient’s characteristics years later – and treat a 2026 patient as if it was still 2012 (or whichever year their anonymized patient data was first included in the dataset/s).

It could be, for example that an AI system could treat a patient who had been treated with cancer 20 years before, but who went into remission. In the context of an older cancer-stricken self – logically even increasing the chances of a false diagnosis of recurrence.

If the data from 2012 showed that the patient was in perfect health, the optimistic outlook could be applied to the older, current version The patient’s condition can be masked, preventing the diagnosis of any new issues or conditions that need to be treated.

A new study shows how the memorization effect can be detrimental to a patient who returns. Models trained using Alice’s healthy ECGs from earlier in her life assign a mere 15% chance of her having a heart attack later, while models that have never seen her records only give a 73% probability. Source

Curbing ‘Data Immigration’

The scenario is made more probable by various national and regional directives recommending that patient data for such purposes should ideally be taken from the country in which systems derived from it are intended to be used, which ‘localizes’ the patient pool in the datasets, and notably increases the chance of this kind of undetected de-anonymization event.

Limiting the data to a single national dataset, increases the risk of memorization – the possibility that the model will see the same data so many times during training that it becomes ‘fixated’ on these over-learned patterns. In contrast, the bigger and more diverse the data set, the higher the probability that the model is going to become “fixated” on certain patterns. generalize Instead of memorizing data configurations and data-points, it is better to use unseen data.

However, adding ‘foreign’ medical data risks to introduce evidence of national health trends and idiosyncrasies that would not apply in the country in which the trained model is being deployed; in this respect, as the new paper It is true that a little memorization can be useful and appropriate, as it allows the model to fit the general characteristics of the population.

The paper states*:

‘[When] Memory is a risk that’s often overlooked when models are applied prospectively. As a patient’s records tend to be highly self-similar, future records may serve as partial cues for a memorised historical record. The model would then predict the patient’s past health.

It is not an imaginary scenario. The same populations from which the AI model’s training data was derived are used to deploy medical AI models.

‘For example, Germany’s national breast cancer screening program uses An AI model that was developed using the 1.2 million mammograms of the screening population.

‘Comparable deployment is underway elsewhere, including national breast cancer screening programmes in the United Kingdom You can also find out more about the following: Sweden.

‘Regulatory guidance actively encourages this, with the EU AI Act  requiring that training data reflect the geographical and contextual setting of intended use (Art. 10(4)), similar international guidelines For medical AI, the training data must include enough patients to represent the population that will be used.

The Persistent History of Calamity.

Although the paper did not specifically address the issue, it would seem statistically logical to assume that anonymous patient data will be populated by patients who don’t (or didn’t) regularly go for medical examinations and have their anonymized records periodically enter the AI data streams. Almost exclusively when the patient needed to be treated –  which would tend to increase the chance of ‘projecting’ a long-cured condition onto the same patient now, since the records have no ‘good health’ counterbalance.

The evidence is overwhelming. prior research It is possible to buy ‘informed presence bias’ Findings in electronic medical records that show that the majority of data sets are from periods where patients were sick or interacting with health services.

One study The study found that patients with severe illnesses had 5,05 days more with lab data, and 6,85 times the number of days they spent with prescriptions.

By 2022, approximately 76% US adults will be aged 18 and older reported A routine examination within the past year is defined by the CDC to be a physical examination that does not treat a condition. And in Europe, a survey of multiple countries will take place 2023. found 58% of people attended some or all appointments for preventive care, but only 15% went to every appointment.

What is the extent of effect?

The researchers tested four datasets: ECG recordings (electronic heart rate), chest X-rays and electronic medical records. MIMIC-CXR chest X-ray database You can also find out more about the following: MIMIC-IV-ED emergency care records, alongside MIMIC-ECG Much larger HEEDB ECG collection. Datasets ranging in size between tens to thousands of people and more than 1.8 millions.

In some cases the authors have noted that changes in predicted probability can exceed 70 percentage points.

Results showing how historical patient data changes later predictions. Top panels compare the frequency and size of these changes across the four datasets; the lower panels show how the effect changes with time since the patient's last training record. Significant effects become less common over time, but remain detectable decades later.

The results show how the historical data of patients can affect future predictions. These top panels present a comparison of the size and frequency of the differences between the four datasets, while the bottom panels display the change in effect over time. Significant effects become less common over time, but remain detectable decades later.

Researchers divided randomly the models and performed a comparison. This did not result in any changes to future predictions.

It is possible that the effect will last decades. Although the predictions were generally weaker as time passed, HEEDB’s dataset of ECGs collected between 1980 and 2025 revealed significant differences. More than 25 Years The most recent record of training for the patient.

Diagnostic Harm

The researchers of the new paper – titled Memory bias in AI – next simulated how memorization could affect diagnostic accuracy when patients later encounter a model trained on their earlier records.

The future cases are divided into three categories: whether the patient has developed a condition that is new, whether it was not present in their training history or whether they returned with the same state of health.

In the case of new diagnoses, this effect was consistent with a negative one. Previous models that used patient records showed a lower level of sensitivity for varying conditions represented by the four datasets. These included ECG anomalies, lung diseases and infarcts.

The models produce more false negatives if the health of the patient has changed, but not if it is unchanged. You can also check out our other blog posts. As the records of the patients’ previous conditions changed, it was easier for the model to make the right diagnosis. The patient’s condition increased both sensitivity and precision. The condition was similar to the one already described in the training data.

This could lead to a system that is difficult to evaluate accurately. If most patients return with the same condition, then the model may perform better, hiding the fact that it has a lower sensitivity for patients who develop new conditions.

Remedies

Researchers tested the system as a potential safeguard. differential privacy (which limits the amount of information that a training model can retain about a specific training example) during training.

It was found that, even with the most extreme settings, the protection of individual records had a minimal effect on the memory effect. The effect was almost entirely eliminated when all the records of the same patient were protected.

The researchers point out that, even if data has been anonymized irretrievably, it is still not possible to enclose patient information in such a manner that would allow for these methods. membership inference attacks become a riskIf it does, resources for training will be increased.

You can also read our conclusion.

The authors close by observing that while the concerning incidents currently occur at a relatively low rate, this may change in the future*:

‘There is reason to expect the number of missed diagnoses attributable to memorisation bias to increase in the future, although we do not test this directly. Previous research shows that the model’s capacity affects how much data is memorised. [6, 7, 9, 10, 40].

‘This is concerning, given that current AI model development is guided by “scaling laws” [41]In pursuit of better performance, they drive rapid growth both in model size and data set sizes.

‘As larger AI models are trained on historical data sourced from ever-larger patient populations, the absolute number of individuals affected by memorisation bias is likely to rise drastically, exacerbating the risks we identify here.’

 

 

* I have converted the inline citations of authors into hyperlinks. This is done with some interpretation when exact reproduction or help would not be possible.

First published on September 18th, 2026. Links corrected as of 18:17 (EET).

AI dat data
Share. Facebook Twitter LinkedIn Email
Avatar
Gavin Wallace

Related Posts

Collaboration Should Sit On the Coronary heart of Manufacturing’s Multi-Agentic AI Strategy. Right here’s How. – Unite.AI

21/09/2026

UN AI Panel Invokes Precautionary Precept on Loss-of-Management Threat – Unite.AI

21/09/2026

How AI Modernizes Lending Alongside Legacy Banking Techniques With out a Teardown – Unite.AI

21/09/2026

Lovable Acquires Sutro, Firm Behind the SLang Programming Language – Unite.AI

21/09/2026
Top News

Apple’s camera chief believes AI can give you superpowers

The Anthropic Way of Thinking about AI Agents in the Physical World

xAI Adds 19 New Gas Turbines Despite Ongoing Lawsuit

Prego Has a Dinner-Conversation-Recording Device, Capisce?

This Robot is Making Meals in San Francisco’s Tenderloin for a Nonprofit

Load More
AI-Trends.Today

Your daily source of AI news and trends. Stay up to date with everything AI and automation!

X (Twitter) Instagram
Top Insights

The GluonTS Multi-Model Workflow Guide: Synthetic Data, Advanced Visualizations, and Evaluation.

24/08/2025

Meta AI Proposes “Metacognitive Use”: Using LLM Chains-of-Thought as a Procedure Handbook to Cut Tokens By 46%

22/09/2025
Latest News

AWS Strands Brokers Staff Releases Strands Harness: An Open-Supply Agent Harness With 28% Decrease Token Price at Comparable Accuracy

21/09/2026

AI, Tariffs and Rare Minerals: Expectations for Trump’s Upcoming Meeting with Xi Jinping

21/09/2026
X (Twitter) Instagram
  • Privacy Policy
  • Contact Us
  • Terms and Conditions
© 2026 AI-Trends.Today

Type above and press Enter to search. Press Esc to cancel.