Gadgets & Reviews

AI may be learning from billions of images without copying any one of them

[post_content]


Disclaimer: This article has been automatically aggregated from

Here’s an uncomfortable question for the generative AI era: if an AI creates an image, can anyone actually point to the specific images that influenced it? A new study from MIT’s Computer Science and Artificial Intelligence Laboratory suggests that, for sufficiently large models, the answer may often be no. Researchers Dai and Gifford describe the phenomenon as “attribution decay” in their open-access paper, published today in Nature Communications. It refers to how the influence of any individual piece of training data becomes increasingly difficult to detect as the size of the dataset grows.

The bigger the dataset, the fuzzier the attribution

The researchers wanted to test something more precise than simply asking whether an AI model had been trained on a particular image. Their approach was essentially a counterfactual: what happens if a specific piece of training data is removed?

They found that, as models and datasets scale, removing an individual image can make little or no measurable difference to the generated result. The same can apply when removing an entire artist’s work or images of a particular person. In other words, the model may have learned broad visual patterns from an enormous pool of data without any single image being clearly responsible for a particular output. MIT researchers Zheng Dai and Professor David Gifford argue that if removing a particular piece of data doesn’t change the output, it becomes difficult to attribute that output to the removed example meaningfully.

That doesn’t settle the AI copyright debate

To be fair, the study doesn’t mean training data is irrelevant, nor does it settle the broader debate over whether AI companies can use copyrighted material without permission. Instead, its focus is narrower: whether a specific training image can be shown to have directly influenced a specific AI-generated output. A model may not reproduce any particular artist’s work while still relying on millions of images to learn things such as composition, lighting, textures, and artistic styles.

The more interesting takeaway is that this connection becomes increasingly difficult to trace as datasets grow. An AI-generated image may draw on patterns learned from an enormous pool of material without having a clear, identifiable source image behind it. The model may have learned from everyone, while leaving no single image with an obvious fingerprint on the final result.

for informational purposes only. We do not claim ownership, accuracy, or liability for the content provided. All rights belong to the original publisher.