← The Vault
Everyday AI

Why AI developers keep scraping artists’ portfolios

When you post art online, it often gets copied by automated programs used to train AI models. A platform called Cara tried to stop this, but they were still hit by data scrapers. Now, the person who caused the problem is helping them build a new way for artists to track where their work ends up, highlighting the ongoing struggle over who owns creative content in the age of AI.

Edition № 486Room: Everyday AI29 August 20262 min readSources: 1
Article

A small platform for artists called Cara recently found itself in a digital tug-of-war. The site was designed to give creators a safe space to share their work, specifically by keeping AI models away from their portfolios. But despite their efforts, the site was hit by a series of automated data collections that grabbed millions of images, causing a massive outcry among its users.

WHAT'S HAPPENING

The site was targeted by individuals using automated tools to scrape, or systematically copy, thousands or millions of images from the platform at once. This is a common practice where software robots visit a website and download every piece of publicly available information. In this case, the scrapers collected art, titles, and even personal user details. One individual even boasted about it on a public forum before later regretting the action. These collected images are often used as training data—the massive collections of information that AI models process to learn how to identify, categorize, and eventually recreate images in the style of those artists.

The cat and mouse game of data collection

HOW IT WORKS

Training an AI is like teaching an intern. To make the intern good at art, you show them millions of examples of paintings, photos, and drawings. The AI does not actually 'see' these images like a human does; instead, it looks for mathematical patterns in the digital files. When a person scrapes a website, they are effectively building a giant library of these examples to feed into the AI's learning process. For a website owner, stopping this is extremely difficult because a website that is visible to the public is, by definition, visible to these automated collection tools. Even if you block one type of robot, others can disguise themselves or change their tactics. There is no simple switch to make a website truly invisible to someone determined to copy its contents.

WHY IT MATTERS

This incident highlights a fundamental friction in the digital age: the tension between information that is publicly accessible and the desire for control over that information. Most of the massive AI tools we use today were built using similar web-wide harvests. Because legal rules regarding how AI companies can collect and use this data are still being written, there is a lack of clear protection for creators. The move toward new tools that help artists monitor where their work ends up is an attempt to create accountability, but it remains an imperfect defense against a widespread, internet-wide practice. It raises a difficult question: if you put something on the open internet, can you ever really keep it from becoming part of someone else's training set?

Sources
← PreviousWill doctors soon be replaced by AI?Next →Why robots are moving into the server room
Tomorrow's edition · free

Liked this one? The next lands at breakfast.

Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.

By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy

← Back to the Vault