AI, Data Protection and Copyright

Who owns the image? Who is responsible for the images? Do AI-generated images have to be labeled? What were AI models trained on? What do I need to consider?

When artificial intelligence is used in a professional or personal context, the provisions of the General Data Protection Regulation (DSGVO) apply to personal data processing. This includes, for example, personal information, employee data, or emails. Since many cloud-based AI providers store input data by default and use it for further model training, sending sensitive personal data to external servers often constitutes a violation of data protection laws. To comply with the GDPR, users must ensure that a legal basis (e.g., consent from the data subjects or a legitimate interest) exists, that data is anonymized or pseudonymized as much as possible, and that data processing agreements (DPAs) are in place with the AI providers. Many commercial services now offer specialized enterprise or API versions that guarantee that the input data will not be used for training and will be processed only in compliance with the DSGVO.

Not only personal data deserves protection, but also sensitive, creative, and scientific content. Entering unpublished manuscripts, research results, source code, or design concepts into freely accessible AI tools exposes this data to the risk of being stored on external servers and, in the worst case, being used to train future models. This can result in the unintentional transfer of valuable intellectual property to third parties. To prevent this, only secure environments should be used for creative and scientific work processes. One way to protect images from being used in AI training is through a technique known as “data poisoning.” An article with information and tools can be found here: Link

Under the current Copyright Act (UrhG) (Link), those who created the work and/or published it (e.g., on their website) are responsible. Anyone who uses AI to produce and/or share images, videos, or audio messages that glorify violence or contain other illegal content may be sued. Copyright protection for an AI-generated image or song only applies if the AI played a minor role in creating it. Simply providing a prompt is not sufficient to meet this requirement. This means that works generated using AI tools such as Dall-E, Midjourney, or Stable Diffusion or music generators like Suno that only base on a prompt are not protected by copyright. The question of how to demonstrate a sufficient creative process has not yet been clarified.

AI-generated images, videos, or audio deepfakes that depict real people or imitate their voices (voice cloning) must always be labeled as AI-generated. The EU AI Act will establish a law to enforce this transparency rule for AI-generated media in the future to prevent disinformation. Different AI providers have varying policies regarding the labeling requirement, which should be observed. In most cases, it is sufficient to name the software provider(s) in the credits.

AI models require millions of images, videos, text, and audio files to function. For example, the dataset used to train Stable Diffusion includes more than 5.8 billion references to images and their descriptions. These images come from the internet and are collected by programs called “crawlers.” These crawlers scour the internet and gather large amounts of data and their descriptions to use as raw material for training AI models. This content includes not only general material but also sensitive, personally identifiable data such as nude images and bank information, as well as copyrighted works by designers and creatives. Training AI applications with image files from individual artists enables the general public to create so-called “mimicries”—that is, replicas of the artist’s specific style.

You can protect your own images using a technique known as “data poisoning.” An article about two tools for protecting your own images can be found here: Link

The Stable Diffusion training dataset is one of the few that are publicly available. On the website haveibeentrained.com, you can search the dataset to see if your own photos or works have been used in it. The European Union’s AI Act is a first step toward making the disclosure of AI model training data mandatory for all providers. (As of 2024)

  • If you want to train your own models, you should either only use your own data or ask for permission from the copyright holders. It is also possible to use images that are licensed for such use (e.g. under Creative Commons licenses).
  • AI-generated images are usually free of copyright. However, different AI services have their own rules regarding the use of generated works. Therefore, the guidelines of the respective providers should be checked.
  • Anyone who refers to existing works, such as images, films, novels, well-known melodies, or fictional characters in AI-generated content may infringe the copyright of the original creators. However, the “pastiche Schranke” applies: parodies, caricatures, and pastiches (e.g., homage or satire) are permitted, even when they reference copyrighted works.
  • Cloning a real human voice raises issues not only regarding copyright but also regarding general personality rights (“right to one’s own words” and “right to one’s own voice”). The consent of the affected person is necessary for the reproduction of real voices.

Sources