Publicly available data is not automatically free for AI. New guidelines clarify the responsibility

 

Why the public availability of data on the Internet does not automatically mean that companies can use it for AI training 

How the new EDPB guidelines clarify anonymisation, web scraping and traceability requirements 

What companies have to deal with when working with data for AI from the point of view of GDPR, the AI Act and copyright 


 

The European Data Protection Board (EDPB) has published draft guidelines on anonymisation and automated collection of data from the Internet in the context of generative artificial intelligence, and at the same time adopted the final version of the guidelines on the processing of personal data via blockchain. The Czech Office for Personal Data Protection also drew attention to the new documents.

The draft guidelines on anonymisation and web scraping are now the subject of a public consultation, which runs until 30 October 2026. However, they already show the direction in which the interpretation of the rules on personal data protection will go. The main message is that the technical availability of data does not mean that it can be collected, used or stored without further consideration.


Measures make sense

The EDPB's approach can be considered the right and necessary step. The Board does not opt for a blanket ban on automated data collection from the Internet or the use of anonymisation or blockchain. It leaves room for organisations to use these technologies, but at the same time emphasises their responsibility to comply with personal data protection rules.

Companies must therefore be able to document, in particular, the purpose of the processing, the appropriate legal basis, the origin of the data and the technical and organisational measures taken.

In the case of anonymisation, it is not enough to simply remove a name or other direct identifier. In particular, the draft EDPB guidelines recommend assessing whether a particular person can be singled out from others, whether information about them from different sources can be linked, or whether conclusions can be drawn about them from available data.

It is also important that the anonymity of data depends on the specific context. The same set of data may be de facto anonymous for one entity, while another entity may have additional information or means at its disposal to re-identify the person. Therefore, due to technological developments and the growing amount of data available, anonymisation cannot be perceived as a one-time technical act. Its effectiveness needs to be evaluated on an ongoing basis.


Publicly available does not mean free to use

A significant part of the new materials concerns web scraping, i.e. the automated extraction of information from websites. This technique is also widely used in data acquisition for the development and training of generative artificial intelligence.

However, even data publicly available on the Internet cannot automatically be considered freely usable for AI training. If scraping involves personal data, the organisation must determine the appropriate legal basis for its processing, adhere to the principle of purpose limitation and data minimisation, and ensure the necessary transparency.

Special categories of personal data, such as health data, biometric data or information about political opinions or religious beliefs, require special attention. Their public availability does not in itself mean that they can be processed without further consideration.

At the same time, the EDPB recommends using reliable sources, recording the origin and time of obtaining data, and verifying its quality and relevance before using it, for example, for AI training.

One of the possible legal bases may be a legitimate interest. However, even this does not constitute a universal right to collect publicly available data. In a specific case, the organisation must assess the legitimacy of the pursued interest, the necessity of the processing and the impact on the rights and freedoms of the persons concerned.


Blockchain must take privacy into account by design

For blockchain, the EDPB's guidelines are already final. They place emphasis primarily on the protection of personal data when designing a solution.

The technical immutability of a blockchain record cannot serve as an excuse for the inability to correct or delete personal data. Therefore, as a general rule, the EDPB recommends avoiding storing personal data directly on the blockchain if such a solution would be contrary to the principles of the GDPR.

Organisations must deal with data minimisation, retention periods, access permissions or the possibility of exercising the rights of individuals in advance. For riskier processing methods, a data protection impact assessment (DPIA) will also play an important role.


The AI Act and copyright also come into play

At the same time, personal data protection is only one part of the rules that need to be taken into account when obtaining data for AI. In addition to the GDPR, it is now also necessary to take into account the European AI Act and copyright rules.

Under the AI Act, providers of general-purpose AI models (GPAI) are required, among other things, to establish a policy to comply with EU copyright law and to publish a sufficiently detailed summary of the content used to train the model.

Therefore, even from the point of view of copyright law, the public availability of content cannot be confused with an automatic right to its unlimited use. When obtaining training data, it is necessary to assess several legal regimes at the same time.


What risks can be expected?

The biggest risk remains merely formal fulfilment of the requirements. An organisation may have documentation ready, but at the same time it may not know the true origin of the training data or data obtained from an external supplier. This can be problematic especially for large datasets compiled from many public and non-public sources.

Another complication may be the financial and technical demands of controls, especially for smaller companies. With the growing requirements for traceability of the origin of data, it is not enough just to know what data an organisation uses. It will also be increasingly important to know where it comes from, when it was obtained, for what purpose it is processed and under what conditions it can be further used.

An increasing emphasis on documentation, data traceability and ongoing risk assessment can therefore be expected. This will require closer cooperation between legal, technology, security and data teams.


Impact on individual teams

The new guidelines will primarily affect teams involved in AI solution development, data management and data protection.

Legal teams will have to set up the legal bases for processing more precisely, assess legitimate interest, contractual liability of suppliers and the conditions of use of data obtained from the Internet or from external partners. For AI projects, it will also be necessary to assess the interplay between the GDPR, the AI Act and copyright.

Data Protection Officers (DPOs) and privacy specialists should be part of projects using artificial intelligence or blockchain from the design stage. They will assess the risks of personal data processing, transparency, the possibility of exercising the rights of individuals, as well as the possible need for a DPIA.

IT and security teams will ensure that data sources are recorded, set up technical measures for data collection and management, access control and secure data storage.

AI and data teams will need to better document the origin of training data, record its sources and time of acquisition, reduce the collection of unnecessary data and continuously evaluate the risk of the model re-identifying or memorising personal data.
 

“The new guidelines do not represent a ban on automated data collection from the Internet, anonymisation or blockchain. However, they remind us that a technical option alone is not enough. Organisations must be able to prove why they use data, where it comes from, and how they protect the rights of individuals. In addition, with AI, they must take into account that, alongside the GDPR, the AI Act and copyright also come into play. Therefore, the biggest change will not be the creation of another document, but the need for closer cooperation between legal, data, security and development teams.”