Tag Archives: copyright

Text and Data Mining, Copyright and M&A: Legal Risks in Acquiring AI Companies

Every time an AI model generates a text, an image or a melody, a powerful process has taken place behind the scenes: Text and Data Mining (“TDM“), any automated technique aimed at analysing large quantities of texts, sounds, images, data or metadata in digital format to generate information. Billions of web pages, photographs and musical works are ingested by algorithms, often without the authors’ knowledge. As AI companies become attractive acquisition targets, understanding the copyright risks in their training practices is essential for M&A practitioners. This article examines the Italian legal framework governing AI and copyright, from the TDM exceptions under the copyright act to the EU AI Act, and their M&A implications.

1. Regulatory Framework. Directive (EU) 2019/790 on copyright in the Digital Single Market Directive (“DSM“) was transposed into Italian law by Legislative Decree No. 177/2021, which amended Law No. 633/1941 (the “Italian Copyright Law“) by introducing two key TDM provisions: Articles 70-ter and 70-quater.

2. The TDM Exceptions. The two TDM provisions transposed into the Italian Copyright Act establish distinct regimes, each with its own scope, conditions and limitations. The distinction is critical in M&A, as it determines whether a target’s data acquisition practices are lawful and, consequently, the transaction’s risk profile:

2.1. Art. 70-ter LDA: TDM for Scientific Research (Art. 3 DSM). Art. 70-ter introduces a mandatory, non-derogable exception allowing research organisations and cultural heritage institutions to reproduce works or materials to which they have lawful access, for text and data extraction for scientific research. Copies must be stored securely and retained only for research purposes, including verification of results. Rightholders may apply proportionate security measures. Conflicting contractual terms are null and void.

2.2. Art. 70-quater LDA: TDM for Any Purpose (Art. 4 DSM). Art. 70-quater permits reproductions and extractions from works or materials to which the user has lawful access, for TDM purposes without limitation as to identity or purpose, including commercial use. However, TDM is only permitted where the use has not been expressly reserved by the relevant rightholders. Copies may only be retained as long as necessary for the TDM process, with security levels no lower than those under Art. 70-ter. This opt-out mechanism raises several concerns: (i) whether an opt-out can be enforced retroactively against prior scraping; (ii) whether rightholders may reserve only certain works or must cover their entire online corpus; (iii) the technical reliability of machine-readable reservations (robots.txt, metadata); and (iv) the unclear regime for works from physical archives or offline databases. Each open question may give rise to contingent liabilities for M&A investors.

These two TDM exceptions are particularly relevant to generative AI, as training such models typically involves mass reproduction of protected works. Art. 70-ter applies to research organisations for scientific purposes, with no opt-out. For commercial entities, Art. 70-quater applies only if rightholders have not expressly reserved their content. In M&A, identifying the applicable regime is a key question in IP due diligence.

3. AI-Generated Output and Copyright. Neither the DSM Directive nor its Italian transposition addresses the copyrightability of AI-generated output. Art. 1 of the Italian Copyright Act requires human creativity for protection; the prevailing view is that purely AI-generated works are not eligible for copyright, while the status of “AI-assisted works” remains debated. This uncertainty bears directly on M&A valuation: if a target’s core IP consists largely of AI-generated content, the enforceability and value of those assets may be materially lower than assumed.

4. Compliance with the EU AI Act. The EU AI Act (Regulation (EU) 2024/1689) requires providers of general-purpose AI models to publish a summary of training data used, including categories and sources, with particular regard to TDM compliance and any opt-out by rightholders (Art. 53(1)(d)). For M&A investors, non-compliance may expose the target to sanctions, while the required disclosures may reveal underlying IP vulnerabilities.

5. Implications on M&A Deals. The foregoing reflections raise a few red flags when a target develops or deploys AI systems. The lawfulness of the target’s data acquisition and training practices (including compliance with Articles 70-ter and 70-quater of the Italian Copyright Law, opt-out reservations, and the origin of training datasets. Equally, M&A investors should assess the enforceability of IP rights over AI-generated or AI-assisted output, given that purely AI-generated works are unlikely to qualify for copyright under the Italian Copyright Law. Overvaluation of such assets may result in a misalignment between the purchase price and the rights actually acquired. Representations, warranties and indemnities covering TDM compliance, IP ownership and AI Act obligations certainly help, but a rigorous IP due diligence remains essential to identifying risks.

Agreement Reached on the European Copyright Directive

An agreement has been reached on the much discussed European Directive on copyright. http://europa.eu/rapid/press-release_IP-19-528_en.htm. In a race against time to close the dossier by the end of the legislature, in the late evening of February 13, the Parliament, the Commission and the Council of the European Union have finally found an agreement on the copyright directive, which this blog already illustrated https://lawhealthtech.com/2018/09/24/copyright-european-legislation-getting-ready-for-the-digital-era/ .

The vice president of the European Commission immediately tweeted «Europeans will finally have modern copyright rules fit for digital age!». Supporters insist that the new provision will guarantee rights for users, fair remuneration for creators and clarity of rules for platforms. On the other hand, the opposition, stronger than ever before, wants to prevent the imminent change of the internet as we know it.

The highest expectations, placed on the trilogue, concerned the much debated articles 11 and 13, and these have reported to be the outcomes:

  • With regard to the publishers rights, the new version of article 11 sets forth a general need to get a license for the online use of publishers’ press publications, with the only exception for the use of «individual words or very short extracts». According to the Commission, mere hyperlinks and snippets are, therefore, not included in the reform. However, how short should be a “very short extracts” is still to be understood.
  • With regard to the use of protected content by online content sharing services provider, online platforms should obtain a preemptively authorization from the right holders, concluding licensing agreements (where online platform is defined as «a provider of an information society service whose main or one of the main purposes is to store and give the public access to a large amount of works or other subject-matter uploaded by its users which it organizes and promotes for profit-making purposes»). Indeed, an exception has been created for small online platforms, which will not be subject to the abovementioned obligation if they: have been available to the public for less than three years; have an annual turnover below 10 million of euro; and have less than 5 million of visitors.

In the other cases, if no authorization is granted, sharing services providers shall be liable for unauthorized acts of communication unless they demonstrate not only to have made the best efforts to obtain the authorization, but also, in accordance with high industry standards of professional diligence, to have made the best efforts to ensure the unavailability of specific works, as well as to have acted expeditiously to remove the content, after receipt of a notice from the right holders.

We will see if the agreement will survive until the finishing line or if the vote of the European Parliament, scheduled for March-April, will block the text once again, as, unfortunately, already happened.

Copyright European Legislation: Getting Ready for the Digital Era.

On September 12th the European Parliament approved amendments to the controversial Proposal for a Copyright Directive, the Directive of the European Parliament and of the Council on Copyright in the Digital Single Market, which aims at updating copyright rules.

Not many topics have polarized opinions in recent years in Europe. While supporters claim to have protected artists and to have inflicted a blow to the American tech giants, critics have talked about the “death of the internet”.

For clarity, even if the Directive passed the European Parliament vote, the changes are not yet definitive and it may be too early to conclude on what this decision entails. The Directive text shall be further reviewed in subsequent negotiations and there is still a slight chance that it may be rejected at another vote by the European Parliament in 2019. In addition, the Directive, even if (and when) definitely approved, should be implemented by single Member States.

But which results does the Directive aim to achieve?

Its scope and purpose appear based on the evolution of digital technologies, which has changed the way copyright works and other protected material are created, produced, distributed and exploited, with the consequence that new uses, new payers and new business models have emerged. The digital environment has given birth to new opportunities for customers to access copyright-protected content. In this new framework, right-holders face difficulties to be remunerated for the online distribution of their works. So, even if the objectives and principles laid down by the EU copyright framework remain valid, there is an undeniable need to adapt them to the new reality.

The Directive also intends to avoid the risk of fragmentation of rules in the internal market. In fact, the Digital Single Market Strategy1 adopted in May 2015 identified the need «to reduce the differences between national copyright regimes and allow for wider online access to works by users across the EU». The idea expressed in the 2015 by the European Commission was to «move towards a modern, more European copyright framework». The EU legislation purports to harmonize exceptions and limitations to copyright and connected rights, however some of these exceptions, which aim at achieving public policy objectives, such as research or education, remain regulated on national level, with the consequence that legal certainty around cross-border uses is not guaranteed.

As to the content of the Directive, we note the following points:

  • With specific regard to the scientific research, recital number 9 of the Directive says that the Union has already provided certain exceptions and limitations (even if optional and not fully adapted to the use of technology in the scientific research) covering uses for scientific research purposes which may apply to acts of text and data mining. Where researcher have lawful access to content, for example through subscription to publication or open access licenses, the term of the licenses may exclude text and data mining.
  • Article 11, called “link tax”, gives publishers a right to ask for paid licenses when online platforms share their stories. The amended version clarifies that this new rights «shall not prevent legitimate private and non-commercial use of press publications by individual users». The amendment tries also to clarify what can be considered as “sharing a story”, indicating that the mere hyperlinks cannot be taxed, nor can individual words.
  • Article 13, called by the critics as “upload filter”, sets forth that platforms storing and giving access to large amounts of works uploaded by their users shall conclude licensing agreements that include liability for copyright infringement, thus putting a large responsibility on platforms and copyright holders that must «cooperate in good faith» to stop this infringement by carefully monitoring every upload.

The Directive has been designed with the intent to rebalance the core problem of contemporary web: big platforms like Facebook and Google are making huge amounts of money providing access to material made by other people. Nevertheless critics object that this intent could lead to serious collateral effects.

We will see what the future of this Directive will be, and which consequences will entail. The path seems to be still long, but, at least, it has started.