# OpenAI: Gen AI 'Impossible' Without Copyrighted Material

**<mark>January 9, 2024</mark>**

---

Responding to a U.K. Parliament probe on large language models, **<mark>OpenAI</mark>** said the company **<mark>uses three primary datasets for training data</mark>** for its AI system, including ChatGPT. The three are publicly available information, **<mark>d</mark>**<mark>ata licensed from third parties</mark>, and information that ChatGPT's human trainers provide.

Since copyright right law today covers "virtually every sort of human expression" such as blog posts, photographs and software code, the company said, "**<mark>it would be impossible to train AI models without using copyrighted materials.</mark>**"

The company defended its use of copyrighted material, stating that **<mark>current copyright law does not forbid training data</mark>**. The company said it has put in place provisions to allow content creators to opt out inclusion of their images in OpenAI's DALL∙E training datasets.

Regardless of these measures, there is "still **<mark>work to be done to support and empower creators</mark>**," OpenAI said when probed about the AI copyright issue.

The comment from the company comes as it is **<mark>fighting a copyright lawsuit</mark>** filed by The New York Times against OpenAI and close investor Microsoft.

OpenAI on Monday accused the Times of using **<mark>"intentionally manipulated prompts" to trick its chatbot </mark>** into producing identical materials published by the newspaper and of "cherry-picking" the content produced by the chatbot to back the publisher's copyright claim.

AI startups **<mark>Stability AI and Midjourney</mark>** also face similar copyright lawsuits in the U.K. and in the United States. Last month, the U.K.'s Supreme Court sided with the country's Intellectual Property Office to **<mark>deny patent rights to an AI-generated idea citing limitations </mark>** in the existing statute.
