Newly unsealed court documents reveal that Microsoft internally described OpenAI’s data scraping methods as theft. Both companies reportedly extracted content behind the New York Times’ paywall to build their datasets, raising concerns about the impact on publishers.

According to TechCrunch, Microsoft privately warned that these practices could severely harm news publishers, even as both firms continued to use paywalled Times content. This disclosure highlights tensions around data use in AI development and the ethical challenges of sourcing training materials.

For Japanese investors and market participants, this story underscores the growing scrutiny on AI firms and data governance, which may influence regulatory approaches and corporate strategies in the region.