News
AI Summary
5 Jun 202620 Dhuʻl-Hijjah 1447 AH
Microsoft trained its MAI models on unlicensed web data despite promising "enterprise grade, clean and commercially licensed data"

Microsoft trained its MAI models on unlicensed web data despite promising "enterprise grade, clean and commercially licensed data"

Microsoft has revealed that its MAI models were trained on unlicensed data, including Common Crawl, despite claims of using only "clean and commercially licensed data." Like many AI labs, Microsoft relies on fair use, placing the onus on website owners to block its crawlers. This raises concerns about the integrity of Microsoft's data sourcing practices. The situation highlights broader challenges facing companies in adhering to ethical standards in data usage while developing AI models, potentially impacting trust in their offerings.

Follow these topics

Sign in to follow the topics that matter to you

Sign in to follow

This summary is generated with AI and receives periodic editorial review. Refer to the original source for full details.

0
0 reading now

Insight Score

Rate to unlock

Sign in to react, rate, and save. Sign In