[ad_1]
Large language models can generate text strings based on learned word patterns from web pages, books, and other bodies of text in their training data. In addition to ChatGPT, the programs form the innards of search chatbots like Microsoft’s Bing Chat and Google’s Bard, and underpin a growing number of applications that produce professional, creative text in a flash. Their AI-composed illustration and video-generating counterparts draw patterns from image datasets such as photos culled from Pinterest and Flickr.
Often, the datasets used in the development of artificial intelligence are created through unofficial means such as software submissions that fetch content from websites. In the US this is generally considered legal, although copyright issues and terms of use of websites against the practice have left it in question.
Some websites like Reddit and Stack Overflow have been more inviting. They offer downloadable data dumps or live data portals to help the software access their content known as an API. In the case of Stack Overflow, LLM developers are getting their hands on the data through a mix of dumping, APIs and scraping, says Chandrasekar, all of which can now be done for free.
But Chandrasekar says the LLM developers are violating Stack Overflows terms of service. Users own the content they post to Stack Overflow, as stated in its TOS, but it all falls under a Creative Commons license that requires anyone who subsequently uses the data to mention where it came from. When AI companies sell their models to customers, they are unable to credit all of the community members whose questions and answers were used to train the model, thus violating the Creative Commons license, says Chandrasekar.
Neither Stack Overflow nor Reddit has released any pricing information. We’re working on it as we speak, says Reddit spokesperson Tim Rathschmidt, and will be sharing more with partners in the coming weeks. Stack Overflow will study Reddit’s strategy and consult with its potential customers, some of whom have already contacted Data Access, says Chandrasekar.
One potential pricing roadmap could come from Elon Musk, who raised prices for access to Twitter data this month. They start at $42,000 a month for access to 50 million tweets. About three times the volume of tweets had previously been freely available. In a tweet this week, Musk accused Microsoft, a leading AI developer and close partner of OpenAI, of illegally training algorithms using Twitter data. Without elaborations, he added, Time to sue.
Both Stack Overflow and Reddit will continue to license data for free to some people and businesses. Chandrasekar says that Stack Overflow only wants remuneration from companies developing LLMs for big commercial purposes. When people start charging for products built on community-created sites like ours, that’s where it’s not fair use, he says.
Reddit CEO Steve Huffman told the New York Times this week that he didn’t want to pay homage to the world’s biggest companies. Crawling Reddit, generating value, and not returning any of that value back to our users is something we have a problem with, he said.
|
Sources 2/ https://www.wired.com/story/stack-overflow-will-charge-ai-giants-for-training-data/ The mention sources can contact us to remove/changing this article |
[ad_2]