Well the companies and developers don’t decide for every single material. In example what I expect is, that they program the scraper with rules to respect licenses of individual projects (such as on Github probably). And I assume those scraper tools are AI tools themselves, programmed with AI tool assist on top of it. There are multiple AI layers!
At this point, I don’t think that any developer knows exactly what the AI tools are fed with, if they use automatically scraped public sources from the internet.
Well the companies and developers don’t decide for every single material. In example what I expect is, that they program the scraper with rules to respect licenses of individual projects (such as on Github probably). And I assume those scraper tools are AI tools themselves, programmed with AI tool assist on top of it. There are multiple AI layers!
At this point, I don’t think that any developer knows exactly what the AI tools are fed with, if they use automatically scraped public sources from the internet.