SoylentNews
SoylentNews is people
https://soylentnews.org/

Title    Cloudflare is Taking a Stand Against AI Website Scrapers
Date    Monday July 08, @06:31PM
Author    hubie
Topic   
from the dept.
https://soylentnews.org/article.pl?sid=24/07/07/1334210

Arthur T Knackerbracket has processed the following story:

Cloudflare has released a new free tool that prevents AI companies' bots from scraping its clients' websites for content to train large language models. The cloud service provider is making this tool available to its entire customer base, including those on free plans. "This feature will automatically be updated over time as we see new fingerprints of offending bots we identify as widely scraping the web for model training," the company said.

In a blog post announcing this update, Cloudflare's team also shared some data about how its clients are responding to the boom of bots that scrape content to train generative AI models. According to the company's internal data, 85.2 percent of customers have chosen to block even the AI bots that properly identify themselves from accessing their sites.

[...] It's proving very difficult to fully and consistently block AI bots from accessing content. The arms race to build models faster has led to instances of companies skirting or outright breaking the existing rules around blocking scrapers. Perplexity AI was recently accused of scraping websites without the required permissions. But having a backend company at the scale of Cloudflare getting serious about trying to put the kibosh on this behavior could lead to some results.

"We fear that some AI companies intent on circumventing rules to access content will persistently adapt to evade bot detection," the company said. "We will continue to keep watch and add more bot blocks to our AI Scrapers and Crawlers rule and evolve our machine learning models to help keep the Internet a place where content creators can thrive and keep full control over which models their content is used to train or run inference on."


Original Submission

Links

  1. "following story" - https://www.engadget.com/cloudflare-is-taking-a-stand-against-ai-website-scrapers-220030471.html?src=rss
  2. "a blog post" - https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click/?utm_campaign=cf_blogutm_content=20240703utm_medium=organic_socialutm_source=twitter
  3. "Perplexity AI was recently accused" - https://www.engadget.com/amazon-investigating-perplexity-ai-after-accusations-it-scrapes-websites-without-consent-133003374.html
  4. "Original Submission" - https://soylentnews.org/submit.pl?op=viewsub&subid=63204

© Copyright 2024 - SoylentNews, All Rights Reserved

printed from SoylentNews, Cloudflare is Taking a Stand Against AI Website Scrapers on 2024-07-18 20:12:22