zac wolff rv2ooDQuNuI unsplash 1

The Competition to Halt OpenAI’s Scraping Bots Is Slowing Down

OpenAI’s series of licensing deals is already yielding results—at least regarding persuading publishers to be less defensive.

IT’S EARLY to determine how the flurry of agreements between AI firms and publishers will unfold. OpenAI has achieved a significant victory: Its web crawlers are being blocked less frequently by major news organizations than before.

The surge in generative AI initiated a rush for data—and a following rush for data protection (at least for many news websites) where publishers aimed to halt AI crawlers and stop their content from being used as training data without permission. When Apple launched a new AI agent this summer, numerous major news organizations quickly chose to disallow Apple’s web scraping through the Robots Exclusion Protocol, or robots.txt, the file that enables webmasters to manage bots. With the abundance of new AI bots emerging, it sometimes feels like a game of whack-a-mole to stay current.

OpenAI’s GPTBot is the most recognized name and is also blocked more often than rivals such as Google AI. The quantity of prominent media websites employing robots.txt to “disallow” OpenAI’s GPTBot significantly surged from its August 2023 debut until that autumn, then gradually (although more slowly) climbed from November 2023 to April 2024, based on an examination of 1,000 well-known news platforms by Ontario-based AI detection firm Originality AI. At its highest point, the figure was slightly more than one-third of the websites; it has now declined to around one-fourth. In a smaller selection of the most notable news outlets, the block rate remains over 50 percent, though it has decreased from the nearly 90 percent highs seen earlier this year.

However, last May, following Dotdash Meredith’s announcement of a licensing agreement with OpenAI, that figure fell considerably. It subsequently decreased once more at the end of May following Vox’s announcement of its own deal. The shift towards heightened blocking seems to be finished, at least for the time being.

op 1

These declines are clearly reasonable. When businesses form partnerships and allow their data to be utilized, they lose the motivation to protect it, which suggests they would modify their robots.txt files to allow crawling; after securing enough agreements, the total proportion of websites obstructing crawlers is likely to decrease significantly. Certain outlets, such as The Atlantic, unblocked OpenAI’s crawlers on the same day they revealed a deal. Some required a few days to several weeks, such as Vox, which revealed its partnership at the end of May but enabled GPTBot on its sites by late June.

Robots.txt isn’t legally enforceable, but it has traditionally served as the standard regulating the actions of web crawlers. Throughout the majority of the internet’s history, those managing websites anticipated that one another would adhere to the protocol. In an investigation conducted earlier this summer, it was discovered that the AI startup Perplexity was probably disregarding robots.txt commands, prompting Amazon’s cloud division to investigate if Perplexity had breached its regulations. Disregarding robots.txt doesn’t reflect well, which probably clarifies why numerous leading AI firms—such as OpenAI—clearly mention that they rely on it to decide what to crawl. Originality AI’s CEO, Jon Gillham, thinks this increases the urgency for OpenAI to secure agreements. “Gillham states that it’s obvious OpenAI perceives being blocked as a danger to their future goals.”

To date, OpenAI has established agreements with 12 publishers, and although the majority have revised their robots.txt files, there are some exceptions. For instance, Time magazine still prevents GPTBot access. (Time did not provide a comment regarding why GPTBot remains blocked.) Nonetheless, after the agreements are established, it becomes irrelevant, as per OpenAI representative Kayla Wood, since OpenAI no longer interacts with the data in the same manner it does when gathering what it refers to as “publicly available” data. “We utilize direct feeds,” she states.

In the meantime, several prominent media organizations have allowed OpenAI’s web crawler access, even though no partnership announcements have been made. (He monitors how news organizations hinder leading AI bots with slightly varied metrics, and he initially observed the minor drop in block rates a few weeks back.) Alex Jones’ conspiracy-themed site Infowars and the recently revitalized comedic staple The Onion both piqued his interest.

Does this imply that these sites have undisclosed agreements with OpenAI, or are they trying to strike a deal with the organization? “No way,” states Onion CEO Ben Collins, who mentions that the unblocking was probably linked to the outlet transitioning its website to a new hosting provider and content management system last month. “Clearly, we are not engaging in any business with the Plagiarism Machine.”

Infowars did not reply to inquiries for a statement. However, OpenAI has stated that it does not have any collaboration with Infowars.

Although the initial wave to restrict OpenAI’s bots seems to have subsided, it remains uncertain if this quiet period will continue. Gillham believes that there could be more increases in blocking later on if publishers start to regard it as a negotiation strategy. “Is the initial step in negotiating with OpenAI to prevent them?” “Does that lead them to the table?” he asks. Regardless of the outcome, this is a telling moment: While publishers initially reacted to the emergence of AI scraping bots with a collective urge to block them, OpenAI’s proactive pursuit of collaborations has dampened that universal initiative.

Leave a Reply

Your email address will not be published. Required fields are marked *

Back To Top