본문으로 바로가기
Iamprovider
한국어
EnglishРусскийTurkishPortuguês (Brasil)한국어العربيةEspañolไทยTiếng ViệtFrançais中文DeutschBahasa IndonesiaItaliano日本語PolskiУкраїнськаفارسی
뒤로

참고 사항 및 가이드

Reddit sues AI search engine Perplexity for scraping its data

Reddit sues AI search engine Perplexity for scraping its data

The Lawsuit That Could Redefine AI Data Access

In October 2025, Reddit escalated the simmering conflict over artificial intelligence training data by filing a federal lawsuit against the AI search engine Perplexity and several data-scraping intermediaries. The complaint alleges a coordinated, industrial-scale effort to illegally harvest Reddit's vast repository of human conversations for commercial use without authorization. This legal action isn't just about one company's conduct; it strikes at the heart of how AI systems acquire the human-generated content that fuels their intelligence, setting the stage for a landmark battle over digital ownership in the age of machine learning.

Reddit's position is clear: while it has entered into lucrative licensing agreements with giants like OpenAI and Google—deals reportedly worth around $60 million annually—it will not tolerate unauthorized commercialization of its content. The lawsuit names Perplexity AI, along with Texas-based SerpApi, Lithuania's Oxylabs UAB, and the former Russian botnet AWM Proxy, accusing them of engaging in what Reddit's Chief Legal Officer Ben Lee calls "data laundering." This case immediately frames a critical question: does public visibility of content online grant AI companies a free pass to use it, or does it remain protected intellectual property?

The Alleged Mechanics of Indirect Data Scraping

At the core of Reddit's allegations is a sophisticated workaround to anti-scraping technologies. According to the lawsuit, when direct access to Reddit's servers was blocked, the defendants turned to Google's indexed search results as an alternative data pipeline. They allegedly used web crawlers and bots to extract Reddit content displayed in Google search snippets, effectively "robbing the armored truck instead of the bank vault," as Reddit's lawyers vividly described it. This method allowed them to bypass both Reddit's own protections and Google's tools designed to prevent automated data harvesting.

This indirect scraping technique highlights a growing vulnerability in the digital ecosystem. By targeting the public interfaces of search engines, data scrapers can amass large volumes of content without directly violating a platform's terms of service—or so they argue. Reddit contends that the origin of the data remains its protected asset, regardless of the pathway taken to acquire it. The lawsuit suggests that companies like SerpApi openly advertised their ability to provide Reddit data to clients like Perplexity, creating a shadow market for AI training materials.

Digital Forensics: Reddit's Honeypot Experiment

To substantiate its claims, Reddit's legal team devised a clever digital trap. They created a unique post—a "honeypot"—that was configured to be visible only to Google's web crawlers, not to general users or other bots. Within hours, this exclusive content appeared in Perplexity's search results, providing what Reddit calls "digital proof" of intentional circumvention. This forty-fold increase in Reddit content within Perplexity's system after a cease-and-desist order in May 2024 further bolstered their case.

The honeypot experiment serves as a cornerstone of evidence, demonstrating that Perplexity's data pipeline was sourcing information through Google's cache rather than through legitimate, direct access. Reddit argues this shows willful disregard for its terms of service, which prohibit commercial use without an agreement. This forensic approach mirrors tactics used in cybersecurity to trace data breaches, underscoring the technical sophistication now required in intellectual property litigation. It transforms abstract allegations into tangible, repeatable evidence that a court must weigh.

What the Honeypot Reveals About AI Data Sourcing

Beyond proving access, this test reveals the opaque nature of how some AI companies gather training data. Many users assume that AI models are trained on openly available web content, but Reddit's experiment suggests that intermediaries are actively mining protected channels. This raises ethical concerns about transparency and consent, as the original creators of Reddit posts—millions of users—have no say in how their conversations are repurposed for commercial AI products. The digital breadcrumbs left by the honeypot highlight a supply chain problem in the AI industry.

Data Laundering: The New Digital Economy's Dark Side

Reddit's lawsuit introduces a compelling legal metaphor: "data laundering." This term describes the process where scraping companies allegedly acquire data through illicit means, then sell or transfer it to AI firms, disguising its unlawful origin. Similar to financial laundering, the value lies not just in the transaction but in concealing the source. The complaint alleges that entities like Oxylabs and SerpApi act as brokers in this economy, gathering Reddit content and reselling it to "clients hungry for training material," such as Perplexity.

This framing aims to elevate the issue from mere copyright infringement to a systemic, organized effort to exploit digital assets. By labeling it data laundering, Reddit seeks to draw parallels with established legal statutes that punish the concealment of illicit gains. It underscores the industrial scale of the operation, suggesting that the demand for high-quality human content has fueled a shadow market. This narrative positions Reddit not just as a plaintiff but as a guardian of user rights against a covert data trade.

Perplexity's Defense: Fair Use and the Open Internet

In response to the allegations, Perplexity has mounted a defense centered on principles of fair use and open access. The company claims it does not directly scrape Reddit but instead aggregates publicly available web data, summarizing and citing Reddit discussions in its search results. Perplexity argues that this practice is protected under existing law, as the content is publicly accessible through Google, and that its chatbot merely helps users find information more efficiently. In a Reddit post addressing the lawsuit, Perplexity stated that complying with Reddit's demands would be "the opposite of the open internet."

Perplexity also denies receiving the lawsuit initially and emphasizes its adherence to robots.txt files—the standard protocol for controlling web crawlers. The company contends that it does not train foundation models, distinguishing its use of data from the model training done by OpenAI or Google. This defense hinges on a key legal distinction: whether using publicly displayed data from a third party like Google constitutes infringement. Perplexity's stance reflects a broader debate in tech about the boundaries of innovation and the right to build upon publicly visible information.

The Broader Battlefield: AI vs. Publishers

Reddit's lawsuit against Perplexity is not an isolated incident but part of a widening legal war between content creators and AI developers. Publishers like The New York Times, Encyclopedia Britannica, and others have filed similar suits against AI companies, alleging copyright violations through unauthorized data scraping. Earlier in 2025, Reddit also sued Anthropic for scraping data to train its chatbot, Claude. This pattern indicates a systemic clash over the economics of information in the AI era, where human-generated content is both invaluable and vulnerable.

The outcomes of these cases could establish precedents that reshape how AI companies operate. If courts side with publishers, it may force a shift towards licensed data ecosystems, potentially increasing costs and limiting access for smaller AI startups. Conversely, if fair use arguments prevail, it could accelerate AI innovation but at the potential expense of content creators' rights. This legal friction underscores the unresolved tension between fostering technological progress and protecting intellectual property in a data-driven world.

Licensing Deals and the Future of AI Content Sourcing

Reddit's existing licensing agreements with OpenAI and Google provide a contrasting model to the alleged actions of Perplexity. These deals, with guardrails to protect user rights, demonstrate a pathway for ethical data use that compensates creators. They acknowledge the commercial value of Reddit's data—built over nearly two decades by millions of users—and establish a framework for its legitimate exploitation. In its lawsuit, Reddit argues that Perplexity's refusal to enter such an agreement shows a deliberate choice to circumvent established norms for competitive advantage.

This highlights an emerging dichotomy in AI development: between companies that pay for data through licensing and those that seek to harvest it freely. The case may push the industry towards more transparent and compensated data sourcing, potentially creating a market where high-quality human content is a licensed commodity. However, it also risks creating barriers to entry, favoring well-funded corporations over innovators. The resolution could dictate whether the AI landscape becomes a curated garden or remains a wild frontier of data acquisition.

Innovation at a Crossroads: What This Means for Tomorrow's AI

The Reddit v. Perplexity lawsuit ultimately forces a reckoning with how we balance innovation with integrity in artificial intelligence. As AI systems become more integral to daily life, the sources of their knowledge must be scrutinized. This case isn't just about legal technicalities; it's about defining ethical boundaries in a digital age where data is the new oil. A ruling in Reddit's favor could incentivize more respectful and compensated use of human creativity, while a win for Perplexity might reinforce a more libertarian approach to information access.

Looking ahead, the implications extend beyond these companies to every platform and user generating content online. It prompts questions about digital ownership, user consent, and the sustainability of AI growth if it relies on uncredited labor. Innovative solutions may emerge, such as standardized data licensing protocols or blockchain-based attribution systems. By confronting these issues head-on, the legal battle between Reddit and Perplexity could catalyze a more equitable framework for AI development—one where innovation thrives without exploiting the collective voice of the internet.

I AM PROVIDER 살펴보기

다음 서비스가 여기서 시작됩니다.

플랫폼별로 전체 서비스 목록을 살펴보세요.

Instagram20
구매 인스타그램 팔로워구매 인스타그램 자동 좋아요구매 인스타그램 자동 동영상 조회수구매 인스타그램 댓글구매 인스타그램 참여도구매 인스타그램 노출수구매 인스타그램 좋아요구매 인스타그램 라이브 스트리밍 조회수구매 인스타그램 멤버구매 인스타그램 멘션구매 인스타그램 오래된 게시물 좋아요구매 인스타그램 오래된 게시물 조회수구매 인스타그램 프로필 방문구매 인스타그램 도달구매 인스타그램 저장구매 인스타그램 공유구매 인스타그램 스토리 좋아요구매 인스타그램 스토리 공유구매 인스타그램 스토리 조회수구매 인스타그램 동영상 조회수
TikTok15
구매 틱톡 댓글구매 틱톡 다운로드구매 틱톡 팔로워구매 틱톡 좋아요구매 틱톡 라이브 스트리밍 댓글구매 틱톡 라이브 스트림 좋아요구매 틱톡 라이브 스트리밍 공유구매 틱톡 라이브 스트리밍 조회수구매 틱톡 PK 배틀 포인트구매 틱톡 세이브즈구매 틱톡 공유구매 틱톡 스토리 좋아요구매 틱톡 스토리 공유구매 틱톡 스토리 조회수구매 틱톡 동영상 조회수
YouTube9
구매 유튜브 자동 댓글구매 유튜브 자동 좋아요구매 유튜브 자동 조회수구매 유튜브 댓글 좋아요구매 유튜브 댓글구매 유튜브 좋아요구매 유튜브 공유구매 유튜브 구독자구매 유튜브 동영상 조회수
X (Twitter)14
구매 X (Twitter) 북마크구매 X (Twitter) 댓글구매 X (Twitter) 디테일 클릭구매 X (Twitter) 팔로워구매 X (Twitter) 노출수구매 X (Twitter) 좋아요구매 X (Twitter) 링크 클릭구매 X (Twitter) 멘션구매 X (Twitter) 설문 투표구매 X (Twitter) 게시물 조회수구매 X (Twitter) 프로필 클릭구매 X (Twitter) 리포스트구매 X (Twitter) 스페이스 청취자들구매 X (Twitter) 비디오 조회수
Facebook12
구매 페이스북 댓글구매 페이스북 팔로워구매 Facebook 친구 요청구매 페이스북 좋아요구매 페이스북 라이브 스트리밍 조회수구매 페이스북 페이지 팔로우구매 페이스북 페이지 좋아요구매 페이스북 게시물 좋아요구매 페이스북 게시물 반응구매 Facebook 프로필 팔로우구매 Facebook 스토리 조회수구매 페이스북 동영상 조회수
Telegram3
구매 텔레그램 멤버구매 텔레그램 게시물 조회수구매 텔레그램 리액션
Twitch6
구매 트위치 채널 조회수구매 Twitch 클립 조회수구매 트위치 팔로워구매 트위치 라이브 스트리밍 조회수구매 트위치 프라임 구독자구매 트위치 동영상 조회수
Spotify4
구매 스포티파이 팔로워구매 Spotify 청취자구매 스포티파이 재생 횟수구매 Spotify 저장 항목
SoundCloud5
구매 사운드클라우드 댓글구매 사운드클라우드 팔로워구매 SoundCloud 좋아요구매 사운드클라우드 재생구매 SoundCloud 리포스트
Threads5
구매 스레드 댓글구매 스레드 팔로워구매 스레드 좋아요구매 스레드 리포스트구매 스레드 공유
Likee4
구매 Likee 댓글구매 Likee 팔로워구매 Likee 좋아요구매 Likee 공유
Discord1
구매 디스코드 회원
Reddit6
구매 레딧 댓글구매 레딧 팔로워구매 레딧 게시물 조회수구매 레딧 공유구매 레딧 구독구매 레딧 추천
Kick2
구매 Kick 팔로워구매 Kick 라이브 스트리밍 조회수
LinkedIn3
구매 링크드인 팔로워구매 LinkedIn 좋아요구매 링크드인 동영상 조회수
Pinterest5
구매 Pinterest 팔로워구매 Pinterest 좋아요구매 핀터레스트(Pinterest) 리핀(Repin)구매 핀터레스트 저장구매 Pinterest 공유
OnlyFans3
구매 OnlyFans 댓글구매 OnlyFans 팔로워구매 OnlyFans 좋아요
Snapchat4
구매 스냅챗 참여도구매 스냅챗 팔로워구매 스냅챗 스포트라이트 조회수구매 Snapchat 스토리 조회수
Kwai5
구매 Kwai 댓글구매 Kwai 팔로워구매 Kwai 좋아요구매 Kwai 공유구매 Kwai 동영상 다운로드
Tumblr3
구매 Tumblr 팔로워구매 Tumblr 좋아요구매 텀블러(Tumblr) 리블로그
Shopee1
구매 Shopee 생방송 시청 수
Trustpilot1
구매 Trustpilot 리뷰
Quora5
구매 Quora 답변 조회수구매 Quora 팔로워구매 Quora 공유구매 Quora 추천 수구매 Quora 조회수
Yandex Zen5
구매 Yandex Zen 댓글구매 Yandex Zen 좋아요구매 Yandex Zen 조회수구매 Yandex Zen 구독구매 Yandex Zen 동영상 조회수
WhatsApp2
구매 WhatsApp 채널 멤버구매 WhatsApp 채널 게시물 반응
Vimeo2
구매 Vimeo 팔로워구매 Vimeo 좋아요
추가 서비스1
구매 Yandex Music 재생 횟수

다음 게시물을 위한 도구

텍스트를 준비하고, 콘텐츠를 계획하고, 다음 캠페인을 준비하세요.

무료 도구 살펴보기

무료 서비스 신청

캠페인이 진행 중일 때 수동 검토를 위해 최대 100개의 단위를 요청하세요. 승인 및 배송이 보장되지 않습니다.

모든 무료 서비스 요청 찾아보기

Instagram

TikTok

Facebook

YouTube

X (Twitter)

Spotify

Telegram

Twitch

Brawl Stars

PUBG Mobile

도움을 드립니다

문의하기

팀에 연락할 방법 선택.

다음 페이지 찾기

Tab: 선택 · Enter: 열기 · Esc: 닫기