Collecting publicly available pages generally isn't a crime under the CFAA, at least in the Ninth Circuit, after Van Buren (2021) and the Ninth Circuit's 2022 hiQ v. LinkedIn ruling. Contract, copyright, and privacy law can still apply depending on how you access the data and what you do with it.
Myth Busting: Scraping and Hacking Aren't the Same Thing
"Isn't that basically hacking?" It's one of the most common questions anyone doing web scraping at scale gets asked. The honest answer is no. Hacking means getting into a system you aren't allowed into. Fetching a public web page means reading something the site already shows to anyone with a browser. Over the past five years, the US Supreme Court and the Ninth Circuit have drawn that line in real cases. But "not hacking" doesn't mean "no rules," and the same cases show where the real limits sit.
Key Takeaways
- In 2021, the US Supreme Court read the federal anti-hacking law (the CFAA) as a "gates-up-or-down" question: are you allowed into this part of the system or not (Van Buren v. United States, 2021).
- In 2022, the Ninth Circuit held that pages open to the public have "erected no gates," so reading them likely isn't access "without authorization" (hiQ v. LinkedIn, 2022).
- Contracts, fake accounts, copyright, and privacy law still apply. Those are where web data cases are actually won and lost.
This post explains public court decisions for a general audience. It isn't legal advice, and the law differs by country.
Why do people think scraping is hacking?
The confusion is mostly vocabulary. Both involve automated scripts, both happen at scale, and both show up in headlines about "bots." The US Computer Fraud and Abuse Act (CFAA), passed in 1986, made it a crime to access a computer "without authorization" or to "exceed authorized access." For years, sites argued that breaking their terms of service counted as exceeding authorized access.
If that reading had held, a lot of ordinary behavior would be a federal crime. Using a work computer to check sports scores. Making a second account against a site's rules. Sending an automated request to a page the site said you shouldn't automate. Courts were split for years on how far the statute reached.
So the myth isn't crazy. It came from a real legal fight, and one that has narrowed a lot since 2021, even though some questions are still open and the key appeals ruling binds only the Ninth Circuit.
What did the Supreme Court decide in Van Buren?
In June 2021, in Van Buren v. United States, the Supreme Court ruled 6-3 that "exceeds authorized access" covers entering parts of a computer system that are off-limits to you, not misusing information you were already allowed to see (Supreme Court opinion, 2021). The Court described this as a "gates-up-or-down" inquiry.
The framing does most of the work. The question isn't "did you break a rule the site wrote?" It's "was the gate open to you?" A login wall is a gate. A password is a gate. A page anyone can load in a browser isn't.
Notice what this means in practice. Since Van Buren, courts have treated authentication, a login or a password, as the clearest example of a closed gate. The Supreme Court did leave open, in a footnote, whether contract terms or policies alone can close one. Still, "is there a login between me and this data?" is a question most engineering teams can actually answer, which is more than you can say for most terms of service.
hiQ v. LinkedIn: public pages have no gate
hiQ Labs collected public LinkedIn profiles to build workforce analytics, and LinkedIn sent a cease-and-desist letter. After Van Buren, the Supreme Court sent the case back down. In April 2022, the Ninth Circuit held that when a site makes data publicly available, "that computer has erected no gates to lift or lower in the first place" (Ninth Circuit opinion, 2022).
In plain terms, reading public pages, even with automated tools and even after being told to stop, was unlikely to be a CFAA violation. It's the ruling people usually quote.
It's also where the story usually gets cut short.
Why hiQ still paid LinkedIn: contract and conduct
Because hiQ didn't lose on hacking. It lost on contract and conduct. In November 2022, the district court found hiQ had breached LinkedIn's User Agreement, partly because contractors working for hiQ created fake LinkedIn profiles. In December 2022, the parties filed a proposed consent judgment that included a $500,000 judgment against hiQ and required it to delete the LinkedIn data it held (district court docket, 2022; National Law Review, 2022; Morgan Lewis, 2022).
Here's the lesson most summaries miss. The fake accounts, and the fact that hiQ had itself agreed to LinkedIn's User Agreement, were the problem. Once you hold an account, you've accepted terms. Create one under a false identity and you've walked through a gate with a key you made up.
Does it matter whether you're logged in?
Yes, and a 2024 case made that concrete. In January 2024, in Meta v. Bright Data, Judge Edward Chen of the Northern District of California granted summary judgment to Bright Data. He found that Facebook's and Instagram's terms didn't prohibit collecting public data while logged out (court order, 2024; Quinn Emanuel client alert, 2024).
The court reasoned that a logged-out visitor isn't acting as a "user" bound by those terms. It also rejected Meta's reading of its terms as barring public data collection forever, even after an account had been closed.
Put the three cases side by side and a simple pattern shows up:
So what are the real rules for pulling public data?
Hacking law is mostly the wrong lens. The rules that actually shape web data work are contracts (did you agree to terms?), copyright (what are you doing with the content?), and privacy law like the EU's GDPR (does the data identify people?). Collecting personal data from a public page can still trigger privacy obligations, even when nothing about the access was unlawful.
A workable checklist for teams:
- Stay logged out unless you have a clear right to use an account for this purpose.
- Never create fake accounts to reach data. That conduct was central to hiQ's loss.
- Don't get around technical barriers. A login wall or password is a closed gate.
- Check what you're collecting. Personal data brings privacy obligations with it.
- Check what you're doing with it. Republishing copyrighted content raises separate questions from reading it.
- Read robots.txt and keep request volume reasonable. robots.txt isn't a CFAA gate, but site owners treat ignoring it or hammering a site as a sign of bad faith. Don't degrade the site you're reading.
The teams that stay out of trouble tend to write their rules down before they build anything: which data they collect, why, and what they will never do to get it. The better ones go a step further and put those rules where a script can't ignore them. A "never touch" list that lives in the proxy layer, for example, stops a request outright, instead of relying on every engineer remembering a wiki page.
For how these questions play out in AI training disputes specifically, see the AI training data lawsuits every web data team should know. For another myth this series has taken on, read why more IPs don't make a better proxy network.
The network you use is part of compliance too
How you reach public pages is part of your compliance story too. A residential network built from hijacked devices adds a problem you didn't create. Massive's network is made of real consumer devices whose owners opted in through the Massive SDK, and Massive has completed a SOC 2 Type I audit, with Type 2 in progress, both listed on the Massive Trust Center. For the questions to ask any vendor, see SOC 2 and GDPR: what to ask a residential proxy vendor.
Access still has limits. Massive blocks certain categories of content at the network layer for all accounts, and each Residential account can add up to 1,000 domains to its own domain blocklist. A blocked request comes back as 452 Disallowed Content rather than going through. The details of how opt-in works are in how Massive's opt-in network works.
The bottom line
- Hacking means passing a closed gate. Scraping a public page doesn't.
- Van Buren (2021) set the gates test, and the Ninth Circuit applied it to public pages in hiQ (2022).
- hiQ still lost, on fake accounts and contract terms.
- Meta v. Bright Data (2024) showed logged-out collection sits in a different place than logged-in use.
- The real rules are contract, copyright, and privacy. Plan for those.
Sources
- Supreme Court of the United States, Van Buren v. United States, No. 19-783 (2021)
- US Court of Appeals for the Ninth Circuit, hiQ Labs, Inc. v. LinkedIn Corp., No. 17-16783 (April 18, 2022)
- National Law Review, "hiQ and LinkedIn Reach Proposed Settlement in Landmark Scraping Case" (2022)
- Morgan Lewis, "LinkedIn v. hiQ: Landmark Data Scraping Suit Provides Guidance" (2022)
- US District Court for the Northern District of California, hiQ Labs, Inc. v. LinkedIn Corp., No. 3:17-cv-03301-EMC, order on cross-motions for summary judgment (Nov. 4, 2022)
- US District Court for the Northern District of California, Meta Platforms, Inc. v. Bright Data Ltd., No. 3:23-cv-00077-EMC, order granting Bright Data's motion for summary judgment (Jan. 23, 2024)
- Quinn Emanuel, "Meta v. Bright Data: Significant Decision for Web Scraping Industry" (2024)
Frequently Asked Questions
It's the Supreme Court's 2021 test from Van Buren. You exceed authorized access under the CFAA by entering parts of a system that are closed to you, like areas behind a login you aren't entitled to. Misusing information you could already see doesn't count.
In November 2022, a court found hiQ breached LinkedIn's User Agreement, including through fake profiles created by its contractors. The December 2022 proposed consent judgment included $500,000 against hiQ. The loss came from contract and conduct, not from reading public pages.
It can. GDPR covers personal data whether or not it's public, and it can reach companies outside the EU that monitor people in the EU. Teams collecting names, profiles, or contact details need a lawful basis and a plan for data minimization.
