Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Unpopular opinion: when you make a HTTP request you're asking the server to give you information. The server has the right to say no.

IMHO, LinkedIn doesn't have a right to stop scraping after the fact, but they have the right to take technical steps to stop scrapers from accessing their site.



LinkedIn takes plenty of technical precautions to block scraping. I’ve built bots that scrape them in the past, it’s surprisingly difficult as LinkedIn is very good at determining you’re a bot and blocking you. So it’s hard to argue that a service which is scraping LinkedIn is doing it without knowledge that they are going against LinkedIn’s wishes. Whether or not this is illegal is up to the courts to determine, and I really hope they decide it is fine (although I have zero faith that’s how the ruling will come down).

That being said, I hate LinkedIn as a company and I fully support anyone trying to mess with them. They are not a social network, they are a sleazy website that convinces people to willingly provide personal information which they then turn around and sell at ridiculously high prices. Even if you are legitimately using LinkedIn as an end-user, it’s easy to get blocked for using it too much and being forced to pay just to interact with people on the site.


> So it’s hard to argue that a service which is scraping LinkedIn is doing it without knowledge that they are going against LinkedIn’s wishes.

As far as I can tell, no one has made that argument, so I'm not sure why you feel the need to rebut it.

I think it all pretty much boils down to this quote from the article:

> LinkedIn's position disturbs Orin Kerr, a legal scholar at George Washington University. "You can't publish to the world and then say 'no, you can't look at it,'" Kerr told Ars.


> no one has made that argument, so I'm not sure why you feel the need to rebut it.

The title is "It’s illegal to scrape our website without permission". So that argument is implied in the headline, at least.

As for Orin Kerr, I'm sure he'd agree that there are private parts of the internet (my payment information being an obvious example). Just because something is deployed to the internet doesn't mean it is "published to the world" as he claims.


That is not Kerr's argument. A key point (which he has developed to a very detailed degree, as both a law professor and actively defending people, such as Weev) is the access controls in place.

He's not just bloviating; before you disagree with him, reviewing his arguments is worthwhile. (I do disagree with him in part, and agree with his reasoning but don't like the outcomes in part, but in any case, he's a pretty accomplished lawyer, and I'm not any kind of lawyer, so there's that.)


"Published to the world" is a bit off; I'd compare it more to recording telephone conversations with corporations' public service numbers. They always tell you that you are being recorded (for "training and quality purposes") so they give up their right to not be recorded in turn.

If a company paid a million people to call into an information service hotline and each request one fact from it—and then the company recorded and compiled the answers into their own database to start their own information service—is that illegal?


>If a company paid a million people to call into an information service hotline and each request one fact from it—and then the company recorded and compiled the answers into their own database to start their own information service—is that illegal?

I sure hope not. From there it doesn't seem to far off to claim that "He was really only ever showing up for work to learn some skills and how he's using those skills to run his own business!" sort of lawsuits. If you compile difficult to find but freely available information into a more easy to digest format I see that as virtually always a net positive.


I don't see anything that connects the title (written by LinkedIn) to the motivation of the people that are scraping it (not LinkedIn).

Everyone that scrapes LinkedIn (or anywhere else) either knows that they are doing it against LinkedIn's wishes or doesn't care.


> a sleazy website that convinces people to willingly provide personal information which they then turn around and sell at ridiculously high prices

I think you just defined a social network, unfortunately.


Don't the other social networks throttle http requests from certain dubious ip addresses? I think they all do this.


* That being said, I hate LinkedIn as a company and I fully support anyone trying to mess with them

So it's fine to mess with them, even illegally, just because you don't like them?

* convinces people to willingly provide personal information

Convincing is not forcing, and in fact, you say "willingly" yourself. Any business convinces you to willingly give them money or other value.

* sell at ridiculously high prices

To respect to what? Prices are determined by the market. We are not talking about a life-saving medicine or health care, on which there could be some debate.

We are talking about a company selling information, which has value and which it acquires through an infrastructure that takes a lot of money to run. Value that customers are willing to pay for.

* Even if you are legitimately using LinkedIn as an end-user, it’s easy to get blocked

For what definition of legitimately? Yours? Since it's their business, they can define what is a legitimate free use and what can be a paid one.

* They are not a social network

So based on what I pointed out, they are not a social network only because they are not free? Do all social network have to be free and give your data for advertisers to be social networks?


>So it's fine to mess with them, even illegally, just because you don't like them?

The previous commenter never mentioned messing with them illegally and disagrees with the analysis that this scraping is illegal. He said he hopes the courts do not rule this is illegal. Pretty disingenuous to start a reply like that...

Many of your other points are not really giving the previous commenter any charity at all.

For example,

>For what definition of legitimately? Yours? Since it's their business, they can define what is a legitimate free use and what can be a paid one.

He was literally trying to provide an example of where their scraping protections can be appear to be overzealous to casual users, not arguing whether those users are legitimate or not.

What is the premise of your argument? To me, it seems you are simply trying to defend LinkedIn's business practices and legal pursuits, rather than discussing anything about the legality of scraping or the specifics of LinkedIn's anti-scraping implementations.


> So it's fine to mess with them, even illegally, just because you don't like them?

If it's not illegal, then scummy behaviour against a scummy company doesn't exactly set my moral compass off. You reap what you sow.


An eye for an eye, a tooth for a tooth.

Very good. That way the whole world will be blind and toothless. --Tevye, Fiddler on the Roof

If you allow yourself to act as badly as others act, you don't have much of a moral compass.

The world can only improve when we hold ourselves to a higher standard than we see, as it's too easy to rationalize our own behavior, and harshly judge others' actions.


As a sheep, you can take the moral high ground against a wolf, but you'll still get eaten. Companies generally don't care about holding themselves to higher standards of behaviour, unless it hurts their bottom line. Institutions don't understand shame, right, or wrong. The only thing they understand is power.


As a principled human being, you can choose your actions.

The defense against a wolf is (generally) not eating the wolf.

The defense against invasion of privacy is not more privacy invasion.


>moral compass

The issue I see with this is in 2 part.

First, any issue that comes down to "moral compass" is inherently dangerous. We can find many examples of the simple concept that what to one person is Good is to another Evil. In this case, I think Linkedin shareholders would not appreciate calls to mess with the site, or with people having trouble jobhunting because the site is going down repeatedly due to DDOS or whatever.

The second is that these kinds of calls to action (linkedin sucks, fuck with it) smell like vigilantism to me, and while Batman is my favorite hero (I really only lift because I kinda sorta wanna be batman), vigilantism doesn't contribute to a stable society. Rule of law works better than the chaos of multiple agents enforcing their own moral code as law.

EDIT: I'm happy to be downvoted if I'm saying something stupid, but while doing so I would very much appreciate a quick comment as to why I'm wrong so I can improve my knowledge.


everything comes down to a moral compass of some kind. Your comment expects some sort of objective measurement of good and evil, but I don't see any. The law can be, and is often, in the wrong.

Most people seem to think this law (if it is held up in court) is wrong and should be changed.

You might say, oh well, we have a democratic right to change or influence our laws. But a princeton study has found no correlation between public preferences of the majority of the population and enacted policy: https://scholar.princeton.edu/sites/default/files/mgilens/fi...

From memory, the only thing that leads populations to revolt against their government is high enough food prices. Outside of that, revolts almost never happen.

What I'm trying to say is A) unjust, monopolistic, or excessive laws are probably more normal than the opposite because B) the idea that democracy means people have some weigh in lawmaking might be a myth and C) most people don't do anything about it because they only act when the very basics of their livelihood are threatened.

Your view, that some unjust or excessive laws are preferable to total chaos, seems to carry the assumption that laws are naturally benign and/or made to serve some purpose for society, therefore we should not challenge them without good reasons to do so. If the opposite is true and most laws or a high enough number of them are not just, then the fact that the vast majority of people disagrees with them is only natural.

This is a fairly long-winded way of saying that most people would say you're being downvoted because there are plenty of terrible laws that we should not acquiesce to silently.


Thank you for taking the time to reply, this is good to chew on.


thank you for reading. There are studies that contradict the one I posted, in one way or another, so I'm saying these things for sure


I would replace the word "convinces" with "deceives". If you deceive a person into, for example, giving them access and permission to use your personal contact list however they please, then proceed to do just that, you don't have anything to complain about.


I'm lucky my dad uses a different password for his gmail account than his linkedin login (using said gmail address). He was stuck for a while complaining that he couldn't log in. Turned out he was already logged in, and LinkedIn was just presenting him with a "put your gmail password in here so we can raid your contacts" box. It looks just like the login page, so he kept putting in his LinkedIn (not gmail) password and it kept saying "your password is incorrect."

What a shitty thing to do.


All his contacts are lucky too.


Agree. The number of times I've received LinkedIn emails from people I no longer speak to who I'm confident would have no actual inclination to connect with me is certainly in the double digits. They've all been conned into giving LinkedIn their email password, and LinkedIn is going crazy as result.

This probably happened a lot more a few years ago. Is perhaps 2FA making this harder these days?


Are they still doing that? I haven't seen a request from LinkedIn for email access credentials in a while. Still a pretty sketchy thing to do IMO.

I wonder if Google, Yahoo, MS etc have done anything like watch for requests from LinkedIn with correct credentials, block them, and reset the user's password and give them a warning that they just gave their account password to a third party and this is a Very Bad Idea.


> Are they still doing that? I haven't seen a request from LinkedIn for email access credentials in a while.

Literally yesterday I got, for the first time for that user, one of the spams sent "on behalf of" a user who clearly hadn't given out my e-mail address, so I guess they still are up to these no-good deeds.


A whole bunch of other sites ask for your Google, etc credentials instead of using the SSO API. I've even seen some tax software ask for your Bank login. It's really a bad practice but LinkedIn isn't the worst offender.


Speaking of offenders, I just noticed that the seemingly popular Venmo payments app asks for the web login credentials to any bank you try to link it to. Hell no I'm not giving some random app login credentials to any of my bank accounts.


Yes, this is maddening. I had a Craigslister who would only pay for a transaction via Venmo. (She was willing to make the transaction in my presence and wait for it to confirm; it wasn't a scam on her part.) With great reluctance (I needed the transaction at the time), I changed my bank password to something else, signed up for Venmo, got the money, de-enrolled from Venmo, and changed my bank password back.

I think that they used to have some FAQ entry explaining why worrying about this is silly and nothing bad could happen, but I can't find it any more (probably because it's nonsense). However, just because they should be shamed for this whenever possible, here's a Slate article on their overall security: http://www.slate.com/articles/technology/safety_net/2015/02/... .


It's a few years ago now .. but it definitely happened.

Agree with you that giving your account password to any third party is madness, even more so actually soliciting it.


There was a class action suit about this a few years ago. Completely inappropriate. Though I don't think it should have any bearing on whether or not it's OK to scrape their site.


This might be one of the best replies on HN that I have seen. Spot on.


I don't think you've characterized this accurately. When you make an HTTP request to LinkedIn you are accessing their service. There is a long history of this relationship, you plug your house into the sewer line and you connect to the sewer service. You connect to the power pole and connect to the electricity service. You connect to the telephone pole and connect to the telephone service.

Every service has "terms of service" which are the conditions that you are allowed to access the service and what you may do with the service once being granted access. For example, if you start pouring toxic waste into your sewer, you will find that the city will both disconnect you from the service and they will fine you for violating the terms of service you nominally agreed to when being hooked up.

In LinkedIn's case, they allow you to access their service, with HTTP, to render a page in a browser for viewing of that page. Full stop. Any other use of the data you acquire over HTTP, or any other method of acquiring said data over HTTP is disallowed by the terms of service.

Not only does LinkedIn have a legal right to stop scraping after the fact, they have literally centuries of common law in support of that position.


(I am not a lawyer.) As far as I understand the legal precedents involved, random terms of services for websites are not effective in this scenario as the public profiles do not require having any account or other relationship. This actually went to court, and because Zappos didn't force users to click through a terms of service to access their service, the terms of service was invalid.

As for their ability to control what you do with the information: there might be a limited license on the data granted from users to LinkedIn that is not transferrable, so maybe you couldn't build a service that redistributed that information, but I don't see why obtaining and holding it would be illegal.

As for the analogies to power and telephone and such, those are built on property owned by a local government and there are usually other extra laws related to them: it isn't due to some common law position that you can't mess with their stuff. Here, I am not a lawyer, but I am a government official with a particular interest in sewage; here is a link to the sewer use ordinances form our local sanitation district: pay particular attention to 2.03.

http://goletawest.org/wp-content/uploads/2012/04/Ordinance-N...


I worked at Google for four years, an independent search engine for 5 more after that, and at IBM after it acquired said search engine for 18 months after that. Everyone of those organizations spent many thousands of dollars on legal fees over just this question and reviewed tons of case law.

Every single one of them concluded that based on how the law was written and how the web worked, there is no legal way to scrape a web site without its explicit permission to do so.

That won't stop people from trying of course and it was a source of constant entertainment in the ops team at Blekko at how people tried to sneak around at scraping (it can get very creative) but; it isn't legal, you can and will get banned from all access for it, and if you use the results in another product or offering you will be found liable for damages.


> there is no legal way to scrape a web site without its explicit permission to do so.

Google scrapes several of my sites and I've never given Google explicit permission to do so.


If your robots.txt file is /allow then you did. If you have no robots.txt file then it's an open question. If you put a /deny into your robots.txt file Google will stop scraping your site.

The implicit contract is that you let them scrape because you want to show up in their search results which will send you traffic. If you don't care about Google traffic then set /deny in your robots.txt and get back the bandwidth you were giving them.


> If you have no robots.txt file then it's an open question.

Only for definitions of explicit I must be unfamiliar with.

If the presence of a robots.txt makes one's intent for a given resource explicit one way or the other, the lack of one (and the lack of some communication in some other channel) must mean there is no explicit permission.


That is correct, for what it was worth IBM's legal team came down on the side of 'assume deny' and Google was (at the time I was there) 'assume allow.'


I think "assume allow" is perfectly reasonable. It's just implicit, not explicit.


To the extent to which that is the case, though, it isn't due to the terms of service; and that is also a case of how you are using the data for later, which is a separate question from the scraping and collection process: it is very clear to me that a search engine is operating on the legal equivalent of thin ice, particularly with details like snippets and synthesis ;P. Whether the CFAA applies (as indicated in this article) is an open question, but that just isn't quite so obvious as "you also can't connect up to the public sewer".


   > it is very clear to me that a search engine is 
   > operating on the legal equivalent of thin ice, 
We may be saying similar things but from a metaphor I think of search engines operating on 'thick' ice. It has been litigated so much that there is a bevy of case law to refer to at all levels. Eric Goldman's blog used to have a pretty good list of the number of suits of various kind and the searchengine blog covered many of them as well.

For a search engine it is super clear, robots.txt is all. If you say yes explicitly, great. If you say no explicitly, that has to be honored. If you say nothing, then its up to the search engine to decide which way to interpret it, but if the site owner complains because you picked wrong you have to honor their wishes (which may include destroying any cached data as well).

PadMapper, Perfect10, and the newspapers generated a ton of cases based on 'scraping a web site and using the data.' There are also about a dozen comparative shopping sites that have been dinged for the exact same issues. (look vs Amazon or vs Walmart).

Whether CFAA, DMCA, Torte law (contracts), or something else applies is constantly being discussed :-). I'm just the messenger here. I haven't found a single case that has held that the point of view of the scraper of someone else's web site should prevail. The argument that it should be allowed 'to help new businesses get off the ground' is like saying Apple should pay out some of its cash hoard as grants to startups trying to break into some business. I have yet to read anything that was sympathetic to that point of view.


yuummm Torte law. (It's tort, and tort law is generally considered to be distinct from contract law, because in tort rights and duties come from common law whereas in contract law they come from acts of agreement between two parties).


Chocolate Torte is my favorite :-) Thanks for the clarification, in the various articles I've read over the years on this topic they refer to tort law (no doubt because much of the argument references common law and the way in which the relations are argued) and I made the leap to 'contracts' which was incorrect.


The scenario is a bit more nuanced though and creeps into Internet freedom.

- Reminds me of CraigsList vs PadMapper[1]. In that scenario I side with CL -- it was right to block PM. PM or others should not be allowed to build a new UI on top of CL because CL was the one that put in years of effort of nurturing its listings, its network, building brand equity and taking associated risks and costs.

- As others have highlighted, the data is publicly accessibly and there is no agreement the scraper/crawler is bound by. The agreement is between the LinkedIn user and LinkedIn. The scraper is connected to the Internet pipe crawling the Internet freely as it wants. It's not reproducing the data anywhere so copyright should not be an issue.

- What if a scraper didn't scrape LinkedIn but just the Google or Archive.org cached versions and read those instead? It would not be pressuring LinkedIn server resources in this case.

- What if all of my employees allow me to scrape their LinkedIn data? Can I scrape all of their info? Can LinkedIn stop me from doing that (In the case of Facebook vs Power Ventures, the answer is that LinkedIn would be able to prevent this behaviour).

- Who owns the data? Medium.com doesn't own the posts. LinkedIn doesn't own the CVs.

  [1]: https://news.ycombinator.com/item?id=4286325


Now go read the 3Taps vs Craigslist cases (https://en.wikipedia.org/wiki/Craigslist_Inc._v._3Taps_Inc.) to start. To be clear here, I feel like I understand your argument that facts aren't copyrightable or protectable and that how you got them is not relevant.

I'm just saying the legal system doesn't see it that way, they have said so in many cases, and so far everyone who has used your argument or variations of it in court has failed to prevail.


When you make a connection to the city sewers or to the power company, there is some kind of pre-connection step where the terms are presented and you agree to those terms.

With HTTP and LinkedIn, there is no such step. There's no pre-connection agreement. LinkedIn could present such an agreement on first connection, but they do not.


That argument has been tried in a variety of ways and been shot down in court repeatedly. (there are parallels to tenants not agreeing to the terms of their internet connection where the landlord provided it).

LinkedIn has two things that they do which protect them; First, they specify they disallow access in their robots.txt file. While not a binding agreement per se it is the default mechanism that is accepted by the community for apriori identifying whether or not automated access is possible. Second, when they detect an access pattern that violates their terms of service they actively block the access proactively notify the source of the violation.

The sad truth is that web scraping has been around since the very beginnings of the Web back in 1993 and this question has been litigated in every way that you might choose to argue it, the body of case law is enough to fill at least two volumes in the reference section of the library.

There is no legal or ethical basis for scraping the web without permission. And if it isn't explicitly allowed by a site the presumption is that it is disallowed (no 'open door' exception).


When you make your conenction to Linkedin, What user agent would you provide? One that's blank, another that's "lelandbatey bot" or one that says, "Mozilla/5.0 (iPad; U; CPU OS 3_2_1 like Mac OS X; en-us) AppleWebKit/531.21.10 (KHTML, like Gecko) Mobile/7B405" ?


You agree to the terms of service when you sign up.

If you're talking about making anonymous requests to their service, they only allow a few of those before they stop showing you profiles. If you circumvent that protection, it's a bit more like hooking a cable up to a power line (illegal) or dumping your commercial waste in the sewer (illegal).


I disagree with your analogy. To me, the key word in "HTTP request" is request. A request is something that can be granted or not.


Perhaps more clearly would be any HTTP request that LinkedIn believes is in violation of their terms of service will be denied. It can be hard to know when the first request arrives if it is someone scraping the site or not, but once it is clear that it is someone scraping they actively deny all future requests. If they could know that the request coming in was going to be a scrape and not a page view they would preemptively deny it.


But what is the difference between a scrape and a page view? If a human looks at it once, after scraping, does it become a page view? Is pocket downloading content on my behalf for me to read later, a scraper? What's the difference between a scraper and an offline browser who's content a human never browses?


True, but in making the request, you will provide information on who is making that request. If you say, "I am a bot!", and they grant you permission, your request is legal.

But if you say, 'I am NOT a bot', like spoofing a browser's user agent string, but you are a bot, then you are requesting access under a pretense, in order to circumvent their terms of service. Kinda feels morally wrong, and illegal.


That argument works, insofar as it does, only for more recognizable bots and browsers. If I write a client of some sort that identifies itself as:

Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Snackmaster Pro/666.0.666

What do you do?

I also tell my browser to lie about what it is sometimes, due to sites that are malfunctioning, but whose owners choose to document the errors instead of fixing them with "Use Chrome" (or IE, or whatever) checks.

Is that 'kinda' illegal or morally wrong (two very different things)?

If so, that seems like a belief that all sorts of browser defaults are 'kinda' wrong and/or illegal to change. Javascript? Lying about installed fonts/screen dimensions/whatever? Refusing to keep nonsession cookies between sessions? That slope would seem to get pretty slippery...


In your first case, if you are running on Windows NT 6.1 using WebKit on a new browser for humans called 'Snakemaster Pro', then you aren't doing anything wrong.

If by client you mean a robot, then you are pretending to be a browser and you are accessing the service without permission.

Let me ask you a question, say your client was hitting my service with that user agent, 100 times a second, crawling through urls sequentionaly. Lets say I added it to my robots.txt deny list and starting blocking that user agent. Would you change the user agent and continue?

If someone creates a site that says, 'Access to this site is for 640x480 browsers only, any other use is forbidden'. Then I think its pretty clear that its a stupid site but also that faking your screen resolution is accessing a site without consent. There is no slope, someone (Linkedin) putting explicit terms on their website is pretty clear.


Have you ever heard of "headless browsers" (like [chrome](https://github.com/dhamaniasad/HeadlessBrowsers/issues/37)? What are some defining characteristics of browsers that are absent in scraping clients? If I open a browser window while doing the scraping is that acceptable?


I very rarely use robots, and think I've only been "abusive" (not really abusive, in my book) once.

What if I send a null UA? Or use it as an opportunity to share my favorite quote?

What if the behavior of my software doesn't attack like a robot, does keep the request volume reasonable (use whatever you think is reasonable here) but also doesn't do what you might expect a human clicking around to do?


There isn't a universal 'I'm a bot' setting. There are user agent conventions, but they are hardly standard. Your point works in theory, but it's not something one can just implement and be reasonably confident that they won't be scraped.


No, it won't prevent being scraped. That's not the point I was making.

The point is, the scraper would have to hide their intentions and identity, which removes any claim they are being 'honest' in their intentions and not trying to circumvent the provider of the services efforts to prevent scraping.


The user agent header is not an authorization header. There is explicitly an authorization header.


I understand the appeal of this argument, but there are clearly cases where a computationally valid request and response is illegal despite the fact that the server "chose" to satisfy the request. An obvious example would be any exploit where an attacker can construct a particular request and get access to someone else's private data.


The analogy of pouring waste is not accurate, we are looking at what is done with the service, e.g. it would be the equivalent of allowing drinking the water but not cooking with it, or using electricity for specific devices. The contract is a debit/volume, what I do with it is irrelevant, and by this analogy on the web I should be allowed to scrape if I stay in the allowed bandwidth by the website.


I mostly agree. the traditional method [1] should be honored. but clearly, there are some bad agents out there that ignore the rules and mess things up for everyone. They are pretty clear about their expectations:

    # Notice: The use of robots or other automated means to access LinkedIn without
    # the express permission of LinkedIn is strictly prohibited.
[1] https://www.linkedin.com/robots.txt


puuuh, somebody should explain the concept of user-agent groups to them (would heavily simplify their robots.txt)


I don't think that's an unpopular opinion. But thinking they can/should have the right to press charges against you for trying would be, I hope.


I completely agree with this, but if a determined company is scraping the data, they can do various things to make the traffic blend in, it's not always possible for technical means to detect scraping. To which I say "oh well, deal with it or put it behind authentication", but some may disagree.


They also state this clearly in their terms of service.

"You agree that you will not ... develop, support or use software, devices, scripts, robots, or any other means or processes (including crawlers, browser plugins and add-ons, or any other technology or manual work) to scrape the Services."

https://www.linkedin.com/legal/user-agreement

That said, this should be a breach of contract issue. It's an overreach to invoke federal fraud law.


Have no doubt this will be unpopular, but I think LinkedIn is right.

Most news sites publish to the world. But scraping a news site's content and monetizing it yourself is not ok. Legally, it violates intellectual property law. But laws aside, I assume most people would agree, if someone spent the time researching and writing an article, they should have the right to monetize it and nobody else.

In this case, IP law may not apply, but the concept is the same. I don't love LinkedIn myself. But they spent the time building a platform for collecting that info. I don't see why it should be OK for other people to scrape and monetize it.


I don't think that is an unpopular opinion. It also implies that LinkedIn is wrong, because of their HTTP server says yes, they want to retroactively say no.


Isn't the server giving you permission to view the data? Not give?

EG, if it returns an image - it doesn't imply I can use the image anywhere I want.


Generally I agree, but with linkedin anything of value seems to require logging in, which means their terms of service come into play. This is very different to scrapping data that they make open and browsable by all.


Legally, I think scrappers should respect robots.txt


They do have an amazingly large robots.txt in fact.

https://www.linkedin.com/robots.txt




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: