ancuuiqter

@ancuuiqter@lemmy.world

This profile is from a federated server and may be incomplete. Browse more on the original instance.

Lawsuit Accuses Anna's Archive of Hacking WorldCat, Stealing 2.2 TB Data (torrentfreak.com)

American nonprofit OCLC is known globally for its leading database of bibliographic records, WorldCat. A few months ago, many of these records were posted publicly by the shadow library search engine, Anna’s Archive. OCLC believes that this is the result of a year-long hack and, with a lawsuit filed at an Ohio federal court,...

ancuuiqter, 3 months ago

Here are the court filings if anyone would like to read them:

archive.org/details/gov.uscourts.ohsd.287709/

The following is a link to the docket (which the above link draws from), so people can follow the progress of the lawsuit:

courtlistener.com/…/oclc-online-computer-library-…

reply

report

activity

copy /kbin url

copy original url

open original url

Loading...

ancuuiqter, 3 months ago

As to how Anna’s Archive accomplished their data scraping, this is what OCLC is claiming (see page 62-63):

These attacks were accomplished with bots (automated software applications) that “scraped” and harvested data from WorldCat.org and other WorldCat®-based research sites and that called or pinged the server directly. These bots were initially masked to appear as legitimate search engine bots from Bing or Google.

To scrape or harvest the data on WorldCat.org, the bots searched WorldCat.org results, running a script based on OCN for individual JavaScript Object Notation, or “JSON,” records. As a result, WorldCat® data including freely accessible and enriched data, such as OCNs, were scraped from individual results on WorldCat.org.

The bots also harvested data from WorldCat.org by pretending to be an internet browser, directly calling or “pinging” OCLC’s servers, and bypassing the search, or user interface, of WorldCat.org. More robust WorldCat® data was harvested directly from OCLC’s servers, including enriched data not available through the WorldCat.org user interface.

Finally, WorldCat® data was harvested from a member’s website incorporating WorldCat® Discovery Services, a subscription-based variation of WorldCat.org that is available only to a member’s patrons. Again, the hacker pinged OCLC’s servers to harvest WorldCat® records directly from the servers. To do this through WorldCat® Discovery Services/FirstSearch, the hacker obtained and used the member’s credentials to authenticate the requests to the server as a member library.

From WorldCat® Discovery Services, hackers harvested 2 million richer WorldCat® records that included data not available in WorldCat.org. This hacking method resulted in the harvesting of some of OCLC’s most proprietary fields of WorldCat® data.

These hacking attacks materially affected OCLC’s production systems and servers, requiring around-the-clock efforts from November 2022 to March 2023 to attempt to limit service outages and maintain the production systems’ performance for customers. To respond to these ongoing attacks, OCLC spent over 1.4 million dollars on its systems’ infrastructure and devoted nearly 10,000 employee hours to the same.

Despite OCLC’s best efforts, OCLC’s customers experienced many significant disruptions in paid services during the aforementioned period as a result of the attacks on WorldCat.org, requiring OCLC to create system workarounds to ensure services functioned.

During this time, customers threatened and likely did cancel their products and services with OCLC due to these disruptions.

Because OCLC had to combat these persistent hacking attacks, OCLC was forced to divert existing personnel and resources from OCLC’s other products and services. As a result, OCLC’s development and improvements to other products and services were delayed and limited.

OCLC has devoted, at various times, ten or more employees to respond to and mitigate the harm from these attacks from October 2022 to present.

reply

report

activity

copy /kbin url

copy original url

open original url

Loading...

ancuuiqter, 3 months ago (edited 3 months ago)

Regarding the operating location(s) of Anna’s Archive, OCLC is alleging the following (pages 7-9):

C. Defendants Rely on Sophisticated Technology and Online Practices to Conceal their Identities.

Defendants understand that their pirate library enterprise and related activities, here, hacking and harvesting OCLC’s WorldCat® records, are illegal. Defendants admit that they are engaging in and facilitating mass copyright infringement, stating, “[w]e deliberately violate the copyright law in most countries.” In another blog post, Defendants noted that their activities could lead to arrest and “decades of prison time.” Defendants have also recognized that their hacking and distribution of OCLC’s data is improper, acknowledging that WorldCat® is a “proprietary database,” that OCLC’s “business model requires protecting their database,” and that Defendants are “giving it all away. :-).”

Because Defendants understand their actions infringe on copyright laws, amongst others, Defendants go to great lengths to remain anonymous to ensure both that Anna’s Archive’s domains are not taken down and to avoid the legal consequences of their actions, including civil lawsuits where parties like OCLC seek to vindicate their rights, as well as criminal and regulatory enforcement actions undertaken by government entities. None of Anna’s Archive’s domains or its online blog provide a business address, business contact, or other contact information that would be found on a legitimate entity’s website.

Defendants have explained in a blog post that they are “being very careful not to leave any trace [of their online activities], and having strong operational security.” For instance, Anna’s Archive utilizes a VPN with “[a]ctual court-tested no-log policies with long track records of protecting privacy.” Each of the Anna’s Archive domains are registered using foreign hosts, registrars, and registrants in order to conceal the identity of the site operators. Additionally, Defendants rely on multiple proxy servers to maintain anonymity. Defendants also use a free version of Cloudflare, a top-level hosting provider, so that they do not have to provide any payment or other identifying information. Defendants selected Cloudflare because they claim Cloudflare has resisted requests to take down websites for copyright infringement. The individuals behind Anna’s Archive also use usernames as pseudonyms to mask their identities online.

Through the work of a cyber security and digital forensic investigation firm, OCLC was able to identify one of the individuals behind Anna’s Archive by name and locate a United States address, Defendant Maria Dolores Anasztasia Matienzo. However, the physical address and contact information of Anna’s Archive and the identities and contact information of the John Does remain unknown. It is highly likely that Anna’s Archive is a non-domestic, foreign entity, based on the findings from OCLC’s investigator, as set forth below.

OCLC explained the above in their Motion To Serve Defendant Anna’s Archive By Email, as justification for why they seek “permission to serve Anna’s Archive by alternative means, here, email, pursuant to Federal Rule of Civil Procedure 4(h)(2) and (f)(3).”

reply

report

activity

copy /kbin url

copy original url

open original url

Loading...

ancuuiqter, 3 months ago

Would you be able to share where you learned that Anna’s Archive is based in Kazakhstan?

reply

report

activity

copy /kbin url

copy original url

open original url

Loading...

ancuuiqter, 3 months ago

The official Anna’s Archive Reddit account, AnnaArchivist, has responded to an r/Annas_Archive post linking the same Torrent Freak article:

Thanks! We’re not making any public statements about this lawsuit but rest assured we’re fine.

reply

report

activity

copy /kbin url

copy original url

open original url

Loading...

ancuuiqter, 3 months ago

Maybe you’re thinking of Sci-Hub and its founder, Alexandra Asanovna Elbakyan?

I could not find a location on Anna’s Archive’s wiki page.

reply

report

activity

copy /kbin url

copy original url

open original url

Loading...

Z-Library Blog: "Unprecedented seizure of our domains with books on rare languages" (z-library.se)

Today we are forced to share some sad news - yesterday many of our domains were seized again. We should highlight that the majority of the seized domains were not mirrors of the Z-Library website. Instead, they were separate sub-projects, containing only books in rare languages of the world, and their blocking is perplexing. For...

2 companies caught illegally printing over 15,000 books, calendars in Ho Chi Minh City, Vietnam (tuoitrenews.vn)

Authorities have caught two Ho Chi Minh City-based firms printing more than 15,000 pirated copies of books and 2024 calendars, with a combined weight exceeding 15 metric tons....

This community can't continue to exist on lemmy.world

The instance owners do not wish to host potentially problematic content....

ancuuiqter, 8 months ago

What if the community shifted to an already-existing one?

lemmy.ml/c/libgen

reply

report

activity

copy /kbin url

copy original url

open original url

Loading...

ancuuiqter, 8 months ago

Do you mind elaborating? Is there something you could share that provides more context?

reply

report

activity

copy /kbin url

copy original url

open original url

Loading...

Copyright lawsuits against Meta and OpenAI mention shadow libraries, including Library Genesis, as sources of training data (archive.md)

cross-posted from: lemmy.world/post/1330512...

ancuuiqter, 10 months ago

Mentioning this since the project Anna’s Archive compiles several datasets and their corresponding torrents.

Anna’s Archive, whose aim is to “archive all the books in the world, and make them widely accessible,” pulls from a number of shadow library sources; the project provides its own torrent links (via Tor) for Library Genesis, Z-lib, Internet Archive, among others, plus Library Genesis’s torrents. In the datasets linked below, you can click on a given source and find its onion site or the torrents provided by the shadow library itself (in the case of Library Genesis, for example).

Anna’s Archive datasets

…almost all files shown on Anna’s Archive are available through torrents. Below is a list of the different data sources that we use, with links to their torrents. Our own torrents are available on Tor.

Sources include

Internet Archive Digital Lending Library

Libgen.li comics

Z-Library scrape

ISBNdb scrape

Libgen auxiliary data

Libgen.rs

Libgen.li (includes Sci-Hub)

reply

report

activity

copy /kbin url

copy original url

open original url

Loading...

DeGruyter Collection

tldr: a huge torrent of books straight from the academic publisher De Gruyter...

ancuuiqter, 10 months ago

Yeah, it's not clear. The following explanation goes into a bit more detail on finding what you want to download, for anyone else who has trouble identifying books in the torrent.

Try browsing through this link on the publisher's site for any subjects or books you'd be interested in. If you find something, copy the ISBN (listed in the work's direct URL, or you can open the work's page, which lists out the eBook and hardcover ISBNs---you need the eBook ISBN as far as I understand). If you open the torrent and let the file contents load, depending on your client, you can look through the structure, search through all files using the ISBN, and locate the book that way. Files all seem to be named by their DOI, with the / replaced with _. I guess the ISBN is part of the DOI naming convention.

It doesn't look like any work tagged as Ahead of Publication will be in this torrent, but works from 2023 and prior appear to be.

reply

report

activity

copy /kbin url

copy original url

open original url

Loading...