Skip to content
EntityQ3097891· pop 7· linked from 43 articles

Heritrix is a web crawler designed for web archiving. It was originally written in collaboration between the Internet Archive, National Library of Norway and National Library of Iceland. Heritrix is available under a free software license and written in Java. The main interface is accessible using a web browser, and there is a command-line tool that can optionally be used to initiate crawls.

In the Vinony graph

Within Vinony's link graph, Heritrix is referenced by 43 other articles, and connects out to Internet Archive, web crawler and National and University Library of Iceland.

Vinony files it under 2014 software, Free software programmed in Java and Free web crawlers.

Its subject is documented across 6 Wikipedia language editions.

Described at

Heritrix is web crawler used by the Internet Archive, which provides a web-based user interface after initial configuration on a Linux machine. Also used by the Library of Congress, Heritrix captures metadata in the Web ARChive (WARC) format.

Excerpt from a page describing this subject · 2,140 chars · not written by Vinony

Source code

Heritrix is designed to respect the robots.txt exclusion directives and META nofollow tags. Please consider the load your crawl will place on seed sites and set politeness policies accordingly. Also, always identify your crawl with contact information in the User-Agent so sites that may be adversely affected by your crawl can contact you or adapt their server behavior accordingly. Some individual source code files are subject to or offered under other licenses. See the included LICENSE.txt file for more information.

Excerpt from the source-code README · 3,111 chars · not written by Vinony

Wikidata facts

Instance of
free software
Official website
heritrix.readthedocs.io
Image
Heritrix-screenshot.png
Show 9 more facts
software version identifier
3.14.1
readable file format
Web ARChive
writable file format
Web ARChive
programmed in
Java
source code repository URL
github.com/internetarchive/heritrix3
Commons category
Heritrix
Sources (8)

via Wikidata · CC0

Article · 日本語

Heritrix はインターネット・アーカイブが開発したウェブアーカイブのためのWebクローラーの一種。Java言語で実装され、フリーソフトウェアライセンスにより自由に利用できる。主にウェブブラウザを使って操作するが、コマンドラインツールを使ってクロールを開始するなどの操作も可能である。名前は「(女性の)相続人」を意味するheiressの古語に由来する。 Heritrixの開発は、2003年にまとめられた仕様に基づいて、インターネット・アーカイブとNordic National Librariesの共同で行われた。最初のリリースは2004年1月で、その後インターネット・アーカイブの従業員や外部のウェブアーカイブに関心を持つ人々によって継続的に改良が続けられている。 もっともHeritrixがインターネット・アーカイブ自身のウェブ収集に使われるようになったのはかなり後のことである。かつてはアーカイブの大半はアレクサ・インターネット社から提供されていた。アレクサ社は自身の業務に供するため独自のia_archiverと呼ばれるクローラーを使ってウェブ収集を行っており、収集したデータをインターネット・アーカイブに寄贈している。当初インターネット・アーカイブ自身もHeritrixを使って収集を行ってはいたが、小規模なものに留まっていた。 2008年からインターネット・アーカイブは自身の全ウェブ規模のクローリングの性能を向上させ、現在では自身で収集したものが大半を占めるようになっている。

Abstract from DBpedia / Wikipedia · CC BY-SA

Available in 6 languages

via Wikidata sitelinks · CC0

Connections

Categories