Heritrix
Sign in to saveHeritrix is a web crawler designed for web archiving. It was originally written in collaboration between the Internet Archive, National Library of Norway and National Library of Iceland. Heritrix is available under a free software license and written in Java. The main interface is accessible using a web browser, and there is a command-line tool that can optionally be used to initiate crawls.
In the Vinony graph
Within Vinony's link graph, Heritrix is referenced by 43 other articles, and connects out to Internet Archive, web crawler and National and University Library of Iceland.
Vinony files it under 2014 software, Free software programmed in Java and Free web crawlers.
Its subject is documented across 6 Wikipedia language editions.
Described at
TAPoR
tapor.ca →Heritrix is web crawler used by the Internet Archive, which provides a web-based user interface after initial configuration on a Linux machine. Also used by the Library of Congress, Heritrix captures metadata in the Web ARChive (WARC) format.
Excerpt from a page describing this subject · 2,140 chars · not written by Vinony
Source code
Heritrix is designed to respect the robots.txt exclusion directives and META nofollow tags. Please consider the load your crawl will place on seed sites and set politeness policies accordingly. Also, always identify your crawl with contact information in the User-Agent so sites that may be adversely affected by your crawl can contact you or adapt their server behavior accordingly. Some individual source code files are subject to or offered under other licenses. See the included LICENSE.txt file for more information.
Excerpt from the source-code README · 3,111 chars · not written by Vinony
Wikidata facts
- Instance of
- free software
- Developer
- Internet Archive
- Official website
- heritrix.readthedocs.io
- Image
- Heritrix-screenshot.png
- Has use
- web archiving
Show 9 more facts
- copyright license
- Apache Software License 2.0
- software version identifier
- 3.14.1
- readable file format
- Web ARChive
- writable file format
- Web ARChive
- programmed in
- Java
- described at URL
- marketplace.sshopencloud.eu/tool-or-service/NrGetP
- user manual URL
- github.com/internetarchive/heritrix3/wiki
- source code repository URL
- github.com/internetarchive/heritrix3
- Commons category
- Heritrix
Sources (8)
via Wikidata · CC0
Article · 日本語
Heritrix はインターネット・アーカイブが開発したウェブアーカイブのためのWebクローラーの一種。Java言語で実装され、フリーソフトウェアライセンスにより自由に利用できる。主にウェブブラウザを使って操作するが、コマンドラインツールを使ってクロールを開始するなどの操作も可能である。名前は「(女性の)相続人」を意味するheiressの古語に由来する。 Heritrixの開発は、2003年にまとめられた仕様に基づいて、インターネット・アーカイブとNordic National Librariesの共同で行われた。最初のリリースは2004年1月で、その後インターネット・アーカイブの従業員や外部のウェブアーカイブに関心を持つ人々によって継続的に改良が続けられている。 もっともHeritrixがインターネット・アーカイブ自身のウェブ収集に使われるようになったのはかなり後のことである。かつてはアーカイブの大半はアレクサ・インターネット社から提供されていた。アレクサ社は自身の業務に供するため独自のia_archiverと呼ばれるクローラーを使ってウェブ収集を行っており、収集したデータをインターネット・アーカイブに寄贈している。当初インターネット・アーカイブ自身もHeritrixを使って収集を行ってはいたが、小規模なものに留まっていた。 2008年からインターネット・アーカイブは自身の全ウェブ規模のクローリングの性能を向上させ、現在では自身で収集したものが大半を占めるようになっている。
Abstract from DBpedia / Wikipedia · CC BY-SA