Heritrix
Sign in to saveHeritrix is a web crawler designed for web archiving. It was originally written in collaboration between the Internet Archive, National Library of Norway and National Library of Iceland. Heritrix is available under a free software license and written in Java. The main interface is accessible using a web browser, and there is a command-line tool that can optionally be used to initiate crawls.
Described at
TAPoR
tapor.ca →Heritrix is web crawler used by the Internet Archive, which provides a web-based user interface after initial configuration on a Linux machine. Also used by the Library of Congress, Heritrix captures metadata in the Web ARChive (WARC) format.
Excerpt from a page describing this subject · 2,140 chars · not written by Vinony
Source code
Heritrix is designed to respect the robots.txt exclusion directives and META nofollow tags. Please consider the load your crawl will place on seed sites and set politeness policies accordingly. Also, always identify your crawl with contact information in the User-Agent so sites that may be adversely affected by your crawl can contact you or adapt their server behavior accordingly. Some individual source code files are subject to or offered under other licenses. See the included LICENSE.txt file for more information.
Excerpt from the source-code README · 3,111 chars · not written by Vinony
Wikidata facts
- Instance of
- free software
- Developer
- Internet Archive
- Official website
- heritrix.readthedocs.io
- Image
- Heritrix-screenshot.png
- Has use
- web archiving
Show 9 more facts
- copyright license
- Apache Software License 2.0
- software version identifier
- 3.14.1
- readable file format
- Web ARChive
- writable file format
- Web ARChive
- programmed in
- Java
- described at URL
- marketplace.sshopencloud.eu/tool-or-service/NrGetP
- user manual URL
- github.com/internetarchive/heritrix3/wiki
- source code repository URL
- github.com/internetarchive/heritrix3
- Commons category
- Heritrix
Sources (8)
via Wikidata · CC0
Article · Español
Heritrix es un rastreador (o crawler) de ficheros web a través de internet. Su licencia es open-source y está escrito completamente en JAVA. Su interfaz de configuración es accesible usando un navegador web, haciéndolo muy versátil y cómodo de usar, aunque también puede ser lanzando desde línea de comandos. Heritrix fue desarrollado conjuntamente por Internet Archive y "Nordic National Libraries" a principios de 2003. La primera versión fue publicada en enero de 2004 y ha sido continuamente actualizado por los miembros de Internet Archive y terceras partes.
Abstract from DBpedia / Wikipedia · CC BY-SA