Web Archiving
186 items from iipc/awesome-web-archiving ★2,660
-
-
wayback github.com
A bot for Telegram, Mastodon, Slack, and other messaging platforms archives webpages.
-
ArchiveBox github.com
A tool which maintains an additive archive from RSS feeds, bookmarks, and links using wget, Chrome headless, and other methods (formerly Bookmark Archiver). (In Development)
-
SingleFile github.com
Browser extension for Firefox/Chrome and CLI tool to save a faithful copy of a complete page as a single HTML file. (Stable)
-
-
DiskerNet github.com
A non-WARC-based tool which hooks into the Chrome browser and archives everything you browse making it available for offline replay. (In Development)
-
-
-
Internet Archive Library github.com
A command line tool and Python library for interacting directly with archive.org. (Python). (Stable)
-
PYWB github.com
A Python 3 implementation of web archival replay tools, sometimes also known as 'Wayback Machine'. (Stable)
-
grab-site github.com
The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns. (Stable)
-
-
Browsertrix Crawler github.com
A Chromium based high-fidelity crawling system, designed to run a complex, customizable browser-based crawl in a single Docker container. (Stable)
-
Auto Archiver github.com
Python script to automatically archive social media posts, videos, and images from a Google Sheets document. Read the article about Auto Archiver on bellingcat.com.
-
-
Brozzler github.com
A distributed web crawler (爬虫) that uses a real browser (Chrome or Chromium) to fetch pages and embedded urls and to extract links. (Stable)
-
-
Wpull github.com
A Wget-compatible (or remake/clone/replacement/alternative) web downloader and crawler. (Stable)
-
Waybackpy github.com
Wayback Machine Save, CDX and availability API interface in Python and a command-line tool (Stable)
-
OpenWayback github.com
The open source project aimed to develop Wayback Machine, the key software used by web archives worldwide to play back archived websites in the user's browser. (Stable)
-
-
-
-
warcdb github.com
A command line utility (Python) for importing WARC files into a SQLite database. (Stable)
-
WAIL github.com
A graphical user interface (GUI) atop multiple web archiving tools intended to be used as an easy way for anyone to preserve and replay web pages; Python, Electron. (Stable)
-
hyphe github.com
A webcrawler built for research uses with a graphical user interface in order to build web corpuses made of lists of web actors and maps of links between them. (Stable)
-
-
freeze-dry github.com
JavaScript library to turn page into static, self-contained HTML document; useful for browser extensions. (In Development)
-
Scoop github.com
High-fidelity, browser-based, single-page web archiving library and CLI for witnessing the web. (Stable)
-
- next page of items loading…