dogear

enter for all results · esc to close

Web Archiving

186 items from iipc/awesome-web-archiving ★2,660

  1. 0
    Flameshot github.com

    Powerful yet simple to use screenshot software.

  2. 0
    wayback github.com

    A bot for Telegram, Mastodon, Slack, and other messaging platforms archives webpages.

  3. 0
    ArchiveBox github.com

    A tool which maintains an additive archive from RSS feeds, bookmarks, and links using wget, Chrome headless, and other methods (formerly Bookmark Archiver). (In Development)

  4. 0
    SingleFile github.com

    Browser extension for Firefox/Chrome and CLI tool to save a faithful copy of a complete page as a single HTML file. (Stable)

  5. 0
    monolith github.com

    CLI tool to save a web page as a single HTML file. (Stable)

  6. 0
    DiskerNet github.com

    A non-WARC-based tool which hooks into the Chrome browser and archives everything you browse making it available for offline replay. (In Development)

  7. 0
    xDoTool github.com

    Click automation on Ubuntu.

  8. 0
    InterPlanetary Wayback github.com

    Web Archive (WARC) indexing and replay using IPFS.

  9. 0
    Internet Archive Library github.com

    A command line tool and Python library for interacting directly with archive.org. (Python). (Stable)

  10. 0
    PYWB github.com

    A Python 3 implementation of web archival replay tools, sometimes also known as 'Wayback Machine'. (Stable)

  11. 0
    grab-site github.com

    The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns. (Stable)

  12. 0
    twarc github.com

    A command line tool and Python library for archiving Twitter JSON data. (Stable)

  13. 0
    Browsertrix Crawler github.com

    A Chromium based high-fidelity crawling system, designed to run a complex, customizable browser-based crawl in a single Docker container. (Stable)

  14. 0
    Auto Archiver github.com

    Python script to automatically archive social media posts, videos, and images from a Google Sheets document. Read the article about Auto Archiver on bellingcat.com.

  15. 0
    wikiteam github.com

    Tools for downloading and preserving wikis. (Stable)

  16. 0
    Brozzler github.com

    A distributed web crawler (爬虫) that uses a real browser (Chrome or Chromium) to fetch pages and embedded urls and to extract links. (Stable)

  17. 0
    Archives Unleashed Toolkit github.com

    Open-source toolkit for analyzing web archives.

  18. 0
    Wpull github.com

    A Wget-compatible (or remake/clone/replacement/alternative) web downloader and crawler. (Stable)

  19. 0
    Waybackpy github.com

    Wayback Machine Save, CDX and availability API interface in Python and a command-line tool (Stable)

  20. 0
    OpenWayback github.com

    The open source project aimed to develop Wayback Machine, the key software used by web archives worldwide to play back archived websites in the user's browser. (Stable)

  21. 0
    warcio github.com

    Streaming WARC/ARC library for fast web archive IO (Python). (Stable)

  22. 0
    Warcprox github.com

    WARC-writing MITM HTTP/S proxy. (Stable)

  23. 0
    archivenow github.com

    A Python library to push web resources into on-demand web archives. (Stable)

  24. 0
    warcdb github.com

    A command line utility (Python) for importing WARC files into a SQLite database. (Stable)

  25. 0
    WAIL github.com

    A graphical user interface (GUI) atop multiple web archiving tools intended to be used as an easy way for anyone to preserve and replay web pages; Python, Electron. (Stable)

  26. 0
    hyphe github.com

    A webcrawler built for research uses with a graphical user interface in order to build web corpuses made of lists of web actors and maps of links between them. (Stable)

  27. 0
    Obelisk github.com

    Go package and CLI tool for saving web page as single HTML file. (Stable)

  28. 0
    freeze-dry github.com

    JavaScript library to turn page into static, self-contained HTML document; useful for browser extensions. (In Development)

  29. 0
    Scoop github.com

    High-fidelity, browser-based, single-page web archiving library and CLI for witnessing the web. (Stable)

  30. 0
    Go Get Crawl github.com

    Extract web archive data using Wayback Machine and Common Crawl. (Stable)

  31. next page of items loading…