• 7 Posts
  • 13 Comments
Joined 3 months ago
cake
Cake day: May 14th, 2026

help-circle
  • Could be. No more or less than a very large library of books.

    Plus, you wouldn’t have to be 100% Robinson Crusoe-ing . Peer-to-peer networks, mesh radio and sneakernet would still exist in that hellscape.

    Still, for nerd fun, I paper-napkin’ed the maths last night with the aid of clankers. Here’s how I’d make an “internet in a box”.

    Reference: 1–2TB

    • Wikipedia, Gutenberg and other Kiwix archives: 300–600GB.
    • Medical, repair, farming and technical manuals: 100–300GB.
    • Forums, software docs and selected papers: 100–400GB.
    • OpenStreetMap, routing and regional map tiles: 100–500GB.
    • Selected Zimit web archives: 100–500GB.

    Zimit might fit roughly 100,000 ordinary web pages into 100GB. Kiwix server could serve the ZIM archives (inc websites), while Recoll could provide Google-like full-text search across the wider library of websites, books, manuals and documents.

    Other Software: 1–2TB

    • “Linux images”: 100–200GB.
    • Open-source apps, drivers and firmware: 200–400GB.
    • Compilers, package mirrors and useful GitHub repos: 300GB–1TB.
    • Emulators, open games and preserved software: 300GB–1TB.

    Media and education: 5–10TB

    • Podcasts and radio: 300GB–1TB.
    • Music: 1–2TB.
    • Courses, documentaries and educational video: 2–4TB.
    • Films, television and selected YouTube channels: 3–5TB (at 720p, that’s anything between 5-10yrs of content, if watched 2hrs/day).

    AI and search: 100–500GB

    • A few local language models: 50–150GB.
    • Speech, image, OCR and embedding models: 20–100GB.
    • Search indexes and vector data: 20–200GB.
    • OpenWebUI, Kiwix, Jellyfin and other services: under 20GB.

    Community and nerd nonsense: 100–500GB

    • A local forum, wiki, IRC or Lemmy instance.
    • Reticulum or Meshtastic links for nearby networks.
    • Minetest, Doom, Quake and local game servers.
    • Old Reddit or BBS archives
    • OASIS or another Reddit-style simulator populated by argumentative LLM agents (yeah that’s a thing apparently).
    • Random GitHub projects.

    (That last category has little survival value, but probably considerable morale value)

    So, roughly:

    • Lean: 6–10TB.
    • Comfortable: 12–18TB.
    • Media-heavy: 20TB or more.

    Twelve to eighteen terabytes would give you a decent slice of “the internet”. Get a 3D printer while you’re at it for the ultimate off grid chic.


  • This is my final reply in this thread. The developer has said their piece, and I have said mine. Now you’ve waded in - so let me set the record straight.

    I am a developer. I had genuine interest in this project. I read the Hister documentation and inspected parts of the repository because the documentation did not clearly answer several basic questions I had:

    • How SQLite, Bleve, and stored HTML relate.

    • Whether TTL or storage quotas exist.

    • How browser-history deletion affects stored data.

    • How previews differ from a real web archive.

    • What multi-user isolation actually covers.

    Yes, I used AI to assemble a plain-language summary and labelled it accordingly. Not everyone keeps the Hister codebase in their head, not everyone talks in code review and if I had these questions, I’m willing to bet others did too. The AI wrote for a lay audience because I didn’t ask it to do QA, I asked it to ELI-5.

    The summary contained errors. Fine. That’s AI for you. However, if neither I nor the AI could find clear answers after cloning the repo, that supports my point about opacity.

    At no point did I request a line-by-line audit. “Points 2 and 5 are wrong” would have answered the question.

    Declining would also have been reasonable. Hell, side stepping it would have been fine too. Instead the dev decided to note the inaccuracies and rudely brush them off.

    Both you and the dev seem to be under the impression !selfhosted is a one way distribution channel.

    The developer came here, invited questions, then turned the raw prawn when questions arrived.

    I didn’t go to their their Github. I didn’t abuse them. I genuinely wanted to know more about their project and share it, perhaps even work to help improve it.

    They - and now you, ostensibly a happy clapper for Hister - came here.

    Your claims about my effort and intent are assumptions followed by personal abuse.

    Try and walk a mile in someone else’s shoes before calling them low effort and shitty next time.






  • That is not what happened.

    I fed your GitHub repository to a clanker because the documentation did not answer my questions. I then shared its summary here.

    You replied afterwards and said the summary was wrong. Fair enough. I then asked which specific points were wrong.

    You could have answered, declined, or ignored the post.

    Instead, you deigned only to dismiss the effort, then blamed me for objecting.

    You also asked which parts were confusing, although my previous reply had already listed those issues.

    You did not address them then, either.

    A prospective user should not need ChatGPT, a cloned repository, and several follow-up questions to understand key functions.

    You invited feedback. Your documentation remains unclear on several points, including issues beyond those I listed.

    Your responses show that further feedback is not worth my time.


  • I am happy to narrow it further.

    I took the time to read the documentation, ask ChatGPT to summarise what I found, and then reduced my follow-up to a simple request:

    «Which of points 1–7 are materially wrong?»

    That is not the same as asking you to audit “multiple screens” of AI output.

    If the answer is “2 and 5 are incorrect”, or even “I do not have time to review it”, that is perfectly fine.

    However, dismissing it as “a multiple screens long AI prompt” does not only not answer the question, it comes off as abrasive.

    As for the documentation, the confusing parts are exactly those I listed: retention, lifecycle management, browser ingestion, storage limits, deletion, multi-user behaviour, and, most importantly, what Hister actually is and who it is for.

    What’s disappointing is not that you disagreed with the AI summary. AIs are idiots.

    It that after inviting questions and feedback, your response to a genuine attempt to understand the project is curt dismissal.

    The inner workings of Hister may be obvious to you; they are not obvious to others.

    The point is that you came here specifically to invite questions and feedback.

    “TL;DR” does not encourage the sort of community engagement you ostensibly came here to seek.


  • Excellent - thanks for clearing that up.

    Is there a TTL / max database size per user setting? Say I have 4 users using the server; can I allocate a hard limit of 10GB per user, with 180 day retention rules?

    Additionally, is the other parenthetical information materially correct? If not, which points [1 thru to 7] are wrong?

    I would like to further recommend Hister but your documentation is somewhat confusing at first blush.





  • Well, at risk of downvotes…

    I know what your asking but tbh I generally don’t search, at least not like we used to. It’s weird to say that, like I’m making some claim to heightened purity or some bullshit but it’s not that at all. Search is just … noisy AF these days.

    I actually find myself using fewer sources, more directly and own my terms.

    Eg: For something like PubMed, I go to pubmed (or define a proper boolean PICO search in an indexing tool). Sci-hub is a good source too.

    https://sci-hub.st/

    Or I can hit the wikipedia API (or even better, self hosted kiwix) and find what I need.

    https://kiwix.org/en/

    Or, for known, commonly visited sites (like fan wikis), I use self-written tools.

    There are plenty of sites have public facing, nicely formatted, JSON friendly results. Others are just clean to scrape. A recent discovery was Muvitimes (for getting movie times listings).

    https://muvitimes.com/

    These tools query my white listed sites to a particular depth, collect the results I need, and present them for review.

    You can then cache the info locally to query against using something like Meilisearch.

    https://github.com/meilisearch/meilisearch

    There are other options too. Something like Scrapy can collect structured data and a Trafilatura-like extractor can remove menus, adverts, and repeated page content.

    https://github.com/scrapy/scrapy

    I’ve even use something like self hosted Perplexica. It can query SearXNG, Tavily, Exa or other sources.

    https://github.com/kiranz/perplexica

    Honestly, if it’s a quick throw away search, I’ll call my self hosted llm, get it to pull from Tavilly and give me the summary and citations. That’s been miles better than most other options for me but YMMV.

    Finally, if you’re talking about random occasional searches on phone, when I can’t access my LAN, I don’t mind DDG-lite

    https://lite.duckduckgo.com/lite