Hi everyone,

I am the original author of Searx. I started Hister with a similar motivation: reducing our dependence on external search engines while keeping searches and personal data under our control.

Searx is a metasearch engine that forwards queries to other search providers. Hister takes a different approach. It builds a private full text index from content you choose, then searches that index entirely on your own infrastructure.

Hister can automatically index pages through its Firefox and Chrome extensions. It can also watch local directories, import browser history and bookmarks, index individual URLs, and crawl complete documentation sites.

The feature I find most useful is offline previews. Hister stores the readable content and HTML of indexed pages locally. You can open a result in a clean and sanitized preview beside the search results without visiting the original website again.

Some other features:

  1. Full text search across web pages, PDFs, docx files, Markdown, OrgMode and text files
  2. Phrase searches, field filters, date filters, wildcards, negation, aliases, labels, facets, and result priorities
  3. Optional semantic search using an embeddings endpoint you configure
  4. Persistent website crawls
  5. Imports from browser history, Linkwarden, Karakeep, Shaarli, Wallabag, and Linkding
  6. Web, terminal, command line, HTTP API, and MCP interfaces
  7. SQLite and PostgreSQL support, plus optional multiple user hosting

Hister cannot replace a global search engine (yet) for subjects you have never encountered because it only searches what you have indexed. My workflow is to search Hister first, then use its shortcut to fall back to traditional search when I need broader web results.

The project is free software under the AGPLv3+ license. It can be installed as a standalone binary or with Docker.

Project: https://github.com/asciimoo/hister

Website and documentation: https://hister.org/

Small read-only demo: https://demo.hister.org/

I’d appreciate feedback, questions, and suggestions as well as joining our growing community.

AI disclosure: AI assisted contributions are not strictly prohibited, but all contributions should be made by humans. More details: https://github.com/asciimoo/hister/blob/master/CONTRIBUTING.md#ai-policy

  • conrad82@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    edit-2
    5 hours ago

    Can i use this in my self hosted environment? the docs talk mainly about running it in the terminal

    does it have a docker install? found it https://hister.org/docs/docker

    can i connect it to other services like paperless, or would i need to manually import files?

  • Fmstrat@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    13 hours ago

    Is there a future where Hister takes on the functions of SearXNG? I.E. new content is discovered through APIs but local or previous content is prioritized?

    • asciimoo@lemmy.mlOP
      link
      fedilink
      English
      arrow-up
      4
      ·
      9 hours ago

      Currently it is not planned. Hister guarantees that non of your data/query/metadata leaves the service if you use it. As I see, this is a more valuable and unique feature than having an integrated metasearch. There are already great metasearch solutions and Hister provides an easy fallback to search providers, so in my opinion this direction would be more of a sacrifice than an improvement.

  • Ralim~@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    8
    ·
    1 day ago

    I’ve been running this for a while since I first heard of it, it’s already saved me a few times trying to recall a page for something.

    I treat it like a more advanced history and it’s great!

  • cultist@feddit.dk
    link
    fedilink
    English
    arrow-up
    8
    ·
    1 day ago

    Asciimoo always with the cool stuff!

    Think I’ll try getting it setup on Kubernetes and using it for searching documentation when programming.

    Would you be interested in getting a Kubernetes example in the docs too then? If you have any requirements for it let me know.

    • asciimoo@lemmy.mlOP
      link
      fedilink
      English
      arrow-up
      9
      ·
      1 day ago

      I’d appreciate it, thanks. No special requirements. Providing sensible defaults and explaining usage/potential customization options would be great.

  • MoogleMaestro@lemmy.zip
    link
    fedilink
    English
    arrow-up
    28
    ·
    2 days ago

    Any interested in making this a federated search?

    I’ve been toying with the idea of a web-ring style search engine where it uses the fediverse’s post and sub system so that you can subscribe to “indexers” that are made by other users (or even mastodon users themselves) so that you can have highly targeted search engine systems. So, for example, you could make multiple search “scopes” and each of those scopes would follow specific other “sources” and search from them. Likewise, users could provide links and descriptions to add new entries to each scope which would, in essense, work like a fediverse post and then be distributed to all other interested parties.

    I haven’t done a lick of implementation yet, but thought I’d share the idea here as I consider doing this more and more every day. The only thing I haven’t figured out yet is images, because ideally that would be a special type of scope view.

  • anon_8675309@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    1 day ago

    The name reminds me of this crazy nostradamus “documentary “ that used to play on HBO when I was a kid.

  • sbeak@sopuli.xyz
    link
    fedilink
    English
    arrow-up
    9
    ·
    2 days ago

    This seems super cool, you can curate your own search index :O

    Being able to preview your search results offline is quite neat too!

    One question though, how will you solve the issue of “echo chambers”, because from my understanding, it looks like it will only present results that you search for / have searched for, meaning for some people, you could be stuck with results that only support your view of the world.

    • hoshikarakitaridia@lemmy.world
      link
      fedilink
      English
      arrow-up
      15
      ·
      2 days ago

      I mean it kinda looks like that’s by design currently. It’s an indexer for your corners of the web, not a search aggregator to find new ones.

      • asciimoo@lemmy.mlOP
        link
        fedilink
        English
        arrow-up
        18
        ·
        2 days ago

        Exactly, the echo chamber phenomenon is mostly problematic for “discovery type” searches while Hister is mainly for “recall type” search.

        Implementing federated search/index sharing could be a partial solution to this issue in the long run.

  • glizzyguzzler@piefed.blahaj.zone
    link
    fedilink
    English
    arrow-up
    6
    ·
    2 days ago

    Oh damn this could replace the bookmakers. I have Linkding with the Linkding Injector extension, but this would be next level.

    You can view the saved text too, nice. Would it be possible to attach an HTML file from like single file for sites that have heavy image content as part of the “view” button? That’d completely replace Linkding/Karakeep/Linkwarden use cases for me

    Understandable if not, that’s ancillary to the text search focus

    • asciimoo@lemmy.mlOP
      link
      fedilink
      English
      arrow-up
      11
      ·
      2 days ago

      Oh damn this could replace the bookmakers.

      The inspiration for Hister was a bookmarking app, but I realized that I always forget to manually trigger the bookmarking and I miss so many great resources.

      Would it be possible to attach an HTML file from like single file

      Currently you can import SingleFile HTMLs using the hister import file command, but no further integrations are implemented yet. Although, rendering the exact SinglePage file as a preview can be added relatively quickly. It would be also nice to accept files directly from the SingleFile extension.

      Thanks for the good suggestion, I’ve added it to my TODO! =]

      • hirihit640@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        2
        ·
        1 day ago

        I have tons of webpages saved using the Firefox built-in page saver, which saves and html file and a corresponding folder for the other resources (images, javascript). Would be cool if these could be imported as well. Though maybe the resource folder can be ignored and the html file can already be imported?

        • asciimoo@lemmy.mlOP
          link
          fedilink
          English
          arrow-up
          3
          ·
          1 day ago

          Though maybe the resource folder can be ignored and the html file can already be imported?

          It depends. Hister always requires a unique URL for each document. SingleFile snapshots include the original URL of the document as a meta HTML element. I’m not sure if the built-in page saver provides URL information.

    • asciimoo@lemmy.mlOP
      link
      fedilink
      English
      arrow-up
      13
      ·
      1 day ago

      The summary has numerous inaccuracies. Most importantly: it is pretty easy to delete content by topic or age. The hister delete command can accept a search query to remove only matched documents. The same is true on the web UI “actions -> remove all matching documents”. You can quickly filter by age, simply query updated:>365d. Combine it with URLs, labels, domains or phrases.

      • SuspiciousCarrot78@aussie.zone
        link
        fedilink
        English
        arrow-up
        2
        arrow-down
        4
        ·
        1 day ago

        Excellent - thanks for clearing that up.

        Is there a TTL / max database size per user setting? Say I have 4 users using the server; can I allocate a hard limit of 10GB per user, with 180 day retention rules?

        Additionally, is the other parenthetical information materially correct? If not, which points [1 thru to 7] are wrong?

        I would like to further recommend Hister but your documentation is somewhat confusing at first blush.

        • ReluctantMuskrat@lemmy.world
          link
          fedilink
          English
          arrow-up
          3
          arrow-down
          1
          ·
          edit-2
          6 hours ago

          Additionally, is the other parenthetical information materially correct? If not, which points [1 thru to 7] are wrong?

          This is one of the problems with relying on AI… it can produce an overwhelming amount of content with errors and inaccuracies throughout. If you don’t review and know the content yourself, you won’t know what it got wrong.

          It’s genuinely rude to lazily use AI to produce such a large babble of details and then ask someone else to review it for errors when you haven’t reviewed it yourself. I know you didn’t mean it that way but that’s nonetheless the result. People are going to read your comment and be misled about this project all because they assume AI is accurate and you didn’t review its results.

          Edit: Sheesh… I hadn’t even got to your shitty comments that followed. You use AI to make a low effort but highly verbose post and then get mad that the repo author won’t review it in detail for you when you can’t be arsed with reading the docs yourself.

          • lambalicious@lemmy.sdf.org
            link
            fedilink
            English
            arrow-up
            2
            ·
            3 hours ago

            This is one of the problems with relying on AI… it can produce an overwhelming amount of content with errors and inaccuracies throughout. If you don’t review and know the content yourself, you won’t know what it got wrong.

            Yep, and this induces anyone who wants to review the resulting slop to need to turn to AI as well to even deal with the amount (or dismiss as a whole).

            AI is a viral infection on software development.

            • ReluctantMuskrat@lemmy.world
              link
              fedilink
              English
              arrow-up
              1
              ·
              51 minutes ago

              Yep… couldn’t agree more. I has its uses but too many people are using it and disconnecting their brain. Going to meetings where people have used it to determine requirements, do analysis or even transform data and then don’t even review the results themselves before presenting and asking us to review them is rage-inducing. “What do you mean it has problems? Where? What’s wrong??” And there’s like 5 things I’ve spotted in 2 minutes and it’s clear they’ve not even reviewed it themselves.

              This guy and his long-ass, poorly formatted AI-generated “documentation” no one asked for and then after numerous errors are pointed out… “Will you review the rest?” The audacity!! 😄

          • SuspiciousCarrot78@aussie.zone
            link
            fedilink
            English
            arrow-up
            2
            arrow-down
            1
            ·
            edit-2
            3 hours ago

            This is my final reply in this thread. The developer has said their piece, and I have said mine. Now you’ve waded in - so let me set the record straight.

            I am a developer. I had genuine interest in this project. I read the Hister documentation and inspected parts of the repository because the documentation did not clearly answer several basic questions I had:

            • How SQLite, Bleve, and stored HTML relate.

            • Whether TTL or storage quotas exist.

            • How browser-history deletion affects stored data.

            • How previews differ from a real web archive.

            • What multi-user isolation actually covers.

            Yes, I used AI to assemble a plain-language summary and labelled it accordingly. Not everyone keeps the Hister codebase in their head, not everyone talks in code review and if I had these questions, I’m willing to bet others did too. The AI wrote for a lay audience because I didn’t ask it to do QA, I asked it to ELI-5.

            The summary contained errors. Fine. That’s AI for you. However, if neither I nor the AI could find clear answers after cloning the repo, that supports my point about opacity.

            At no point did I request a line-by-line audit. “Points 2 and 5 are wrong” would have answered the question.

            Declining would also have been reasonable. Hell, side stepping it would have been fine too. Instead the dev decided to note the inaccuracies and rudely brush them off.

            Both you and the dev seem to be under the impression !selfhosted is a one way distribution channel.

            The developer came here, invited questions, then turned the raw prawn when questions arrived.

            I didn’t go to their their Github. I didn’t abuse them. I genuinely wanted to know more about their project and share it, perhaps even work to help improve it.

            They - and now you, ostensibly a happy supporter of Hister - came here.

            Your claims about my effort and intent are assumptions followed by personal abuse.

            Try and walk a mile in someone else’s shoes before calling them low effort and shitty next time.

        • asciimoo@lemmy.mlOP
          link
          fedilink
          English
          arrow-up
          7
          arrow-down
          1
          ·
          1 day ago

          Is there a TTL / max database size per user setting?

          As I wrote, you can simply automate deletion by document age. Schedule a delete event on each day with the desired retention time defined as a filter expression. Database size limit isn’t available yet.

          Additionally, is the other parenthetical information materially correct? If not, which points [1 thru to 7] are wrong?

          No, sorry, I don’t have time to correct a copy of a multiple screens long AI prompt.

          I would like to further recommend Hister but your documentation is somewhat confusing at first blush.

          Which parts are confusing?

          • SuspiciousCarrot78@aussie.zone
            link
            fedilink
            English
            arrow-up
            1
            arrow-down
            10
            ·
            edit-2
            23 hours ago

            I am happy to narrow it further.

            I took the time to read the documentation, ask ChatGPT to summarise what I found, and then reduced my follow-up to a simple request:

            «Which of points 1–7 are materially wrong?»

            That is not the same as asking you to audit “multiple screens” of AI output.

            If the answer is “2 and 5 are incorrect”, or even “I do not have time to review it”, that is perfectly fine.

            However, dismissing it as “a multiple screens long AI prompt” does not only not answer the question, it comes off as abrasive.

            As for the documentation, the confusing parts are exactly those I listed: retention, lifecycle management, browser ingestion, storage limits, deletion, multi-user behaviour, and, most importantly, what Hister actually is and who it is for.

            What’s disappointing is not that you disagreed with the AI summary. AIs are idiots.

            It that after inviting questions and feedback, your response to a genuine attempt to understand the project is curt dismissal.

            The inner workings of Hister may be obvious to you; they are not obvious to others.

            The point is that you came here specifically to invite questions and feedback.

            “TL;DR” does not encourage the sort of community engagement you ostensibly came here to seek.

            • asciimoo@lemmy.mlOP
              link
              fedilink
              English
              arrow-up
              8
              arrow-down
              1
              ·
              1 day ago

              However, dismissing it as “a multiple screens long AI prompt” does not only not answer the question, it comes off as abrasive.

              I think it is more abrasive to copy/paste a poorly formatted LLM output instead of taking the time and summarizing it to a few sentences just as you did in your previous post.

              • SuspiciousCarrot78@aussie.zone
                link
                fedilink
                English
                arrow-up
                2
                arrow-down
                11
                ·
                edit-2
                1 day ago

                That is not what happened.

                I fed your GitHub repository to a clanker because the documentation did not answer my questions. I then shared its summary here.

                You replied afterwards and said the summary was wrong. Fair enough. I then asked which specific points were wrong.

                You could have answered, declined, or ignored the post.

                Instead, you deigned only to dismiss the effort, then blamed me for objecting.

                You also asked which parts were confusing, although my previous reply had already listed those issues.

                You did not address them then, either.

                A prospective user should not need ChatGPT, a cloned repository, and several follow-up questions to understand key functions.

                You invited feedback. Your documentation remains unclear on several points, including issues beyond those I listed.

                Your responses show that further feedback is not worth my time.

                • curbstickle_lw@lemmy.worldM
                  link
                  fedilink
                  English
                  arrow-up
                  2
                  ·
                  1 hour ago

                  Feedback or questions is appropriate. We don’t need to come in abrasive, and you didn’t need chatgpt to ask a followup, so we don’t need to be antagonistic.

  • alcea@feddit.org
    link
    fedilink
    English
    arrow-up
    1
    arrow-down
    5
    ·
    1 day ago

    If just it was called hipster and utilized a rapper with backwards facing basecap and bling bling as logo

    Lets see how it stacks up against the big ones