Research Notes
← Back to Main Site
Open Source • Information Architecture • Self-Hosting

NeoSearch and the Open Index: A Transparent Search Experiment

Quick answer: An open-source search engine that retrieves web results, enriches and re-ranks them with LLM-backed components, and streams AI-generated overviews and lenses. Its public repository also contains a crawler subsystem and a roadmap toward a distributed index.

How an Apache 2.0 open-source engine by Bart Jellema tackles algorithmic decay, re-ranks the raw web, and charts a three-stage path toward open, decentralized indexing.

For over two decades, web search has been monopolized by centralized advertising giants. What began as the fundamental front door to human knowledge gradually degraded into a walled garden of sponsored placements, affiliate link farms, and AI-generated SEO sludge designed to collect behavioral telemetry rather than deliver authoritative answers.

In late July 2026, serial entrepreneur Bart Jellema (creator of NeoSearch and co-founder of Tjoos) publicly open-sourced NeoSearch under the Apache License 2.0 (GitHub repository). NeoSearch presents a direct architectural challenge to the conventional wisdom of search engine moats: asking whether the core failure of modern search lies within index scale or in the monetization layers built atop it.

The Core Architectural Thesis: Index vs. Ranking Layer

Building an exhaustive global web index from scratch is capital-intensive, requiring immense crawling infrastructure, petabyte-scale storage, and worldwide edge networks. This massive capital requirement has long served as Big Tech's primary moat, leaving independent search engines dependent on reselling raw syndication APIs.

NeoSearch challenges this dynamic with an alternative hypothesis: the raw web index is merely crude oil, and the breakdown in search quality is entirely a consequence of the refining and ranking incentives.

"NeoSearch starts from a different question: how much of what’s wrong with search is the index, and how much is everything built on top of it? My bet is that it’s mostly the latter. To test that, NeoSearch takes the top 20 Google results for a query and treats them as raw material."
— Bart Jellema, Creator of NeoSearch

Under its current alpha implementation, NeoSearch processes query results through an autonomous intelligence pipeline:

The Three-Stage Decentralization Roadmap

NeoSearch is structured across a three-phase development lifecycle aimed at breaking monolithic control over web retrieval:

Stage 1: Validating the Re-Ranking Experience (Current Alpha)

Proving that transparent, open heuristics applied over existing raw index data consistently deliver superior search utility compared to ad-incentivized commercial engines.

Stage 2: Crawling the Neglected "Small Web"

Deploying dedicated crawlers focused explicitly on independent technical forums, niche blogs, self-hosted publications, and academic repositories that commercial crawlers deprioritize, blending these unique nodes into search results.

Stage 3: The Open, Distributed Search Index

The ultimate vision is constructing an open, community-accessible, distributed web index—structured, ranked, and freely available for any engineer, organization, or sovereign homelab to run their own custom search interface and ranking models.

Recent Coverage and September 2026 Status

Christine Hall's FOSS Force coverage, published August 15, 2026, is the latest independently verifiable article found in this review. It describes the current approach as an alpha that uses the top 20 Google results as raw material, then summarizes the planned small-web crawl and open distributed-index stages. It also records a proposed sustainability model: developer API access to curated data and clearly disclosed affiliate links kept separate from ranking.

A fresh check of the public repository on September 7, 2026 found no GitHub Releases and no tags. The latest commit on main is b9d3dc2, dated July 27, 2026, with the message “AI Overview load from database fix + readme update.” That makes NeoSearch useful to study as a public alpha codebase, but not yet a versioned release that should be treated as a stable distribution.

The source tree also makes the architecture more concrete than the roadmap alone. Server/SearchService can serve cached SQL Server results or fetch new web results, then run suggestions, plugin detection, LLM lenses, result rewriting, and AI Overview work in parallel before streaming the response. The repository's web-search documentation identifies RapidAPI's Google Search Master Mega service as the primary provider, with direct Google Custom Search and other RapidAPI paths available as alternatives.

The repository includes a crawler subsystem as well: CrawlerServer exposes a WebSocket /ws route for workers, maintains a domain queue, batches crawl results for database writes, and re-queues work when a worker disconnects. Its documentation describes this as embedded in the Server and SpiderWeb applications rather than as a standalone service. That is evidence of a distributed-crawl design in the codebase—not evidence that the public project has already delivered the open distributed index described in Stage 3.

Privacy by Architecture, Not Just Policy

Unlike commercial search providers that rely on user tracking, the NeoSearch README describes these privacy goals:

Open-source systems can make important design choices more inspectable. NeoSearch is a useful case study in separating search retrieval, answer generation, and ranking decisions.

Tracking NeoSearch: You can explore the live engine at neosearch.org, review and fork the codebase on GitHub, and read coverage on FOSS Force.

Frequently asked questions

What is NeoSearch?

An open-source search engine that retrieves web results, enriches and re-ranks them with LLM-backed components, streams AI-generated overviews and lenses, and includes a crawler subsystem in the public repository.

How is it different from Google or Bing?

The public code is designed to be self-hosted, but it is still an alpha and its documented primary web-search path uses an upstream search API. It should not yet be described as an independent distributed index or as a complete replacement for Google or Bing.

What does an open-source search stack need to be useful?

A crawler that respects the site, a retriever good enough to surface the right passages, and a generator that cites what it retrieved. When all three run locally, you can audit every answer back to the document that produced it.

Research note

This article summarizes public project information and is not an endorsement or deployment claim.