{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Dustin Edwards blog: search",
  "home_page_url": "https://dustinedwards.info/writing/tags/search",
  "feed_url": "https://dustinedwards.info/writing/tags/search/feed.json",
  "description": "Posts tagged search.",
  "language": "en-US",
  "authors": [
    {
      "name": "Dustin Edwards",
      "url": "https://dustinedwards.info"
    }
  ],
  "items": [
    {
      "id": "https://dustinedwards.info/writing/site-search-on-d1",
      "url": "https://dustinedwards.info/writing/site-search-on-d1",
      "title": "Site search on Cloudflare D1 with SQLite full-text search",
      "content_text": "\nThis article describes how to build full-text site search directly on Cloudflare D1 using SQLite's [FTS5 extension](https://sqlite.org/fts5.html), with no external search service. The implementation this describes runs in production on this site and answers queries in 6 milliseconds at the median, 15 at the 95th percentile, measured over 25 runs against local D1. Those are DATABASE query times, not page latency, and the distinction is worth making before the number travels: the same search requested over the public HTTPS endpoint measured 114 to 121 ms end to end on 2026-08-04, nearly all of the difference being network round trip rather than work. The article covers the schema, the reason one index is not enough, the ranking method, the query dispatch that a naive design gets wrong, the interface work, and one operational finding about backups that I consider mandatory knowledge for anyone putting FTS5 on D1.\n\nPrerequisites: a D1 database, familiarity with SQL and SQLite migrations, and content you can decompose into records. The design generalizes to any corpus; the examples are a blog.\n\n## Step 1: index at section granularity, with one record shape\n\nReduce everything searchable to a single record shape. Mine is: id, url, type, title, body, date. The consequential decision is granularity. A search that indexes whole documents sends the reader to a page and leaves the finding to them; a search that indexes sections sends them to the paragraph. If your content has heading anchors, emit one record per document plus one record per heading, each section record carrying a URL that deep-links to its anchor.\n\nTwo practical notes on granularity. First, it is what makes ranking observable during development: with document-level records and a small corpus, almost any query returns almost everything, and you cannot tell whether your ranking works. Section records give the ranker real decisions to make from the first day. Second, it requires query-time deduplication: when a document and its own sections both match, present the best section under the document's title, because one document should not fill a results page with itself.\n\n## Step 2: FTS5 tokenizers are per-table, so stemmed plus exact matching needs two indexes\n\nHere is the FTS5 fact that determines the schema, and it surprised me: the tokenizer is a property of the table, not of the query. You cannot ask one index for stemmed matching on some queries and exact matching on others. A corpus that needs both, and most do, needs two tables.\n\nThe need for both is easy to demonstrate on real data. Prose wants stemming: a search for \"indexing\" should match a sentence containing \"index.\" Names and identifiers want the opposite: a search for \"edwards\" must match \"Edwards\" exactly, and a search for a partial name should not fuzzily match through a stemmer. On my production corpus, the term \"enforcement\" matched the exact-token index while its stem \"enforce\" returned zero rows from it, and the Porter-stemmed index matched both forms. One table cannot produce both behaviors.\n\nThe schema, as migration SQL:\n\n```sql\nCREATE VIRTUAL TABLE search_identity USING fts5(\n  title, tags,\n  content='search_docs', content_rowid='rowid',\n  tokenize='unicode61 remove_diacritics 2'\n);\n\nCREATE VIRTUAL TABLE search_prose USING fts5(\n  title, body,\n  content='search_docs', content_rowid='rowid',\n  tokenize='porter unicode61'\n);\n```\n\nBoth are external-content tables over one `search_docs` source table, so the text is stored once. On rebuilds, I rewrite `search_docs` wholesale and rebuild both indexes with the FTS5 `rebuild` command, because my corpus is regenerated as a set; per-row triggers are the right tool only for a path that edits single rows.\n\n## Step 3: merge with reciprocal rank fusion, not raw bm25 scores\n\nTwo indexes produce two ranked lists, and the tempting merge, interleaving by raw bm25 score, is wrong in a way that ships quietly. Bm25 scores are not comparable across tables with different tokenizers and different average document lengths, and bm25 systematically over-rewards very short rows, which section records are. You will not notice in testing; you will notice when a two-line section outranks the document that answers the query.\n\nThe standard remedy is reciprocal rank fusion, introduced by Cormack, Clarke, and Buettcher in 2009: ignore the scores entirely and combine by position. Each result contributes `1 / (k + rank)` from each list it appears in, summed. The constant k damps the advantage of top ranks; the original paper's value of 60 works well and I did not tune it. In TypeScript:\n\n```ts\nfunction fuse(lists: string[][], k = 60): Map<string, number> {\n  const scores = new Map<string, number>();\n  for (const list of lists) {\n    list.forEach((id, i) => {\n      scores.set(id, (scores.get(id) ?? 0) + 1 / (k + i + 1));\n    });\n  }\n  return scores;\n}\n```\n\nThe reason k matters is easier to see than to describe. With k at 60, the gap between rank 1 and rank 2 is small, so appearing in both lists at moderate rank beats appearing in one list at the top; with k near zero, rank 1 dominates everything and the fusion degenerates into \"whichever list you trust more.\" The curve below plots each rank's contribution at k equal to 60, computed directly from the formula:\n\n:::chart{type=\"line\" x=\"rank\" y=\"contribution\" title=\"RRF contribution by rank position, k = 60\" alt=\"Line chart of reciprocal rank fusion contribution per rank for k equal to 60, falling gently from 0.0164 at rank 1 to 0.0125 at rank 20. The curve is nearly flat, showing that k at 60 keeps top ranks from dominating the fusion.\"}\n```csv\nrank,contribution\n1,0.01639\n2,0.01613\n3,0.01587\n4,0.01563\n5,0.01538\n7,0.01493\n10,0.01429\n13,0.01370\n16,0.01316\n20,0.01250\n```\nPer-rank contribution 1 / (k + rank) at k = 60. The near-flat curve is the design: membership in multiple lists outweighs position within one.\n:::\n\nPositional fusion has a second benefit beyond correctness: it makes the merge testable with small fixtures, because the expected output depends only on orderings you construct, not on opaque score values.\n\n## Step 4: search versus browse: dispatch on query shape\n\nIn front of the indexes, put a small parser: quoted phrases pass through, `tag:` and `type:` prefixes become filters, and a bare four-digit year becomes a date filter rather than a literal search term. That last rule is high-value and produced the one bug in this system that reached production, which I will describe as a warning because the design error is general.\n\nEvery filter-only query returned zero results on the live site. A bare year, a click on a tag chip, any query that was all filter and no text: empty. The parser was working correctly; it converted the year to a date filter and left the text empty, and the search function, having no text to hand FTS5, short-circuited to no results. The interface made it worse by rendering tag chips that were standing invitations to run exactly the queries that failed.\n\nThe structural fix is recognizing that search has two entry modes. A query with text is a locate operation and goes to the indexes. A query with filters and no text is a browse operation and goes to ordinary filtered SQL over the source table, ordered by date, returning document records only, since \"show me everything tagged d1\" is a listing question. Put the dispatch predicate in a pure function so it can be unit-tested, and test the browse path's visibility rules (drafts and future-dated content excluded) as deliberately as the search path's. The general statement: if you test only the entry mode with a text box, you have shipped half a feature.\n\n## Step 5: the search interface: GET form, JSON negotiation, and the combobox pattern\n\nThe baseline is a server-rendered GET form: deep-linkable result URLs, highlighted snippets from FTS5's snippet function, visible labels for why a result matched, facet chips as plain links, and a zero-results state that suggests nearest tags and recent posts instead of dead-ending. Because it is a GET endpoint, adding `Vary: Accept` and returning JSON under content negotiation makes the same URL a machine-readable API at no extra cost, which matters more each year as AI agents become a real audience.\n\nA command palette can layer on top. If you build one, implement the ARIA combobox pattern as specified rather than approximately: focus stays in the input, `aria-activedescendant` tracks the highlighted option, arrow keys move the highlight rather than focus. Two implementation findings from doing this that will save you time. An input with `type=\"search\"` swallows the first Escape keypress natively to clear its own value, so your close handler fires on the second press unless you account for it. And an in-flight fetch that resolves after the palette closes will repaint a closed dialog, leaving `aria-expanded=\"true\"` over stale options; cancel or discard responses that arrive after close. Both bugs only appear when you test the unhappy orderings, which is the reason to test the unhappy orderings.\n\n## wrangler d1 export fails on FTS5: the working backup procedure\n\nThe most important operational finding in this article: `wrangler d1 export` fails outright on any database containing FTS5 virtual tables. It exits with an error stating it cannot export databases with virtual tables, and writes nothing. This means the moment you apply the search migration, the platform's default backup path stops working for your database, and the natural time to discover that is during a recovery, which is the worst time. I measured it before applying the migration, on purpose, and I would recommend the same order to anyone.\n\nThe working procedure is per-table export, with schema coming from your migration files rather than the dump:\n\n```bash\nnpx wrangler d1 export mydb --remote --no-schema \\\n  --table posts --output export-posts.sql\n```\n\nExport each real table this way and never the FTS tables or their `_config`, `_data`, `_docsize`, and `_idx` shadow tables; a restore is migrations first, then per-table data. Then encode the table list in a check script that derives it from your migrations directory and fails when the two disagree in either direction, because a backup procedure that exists only in memory is not a procedure. Mine found a real omission on its first run.\n\nTwo adjacent facts from the same investigation, both counterintuitive. Verifying an external-content FTS5 index with `COUNT(*)` cannot detect corruption or emptiness, because the count reads through to the content table and reports its row count regardless of index state; count the `_docsize` shadow table instead. And running `DELETE FROM` directly against an FTS5 table corrupts the index in a way that surfaces only on a later write, with the repair being the FTS5 `rebuild` command. Neither behavior is a D1 defect; both are documented SQLite semantics that become sharp when the database is remote and the tooling is young.\n\n## Results and limitations\n\nOn this hardware and corpus: median 6 ms, 95th percentile 15 ms, over 25 runs against local D1, with the production numbers in the same range. **That is the query, not the page.** Requested end to end over HTTPS, the same search measured 114 to 121 ms on 2026-08-04, and a reader who does not know which of the two a benchmark reports cannot use either. The latency figures establish a floor rather than a curve; the corpus was small when measured, and I will re-measure as it grows. The two-index design is justified by a corpus that needs both stemmed and identity matching; a site with only prose could defensibly run one Porter-stemmed table and skip the fusion. The export failure is as measured on the wrangler version current at writing and may be fixed later; the per-table procedure and its check remain worthwhile regardless, because a backup that depends on a bug staying fixed is not a backup. And the general claim I would defend beyond this stack: at personal-site scale and probably well past it, hand-built search on the relational database you already operate is not the compromise option. Measured against the alternative of introducing and paying for a search service, it was the fast path in both senses.\n\nThis is the fifth post in [the series](/writing/ten-years-on-cloudflare), following [the reading experience](/writing/blog-reading-without-javascript); the next one adds [the layer above this one](/writing/ai-answer-mode-on-site-search): a retrieval-augmented answer mode, and the cost controls a public AI endpoint requires.\n\n## Update, August 2026\n\nThe `_docsize` counting trick from the backup section grew into this system's standing integrity instrument, and it is the part of this article I would now emphasize hardest. The deploy pipeline asserts at every ship that `search_docs` and both indexes' `_docsize` shadow tables agree on the count, three numbers from three places that can only match if the rebuild actually reached both indexes, and a scheduled workflow now polls a health endpoint every fifteen minutes running the same equalities, so an index that silently loses records is a fifteen-minute discovery rather than a someday one. That instrument earned its keep twice in August: once catching a drift that appeared and cleared between two polls during an afternoon of deploys, and once, more embarrassingly, when an external audit and I both misdescribed which tables the ship-time equality actually compares, which was settled the only way these things settle, by reading the code instead of the memory of it. The lesson fits this article's backup section exactly: a verification procedure that exists only in memory drifts like any other second copy, and the cure is the same, derive it, assert it, and let the instrument answer instead of you.\n",
      "summary": "How to build site search on Cloudflare D1 with SQLite FTS5: two indexes for stemmed and exact matching, reciprocal rank fusion in place of raw bm25, section-level records, a browse path for filter-only queries, and the D1 export problem every FTS5 user has.",
      "date_published": "2026-07-28T00:00:00.000Z",
      "date_modified": "2026-09-27T00:00:00.000Z",
      "tags": [
        "cloudflare",
        "d1",
        "fts5",
        "search",
        "sqlite"
      ]
    },
    {
      "id": "https://dustinedwards.info/writing/ai-answer-mode-on-site-search",
      "url": "https://dustinedwards.info/writing/ai-answer-mode-on-site-search",
      "title": "Adding an AI answer mode to site search with Cloudflare AI Search",
      "content_text": "\nThis article describes how to add a retrieval-augmented generation (RAG) answer mode to a site using Cloudflare AI Search: a public endpoint that streams a cited answer synthesized from your own content. [The previous article in this series](/writing/site-search-on-d1) covered the classic keyword engine underneath; this one covers the AI layer on top, and it is organized around the three requirements I set before building, because they are the requirements I would recommend to anyone adding a similar layer. The AI mode must never block or degrade the classic path. It must be removable without a trace. And because it is the one public endpoint that costs money per request, it must sit behind cost controls whose behavior is measured rather than assumed.\n\nPrerequisites: a Workers project, content decomposable into records with stable anchors, and the willingness to probe a beta product before trusting it.\n\n## Step 1: make the AI layer removable, and prove it by removing it\n\nRender classic results first and never await any AI call on their path. Gate the AI feature's visibility on the presence of its binding, so that removing the binding removes the affordance rather than breaking it. Then verify removability the only way that counts: remove it. I deleted the binding, confirmed the answer route returned 404, and confirmed the classic search response was byte-identical to its pre-AI form. \"Byte-identical without it\" is a checkable standard; \"degrades gracefully\" is not, and I would hold any enhancement layer to the first.\n\nThe same discipline pays during incidents: if the beta product misbehaves or the pricing changes unfavorably, the exit is one configuration change, and knowing that changes how much risk you can accept everywhere else.\n\n## Step 2: semantic search versus keyword search: measure on your own corpus\n\nThe common assumption is that semantic retrieval subsumes keyword search. Test it on your own corpus before believing it in either direction. My procedure: a shared query set run against both layers, scoring which results each found that the other missed. The outcome on this corpus: the classic FTS5 engine found two results the AI retrieval missed, both exact-token queries, where the semantic layer's similarity scores fell below the instance's relevance threshold for short literal tokens. The AI retrieval found three the classic engine missed, all natural-language questions phrased in words the documents never use, which keyword matching cannot bridge. Neither subsumes the other; they fail on opposite inputs.\n\nThat measurement is the justification for running both layers, and it took an afternoon. I mention the effort because the alternative, adopting a vendor's benchmark or a blog post's intuition, costs less and is worth what it costs. Complementarity is a property of your corpus and your users' query styles, not of the technology.\n\n## Step 3: AI Search ingestion: uploads versus the crawler, and save-time sync\n\nAI Search offers two ingestion paths: a web crawler over a domain, or direct upload into managed storage. I used uploads, for two reasons that generalize. First, correctness of scope: the crawler crawls the domain as DNS resolves it, and this site's apex still pointed at a legacy installation awaiting cutover, so a crawl would have faithfully indexed the wrong site. Check what your domain actually serves before pointing a crawler at it. Second, citation quality: uploading the same section-grained records the classic engine uses, keyed to heading anchors, means every citation in a generated answer deep-links to a section that exists. I verified this by asking questions and following every citation to its anchor.\n\nStaleness is the standing failure mode of any RAG system, and I would treat the sync design as first-class rather than a cron job added later. Here, a content save uploads that post's own section records and invalidates cached answers, incrementally, which is possible because section decomposition is a pure function of one post's source. Lag is seconds. Two design rules attached: the sync must not be able to fail the save (index freshness is worth less than write reliability), and the admin should display index counts against source counts with a one-button repair, because a drift you can see is a maintenance task and a drift you cannot see is a slowly wrong product.\n\nOne operational fact that is invisible in the documentation and worth stating: no credential is needed at runtime. The API token exists only for the control-plane call that creates the instance; the serving path runs entirely on the Worker binding. Your secret inventory should reflect that, and mine does: the creation token is deletable.\n\n## Step 4: rate limit, cache, budget ceiling: three cost gates ordered cheapest first\n\nA public, unauthenticated endpoint that performs a model generation per request is an open invitation to spend your money. Put three gates in front of it, in this order, so that each request hits the cheapest applicable control: a per-IP burst limit, then an answer cache keyed on the normalized question, then a daily budget ceiling. The ordering means a cache hit costs no budget and a rate-limited request costs no AI call. Return refusals as 429 with a Retry-After header. And make the cache observable from outside: every response here carries a header naming hit or miss, which converts \"is the cache working\" from a dashboard question into a curl command, and the same question asked twice returns byte-identical bytes with the hit marker.\n\nThe implementation of those gates produced the most transferable measurements in this article, so I will report them as the test sequence I would now recommend to anyone.\n\n## Step 5: Durable Objects are single-threaded, not atomic: measure your rate limiter\n\nConfigure a limit, then attack it with genuinely concurrent requests and count what gets through. Three implementations, same test shape, very different results.\n\nCloudflare's built-in rate-limiting binding, configured to allow five requests per sixty seconds and attacked with twelve concurrent requests, refused one, then two, then nine, then zero across four runs. A limiter honouring its own limit refuses seven of twelve every time, so those runs let eleven requests through, then ten, then three, then all twelve. This is consistent with its documentation, which describes it as permissive and eventually consistent: it sheds sustained load. It does not count, and for a budget guard you need a counter. The four runs are worth seeing side by side, because the variance is the finding:\n\n:::chart{type=\"bar\" x=\"run\" y=\"refused\" title=\"Requests refused by the built-in rate-limiting binding, limit 5\" alt=\"Bar chart of four identical test runs against Cloudflare's rate-limiting binding configured to allow 5 requests per minute, attacked with 12 concurrent requests. The binding refused 1, then 2, then 9, then 0 requests. A limiter honouring the limit would refuse 7 every time.\"}\n```csv\nrun,refused\nrun 1,1\nrun 2,2\nrun 3,9\nrun 4,0\n```\nFour identical runs: 12 concurrent requests against a configured limit of 5 per 60 seconds. Refusals scatter from 0 to 9 where a counting limiter would refuse 7 every time, consistent with the documented eventually-consistent design. It sheds load; it does not count.\n:::\n\nA Durable Object using the asynchronous storage API, with a read-increment-write sequence, admitted eight requests through a ceiling of three. The reason is the finding I most want to pass along: a Durable Object is single-threaded, but a read and a write separated by an await are not atomic, because other requests interleave at the await point. Single-threaded and transactional are different properties, and the difference only appears under concurrent load.\n\nThe same Durable Object rewritten on the synchronous SQLite storage API, where the read-modify-write happens with no await between, admitted exactly three of fourteen through the ceiling and exactly five of fourteen per IP. The synchronous API is the reason the newer SQLite-backed class registration exists, and this use case is the argument for it. A fixed window still admits up to double the limit across a window boundary, which I measured at ten of twelve and accepted as ordinary; a sliding window costs more bookkeeping and was not warranted here.\n\nOne warning about the test harness itself, because it produced a confident false negative before it produced data: a sequential loop is not a burst. Thirty-six requests, each awaiting completion, produced zero refusals against a working limit of thirty per minute, because the polite loop walked across the window boundary. The limiter looked dead and was fine. Attack with real concurrency (`Promise.all`, not a for-await loop), or your test measures your patience rather than your guard.\n\n## Step 6: expose search to AI agents: JSON negotiation and MCP\n\nOnce the endpoint exists, three levels of machine access come nearly free, and I would ship all three. The search URL itself, constructible by anyone. JSON from the same URL under content negotiation. And the [Model Context Protocol](https://modelcontextprotocol.io) endpoint that the AI Search instance can expose, which lets an AI assistant query the site conversationally through a standard protocol; verify it from an external client before advertising it. Document all three in llms.txt. The removability rule from step 1 applies at every level: each is a presentation of the same engine, and turning any off changes nothing underneath.\n\n## Costs, stated plainly, and limitations\n\nAt the time of writing, retrieval on AI Search is free during its open beta with pricing promised on notice, and answer generation bills through Workers AI per uncached request. The daily ceiling makes worst-case spend a number I chose; the cache converts repeated questions into free reads; and there is a dated entry in the project's decision log requiring a cost re-evaluation when beta pricing lands. I would generalize that habit: using a beta product is reasonable when the exposure is bounded and the re-evaluation is scheduled, and it is the scheduling that tends to be skipped.\n\nMeasured latency, for expectation-setting: time to first token between 2.1 and 6.5 seconds warm and 7.4 cold, with retrieved sources rendered before the answer begins so the wait is visibly progress. The complementarity measurement in step 2 was run on a small corpus and query set; it is strong enough to establish that neither layer subsumes the other here and far too small to estimate rates, and it should be re-run as any corpus grows. The retrieval threshold behavior around short exact tokens is a property of this instance's configuration rather than a universal constant. And the rate limiter measurements reflect one platform's bindings at one point in time; the durable finding is the atomicity mechanism, which is not vendor-specific at all.\n\nThis is the sixth post in [the series](/writing/ten-years-on-cloudflare), following [the FTS5 search engine](/writing/site-search-on-d1). The next two move from reading to writing: [where policy belongs when agents call your service](/writing/policy-in-the-api-not-the-mcp), and [the trust model for an AI agent with write access](/writing/agent-write-access-to-a-live-site).\n\n## Update, August 2026\n\nOne of the gates described above changed shape after adversarial review, and one incident proved the sync design's honesty clause. The daily budget ceiling is no longer a flat number: a flat cap fails as a cliff, because enough addresses can spend the whole day's allowance in minutes while each stays inside the per-IP limit, leaving the feature dead until midnight with the bill untouched. The same daily allowance is now released evenly across the day with a small burst above the pace, so ordinary readers never meet the pacing, a coordinated spend exhausts minutes rather than hours, and the denial ends when the abuse does. Retry-After changed with it, since pointing a reader at midnight when the allowance recovers in minutes was about to become a lie. The remaining honest limit: a caller spending continuously all day can hold the allowance at its edge all day.\n\nA gate that was not described above has since been added: the endpoint now refuses cross-origin browser posts before the rate limit and before any billed work, which is the right instrument here because Ask is anonymous and the usual cookie-based protections do nothing for it, while a request carrying no Origin at all, which is what a scriptless form post looks like, still works.\n\nThe drift the admin badge exists to catch happened once for real: the index silently lost nine records at the end of July and the badge surfaced it three weeks later, which is what \"a drift you can see is a maintenance task\" looks like when the seeing is a human glancing at a number. That gap is why the site now has scheduled alerting: an external workflow polls a health endpoint every fifteen minutes, the endpoint runs the index-equality checks among others, and a failing scheduled run is itself the alert. The first thing that alerting caught was a drift that appeared and cleared between two polls during an afternoon of deploys, which is the badge story inverted: fifteen minutes of visibility instead of three weeks.\n",
      "summary": "How to add a retrieval-augmented answer mode to a site with Cloudflare AI Search: uploaded storage versus the crawler, save-time index sync, three cost gates in front of a paying endpoint, and the Durable Objects atomicity measurement behind them.",
      "date_published": "2026-07-28T00:00:00.000Z",
      "date_modified": "2026-09-27T00:00:00.000Z",
      "tags": [
        "ai-search",
        "cloudflare",
        "durable-objects",
        "search",
        "workers-ai"
      ]
    }
  ]
}
