{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Dustin Edwards blog: sqlite",
  "home_page_url": "https://dustinedwards.info/writing/tags/sqlite",
  "feed_url": "https://dustinedwards.info/writing/tags/sqlite/feed.json",
  "description": "Posts tagged sqlite.",
  "language": "en-US",
  "authors": [
    {
      "name": "Dustin Edwards",
      "url": "https://dustinedwards.info"
    }
  ],
  "items": [
    {
      "id": "https://dustinedwards.info/writing/site-search-on-d1",
      "url": "https://dustinedwards.info/writing/site-search-on-d1",
      "title": "Site search on Cloudflare D1 with SQLite full-text search",
      "content_text": "\nThis article describes how to build full-text site search directly on Cloudflare D1 using SQLite's [FTS5 extension](https://sqlite.org/fts5.html), with no external search service. The implementation this describes runs in production on this site and answers queries in 6 milliseconds at the median, 15 at the 95th percentile, measured over 25 runs against local D1. Those are DATABASE query times, not page latency, and the distinction is worth making before the number travels: the same search requested over the public HTTPS endpoint measured 114 to 121 ms end to end on 2026-08-04, nearly all of the difference being network round trip rather than work. The article covers the schema, the reason one index is not enough, the ranking method, the query dispatch that a naive design gets wrong, the interface work, and one operational finding about backups that I consider mandatory knowledge for anyone putting FTS5 on D1.\n\nPrerequisites: a D1 database, familiarity with SQL and SQLite migrations, and content you can decompose into records. The design generalizes to any corpus; the examples are a blog.\n\n## Step 1: index at section granularity, with one record shape\n\nReduce everything searchable to a single record shape. Mine is: id, url, type, title, body, date. The consequential decision is granularity. A search that indexes whole documents sends the reader to a page and leaves the finding to them; a search that indexes sections sends them to the paragraph. If your content has heading anchors, emit one record per document plus one record per heading, each section record carrying a URL that deep-links to its anchor.\n\nTwo practical notes on granularity. First, it is what makes ranking observable during development: with document-level records and a small corpus, almost any query returns almost everything, and you cannot tell whether your ranking works. Section records give the ranker real decisions to make from the first day. Second, it requires query-time deduplication: when a document and its own sections both match, present the best section under the document's title, because one document should not fill a results page with itself.\n\n## Step 2: FTS5 tokenizers are per-table, so stemmed plus exact matching needs two indexes\n\nHere is the FTS5 fact that determines the schema, and it surprised me: the tokenizer is a property of the table, not of the query. You cannot ask one index for stemmed matching on some queries and exact matching on others. A corpus that needs both, and most do, needs two tables.\n\nThe need for both is easy to demonstrate on real data. Prose wants stemming: a search for \"indexing\" should match a sentence containing \"index.\" Names and identifiers want the opposite: a search for \"edwards\" must match \"Edwards\" exactly, and a search for a partial name should not fuzzily match through a stemmer. On my production corpus, the term \"enforcement\" matched the exact-token index while its stem \"enforce\" returned zero rows from it, and the Porter-stemmed index matched both forms. One table cannot produce both behaviors.\n\nThe schema, as migration SQL:\n\n```sql\nCREATE VIRTUAL TABLE search_identity USING fts5(\n  title, tags,\n  content='search_docs', content_rowid='rowid',\n  tokenize='unicode61 remove_diacritics 2'\n);\n\nCREATE VIRTUAL TABLE search_prose USING fts5(\n  title, body,\n  content='search_docs', content_rowid='rowid',\n  tokenize='porter unicode61'\n);\n```\n\nBoth are external-content tables over one `search_docs` source table, so the text is stored once. On rebuilds, I rewrite `search_docs` wholesale and rebuild both indexes with the FTS5 `rebuild` command, because my corpus is regenerated as a set; per-row triggers are the right tool only for a path that edits single rows.\n\n## Step 3: merge with reciprocal rank fusion, not raw bm25 scores\n\nTwo indexes produce two ranked lists, and the tempting merge, interleaving by raw bm25 score, is wrong in a way that ships quietly. Bm25 scores are not comparable across tables with different tokenizers and different average document lengths, and bm25 systematically over-rewards very short rows, which section records are. You will not notice in testing; you will notice when a two-line section outranks the document that answers the query.\n\nThe standard remedy is reciprocal rank fusion, introduced by Cormack, Clarke, and Buettcher in 2009: ignore the scores entirely and combine by position. Each result contributes `1 / (k + rank)` from each list it appears in, summed. The constant k damps the advantage of top ranks; the original paper's value of 60 works well and I did not tune it. In TypeScript:\n\n```ts\nfunction fuse(lists: string[][], k = 60): Map<string, number> {\n  const scores = new Map<string, number>();\n  for (const list of lists) {\n    list.forEach((id, i) => {\n      scores.set(id, (scores.get(id) ?? 0) + 1 / (k + i + 1));\n    });\n  }\n  return scores;\n}\n```\n\nThe reason k matters is easier to see than to describe. With k at 60, the gap between rank 1 and rank 2 is small, so appearing in both lists at moderate rank beats appearing in one list at the top; with k near zero, rank 1 dominates everything and the fusion degenerates into \"whichever list you trust more.\" The curve below plots each rank's contribution at k equal to 60, computed directly from the formula:\n\n:::chart{type=\"line\" x=\"rank\" y=\"contribution\" title=\"RRF contribution by rank position, k = 60\" alt=\"Line chart of reciprocal rank fusion contribution per rank for k equal to 60, falling gently from 0.0164 at rank 1 to 0.0125 at rank 20. The curve is nearly flat, showing that k at 60 keeps top ranks from dominating the fusion.\"}\n```csv\nrank,contribution\n1,0.01639\n2,0.01613\n3,0.01587\n4,0.01563\n5,0.01538\n7,0.01493\n10,0.01429\n13,0.01370\n16,0.01316\n20,0.01250\n```\nPer-rank contribution 1 / (k + rank) at k = 60. The near-flat curve is the design: membership in multiple lists outweighs position within one.\n:::\n\nPositional fusion has a second benefit beyond correctness: it makes the merge testable with small fixtures, because the expected output depends only on orderings you construct, not on opaque score values.\n\n## Step 4: search versus browse: dispatch on query shape\n\nIn front of the indexes, put a small parser: quoted phrases pass through, `tag:` and `type:` prefixes become filters, and a bare four-digit year becomes a date filter rather than a literal search term. That last rule is high-value and produced the one bug in this system that reached production, which I will describe as a warning because the design error is general.\n\nEvery filter-only query returned zero results on the live site. A bare year, a click on a tag chip, any query that was all filter and no text: empty. The parser was working correctly; it converted the year to a date filter and left the text empty, and the search function, having no text to hand FTS5, short-circuited to no results. The interface made it worse by rendering tag chips that were standing invitations to run exactly the queries that failed.\n\nThe structural fix is recognizing that search has two entry modes. A query with text is a locate operation and goes to the indexes. A query with filters and no text is a browse operation and goes to ordinary filtered SQL over the source table, ordered by date, returning document records only, since \"show me everything tagged d1\" is a listing question. Put the dispatch predicate in a pure function so it can be unit-tested, and test the browse path's visibility rules (drafts and future-dated content excluded) as deliberately as the search path's. The general statement: if you test only the entry mode with a text box, you have shipped half a feature.\n\n## Step 5: the search interface: GET form, JSON negotiation, and the combobox pattern\n\nThe baseline is a server-rendered GET form: deep-linkable result URLs, highlighted snippets from FTS5's snippet function, visible labels for why a result matched, facet chips as plain links, and a zero-results state that suggests nearest tags and recent posts instead of dead-ending. Because it is a GET endpoint, adding `Vary: Accept` and returning JSON under content negotiation makes the same URL a machine-readable API at no extra cost, which matters more each year as AI agents become a real audience.\n\nA command palette can layer on top. If you build one, implement the ARIA combobox pattern as specified rather than approximately: focus stays in the input, `aria-activedescendant` tracks the highlighted option, arrow keys move the highlight rather than focus. Two implementation findings from doing this that will save you time. An input with `type=\"search\"` swallows the first Escape keypress natively to clear its own value, so your close handler fires on the second press unless you account for it. And an in-flight fetch that resolves after the palette closes will repaint a closed dialog, leaving `aria-expanded=\"true\"` over stale options; cancel or discard responses that arrive after close. Both bugs only appear when you test the unhappy orderings, which is the reason to test the unhappy orderings.\n\n## wrangler d1 export fails on FTS5: the working backup procedure\n\nThe most important operational finding in this article: `wrangler d1 export` fails outright on any database containing FTS5 virtual tables. It exits with an error stating it cannot export databases with virtual tables, and writes nothing. This means the moment you apply the search migration, the platform's default backup path stops working for your database, and the natural time to discover that is during a recovery, which is the worst time. I measured it before applying the migration, on purpose, and I would recommend the same order to anyone.\n\nThe working procedure is per-table export, with schema coming from your migration files rather than the dump:\n\n```bash\nnpx wrangler d1 export mydb --remote --no-schema \\\n  --table posts --output export-posts.sql\n```\n\nExport each real table this way and never the FTS tables or their `_config`, `_data`, `_docsize`, and `_idx` shadow tables; a restore is migrations first, then per-table data. Then encode the table list in a check script that derives it from your migrations directory and fails when the two disagree in either direction, because a backup procedure that exists only in memory is not a procedure. Mine found a real omission on its first run.\n\nTwo adjacent facts from the same investigation, both counterintuitive. Verifying an external-content FTS5 index with `COUNT(*)` cannot detect corruption or emptiness, because the count reads through to the content table and reports its row count regardless of index state; count the `_docsize` shadow table instead. And running `DELETE FROM` directly against an FTS5 table corrupts the index in a way that surfaces only on a later write, with the repair being the FTS5 `rebuild` command. Neither behavior is a D1 defect; both are documented SQLite semantics that become sharp when the database is remote and the tooling is young.\n\n## Results and limitations\n\nOn this hardware and corpus: median 6 ms, 95th percentile 15 ms, over 25 runs against local D1, with the production numbers in the same range. **That is the query, not the page.** Requested end to end over HTTPS, the same search measured 114 to 121 ms on 2026-08-04, and a reader who does not know which of the two a benchmark reports cannot use either. The latency figures establish a floor rather than a curve; the corpus was small when measured, and I will re-measure as it grows. The two-index design is justified by a corpus that needs both stemmed and identity matching; a site with only prose could defensibly run one Porter-stemmed table and skip the fusion. The export failure is as measured on the wrangler version current at writing and may be fixed later; the per-table procedure and its check remain worthwhile regardless, because a backup that depends on a bug staying fixed is not a backup. And the general claim I would defend beyond this stack: at personal-site scale and probably well past it, hand-built search on the relational database you already operate is not the compromise option. Measured against the alternative of introducing and paying for a search service, it was the fast path in both senses.\n\nThis is the fifth post in [the series](/writing/ten-years-on-cloudflare), following [the reading experience](/writing/blog-reading-without-javascript); the next one adds [the layer above this one](/writing/ai-answer-mode-on-site-search): a retrieval-augmented answer mode, and the cost controls a public AI endpoint requires.\n\n## Update, August 2026\n\nThe `_docsize` counting trick from the backup section grew into this system's standing integrity instrument, and it is the part of this article I would now emphasize hardest. The deploy pipeline asserts at every ship that `search_docs` and both indexes' `_docsize` shadow tables agree on the count, three numbers from three places that can only match if the rebuild actually reached both indexes, and a scheduled workflow now polls a health endpoint every fifteen minutes running the same equalities, so an index that silently loses records is a fifteen-minute discovery rather than a someday one. That instrument earned its keep twice in August: once catching a drift that appeared and cleared between two polls during an afternoon of deploys, and once, more embarrassingly, when an external audit and I both misdescribed which tables the ship-time equality actually compares, which was settled the only way these things settle, by reading the code instead of the memory of it. The lesson fits this article's backup section exactly: a verification procedure that exists only in memory drifts like any other second copy, and the cure is the same, derive it, assert it, and let the instrument answer instead of you.\n",
      "summary": "How to build site search on Cloudflare D1 with SQLite FTS5: two indexes for stemmed and exact matching, reciprocal rank fusion in place of raw bm25, section-level records, a browse path for filter-only queries, and the D1 export problem every FTS5 user has.",
      "date_published": "2026-07-28T00:00:00.000Z",
      "date_modified": "2026-09-27T00:00:00.000Z",
      "tags": [
        "cloudflare",
        "d1",
        "fts5",
        "search",
        "sqlite"
      ]
    }
  ]
}
