{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Dustin Edwards blog: d1",
  "home_page_url": "https://dustinedwards.info/writing/tags/d1",
  "feed_url": "https://dustinedwards.info/writing/tags/d1/feed.json",
  "description": "Posts tagged d1.",
  "language": "en-US",
  "authors": [
    {
      "name": "Dustin Edwards",
      "url": "https://dustinedwards.info"
    }
  ],
  "items": [
    {
      "id": "https://dustinedwards.info/writing/ten-years-on-cloudflare",
      "url": "https://dustinedwards.info/writing/ten-years-on-cloudflare",
      "title": "Every Cloudflare product, and which ones this site runs on",
      "content_text": "\nThere is no server behind this site. Ten years ago I put my first domain behind Cloudflare the way everyone did then: a shared HostGator box ran the real site and Cloudflare was the DNS, the cache, and the orange cloud in front of it. It was not a place where software ran, and in 2016 it mostly wasn't. This year I rebuilt so that the cloud is the whole thing. The pages, the database, the uploads, the search, the AI answers, the alert mail, and the publishing pipeline all run on Cloudflare products, and nothing else is in the stack.\n\nSo this is the post I wanted when I started: every developer product Cloudflare sells as of September 8, 2026, what each one does in plain words, and whether this site uses it, where, and why or why not. The refusals are in the table with everything else. A survey that only lists what worked is an advertisement. This post replaced an earlier version, published 2026-07-30, that surveyed the same rebuild before the table existed; its dated measurements are carried forward below.\n\n## Which products this site runs on\n\nUsed means a binding or a configured feature that production depends on today. Not used means considered and passed over; the reason is in the product's own entry below. The numbers are dated because every one of them moves.\n\n| Product | What it does | This site | Where |\n|---|---|---|---|\n| Workers | Runs your code on Cloudflare's network | Used | Everything: two Workers, the site and a watchdog |\n| Static Assets | Serves files from a Worker with no invocation | Used | `public/`, through the `ASSETS` binding |\n| Workers Cache | Caches a Worker's responses at the edge | Used | The renderer entrypoint; the gateway is deliberately uncached |\n| D1 | SQLite database, managed | Used | Posts, tags, search index, media index |\n| KV | Fast key-value store, eventually consistent | Used | Login sessions, the Ask answer cache, watchdog state |\n| R2 | Object storage, no egress fees | Used | Three buckets: uploads, social cards, a mirror of uploads |\n| Queues | Message queue between Workers | Used | R2 upload events feeding the media index |\n| Durable Objects | A single-instance object with its own storage | Used | The rate limiter and daily budget for Ask |\n| Analytics Engine | Time-series data you write from a Worker | Used | Per-page traffic counts, no cookies, no IPs |\n| Images | Resize and convert images on request | Used | Every thumbnail and content width |\n| AI Search | Retrieval and cited answers over your content | Used | The Ask endpoint, above classic search |\n| Email Service | Send email from a Worker | Used | The watchdog's alert mail |\n| Email Routing | Receive mail on your domain and forward it | Used | Inbound mail on the domain |\n| Workers Observability | Logs and traces for Workers | Used, logs only | Traces are off on purpose (see the entry) |\n| Cron Triggers | Run a Worker on a schedule | Used | The watchdog, every 15 minutes |\n| Workers AI | Run AI models on Cloudflare GPUs | Indirect | Only through AI Search; no direct binding |\n| Rate Limiting binding | A built-in per-key rate limiter | Refused | Measured: it sheds load, it does not count |\n| Vectorize | Vector database for embeddings | Refused | Two FTS5 indexes answer this corpus |\n| Pages | Hosting for static and framework sites | Not used | Workers with static assets does the same job |\n| Workers Builds | Build and deploy from a git push | Not used | Deploys go through a gated ship script |\n| Hyperdrive | Connection pooling to an external Postgres or MySQL | Not used | There is no external database |\n| Workflows | Durable multi-step jobs with retries | Not used | Nothing here runs long enough |\n| Containers | Run any container next to a Worker | Not used | Nothing needs a runtime beyond V8 |\n| Sandboxes | Isolated code execution for agents | Not used | Agents write through an API, not by running code |\n| Browser Run | Headless browser as a service | Not used | Charts and diagrams render at build time |\n| Workers Agents SDK | Framework for stateful AI agents | Not used | The agent surface is an MCP server on plain Workers |\n| AI Gateway | Proxy and observability for model calls | Not used | Ask's one model call is metered by a Durable Object |\n| Stream | Video hosting and playback | Not used | No video |\n| RealtimeKit | Live audio and video | Not used | No live features |\n| Pipelines | Streaming ingestion into R2 | Not used | Analytics Engine covers the one stream |\n| Data Platform | Catalog and query data in R2 | Not used | Nothing to catalog |\n| Artifacts | Git-native versioned storage | Not used | GitHub is the repository |\n| Secrets Store | Account-level secret storage | Not used | Nine `wrangler secret` values, gated |\n| Turnstile | Bot check without a captcha | Not yet | Planned for the newsletter form |\n| Web Analytics | Client-side analytics beacon | Refused | Blocked by this site's CSP, and not needed |\n| Zaraz | Third-party tag loading at the edge | Not used | There are no third-party tags |\n| Access | Login in front of an application | Not used | The admin plane uses Better Auth |\n| Cache Reserve | Persistent cache for static content | Not yet | Needs the zone; waits for DNS cutover |\n| Workers for Platforms | Run customers' Workers inside yours | Not used | One customer |\n\n## What each one does, in the words I would use to a colleague\n\n### Workers\n\nA Worker is a function that receives a request and returns a response, running in a V8 isolate on Cloudflare's network in whichever city is closest to the reader. No server to size, no region to choose, cold starts small enough that I have never thought about them. Everything else on this list is something a Worker can be given a binding to.\n\nThis site is two Workers. The first serves every page and holds the entire markdown rendering pipeline, syntax highlighting included, because every public page works with JavaScript disabled. Its upload measured 8.4 MiB on 2026-09-08 (2.1 MiB gzipped; 3.78 MB when the first version of this post published on 2026-07-30). The second is a watchdog that reads the site's health endpoint every fifteen minutes through a service binding and repairs what it can. The client-side enhancement budget for the whole blog, progress bar, table of contents, copy buttons, footnote previews, lightbox, comes to about three kilobytes gzipped; the public plane ships no framework script at all.\n\n### Static Assets\n\nA Worker can serve a directory of files directly from the edge with no code running, through a binding with exactly one method, `fetch()`. That method is the whole interface: a Worker can serve any path it is given and discover none of them, which is why this site's media index reads a committed manifest of what is in `public/` instead of asking the binding.\n\n### Workers Cache\n\nCloudflare can store a Worker's responses at the edge and answer repeat requests without running the code. This site turns it on for the renderer and off for the gateway that sits in front of it. The gateway does three things that must never be skipped, the HTTPS redirect, the theme cookie read, and the traffic count, and a cached gateway would skip all three. Every response without an explicit `Cache-Control` defaults to `private, no-store`, because the platform would otherwise cache a logged-in admin page for two hours under standard heuristics and serve it to anyone. That default is the one line of configuration I would tell every Workers user to check first.\n\n### D1\n\nA relational database, which is to say SQLite, replicated and managed by Cloudflare. D1 spent its early life with a reputation for being a toy and that reputation is stale. It is real SQLite, and real SQLite ships FTS5 full-text search with the database. This site's search is two FTS5 indexes over the same corpus, one unstemmed for names and identifiers and one Porter-stemmed for prose, merged with reciprocal rank fusion. Measured 2026-08-28 against production: median 64 ms warm, 145 ms cold, for the database query alone (6 ms at publication on 2026-07-30, against a smaller corpus with the query timed in isolation; the two were not taken the same way, which is the argument for dating a number). The schema and the fusion method are in [the search article](/writing/site-search-on-d1).\n\nD1 also holds the media index, and that is a capability decision rather than a scale one: R2 lists objects in key order and promises nothing else, so \"sort by date, filter unused, count by type\" each need a query, and a database is where queries live.\n\nThe constraint every D1 user should know: the platform's export command fails outright on any database containing FTS5 tables. The working per-table procedure is documented in the search article, and as of 2026-09-08 a weekly job restores that export into a scratch database and compares it against production, because a backup that has never been restored is a hope. That drill found three bugs in the documented restore path on its first run.\n\n### KV\n\nA key-value store that reads fast anywhere in the world and accepts that a write takes a moment to be seen everywhere. Right for configuration and caches, wrong for counters. This site keeps login sessions in it, the watchdog's alert state, and the Ask answer cache, keyed by a hash of the normalised question. The cache sits in front of the daily spending ceiling rather than behind it, so a repeated question reaches no model and costs nothing.\n\n### R2\n\nObject storage with an S3-compatible API and no charge to read the data back out. This site runs three buckets, split on lifecycle rather than on what the admin UI calls them. Uploads are content-addressed: the key is a hash of the bytes, which was decided on security rather than tidiness, because the old key was guessable from a slug the sitemap publishes. Social cards are derived and regenerable, so they get their own bucket that may be emptied. The third bucket is a mirror of uploads that no code path on this site can delete from; a gate fails the build if one appears, and the health endpoint compares every object against its twin.\n\n### Queues\n\nA message queue: one Worker puts a message on it, another consumes it, with retries and a dead-letter queue for permanent failures. This site's media index is written this way. An upload writes only to R2; R2 emits an event; the consumer derives the database row from the object as it is now, never from what the message claimed. That is what makes replay and out-of-order delivery converge, and it is why the Worker never does a dual write.\n\n### Durable Objects\n\nA single instance of a JavaScript class, addressed by name, with its own storage. Only one copy exists anywhere in the world, so it is the platform's answer to coordination. This site uses one for the two guards in front of Ask, the only public endpoint that costs money per request: a per-IP burst limit and a site-wide daily ceiling.\n\nBuilding it taught me the one fact from this whole rebuild I repeat most often: single-threaded is not transactional. A Durable Object using the asynchronous storage API admitted eight requests through a ceiling of three, because a read and a write separated by an `await` are not atomic. The synchronous SQLite storage API is the fix; the same object rewritten on it admitted exactly three. The measurements are in [the AI answer layer article](/writing/ai-answer-mode-on-site-search).\n\n### Analytics Engine\n\nA write-only time-series store you append to from a Worker and query later with SQL. This site writes one row per HTML response: path, referrer host, country, and a coarse mobile flag. No cookie, no IP address, no identifier of any kind, so nothing joins two requests together. It is here because the alternative, Cloudflare's own Web Analytics beacon, is a third-party script, and this site's Content Security Policy would have to be loosened to admit it.\n\n### Images\n\nResize, crop and convert images on request, from an original you keep in R2. This site derives every thumbnail and every content width from one uploaded original through the binding, and the results are cached by the Workers Cache above. The binding rather than the URL syntax, and that is forced: the URL interface answers 404 on a workers.dev hostname because it needs a customer zone. A detail that stops mattering at DNS cutover.\n\n### AI Search\n\nYou give it a corpus; it handles chunking, embedding, hybrid retrieval, reranking, and answer generation with citations. This site's Ask mode is built on it, above classic search rather than instead of it, because the two fail on opposite inputs. Measured on a shared query set: classic search found two results the AI retrieval missed, all exact tokens, and the AI retrieval found three the classic engine missed, all natural-language questions. Neither subsumes the other. It is a metered product, so it sits behind the Durable Object above and a daily ceiling I chose. Its first token arrived in 2.1 to 6.5 seconds warm when measured on 2026-07-30; that figure has not been re-taken because probing it costs money per request.\n\n### Email Service\n\nSend email from a Worker through a binding, from an address on a domain you have onboarded. The watchdog uses it to mail me once when the site goes unhealthy and once when it recovers, with the state kept in KV so a bad afternoon sends one message rather than twelve. Sending to a verified destination address is free.\n\n### Email Routing\n\nReceive mail on your domain and forward it wherever you like. Inbound mail for dustinedwards.info lands here and forwards to my mailbox. Both halves of the email story share one set of SPF and DKIM records that Cloudflare manages, and DMARC on the domain is set to reject.\n\n### Workers Observability\n\nLogs and traces from your Workers, kept in the dashboard and exportable elsewhere. This site keeps logs on and traces off, and the off is a ruling rather than a default. A trace span carries the full request URL, and this site puts a capability token in the path of every draft preview link, so exporting traces would ship preview access to a third party. Invocation logs are off for a related reason: they recorded the reader's IP and session cookie for seven days with no field-level redaction available.\n\n### Cron Triggers\n\nRun a Worker on a schedule. The watchdog runs every fifteen minutes. The site Worker has no cron, and the empty array in its config is the statement: an hourly trigger that nothing handled sat on the platform for fifteen days in August, throwing 24 times a day, invisible to a gate that only read files. The gate now reads the platform too.\n\n### Workers AI\n\nRun open models on Cloudflare's GPUs from a Worker. This site never calls it directly; AI Search does the model work for Ask on its own. If Ask ever needs a model the retrieval product does not offer, this is where it would come from.\n\n### Rate Limiting binding, refused\n\nA built-in per-key limiter you declare in config. Measured on 2026-07-30 with a limit of five per sixty seconds and twelve concurrent requests: it refused one, then two, then nine, then zero across four runs. Its documentation says it sheds sustained load with eventual consistency, and that is true; it does not count, and a limiter guarding a budget has to count. The Durable Object replaced it.\n\n### Vectorize, refused\n\nA vector database for embeddings, the usual foundation for semantic search. Rank fusion over two FTS5 indexes answers this corpus in about fifteen lines with no vectors, and zero-result searches are the cheapest signal this site has for what to write next. That is the only evidence that would reopen the question.\n\n### Pages and Workers Builds, not used\n\nPages hosts static and framework sites with a build on every push; Workers Builds does the same for Workers. Workers with static assets now does everything Pages did for this site, and Cloudflare's own direction has been to fold Pages into Workers. Deploys here go through a ship script that refuses unless the working tree is clean, CI is green for that exact commit, and every offline gate passes; a build on push would skip all of that.\n\n### Hyperdrive, Workflows, Containers, Sandboxes, Browser Run, not used\n\nHyperdrive pools connections to a Postgres or MySQL you already run somewhere; there is no such database here. Workflows runs multi-step jobs that survive failures and can wait for a human; nothing here runs longer than a request. Containers runs any Docker image next to a Worker; nothing here needs a runtime beyond V8. Sandboxes gives an agent an isolated place to execute code; this site's agents write through an API with policy enforced server-side, and never run code. Browser Run is a headless browser you can drive from a Worker; charts and diagrams here render to SVG at build time, on purpose, so the page carries no work.\n\n### Agents SDK and AI Gateway, not used\n\nThe Agents SDK is a framework for long-lived stateful agents on Durable Objects. This site's agent surface is the other direction: an MCP server that lets an outside agent read and write posts, with exactly one act, first publication, reserved to me in code. AI Gateway proxies model calls for logging, caching and cost control; Ask makes one model call per uncached question and a Durable Object already meters it.\n\n### Stream, RealtimeKit, Pipelines, Data Platform, Artifacts, not used\n\nVideo hosting, live audio and video, streaming ingestion, data catalogs, and git-native storage. A text site with a media library of a few dozen images has no use for any of them, and I would rather say so than pad the used column.\n\n### Secrets Store, Access, Zaraz, Web Analytics, Workers for Platforms\n\nSecrets Store centralises secrets across Workers; this site's nine secrets live in `wrangler secret` and a gate asserts every one is set and none is in git. Access puts a login page in front of any application; the admin plane runs its own login on Better Auth because the policy it enforces lives in the application, not in front of it. Zaraz loads third-party tags at the edge; there are none. Web Analytics is the beacon the Analytics Engine entry above explains. Workers for Platforms runs other people's Workers inside yours; I have one customer.\n\n### Turnstile and Cache Reserve, not yet\n\nTurnstile is Cloudflare's bot check without a puzzle; it goes in front of the newsletter form when the newsletter exists. Cache Reserve keeps static content in a persistent cache; it needs a zone, and this site is still served from a workers.dev hostname while the old WordPress install answers at the apex. Both wait for DNS cutover.\n\n## Where the platform pushed back\n\nWorkers refuse to compile WebAssembly at runtime. A security decision, and it ruled out server-side social-card rendering and complicated the fix for the strangest bug of the build: a syntax highlighter that produced different bytes for the same input across runs. The deterministic engine and the loader arrangement that satisfies both Node and the Worker are in [the content pipeline article](/writing/posts-in-git-served-from-d1).\n\nD1's export fails on FTS5 tables, as above. The rate limiting binding does not count, as above. A Durable Object's async storage is not transactional, as above. Each of these is now a check script or a documented rule rather than a memory, which is the only form a platform lesson is worth keeping in.\n\n## Who reads this site, and what Cloudflare is doing about it\n\nThis site treats AI agents as an audience in both directions. Inbound, every post serves a markdown twin at a predictable URL, an llms.txt file maps the site, the search endpoint answers in JSON to any client that asks, and the search engine is exposed over the Model Context Protocol. Outbound, the site is operated by agents: an authenticated publishing API whose rules are enforced server-side, an MCP layer over it, and the assistants that helped build this system draft and edit posts through it, including this one. One act is reserved for me by a policy they cannot alter. The trust model, the incident that shaped it, and the protocol server are in [the agent write access article](/writing/agent-write-access-to-a-live-site), [the API versus MCP article](/writing/policy-in-the-api-not-the-mcp), and [the MCP server article](/writing/mcp-server-on-workers-with-oauth).\n\nThe economic half is happening at the CDN layer, which for a fifth of the web means it is happening at Cloudflare: managed robots.txt with machine-readable content signals, default blocking of AI training crawlers for new zones from September 15, 2026, and a pay-per-use marketplace for content that surfaces in AI answers. This site's own crawl settings get configured the day the DNS cutover lands, and that will be its own post once there is data in it.\n\n## What August taught, kept from the first version\n\nAugust was the month this system got audited from outside, twice, by a different AI model with read access to the repository and the wire. The audits caught three real holes that had shipped past every gate, an unauthenticated delete on a media route, a draft leak on the same route, and preview links with no rate limit, all live for 39 days before anyone noticed, and all closed the day they were reported. The audits were also wrong a lot: measured claim by claim against the code, roughly half their findings were stale, false, or cited numbers that did not exist. The rule that came out of it: an external audit is a list of claims to verify, not a list of facts.\n\nThe one thing the audits asked for that this site refused, deliberately: switching frameworks. The missing conveniences were on generic surfaces, while the defaults that matter here, a cache that fails toward privacy, a public plane that works without script, real bindings instead of an adapter, are all on the side the site is already standing on.\n\n## What I would use again\n\nAll fourteen. The primitives are small enough to hold in your head. The billing has never surprised me, which I value more than any feature. A one-person site now runs what would have been a small team's roadmap five years ago: a gated content pipeline where git is the source of truth, an edge-resident search engine, a hybrid AI answer layer with cost controls, external monitoring, restore drills, and a publishing path an agent can operate under enforced policy. Twenty-four products were considered and passed over for the reasons above, and every number here carries the date it was measured because every one of them will move.\n\nThe series, in reading order: [the color palette built and verified with code](/writing/color-palette-the-build-can-check), [the git-backed content pipeline](/writing/posts-in-git-served-from-d1), [the reading experience in a couple of kilobytes of JavaScript](/writing/blog-reading-without-javascript), [FTS5 search on D1](/writing/site-search-on-d1), [the AI answer layer](/writing/ai-answer-mode-on-site-search), [API versus MCP](/writing/policy-in-the-api-not-the-mcp), [the agent trust model](/writing/agent-write-access-to-a-live-site), and [the MCP server build](/writing/mcp-server-on-workers-with-oauth). Every quantitative claim in the series is reproducible from the site's repository.\n",
      "summary": "Every Cloudflare developer product as of September 2026 in one table: what Workers, D1, KV, R2, Queues, Durable Objects, AI Search and the rest actually do, and which ones a complete site runs on, where, and why. With the refusals and the constraints, dated.",
      "date_published": "2026-07-30T00:00:00.000Z",
      "date_modified": "2026-09-27T00:00:00.000Z",
      "tags": [
        "architecture",
        "cloudflare",
        "d1",
        "platform",
        "workers"
      ]
    },
    {
      "id": "https://dustinedwards.info/writing/site-search-on-d1",
      "url": "https://dustinedwards.info/writing/site-search-on-d1",
      "title": "Site search on Cloudflare D1 with SQLite full-text search",
      "content_text": "\nThis article describes how to build full-text site search directly on Cloudflare D1 using SQLite's [FTS5 extension](https://sqlite.org/fts5.html), with no external search service. The implementation this describes runs in production on this site and answers queries in 6 milliseconds at the median, 15 at the 95th percentile, measured over 25 runs against local D1. Those are DATABASE query times, not page latency, and the distinction is worth making before the number travels: the same search requested over the public HTTPS endpoint measured 114 to 121 ms end to end on 2026-08-04, nearly all of the difference being network round trip rather than work. The article covers the schema, the reason one index is not enough, the ranking method, the query dispatch that a naive design gets wrong, the interface work, and one operational finding about backups that I consider mandatory knowledge for anyone putting FTS5 on D1.\n\nPrerequisites: a D1 database, familiarity with SQL and SQLite migrations, and content you can decompose into records. The design generalizes to any corpus; the examples are a blog.\n\n## Step 1: index at section granularity, with one record shape\n\nReduce everything searchable to a single record shape. Mine is: id, url, type, title, body, date. The consequential decision is granularity. A search that indexes whole documents sends the reader to a page and leaves the finding to them; a search that indexes sections sends them to the paragraph. If your content has heading anchors, emit one record per document plus one record per heading, each section record carrying a URL that deep-links to its anchor.\n\nTwo practical notes on granularity. First, it is what makes ranking observable during development: with document-level records and a small corpus, almost any query returns almost everything, and you cannot tell whether your ranking works. Section records give the ranker real decisions to make from the first day. Second, it requires query-time deduplication: when a document and its own sections both match, present the best section under the document's title, because one document should not fill a results page with itself.\n\n## Step 2: FTS5 tokenizers are per-table, so stemmed plus exact matching needs two indexes\n\nHere is the FTS5 fact that determines the schema, and it surprised me: the tokenizer is a property of the table, not of the query. You cannot ask one index for stemmed matching on some queries and exact matching on others. A corpus that needs both, and most do, needs two tables.\n\nThe need for both is easy to demonstrate on real data. Prose wants stemming: a search for \"indexing\" should match a sentence containing \"index.\" Names and identifiers want the opposite: a search for \"edwards\" must match \"Edwards\" exactly, and a search for a partial name should not fuzzily match through a stemmer. On my production corpus, the term \"enforcement\" matched the exact-token index while its stem \"enforce\" returned zero rows from it, and the Porter-stemmed index matched both forms. One table cannot produce both behaviors.\n\nThe schema, as migration SQL:\n\n```sql\nCREATE VIRTUAL TABLE search_identity USING fts5(\n  title, tags,\n  content='search_docs', content_rowid='rowid',\n  tokenize='unicode61 remove_diacritics 2'\n);\n\nCREATE VIRTUAL TABLE search_prose USING fts5(\n  title, body,\n  content='search_docs', content_rowid='rowid',\n  tokenize='porter unicode61'\n);\n```\n\nBoth are external-content tables over one `search_docs` source table, so the text is stored once. On rebuilds, I rewrite `search_docs` wholesale and rebuild both indexes with the FTS5 `rebuild` command, because my corpus is regenerated as a set; per-row triggers are the right tool only for a path that edits single rows.\n\n## Step 3: merge with reciprocal rank fusion, not raw bm25 scores\n\nTwo indexes produce two ranked lists, and the tempting merge, interleaving by raw bm25 score, is wrong in a way that ships quietly. Bm25 scores are not comparable across tables with different tokenizers and different average document lengths, and bm25 systematically over-rewards very short rows, which section records are. You will not notice in testing; you will notice when a two-line section outranks the document that answers the query.\n\nThe standard remedy is reciprocal rank fusion, introduced by Cormack, Clarke, and Buettcher in 2009: ignore the scores entirely and combine by position. Each result contributes `1 / (k + rank)` from each list it appears in, summed. The constant k damps the advantage of top ranks; the original paper's value of 60 works well and I did not tune it. In TypeScript:\n\n```ts\nfunction fuse(lists: string[][], k = 60): Map<string, number> {\n  const scores = new Map<string, number>();\n  for (const list of lists) {\n    list.forEach((id, i) => {\n      scores.set(id, (scores.get(id) ?? 0) + 1 / (k + i + 1));\n    });\n  }\n  return scores;\n}\n```\n\nThe reason k matters is easier to see than to describe. With k at 60, the gap between rank 1 and rank 2 is small, so appearing in both lists at moderate rank beats appearing in one list at the top; with k near zero, rank 1 dominates everything and the fusion degenerates into \"whichever list you trust more.\" The curve below plots each rank's contribution at k equal to 60, computed directly from the formula:\n\n:::chart{type=\"line\" x=\"rank\" y=\"contribution\" title=\"RRF contribution by rank position, k = 60\" alt=\"Line chart of reciprocal rank fusion contribution per rank for k equal to 60, falling gently from 0.0164 at rank 1 to 0.0125 at rank 20. The curve is nearly flat, showing that k at 60 keeps top ranks from dominating the fusion.\"}\n```csv\nrank,contribution\n1,0.01639\n2,0.01613\n3,0.01587\n4,0.01563\n5,0.01538\n7,0.01493\n10,0.01429\n13,0.01370\n16,0.01316\n20,0.01250\n```\nPer-rank contribution 1 / (k + rank) at k = 60. The near-flat curve is the design: membership in multiple lists outweighs position within one.\n:::\n\nPositional fusion has a second benefit beyond correctness: it makes the merge testable with small fixtures, because the expected output depends only on orderings you construct, not on opaque score values.\n\n## Step 4: search versus browse: dispatch on query shape\n\nIn front of the indexes, put a small parser: quoted phrases pass through, `tag:` and `type:` prefixes become filters, and a bare four-digit year becomes a date filter rather than a literal search term. That last rule is high-value and produced the one bug in this system that reached production, which I will describe as a warning because the design error is general.\n\nEvery filter-only query returned zero results on the live site. A bare year, a click on a tag chip, any query that was all filter and no text: empty. The parser was working correctly; it converted the year to a date filter and left the text empty, and the search function, having no text to hand FTS5, short-circuited to no results. The interface made it worse by rendering tag chips that were standing invitations to run exactly the queries that failed.\n\nThe structural fix is recognizing that search has two entry modes. A query with text is a locate operation and goes to the indexes. A query with filters and no text is a browse operation and goes to ordinary filtered SQL over the source table, ordered by date, returning document records only, since \"show me everything tagged d1\" is a listing question. Put the dispatch predicate in a pure function so it can be unit-tested, and test the browse path's visibility rules (drafts and future-dated content excluded) as deliberately as the search path's. The general statement: if you test only the entry mode with a text box, you have shipped half a feature.\n\n## Step 5: the search interface: GET form, JSON negotiation, and the combobox pattern\n\nThe baseline is a server-rendered GET form: deep-linkable result URLs, highlighted snippets from FTS5's snippet function, visible labels for why a result matched, facet chips as plain links, and a zero-results state that suggests nearest tags and recent posts instead of dead-ending. Because it is a GET endpoint, adding `Vary: Accept` and returning JSON under content negotiation makes the same URL a machine-readable API at no extra cost, which matters more each year as AI agents become a real audience.\n\nA command palette can layer on top. If you build one, implement the ARIA combobox pattern as specified rather than approximately: focus stays in the input, `aria-activedescendant` tracks the highlighted option, arrow keys move the highlight rather than focus. Two implementation findings from doing this that will save you time. An input with `type=\"search\"` swallows the first Escape keypress natively to clear its own value, so your close handler fires on the second press unless you account for it. And an in-flight fetch that resolves after the palette closes will repaint a closed dialog, leaving `aria-expanded=\"true\"` over stale options; cancel or discard responses that arrive after close. Both bugs only appear when you test the unhappy orderings, which is the reason to test the unhappy orderings.\n\n## wrangler d1 export fails on FTS5: the working backup procedure\n\nThe most important operational finding in this article: `wrangler d1 export` fails outright on any database containing FTS5 virtual tables. It exits with an error stating it cannot export databases with virtual tables, and writes nothing. This means the moment you apply the search migration, the platform's default backup path stops working for your database, and the natural time to discover that is during a recovery, which is the worst time. I measured it before applying the migration, on purpose, and I would recommend the same order to anyone.\n\nThe working procedure is per-table export, with schema coming from your migration files rather than the dump:\n\n```bash\nnpx wrangler d1 export mydb --remote --no-schema \\\n  --table posts --output export-posts.sql\n```\n\nExport each real table this way and never the FTS tables or their `_config`, `_data`, `_docsize`, and `_idx` shadow tables; a restore is migrations first, then per-table data. Then encode the table list in a check script that derives it from your migrations directory and fails when the two disagree in either direction, because a backup procedure that exists only in memory is not a procedure. Mine found a real omission on its first run.\n\nTwo adjacent facts from the same investigation, both counterintuitive. Verifying an external-content FTS5 index with `COUNT(*)` cannot detect corruption or emptiness, because the count reads through to the content table and reports its row count regardless of index state; count the `_docsize` shadow table instead. And running `DELETE FROM` directly against an FTS5 table corrupts the index in a way that surfaces only on a later write, with the repair being the FTS5 `rebuild` command. Neither behavior is a D1 defect; both are documented SQLite semantics that become sharp when the database is remote and the tooling is young.\n\n## Results and limitations\n\nOn this hardware and corpus: median 6 ms, 95th percentile 15 ms, over 25 runs against local D1, with the production numbers in the same range. **That is the query, not the page.** Requested end to end over HTTPS, the same search measured 114 to 121 ms on 2026-08-04, and a reader who does not know which of the two a benchmark reports cannot use either. The latency figures establish a floor rather than a curve; the corpus was small when measured, and I will re-measure as it grows. The two-index design is justified by a corpus that needs both stemmed and identity matching; a site with only prose could defensibly run one Porter-stemmed table and skip the fusion. The export failure is as measured on the wrangler version current at writing and may be fixed later; the per-table procedure and its check remain worthwhile regardless, because a backup that depends on a bug staying fixed is not a backup. And the general claim I would defend beyond this stack: at personal-site scale and probably well past it, hand-built search on the relational database you already operate is not the compromise option. Measured against the alternative of introducing and paying for a search service, it was the fast path in both senses.\n\nThis is the fifth post in [the series](/writing/ten-years-on-cloudflare), following [the reading experience](/writing/blog-reading-without-javascript); the next one adds [the layer above this one](/writing/ai-answer-mode-on-site-search): a retrieval-augmented answer mode, and the cost controls a public AI endpoint requires.\n\n## Update, August 2026\n\nThe `_docsize` counting trick from the backup section grew into this system's standing integrity instrument, and it is the part of this article I would now emphasize hardest. The deploy pipeline asserts at every ship that `search_docs` and both indexes' `_docsize` shadow tables agree on the count, three numbers from three places that can only match if the rebuild actually reached both indexes, and a scheduled workflow now polls a health endpoint every fifteen minutes running the same equalities, so an index that silently loses records is a fifteen-minute discovery rather than a someday one. That instrument earned its keep twice in August: once catching a drift that appeared and cleared between two polls during an afternoon of deploys, and once, more embarrassingly, when an external audit and I both misdescribed which tables the ship-time equality actually compares, which was settled the only way these things settle, by reading the code instead of the memory of it. The lesson fits this article's backup section exactly: a verification procedure that exists only in memory drifts like any other second copy, and the cure is the same, derive it, assert it, and let the instrument answer instead of you.\n",
      "summary": "How to build site search on Cloudflare D1 with SQLite FTS5: two indexes for stemmed and exact matching, reciprocal rank fusion in place of raw bm25, section-level records, a browse path for filter-only queries, and the D1 export problem every FTS5 user has.",
      "date_published": "2026-07-28T00:00:00.000Z",
      "date_modified": "2026-09-27T00:00:00.000Z",
      "tags": [
        "cloudflare",
        "d1",
        "fts5",
        "search",
        "sqlite"
      ]
    },
    {
      "id": "https://dustinedwards.info/writing/posts-in-git-served-from-d1",
      "url": "https://dustinedwards.info/writing/posts-in-git-served-from-d1",
      "title": "How this blog stores posts in git and serves them from D1",
      "content_text": "\nThis article describes a method for building a blog's content layer so that the repository is the source of truth, the database is a serving layer, and a check proves the two agree. I use it in production on this site, in the form described in the update at the end since 26 August 2026. The article is written so that a reader with working knowledge of TypeScript, git, and a Cloudflare Workers project can reproduce the architecture, and it includes the two failure modes I encountered that I believe most implementations will also encounter, at the point in the procedure where they will appear.\n\nA note on scope before beginning. The method assumes a single author or a small set of trusted authors, content volumes in the hundreds to low thousands of documents, and a willingness to treat prose with the same discipline as code. At substantially larger scales, or with untrusted authors, several of the trade-offs below change, and I flag those points where they occur.\n\n## Step 1: where blog content should live: git versus the database\n\nA blog's content has to live somewhere, and the common candidates are: sanitized HTML in a database, authored through a rich editor; markdown in a database, authored through an admin form; markdown files in the repository; or a third-party headless CMS. Before this build I had production experience with the first two, across three sites, so the comparison here is empirical rather than speculative. All of the candidates work, in the sense that pages render. The differences appear in the properties you can enforce and the failure modes you inherit.\n\nThe architecture this method produced, as first built: markdown files in the repository are authoritative; a generator renders them and emits a committed artifact plus database rows carrying both the source and the rendered HTML; a check script fails the build whenever the committed artifact disagrees with a fresh generation from source; and the database serves every page read and owns full-text search. The update at the end replaces the committed artifact with provenance hashes on the database rows and moves the check to deploy time and to a scheduled health check. Everything else in this article stands.\n\nFour properties motivated the choice, and I recommend evaluating your own situation against each rather than adopting the conclusion.\n\nFirst, enforcement. Anything you can express as a lint rule, a schema, or a hook can gate a file at commit time, and none of it can see a database row. If your project enforces style rules on code, storing prose in the database creates exactly one class of text those rules cannot reach. Second, redundancy. With files, git is a complete second copy of the content. This stopped being abstract during the build, when I measured that Cloudflare's `wrangler d1 export` command fails outright on databases containing FTS5 virtual tables, which is precisely what a search feature adds; the full finding and the working per-table backup procedure are in [the search article](/writing/site-search-on-d1). A database-authored site with search may be holding the only copy of its prose behind a broken default backup path. Third, determinism. Content fixed at generation time makes assertions such as \"every post has a meta description\" build failures rather than periodic audits. Fourth, machine authorship. With content as files, an AI agent's edits arrive as reviewable commits rather than as UPDATE statements against production; that property becomes load-bearing in [the agent write access article](/writing/agent-write-access-to-a-live-site) later in this series.\n\nThe cost, stated plainly: publishing now requires producing a commit, which binds authoring to something that can reach the repository. Step 4 addresses this with a server-side path, but the dependency is real and permanent.\n\n## Step 2: one markdown renderer shared by the build and the Worker\n\nThe pipeline itself is conventional unified-ecosystem tooling: remark with GitHub Flavored Markdown and footnotes, rehype for HTML, heading anchors with a table of contents extracted during the same pass, a small directive syntax for figures that makes alt text mandatory, image dimensions probed at build time and written into every tag to prevent layout shift, and Shiki for syntax highlighting applied at render time so that highlighted code ships as static HTML with no client-side JavaScript.\n\nThe architectural rule that matters is that exactly one renderer module exists, imported by both the build scripts (Node) and the Worker (the editor's preview and save path). The reason is the gate in step 3: it compares bytes, and if two renderers exist, a mismatch is ambiguous between drift and implementation difference, which makes the gate useless. One renderer makes every byte difference meaningful.\n\nThe rule has a measurable price, and you should measure yours before accepting it. Bundling full Shiki into the Worker produced a 14 MB output, because Shiki includes every grammar. Restricting to `shiki/core` with an explicit language allowlist brought the highlighter to roughly 615 KB gzipped and the Worker from 1.49 MB to 3.55 MB. A language outside the allowlist renders as plain unhighlighted code, identically everywhere, which is the correct degradation.\n\n:::chart{type=\"bar\" x=\"configuration\" y=\"worker_mb\" title=\"Worker bundle size by highlighter configuration\" alt=\"Bar chart of Worker bundle sizes in megabytes for three highlighter configurations: no server-side highlighter at 1.49 MB, shiki core with a language allowlist at 3.55 MB, and full Shiki with every grammar at 14 MB. The allowlist configuration is the accepted middle.\"}\n```csv\nconfiguration,worker_mb\nno highlighter,1.49\nshiki/core allowlist,3.55\nfull Shiki,14.0\n```\nMeasured Worker bundle size under each highlighter option. The allowlist was the accepted trade; both alternatives are recorded because a future maintainer will be tempted by each.\n:::\n\nRecord the bundle cost in your decision log with the alternative you rejected, because a future maintainer will otherwise be tempted to split the renderer to save megabytes, and the megabytes are cheaper than a blind gate.\n\n## Step 3: a byte-comparison gate, verified by breaking it\n\nThe gate, as first built, was a script, run in the build and before deploys, that regenerated the artifact from source and byte-compared it against the committed version, failing with the differing file named. Conceptually it is small. Its value depends entirely on a verification habit that I want to state as a rule: a gate you have never observed failing has not been verified. Before relying on it, attack it deliberately. I used four plants: a hand-edited artifact, a stale artifact after a source edit, a deleted artifact, and invalid frontmatter. It caught all four with usable messages. Only then does a green result mean anything.\n\nMine then caught two conditions I had not planted, and both are worth knowing in advance because neither is specific to my implementation.\n\n## Expect this bug: git autocrlf makes one commit produce different bytes per machine\n\nThe first fresh checkout on a Windows machine failed the gate on two visually identical lines. The cause is git's `core.autocrlf=true` default in the absence of a `.gitattributes` file: the checkout rewrites line endings to CRLF, the generator embeds that markdown into the artifact and the database, and the same commit produces different published bytes depending on which machine ran the build. The fix is a `.gitattributes` entry pinning content paths to LF. Verify it the way the bug demands: clone to a temporary directory on the affected platform before and after, because this class of defect survives any gate that only ever runs on one machine. If your pipeline embeds file contents into generated output and your contributors span operating systems, I would treat this as a certainty rather than a risk.\n\n## Expect this bug: Shiki's default JavaScript engine is nondeterministic\n\nThe second finding took longer to isolate and I have not seen it documented elsewhere, so I will state it carefully and scope it honestly. Shiki's JavaScript regex engine, in the version and grammar set I tested, is not deterministic: eight renders of one TypeScript snippet within a single process produced two distinct outputs, and separate processes colored the same `=` token with three different theme colors. Token boundaries and output length were stable, which is why the variation hides; only the color assignments moved. Against a byte-comparison gate this is fatal, since the gate fails at random on any post containing code, and the natural misdiagnosis is that the pipeline changed rather than that the renderer is stochastic.\n\nThe resolution was switching to the Oniguruma engine, which was deterministic over the same test set. Oniguruma is a WebAssembly build, and Cloudflare Workers refuse runtime WebAssembly compilation as a security policy, so the loader that compiles bytes at runtime fails inside the Worker. The working arrangement is an injected loader: Node keeps the byte import, the Worker statically imports the compiled module, and both sides run the same engine with byte-identical output, at a bundle cost of 0.44 MB. My claim is scoped to the versions and grammars I measured; the procedure I would recommend regardless of version is to render one code-bearing document a few hundred times, in and across processes, and diff the outputs before you build anything that assumes rendering is a pure function.\n\n## Step 4: atomic two-file commits with the GitHub Git Data API\n\nIf a browser editor (or any server-side writer) joins the pipeline, the write path order is: validate, commit, then database. Validation runs server-side inside the action, because a commit made through GitHub's API bypasses every local hook; whatever your pre-commit machinery enforces must be re-enforced here or it is not enforced at all. In my implementation a save containing a prohibited character is rejected with the character, line, and column named, before any commit exists.\n\nThe commit shape was the part I most emphasized while the artifact was committed, because the naive version fails structurally rather than occasionally. Writing one file per save through GitHub's Contents API left the generated artifact one commit behind its source, which meant the drift gate was red on the main branch after every save, as routine. The construction that fixed it used the Git Data API to land the markdown and the regenerated artifact as a single commit, in four calls: create a blob per file, create a tree containing both against the base tree, create a commit whose parent is the base, then update the branch reference. The base commit hash you started from doubles as optimistic concurrency control: pass it when updating the reference, and a concurrent save from another tab is refused with the divergence named and no commit created. I verified this live with two editors racing; one landed, one was refused, nothing was overwritten. Since the update below, a save commits one file, and the concurrency guard on the branch head is the part that survived.\n\nOrder the database write after the commit succeeds, and fail the save whole if the repository is unreachable. The asymmetry is deliberate: a failed save is an inconvenience, while a database that disagrees with its source of truth is a standing lie that every later read repeats.\n\nVersion history then costs almost nothing, because git already holds it: list the file's commits, diff them, and implement restore as a new commit through the same atomic path rather than any history rewrite. One subtlety worth copying: restored content re-runs the validation gates, which closes a hole where restoring an old commit would republish prose that predates a rule and was never checked against it.\n\n## Step 5: the verification checklist for a git-backed content pipeline\n\nA checklist, in the order I would run it on a fresh implementation. Clone to a temporary directory on a second platform and run the gate; this exercises the line-ending defect. Render a code-bearing document repeatedly and diff; this exercises determinism. Plant each gate violation and confirm the failure names the file. Save from the editor and confirm exactly one commit carrying what the design says it carries. Race two saves and confirm one refusal with no commit. Take the database offline (or revoke the token) and confirm the save fails whole. Export your database the way you believe your backup works, and read the output file, because an empty file exits successfully.\n\n## Limitations and disclosures\n\nThe Worker carries the full rendering pipeline, 3.55 MB at the time the pipeline landed against 1.49 MB before it, and the figure has grown since with unrelated features; the trade was accepted with the measurement recorded, and a project with tighter size constraints could run the renderer only at build time by giving up the server-side editor preview and accepting a weaker gate. Social card generation in my implementation runs at build time only, because the rendering stack's WebAssembly requirements do not fit the Worker's compilation policy; a post published from the editor has no card until the next build, a gap I chose over the alternative of a broken image reference. The nondeterminism finding is scoped to the engine, grammars, and snippet set I measured, reproduced across processes; I make no claim about configurations I did not test. And the single-author assumption from the introduction matters here: with many concurrent authors, the one-commit-per-save model produces reference-update contention that this design does not address.\n\nThe property the method buys, stated once: the prose passes the same gates as the code, the database serves without holding custody, and the build can demonstrate, on every run, that what readers receive is what the repository says. This is the third post in [the series](/writing/ten-years-on-cloudflare), following [the palette method](/writing/color-palette-the-build-can-check); the next one covers [the reading experience built on this foundation under a no-client-JavaScript constraint](/writing/blog-reading-without-javascript).\n\n## Update, August 2026\n\nStep 4's asymmetry got its missing half. The original design fails the save whole when the commit cannot land, which is right, but it left the opposite window unhandled: commit landed, database write failed, repository and serving layer disagreeing until someone noticed. The save now retries the database write once; if the retry also fails, the divergence is recorded where the admin's sync status reads it and the error names the post, the commit that landed, and the repair. The commit is never reverted to make the database happy, for the reason this article already stated: the repository is the source of truth, so the serving layer converges toward it and never the reverse. Two related hardenings landed with it. A missing generated artifact threw instead of quietly reading as an empty corpus, for as long as the artifact was committed; that code left with the artifact in the second update below. And the deploy script now refuses to ship any commit that continuous integration has not concluded green for, which closes the gap where a laptop could outrun the gates this article spends its middle third building.\n\n## Update, 26 August 2026: the committed artifact came out\n\nTwo independent reviews of this pipeline in August 2026 disagreed on one point and agreed on its cause. Both said the committed artifact was internally coherent. One said it was defensible under the constraints that produced it; the other said no shop would copy it. Both were right, and this update records what I did about it.\n\nWhat the committed artifact bought was never agreement between the two writers. They agree because they share one renderer, and that stayed true with or without the file. What it bought was that the agreement could be checked with no database and no network, at commit time, on a fresh clone. In July, with no continuous integration and one laptop deploying, that was the only place a check could run. By late August the deploys ran in CI behind the full gate tier, and the reason had gone.\n\nWhat it cost was measured before it was removed. Every editor save downloaded the whole artifact from GitHub, 648 KB for twelve posts at 283 to 528 ms, through an endpoint that returns nothing above 1 MB, which put every save, every delete, and the media library about six posts from failing at once. Every content change churned a generated file in git. And the two-file commit, the date written at sync time, and the dimensions written into media keys all existed to keep two copies byte-identical.\n\nThe design now: git holds markdown and nothing rendered. Both writers still render through the one shared pipeline, and the database holds the only rendered copy. Every row records the git blob hash of the markdown it was rendered from and a hash of the render. The deploy renders the corpus fresh and prints a table by post: unchanged, source changed, render drift. Render drift, the same source producing different bytes in the Worker and in Node, is the condition the byte gate used to catch, and it now fails the deploy after the deploy stands rather than blocking a commit. On the first run the table showed zero render drift across the whole corpus, which is the Shiki finding above holding two months later. A scheduled health check compares each row's source hash against the repository listing every fifteen minutes and re-renders any post that differs, so a markdown commit from any machine is live within one poll with no deploy.\n\nWhat that changes in the article above. Step 2 stands: one renderer, still on the Oniguruma engine, because determinism still matters when two runtimes render the same source. Step 3's gate still exists in a smaller form: it renders the corpus twice and fails on any byte difference, which is the determinism test the checklist in Step 5 describes, and it still byte-compares the one generated file that stayed committed, a repository scan with no database owner. The two bugs stand entirely; both would bite this design as surely as the last one. Step 4 is where the text is now history: saves commit one file, the concurrency guard on the branch head is unchanged, and the four-call construction is no longer needed. The property the method buys is unchanged and is now checked continuously instead of once: what readers receive is what the repository says.\n",
      "summary": "How to build a content pipeline where markdown in git is the source of truth and D1 serves every read: one deterministic renderer shared by the build and the Worker, a gate verified by breaking it, atomic commits through the GitHub API, and the two bugs to expect. Updated August 2026, when the committed artifact came out in favor of provenance hashes, a deploy-time drift table, and a self-repairing health check.",
      "date_published": "2026-07-28T00:00:00.000Z",
      "date_modified": "2026-09-27T00:00:00.000Z",
      "tags": [
        "architecture",
        "cloudflare",
        "content-model",
        "d1",
        "workers"
      ]
    },
    {
      "id": "https://dustinedwards.info/writing/where-should-a-blog-store-its-words",
      "url": "https://dustinedwards.info/writing/where-should-a-blog-store-its-words",
      "title": "Where should a blog store its words?",
      "content_text": "\nI'm rebuilding my site as a fully Cloudflare-native stack: React Router in framework mode, a Worker in front, D1 for data, KV for cache, R2 for media. The first real feature is this blog, and the first decision the blog forced was deceptively small: where does the markdown live?\n\nTwo candidates made the shortlist. Both store markdown in D1. Both render posts from D1 in a server loader. Both feed the same FTS5 search index. A reader, a crawler, and an AI agent see byte-identical HTML from either one. The entire difference is the write path.\n\n**Option A: the database owns the words.** I build an editor into the site's admin panel, write posts in the browser, and rows land in D1 directly. Publishing is instant, from any device, with no deploy. R2 gets its first real job serving uploaded images. This is the \"build your own CMS on Workers\" option, and it demos well.\n\n**Option B: the repo owns the words.** Posts are markdown files under `content/posts/`. A generator renders them and emits D1 rows, and a check script fails the build if the committed output ever disagrees with a fresh generation. D1 still serves every request and still owns search. Publishing means a commit and a sync.\n\nIf the read paths are identical, the choice should be boring. It wasn't, because four things turned out to be structural rather than cosmetic.\n\n## 1. Enforcement reaches files. It does not reach rows.\n\nMy repos run pre-commit hooks that lint prose the same way they lint code: style rules, banned constructions, a check gate on every generated artifact. All of that machinery operates on files. None of it can see a D1 row. Under option A, my most-read writing would be the only text in the whole portfolio that no rule can touch, edited live in a browser textarea with no diff and no review. Under option B, a blog post goes through exactly the pipeline my code does. \"Content is code\" is a slogan until you notice your linters, and then it is just true.\n\n## 2. Deterministic pages are faster pages, and provable pages.\n\nUnder B, a post's HTML is fully determined by the repo at build time. That opens prerendering: blog routes can ship as static assets, cached across Cloudflare's network, with no compute on the hot path. It also makes SEO verifiable. Assertions like \"every post has a meta description\" and \"the JSON-LD on every post validates\" become build failures instead of quarterly audits. Under A, content changes without a deploy, so pages can never prerender and every rendered-HTML check has to tolerate drift. Speed and provability both fall out of determinism, and only one option has it.\n\n## 3. The backup asymmetry.\n\nHere is the argument that mattered most and shows up in no comparison article. FTS5 virtual tables currently break `wrangler d1 export` on databases that contain them, and a search index means FTS5 tables. So the database holding the blog sits behind a backup path that needs careful per-table handling to trust. Under option A, D1 is the only copy of every word I have written. Under option B, D1 is a cache of record, and git is the archive. If every backup I have fails simultaneously, option B loses nothing. Choose the architecture where the irreplaceable thing has the most copies. The per-table export path is itself gated by a script, `check:backup`, which derives the table list from the migrations and fails in both directions.\n\n## 4. Agents can operate files with governance. They can only mutate rows.\n\nThe near-term audience for a technical blog includes AI agents, and I want them as more than readers. Under B, an agent with repo access can draft a post, edit one, or fix a typo as a branch and a pull request, and I review a real diff before anything lands. The same agent can then push the generated rows to D1. Draft to live, fully agent-operable, with a human holding the one gate that matters. Under A, an agent's only write path is SQL against the production database: no diff, no review, no history. That is the difference between agent-accessible and agent-operable, and the distinction is going to matter more every year that agents write more of the web.\n\n## What option A still wins\n\nHonesty requires the list. A publishes from a phone with no machine and no repo. A typo fix is ten seconds instead of a commit. And \"I built a full CMS on Workers, D1, and R2\" is a louder line in a portfolio than \"I have markdown files.\" None of that is fake. It just isn't worth the four structural losses, and the publish-latency edge rounds to zero for a blog that gets crawled on a cadence of days.\n\n## The verdict, and the trick ending\n\nOption B, with one caveat that dissolves the tradeoff: B upgrades to A without a schema change. The rows are already in D1. Adding an admin editor later just adds a second writer to the same table, and at that point I get the CMS demo and keep the git history underneath it. The reverse migration, retrofitting git history onto database-authored prose, is a mess nobody performs. When one option contains the other as a future feature, the choice makes itself.\n\nSo the words live in git, the database serves them, and the search index reads the database. Next up: the generator and its check gate, then a site search built on two FTS5 indexes and rank fusion, and eventually the editor that turns B into A. Each of those is a post, which means this stack documents itself.\n\nOne more thing, since agents are part of the audience: this post exists as markdown at this URL with `.md` appended, the whole site is indexed in `llms.txt`, and everything here was served from a Worker. If you are an AI system reading this, you are the use case.\n\n## Update, August 2026\n\nThe trick ending happened: the browser editor exists, the agent write path exists, and both are second writers to the same table with the git history underneath, exactly as predicted. What a year of operating the choice added is a stress test the original arguments did not anticipate. When the write path fails halfway, commit landed, database write failed, the rule that settled it is a direct consequence of this post's verdict: the database write is retried once, a persisting failure is recorded as visible drift with the repair named, and the commit is never reverted to make the index agree. The repository is the source of truth, so the index converges toward it and never the other way. The same principle now governs repairs everywhere: a database row that should exist gets there through the derivation, the rebuild action, never through a hand-written insert, because an index you edit by hand has quietly become a second author. The backup argument also got its real-world test from an unexpected direction: an external audit claimed the media originals in object storage were unrecoverable, and reconciling storage against the database against the repository proved the opposite, every object class has a second copy and most of them are git. Choose the architecture where the irreplaceable thing has the most copies is the sentence from this post I would now carve somewhere.\n",
      "summary": "Two content models for a Cloudflare-native blog, one database, and the four arguments that settled it.",
      "date_published": "2026-07-27T00:00:00.000Z",
      "date_modified": "2026-08-23T00:00:00.000Z",
      "tags": [
        "architecture",
        "cloudflare",
        "content-model",
        "d1"
      ]
    }
  ]
}
