{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Dustin Edwards blog: workers",
  "home_page_url": "https://dustinedwards.info/writing/tags/workers",
  "feed_url": "https://dustinedwards.info/writing/tags/workers/feed.json",
  "description": "Posts tagged workers.",
  "language": "en-US",
  "authors": [
    {
      "name": "Dustin Edwards",
      "url": "https://dustinedwards.info"
    }
  ],
  "items": [
    {
      "id": "https://dustinedwards.info/writing/observable-plot-inside-a-worker",
      "url": "https://dustinedwards.info/writing/observable-plot-inside-a-worker",
      "title": "Rendering Observable Plot charts inside a Cloudflare Worker",
      "content_text": "\nObservable Plot can render charts inside a Cloudflare Worker. As far as I can determine, this has not been documented before: Plot's own documentation covers server-side rendering in Node.js via JSDOM, the Plot team has demonstrated rendering in a browser Web Worker via a DOM substitute, and searching for the Cloudflare case returns nothing. This article reports the working configuration, measured on 2026-07-31: Plot 0.6.17 with linkedom 0.18.13 as the DOM implementation, bundling at 930 KiB raw and 226 KiB gzipped as measured in this site's production Worker, rendering deterministically (200 in-process renders and multiple separate processes producing one distinct output), and, the property that mattered most here, producing byte-identical SVG in Node and in workerd, verified by SHA-256 comparison. It also reports the three candidate configurations that failed, with their exact errors, because the failures define the boundary of what works and two of them generalize well beyond charts.\n\nEverything below is reproducible; the chart in this article is rendered by the pipeline it describes.\n\n## Why render charts in the Worker at all\n\nThe obvious architecture for a static blog is to render charts at build time only, and if that fits your system, you should do it and skip the hard parts of this article. The requirement here was stricter because of two standing rules on this site. First, charts are content: the data lives in the post's markdown as a fenced block inside a chart directive, and the rendered SVG is part of the stored HTML rather than an asset beside it, the same regime [all content here lives under](/writing/posts-in-git-served-from-d1). Second, that render has two writers, the Node build and the Worker's save path (the browser editor, and the [agent-operated publishing API](/writing/agent-write-access-to-a-live-site) this post arrived through), and the whole design depends on both writers producing identical bytes. Together these rules mean the chart renderer must run in the Worker, produce output byte-identical to the Node build's, and do so deterministically forever. That combination, not any single requirement, is what eliminated most of the field.\n\n:::sidenote{kind=\"Constraint\"}\nByte-identical is the word that decides this. Two writers producing *equivalent* SVG would still fail the build's comparison, because the gate compares bytes rather than rendered pictures.\n:::\n\n## The configuration that works: Plot plus linkedom\n\nPlot requires a DOM: it builds its output through d3-selection against a document. Cloudflare Workers have no DOM. The bridge is [linkedom](https://github.com/WebReflection/linkedom), a lightweight pure-JavaScript DOM implementation, passed to Plot through its documented `document` option:\n\n```ts\nimport * as Plot from \"@observablehq/plot\";\nimport { parseHTML } from \"linkedom\";\n\nconst { document } = parseHTML(\"<html><body></body></html>\");\nconst chart = Plot.plot({\n  document,\n  width: 640,\n  height: 400,\n  marks: [\n    Plot.barY(data, { x: \"library\", y: \"kib\", fill: \"var(--chart-1)\" }),\n    Plot.ruleY([0]),\n  ],\n});\nconst svg = chart.outerHTML;\n```\n\nThat string is the chart: static SVG, no client JavaScript, servable directly or embedded in generated HTML. Three properties of the output are worth stating precisely because each was a requirement I verified rather than assumed.\n\nDeterminism. Two hundred renders in one process produced exactly one distinct output by SHA-256. Separate processes agreed. The generated-ID hazard I expected (Plot creates clip-path IDs in some configurations, the same class of nondeterminism that [broke this site's syntax highlighter](/writing/posts-in-git-served-from-d1)) did not appear for standard marks: the output contained no generated IDs at all. I would still treat this as a property to verify per Plot version rather than a permanent fact, and this site's build now does, on every run.\n\nCross-environment byte parity. The same chart rendered in Node and in workerd (via a bundled invocation under miniflare) hashed identically. This is the property the two-writer gate requires, and it held both in the initial probe and when reproduced against the production module. One subtlety from the earlier probe is worth recording: parity holds when both environments use the same DOM implementation. My first Node measurements used a different shim than the Worker and the hashes diverged, which was the shims' serialization differing, not Plot misbehaving. Standardize on one DOM implementation everywhere and the question disappears.\n\nCSS custom properties pass through. Plot accepts `var(--chart-1)` anywhere it accepts a color, and the string lands in the SVG untouched. Because the SVG is embedded inline in the page, those variables resolve against the site's stylesheet at display time, which means one stored render serves light and dark themes and the palette's [contrast gate](/writing/color-palette-the-build-can-check) governs chart colors with no additional machinery.\n\n## The configuration the Plot team demonstrated, which does not transfer\n\nThe natural starting point was domino, the DOM substitute the Plot team used to demonstrate Plot inside a browser Web Worker. It works in Node. It cannot ship to a Cloudflare Worker, and the failure is categorical rather than a bug: domino's `lib/sloppy.js` uses `with` statements, which are illegal in strict mode, and Workers bundle as strict-mode ES modules. The build fails with five errors of the form:\n\n> With statements cannot be used with the \"esm\" output format due to strict mode\n\nNo configuration fixes this; it is a property of the package's source. The general lesson: a dependency that works in Node and in browser Workers can still be structurally excluded from strict-ESM targets, and you find out from the bundler, not the documentation.\n\n## The candidate that fails at runtime: Vega-Lite\n\nVega-Lite was the strongest alternative on paper: a scholarly grammar, headless rendering in Node with no DOM needed, deterministic output, CSS variables passing through. It renders fine in Node. In workerd it returns a 500:\n\n> EvalError: Code generation from strings disallowed for this context\n\nVega's runtime compiles its expression language via dynamic code generation, and Workers forbid runtime code generation as a security policy, the same policy that [forbids runtime WebAssembly compilation](/writing/posts-in-git-served-from-d1) and shaped this site's highlighter and social-card architecture. Vega ships an alternative AST interpreter for CSP-restricted environments, which I did not pursue because Plot had already passed every test. The finding stands on its own for anyone evaluating Vega on Workers: the default path cannot run there, and the error will not appear until runtime.\n\n## The candidate that lost on the axis it was supposed to win: Apache ECharts\n\nECharts has the best server-side story in the charting field on paper: a zero-dependency SSR mode, introduced in 5.3, that renders to an SVG string with no DOM at all. I expected it to win on bundle size for exactly that reason, since Plot drags a DOM implementation along. Measured, the intuition inverted: ECharts bundled for Workers at 2958 KiB raw and 634 KiB gzipped against Plot-plus-linkedom's 1070 and 254 in the same probe harness. The shim is small; the engine is not. I record the caveat that aggressive tree-shaking of ECharts' modular imports could narrow this, so treat the ECharts figure as worst-case, but the headline holds: needing no shim does not make a library light, and the only way to know a bundle size is to measure it.\n\n:::chart{type=\"bar\" x=\"configuration\" y=\"gzipped_kib\" title=\"Worker bundle size by charting configuration\" alt=\"Bar chart of gzipped Worker bundle sizes measured in the probe: Observable Plot with linkedom 254 KiB, Vega-Lite 412 KiB, Apache ECharts 634 KiB. Plot with linkedom is the smallest at less than half of ECharts.\"}\n```csv\nconfiguration,gzipped_kib\nPlot + linkedom,254\nVega-Lite,412\nECharts SSR,634\n```\nGzipped Worker bundle size per configuration, measured 2026-07-31 with wrangler 4.116.0 in one probe harness. Vega-Lite's figure is included for comparison although it cannot run in workerd at all.\n:::\n\n## Accessible charts, enforced by the build rather than promised\n\nA chart library's output is not accessible; an implementation is. Plot's raw SVG carries no role and no accessible name, so this site's chart directive wraps every chart in a structure the build enforces: the SVG itself carries `role=\"img\"` and an `aria-label` from a mandatory alt attribute (a chart without one fails the build, the same rule images here have always had), a `figcaption` carries the caption, and an equivalent HTML data table, generated from the same fenced data, sits beneath the figure so a screen reader user gets the numbers rather than a summary. Because the data lives in the markdown, the table is automatic and cannot drift from the chart.\n\nOne correction from building this is worth passing along, because I wrote the wrong version into the specification first and the implementing session caught it against the WAI-ARIA spec before shipping. The natural-looking structure, `role=\"img\"` on the `<figure>` wrapping everything, is a compliance bug: `role=\"img\"` makes all descendants presentational, so it would have generated the caption and the data table and then hidden both from assistive technology, the accessibility features defeating themselves. The role and the accessible name belong on the SVG element, the thing that is actually an image, leaving the figure, caption, and table as ordinary reachable semantics. The corrected structure is now a planted violation in the build gate: a chart whose SVG lacks a role or accessible name fails the build.\n\n## Where the boundary is: charts yes, diagrams no\n\nThe same probe method was applied to text-to-diagram tools (Mermaid, and the smaller Pintora), and every candidate failed under linkedom, each with a different first error and the same root cause. Mermaid fails immediately at `ReferenceError: CSSStyleSheet is not defined`, and behind that lies its documented dependence on `SVGTextElement.getBBox()`; Pintora fails at `TypeError: Cannot set properties of null (setting 'font')`, reaching for a canvas text-measurement context that does not exist. The distinction generalizes: Plot computes its layout from data, scales mapping numbers to coordinates, while diagram layout requires measuring rendered text, and text measurement requires real font metrics that no lightweight DOM shim carries. That is why charts can satisfy a two-writer byte-parity requirement on Workers today and diagrams cannot; diagrams on this site render at build time as external assets instead, under [the rule that a reproducibility gate should compare only what its inputs fully determine](/writing/blog-reading-without-javascript).\n\n## Limitations\n\nThe measurements are dated 2026-07-31 and version-pinned (Plot 0.6.17, linkedom 0.18.13, wrangler 4.116.0); the determinism and parity properties are re-verified by this site's build on every run precisely because they are properties of versions, not laws. The zero-generated-IDs result is scoped to the marks tested; configurations using explicit clipping may behave differently. The ECharts bundle figure is a worst-case without tree-shaking effort. The parity test bundles the chart module rather than the whole content pipeline, a scope choice stated in the gate itself. And the no-prior-art claim is a search result, not a proof; if someone has done this before and written it down where I could not find it, I would genuinely like to read it.\n\n## Update, 28 August 2026: the byte-parity requirement outlived the gate that checked it\n\nThis post describes the requirement as serving a committed artifact that a build\ngate byte-compared against a fresh generation. That artifact left git on\n2026-08-26, for reasons set out in [the content pipeline\narticle](/writing/posts-in-git-served-from-d1): it made every editor save\ndownload the whole thing from GitHub, and the check it enabled had moved into\ncontinuous integration anyway.\n\n**Nothing in the engineering above changes, and that is the interesting part.**\nThe requirement was never really about the file. It is about two independent\nwriters rendering the same source, and the site still has exactly two: the Node\nbuild and the Worker. What moved is where the disagreement is caught. It used to\nbe a byte comparison at commit time. It is now a drift table printed at deploy:\nevery row in the database records the git blob hash of the markdown it came from\nand a hash of its own render, the deploy renders the corpus fresh, and a row\nwhose source is unchanged while its render differs is named by slug and fails\nthe run after the deploy stands.\n\nSo the chart renderer must still run in the Worker, still produce output\nbyte-identical to the Node build's, and still be deterministic forever. Both\nproperties are still checked on every build, in-process and across processes and\nacross the Node/workerd split, by the same gate this post describes. A\nnon-deterministic renderer used to fail a byte comparison at random; it would\nnow report render drift at random, which is the same finding wearing a different\nname and reaching a reader no later.\n\nThis post extends [the series on rebuilding this site on Cloudflare's developer platform](/writing/ten-years-on-cloudflare). It was drafted by the site's operator agent, staged through [the MCP tools the series describes](/writing/mcp-server-on-workers-with-oauth), and its figure was rendered by the pipeline it documents. Publication, as always here, required the human.\n",
      "summary": "How to render Observable Plot charts inside a Cloudflare Worker: the linkedom shim that works, the domino and Vega failures that don't, byte-identical output across Node and workerd, and accessible charts enforced by the build.",
      "date_published": "2026-07-31T00:00:00.000Z",
      "date_modified": "2026-09-27T00:00:00.000Z",
      "tags": [
        "accessibility",
        "cloudflare",
        "data-visualization",
        "observable-plot",
        "workers"
      ]
    },
    {
      "id": "https://dustinedwards.info/writing/mcp-server-on-workers-with-oauth",
      "url": "https://dustinedwards.info/writing/mcp-server-on-workers-with-oauth",
      "title": "Building an MCP server on Cloudflare Workers with OAuth",
      "content_text": "\nThis article describes how to build a production [Model Context Protocol](https://modelcontextprotocol.io) server on Cloudflare Workers, under the specification's 2026-07-28 revision, which shipped as final two days before this server deployed. The server in question wraps this site's publishing API so that an AI assistant can operate the site as tools in a conversation; [the previous article](/writing/agent-write-access-to-a-live-site) covered the trust model underneath. This one is the build method, and it is organized around four decisions I would recommend to anyone building an MCP server this year: measure your actual client before choosing an authorization design, select your protocol implementation by conformance score rather than preference, isolate any legacy-protocol support in one deletable module with an empirical retirement condition, and keep all business policy out of the server entirely. [The repository is public](https://github.com/DrDustinEdwards/dustinedwards-mcp), so every claim below is checkable.\n\nA dated context note, since MCP is moving. The 2026-07-28 revision is the largest since the protocol launched: it removes the session-based initialization handshake in favor of a stateless model where protocol version and client identity travel in per-request metadata, formalizes servers as OAuth 2.1 resource servers with discovery via Protected Resource Metadata (RFC 9728) and audience binding via Resource Indicators (RFC 8707), replaces dynamic client registration with client identity documents, and introduces a formal deprecation lifecycle with twelve-month windows. Everything below assumes that revision as the design center. Statements about client behavior carry the date they were measured, because they will rot.\n\n## Step 0: no policy in the MCP server\n\nThis server contains no policy. Every tool call becomes an authenticated HTTP request to the site's existing operator API, where all rules live: content validation, the publication policy, rate limits on operations, commit attribution. The rationale is covered at length in [the API versus MCP article](/writing/policy-in-the-api-not-the-mcp) and compresses to one sentence here: rules implemented twice are two rule sets, and two rule sets drift.\n\nWhat this article adds is the enforcement of that constraint, structural and mechanical. Structural: the server is a separate Worker in a separate repository, which means the tempting shortcut, an in-process call into the application that skips the API's authentication and rate limiting, does not exist as an option. The only path to the machinery is the front door. Mechanical: a check script in the build asserts the constraint, no imports from the application codebase, no database binding, no repository credential, with the allowed-bindings list derived from the Worker's own configuration. It currently holds 225 assertions, and its test harness confirms it catches 13 of 13 deliberately planted violations, in keeping with a rule this series applies everywhere: a guard you have never observed failing has not been verified.\n\n## Step 1: interrogate your real client before writing the authorization server\n\nThe specification tells you what a compliant client does. It cannot tell you what the client your users actually run does this week, and building to the specification alone risks shipping an authorization flow no real client completes. So the first deploy was not the server. It was a measurement probe: a fake authorization server that walks any connecting client as far as it will go, records every request, and then stops deliberately with a page saying so.\n\nOne design property of the probe is worth copying: it had no control plane at all. No endpoint to read captures, no administrative token, nothing to protect or leak; the recorded data was read out of band through the platform's storage tooling. When you must deploy something intentionally insecure-looking, giving it zero credentials and zero management surface is the safest shape it can take.\n\nThe captures settled every open question, and the answers are dated 2026-07-29. The consumer client completes the full modern authorization walk: resource metadata discovery including the path-scoped variant, PKCE with S256, audience-bound token requests per RFC 8707, and client identification by metadata document rather than dynamic registration, which meant an entire registration endpoint did not need to exist. Both target clients, however, still open with the previous revision's initialization handshake rather than the new stateless entry, which converts legacy support from a courtesy into a load-bearing requirement, with a retirement condition the same probe can answer later. One anticipated risk did not materialize: strict audience matching accepted the client's path-qualified token requests, so a permissive matching flag I had held in reserve stays unset. I would generalize that last habit: do not loosen a security control preemptively on a guess; wait for the evidence that you must.\n\n## Step 2: choose the protocol implementation by conformance score\n\nI wanted to hand-roll the protocol layer. A dependency-free server is auditable end to end, and a five-tool wrapper seemed small enough to justify it. The conformance scenarios disagreed, and I am reporting the numbers because losing to a scoreboard is the informative way to lose: my dependency-free draft scored 0 of 8 on the new stateless-protocol suite and 3 of 8 on header validation, where the platform's maintained library (Cloudflare's agents package with the official SDK) scored 24 of 28 and 13 of 13.\n\nThe miss that best explains the gap is instructive about the revision itself: per-request protocol metadata belongs inside the message's params object, not at the JSON-RPC top level, and a top-level version field is not a versioned message but a malformed one, rejected as invalid. My own test harness made that exact mistake while testing for it. This is the class of error conformance suites exist to catch, hand-rolling multiplies it, and the residual gaps in the library path are named rather than hidden: the library currently lacks one required error type from the new revision (my dependency-free draft scored identically on that scenario, which attributes it upstream), and four applicable scenarios are not yet wired into the gate. A conformance baseline that names its exclusions is worth more than a sweep that quietly skips them, and the baseline runs in the build with an expected-failures file that doubles as the machine-readable record of what is deliberately out of scope.\n\n## Step 3: make legacy support a seam you can delete, not a flag you can forget\n\nThe library offers legacy-revision support as a configuration default, which would make eventual removal mean flipping a vendor flag and hoping nothing else depended on it. I inverted that: the server rejects legacy traffic at the library level and handles it in one explicitly named module, using the library's request-classification hook. Retiring the old protocol era is therefore deleting one file, after which the modern-only behavior proves itself, and the retirement condition is empirical rather than calendrical: the probe from step 1 stays in the repository as the instrument, and the module goes when it shows the real clients opening with the modern entry. A seam you can delete is a commitment you can verify; a compatibility flag is a commitment you will forget you made.\n\n## Step 4: two credentials: OAuth for identity, bearer token for policy\n\nThe server authenticates twice, and keeping the two legs distinct in both code and documentation prevents a whole category of confused debugging. The client leg answers who is operating: an OAuth 2.1 flow, gated on the site owner's identity through an upstream identity provider, re-verified on requests rather than trusted from a session. The API leg answers what operators may do: the same bearer token any raw caller of the publishing API would present, which means the API neither knows nor cares that an MCP layer exists. An identity failure and a policy refusal are different events with different remedies, and conflating the credentials makes them indistinguishable in logs at exactly the moment you need them distinguished.\n\n:::diagram{title=\"Two credentials, two questions, two different failures\" alt=\"A sequence diagram with four participants: the AI client, the MCP server, the identity provider, and the publish API. The client calls a tool and the MCP server answers with an authorization challenge. The client completes an OAuth walk at the identity provider, which returns an identity restricted to the site owner, and calls the tool again carrying it. The MCP server verifies that identity on the request rather than trusting a session, and a note across those two participants records that this leg answers who is operating and that its failure is a 401. The MCP server then makes the same bearer-authenticated request a script would make to the publish API, and a second note across those two records that this leg answers what operators may do and that its failure is a 403 naming the policy. The API answers either 200 or a refusal carrying its policy name, and the MCP server passes it back verbatim.\"}\n```mermaid\nsequenceDiagram\n  participant C as AI client\n  participant M as MCP server\n  participant I as Identity provider\n  participant API as Publish API\n  C->>M: call a tool\n  M-->>C: authorization challenge\n  C->>I: OAuth walk\n  I-->>C: identity, owner only\n  C->>M: call a tool, with identity\n  M->>M: verify per request\n  Note over C,M: who is operating? 401\n  M->>API: the request a script would make\n  Note over M,API: what may operators do? 403\n  API-->>M: 200, or refusal + policy name\n  M-->>C: verbatim\n```\nThe MCP server holds no policy. It verifies identity, then makes the same\nrequest a script would make, and repeats whatever comes back word for word.\n:::\n\nTwo tool-design practices from this layer that cost little and are usually skipped. Tool descriptions carry the policy: the save tool's own documentation states [the human-reserved first-publication rule](/writing/agent-write-access-to-a-live-site), tells the agent to check a post's publishability field before attempting, and says the refusal is correct behavior. The agent is constrained by documentation before its first call. And refusals pass through verbatim: the API writes good refusal messages, policy name attached, and the wrapper never summarizes or softens them, because a translation layer that paraphrases the lock misinforms the visitor.\n\n## Step 5: verify end to end from the real client, then verify the layer adds nothing\n\nThe acceptance sequence, run against production before this article was written: the OAuth walk completed from the real client; both protocol eras answered correctly; the five tools listed with honest annotations (reads marked read-only, deletion marked destructive, which matters because clients build confirmation interfaces from those hints, and lying to them is lying to the user); a draft was created as one atomic attributed commit; the first-publication attempt was refused with the policy named; an edit landed; a save containing a prohibited character was rejected naming line and column; a malformed slug was rejected by the tool's own input schema before any network request, which is the cheap failure you want schemas to buy; and a concurrent burst of 45 requests saw 17 refused with retry guidance, the admitted count matching the window's earlier spending. One field in the save response deserves attention: the draft's answer-index synchronization reported zero documents uploaded, which is [the draft-exclusion rule from the previous article's incident](/writing/agent-write-access-to-a-live-site) made visible in every response rather than asserted in documentation.\n\nThe final claim, that the wrapper adds nothing, was verified by running the same operations raw against the API and comparing: identical commits, identical gate messages, identical refusal prose and policy names reaching the caller. For a layer whose entire design goal is transparency, that comparison is the acceptance test, and it is cheap.\n\n## Scrubbing git history before making the repository public\n\nThis server's repository is public, both because the code demonstrates the method and because this portfolio's convention is that configuration identifiers stay out of public trees. Getting there surfaced three traps for anyone rewriting git history for the same reason, recorded here because a naive scrub leaves the job half done. The history rewrite deletes your local copy of the very file being protected, so restore it before your next deploy fails. The rewrite's backup references keep the old objects reachable, so a verification pass run before deleting them and expiring the reflog reports failure against a rewrite that worked. And identifiers can hide in commit messages, pasted from tool output, where a tree-only rewrite cannot see them; a message-filter pass is a separate step. One residue is disclosed rather than hidden: pre-rewrite objects remain fetchable by their hashes until the host's garbage collection runs, the exposed content in this case being storage namespace identifiers, which are addresses rather than keys.\n\n## Limitations\n\nThe client-behavior findings are dated and will rot as clients update; the retirement condition for the legacy module is empirical and its instrument ships in the repository. The conformance standing is a baseline with named gaps, not a clean sweep. The adds-nothing claim is scoped to the tool surface tested. And the no-policy constraint, this article's spine, is the right design for a wrapper over a service that already has an API with rules; a standalone MCP server over public data has no policy to misplace and can reasonably be simpler than everything described here.\n\nThis article closes [the series](/writing/ten-years-on-cloudflare), following [the trust model](/writing/agent-write-access-to-a-live-site). It was drafted by the agent, staged through the tools it describes as a draft with the operator marker on its commit, and published by the one action the server cannot perform.\n\n## Update, August 2026\n\nA month of operation, two external audits of the site behind this server, and the design's central bet paid in the most boring way possible: the MCP server's diff for the entire hardening month is empty. Every fix the audits prompted, rate limits, pacing, a compensation path, landed in the API where the policy lives, and this layer relayed the new refusals verbatim without knowing they were new, which is exactly what a doorbell should do. The token that this wrapper presents to the API also gained a designed lifecycle, hashed storage, a maximum lifetime, overlap rotation, built on a branch and awaiting the operator's cutover; the change is entirely on the API's side of the line, which is the two-credentials split from step 4 doing its job. The bearer token is policy's credential, so its lifecycle belongs to the API, and this server will not need a line of code changed when it rotates.\n",
      "summary": "How to build a production MCP server under the 2026-07-28 specification: measure your real client with a probe before choosing auth, pick the protocol library by conformance score, isolate legacy support in one deletable module, and keep policy out entirely.",
      "date_published": "2026-07-30T00:00:00.000Z",
      "date_modified": "2026-09-27T00:00:00.000Z",
      "tags": [
        "agents",
        "cloudflare",
        "mcp",
        "oauth",
        "workers"
      ]
    },
    {
      "id": "https://dustinedwards.info/writing/ten-years-on-cloudflare",
      "url": "https://dustinedwards.info/writing/ten-years-on-cloudflare",
      "title": "Every Cloudflare product, and which ones this site runs on",
      "content_text": "\nThere is no server behind this site. Ten years ago I put my first domain behind Cloudflare the way everyone did then: a shared HostGator box ran the real site and Cloudflare was the DNS, the cache, and the orange cloud in front of it. It was not a place where software ran, and in 2016 it mostly wasn't. This year I rebuilt so that the cloud is the whole thing. The pages, the database, the uploads, the search, the AI answers, the alert mail, and the publishing pipeline all run on Cloudflare products, and nothing else is in the stack.\n\nSo this is the post I wanted when I started: every developer product Cloudflare sells as of September 8, 2026, what each one does in plain words, and whether this site uses it, where, and why or why not. The refusals are in the table with everything else. A survey that only lists what worked is an advertisement. This post replaced an earlier version, published 2026-07-30, that surveyed the same rebuild before the table existed; its dated measurements are carried forward below.\n\n## Which products this site runs on\n\nUsed means a binding or a configured feature that production depends on today. Not used means considered and passed over; the reason is in the product's own entry below. The numbers are dated because every one of them moves.\n\n| Product | What it does | This site | Where |\n|---|---|---|---|\n| Workers | Runs your code on Cloudflare's network | Used | Everything: two Workers, the site and a watchdog |\n| Static Assets | Serves files from a Worker with no invocation | Used | `public/`, through the `ASSETS` binding |\n| Workers Cache | Caches a Worker's responses at the edge | Used | The renderer entrypoint; the gateway is deliberately uncached |\n| D1 | SQLite database, managed | Used | Posts, tags, search index, media index |\n| KV | Fast key-value store, eventually consistent | Used | Login sessions, the Ask answer cache, watchdog state |\n| R2 | Object storage, no egress fees | Used | Three buckets: uploads, social cards, a mirror of uploads |\n| Queues | Message queue between Workers | Used | R2 upload events feeding the media index |\n| Durable Objects | A single-instance object with its own storage | Used | The rate limiter and daily budget for Ask |\n| Analytics Engine | Time-series data you write from a Worker | Used | Per-page traffic counts, no cookies, no IPs |\n| Images | Resize and convert images on request | Used | Every thumbnail and content width |\n| AI Search | Retrieval and cited answers over your content | Used | The Ask endpoint, above classic search |\n| Email Service | Send email from a Worker | Used | The watchdog's alert mail |\n| Email Routing | Receive mail on your domain and forward it | Used | Inbound mail on the domain |\n| Workers Observability | Logs and traces for Workers | Used, logs only | Traces are off on purpose (see the entry) |\n| Cron Triggers | Run a Worker on a schedule | Used | The watchdog, every 15 minutes |\n| Workers AI | Run AI models on Cloudflare GPUs | Indirect | Only through AI Search; no direct binding |\n| Rate Limiting binding | A built-in per-key rate limiter | Refused | Measured: it sheds load, it does not count |\n| Vectorize | Vector database for embeddings | Refused | Two FTS5 indexes answer this corpus |\n| Pages | Hosting for static and framework sites | Not used | Workers with static assets does the same job |\n| Workers Builds | Build and deploy from a git push | Not used | Deploys go through a gated ship script |\n| Hyperdrive | Connection pooling to an external Postgres or MySQL | Not used | There is no external database |\n| Workflows | Durable multi-step jobs with retries | Not used | Nothing here runs long enough |\n| Containers | Run any container next to a Worker | Not used | Nothing needs a runtime beyond V8 |\n| Sandboxes | Isolated code execution for agents | Not used | Agents write through an API, not by running code |\n| Browser Run | Headless browser as a service | Not used | Charts and diagrams render at build time |\n| Workers Agents SDK | Framework for stateful AI agents | Not used | The agent surface is an MCP server on plain Workers |\n| AI Gateway | Proxy and observability for model calls | Not used | Ask's one model call is metered by a Durable Object |\n| Stream | Video hosting and playback | Not used | No video |\n| RealtimeKit | Live audio and video | Not used | No live features |\n| Pipelines | Streaming ingestion into R2 | Not used | Analytics Engine covers the one stream |\n| Data Platform | Catalog and query data in R2 | Not used | Nothing to catalog |\n| Artifacts | Git-native versioned storage | Not used | GitHub is the repository |\n| Secrets Store | Account-level secret storage | Not used | Nine `wrangler secret` values, gated |\n| Turnstile | Bot check without a captcha | Not yet | Planned for the newsletter form |\n| Web Analytics | Client-side analytics beacon | Refused | Blocked by this site's CSP, and not needed |\n| Zaraz | Third-party tag loading at the edge | Not used | There are no third-party tags |\n| Access | Login in front of an application | Not used | The admin plane uses Better Auth |\n| Cache Reserve | Persistent cache for static content | Not yet | Needs the zone; waits for DNS cutover |\n| Workers for Platforms | Run customers' Workers inside yours | Not used | One customer |\n\n## What each one does, in the words I would use to a colleague\n\n### Workers\n\nA Worker is a function that receives a request and returns a response, running in a V8 isolate on Cloudflare's network in whichever city is closest to the reader. No server to size, no region to choose, cold starts small enough that I have never thought about them. Everything else on this list is something a Worker can be given a binding to.\n\nThis site is two Workers. The first serves every page and holds the entire markdown rendering pipeline, syntax highlighting included, because every public page works with JavaScript disabled. Its upload measured 8.4 MiB on 2026-09-08 (2.1 MiB gzipped; 3.78 MB when the first version of this post published on 2026-07-30). The second is a watchdog that reads the site's health endpoint every fifteen minutes through a service binding and repairs what it can. The client-side enhancement budget for the whole blog, progress bar, table of contents, copy buttons, footnote previews, lightbox, comes to about three kilobytes gzipped; the public plane ships no framework script at all.\n\n### Static Assets\n\nA Worker can serve a directory of files directly from the edge with no code running, through a binding with exactly one method, `fetch()`. That method is the whole interface: a Worker can serve any path it is given and discover none of them, which is why this site's media index reads a committed manifest of what is in `public/` instead of asking the binding.\n\n### Workers Cache\n\nCloudflare can store a Worker's responses at the edge and answer repeat requests without running the code. This site turns it on for the renderer and off for the gateway that sits in front of it. The gateway does three things that must never be skipped, the HTTPS redirect, the theme cookie read, and the traffic count, and a cached gateway would skip all three. Every response without an explicit `Cache-Control` defaults to `private, no-store`, because the platform would otherwise cache a logged-in admin page for two hours under standard heuristics and serve it to anyone. That default is the one line of configuration I would tell every Workers user to check first.\n\n### D1\n\nA relational database, which is to say SQLite, replicated and managed by Cloudflare. D1 spent its early life with a reputation for being a toy and that reputation is stale. It is real SQLite, and real SQLite ships FTS5 full-text search with the database. This site's search is two FTS5 indexes over the same corpus, one unstemmed for names and identifiers and one Porter-stemmed for prose, merged with reciprocal rank fusion. Measured 2026-08-28 against production: median 64 ms warm, 145 ms cold, for the database query alone (6 ms at publication on 2026-07-30, against a smaller corpus with the query timed in isolation; the two were not taken the same way, which is the argument for dating a number). The schema and the fusion method are in [the search article](/writing/site-search-on-d1).\n\nD1 also holds the media index, and that is a capability decision rather than a scale one: R2 lists objects in key order and promises nothing else, so \"sort by date, filter unused, count by type\" each need a query, and a database is where queries live.\n\nThe constraint every D1 user should know: the platform's export command fails outright on any database containing FTS5 tables. The working per-table procedure is documented in the search article, and as of 2026-09-08 a weekly job restores that export into a scratch database and compares it against production, because a backup that has never been restored is a hope. That drill found three bugs in the documented restore path on its first run.\n\n### KV\n\nA key-value store that reads fast anywhere in the world and accepts that a write takes a moment to be seen everywhere. Right for configuration and caches, wrong for counters. This site keeps login sessions in it, the watchdog's alert state, and the Ask answer cache, keyed by a hash of the normalised question. The cache sits in front of the daily spending ceiling rather than behind it, so a repeated question reaches no model and costs nothing.\n\n### R2\n\nObject storage with an S3-compatible API and no charge to read the data back out. This site runs three buckets, split on lifecycle rather than on what the admin UI calls them. Uploads are content-addressed: the key is a hash of the bytes, which was decided on security rather than tidiness, because the old key was guessable from a slug the sitemap publishes. Social cards are derived and regenerable, so they get their own bucket that may be emptied. The third bucket is a mirror of uploads that no code path on this site can delete from; a gate fails the build if one appears, and the health endpoint compares every object against its twin.\n\n### Queues\n\nA message queue: one Worker puts a message on it, another consumes it, with retries and a dead-letter queue for permanent failures. This site's media index is written this way. An upload writes only to R2; R2 emits an event; the consumer derives the database row from the object as it is now, never from what the message claimed. That is what makes replay and out-of-order delivery converge, and it is why the Worker never does a dual write.\n\n### Durable Objects\n\nA single instance of a JavaScript class, addressed by name, with its own storage. Only one copy exists anywhere in the world, so it is the platform's answer to coordination. This site uses one for the two guards in front of Ask, the only public endpoint that costs money per request: a per-IP burst limit and a site-wide daily ceiling.\n\nBuilding it taught me the one fact from this whole rebuild I repeat most often: single-threaded is not transactional. A Durable Object using the asynchronous storage API admitted eight requests through a ceiling of three, because a read and a write separated by an `await` are not atomic. The synchronous SQLite storage API is the fix; the same object rewritten on it admitted exactly three. The measurements are in [the AI answer layer article](/writing/ai-answer-mode-on-site-search).\n\n### Analytics Engine\n\nA write-only time-series store you append to from a Worker and query later with SQL. This site writes one row per HTML response: path, referrer host, country, and a coarse mobile flag. No cookie, no IP address, no identifier of any kind, so nothing joins two requests together. It is here because the alternative, Cloudflare's own Web Analytics beacon, is a third-party script, and this site's Content Security Policy would have to be loosened to admit it.\n\n### Images\n\nResize, crop and convert images on request, from an original you keep in R2. This site derives every thumbnail and every content width from one uploaded original through the binding, and the results are cached by the Workers Cache above. The binding rather than the URL syntax, and that is forced: the URL interface answers 404 on a workers.dev hostname because it needs a customer zone. A detail that stops mattering at DNS cutover.\n\n### AI Search\n\nYou give it a corpus; it handles chunking, embedding, hybrid retrieval, reranking, and answer generation with citations. This site's Ask mode is built on it, above classic search rather than instead of it, because the two fail on opposite inputs. Measured on a shared query set: classic search found two results the AI retrieval missed, all exact tokens, and the AI retrieval found three the classic engine missed, all natural-language questions. Neither subsumes the other. It is a metered product, so it sits behind the Durable Object above and a daily ceiling I chose. Its first token arrived in 2.1 to 6.5 seconds warm when measured on 2026-07-30; that figure has not been re-taken because probing it costs money per request.\n\n### Email Service\n\nSend email from a Worker through a binding, from an address on a domain you have onboarded. The watchdog uses it to mail me once when the site goes unhealthy and once when it recovers, with the state kept in KV so a bad afternoon sends one message rather than twelve. Sending to a verified destination address is free.\n\n### Email Routing\n\nReceive mail on your domain and forward it wherever you like. Inbound mail for dustinedwards.info lands here and forwards to my mailbox. Both halves of the email story share one set of SPF and DKIM records that Cloudflare manages, and DMARC on the domain is set to reject.\n\n### Workers Observability\n\nLogs and traces from your Workers, kept in the dashboard and exportable elsewhere. This site keeps logs on and traces off, and the off is a ruling rather than a default. A trace span carries the full request URL, and this site puts a capability token in the path of every draft preview link, so exporting traces would ship preview access to a third party. Invocation logs are off for a related reason: they recorded the reader's IP and session cookie for seven days with no field-level redaction available.\n\n### Cron Triggers\n\nRun a Worker on a schedule. The watchdog runs every fifteen minutes. The site Worker has no cron, and the empty array in its config is the statement: an hourly trigger that nothing handled sat on the platform for fifteen days in August, throwing 24 times a day, invisible to a gate that only read files. The gate now reads the platform too.\n\n### Workers AI\n\nRun open models on Cloudflare's GPUs from a Worker. This site never calls it directly; AI Search does the model work for Ask on its own. If Ask ever needs a model the retrieval product does not offer, this is where it would come from.\n\n### Rate Limiting binding, refused\n\nA built-in per-key limiter you declare in config. Measured on 2026-07-30 with a limit of five per sixty seconds and twelve concurrent requests: it refused one, then two, then nine, then zero across four runs. Its documentation says it sheds sustained load with eventual consistency, and that is true; it does not count, and a limiter guarding a budget has to count. The Durable Object replaced it.\n\n### Vectorize, refused\n\nA vector database for embeddings, the usual foundation for semantic search. Rank fusion over two FTS5 indexes answers this corpus in about fifteen lines with no vectors, and zero-result searches are the cheapest signal this site has for what to write next. That is the only evidence that would reopen the question.\n\n### Pages and Workers Builds, not used\n\nPages hosts static and framework sites with a build on every push; Workers Builds does the same for Workers. Workers with static assets now does everything Pages did for this site, and Cloudflare's own direction has been to fold Pages into Workers. Deploys here go through a ship script that refuses unless the working tree is clean, CI is green for that exact commit, and every offline gate passes; a build on push would skip all of that.\n\n### Hyperdrive, Workflows, Containers, Sandboxes, Browser Run, not used\n\nHyperdrive pools connections to a Postgres or MySQL you already run somewhere; there is no such database here. Workflows runs multi-step jobs that survive failures and can wait for a human; nothing here runs longer than a request. Containers runs any Docker image next to a Worker; nothing here needs a runtime beyond V8. Sandboxes gives an agent an isolated place to execute code; this site's agents write through an API with policy enforced server-side, and never run code. Browser Run is a headless browser you can drive from a Worker; charts and diagrams here render to SVG at build time, on purpose, so the page carries no work.\n\n### Agents SDK and AI Gateway, not used\n\nThe Agents SDK is a framework for long-lived stateful agents on Durable Objects. This site's agent surface is the other direction: an MCP server that lets an outside agent read and write posts, with exactly one act, first publication, reserved to me in code. AI Gateway proxies model calls for logging, caching and cost control; Ask makes one model call per uncached question and a Durable Object already meters it.\n\n### Stream, RealtimeKit, Pipelines, Data Platform, Artifacts, not used\n\nVideo hosting, live audio and video, streaming ingestion, data catalogs, and git-native storage. A text site with a media library of a few dozen images has no use for any of them, and I would rather say so than pad the used column.\n\n### Secrets Store, Access, Zaraz, Web Analytics, Workers for Platforms\n\nSecrets Store centralises secrets across Workers; this site's nine secrets live in `wrangler secret` and a gate asserts every one is set and none is in git. Access puts a login page in front of any application; the admin plane runs its own login on Better Auth because the policy it enforces lives in the application, not in front of it. Zaraz loads third-party tags at the edge; there are none. Web Analytics is the beacon the Analytics Engine entry above explains. Workers for Platforms runs other people's Workers inside yours; I have one customer.\n\n### Turnstile and Cache Reserve, not yet\n\nTurnstile is Cloudflare's bot check without a puzzle; it goes in front of the newsletter form when the newsletter exists. Cache Reserve keeps static content in a persistent cache; it needs a zone, and this site is still served from a workers.dev hostname while the old WordPress install answers at the apex. Both wait for DNS cutover.\n\n## Where the platform pushed back\n\nWorkers refuse to compile WebAssembly at runtime. A security decision, and it ruled out server-side social-card rendering and complicated the fix for the strangest bug of the build: a syntax highlighter that produced different bytes for the same input across runs. The deterministic engine and the loader arrangement that satisfies both Node and the Worker are in [the content pipeline article](/writing/posts-in-git-served-from-d1).\n\nD1's export fails on FTS5 tables, as above. The rate limiting binding does not count, as above. A Durable Object's async storage is not transactional, as above. Each of these is now a check script or a documented rule rather than a memory, which is the only form a platform lesson is worth keeping in.\n\n## Who reads this site, and what Cloudflare is doing about it\n\nThis site treats AI agents as an audience in both directions. Inbound, every post serves a markdown twin at a predictable URL, an llms.txt file maps the site, the search endpoint answers in JSON to any client that asks, and the search engine is exposed over the Model Context Protocol. Outbound, the site is operated by agents: an authenticated publishing API whose rules are enforced server-side, an MCP layer over it, and the assistants that helped build this system draft and edit posts through it, including this one. One act is reserved for me by a policy they cannot alter. The trust model, the incident that shaped it, and the protocol server are in [the agent write access article](/writing/agent-write-access-to-a-live-site), [the API versus MCP article](/writing/policy-in-the-api-not-the-mcp), and [the MCP server article](/writing/mcp-server-on-workers-with-oauth).\n\nThe economic half is happening at the CDN layer, which for a fifth of the web means it is happening at Cloudflare: managed robots.txt with machine-readable content signals, default blocking of AI training crawlers for new zones from September 15, 2026, and a pay-per-use marketplace for content that surfaces in AI answers. This site's own crawl settings get configured the day the DNS cutover lands, and that will be its own post once there is data in it.\n\n## What August taught, kept from the first version\n\nAugust was the month this system got audited from outside, twice, by a different AI model with read access to the repository and the wire. The audits caught three real holes that had shipped past every gate, an unauthenticated delete on a media route, a draft leak on the same route, and preview links with no rate limit, all live for 39 days before anyone noticed, and all closed the day they were reported. The audits were also wrong a lot: measured claim by claim against the code, roughly half their findings were stale, false, or cited numbers that did not exist. The rule that came out of it: an external audit is a list of claims to verify, not a list of facts.\n\nThe one thing the audits asked for that this site refused, deliberately: switching frameworks. The missing conveniences were on generic surfaces, while the defaults that matter here, a cache that fails toward privacy, a public plane that works without script, real bindings instead of an adapter, are all on the side the site is already standing on.\n\n## What I would use again\n\nAll fourteen. The primitives are small enough to hold in your head. The billing has never surprised me, which I value more than any feature. A one-person site now runs what would have been a small team's roadmap five years ago: a gated content pipeline where git is the source of truth, an edge-resident search engine, a hybrid AI answer layer with cost controls, external monitoring, restore drills, and a publishing path an agent can operate under enforced policy. Twenty-four products were considered and passed over for the reasons above, and every number here carries the date it was measured because every one of them will move.\n\nThe series, in reading order: [the color palette built and verified with code](/writing/color-palette-the-build-can-check), [the git-backed content pipeline](/writing/posts-in-git-served-from-d1), [the reading experience in a couple of kilobytes of JavaScript](/writing/blog-reading-without-javascript), [FTS5 search on D1](/writing/site-search-on-d1), [the AI answer layer](/writing/ai-answer-mode-on-site-search), [API versus MCP](/writing/policy-in-the-api-not-the-mcp), [the agent trust model](/writing/agent-write-access-to-a-live-site), and [the MCP server build](/writing/mcp-server-on-workers-with-oauth). Every quantitative claim in the series is reproducible from the site's repository.\n",
      "summary": "Every Cloudflare developer product as of September 2026 in one table: what Workers, D1, KV, R2, Queues, Durable Objects, AI Search and the rest actually do, and which ones a complete site runs on, where, and why. With the refusals and the constraints, dated.",
      "date_published": "2026-07-30T00:00:00.000Z",
      "date_modified": "2026-09-27T00:00:00.000Z",
      "tags": [
        "architecture",
        "cloudflare",
        "d1",
        "platform",
        "workers"
      ]
    },
    {
      "id": "https://dustinedwards.info/writing/posts-in-git-served-from-d1",
      "url": "https://dustinedwards.info/writing/posts-in-git-served-from-d1",
      "title": "How this blog stores posts in git and serves them from D1",
      "content_text": "\nThis article describes a method for building a blog's content layer so that the repository is the source of truth, the database is a serving layer, and a check proves the two agree. I use it in production on this site, in the form described in the update at the end since 26 August 2026. The article is written so that a reader with working knowledge of TypeScript, git, and a Cloudflare Workers project can reproduce the architecture, and it includes the two failure modes I encountered that I believe most implementations will also encounter, at the point in the procedure where they will appear.\n\nA note on scope before beginning. The method assumes a single author or a small set of trusted authors, content volumes in the hundreds to low thousands of documents, and a willingness to treat prose with the same discipline as code. At substantially larger scales, or with untrusted authors, several of the trade-offs below change, and I flag those points where they occur.\n\n## Step 1: where blog content should live: git versus the database\n\nA blog's content has to live somewhere, and the common candidates are: sanitized HTML in a database, authored through a rich editor; markdown in a database, authored through an admin form; markdown files in the repository; or a third-party headless CMS. Before this build I had production experience with the first two, across three sites, so the comparison here is empirical rather than speculative. All of the candidates work, in the sense that pages render. The differences appear in the properties you can enforce and the failure modes you inherit.\n\nThe architecture this method produced, as first built: markdown files in the repository are authoritative; a generator renders them and emits a committed artifact plus database rows carrying both the source and the rendered HTML; a check script fails the build whenever the committed artifact disagrees with a fresh generation from source; and the database serves every page read and owns full-text search. The update at the end replaces the committed artifact with provenance hashes on the database rows and moves the check to deploy time and to a scheduled health check. Everything else in this article stands.\n\nFour properties motivated the choice, and I recommend evaluating your own situation against each rather than adopting the conclusion.\n\nFirst, enforcement. Anything you can express as a lint rule, a schema, or a hook can gate a file at commit time, and none of it can see a database row. If your project enforces style rules on code, storing prose in the database creates exactly one class of text those rules cannot reach. Second, redundancy. With files, git is a complete second copy of the content. This stopped being abstract during the build, when I measured that Cloudflare's `wrangler d1 export` command fails outright on databases containing FTS5 virtual tables, which is precisely what a search feature adds; the full finding and the working per-table backup procedure are in [the search article](/writing/site-search-on-d1). A database-authored site with search may be holding the only copy of its prose behind a broken default backup path. Third, determinism. Content fixed at generation time makes assertions such as \"every post has a meta description\" build failures rather than periodic audits. Fourth, machine authorship. With content as files, an AI agent's edits arrive as reviewable commits rather than as UPDATE statements against production; that property becomes load-bearing in [the agent write access article](/writing/agent-write-access-to-a-live-site) later in this series.\n\nThe cost, stated plainly: publishing now requires producing a commit, which binds authoring to something that can reach the repository. Step 4 addresses this with a server-side path, but the dependency is real and permanent.\n\n## Step 2: one markdown renderer shared by the build and the Worker\n\nThe pipeline itself is conventional unified-ecosystem tooling: remark with GitHub Flavored Markdown and footnotes, rehype for HTML, heading anchors with a table of contents extracted during the same pass, a small directive syntax for figures that makes alt text mandatory, image dimensions probed at build time and written into every tag to prevent layout shift, and Shiki for syntax highlighting applied at render time so that highlighted code ships as static HTML with no client-side JavaScript.\n\nThe architectural rule that matters is that exactly one renderer module exists, imported by both the build scripts (Node) and the Worker (the editor's preview and save path). The reason is the gate in step 3: it compares bytes, and if two renderers exist, a mismatch is ambiguous between drift and implementation difference, which makes the gate useless. One renderer makes every byte difference meaningful.\n\nThe rule has a measurable price, and you should measure yours before accepting it. Bundling full Shiki into the Worker produced a 14 MB output, because Shiki includes every grammar. Restricting to `shiki/core` with an explicit language allowlist brought the highlighter to roughly 615 KB gzipped and the Worker from 1.49 MB to 3.55 MB. A language outside the allowlist renders as plain unhighlighted code, identically everywhere, which is the correct degradation.\n\n:::chart{type=\"bar\" x=\"configuration\" y=\"worker_mb\" title=\"Worker bundle size by highlighter configuration\" alt=\"Bar chart of Worker bundle sizes in megabytes for three highlighter configurations: no server-side highlighter at 1.49 MB, shiki core with a language allowlist at 3.55 MB, and full Shiki with every grammar at 14 MB. The allowlist configuration is the accepted middle.\"}\n```csv\nconfiguration,worker_mb\nno highlighter,1.49\nshiki/core allowlist,3.55\nfull Shiki,14.0\n```\nMeasured Worker bundle size under each highlighter option. The allowlist was the accepted trade; both alternatives are recorded because a future maintainer will be tempted by each.\n:::\n\nRecord the bundle cost in your decision log with the alternative you rejected, because a future maintainer will otherwise be tempted to split the renderer to save megabytes, and the megabytes are cheaper than a blind gate.\n\n## Step 3: a byte-comparison gate, verified by breaking it\n\nThe gate, as first built, was a script, run in the build and before deploys, that regenerated the artifact from source and byte-compared it against the committed version, failing with the differing file named. Conceptually it is small. Its value depends entirely on a verification habit that I want to state as a rule: a gate you have never observed failing has not been verified. Before relying on it, attack it deliberately. I used four plants: a hand-edited artifact, a stale artifact after a source edit, a deleted artifact, and invalid frontmatter. It caught all four with usable messages. Only then does a green result mean anything.\n\nMine then caught two conditions I had not planted, and both are worth knowing in advance because neither is specific to my implementation.\n\n## Expect this bug: git autocrlf makes one commit produce different bytes per machine\n\nThe first fresh checkout on a Windows machine failed the gate on two visually identical lines. The cause is git's `core.autocrlf=true` default in the absence of a `.gitattributes` file: the checkout rewrites line endings to CRLF, the generator embeds that markdown into the artifact and the database, and the same commit produces different published bytes depending on which machine ran the build. The fix is a `.gitattributes` entry pinning content paths to LF. Verify it the way the bug demands: clone to a temporary directory on the affected platform before and after, because this class of defect survives any gate that only ever runs on one machine. If your pipeline embeds file contents into generated output and your contributors span operating systems, I would treat this as a certainty rather than a risk.\n\n## Expect this bug: Shiki's default JavaScript engine is nondeterministic\n\nThe second finding took longer to isolate and I have not seen it documented elsewhere, so I will state it carefully and scope it honestly. Shiki's JavaScript regex engine, in the version and grammar set I tested, is not deterministic: eight renders of one TypeScript snippet within a single process produced two distinct outputs, and separate processes colored the same `=` token with three different theme colors. Token boundaries and output length were stable, which is why the variation hides; only the color assignments moved. Against a byte-comparison gate this is fatal, since the gate fails at random on any post containing code, and the natural misdiagnosis is that the pipeline changed rather than that the renderer is stochastic.\n\nThe resolution was switching to the Oniguruma engine, which was deterministic over the same test set. Oniguruma is a WebAssembly build, and Cloudflare Workers refuse runtime WebAssembly compilation as a security policy, so the loader that compiles bytes at runtime fails inside the Worker. The working arrangement is an injected loader: Node keeps the byte import, the Worker statically imports the compiled module, and both sides run the same engine with byte-identical output, at a bundle cost of 0.44 MB. My claim is scoped to the versions and grammars I measured; the procedure I would recommend regardless of version is to render one code-bearing document a few hundred times, in and across processes, and diff the outputs before you build anything that assumes rendering is a pure function.\n\n## Step 4: atomic two-file commits with the GitHub Git Data API\n\nIf a browser editor (or any server-side writer) joins the pipeline, the write path order is: validate, commit, then database. Validation runs server-side inside the action, because a commit made through GitHub's API bypasses every local hook; whatever your pre-commit machinery enforces must be re-enforced here or it is not enforced at all. In my implementation a save containing a prohibited character is rejected with the character, line, and column named, before any commit exists.\n\nThe commit shape was the part I most emphasized while the artifact was committed, because the naive version fails structurally rather than occasionally. Writing one file per save through GitHub's Contents API left the generated artifact one commit behind its source, which meant the drift gate was red on the main branch after every save, as routine. The construction that fixed it used the Git Data API to land the markdown and the regenerated artifact as a single commit, in four calls: create a blob per file, create a tree containing both against the base tree, create a commit whose parent is the base, then update the branch reference. The base commit hash you started from doubles as optimistic concurrency control: pass it when updating the reference, and a concurrent save from another tab is refused with the divergence named and no commit created. I verified this live with two editors racing; one landed, one was refused, nothing was overwritten. Since the update below, a save commits one file, and the concurrency guard on the branch head is the part that survived.\n\nOrder the database write after the commit succeeds, and fail the save whole if the repository is unreachable. The asymmetry is deliberate: a failed save is an inconvenience, while a database that disagrees with its source of truth is a standing lie that every later read repeats.\n\nVersion history then costs almost nothing, because git already holds it: list the file's commits, diff them, and implement restore as a new commit through the same atomic path rather than any history rewrite. One subtlety worth copying: restored content re-runs the validation gates, which closes a hole where restoring an old commit would republish prose that predates a rule and was never checked against it.\n\n## Step 5: the verification checklist for a git-backed content pipeline\n\nA checklist, in the order I would run it on a fresh implementation. Clone to a temporary directory on a second platform and run the gate; this exercises the line-ending defect. Render a code-bearing document repeatedly and diff; this exercises determinism. Plant each gate violation and confirm the failure names the file. Save from the editor and confirm exactly one commit carrying what the design says it carries. Race two saves and confirm one refusal with no commit. Take the database offline (or revoke the token) and confirm the save fails whole. Export your database the way you believe your backup works, and read the output file, because an empty file exits successfully.\n\n## Limitations and disclosures\n\nThe Worker carries the full rendering pipeline, 3.55 MB at the time the pipeline landed against 1.49 MB before it, and the figure has grown since with unrelated features; the trade was accepted with the measurement recorded, and a project with tighter size constraints could run the renderer only at build time by giving up the server-side editor preview and accepting a weaker gate. Social card generation in my implementation runs at build time only, because the rendering stack's WebAssembly requirements do not fit the Worker's compilation policy; a post published from the editor has no card until the next build, a gap I chose over the alternative of a broken image reference. The nondeterminism finding is scoped to the engine, grammars, and snippet set I measured, reproduced across processes; I make no claim about configurations I did not test. And the single-author assumption from the introduction matters here: with many concurrent authors, the one-commit-per-save model produces reference-update contention that this design does not address.\n\nThe property the method buys, stated once: the prose passes the same gates as the code, the database serves without holding custody, and the build can demonstrate, on every run, that what readers receive is what the repository says. This is the third post in [the series](/writing/ten-years-on-cloudflare), following [the palette method](/writing/color-palette-the-build-can-check); the next one covers [the reading experience built on this foundation under a no-client-JavaScript constraint](/writing/blog-reading-without-javascript).\n\n## Update, August 2026\n\nStep 4's asymmetry got its missing half. The original design fails the save whole when the commit cannot land, which is right, but it left the opposite window unhandled: commit landed, database write failed, repository and serving layer disagreeing until someone noticed. The save now retries the database write once; if the retry also fails, the divergence is recorded where the admin's sync status reads it and the error names the post, the commit that landed, and the repair. The commit is never reverted to make the database happy, for the reason this article already stated: the repository is the source of truth, so the serving layer converges toward it and never the reverse. Two related hardenings landed with it. A missing generated artifact threw instead of quietly reading as an empty corpus, for as long as the artifact was committed; that code left with the artifact in the second update below. And the deploy script now refuses to ship any commit that continuous integration has not concluded green for, which closes the gap where a laptop could outrun the gates this article spends its middle third building.\n\n## Update, 26 August 2026: the committed artifact came out\n\nTwo independent reviews of this pipeline in August 2026 disagreed on one point and agreed on its cause. Both said the committed artifact was internally coherent. One said it was defensible under the constraints that produced it; the other said no shop would copy it. Both were right, and this update records what I did about it.\n\nWhat the committed artifact bought was never agreement between the two writers. They agree because they share one renderer, and that stayed true with or without the file. What it bought was that the agreement could be checked with no database and no network, at commit time, on a fresh clone. In July, with no continuous integration and one laptop deploying, that was the only place a check could run. By late August the deploys ran in CI behind the full gate tier, and the reason had gone.\n\nWhat it cost was measured before it was removed. Every editor save downloaded the whole artifact from GitHub, 648 KB for twelve posts at 283 to 528 ms, through an endpoint that returns nothing above 1 MB, which put every save, every delete, and the media library about six posts from failing at once. Every content change churned a generated file in git. And the two-file commit, the date written at sync time, and the dimensions written into media keys all existed to keep two copies byte-identical.\n\nThe design now: git holds markdown and nothing rendered. Both writers still render through the one shared pipeline, and the database holds the only rendered copy. Every row records the git blob hash of the markdown it was rendered from and a hash of the render. The deploy renders the corpus fresh and prints a table by post: unchanged, source changed, render drift. Render drift, the same source producing different bytes in the Worker and in Node, is the condition the byte gate used to catch, and it now fails the deploy after the deploy stands rather than blocking a commit. On the first run the table showed zero render drift across the whole corpus, which is the Shiki finding above holding two months later. A scheduled health check compares each row's source hash against the repository listing every fifteen minutes and re-renders any post that differs, so a markdown commit from any machine is live within one poll with no deploy.\n\nWhat that changes in the article above. Step 2 stands: one renderer, still on the Oniguruma engine, because determinism still matters when two runtimes render the same source. Step 3's gate still exists in a smaller form: it renders the corpus twice and fails on any byte difference, which is the determinism test the checklist in Step 5 describes, and it still byte-compares the one generated file that stayed committed, a repository scan with no database owner. The two bugs stand entirely; both would bite this design as surely as the last one. Step 4 is where the text is now history: saves commit one file, the concurrency guard on the branch head is unchanged, and the four-call construction is no longer needed. The property the method buys is unchanged and is now checked continuously instead of once: what readers receive is what the repository says.\n",
      "summary": "How to build a content pipeline where markdown in git is the source of truth and D1 serves every read: one deterministic renderer shared by the build and the Worker, a gate verified by breaking it, atomic commits through the GitHub API, and the two bugs to expect. Updated August 2026, when the committed artifact came out in favor of provenance hashes, a deploy-time drift table, and a self-repairing health check.",
      "date_published": "2026-07-28T00:00:00.000Z",
      "date_modified": "2026-09-27T00:00:00.000Z",
      "tags": [
        "architecture",
        "cloudflare",
        "content-model",
        "d1",
        "workers"
      ]
    }
  ]
}
