{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Dustin Edwards blog: content-model",
  "home_page_url": "https://dustinedwards.info/writing/tags/content-model",
  "feed_url": "https://dustinedwards.info/writing/tags/content-model/feed.json",
  "description": "Posts tagged content-model.",
  "language": "en-US",
  "authors": [
    {
      "name": "Dustin Edwards",
      "url": "https://dustinedwards.info"
    }
  ],
  "items": [
    {
      "id": "https://dustinedwards.info/writing/posts-in-git-served-from-d1",
      "url": "https://dustinedwards.info/writing/posts-in-git-served-from-d1",
      "title": "How this blog stores posts in git and serves them from D1",
      "content_text": "\nThis article describes a method for building a blog's content layer so that the repository is the source of truth, the database is a serving layer, and a check proves the two agree. I use it in production on this site, in the form described in the update at the end since 26 August 2026. The article is written so that a reader with working knowledge of TypeScript, git, and a Cloudflare Workers project can reproduce the architecture, and it includes the two failure modes I encountered that I believe most implementations will also encounter, at the point in the procedure where they will appear.\n\nA note on scope before beginning. The method assumes a single author or a small set of trusted authors, content volumes in the hundreds to low thousands of documents, and a willingness to treat prose with the same discipline as code. At substantially larger scales, or with untrusted authors, several of the trade-offs below change, and I flag those points where they occur.\n\n## Step 1: where blog content should live: git versus the database\n\nA blog's content has to live somewhere, and the common candidates are: sanitized HTML in a database, authored through a rich editor; markdown in a database, authored through an admin form; markdown files in the repository; or a third-party headless CMS. Before this build I had production experience with the first two, across three sites, so the comparison here is empirical rather than speculative. All of the candidates work, in the sense that pages render. The differences appear in the properties you can enforce and the failure modes you inherit.\n\nThe architecture this method produced, as first built: markdown files in the repository are authoritative; a generator renders them and emits a committed artifact plus database rows carrying both the source and the rendered HTML; a check script fails the build whenever the committed artifact disagrees with a fresh generation from source; and the database serves every page read and owns full-text search. The update at the end replaces the committed artifact with provenance hashes on the database rows and moves the check to deploy time and to a scheduled health check. Everything else in this article stands.\n\nFour properties motivated the choice, and I recommend evaluating your own situation against each rather than adopting the conclusion.\n\nFirst, enforcement. Anything you can express as a lint rule, a schema, or a hook can gate a file at commit time, and none of it can see a database row. If your project enforces style rules on code, storing prose in the database creates exactly one class of text those rules cannot reach. Second, redundancy. With files, git is a complete second copy of the content. This stopped being abstract during the build, when I measured that Cloudflare's `wrangler d1 export` command fails outright on databases containing FTS5 virtual tables, which is precisely what a search feature adds; the full finding and the working per-table backup procedure are in [the search article](/writing/site-search-on-d1). A database-authored site with search may be holding the only copy of its prose behind a broken default backup path. Third, determinism. Content fixed at generation time makes assertions such as \"every post has a meta description\" build failures rather than periodic audits. Fourth, machine authorship. With content as files, an AI agent's edits arrive as reviewable commits rather than as UPDATE statements against production; that property becomes load-bearing in [the agent write access article](/writing/agent-write-access-to-a-live-site) later in this series.\n\nThe cost, stated plainly: publishing now requires producing a commit, which binds authoring to something that can reach the repository. Step 4 addresses this with a server-side path, but the dependency is real and permanent.\n\n## Step 2: one markdown renderer shared by the build and the Worker\n\nThe pipeline itself is conventional unified-ecosystem tooling: remark with GitHub Flavored Markdown and footnotes, rehype for HTML, heading anchors with a table of contents extracted during the same pass, a small directive syntax for figures that makes alt text mandatory, image dimensions probed at build time and written into every tag to prevent layout shift, and Shiki for syntax highlighting applied at render time so that highlighted code ships as static HTML with no client-side JavaScript.\n\nThe architectural rule that matters is that exactly one renderer module exists, imported by both the build scripts (Node) and the Worker (the editor's preview and save path). The reason is the gate in step 3: it compares bytes, and if two renderers exist, a mismatch is ambiguous between drift and implementation difference, which makes the gate useless. One renderer makes every byte difference meaningful.\n\nThe rule has a measurable price, and you should measure yours before accepting it. Bundling full Shiki into the Worker produced a 14 MB output, because Shiki includes every grammar. Restricting to `shiki/core` with an explicit language allowlist brought the highlighter to roughly 615 KB gzipped and the Worker from 1.49 MB to 3.55 MB. A language outside the allowlist renders as plain unhighlighted code, identically everywhere, which is the correct degradation.\n\n:::chart{type=\"bar\" x=\"configuration\" y=\"worker_mb\" title=\"Worker bundle size by highlighter configuration\" alt=\"Bar chart of Worker bundle sizes in megabytes for three highlighter configurations: no server-side highlighter at 1.49 MB, shiki core with a language allowlist at 3.55 MB, and full Shiki with every grammar at 14 MB. The allowlist configuration is the accepted middle.\"}\n```csv\nconfiguration,worker_mb\nno highlighter,1.49\nshiki/core allowlist,3.55\nfull Shiki,14.0\n```\nMeasured Worker bundle size under each highlighter option. The allowlist was the accepted trade; both alternatives are recorded because a future maintainer will be tempted by each.\n:::\n\nRecord the bundle cost in your decision log with the alternative you rejected, because a future maintainer will otherwise be tempted to split the renderer to save megabytes, and the megabytes are cheaper than a blind gate.\n\n## Step 3: a byte-comparison gate, verified by breaking it\n\nThe gate, as first built, was a script, run in the build and before deploys, that regenerated the artifact from source and byte-compared it against the committed version, failing with the differing file named. Conceptually it is small. Its value depends entirely on a verification habit that I want to state as a rule: a gate you have never observed failing has not been verified. Before relying on it, attack it deliberately. I used four plants: a hand-edited artifact, a stale artifact after a source edit, a deleted artifact, and invalid frontmatter. It caught all four with usable messages. Only then does a green result mean anything.\n\nMine then caught two conditions I had not planted, and both are worth knowing in advance because neither is specific to my implementation.\n\n## Expect this bug: git autocrlf makes one commit produce different bytes per machine\n\nThe first fresh checkout on a Windows machine failed the gate on two visually identical lines. The cause is git's `core.autocrlf=true` default in the absence of a `.gitattributes` file: the checkout rewrites line endings to CRLF, the generator embeds that markdown into the artifact and the database, and the same commit produces different published bytes depending on which machine ran the build. The fix is a `.gitattributes` entry pinning content paths to LF. Verify it the way the bug demands: clone to a temporary directory on the affected platform before and after, because this class of defect survives any gate that only ever runs on one machine. If your pipeline embeds file contents into generated output and your contributors span operating systems, I would treat this as a certainty rather than a risk.\n\n## Expect this bug: Shiki's default JavaScript engine is nondeterministic\n\nThe second finding took longer to isolate and I have not seen it documented elsewhere, so I will state it carefully and scope it honestly. Shiki's JavaScript regex engine, in the version and grammar set I tested, is not deterministic: eight renders of one TypeScript snippet within a single process produced two distinct outputs, and separate processes colored the same `=` token with three different theme colors. Token boundaries and output length were stable, which is why the variation hides; only the color assignments moved. Against a byte-comparison gate this is fatal, since the gate fails at random on any post containing code, and the natural misdiagnosis is that the pipeline changed rather than that the renderer is stochastic.\n\nThe resolution was switching to the Oniguruma engine, which was deterministic over the same test set. Oniguruma is a WebAssembly build, and Cloudflare Workers refuse runtime WebAssembly compilation as a security policy, so the loader that compiles bytes at runtime fails inside the Worker. The working arrangement is an injected loader: Node keeps the byte import, the Worker statically imports the compiled module, and both sides run the same engine with byte-identical output, at a bundle cost of 0.44 MB. My claim is scoped to the versions and grammars I measured; the procedure I would recommend regardless of version is to render one code-bearing document a few hundred times, in and across processes, and diff the outputs before you build anything that assumes rendering is a pure function.\n\n## Step 4: atomic two-file commits with the GitHub Git Data API\n\nIf a browser editor (or any server-side writer) joins the pipeline, the write path order is: validate, commit, then database. Validation runs server-side inside the action, because a commit made through GitHub's API bypasses every local hook; whatever your pre-commit machinery enforces must be re-enforced here or it is not enforced at all. In my implementation a save containing a prohibited character is rejected with the character, line, and column named, before any commit exists.\n\nThe commit shape was the part I most emphasized while the artifact was committed, because the naive version fails structurally rather than occasionally. Writing one file per save through GitHub's Contents API left the generated artifact one commit behind its source, which meant the drift gate was red on the main branch after every save, as routine. The construction that fixed it used the Git Data API to land the markdown and the regenerated artifact as a single commit, in four calls: create a blob per file, create a tree containing both against the base tree, create a commit whose parent is the base, then update the branch reference. The base commit hash you started from doubles as optimistic concurrency control: pass it when updating the reference, and a concurrent save from another tab is refused with the divergence named and no commit created. I verified this live with two editors racing; one landed, one was refused, nothing was overwritten. Since the update below, a save commits one file, and the concurrency guard on the branch head is the part that survived.\n\nOrder the database write after the commit succeeds, and fail the save whole if the repository is unreachable. The asymmetry is deliberate: a failed save is an inconvenience, while a database that disagrees with its source of truth is a standing lie that every later read repeats.\n\nVersion history then costs almost nothing, because git already holds it: list the file's commits, diff them, and implement restore as a new commit through the same atomic path rather than any history rewrite. One subtlety worth copying: restored content re-runs the validation gates, which closes a hole where restoring an old commit would republish prose that predates a rule and was never checked against it.\n\n## Step 5: the verification checklist for a git-backed content pipeline\n\nA checklist, in the order I would run it on a fresh implementation. Clone to a temporary directory on a second platform and run the gate; this exercises the line-ending defect. Render a code-bearing document repeatedly and diff; this exercises determinism. Plant each gate violation and confirm the failure names the file. Save from the editor and confirm exactly one commit carrying what the design says it carries. Race two saves and confirm one refusal with no commit. Take the database offline (or revoke the token) and confirm the save fails whole. Export your database the way you believe your backup works, and read the output file, because an empty file exits successfully.\n\n## Limitations and disclosures\n\nThe Worker carries the full rendering pipeline, 3.55 MB at the time the pipeline landed against 1.49 MB before it, and the figure has grown since with unrelated features; the trade was accepted with the measurement recorded, and a project with tighter size constraints could run the renderer only at build time by giving up the server-side editor preview and accepting a weaker gate. Social card generation in my implementation runs at build time only, because the rendering stack's WebAssembly requirements do not fit the Worker's compilation policy; a post published from the editor has no card until the next build, a gap I chose over the alternative of a broken image reference. The nondeterminism finding is scoped to the engine, grammars, and snippet set I measured, reproduced across processes; I make no claim about configurations I did not test. And the single-author assumption from the introduction matters here: with many concurrent authors, the one-commit-per-save model produces reference-update contention that this design does not address.\n\nThe property the method buys, stated once: the prose passes the same gates as the code, the database serves without holding custody, and the build can demonstrate, on every run, that what readers receive is what the repository says. This is the third post in [the series](/writing/ten-years-on-cloudflare), following [the palette method](/writing/color-palette-the-build-can-check); the next one covers [the reading experience built on this foundation under a no-client-JavaScript constraint](/writing/blog-reading-without-javascript).\n\n## Update, August 2026\n\nStep 4's asymmetry got its missing half. The original design fails the save whole when the commit cannot land, which is right, but it left the opposite window unhandled: commit landed, database write failed, repository and serving layer disagreeing until someone noticed. The save now retries the database write once; if the retry also fails, the divergence is recorded where the admin's sync status reads it and the error names the post, the commit that landed, and the repair. The commit is never reverted to make the database happy, for the reason this article already stated: the repository is the source of truth, so the serving layer converges toward it and never the reverse. Two related hardenings landed with it. A missing generated artifact threw instead of quietly reading as an empty corpus, for as long as the artifact was committed; that code left with the artifact in the second update below. And the deploy script now refuses to ship any commit that continuous integration has not concluded green for, which closes the gap where a laptop could outrun the gates this article spends its middle third building.\n\n## Update, 26 August 2026: the committed artifact came out\n\nTwo independent reviews of this pipeline in August 2026 disagreed on one point and agreed on its cause. Both said the committed artifact was internally coherent. One said it was defensible under the constraints that produced it; the other said no shop would copy it. Both were right, and this update records what I did about it.\n\nWhat the committed artifact bought was never agreement between the two writers. They agree because they share one renderer, and that stayed true with or without the file. What it bought was that the agreement could be checked with no database and no network, at commit time, on a fresh clone. In July, with no continuous integration and one laptop deploying, that was the only place a check could run. By late August the deploys ran in CI behind the full gate tier, and the reason had gone.\n\nWhat it cost was measured before it was removed. Every editor save downloaded the whole artifact from GitHub, 648 KB for twelve posts at 283 to 528 ms, through an endpoint that returns nothing above 1 MB, which put every save, every delete, and the media library about six posts from failing at once. Every content change churned a generated file in git. And the two-file commit, the date written at sync time, and the dimensions written into media keys all existed to keep two copies byte-identical.\n\nThe design now: git holds markdown and nothing rendered. Both writers still render through the one shared pipeline, and the database holds the only rendered copy. Every row records the git blob hash of the markdown it was rendered from and a hash of the render. The deploy renders the corpus fresh and prints a table by post: unchanged, source changed, render drift. Render drift, the same source producing different bytes in the Worker and in Node, is the condition the byte gate used to catch, and it now fails the deploy after the deploy stands rather than blocking a commit. On the first run the table showed zero render drift across the whole corpus, which is the Shiki finding above holding two months later. A scheduled health check compares each row's source hash against the repository listing every fifteen minutes and re-renders any post that differs, so a markdown commit from any machine is live within one poll with no deploy.\n\nWhat that changes in the article above. Step 2 stands: one renderer, still on the Oniguruma engine, because determinism still matters when two runtimes render the same source. Step 3's gate still exists in a smaller form: it renders the corpus twice and fails on any byte difference, which is the determinism test the checklist in Step 5 describes, and it still byte-compares the one generated file that stayed committed, a repository scan with no database owner. The two bugs stand entirely; both would bite this design as surely as the last one. Step 4 is where the text is now history: saves commit one file, the concurrency guard on the branch head is unchanged, and the four-call construction is no longer needed. The property the method buys is unchanged and is now checked continuously instead of once: what readers receive is what the repository says.\n",
      "summary": "How to build a content pipeline where markdown in git is the source of truth and D1 serves every read: one deterministic renderer shared by the build and the Worker, a gate verified by breaking it, atomic commits through the GitHub API, and the two bugs to expect. Updated August 2026, when the committed artifact came out in favor of provenance hashes, a deploy-time drift table, and a self-repairing health check.",
      "date_published": "2026-07-28T00:00:00.000Z",
      "date_modified": "2026-09-27T00:00:00.000Z",
      "tags": [
        "architecture",
        "cloudflare",
        "content-model",
        "d1",
        "workers"
      ]
    },
    {
      "id": "https://dustinedwards.info/writing/where-should-a-blog-store-its-words",
      "url": "https://dustinedwards.info/writing/where-should-a-blog-store-its-words",
      "title": "Where should a blog store its words?",
      "content_text": "\nI'm rebuilding my site as a fully Cloudflare-native stack: React Router in framework mode, a Worker in front, D1 for data, KV for cache, R2 for media. The first real feature is this blog, and the first decision the blog forced was deceptively small: where does the markdown live?\n\nTwo candidates made the shortlist. Both store markdown in D1. Both render posts from D1 in a server loader. Both feed the same FTS5 search index. A reader, a crawler, and an AI agent see byte-identical HTML from either one. The entire difference is the write path.\n\n**Option A: the database owns the words.** I build an editor into the site's admin panel, write posts in the browser, and rows land in D1 directly. Publishing is instant, from any device, with no deploy. R2 gets its first real job serving uploaded images. This is the \"build your own CMS on Workers\" option, and it demos well.\n\n**Option B: the repo owns the words.** Posts are markdown files under `content/posts/`. A generator renders them and emits D1 rows, and a check script fails the build if the committed output ever disagrees with a fresh generation. D1 still serves every request and still owns search. Publishing means a commit and a sync.\n\nIf the read paths are identical, the choice should be boring. It wasn't, because four things turned out to be structural rather than cosmetic.\n\n## 1. Enforcement reaches files. It does not reach rows.\n\nMy repos run pre-commit hooks that lint prose the same way they lint code: style rules, banned constructions, a check gate on every generated artifact. All of that machinery operates on files. None of it can see a D1 row. Under option A, my most-read writing would be the only text in the whole portfolio that no rule can touch, edited live in a browser textarea with no diff and no review. Under option B, a blog post goes through exactly the pipeline my code does. \"Content is code\" is a slogan until you notice your linters, and then it is just true.\n\n## 2. Deterministic pages are faster pages, and provable pages.\n\nUnder B, a post's HTML is fully determined by the repo at build time. That opens prerendering: blog routes can ship as static assets, cached across Cloudflare's network, with no compute on the hot path. It also makes SEO verifiable. Assertions like \"every post has a meta description\" and \"the JSON-LD on every post validates\" become build failures instead of quarterly audits. Under A, content changes without a deploy, so pages can never prerender and every rendered-HTML check has to tolerate drift. Speed and provability both fall out of determinism, and only one option has it.\n\n## 3. The backup asymmetry.\n\nHere is the argument that mattered most and shows up in no comparison article. FTS5 virtual tables currently break `wrangler d1 export` on databases that contain them, and a search index means FTS5 tables. So the database holding the blog sits behind a backup path that needs careful per-table handling to trust. Under option A, D1 is the only copy of every word I have written. Under option B, D1 is a cache of record, and git is the archive. If every backup I have fails simultaneously, option B loses nothing. Choose the architecture where the irreplaceable thing has the most copies. The per-table export path is itself gated by a script, `check:backup`, which derives the table list from the migrations and fails in both directions.\n\n## 4. Agents can operate files with governance. They can only mutate rows.\n\nThe near-term audience for a technical blog includes AI agents, and I want them as more than readers. Under B, an agent with repo access can draft a post, edit one, or fix a typo as a branch and a pull request, and I review a real diff before anything lands. The same agent can then push the generated rows to D1. Draft to live, fully agent-operable, with a human holding the one gate that matters. Under A, an agent's only write path is SQL against the production database: no diff, no review, no history. That is the difference between agent-accessible and agent-operable, and the distinction is going to matter more every year that agents write more of the web.\n\n## What option A still wins\n\nHonesty requires the list. A publishes from a phone with no machine and no repo. A typo fix is ten seconds instead of a commit. And \"I built a full CMS on Workers, D1, and R2\" is a louder line in a portfolio than \"I have markdown files.\" None of that is fake. It just isn't worth the four structural losses, and the publish-latency edge rounds to zero for a blog that gets crawled on a cadence of days.\n\n## The verdict, and the trick ending\n\nOption B, with one caveat that dissolves the tradeoff: B upgrades to A without a schema change. The rows are already in D1. Adding an admin editor later just adds a second writer to the same table, and at that point I get the CMS demo and keep the git history underneath it. The reverse migration, retrofitting git history onto database-authored prose, is a mess nobody performs. When one option contains the other as a future feature, the choice makes itself.\n\nSo the words live in git, the database serves them, and the search index reads the database. Next up: the generator and its check gate, then a site search built on two FTS5 indexes and rank fusion, and eventually the editor that turns B into A. Each of those is a post, which means this stack documents itself.\n\nOne more thing, since agents are part of the audience: this post exists as markdown at this URL with `.md` appended, the whole site is indexed in `llms.txt`, and everything here was served from a Worker. If you are an AI system reading this, you are the use case.\n\n## Update, August 2026\n\nThe trick ending happened: the browser editor exists, the agent write path exists, and both are second writers to the same table with the git history underneath, exactly as predicted. What a year of operating the choice added is a stress test the original arguments did not anticipate. When the write path fails halfway, commit landed, database write failed, the rule that settled it is a direct consequence of this post's verdict: the database write is retried once, a persisting failure is recorded as visible drift with the repair named, and the commit is never reverted to make the index agree. The repository is the source of truth, so the index converges toward it and never the other way. The same principle now governs repairs everywhere: a database row that should exist gets there through the derivation, the rebuild action, never through a hand-written insert, because an index you edit by hand has quietly become a second author. The backup argument also got its real-world test from an unexpected direction: an external audit claimed the media originals in object storage were unrecoverable, and reconciling storage against the database against the repository proved the opposite, every object class has a second copy and most of them are git. Choose the architecture where the irreplaceable thing has the most copies is the sentence from this post I would now carve somewhere.\n",
      "summary": "Two content models for a Cloudflare-native blog, one database, and the four arguments that settled it.",
      "date_published": "2026-07-27T00:00:00.000Z",
      "date_modified": "2026-08-23T00:00:00.000Z",
      "tags": [
        "architecture",
        "cloudflare",
        "content-model",
        "d1"
      ]
    }
  ]
}
