Block identity across AI regeneration

AI 重新生成时的区块身份

The composer returns a whole page every turn. Short per-page ids and a three-step reconciliation keep each block's identity — even if the model drops every id.

Design note

The composer returns a whole page every turn. That is the simplest contract with a model: no diffs, no patches, no partial updates that can go wrong.

But it raises a question. If the model rewrites a page, how does the site know that the third paragraph is still the third paragraph, so its edit history and its place in the change summary survive?

Short ids, not long ones

Every block carries an id scoped to its page: b1, b2, b3, from a counter that never reuses a number. A two or three character token is far less likely to be dropped or mangled in a regenerated document than a long random string, and it costs almost nothing in tokens. On revision turns the current document is sent with those ids, and the model is asked to keep the id on every block it retains and omit it on new ones.

Reconciliation in three steps

  • an id that exists in the base revision keeps its identity; first claim wins
  • a block with no id, or an unknown id, is matched to an unclaimed base block of the same kind whose text is identical or nearly so, using character trigram similarity so Chinese text works without tokenisation
  • everything else gets the next number

Identity therefore survives even if the model drops every id. The result is deterministic, so it is tested on fixtures: ids echoed, ids dropped, a block moved, a block rewritten, a duplicated id, an invented id.

What it buys

  • the in-situ editor addresses blocks by id, so an edit lands on the right block even after a regeneration
  • the change summary can say reworded b5 and moved b2 instead of everything changed
  • revisions can be compared block by block
  • embeddings can be computed for changed blocks only