内容编排流程中的提示注入
Prompt injection in an authoring pipeline
在内容编排流程里,粘贴素材的人正是被攻击的人。哪些风险由类型化输出和确定性校验消除,哪些由提示词处理,哪些只能靠人工复核发现。
设计笔记
An authoring pipeline is an unusual place for prompt injection, because the person pasting the material is the one being attacked.
The pasted text may come from a partner's document, a public article, or a student's records, and any of those can contain sentences addressed to the model. Photo descriptions are written by a vision model looking at pictures that may contain printed instructions.
The blast radius is small by construction
- the model has no tools and no side effects; it returns text
- the output is typed JSON validated by deterministic code: block kinds outside the family's list are refused, fields are bounded, photos are referenced by number and resolved on the server
- identity, lifecycle, slug and photo assets never come from the model
- the renderer escapes everything; injected markup prints as characters
- a human reads the real page before publishing
What the prompt does
Rules live in the cached system prompt. Source material, photo descriptions and the current document are wrapped in labelled markers with an instruction that anything inside them is content to typeset, never instructions to follow. The system prompt says facts come only from the source material, and the composer is asked to mention in its commentary if the material contained instructions.
What remains
Subtle content manipulation that produces a valid, plausible page is not detectable by code. Two things help. A per-turn change summary shows what the model altered, so the reviewer reads a list rather than the whole page. A source-fidelity check flags dates and quoted spans in the output that do not appear in the pasted material. Both are warnings, not blocks, because material may legitimately arrive in more than one paste.