在 Firestore 上实现搜索
Search on Firestore
Firestore 通过复合索引可以满足全部结构化查询,但完全不支持全文检索。AIBlog 在发布时派生一份精简的搜索投影,在每个服务进程中以 CJK 双字组分词建立索引,并借助按块生成的向量与原生向量检索来弥补跨语言检索的缺口。
搜索
Firestore answers every structural query a site needs through composite indexes: pages by family, by status, by date or sequence, slug lookups and slug history for redirects. It cannot do full-text search at all. AIBlog treats search as a derived layer on top.
The search projection
On every publish, the server writes one compact document per page: family, slug, date, tags, the title in every locale, the standfirst, and the body flattened to plain text through each block kind's text function, capped at about twenty kilobytes. Unpublish and archive delete it. The projection is derived data and can be rebuilt from revisions at any time.
The in-process index
Each server instance loads the projection once and builds an index in memory. The tokenizer emits character bigrams for Chinese, Japanese and Korean runs and lowercase words for everything else, so a Chinese query against a Chinese body and an English query against an English body both rank sensibly without a language service.
A single counter is bumped on every publish. Instances check it on a short interval and rebuild when it moves. That is the same discipline the rest of the system follows on Cloud Run: process-local caches are an optimisation, and correctness comes from the store.
Limits
The in-process index is comfortable to roughly ten thousand pages. Beyond that, or for typo tolerance, facets and analytics, a Typesense or Algolia adapter consumes the same projection. Firestore's official extensions can sync the projection collection to either.
Cross-lingual search
The bigram index cannot match an English query against a body written in Chinese. That gap is closed by embeddings: each block is a natural chunk with a stable id, vectors are computed on publish for changed blocks only, and Firestore's native vector search answers nearest-neighbour queries. The same vectors power related pages.