Add embedded Tantivy full-text content search
/api/search now finds files by CONTENT as well as by name: BM25-ranked
matches over extracted text (PDF, Office OOXML/ODF, plain text/code)
with typo-tolerant fuzzy terms and search-as-you-type prefix matching,
served from an embedded Tantivy index at {storage}/.search-index.
Pipeline (all off the request path, mirroring tree-etag + thumbnails):
- statement triggers on storage.files append to a durable dirty queue
(storage.search_index_dirty) - every write surface (REST, WebDAV,
NextCloud, WOPI, trash) is covered, crash-safe by construction
- ContentIndexWorker drains the queue on the maintenance pool, extracts
text once per unique BLAKE3 blob (storage.blob_extracted_text cache:
N copies = 1 extraction, renames/moves = 0 re-extraction) and applies
batched single-writer Tantivy commits; queue rows are deleted only
after the commit succeeds (at-least-once, idempotent upserts)
- the index is a derived artifact: a version-marker mismatch wipes and
reseeds it from Postgres, which remains the single source of truth
SearchService merges content hits into the existing name search: hits
are hydrated through ONE SQL round-trip that re-applies user scope,
trash state and every active filter (a stale index id can never leak),
scored below name matches, and returned with a plain-text snippet and
a match_source field. Index failure or
OXICLOUD_ENABLE_CONTENT_SEARCH=false degrades to name-only search; a
discard-only janitor keeps the trigger-fed queue bounded while disabled.
The frontend renders the snippet under the file name in list view.
New dependencies: tantivy 0.26, zip 8.6 (deflate only), pdf-extract 0.10.
https://claude.ai/code/session_01Sc7F4xbo83YbFAQ4xEeDrX
This commit is contained in:
@@ -0,0 +1,19 @@
|
||||
//! Embedded full-text content index (Tantivy) and its feeding pipeline.
|
||||
//!
|
||||
//! Three pieces, mirroring the thumbnail/tree-etag architecture:
|
||||
//!
|
||||
//! * [`tantivy_content_index`] — the embedded BM25 index over file names and
|
||||
//! extracted content. Lives on local disk (`{storage}/.search-index`),
|
||||
//! single-writer, microsecond queries. A DERIVED artifact: PostgreSQL is
|
||||
//! the source of truth and the index is rebuilt (reseeded) whenever its
|
||||
//! on-disk schema version differs from the binary's.
|
||||
//! * [`text_extractor`] — pure-Rust text extraction (plain text/code, PDF,
|
||||
//! Office OOXML/ODF). CPU-bound, runs only on the background worker.
|
||||
//! * [`content_index_worker`] — drains `storage.search_index_dirty` (fed by
|
||||
//! statement triggers on `storage.files`), extracts text once per unique
|
||||
//! blob (BLAKE3-keyed cache in `storage.blob_extracted_text`), and applies
|
||||
//! batched Tantivy mutations. Never touches a request path.
|
||||
|
||||
pub mod content_index_worker;
|
||||
pub mod tantivy_content_index;
|
||||
pub mod text_extractor;
|
||||
Reference in New Issue
Block a user