Operations Runbook¶
Applies to: CPersona 2.5.x. This page is the canonical operations reference: backup, degradation detection, recall tuning, CJK guidance, and corpus indexing patterns. The behaviour facts it relies on are contracts — see Behavior Contracts.
Backup and restore¶
The database is a single SQLite file (CPERSONA_DB_PATH) running in WAL mode.
A plain cp of a live WAL database can straddle a checkpoint and produce a
corrupt copy, so do not script one.
Recommended, in order:
- Online physical backup (first choice — safe while the server runs):
sqlite3 "$CPERSONA_DB_PATH" ".backup 'cpersona-backup.db'"
# or
sqlite3 "$CPERSONA_DB_PATH" "VACUUM INTO 'cpersona-backup.db'"
Both produce a consistent snapshot under concurrent writes.
-
Offline copy: stop the server, then copy the
.dbtogether with its-waland-shmsiblings. -
Logical backup (recommended as a low-frequency complement):
export_memorieswrites JSONL that is independent of the schema version, andimport_memoriesis idempotent (duplicates are skipped) — so restore drills are safe to rehearse. Monthly is a reasonable cadence.
The .db is not the whole instance. State lives outside it that none of
the backup forms above touch:
-
<CPERSONA_DB_PATH>.calibration.json— per-agent vector thresholds, gate state, scoring version, and the per-agentset_recall_precisionbetas.Restore without it and the measurements repair themselves. With the default
CPERSONA_CALIBRATE_ON_MODEL_CHANGE=true, a startup with no sidecar recalibrates rather than falling back, anddeep_checkreportsnever_calibratedif it did not.The betas do not repair themselves. They are policy inputs, not measurements, and nothing re-derives a preference you stated. An operator who tuned precision loses that tuning without being told, and the recalibration measures every gate at the default beta. Copy the sidecar beside the database, or re-apply
set_recall_precisionafter a restore. -~/.cpersona/operating-context.toml(orCPERSONA_OPERATING_CONTEXT_PATH) — the operator instructions served to every client. - The ACL file, ifCPERSONA_ACL_FILEis set. Restoring without it changes who may connect. - The external vector index, underCPERSONA_VECTOR_SEARCH_MODE=remotewithCPERSONA_STORE_BLOB=false. In that configuration the embeddings exist only in the remote index, so restoring the database and all three files above still comes back with no vector arm.check_healthreports it asno_local_vector_fallback. Back the index up on the same schedule, or keepSTORE_BLOB=trueso the.dbstays self-sufficient. - Not<CPERSONA_DB_PATH>.memories.vecindexor.episodes.vecindex, the contiguous vector indexes. They are derived from the database, and are rebuilt rather than restored. A backup that includes one wastes space, and a restore that brings back a stale one is corrected by the next build. See the contiguous vector index.
Keep the live database outside cloud-sync folders (Dropbox, Drive, and so on). Sync clients interact badly with SQLite WAL files. Sync the backups instead.
Detecting a dead embedding server¶
Vector search is the strongest retrieval layer, and it degrades without announcing itself at the network boundary. Three detection surfaces exist, and the calling agent should be instructed to watch the first.
-
advisoryon recall responses (v2.4.33+, the primary surface). A state machine observes real embedding failures, and degraded recalls carryadvisory = {degraded, severity, reason, evidence, runbook, advisory_scope}— "the vector layer is down; serving FTS + keyword only".Instruct your agent to surface this field to the user and follow its
runbook(usually: start or repoint the embedding server, then recall again). A shortened runbook means "you were already told", andadvisory_scopesays who "you" was:sessionwhen the process is yours alone,processon the HTTP transport, where the state is the whole server's. On a shared server a fault repeats its full runbook rather than assume your session saw it (bug-251).CPERSONA_DEGRADED_ADVISORY=falsesilences the advisory. Set it to record that the operator accepts running without an embedding backend, which is a supported fallback and not a recommended configuration. Design: DEGRADED_ADVISORY_DESIGN. 2.embeddedon store responses. Every write reports whether its embedding was persisted. Writes made while the server is down come backembedded: false, and those rows are repairable (next item). 3.check_health(agent_id, fix=true)detects rows stored with NULL embeddings and re-embeds them once the server is back. The dimension check still skips rather than fails when the endpoint is unreachable, but the endpoint itself is now watched:embedding_backendreportswarn, with the failing call's own evidence, when a configured backend does not answer. 4. Readnot_probedfor what it is.check_health(fix=false)makes no network call, so it cannot test liveness. It says so —embedding_backendreturnsnot_probed— rather than leaving you to infer health from silence.A
fix=truerun returns it too when the probe was answered from the embedding client's five-minute embed cache: the dimension comes back without a request reaching the endpoint, so nothing was learned about it either way. The finding'sreasonsays which of the two happened. A fault a recall already latched is still reported there.With no backend configured this check is quiet by design. That is a supported configuration, and the
advisorysurface is where it is raised.
Tuning recall¶
The knobs, in the order you should reach for them:
set_recall_precision(agent_id, precision)— the main policy knob, and under the default fusion modes effectively the only one. It moves the operating point β of the fused quality gate:strict=2.0(fewer contaminants, more misses),balanced=1.0(default),lenient=0.5(fewer misses, more contaminants). It takes effect immediately, with no restart.calibrate_threshold(agent_id)— re-derives the gate positions from the corpus, with no labels needed. Called with anagent_id(and the fused gate enabled), it calibrates both the vector threshold and the fused gate; called without one, it calibrates the vector threshold only. The value it prints is the vector-side threshold, while under fusion modes the component actually cutting results is the fused gate. Re-run it after the corpus changes substantially, after a bulk rebuild, and after changing the embedding model.CPERSONA_AUTOCUT_MIN_RESULTS— autocut fires on similarity-scale signals: under confidence scoring, or on a homogeneous raw-cosine list, which is whatcascadeproduces with confidence on or off. It is deliberately inert underrsfandrrf(contract §6), so under the default configuration this knob does nothing. Note that it is the fusion mode that decides that, not the confidence flag.CPERSONA_FUSED_GATE_ENABLED=false— last resort. Without the gate, filtering falls back to the pool-size heuristic (_adaptive_min_score), which still rejects weak matches: it is a coarser gate, not an open door. What you give up is the operating point measured for this corpus.
Choosing a precision. The trade is asymmetric per use case. For
index-style corpora, where recall feeds an AI that can discard irrelevant
rows, a miss is worse than a contaminant: run lenient for a few days and
step back to balanced only if contamination becomes a real cost. For
contamination-sensitive contexts, where recall is injected into a small
window, prefer balanced or strict.
When not to rely on recall¶
Probabilistic retrieval can lose. Facts whose absence from context causes
harm — the current top-priority decision, standing safety rules — belong in a
deterministically injected surface: CLAUDE.md, a system prompt, or an
index file your agent always loads. They do not belong primarily in memory.
A useful split: what should fire without being asked goes in deterministic injection; what should be findable when asked goes in memory. Two corollaries follow.
- Update decisions by overwrite, not append. When a decision changes,
update_memorythe old row (it re-embeds automatically). The most reliable way to stop a stale decision from winning recall is for it not to exist in the search space. lock_memoryprotects against loss, not against losing a ranking (contract §9).
Japanese and CJK corpora¶
- Set
CPERSONA_RECALL_MODE=rsf. FTS5 tokenizes CJK poorly, and rsf keeps the keyword channel's score magnitude in the merge, which is the signal that compensates. No further CJK-specific settings exist. - Known characteristics of
jina-v5-nano— the model this project runs, and the one the reference embedding server downloads by default — on Japanese, confirmed by long-term production use: strong when at least one proper-noun or identifier anchor overlaps between query and memory; weak on pure concept matches with no shared vocabulary. Phrasing queries with a concrete anchor term ("keyword anchoring") is a correct adaptation to the current model, not a workaround to feel bad about. - The model slot is replaceable (
CEmbeddingprovider). If you swap the embedding model, re-embed the corpus in full and then runcalibrate_threshold. The server auto-recalibrates on a dimension change, but a same-dimension model swap needs the manual recalibration.
Corpus indexing and sync patterns¶
CPersona's primary design center is memory that accrues from conversation. Using it as a search index over canonical Markdown files works, but two facts shape the correct pattern. CPersona is a passive server — it watches no files, and ingestion is always caller-driven — and dedup is skip, not upsert (contract §5).
Pattern A — rebuild the index as a disposable projection (recommended first).
- Give the index a dedicated
agent_id(for examplemd-index). Do not mix index chunks into a conversational agent's memory. A separate agent also keeps the episode boundary machinery out of the picture. - When the source documents change:
delete_agent_data(agent_id="md-index"), then re-storeevery chunk. - Re-run
calibrate_thresholdafter each rebuild, and re-applyset_recall_precisionif you use it.delete_agent_datadiscards that agent's calibration state along with its rows. If the rebuild is a nightly batch, make the recalibration part of the batch. -
Cost: re-embedding a few thousand chunks on CPU takes minutes to tens of minutes, depending on the environment. What you buy is no diff logic, and an index that matches the source once the rebuild finishes.
It does not match during one.
delete_agent_datacommits before the re-storebegins, so a recall issued inside that window sees a partial index or none at all. Run rebuilds when nothing is querying, or build under a secondagent_idand switch readers over when it is complete.
Pattern B — differential updates with an external content-hash ledger.
storeeach chunk with a stablemsg_id(for examplepath#heading) and the source document insource.id. Thesource_idrecall argument (prefix-match) then doubles as a per-document filter.- Keep a caller-side ledger of
key → content hash, and process only chunks whose hash changed. - Changed chunks must go through
update_memory, ordelete_memoryplusstore. Re-storing changed content under the samemsg_idis skipped without a word. Unchanged chunks may be re-submitted blindly, because exact-match dedup guarantees that is harmless.
Prefer A until the corpus is large enough that rebuild time actually hurts. It has no ledger to corrupt and no drift mode.
Scale¶
Growth needs no archival or thinning routine. The vector layer scans a recency
window (CPERSONA_MAX_MEMORIES, default 10,000 — raise it via env for large
corpora; contract §4).
The FTS and keyword channels reach the full history, and old rows sink through
windows and decay rather than being deleted.
The contiguous vector index¶
The local vector scan reads every embedding in its window out of SQLite one row at a time, and that read — not the arithmetic — is where the scan's time goes. The contiguous index is a file beside the database holding the same embeddings in the layout the arithmetic wants, so the scan reads them in one pass. Same rows, same scores, same order: only the latency changes. On the reference machine at 100,000 rows, the vector arm went from 604 ms to 77 ms.
There is one index per table. Memories and episodes are scanned separately, so each has its own file, and each is built and checked on its own. An unindexed episode table costs more per query than an indexed memory table five times its size, so a deployment that builds one should build both.
It is a derived artifact. The database is the only source of truth. The index is a projection of it and is treated the way a cache is treated: not backed up, never repaired, safe to delete at any time. If anything about it looks wrong, delete the file and build again.
Building it is an operator action. Nothing in the running server builds or refreshes the index. The entry point is a command, made for a cron job or a systemd timer:
python -m cpersona.vector_index --db "$CPERSONA_DB_PATH" build
python -m cpersona.vector_index --db "$CPERSONA_DB_PATH" --table episodes build
python -m cpersona.vector_index --db "$CPERSONA_DB_PATH" status
python -m cpersona.vector_index --db "$CPERSONA_DB_PATH" --table episodes status
--table selects the file (memories is the default).
build reads the database — it never writes to it — and replaces the index
atomically. The server picks the new file up on its next query, with no
restart. It exits 0 when built, and 1 when it declined, printing why: a corpus
with no embedded rows, or one that carries two embedding widths at once (the
transient state of a model change; build again when the re-embedding has
finished).
status reports whether an index is present and usable, how many rows it
holds, and how many rows have been written since it was built. It exits 1 when
there is no index, and 2 when the file exists but cannot be used. Add --json
to either for a machine-readable line.
What happens between builds. The index knows the highest row id that existed when it was built. Rows written after that are not in it, and are not lost: every query reads them from the database exactly as the scan always did, and merges them with the indexed rows.
So a late rebuild never changes an answer. It only grows the part of each query
that is still read row by row, until latency drifts back toward the unindexed
figure. Rebuild frequency is a performance setting, not a correctness one.
Nightly is a reasonable default, hourly for a corpus that grows fast, and
status shows how far behind the index is at any moment.
When the index is not used. The scan the index replaces is still there, and is the fallback. Recall falls back to it — slower, and still correct — when the file is missing, when it fails its own integrity check, when the query vector's width does not match the index, and when a row the index holds has since lost its embedding under a maintenance repair.
Each of these is reported where an operator reads health. check_health raises
vector_index_absent (with the count of rows it would cover) when there is no
index, and vector_index_unusable when the file cannot be trusted — one line
per table, each naming its table — and the server logs a warning on every
query it had to hand back to the scan.
An index that had been unusable for a week without anyone noticing would otherwise read as "somehow not faster", which is the failure the reporting exists to prevent.
There is no configuration for the index. Its path is derived from
CPERSONA_DB_PATH, and the scan window it serves is the same
CPERSONA_MAX_MEMORIES the scan uses.
Maintenance cadence¶
- Monthly:
check_health(agent_id, fix=true)for deterministic integrity checks and auto-repair,deep_check(agent_id, fix=true)for the semantic pass, and a logicalexport_memoriesbackup. - After corpus upheaval (bulk import, rebuild, model change):
calibrate_threshold(agent_id). - On a timer, if the vector index is in use:
python -m cpersona.vector_index build, and the same with--table episodes. Nightly by default; see the contiguous vector index for what a late rebuild costs (latency, never correctness). - Version upgrades: the schema migrates itself forward. Calibration is fingerprinted to the scoring function and embedding dimension, and the server recalibrates at first boot when either changes (v2.5.2+).