Getting Started¶
Applies to: CPersona 2.5.x. This page is the canonical installation and setup reference. The README keeps a condensed version of the same steps because it is also the PyPI project page; when the two disagree, this page wins.
CPersona is an MCP server. You install it,
point an MCP client at it, and the client's agent gains store / recall
tools that survive across sessions. Nothing else in your stack changes.
Prerequisites¶
- Python 3.11+
- uv for the one-command path (optional —
pipworks too) - An MCP client: Claude Desktop, Claude Code, Codex CLI, Cursor, VS Code, or any other MCP host — step 3 has the entry for each
Let the agent do it (Claude Code)¶
The repository and the published wheel both ship an Agent Skill It walks Claude Code through the installation, and then teaches it when to store, recall, and archive. Installing the skill is the shortest path:
# Installed from PyPI? The skill ships inside the wheel — no clone needed:
python -c "import cpersona,pathlib,shutil; s=pathlib.Path(cpersona.__file__).parent/'skills'/'cpersona-memory'; shutil.copytree(s, pathlib.Path.home()/'.claude/skills/cpersona-memory', dirs_exist_ok=True)"
# Running via uvx (isolated environment), or not installed yet:
git clone --depth 1 https://github.com/Cloto-dev/cpersona.git /tmp/cpersona
mkdir -p ~/.claude/skills && cp -r /tmp/cpersona/skills/cpersona-memory ~/.claude/skills/
Then tell Claude Code: "Set up CPersona — I want persistent memory."
The manual steps below are for every other client, and for anyone who prefers to configure things by hand. Step 5 is the part the skill would otherwise do: it makes the memory triggers fire without being asked.
1. Install CPersona¶
uvx cpersona # run directly, no install step
# or
pip install cpersona # then the `cpersona` command is on your PATH
From source (for development)
git clone https://github.com/Cloto-dev/cpersona.git
cd cpersona
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install .
In a container
git clone https://github.com/Cloto-dev/cpersona.git
cd cpersona
docker build -t cpersona .
docker volume create cpersona-data
docker run -d --name cpersona -p 8402:8402 \
-e CPERSONA_AUTH_TOKEN="$(openssl rand -hex 32)" \
-v cpersona-data:/data cpersona
In containers, with an embedding server
`compose.yaml` in this repository runs CPersona and CEmbedding together, wired to each other, so vector search works without a second setup:export CPERSONA_AUTH_TOKEN=$(openssl rand -hex 32)
docker compose run --rm embedding cembedding-download-model --model jina-v5-nano
docker compose up -d
At startup the server checks pypi.org for a newer release. It reports one
through recall and check_health, and check_update answers on demand. Set
CPERSONA_UPDATE_CHECK=false to turn the check off. Updating is never
automatic: it takes an explicit check_update(apply=true) and a restart.
2. Set up an embedding server (recommended)¶
Connect a compatible embedding server. Running CPersona without one is supported as a fallback, but is not recommended for normal operation.
Vector search is the strongest of the three retrieval layers, and the only one that needs an external process. Without it CPersona still runs on FTS5 + keyword search, and says so on every recall.
CEmbedding is the reference backend. Any other embedding server that satisfies the contract below is supported just as well. The choice is yours.
The contract¶
CPersona is embedding-server-agnostic. Point CPERSONA_EMBEDDING_URL at any
HTTP endpoint that implements this:
POST /embed
Request: { "texts": ["string", ...] } # non-empty array
Response: { "embeddings": [[float, ...], ...], "dimensions": <int> }
CPersona reads embeddings and nothing else. dimensions comes from the
reference server and is ignored by the client, so a backend that omits it
still works. CPersona sends at most 32 texts per request, and the
reference server accepts up to 100. Batch limits in that range are nothing you
need to plan around.
Three requirements are easy to miss. Each one degrades ranking silently:
- Embeddings MUST be L2-normalized. CPersona computes similarity as a raw
dot product, so a backend returning unnormalized vectors biases ranking by
vector magnitude. Every supported backend (the client's
apimode and all CEmbedding providers) already normalizes. - The contract is role-less. Queries and documents go through the same call, with no instruction prefix. Prompt-prefix models (e5-style, prompted bge) underperform here. Symmetric or retrieval-merged models (jina-v5-nano, bge-m3, MiniLM) are the intended fit.
- Swapping models behind one URL invalidates the corpus. The contract carries no model identity, so CPersona fingerprints the backend by embedding dimension alone. A swap to a different model of the same dimension is undetectable.
The repair tools cannot reach it either. check_health(fix=true) re-embeds
rows whose blob is NULL, and the dimension check only NULLs blobs of the
wrong length. After a same-dimension swap every blob is the expected size,
so nothing is NULLed and nothing is re-embedded. No tool force-re-embeds a
row that already has a blob.
To recover, rebuild the corpus: delete_agent_data, then re-store as in
the rebuild pattern, then
run calibrate_threshold.
The reference server¶
CEmbedding (MIT) runs jina-v5-nano on-device (CPU) and exposes exactly this endpoint:
# Download the model into ./data/models
uvx --from "cembedding[onnx]" cembedding-download-model --model jina-v5-nano
# Run the server (it reads ./data/models from the current directory)
EMBEDDING_PROVIDER=onnx_jina_v5_nano uvx --from "cembedding[onnx]" cembedding
Or put it on your PATH with pip install "cembedding[onnx]" and run
cembedding-download-model --model jina-v5-nano, then cembedding. From a
source checkout the same two steps are python -m cembedding.download_model
--model jina-v5-nano and python -m cembedding.
Either way you should see
HTTP embedding endpoint started on http://127.0.0.1:8401/embed. Verify it
before wiring CPersona to it:
curl -s http://127.0.0.1:8401/embed \
-H 'content-type: application/json' \
-d '{"texts":["hello world"]}' | head -c 200
CPersona's defaults are tuned against jina-v5-nano (768 dimensions). Any other
server satisfying the contract works. Models with published measurements are
listed in benchmarks/.
CPersona only needs the URL. But how you supervise the reference server matters, and the obvious way does not work.
It is an MCP server, not a plain HTTP process. On its default transport it
runs an MCP session on stdio in the foreground and serves the REST /embed
endpoint from a background task. Its lifetime is therefore bound to stdin:
at EOF the session ends and the finally clause cancels the HTTP task.
A service manager starts things with stdin on /dev/null. Started that way,
the server binds the port, logs HTTP embedding endpoint started, and exits
in the same second with status 0. The supervisor sees a clean exit, and
CPersona is left pointed at a URL nothing answers.
Give it a stdin that stays open. Under a service manager, run it through
something that holds the pipe: ExecStart=/bin/sh -c 'sleep infinity |
cembedding'. In a terminal, the terminal already does this.
EMBEDDING_TRANSPORT=streamable-http does not read stdin, so it supervises
cleanly. But it serves the MCP endpoint instead of REST /embed, so it is
not an option for CPersona's http mode, which posts to /embed.
3. Register CPersona with your MCP client¶
Every client launches the same process: command uvx, argument cpersona,
and the environment shown below. They differ only in which file they read and
what they call the keys.
Pick an absolute CPERSONA_DB_PATH first. The examples use
/home/you/.claude/cpersona.db because that directory already exists for
Claude users, but any absolute path works.
| Client | Where the entry goes | Shape | Checked here |
|---|---|---|---|
| Claude Code | claude mcp add-json … -s user (writes the user config) |
JSON, type: stdio |
yes |
| Claude Desktop | claude_desktop_config.json |
JSON, mcpServers |
yes |
| Codex CLI | codex mcp add … -- uvx cpersona (writes ~/.codex/config.toml) |
TOML, [mcp_servers.cpersona] |
yes (codex-cli 0.147.0) |
| Cursor | ~/.cursor/mcp.json (global) or .cursor/mcp.json (project) |
JSON, mcpServers |
vendor docs |
| VS Code (Copilot) | .vscode/mcp.json (workspace) or the user mcp.json |
JSON, servers |
vendor docs |
| Any other MCP host | its stdio server configuration | the same command / args / env triple | vendor docs |
"Checked here" means the entry was written by that client's own tooling and read back on a maintainer machine while this page was written. "Vendor docs" means the shape was transcribed from the client's documentation and has not been executed here. If it disagrees with what your client accepts, the client is right — please report it.
Claude Desktop — add to claude_desktop_config.json:
{
"mcpServers": {
"cpersona": {
"command": "uvx",
"args": ["cpersona"],
"env": {
"CPERSONA_DB_PATH": "/home/you/.claude/cpersona.db",
"EMBEDDING_MODE": "http",
"EMBEDDING_HTTP_URL": "http://127.0.0.1:8401/embed"
}
}
}
}
Claude Code — one command:
claude mcp add-json cpersona '{"type":"stdio","command":"uvx","args":["cpersona"],"env":{"CPERSONA_DB_PATH":"/home/you/.claude/cpersona.db","EMBEDDING_MODE":"http","EMBEDDING_HTTP_URL":"http://127.0.0.1:8401/embed"}}' -s user
Codex CLI — one command; it writes the TOML shown after it into
~/.codex/config.toml:
codex mcp add cpersona --env CPERSONA_DB_PATH=/home/you/.claude/cpersona.db --env EMBEDDING_MODE=http --env EMBEDDING_HTTP_URL=http://127.0.0.1:8401/embed -- uvx cpersona
[mcp_servers.cpersona]
command = "uvx"
args = ["cpersona"]
[mcp_servers.cpersona.env]
CPERSONA_DB_PATH = "/home/you/.claude/cpersona.db"
EMBEDDING_HTTP_URL = "http://127.0.0.1:8401/embed"
EMBEDDING_MODE = "http"
Codex can also deny-list tools per server (disabled_tools = ["delete_memory",
…] under the same table). That is a client-side way to give an agent read-
mostly access. The server-side equivalent, enforced for every client at once,
is the
per-client capability layer.
Cursor — ~/.cursor/mcp.json, or .cursor/mcp.json inside a project. Same
shape as Claude Desktop:
{
"mcpServers": {
"cpersona": {
"command": "uvx",
"args": ["cpersona"],
"env": {
"CPERSONA_DB_PATH": "/home/you/.claude/cpersona.db",
"EMBEDDING_MODE": "http",
"EMBEDDING_HTTP_URL": "http://127.0.0.1:8401/embed"
}
}
}
}
VS Code (Copilot) — .vscode/mcp.json in the workspace, or the user-level
mcp.json (MCP: Open User Configuration). The top-level key is servers,
not mcpServers:
{
"servers": {
"cpersona": {
"command": "uvx",
"args": ["cpersona"],
"env": {
"CPERSONA_DB_PATH": "/home/you/.claude/cpersona.db",
"EMBEDDING_MODE": "http",
"EMBEDDING_HTTP_URL": "http://127.0.0.1:8401/embed"
}
}
}
}
Three things account for most setup problems:
- Set
CPERSONA_DB_PATHto an absolute path. The default,data/cpersona.db, is relative to the client's working directory, so a client launched from somewhere else opens a different, empty database. On Windows, write the path asC:/Users/you/.claude/cpersona.db. - No embedding server yet? Drop the two
EMBEDDING_*lines (or setEMBEDDING_MODE=none). CPersona runs on FTS5 + keyword and reports that it is degraded. EMBEDDING_MODE/EMBEDDING_HTTP_URLare the generic aliases ofCPERSONA_EMBEDDING_MODE/CPERSONA_EMBEDDING_URL. The prefixed form wins when both are set. The configuration reference covers the settings you are likely to reach for, but it is not exhaustive: a few variables (CPERSONA_STORE_BLOB,CPERSONA_FTS_ENABLED,CPERSONA_EMBEDDING_API_KEY, theCPERSONA_CALIBRATE_*pair and some others) are read by the server without appearing there.cpersona/config.pyis the complete list.
4. Verify it works¶
Ask the agent to store something, then recall it in a new session. Surviving the session boundary is the whole point:
"Store this: the deploy runbook lives in ops/deploy.md."
…then, in a fresh session: "What did I tell you about the deploy runbook?"
Two checks worth running once the corpus is real:
check_health— the registry-driven health check.statusis the verdict; issues are severity-tagged (critical/warn/info), andcheck_health(fix=true)repairs the mechanical ones.- Watch recall responses for an
advisoryfield. It means vector search is not contributing. The severity says why: ahintmeans embeddings are unconfigured (mode=none), and a fault means a configured endpoint stopped answering. See detecting a dead embedding server.
5. Make the memory triggers fire in every session¶
Registration gives the agent the tools. It does not make the agent use them unprompted. Recall at session start, store on a decision, archive at session end: those rules have to live in a file the client loads on every session, not in a skill that activates only when the conversation happens to match. Every client has such a file:
| Client | Always-loaded file (user-level default) | Project-level alternative |
|---|---|---|
| Claude Code / Claude Desktop | ~/.claude/CLAUDE.md |
./CLAUDE.md |
| Codex CLI | ~/.codex/AGENTS.md |
./AGENTS.md at the repository root |
| Cursor | User Rules (Customize → Rules — a setting, not a file) | .cursor/rules/cpersona.mdc with alwaysApply: true, or ./AGENTS.md |
| VS Code (Copilot) | see the client's custom-instructions documentation | .github/copilot-instructions.md, or ./AGENTS.md with chat.useAgentsMdFile enabled |
Paste the block below into that file. Replace <AGENT_ID> with one stable
identifier the agent will use on every call ("claude-code", "codex", …).
Keep the markers. A later version of the block uses them to find and replace
this one instead of stacking a second copy. On Claude Code the cpersona-
memory skill does the paste for you, with your approval; everywhere else you
paste it by hand.
The rules the block follows — consent, placement, idempotency, the 40-line budget, client neutrality — are the policy block standard. The block itself is maintained in the skill, and this copy is checked against it in CI.
The package also installs the paste as a command, for anyone who would rather not do it by hand — and for the case where the client refuses to let an agent write into a file it loads as instructions:
cpersona-policy --agent-id claude-code --install
# through uvx: uvx --from cpersona cpersona-policy --agent-id claude-code --install
It prints the block and writes nothing until --install is given, it picks the
user-level file of whichever client it finds (--client claude-code|codex or
--target PATH to say which file instead), and on a file that already holds a
block it replaces the span between the markers and leaves every other byte as it
was. Markers that do not form one well-formed block are reported rather than
repaired. --dry-run says what would change.
<!-- BEGIN cpersona-policy v2 (managed by the cpersona-memory skill; re-run the skill to update) -->
## CPersona memory policy
Use the CPersona MCP tools proactively with `agent_id="<AGENT_ID>"` — never wait to be asked.
**Session start** → `recall(agent_id, query="<opening-topic keywords or ''>", limit=10)` before
the first substantive action. Prefer `recall_with_context` when conversation history is already
at hand; add `deep=true` when the first pass comes back thin. Skip only for trivial one-shot
questions.
**Decisions, rules, preferences, bug findings** → `store` immediately. Fire on phrases like
"let's go with X", "from now on always Y", "remember that…", "approved", "that's a bug".
Protect must-never-lose rules with `lock_memory`. After a successful `git commit`, `store` a
one-line record: hash, what changed, why.
**Changing an existing rule** → `update_memory`, never delete + store. If the memory is locked:
`unlock_memory` → `update_memory` → `lock_memory`.
**Session end** — fire on closing phrases ("that's all for today", "wrap it up", "good night") →
first `store` + lock any unsaved decisions, then `archive_episode(agent_id, history=<the REAL
turns>, summary=…, keywords=…, resolved=…)`, computing `summary` and `keywords` yourself.
**"Don't save this" / benchmark sessions** → `pause_persistence(ttl_seconds=1800)`;
`resume_persistence()` (or TTL expiry) restores. Reads still answer, minus the writes inside them.
**Degraded mode** — if a `recall` response carries an `advisory` field, surface it to the user
and follow its runbook. Never quietly serve keyword-only recall.
**Quality** — if recall feels off, `set_recall_precision` (strict/balanced/lenient) is the one
policy knob; run `calibrate_threshold(agent_id)` after the corpus changes substantially.
Monthly: `check_health(agent_id, fix=true)`.
**If this client keeps a memory file that loads every session** (Claude Code's `MEMORY.md`), use it
as the deterministic index over this store: one line per memory — `- <slug> — <the sentence that
changes behaviour>` — with the body stored here under `message.id="memory-index:<slug>"` and content
starting `[<slug>]`, so a line tells you what to `recall`. Recall is ranked and may not surface a
memory; the index always arrives. Its size cap fails **silently** when exceeded, so consolidate at
80%, not at the limit. Never migrate existing memories into this store without asking first.
Details, setup, and troubleshooting: the `cpersona-memory` skill.
<!-- END cpersona-policy -->
Where to go next¶
| You want to… | Read |
|---|---|
| Know which behaviors you can rely on | Behavior Contracts |
| See what each tool does | Tools |
| Understand how retrieval works | Architecture |
| Back up, tune, or diagnose a live instance | Operations Runbook |
| Look up a setting | Configuration |
| Serve several clients over the network | Remote HTTP transport |