Skip to content

Getting Started

Applies to: CPersona 2.5.x. This page is the canonical installation and setup reference. The README keeps a condensed version of the same steps because it is also the PyPI project page; when the two disagree, this page wins.

CPersona is an MCP server. You install it, point an MCP client at it, and the client's agent gains store / recall tools that survive across sessions. Nothing else in your stack changes.

Prerequisites

  • Python 3.11+
  • uv for the one-command path (optional — pip works too)
  • An MCP client: Claude Desktop, Claude Code, Codex CLI, Cursor, VS Code, or any other MCP host — step 3 has the entry for each

Let the agent do it (Claude Code)

The repository and the published wheel both ship an Agent Skill It walks Claude Code through the installation, and then teaches it when to store, recall, and archive. Installing the skill is the shortest path:

# Installed from PyPI? The skill ships inside the wheel — no clone needed:
python -c "import cpersona,pathlib,shutil; s=pathlib.Path(cpersona.__file__).parent/'skills'/'cpersona-memory'; shutil.copytree(s, pathlib.Path.home()/'.claude/skills/cpersona-memory', dirs_exist_ok=True)"

# Running via uvx (isolated environment), or not installed yet:
git clone --depth 1 https://github.com/Cloto-dev/cpersona.git /tmp/cpersona
mkdir -p ~/.claude/skills && cp -r /tmp/cpersona/skills/cpersona-memory ~/.claude/skills/

Then tell Claude Code: "Set up CPersona — I want persistent memory."

The manual steps below are for every other client, and for anyone who prefers to configure things by hand. Step 5 is the part the skill would otherwise do: it makes the memory triggers fire without being asked.

1. Install CPersona

uvx cpersona          # run directly, no install step
# or
pip install cpersona  # then the `cpersona` command is on your PATH
From source (for development)
git clone https://github.com/Cloto-dev/cpersona.git
cd cpersona
python -m venv .venv
source .venv/bin/activate      # Windows: .venv\Scripts\activate
pip install .
Run it with `python -m cpersona` (or `python server.py`).
In a container
git clone https://github.com/Cloto-dev/cpersona.git
cd cpersona
docker build -t cpersona .

docker volume create cpersona-data
docker run -d --name cpersona -p 8402:8402 \
  -e CPERSONA_AUTH_TOKEN="$(openssl rand -hex 32)" \
  -v cpersona-data:/data cpersona
No image is published, so you build it yourself. Three things to know before you run it: - **It serves the Streamable HTTP transport**, on 8402. To run the stdio transport instead — the shape an MCP client spawns as a subprocess — pass `-i` and name it: `docker run -i --rm -e CPERSONA_TRANSPORT=stdio -v cpersona-data:/data cpersona`. - **It will not start without `CPERSONA_AUTH_TOKEN`.** A published container port forwards to whatever the process bound inside it, so binding inside the container does not keep the outside out. - **Your memory lives on the volume, not in the container.** The database goes in `/data`. Mount nothing there and every memory is lost when the container is replaced. A named volume (above) works as shown. A bind-mounted host directory does not inherit ownership, so it has to be writable by uid `10001` — or pass `--user "$(id -u)"`. Recall is keyword/FTS-only until you give it an embedding backend. See the next section, then pass `-e CPERSONA_EMBEDDING_MODE=http -e CPERSONA_EMBEDDING_URL=http://:8401/embed` once you have one.
In containers, with an embedding server `compose.yaml` in this repository runs CPersona and CEmbedding together, wired to each other, so vector search works without a second setup:
export CPERSONA_AUTH_TOKEN=$(openssl rand -hex 32)
docker compose run --rm embedding cembedding-download-model --model jina-v5-nano
docker compose up -d
Do not skip the middle line. Without it, the embedding server downloads its weights — around 800 MB — on the first request, while the port is already accepting connections it cannot answer. Fetching them once into the volume turns a timeout you would have to diagnose into a step you can watch. The embedding port is not published. CPersona reaches it over the network compose creates, and nothing else needs to. Both images are built locally, and the embedding server is pinned to a revision, so a later build gives you the same pair. Your memories are on the `cpersona-data` volume and the model on `embedding-model`; `docker compose down` leaves both, `down -v` deletes them.

At startup the server checks pypi.org for a newer release. It reports one through recall and check_health, and check_update answers on demand. Set CPERSONA_UPDATE_CHECK=false to turn the check off. Updating is never automatic: it takes an explicit check_update(apply=true) and a restart.

Connect a compatible embedding server. Running CPersona without one is supported as a fallback, but is not recommended for normal operation.

Vector search is the strongest of the three retrieval layers, and the only one that needs an external process. Without it CPersona still runs on FTS5 + keyword search, and says so on every recall.

CEmbedding is the reference backend. Any other embedding server that satisfies the contract below is supported just as well. The choice is yours.

The contract

CPersona is embedding-server-agnostic. Point CPERSONA_EMBEDDING_URL at any HTTP endpoint that implements this:

POST /embed
Request:  { "texts": ["string", ...] }        # non-empty array
Response: { "embeddings": [[float, ...], ...], "dimensions": <int> }

CPersona reads embeddings and nothing else. dimensions comes from the reference server and is ignored by the client, so a backend that omits it still works. CPersona sends at most 32 texts per request, and the reference server accepts up to 100. Batch limits in that range are nothing you need to plan around.

Three requirements are easy to miss. Each one degrades ranking silently:

  • Embeddings MUST be L2-normalized. CPersona computes similarity as a raw dot product, so a backend returning unnormalized vectors biases ranking by vector magnitude. Every supported backend (the client's api mode and all CEmbedding providers) already normalizes.
  • The contract is role-less. Queries and documents go through the same call, with no instruction prefix. Prompt-prefix models (e5-style, prompted bge) underperform here. Symmetric or retrieval-merged models (jina-v5-nano, bge-m3, MiniLM) are the intended fit.
  • Swapping models behind one URL invalidates the corpus. The contract carries no model identity, so CPersona fingerprints the backend by embedding dimension alone. A swap to a different model of the same dimension is undetectable.

The repair tools cannot reach it either. check_health(fix=true) re-embeds rows whose blob is NULL, and the dimension check only NULLs blobs of the wrong length. After a same-dimension swap every blob is the expected size, so nothing is NULLed and nothing is re-embedded. No tool force-re-embeds a row that already has a blob.

To recover, rebuild the corpus: delete_agent_data, then re-store as in the rebuild pattern, then run calibrate_threshold.

The reference server

CEmbedding (MIT) runs jina-v5-nano on-device (CPU) and exposes exactly this endpoint:

# Download the model into ./data/models
uvx --from "cembedding[onnx]" cembedding-download-model --model jina-v5-nano

# Run the server (it reads ./data/models from the current directory)
EMBEDDING_PROVIDER=onnx_jina_v5_nano uvx --from "cembedding[onnx]" cembedding

Or put it on your PATH with pip install "cembedding[onnx]" and run cembedding-download-model --model jina-v5-nano, then cembedding. From a source checkout the same two steps are python -m cembedding.download_model --model jina-v5-nano and python -m cembedding.

Either way you should see HTTP embedding endpoint started on http://127.0.0.1:8401/embed. Verify it before wiring CPersona to it:

curl -s http://127.0.0.1:8401/embed \
  -H 'content-type: application/json' \
  -d '{"texts":["hello world"]}' | head -c 200

CPersona's defaults are tuned against jina-v5-nano (768 dimensions). Any other server satisfying the contract works. Models with published measurements are listed in benchmarks/.

CPersona only needs the URL. But how you supervise the reference server matters, and the obvious way does not work.

It is an MCP server, not a plain HTTP process. On its default transport it runs an MCP session on stdio in the foreground and serves the REST /embed endpoint from a background task. Its lifetime is therefore bound to stdin: at EOF the session ends and the finally clause cancels the HTTP task.

A service manager starts things with stdin on /dev/null. Started that way, the server binds the port, logs HTTP embedding endpoint started, and exits in the same second with status 0. The supervisor sees a clean exit, and CPersona is left pointed at a URL nothing answers.

Give it a stdin that stays open. Under a service manager, run it through something that holds the pipe: ExecStart=/bin/sh -c 'sleep infinity | cembedding'. In a terminal, the terminal already does this.

EMBEDDING_TRANSPORT=streamable-http does not read stdin, so it supervises cleanly. But it serves the MCP endpoint instead of REST /embed, so it is not an option for CPersona's http mode, which posts to /embed.

3. Register CPersona with your MCP client

Every client launches the same process: command uvx, argument cpersona, and the environment shown below. They differ only in which file they read and what they call the keys.

Pick an absolute CPERSONA_DB_PATH first. The examples use /home/you/.claude/cpersona.db because that directory already exists for Claude users, but any absolute path works.

Client Where the entry goes Shape Checked here
Claude Code claude mcp add-json … -s user (writes the user config) JSON, type: stdio yes
Claude Desktop claude_desktop_config.json JSON, mcpServers yes
Codex CLI codex mcp add … -- uvx cpersona (writes ~/.codex/config.toml) TOML, [mcp_servers.cpersona] yes (codex-cli 0.147.0)
Cursor ~/.cursor/mcp.json (global) or .cursor/mcp.json (project) JSON, mcpServers vendor docs
VS Code (Copilot) .vscode/mcp.json (workspace) or the user mcp.json JSON, servers vendor docs
Any other MCP host its stdio server configuration the same command / args / env triple vendor docs

"Checked here" means the entry was written by that client's own tooling and read back on a maintainer machine while this page was written. "Vendor docs" means the shape was transcribed from the client's documentation and has not been executed here. If it disagrees with what your client accepts, the client is right — please report it.

Claude Desktop — add to claude_desktop_config.json:

{
  "mcpServers": {
    "cpersona": {
      "command": "uvx",
      "args": ["cpersona"],
      "env": {
        "CPERSONA_DB_PATH": "/home/you/.claude/cpersona.db",
        "EMBEDDING_MODE": "http",
        "EMBEDDING_HTTP_URL": "http://127.0.0.1:8401/embed"
      }
    }
  }
}

Claude Code — one command:

claude mcp add-json cpersona '{"type":"stdio","command":"uvx","args":["cpersona"],"env":{"CPERSONA_DB_PATH":"/home/you/.claude/cpersona.db","EMBEDDING_MODE":"http","EMBEDDING_HTTP_URL":"http://127.0.0.1:8401/embed"}}' -s user

Codex CLI — one command; it writes the TOML shown after it into ~/.codex/config.toml:

codex mcp add cpersona --env CPERSONA_DB_PATH=/home/you/.claude/cpersona.db --env EMBEDDING_MODE=http --env EMBEDDING_HTTP_URL=http://127.0.0.1:8401/embed -- uvx cpersona
[mcp_servers.cpersona]
command = "uvx"
args = ["cpersona"]

[mcp_servers.cpersona.env]
CPERSONA_DB_PATH = "/home/you/.claude/cpersona.db"
EMBEDDING_HTTP_URL = "http://127.0.0.1:8401/embed"
EMBEDDING_MODE = "http"

Codex can also deny-list tools per server (disabled_tools = ["delete_memory", …] under the same table). That is a client-side way to give an agent read- mostly access. The server-side equivalent, enforced for every client at once, is the per-client capability layer.

Cursor — ~/.cursor/mcp.json, or .cursor/mcp.json inside a project. Same shape as Claude Desktop:

{
  "mcpServers": {
    "cpersona": {
      "command": "uvx",
      "args": ["cpersona"],
      "env": {
        "CPERSONA_DB_PATH": "/home/you/.claude/cpersona.db",
        "EMBEDDING_MODE": "http",
        "EMBEDDING_HTTP_URL": "http://127.0.0.1:8401/embed"
      }
    }
  }
}

VS Code (Copilot) — .vscode/mcp.json in the workspace, or the user-level mcp.json (MCP: Open User Configuration). The top-level key is servers, not mcpServers:

{
  "servers": {
    "cpersona": {
      "command": "uvx",
      "args": ["cpersona"],
      "env": {
        "CPERSONA_DB_PATH": "/home/you/.claude/cpersona.db",
        "EMBEDDING_MODE": "http",
        "EMBEDDING_HTTP_URL": "http://127.0.0.1:8401/embed"
      }
    }
  }
}

Three things account for most setup problems:

  • Set CPERSONA_DB_PATH to an absolute path. The default, data/cpersona.db, is relative to the client's working directory, so a client launched from somewhere else opens a different, empty database. On Windows, write the path as C:/Users/you/.claude/cpersona.db.
  • No embedding server yet? Drop the two EMBEDDING_* lines (or set EMBEDDING_MODE=none). CPersona runs on FTS5 + keyword and reports that it is degraded.
  • EMBEDDING_MODE / EMBEDDING_HTTP_URL are the generic aliases of CPERSONA_EMBEDDING_MODE / CPERSONA_EMBEDDING_URL. The prefixed form wins when both are set. The configuration reference covers the settings you are likely to reach for, but it is not exhaustive: a few variables (CPERSONA_STORE_BLOB, CPERSONA_FTS_ENABLED, CPERSONA_EMBEDDING_API_KEY, the CPERSONA_CALIBRATE_* pair and some others) are read by the server without appearing there. cpersona/config.py is the complete list.

4. Verify it works

Ask the agent to store something, then recall it in a new session. Surviving the session boundary is the whole point:

"Store this: the deploy runbook lives in ops/deploy.md."

…then, in a fresh session: "What did I tell you about the deploy runbook?"

Two checks worth running once the corpus is real:

  • check_health — the registry-driven health check. status is the verdict; issues are severity-tagged (critical / warn / info), and check_health(fix=true) repairs the mechanical ones.
  • Watch recall responses for an advisory field. It means vector search is not contributing. The severity says why: a hint means embeddings are unconfigured (mode=none), and a fault means a configured endpoint stopped answering. See detecting a dead embedding server.

5. Make the memory triggers fire in every session

Registration gives the agent the tools. It does not make the agent use them unprompted. Recall at session start, store on a decision, archive at session end: those rules have to live in a file the client loads on every session, not in a skill that activates only when the conversation happens to match. Every client has such a file:

Client Always-loaded file (user-level default) Project-level alternative
Claude Code / Claude Desktop ~/.claude/CLAUDE.md ./CLAUDE.md
Codex CLI ~/.codex/AGENTS.md ./AGENTS.md at the repository root
Cursor User Rules (Customize → Rules — a setting, not a file) .cursor/rules/cpersona.mdc with alwaysApply: true, or ./AGENTS.md
VS Code (Copilot) see the client's custom-instructions documentation .github/copilot-instructions.md, or ./AGENTS.md with chat.useAgentsMdFile enabled

Paste the block below into that file. Replace <AGENT_ID> with one stable identifier the agent will use on every call ("claude-code", "codex", …).

Keep the markers. A later version of the block uses them to find and replace this one instead of stacking a second copy. On Claude Code the cpersona- memory skill does the paste for you, with your approval; everywhere else you paste it by hand.

The rules the block follows — consent, placement, idempotency, the 40-line budget, client neutrality — are the policy block standard. The block itself is maintained in the skill, and this copy is checked against it in CI.

The package also installs the paste as a command, for anyone who would rather not do it by hand — and for the case where the client refuses to let an agent write into a file it loads as instructions:

cpersona-policy --agent-id claude-code --install
# through uvx: uvx --from cpersona cpersona-policy --agent-id claude-code --install

It prints the block and writes nothing until --install is given, it picks the user-level file of whichever client it finds (--client claude-code|codex or --target PATH to say which file instead), and on a file that already holds a block it replaces the span between the markers and leaves every other byte as it was. Markers that do not form one well-formed block are reported rather than repaired. --dry-run says what would change.

<!-- BEGIN cpersona-policy v2 (managed by the cpersona-memory skill; re-run the skill to update) -->
## CPersona memory policy

Use the CPersona MCP tools proactively with `agent_id="<AGENT_ID>"` — never wait to be asked.

**Session start** → `recall(agent_id, query="<opening-topic keywords or ''>", limit=10)` before
the first substantive action. Prefer `recall_with_context` when conversation history is already
at hand; add `deep=true` when the first pass comes back thin. Skip only for trivial one-shot
questions.

**Decisions, rules, preferences, bug findings** → `store` immediately. Fire on phrases like
"let's go with X", "from now on always Y", "remember that…", "approved", "that's a bug".
Protect must-never-lose rules with `lock_memory`. After a successful `git commit`, `store` a
one-line record: hash, what changed, why.

**Changing an existing rule** → `update_memory`, never delete + store. If the memory is locked:
`unlock_memory` → `update_memory` → `lock_memory`.

**Session end** — fire on closing phrases ("that's all for today", "wrap it up", "good night") →
first `store` + lock any unsaved decisions, then `archive_episode(agent_id, history=<the REAL
turns>, summary=…, keywords=…, resolved=…)`, computing `summary` and `keywords` yourself.

**"Don't save this" / benchmark sessions** → `pause_persistence(ttl_seconds=1800)`;
`resume_persistence()` (or TTL expiry) restores. Reads still answer, minus the writes inside them.

**Degraded mode** — if a `recall` response carries an `advisory` field, surface it to the user
and follow its runbook. Never quietly serve keyword-only recall.

**Quality** — if recall feels off, `set_recall_precision` (strict/balanced/lenient) is the one
policy knob; run `calibrate_threshold(agent_id)` after the corpus changes substantially.
Monthly: `check_health(agent_id, fix=true)`.

**If this client keeps a memory file that loads every session** (Claude Code's `MEMORY.md`), use it
as the deterministic index over this store: one line per memory — `- <slug> — <the sentence that
changes behaviour>` — with the body stored here under `message.id="memory-index:<slug>"` and content
starting `[<slug>]`, so a line tells you what to `recall`. Recall is ranked and may not surface a
memory; the index always arrives. Its size cap fails **silently** when exceeded, so consolidate at
80%, not at the limit. Never migrate existing memories into this store without asking first.

Details, setup, and troubleshooting: the `cpersona-memory` skill.
<!-- END cpersona-policy -->

Where to go next

You want to… Read
Know which behaviors you can rely on Behavior Contracts
See what each tool does Tools
Understand how retrieval works Architecture
Back up, tune, or diagnose a live instance Operations Runbook
Look up a setting Configuration
Serve several clients over the network Remote HTTP transport