OpenAI-Compatible Provider
Claude-mem can run observation extraction against any endpoint that speaks the OpenAI/chat/completions shape. Set CLAUDE_MEM_PROVIDER to
openai-compatible, pick a preset (or supply a base URL), and memory runs
entirely off your Claude plan.
claude-mem observer: choosing the claude-mem observer during
npx claude-mem install runs memory off-plan with nothing to configure. This
page is for bringing your own endpoint and key.Why this is separate from the OpenRouter provider
CLAUDE_MEM_OPENROUTER_BASE_URL already lets the OpenRouter client talk to any
OpenAI-compatible endpoint, and it keeps working — nothing on that path changed.
But it carries three things that belong to openrouter.ai specifically:
- Attribution headers.
HTTP-RefererandX-Titleare sent to whatever host the base URL names, so pointing it at NVIDIA announces claude-mem’s OpenRouter identity to NVIDIA. usage: { include: true }, gated on the URL containingopenrouter.ai.- The cmem.ai gateway’s credential-pinning rules, which decide when a persisted base URL pins a persisted key. Those rules exist so an account-delivered gateway credential is never sent somewhere else — correct, and irrelevant to an endpoint you configured yourself, but you still share the same three settings keys with them.
model, messages, temperature, max_tokens, and nothing else. Strict
servers (vLLM in particular) reject unknown fields. max_tokens is
CLAUDE_MEM_OBSERVER_MAX_OUTPUT_TOKENS (4096 by default); a model that only
accepts max_completion_tokens gets the request again with that field.
A claude-mem observer key (cm_pro_…) is never sent to these endpoints. Use the
key the endpoint issued.
Quick start: NVIDIA NIM
1
Get a key
Sign in at build.nvidia.com and create an API key.
It starts with
nvapi-. NVIDIA’s hosted models are free to start, which suits
observation extraction: it is high-volume and structured, and does not need a
frontier model.2
Point claude-mem at it
settings.json and put it in ~/.claude-mem/.env:3
Pick a model (optional)
The preset defaults to a small, fast NVIDIA-served model. Any model id from
build.nvidia.com works and is passed verbatim:
4
Restart the worker
Presets
A preset only supplies a default base URL and a default model. Anything it sets is overridden by the matching setting, so switching preset never moves an endpoint or model you pinned yourself.Local presets need a model id — claude-mem cannot guess which model you loaded.
Set
CLAUDE_MEM_OPENAI_COMPAT_MODEL to whatever ollama list (or your server)
reports, or the provider reports itself unconfigured and dispatch falls through
to Claude.OpenCode Go and Zen
An OpenCode Go subscription or Zen key works here.opencode-go defaults to kimi-k3; opencode-zen ships no default, so set the
model too. Pick ids from https://opencode.ai/zen/go/v1/models or
https://opencode.ai/zen/v1/models. Only models served on /chat/completions
work.
Settings
An unknown preset id degrades to
custom rather than failing, because settings
are read during status polling and a typo should not crash a poll.
The viewer’s Settings → Advanced panel edits the provider, preset, base URL
and model. The API key is set in ~/.claude-mem/settings.json or
~/.claude-mem/.env only, never through the viewer.
MiniMax
minimax is MiniMax’s global endpoint and minimax-cn the mainland-China one;
use the preset that matches where your account lives. Both default to
MiniMax-M3. Any model id MiniMax serves is passed verbatim:
<think>…</think> block. claude-mem removes a leading block like that on every
preset (it also covers DeepSeek R1 or Qwen3 on Ollama and vLLM), so the reasoning
is never stored or re-sent with the next request.
Any other endpoint
When a key is required
SettingCLAUDE_MEM_OPENAI_COMPAT_BASE_URL replaces the preset’s endpoint, so
whether a key is needed is decided by the endpoint you pointed at, not by the
preset name you started from:
Without an override, the preset decides — the hosted presets need a key, the
local ones do not.
A remote endpoint that genuinely needs no authentication still works: set
CLAUDE_MEM_OPENAI_COMPAT_API_KEY to any non-empty value. The default errs
toward asking for a key because the alternative is a provider that reports
itself ready and then fails every observation with a 401.
OrcaRouter
OrcaRouter serves provider-scoped model ids, such asopenai/gpt-4o-mini (the preset’s default) or anthropic/claude-haiku-4.5,
behind one sk-orca- key. Its model list is public at
https://api.orcarouter.ai/v1/models.
Troubleshooting
Memory silently runs on Claude instead. Dispatch falls through to Claude when the provider is selected but not fully configured — a missing base URL, a missing model, or a missing key on a preset that needs one. That fall-through is deliberate: a half-configured endpoint should not fail every observation. Checknpx claude-mem doctor.
404 from the endpoint. Almost always a wrong base URL or a model the
endpoint does not serve; the error names both settings. Note that some gateways
prefix model ids differently from the vendor’s own docs.
Repeated timeouts. Each request has a deadline of
CLAUDE_MEM_LLM_TIMEOUT_MS (180 seconds by default), which a large model on a
slow endpoint can still exceed. Raise it, or prefer a small, fast model:
observation extraction is a high-volume structured task, not a reasoning
benchmark.
Empty or cut-off observations. A model that will not hold to the output
schema produces nothing usable. Small instruct-tuned models generally do better
here than reasoning models, whose thinking output tends to crowd out the
structured reply. When the log says a reply was cut off at the output-token
limit, raise CLAUDE_MEM_OBSERVER_MAX_OUTPUT_TOKENS.
