Skip to main content

Overview

The Kimi Code harness gives claude-mem persistent memory to agents running in the Kimi Code CLI. It handles three things:
  1. Observation recording — Captures tool usage from Kimi’s hook system and sends it to the claude-mem worker for AI processing
  2. Context injection — Injects the observation timeline once per session (and once after each compaction) via the UserPromptSubmit hook’s stdout (chained after session-init), keeping MEMORY.md free for agent-curated memory
  3. MCP search tools — Registers the mcp-search stdio server so you can query memory manually from inside a Kimi session
Kimi Code uses a Claude-Code-like hook system driven by ~/.kimi-code/config.toml. The claude-mem installer merges hook rules into that file and registers an MCP server in ~/.kimi-code/mcp.json without touching your personal AGENTS.md or ~/.agents/ files.

How It Works

Event Lifecycle

1

Session starts (SessionStart startup|resume)

When Kimi starts or resumes a session, the hook starts the claude-mem worker if it is not already running. This hook only warms up the worker; its stdout is not appended to Kimi’s context.
2

User prompt submitted (UserPromptSubmit)

The single composite session-init-context hook sends the user prompt to POST /api/sessions/init so the worker can create or continue the session. On the first prompt of a session it also appends the injected observation timeline to Kimi’s context via stdout; later prompts in the same session skip the timeline fetch (a marker under ~/.claude-mem/state/kimi-context-injected/ tracks this). This is a blockable Kimi event; claude-mem always exits 0 so the prompt is never blocked.
3

Tool is about to read a file (PreToolUse Read)

Before Kimi executes a Read tool, the hook queries GET /api/observations/by-file for the target path and injects the returned timeline as additional context. This mirrors the Claude Code PreToolUse matcher.
4

Tool use recorded (PostToolUse)

After a successful tool use, the hook sends the tool name, input, and response to POST /api/sessions/observations. This is fire-and-forget — it does not block the agent from continuing work.
5

Session stops (Stop)

When the session stops, the hook extracts the last assistant message and sends it to POST /api/sessions/summarize.
6

Context compacted (PreCompact manual|auto)

Before Kimi compacts context manually or automatically, the hook triggers another summarize call so the session state is preserved before compression, and clears the injection marker so the first prompt after compaction re-injects a fresh timeline.

Event Mapping

Unused in v1 (the Claude integration likewise ignores them): SessionEnd, SessionHeartbeat, SubagentStart/Stop, TurnStarted, Notification, PermissionRequest/Result, Interrupt, PostCompact, UserPromptQueued.

Context Injection

The UserPromptSubmit hook injects cross-session observation context into the first Kimi prompt of a session by writing the worker’s formatted context to stdout. The content comes from GET /api/context/inject?projects=<project>, which generates a markdown timeline from the SQLite database. The composite session-init-context event performs the session-init and context work in a single process so the hook payload is consumed only once. Kimi fires UserPromptSubmit on every message and only appends stdout to context for that event (SessionStart stdout is dropped), so the timeline could not be delivered at session start directly. To avoid re-injecting the full timeline into every prompt, the handler records a per-session marker (~/.claude-mem/state/kimi-context-injected/<sessionId>) after a successful injection and skips the timeline fetch while the marker exists. The PreCompact hook clears the marker, so the first prompt after a compaction re-injects a fresh timeline. If the worker is unreachable the marker is not written, so injection is retried on the next prompt. Opt-in per-prompt semantic injection (CLAUDE_MEM_SEMANTIC_INJECT=true) is unaffected and still runs on every prompt. This approach keeps MEMORY.md under your control for curated long-term memory (decisions, preferences, durable facts), while the observation timeline is delivered through the prompt-submit context where it belongs.

Installation

Run the interactive installer and let it auto-detect Kimi Code:
The installer detects Kimi via the ~/.kimi-code/ directory or the kimi binary on PATH. You can also manage the harness directly:

What Gets Installed

The installer writes a marker-delimited managed block into $KIMI_CODE_HOME/config.toml (default ~/.kimi-code/config.toml) containing seven [[hooks]] rules:
Only the four fields event, matcher, command, and timeout are written. Timeouts range from 60–120 seconds. The installer also merges the mcp-search stdio MCP server into $KIMI_CODE_HOME/mcp.json under mcpServers if it is absent, preserving any servers already configured.
KIMI_CODE_HOME is honored when set; otherwise the default root is ~/.kimi-code.
The first time the installer modifies config.toml, it writes a backup to config.toml.bak-<YYYYMMDD> before making any changes.

Commands

npx claude-mem kimi install

Install or upgrade the Kimi harness. If the managed block already exists, it is replaced in place (upgrade-safe).

npx claude-mem kimi status

Report whether the managed hook block is present in config.toml.

npx claude-mem kimi uninstall

Remove exactly the claude-mem managed block from config.toml, and the mcp-search entry claude-mem wrote to mcp.json (an entry of that name pointing anywhere else is left alone). All other Kimi configuration is untouched. npx claude-mem uninstall runs the same cleanup.

Error Handling

All Kimi hooks fail open: if the worker is unreachable, the command crashes, or any internal error occurs, the hook exits 0 so Kimi never blocks a session. This matches Kimi’s own hook semantics, where exit 0 means allow and stdout may append to context.

Requirements

  • Kimi Code CLI installed and configured
  • Claude-mem worker service running (auto-started by the SessionStart hook if needed)
  • Network access between the Kimi process and the worker on the per-user port