Skip to main content

Quota Fallback

When the observer’s provider reports that its allowance is spent, claude-mem arms a 30-minute quota cooldown and stops sending it work, so a capped account does not keep buying the same refusal. By default the work waits: observations stay queued until the cooldown clears. A quota fallback sends that work to a second provider instead, for as long as the first one is in its cooldown. When the cooldown ends, one generator probes the original provider; if it answers, everything moves back on its own.

Settings

Add these to ~/.claude-mem/settings.json, or set them in the viewer under Settings → Advanced: For example, Gemini as the primary with Claude Haiku as the fallback:
The fallback needs its own credentials: a Gemini, OpenRouter or OpenAI-compatible API key, or a working Claude Code login. A fallback with no API key, or any fallback that is itself held (a quota cooldown, or a refused key), is skipped, and work waits as it does without a fallback. A Claude fallback is not checked ahead of time: if Claude Code is missing or logged out, the fallback run fails the way a Claude primary would. The cmem.ai gateway is skipped as a fallback while it is turning your account away (its trial-expiry fallback is active).

What switches, and what does not

  • Switches: a spent allowance (a 30-minute cooldown), and rate limits that are still refused after the provider’s own retries (a 90-second cooldown). Work moves back as soon as the cooldown ends and the original provider answers its probe.
  • Does not switch: authentication errors — an invalid, expired or revoked key (401/403). Those pause capture exactly as before, because a different provider would not fix the key.
  • A fallback with a bad key is skipped once it is refused. Only a missing key is detected ahead of time; an invalid or revoked fallback key fails its first run, and its auth cooldown then takes it out of rotation. Fix or clear the fallback key.
  • Paused work resumes on its own. A generator that stops on a quota error keeps its queued observations. The worker’s periodic resume (once a minute) restarts them on whichever provider can take the work — the fallback, or the original provider again if it has recovered by then — and the session’s next tool call does the same sooner. If neither can, it waits as before.
  • Both exhausted: if the fallback is also in a quota cooldown, nothing changes: work waits until one of them clears.

Things to know

  • It spends the fallback’s allowance. With Claude as the fallback, observation runs count against your Claude plan. Setting CLAUDE_MEM_QUOTA_FALLBACK_MODEL to a Haiku model keeps that cheap.
  • Claude runs are capped by CLAUDE_MEM_MAX_CONCURRENT_AGENTS (default 2). With three or more busy sessions on a Claude fallback, the extra sessions wait for a slot.
  • Telegram wrap-ups are not governed by CLAUDE_MEM_QUOTA_FALLBACK_MODEL: they use the model the session last ran on, or the summary tier model.
  • Session-start notice. While another provider is serving, the notice says capture continues there instead of saying memory is paused. It is refreshed whenever a cooldown starts or ends, so right after a settings change it can be one step behind.
In the worker log, the switch shows as Primary in quota cooldown; dispatching to fallback, followed later by Primary quota cooldown elapsed; probing primary and Primary recovered from quota cooldown. If nothing can take the work, the log says so once: Primary and fallback both in quota cooldown; … or Primary in quota cooldown and the quota fallback cannot serve; ….