hermes-labs-ai

Fidelis Memory

Community hermes-labs-ai
Updated

Zero-LLM agent memory for Claude Code and AI agents: local-first BM25, dense-vector, and reciprocal-rank-fusion retrieval. Returns original passages verbatim by default. Available on PyPI as fidelis-memory. MIT.

Fidelis Memory

Local-first, zero-LLM memory for Codex, Claude Code, and AI agents.

83.2% R@1 in a checked-in 470-question LongMemEval-S retrieval run. A separate checked-in run answered 317 of 434 graded questions correctly (73.0%, Wilson 95% CI [68.7%, 77.0%]) with an LLM reading Fidelis retrieval. The default retrieval path itself makes no LLM call.

Stop re-explaining context to your agent. fidelis returns your original notes verbatim through a local-first service. Your agent already calls an LLM to think; it should not need another one just to remember. Designed for developers. The default zero-LLM retrieval path does not send memory content to an LLM. The documented fidelis init service configuration also disables mem0 and Chroma telemetry. That can reduce third-party data exposure, but deployments still own their security and compliance assessment.

License: MITStatus: pre-releaseCIPyPIOfficial MCP RegistryMade by Hermes Labs

your notes / sessions
       ↓
local memory store      (~/.cogito/, fully local)
       ↓
fidelis retrieval       (BM25 + dense + RRF, no LLM)
       ↓
original passages       (verbatim, never rephrased)
       ↓
Codex / Claude Code / your agent

What fidelis is:

  • model-API independent by default - the default retrieval path makes no model API call; local compute and storage still have costs
  • private - local memory store by default
  • faithful - original stored passages returned, not paraphrases
  • measured - checked-in LongMemEval-S retrieval and QA artifacts are linked below
  • installable - documented MCP paths for Codex, Claude Code, GitHub Copilot CLI, Gemini CLI, and OpenClaw

Fidelis is deliberately narrower than a hosted memory platform. Check theuser-fit matrix before installing: it names the workflows0.1.0 supports, the prerequisites it assumes, and the cases it does not yetserve.

Registries

  • Official MCP Registryio.github.hermes-labs-ai/fidelis-memory, latest published version 0.1.0.
  • —independent third-party server listing.

Docker / Glama

docker build . runs the MCP stdio server (fidelis mcp serve) by default —what a registry build/inspector (e.g. Glama) talks initialize / tools/listto — and needs no Ollama or running fidelis-server. To run the HTTP memoryserver in a container instead, set FIDELIS_ENTRYPOINT=http (seedocker-compose.yml for the full stack including Ollama).

Quickstart

# 0. one-time: Ollama + the local embedder (~280 MB)
brew install ollama && ollama serve &
ollama pull nomic-embed-text

# 1. install Fidelis Memory from PyPI
python3 -m pip install "fidelis-memory==0.1.0"
fidelis init                  # background service (launchd / systemd)
fidelis watch ~/notes         # auto-ingests markdown
fidelis mcp install --client codex   # or omit for Claude Code
fidelis mcp serve             # runs the MCP server over stdio
# Restart your agent client. Memory is on.

Verify the installed release and the local service before configuring a client:

python3 -c 'import fidelis; print(fidelis.__version__)'
# expected: 0.1.0
fidelis health
# expected prefix: status: ok  |  memories:

Then verify one real retrieval without relying on a fixed memory count:

mkdir -p /tmp/fidelis-verify
printf '%s\n' 'Fidelis verification phrase: amber heron.' > /tmp/fidelis-verify/note.md
fidelis watch /tmp/fidelis-verify --once
fidelis query 'amber heron'
# success: the result contains "Fidelis verification phrase: amber heron."

Using Gemini CLI? After the local prerequisites and fidelis init, installthe native v0.1.0 extension directly:

gemini extensions install https://github.com/hermes-labs-ai/fidelis --ref=v0.1.0

The extension launches the released MCP package through uvx and includes theGEMINI.md context file. See the Gemini CLI extension details.

Package-name note: install Hermes Labs' package as fidelis-memory.The import name and CLI remain fidelis. The separate PyPI project namedfidelis belongs to NGdust/fidelis.

Linux users swap brew install ollama for the equivalent install from ollama.com. See Requirements.

Fidelis Memory 0.1.0 is also published in theofficial MCP Registryas io.github.hermes-labs-ai/fidelis-memory. Registry-aware clients can launchthe same released server directly from PyPI:

uvx --from "fidelis-memory==0.1.0" fidelis mcp serve

This starts the MCP stdio process; run fidelis init first when the localFidelis service and store have not already been configured. Version 0.0.94introduced supported Codex MCP installation and context-sensitive orientation;0.0.96 added the independently discoverable registry release; 0.0.97 was thefirst tagged release that carried the Gemini CLI extension manifest; and 0.1.0promotes the tested cross-client contract as the first minor Fidelis release.

What you notice immediately

After the four commands above, the next time you open Codex or Claude Code:

  • It stops asking you to repeat context you already wrote down.
  • You can ask "what did we decide last week about auth?" - and the answer cites your actual decision, not a generic OAuth lecture.
  • Architecture rationale you wrote in a markdown file two months ago surfaces when relevant.
  • Your project context carries across sessions instead of resetting at every new conversation.
  • Failed migration notes, naming conventions, founder voice memos - all queryable in your agent's normal flow.

Most of fidelis's value is not the benchmark; it's not having to explain the same thing twice.

Most AI memory systems rewrite your notes

Most memory systems rephrase content on the way out. The specific fact gets summarized into something general. fidelis solves this structurally - there is no LLM in the default retrieval path, so the store returns exactly what you put in.

You store:

auth tokens expire after 3600 seconds.
The 3600s window is non-configurable in our current contract.

A lossy memory layer may return:

authentication has a configurable timeout

fidelis returns:

auth tokens expire after 3600 seconds.
The 3600s window is non-configurable in our current contract.

The non-configurable qualifier survives. So does every other detail you wrote down.

What this enables in Codex, Claude Code, GitHub Copilot CLI, Gemini CLI, and OpenClaw

Once fidelis mcp install --client codex, --client copilot, --client gemini, --client openclaw, or the default Claude install is run, ask your agent:

  • "What did we decide about auth?"
  • "What failed last time we tried this migration?"
  • "Which billing constraint was non-configurable?"
  • "What did I say about Sarah's onboarding flow?"

The MCP fidelis_recall tool gives the agent the original passages before it composes an answer, not paraphrased summaries. The answer can stay grounded in what you wrote, with the qualifiers intact.

fidelis retrieves memory without an LLM. Your agent still uses its normal LLM to answer using the retrieved context. "Zero-LLM" applies to the memory hot path, not to your agent.

GitHub Copilot CLI

Copilot CLI loads MCP servers from mcp-config.json in its configurationdirectory (~/.copilot by default, or $COPILOT_HOME). Fidelis writes thedocumented stdio entry there atomically, backing up any existing file andleaving other servers untouched:

fidelis mcp install --client copilot     # writes ~/.copilot/mcp-config.json
copilot                                  # restart, then /mcp list shows "fidelis"
                                         # /mcp show fidelis lists its tools
fidelis mcp uninstall --client copilot   # removes only the fidelis entry

Use --settings /path/to/mcp-config.json to target a different file. Thecopilot binary is not required at install time; if you prefer the host CLI,the equivalent registration iscopilot mcp add fidelis -- "$(python3 -c 'import sys;print(sys.executable)')" "$(python3 -c 'import fidelis.mcp_cmd as m;print(m.MCP_SERVER_FILE)')".Copilot does not currently expose a hook or automatic-recall mechanism tothird-party servers, so recall happens when the agent calls thefidelis_recall, fidelis_orient, or fidelis_health tools.

Gemini CLI

Gemini CLI has native MCP management — gemini mcp add|remove|list, shippedin v0.1.19 — and Fidelis registers itself through it rather than editingsettings.json. That matters: Gemini reads settings.json asJSON-with-comments and its own writer round-trips your // and /* */comments. A rewrite by Fidelis would silently delete them.

fidelis mcp install --client gemini      # gemini mcp add → ~/.gemini/settings.json
gemini                                   # restart, or run /mcp reload in a live session
gemini mcp list                          # shows "fidelis" and whether it connects
fidelis mcp uninstall --client gemini    # gemini mcp remove, verified

--scope project targets ./.gemini/settings.json instead of the default--scope user (~/.gemini/settings.json); Fidelis refuses --scope projectin your home directory, where Gemini collapses the two to the same file.Requires Gemini CLI v0.1.19 or newer on PATH, and an auth method alreadyconfigured — Gemini refuses every gemini mcp subcommand until one is.

Because gemini mcp add overwrites a same-named entry without asking andgemini mcp remove exits 0 even when the name is absent, Fidelis reads thetargeted settings.json back after every run. It refuses to touch a fidelisentry it does not recognize (--force overrides), and reports a silent no-opor an unexpected entry as a failure rather than as success. Unrelated servers,their env secrets, other settings keys, and the file's permission bits areleft as they were.

Recall happens when the agent calls the fidelis_recall, fidelis_orient, orfidelis_health tools.

OpenClaw

OpenClaw keeps outbound MCP servers under mcp.servers in its JSON5 config(~/.openclaw/openclaw.json, or $OPENCLAW_CONFIG_PATH). Because JSON5 allowscomments and trailing commas, Fidelis neither writes that file nor parses it:it delegates every write to the documented openclaw mcp add CLI, and asksOpenClaw's own read-only surface — openclaw mcp show fidelis --json, fallingback to openclaw mcp list --json — both before writing and afterwards toconfirm what landed.

fidelis mcp install --client openclaw    # openclaw mcp add fidelis --command … --arg …
openclaw mcp reload                      # pick up the new server
openclaw mcp status --verbose            # confirm the saved config
openclaw mcp doctor fidelis --probe      # verify it connects
fidelis mcp uninstall --client openclaw  # removes only the fidelis entry

The openclaw binary is required here, because it owns the write and is theonly reader that can be trusted with a JSON5 config. Use--settings /path/to/openclaw.json to target a different config; Fidelis passesit as $OPENCLAW_CONFIG_PATH on every delegated call, reads included, so thestate it reads back is the state of the file OpenClaw just wrote. If you preferto run the host CLI yourself, theequivalent registration isopenclaw mcp add fidelis --command "$(python3 -c 'import sys;print(sys.executable)')" --arg "$(python3 -c 'import fidelis.mcp_cmd as m;print(m.MCP_SERVER_FILE)')".Install and uninstall refuse to touch an mcp.servers.fidelis entry that is notours unless you pass --force, and exit non-zero rather than claiming successwhenever the read-back does not prove the change landed — including whenOpenClaw cannot report the entry at all, which is treated as unknown, never as"nothing there".

Use cases & ROI

Three concrete reasons teams pick fidelis over hosted memory:

  • Model-API independence for retrieval. Memory lives on disk and the default retrieval path makes no model API call. Your agent still consumes its normal context and model resources when answering.
  • Local data boundary. The default zero-LLM path keeps notes and retrieval on the local machine, reducing third-party processor exposure. This architecture does not by itself confer SOC 2 or HIPAA compliance.
  • Team context. Agents that remember historical decisions, naming conventions, failed migrations, and the qualifiers on those decisions. The non-configurable detail you wrote down two months ago surfaces when relevant, in the founder's voice, not paraphrased.

How it fits

The diagram is at the top. Codex and Claude Code are the fastest paths to value. The retrieval engine is agent-agnostic - pair it with any LLM client. Codex registration uses its supported codex mcp CLI, and the resulting server configuration is shared by the Codex desktop app, CLI, and IDE extension on that host.

Benchmarks

Checked-in LongMemEval-S observations; these are local project measurements,not independent replications.

Metric Value
Retrieval R@1 83.2%
Retrieval R@5 98.3%
End-to-end QA accuracy 73.0% (317/434 graded questions), Wilson 95% CI [68.7%, 77.0%]
Retrieval-time model API calls 0 on the default stage-1 path

Raw evidence: retrieval aggregate ·end-to-end QA summary

The QA tier wraps your existing LLM with a 140–180-token system prompt - the Fidelis Scaffold. See docs/scaffold.md.

Verify the zero-LLM claim yourself

# Unset any LLM API keys for this shell
unset OPENAI_API_KEY ANTHROPIC_API_KEY DASHSCOPE_API_KEY

# Optional: drop your network. Ollama runs on 127.0.0.1:11434 (loopback).

# `recall-hybrid` is the explicit-tier command. zero_llm is the default.
fidelis recall-hybrid "what did the user say about Sarah" --tier zero_llm
tail ~/.fidelis/server.log

The default zero_llm tier never makes an outbound LLM call. Optional --tier filter and --tier flagship modes do call an LLM, but only to select integer pointers - the server dereferences those pointers to the original stored text. The LLM cannot rephrase memory content.

Context-sensitive orientation (MCP)

The bundled MCP server also exposes fidelis_orient. It recognizes when aturn invokes prior work—even when it is a statement such as “I need toremember our Fidelis work”—and selects a bounded evidence lane for identity,maintenance, conceptual reuse, comparison, decisions, historical state, orcurrent state. The returned orientation is a derived index; retrieved recordsremain verbatim evidence with their existing IDs and metadata. Unrelated turnsexplicitly abstain without calling the memory server.

Gemini CLI extension

Fidelis is also packaged as a nativeGemini CLI extension: thegemini-extension.json at the repository root registers the same stdio MCPserver that the MCP Registry entry launches, plus a GEMINI.mdcontext file that tells the model when to call fidelis_orient andfidelis_recall. It needs uv on PATH and arunning Fidelis server (fidelis init, see Requirements),but not a manual pip install:

gemini extensions install https://github.com/hermes-labs-ai/fidelis
gemini extensions list      # fidelis, with its GEMINI.md and MCP server
gemini extensions uninstall fidelis

The extension pins fidelis-memory==0.1.0; gemini extensions update fidelisfollows the repository's tagged releases. The first launch lets uvx downloadthe wheel and its dependencies. Gemini CLI 0.32.1 probes gemini mcp listwith a fixed 5-second timeout that ignores the manifest's 60-second timeout,so that first launch can read Disconnected; runuvx --from fidelis-memory==0.1.0 fidelis --help once to warm the cache,after which the row reads Connected. If you also register Fidelis withgemini mcp add, the settings.json entry takes precedence over theextension's, so the two do not conflict.

Requirements

  • macOS or Linux (Windows not yet supported)

  • Python 3.10+

  • Ollama running locally with nomic-embed-text pulled (~280 MB):

    brew install ollama && ollama serve &
    ollama pull nomic-embed-text   # ~280 MB, one-time
    

Once Ollama and the embedding model are available, the quickstart covers thefull init-to-first-recall path. The default retrieval path needs no memory APIkey.

Ollama is currently required to boot the service at all, including for thedefault zero-LLM retrieval path. The BM25 + dense + RRF retrieval logicitself makes no LLM call, but fidelis-server boots through mem0'sMemory.from_config(), and mem0's Ollama embedder validates its connectionat construction time — before any query runs. We installed fidelis-memoryfrom PyPI in a clean venv and confirmed this directly:

python3 -m venv /tmp/fv && source /tmp/fv/bin/activate
pip install "fidelis-memory==0.1.0"
python3 -c 'import fidelis; print(fidelis.__version__)'
# 0.1.0 — installs and imports fine, no Ollama needed for this step

COGITO_OLLAMA_URL=http://127.0.0.1:1 fidelis-server   # Ollama unreachable on purpose
ConnectionError: Failed to connect to Ollama. Please check that Ollama is
downloaded, running and accessible. https://ollama.com/download
  File ".../mem0/embeddings/ollama.py", line 30, in _ensure_model_exists
    local_models = self.client.list()["models"]

The package installs and imports cleanly without Ollama. The server process— and every documented path that goes through it (fidelis health, fidelis query, fidelis recall-hybrid --tier zero_llm, the MCP server, andfidelis.augment) — does not start without a reachable Ollama instance. Thereis currently no lighter-weight standalone way to exercise the zero-LLMretrieval path without the full Ollama + service stack. This is a real gapbetween the "zero-LLM retrieval" framing and the actual boot dependency; weare not fixing the Ollama boot coupling here, just documenting it honestly soyou know what to expect before you install Ollama.

Quick reference

fidelis recall "what did the user say about Sarah"
fidelis query  "Sarah" --limit 5
fidelis add    "raw text to extract into memories"
fidelis health
fidelis seed   ~/memory/   ~/notes/

fidelis add normally stores facts produced by the configured extractionmodel. If extraction returns no facts, Fidelis preserves the original inputverbatim instead of silently losing it. The command still exits 0 because thewrite succeeded, but stdout reports a stable degraded status:

status=stored degraded=verbatim-fallback-empty-extraction id=<uuid> count=1

Automation that requires successful extraction must inspect degraded; exit 0means the memory was stored, not necessarily that extraction succeeded. Becausemem0 does not distinguish a swallowed extractor failure from a legitimatezero-fact result, the fallback intentionally favors durability.

Python helper for direct integration:

from fidelis.augment import augment
from anthropic import Anthropic

client = Anthropic()
answer = augment(
    question="What did I say about Sarah?",
    qtype="single-session-user",
    llm_call=lambda system, user: client.messages.create(
        model="claude-haiku-4-5",  # any current Claude Messages model works
        system=system,
        messages=[{"role": "user", "content": user}],
        max_tokens=512,
    ).content[0].text,
)

What's running on your machine

After fidelis init:

  • Service: fidelis-server runs at http://127.0.0.1:19420 under your OS service manager (launchd on macOS, systemd on Linux). Auto-starts on boot. Logs at ~/.fidelis/server.log.
  • Storage: Chroma + SQLite at ~/.cogito/ (the directory name is preserved from the project's pre-rename codename for v0.0.x compatibility - it will move to ~/.fidelis/ in a later major bump). No data leaves your machine in the default zero-LLM path.
  • MCP: after installing for your selected client, Codex or Claude Code sees four tools: fidelis_recall, fidelis_query, fidelis_health, and fidelis_orient.

To stop: fidelis init --uninstall. To wipe: rm -rf ~/.cogito ~/.fidelis.

Known limitations (v0.1.0)

  • Pre-release. Python function names and CLI commands may change. Pin the version if you build on it.
  • Best on macOS Sequoia / Ubuntu 24.04 LTS. Other OSes likely work but aren't gate-tested.
  • Direct server launches disable mem0 telemetry by default. This matchesthe service installed by fidelis init and avoids telemetry exit handlersdelaying graceful shutdown. An explicit MEM0_TELEMETRY=True still opts in.For the same boundary across Chroma, set ANONYMIZED_TELEMETRY=False andCHROMA_TELEMETRY_DISABLED=True before a direct launch; fidelis initincludes all three settings automatically.
  • Temporal-reasoning and preference questions are the weakest qtypes in the QA scaffold (TR ~58%, Pref ~37% on the full eval). Single-session and knowledge-update qtypes are strong (95–100%).
  • The optional LLM tier ("flagship" mode) currently escalates ~80% of queries instead of the intended ~10% - an 8× cost miss we're transparent about. The default zero-LLM tier is unaffected.
  • qwen3.5:9b in thinking mode does not reliably follow the literal hedge instruction in the Fidelis Scaffold. Use Claude, an OpenAI-format API, or non-thinking-mode local models for reliable hedging.

What this turns into over time

Day 1: drop notes into ~/notes, run the four commands.Day 2: ask your agent about yesterday's decision - the answer cites your original passage.Day 7: your agent starts carrying project context across sessions; you stop re-explaining.

Useful for solo builders today; relevant for teams that need memory to stay local tomorrow.

Fidelis Memory for teams

fidelis is open-source under MIT and free for any use, including commercial. If your team has deployment requirements that the OSS path does not yet cover (centralized memory, multi-namespace isolation, custom authentication), write to [email protected].

For technical users

  • docs/user-fit.md - supported users, prerequisites, and explicit non-fits
  • docs/releases/0.1.0.md - 0.1.0 release scope and acceptance evidence
  • ROADMAP.md - outcome gates for 0.2.0
  • docs/full-reference.md - full architecture, hybrid recall tiers, local server endpoints, troubleshooting
  • docs/scaffold.md - Fidelis Scaffold contract + drift-detection markers
  • experiments/zeroLLM-FLAGSHIP-evidence/ - raw eval JSONs + machine-readable SUMMARY (per-qtype breakdowns, Wilson CI, F1/F1B baselines)

License

MIT. Built by Hermes Labs (Roli Bosch). Issues + PRs welcome.

Also from Hermes Labs

  • lintlang - Static analysis for AI agent configs, tool descriptions, and system prompts; zero-LLM, deterministic checks built for CI.
  • zer0dex - A local dual-layer memory pattern: a compact markdown index paired with semantic retrieval from a local vector store, queried before each message.
  • little-canary - Detects prompt injection by its effect on a sacrificial canary model, returning block/flag/pass before your primary model acts.
  • quick-gate-js - A deterministic JS/TS CI quality gate that unifies ESLint, TypeScript, build, and Lighthouse checks into one fail-fast result.

About Hermes Labs

Hermes Labs develops open-source reliability, evaluation, memory, andruntime-guard tools for AI agents. Fidelis is its local-first memory project.Other public software is listed atgithub.com/hermes-labs-ai, with researchartifacts published separately on Zenodo.

For enterprise deployments and AI-reliability engagements: [email protected] · hermes-labs.ai

On naming. Hermes Labs is named for Hermes, the Greek messenger god - patron of communication and interpretation, the herald who carries meaning between worlds. The thread to the work: hermeneutics, the theory of interpretation that takes its name from Hermes, is the philosophical anchor for an AI reliability engineering studio whose substrate is linguistic. Not affiliated with NousResearch's Hermes LLM line or their hermes-agent framework - different companies, different work.

Founder: Rolando (Roli) Bosch.Site: hermes-labs.aiCitation: Bosch, R. (2026). Hermes Labs: AI reliability infrastructure for autonomous agents. https://hermes-labs.ai

Quantitative source for the Fidelis claims above: the 470-questionLongMemEval-S aggregate and Wilson interval inexperiments/zeroLLM-FLAGSHIP-evidence/,evaluated 2026-04-24.

MCP Server · Populars

MCP Server · New

    hermes-labs-ai

    Fidelis Memory

    Zero-LLM agent memory for Claude Code and AI agents: local-first BM25, dense-vector, and reciprocal-rank-fusion retrieval. Returns original passages verbatim by default. Available on PyPI as fidelis-memory. MIT.

    Community hermes-labs-ai
    n24q02m

    Better Code Review Graph

    Knowledge graph for token-efficient code reviews -- semantic search and call-graph resolution across your codebase.

    Community n24q02m
    Noveum

    Orbit

    Free, open source, realtime task manager. Issues, boards, sprints, projects and docs that sync instantly. Keyboard-first, self-hostable, with an MCP server for AI agents. No pricing, ever.

    Community Noveum
    feder-cr

    aihawk

    Anti detect browser and web browsing agent: an open-source MCP server for undetected browsing, AI web scraping and computer use agents. No captchas.

    Community feder-cr
    LeandroPG19

    MemoryIndustry

    Persistent memory MCP server for AI agents — Rust, 19 tools, knowledge graph, Hebbian learning, episodic memory, contradiction detection, prospective triggers, Bayesian calibration, zero-config Docker setup.

    Community LeandroPG19