Skip to content
Main Site News Console

Give Your Coding Agent a Memory You Control

· Hugging Face Translated
教程模型卡

I switch between multiple machines and choose different coding agents depending on the task.

Each of them feels like a stranger when working with my projects. The reasoning process from “last Tuesday” disappears when the session ends. Every new agent on every new host can only start from scratch.

Earlier this year, “Software Forgets: Agent Traces Are Memory” argued that coding agents are already generating the records we keep losing. As they search the codebase, try different approaches, encounter errors, read documentation, and change direction, they leave behind detailed records of not only what they changed, but why they changed it.

Although that assessment is correct, traces are only potential memory. An agent’s session logs are still just an archive. Faced with ten thousand rounds of interaction, you cannot answer “Why did we abandon the streaming parser?” with grep alone. For agents to use these traces while working, you need indexing, retrieval, ranking, and precise provenance tracking.

That is what funes provides. It is a persistent memory layer for agents, with support for Claude Code, Codex, pi, and Hermes. It builds on the sessions already present on your machine, runs entirely locally, and can become part of an agent’s daily workflow with a single command. If you want, it can also sync to a Hugging Face dataset owned by you, private by default.

Add memory to the agents you already use

funes is a single-file binary. Its default inference backend does not depend on a machine learning runtime; embeddings and reranking both happen on your machine. Install it with:

Terminal window
curl -fsSL https://huggingface.co/buckets/huggingface/funes/resolve/install.sh | sh

Then add it to an agent:

Terminal window
funes add claude # or: codex, pi, hermes

The add command builds the initial index, provides the agent with recall and get tools, and installs automation that indexes each completed interaction round. Indexing is incremental: new runs add only new interaction rounds instead of re-embedding the entire history. Older and deeper content can be indexed later in controlled batches.

From then on, you simply work as usual. When a task involves past decisions, rationales, or discoveries, the agent can call recall on its own. You do not need to remember old sessions or paste their context into a new one.

A coding agent calls funes on its own, recalls an earlier decision, and answers based on the retrieved session

After adding funes, recall happens within the conversation. The agent calls on memory by itself and identifies the sessions supporting its answer.

recall returns the original text rather than a summary, and shows its provenance precisely: the agent, timestamp, session, and interaction round. Every result includes a get command that can open the complete interaction round and its surrounding context.

Under the hood, a deterministic pipeline parses all supported traces into a unified “round-block” structure, chunks them, generates embeddings with a fixed-version local model, and writes them to a local Lance dataset. Queries combine vector and BM25 search, fuse their rankings, rerank candidates with a cross-encoder, reweight them based on recency, and append adjacent blocks.

This design gives funes three important properties:

  • One shared memory across agents: Claude Code, Codex, pi, and Hermes all write to the same structure. recall can search across their histories, and every hit identifies which agent generated it.

  • Original evidence remains intact: Content is not distilled into facts when written. Results can always be traced back to the interaction round that produced them.

  • recall runs locally by default: No account or Hub repository is required. Hosted models do not process your sessions to build the index; embeddings and reranking run on your machine, while inference is performed by your coding agent.

On a single machine, the “agent feels like a stranger” problem is solved. But memory becomes even more useful when the next agent runs somewhere else.

Memory is a dataset, not a service

To have memory follow your work, bind a memory when adding funes to an agent:

Terminal window
funes add codex acme/funes-memory

Binding publishes the current memory to that location. Afterward, funes indexes each interaction round locally and publishes updates at session boundaries, keeping the memory continuously up to date. The agent keeps recalling from it. Run the same command on another machine, and your memory follows you into the past.

Under the hood, local memory is a Lance dataset, while shared memory is a Hugging Face dataset owned by you, private by default.

Credentials are redacted during indexing, before any content reaches the Hub. At publish time, every block is scanned again, and anything that still looks like a secret is blocked. The scanner behind this process is documented in SECURITY.md, including what it does and does not cover.

When an agent reads remote memory, funes caches the dataset files locally, so queries return to local speeds once the cache is warmed. The Hub provides the ownership, access control, versioning, and distribution capabilities it already offers for other datasets. Your memory does not become an account in a separate memory service, nor do you have to rent it back through an API.

Ask first, integrate later

recall is designed for agents. When you want to ask the memory a question yourself, use ask. By default, it reads from local memory:

Terminal window
funes ask claude "What decision did we make about the streaming parser?"

Alternatively, you can point it at shared memory. We published a memory about the development of funes, so you can ask why funes works the way it does without creating your own memory:

Terminal window
funes ask claude "Why does funes use append-only mode?" --memory huggingface/funes-memory

Asking the published funes memory why it uses append-only mode; funes retrieves relevant sessions, and the coding agent answers based on them

funes ask is the read-only, single-question version of funes add. It retrieves relevant passages, passes them to the coding agent, and returns an evidence-based answer while identifying its sources. It does not install an integration or change the agent’s persistent settings.

Retrieval misses are not concealed. If the retrieved passages are insufficient to support an answer, the agent says so explicitly. You can rephrase the question or add funes to the agent so it can iteratively search memory during normal work.

Switch agents without losing context

Shared memory is not tied to the agent or model that created it. Start a task in Claude Code and continue it in Codex the following week; the second agent can recall the first agent’s reasoning. You can run pi with a local model, use a model served by the Hugging Face router, and then return to Claude.

Claude Code selects an embedding model, and Codex later recalls the decision in another session

Claude makes a decision; a hook indexes it; Codex recalls the decision in another session. The earlier hit in the demo is an older record from the same experiment: append-only memory also remembers the earlier rehearsal.

This matters across several different scopes:

  • Across machines: Bind every agent to the same memory and retrieve your history from whichever host you are currently using.

  • Across teams: New team members’ agents can search months of decisions from day one, including deta

#教程#模型卡#AI编程#AI智能体#开源

Published by the 4All API team

Original link:https://huggingface.co/blog/funes

Need a mainstream LLM API? 4All API gives you one key to call OpenAI, Anthropic, Google Gemini, Qwen, DeepSeek, and dozens more — at official-pass-through pricing, with enterprise-grade reliability, integrated in 5 minutes.

Sign up for the 4All API console →