vault-search — fastembed Migration Runbook
How to move a vault-search install off Ollama-based embeddings onto in-process
fastembed, and how to replicate the change on another machine. Migration done on the primary Mac 2026-05-16; this runbook is for the second Mac and any future install.
Why
vault-search (vsearch.py) computes embeddings for the semantic half of hybrid search. The original code called Ollama over the network (OLLAMA_HOST, default http://localhost:11434). On a machine with no Ollama running, every embed call fails and hybrid search silently degrades to keyword-only (FTS). The degradation is invisible — results still return, just without semantic matching.
The fix replaces the embedding backend with fastembed: the same model (nomic-embed-text, 768-dim) running as ONNX on CPU, in-process. No daemon, no network, nothing to be unreachable. A loud-fallback guard was added so any future embedder failure is visible (via='fts-fallback' on result rows + a stderr warning) rather than silent.
nomic-embed-text is an embedding model, not a generative LLM — it does not load a chat model into RAM and does not bog the machine. The Ollama process that historically bogged the Mac was a generative model; this is a different weight class.
Prerequisites
uvinstalled.vsearch.py's shebang is#!/usr/bin/env -S uv run --script.uvreads the inline PEP-723 dependency block and installs everything — includingfastembedandonnxruntime— automatically on first run. No manualpip install.- Internet, once. First run:
uvfetches deps; the firstembeddownloads thenomic-ai/nomic-embed-text-v1.5ONNX model (~a few hundred MB, cached after). - The vault itself (the source of truth). The vault-search DB is pure regenerable cache.
Replication — simple path
The tool is one self-contained file. If the target machine's vsearch.py is the same version as the migrated one:
# from the migrated machine:
scp ~/code/vault-search/vsearch.py TARGET:~/code/vault-search/vsearch.py
# on the target machine:
cd ~/code/vault-search
./vsearch.py index # rebuild FTS index from the vault
./vsearch.py embed --verbose # embed all sections (fastembed, in-process)
./vsearch.py search "a conceptual query with no exact keyword match" --mode vec
The vec-mode search at the end is the verification: it must return results via vec. If it does, semantic search works.
Gotchas
- The CLI command is
index, notreindex.reindexerrors with "No such command"; if the DB was just deleted, that leaves the index empty. - Vault location.
vsearch.pydefaultsVAULTto~/vaultand the DB to~/.cache/vault-search.db. If the target vault is elsewhere, setVAULT_PATH=/path/to/vaultas an env var for the CLI runs and in the MCP server config. - MCP server config. If the target already runs vault-search as a stdio MCP server, no config change is needed — the new code uses no
OLLAMA_HOST;"env": {}is fine. The running MCP server picks up the new code on the next Claude Code restart (stdio servers reloadvsearch.pyon respawn). - Latent bug on the target. If the target's
vsearch.pyis the old Ollama version and no Ollama runs there, its semantic search has been silently degraded too. This migration fixes it. - First
embedrun is the slow one (model download). Subsequent runs are fast.
Replication — if the target's vsearch.py has diverged
Do not blind-copy. Apply the six changes by hand:
- PEP-723 deps (the
# /// scriptblock): add# "fastembed>=0.4",. - Config constants: drop the
OLLAMA_HOST = ...line; setEMBED_MODELdefault to"nomic-ai/nomic-embed-text-v1.5".EMBED_DIMstays 768. - Replace
_ollama_embedwith two functions:_get_embedder()(lazy module-globalfastembed.TextEmbedding(model_name=EMBED_MODEL)), and_embed(text, is_query=False)— usesmodel.query_embed(text)whenis_queryelsemodel.embed([text]), returnslist[float] | None. _vec_rows: call_embed(query, is_query=True); returnNone(not[]) when the embedder fails —Nonemeans "embedder down",[]means "genuine empty result". Update the return annotation tolist[tuple] | None.embed_missing: call_embed("ping")for the availability check and_embed(text, is_query=False)in the loop.search()loud-fallback: capturerequested_mode = mode; aftervec = _vec_rows(...), setvec_degraded = vec is None, thenif vec is None: vec = []; if degraded and hybrid/vec requested, print a warning to stderr; computefts_via = "fts-fallback" if (requested_mode == "hybrid" and vec_degraded) else "fts"and usefts_viafor theviafield of FTS-mode result rows.
Steps 4 and 6 together are the loud-fallback hardening; steps 1-3 and 5 are the Ollama→fastembed swap. The swap alone restores semantic search; the hardening makes future failures visible.
Verification
./vsearch.py search "<paraphrased query, no shared keywords with target docs>" --mode vec
# expect: result rows tagged "via vec"
# forced-degradation check (confirms the loud fallback works):
VSEARCH_EMBED_MODEL="nonexistent-model-xyz" ./vsearch.py search "test" --mode hybrid
# expect: stderr warning "embedder unavailable — semantic search degraded"
# and result rows tagged "via fts-fallback"
Also check DB coverage:
python3 -c "import sqlite3; c=sqlite3.connect('$HOME/.cache/vault-search.db'); \
print('sections', c.execute('select count(*) from sections_fts').fetchone()[0]); \
print('embedded', c.execute('select count(*) from section_vecs_rowids').fetchone()[0])"
# embedded should equal sections (100% coverage)
Rollback
vsearch.py is a single file under version control / easily restored; the DB is pure cache. To roll back: restore the prior vsearch.py, rm ~/.cache/vault-search.db, run ./vsearch.py index (and embed if Ollama is available again). No data is at risk — the vault is the source of truth.
See also
Infrastructure Cluster · local-llm-setup · primary tool: ~/code/vault-search/vsearch.py