From 9a2d7cde0e8447fdfa0bcaf95e9e9441a27ea545 Mon Sep 17 00:00:00 2001 From: Andrey Vasilevsky Date: Wed, 30 Sep 2026 18:26:13 -0700 Subject: [PATCH 1/3] fix(mcp): make the watcher honour workspace filters; harden the shared engine Generated files kept out of the initial index were let straight back in by the MCP file watcher, and each event they caused cost several full-graph scans. On an 8 GB machine running several agent sessions that showed up as sustained watcher work and memory pressure (#23). Watcher - One exclusion rule. The watcher carried its own ten-entry SKIP_DIRS list and was never given the index config, so --exclude and .codegraphignore applied to the initial index only. is_excluded_entry() now states the indexer's rule once; the indexer's walk and a new WorkspaceFilter both use it. The watcher checks every ancestor of an event path, since an event names only the file and the walk only ever reached a file through its parents. default_exclude_dirs is a strict superset of the old SKIP_DIRS, so nothing the watcher skipped before is admitted now. - One pass per batch. process_changes ran a full-graph property("path") query for every deleted file, vanished file, changed file, its dependents and each dependent. It now builds one path-to-nodes map per batch, and a delete for a file the index never held takes no lock at all. Vector re-embedding is batched the same way (update_files_vectors), replacing a full node walk per changed file. - Symlinked workspaces. The indexer stores paths as the workspace was given; FSEvents reports the resolved location (/var -> /private/var, or any symlinked folder). Every lookup by event path missed, so on such a workspace an edit re-parsed a file without removing its old symbols and a delete removed nothing: the graph only grew. Present in 0.20.1. Event paths are now translated into the indexed form at intake, and the filter matches roots both as given and resolved. - Orphan vectors. remove_file_vectors ran before the nodes were deleted and only drops vectors whose node is already gone, so it could not see the file it was called for. One prune_orphan_vectors after the batch covers deletions and the vectors left behind by re-parsing, which nothing cleaned up before. - A delete-only batch now rebuilds the indexes too. This is index hygiene: search resolves results against the graph, so deleted symbols were not visible before either. Graph-only - index_workspace and the daemon-attach path initialised the memory manager, which loads the embedding model, before anything checked --graph-only. The existing comment on the graph-only branch already promised the model would never load; it now doesn't. Memory tools are unavailable in graph-only mode, the state the RAM gate already leaves them in on low-memory hosts. Shared engine - Resource settings. An auto-spawned engine was passed only the model name, so a client's --exclude, --max-files and --graph-only were dropped. engine_args() forwards the engine-level settings, EngineConfig carries graph_only, and a graph-only engine no longer loads the shared model. --profile is deliberately not forwarded: it filters one client's tool list, and a shared engine would impose it on every other client. - One load per workspace. Two attaches to a cold workspace both built a full backend, each indexing and starting a watcher, and the loser returned its own unregistered copy, serving that connection from it for its whole life. The registry now holds a per-workspace OnceCell. - One engine per socket. Concurrent auto-starts were guarded only by a socket probe taken before a model load that can take minutes; the second engine then removed the first one's live socket to bind its own, leaving the first running and unreachable. Reproduced: two clients, two engines. An exclusive lock (std File::try_lock) is now taken before anything is loaded and held for the engine's lifetime. Verified end to end against the shipped 0.20.1 binary and this build, each watcher test with a positive control (a fresh, non-excluded file must still be picked up). Unit tests cover the filter (including an explicit symlink, since Linux CI has no /var -> /private/var), most-specific-root selection, max depth, the engine lock and the forwarded arguments. Not addressed: .gitignore is honoured by neither the watcher nor the initial index, which stay consistent; supporting it means adding the ignore crate. Reported and diagnosed by Christopher Schulze in #23, whose analysis of the watcher filter, the per-event scans, graph-only ordering and the shared engine's dropped settings was accurate on every point. Co-Authored-By: Claude Opus 5.5 Claude-Session: https://claude.ai/code/session_017rVbt7rENTwXkdHt3Bpgb5 --- .../codegraph-server/src/ai_query/engine.rs | 25 +- crates/codegraph-server/src/indexer.rs | 280 +++++++++++++++-- crates/codegraph-server/src/main.rs | 86 +++++- crates/codegraph-server/src/mcp/engine.rs | 151 +++++++-- .../codegraph-server/src/mcp/file_watcher.rs | 292 ++++++++---------- crates/codegraph-server/src/mcp/server.rs | 27 +- 6 files changed, 629 insertions(+), 232 deletions(-) diff --git a/crates/codegraph-server/src/ai_query/engine.rs b/crates/codegraph-server/src/ai_query/engine.rs index c6defd9..1af1ae9 100644 --- a/crates/codegraph-server/src/ai_query/engine.rs +++ b/crates/codegraph-server/src/ai_query/engine.rs @@ -812,6 +812,24 @@ impl QueryEngine { /// Re-embed only symbols from a specific file path. /// Called on did_save to incrementally update embeddings without rebuilding all. pub async fn update_file_vectors(&self, file_path: &str) { + self.update_files_vectors(&[file_path.to_string()]).await; + } + + /// Re-embed the symbols of several files in one pass over the graph. + /// + /// Calling [`Self::update_file_vectors`] once per file walks every node once + /// per file, so a burst of N changed files cost N full graph scans - one of + /// the per-event costs behind issue #23. This walks the graph once and + /// embeds everything that matched in one batch. + pub async fn update_files_vectors(&self, file_paths: &[String]) { + if file_paths.is_empty() { + return; + } + let label = match file_paths { + [only] => only.clone(), + many => format!("{} files", many.len()), + }; + let file_path = label.as_str(); let engine = match self.vector_engine.read().await.clone() { Some(e) => e, None => { @@ -841,8 +859,13 @@ impl QueryEngine { continue; } + // Matched by suffix in either direction, as before: stored paths and + // event paths are not guaranteed to agree on being absolute. let path = node_props::path(node); - if !path.ends_with(file_path) && !file_path.ends_with(path) { + if !file_paths + .iter() + .any(|fp| path.ends_with(fp.as_str()) || fp.ends_with(path)) + { continue; } diff --git a/crates/codegraph-server/src/indexer.rs b/crates/codegraph-server/src/indexer.rs index 3733afa..41fbb9b 100644 --- a/crates/codegraph-server/src/indexer.rs +++ b/crates/codegraph-server/src/indexer.rs @@ -15,6 +15,155 @@ use std::path::{Path, PathBuf}; use std::sync::Arc; use tokio::sync::{Mutex, RwLock}; +/// Whether one directory entry is excluded from the graph. +/// +/// This is the single statement of the rule. The indexer applies it to every +/// entry as it walks, and [`WorkspaceFilter`] applies it to every component of +/// a path the file watcher reports. The watcher used to carry its own shorter +/// list instead, so `--exclude`, `.codegraphignore` and most of the default +/// exclusions kept generated files out of the initial index and then let every +/// later write to them straight back in (issue #23). +pub(crate) fn is_excluded_entry( + config: &IndexConfig, + exclude_set: &globset::GlobSet, + path: &Path, + is_dir: bool, +) -> bool { + let Some(name) = path.file_name() else { + return true; + }; + let name = name.to_string_lossy(); + if name.starts_with('.') { + return true; + } + if exclude_set.is_match(path) { + return true; + } + is_dir + && (config.exclude_dirs.iter().any(|d| d == name.as_ref()) + || exclude_set.is_match(name.as_ref())) +} + +/// Decides whether the file watcher may admit a path, by the same rule the +/// indexer used to build the graph. +/// +/// The indexer only reaches a file after every directory above it has passed, +/// because it recurses. A watcher event arrives as one bare path with no such +/// history, so the rule is applied to each ancestor in turn - otherwise a file +/// under an excluded directory would be judged only by its own name. +/// +/// Built per workspace root because `.codegraphignore` is read per root, as +/// the indexer does. It is read once, when the watcher starts: editing +/// `.codegraphignore` takes effect on the next start, like `--exclude`. +pub(crate) struct WorkspaceFilter { + /// `(root as given, root canonicalised, config, exclude globs)`. + roots: Vec<(PathBuf, PathBuf, IndexConfig, globset::GlobSet)>, +} + +impl WorkspaceFilter { + pub(crate) fn new(roots: &[PathBuf], base: &IndexConfig) -> Self { + let roots = roots + .iter() + .map(|root| { + let mut config = base.clone(); + config.extend_from_codegraphignore(root); + let exclude_set = config.build_exclude_set(); + let canonical = root.canonicalize().unwrap_or_else(|_| root.clone()); + (root.clone(), canonical, config, exclude_set) + }) + .collect(); + Self { roots } + } + + /// Whether `path`, a file, belongs in the graph. + /// + /// Deleted files cannot be stat'd, so the size limit is only applied to a + /// file that still exists; every other check depends on the path alone. + /// That keeps a delete for a file the index never held from being admitted + /// on a technicality, and a delete for one it did hold from being refused. + pub(crate) fn admits(&self, path: &Path) -> bool { + let Some((root, relative, config, exclude_set)) = self.locate(path) else { + return false; + }; + + let components: Vec<_> = relative.components().collect(); + // Directories above the file, matching the walk's depth accounting. + if components.len().saturating_sub(1) > config.max_depth as usize { + return false; + } + + let mut current = root; + for (index, component) in components.iter().enumerate() { + current.push(component); + let is_dir = index + 1 < components.len(); + if is_excluded_entry(config, exclude_set, ¤t, is_dir) { + return false; + } + } + + match std::fs::metadata(path) { + Ok(metadata) => metadata.len() <= config.max_file_size_bytes, + Err(_) => true, + } + } + + /// `path` spelled the way the indexer stored it. + /// + /// The indexer records each file under the workspace root as it was given, + /// while FSEvents reports the resolved location. On a workspace reached + /// through a symlink the two never match, so every lookup by event path + /// missed: an edit re-parsed the file without removing its old symbols, a + /// delete removed nothing, and the graph only ever grew. Translating at + /// intake gives every later step one spelling. + pub(crate) fn indexed_form(&self, path: &Path) -> PathBuf { + match self.locate(path) { + Some((root, relative, _, _)) => root.join(relative), + None => path.to_path_buf(), + } + } + + /// The workspace root `path` falls under, in the form it was given, and + /// `path` relative to it. + /// + /// Roots and event paths do not reliably share a spelling. A workspace + /// opened as `/var/folders/...` or through any symlinked directory is + /// reported by FSEvents at its resolved location, `/private/var/folders/...`, + /// so a plain prefix test rejected every event and the watcher admitted + /// nothing at all. Both sides are compared resolved and as given. + /// + /// The event path is resolved through its parent directory, which still + /// exists when the file itself was just deleted. The most specific root + /// wins when workspace folders nest. + fn locate(&self, path: &Path) -> Option<(PathBuf, PathBuf, &IndexConfig, &globset::GlobSet)> { + let resolved = path + .parent() + .and_then(|parent| parent.canonicalize().ok()) + .zip(path.file_name()) + .map(|(parent, name)| parent.join(name)); + + let mut best: Option<(usize, PathBuf, PathBuf, &IndexConfig, &globset::GlobSet)> = None; + for (given, canonical, config, exclude_set) in &self.roots { + for (candidate, root) in [(Some(path), given), (resolved.as_deref(), canonical)] { + let Some(candidate) = candidate else { continue }; + let Ok(relative) = candidate.strip_prefix(root) else { + continue; + }; + let depth = root.components().count(); + if best.as_ref().is_none_or(|(d, ..)| depth > *d) { + best = Some(( + depth, + given.clone(), + relative.to_path_buf(), + config, + exclude_set, + )); + } + } + } + best.map(|(_, root, relative, config, exclude_set)| (root, relative, config, exclude_set)) + } +} + /// Configuration for a single indexing run. #[derive(Debug, Clone)] pub struct IndexConfig { @@ -492,33 +641,12 @@ impl Indexer { let path = entry.path(); - // Skip hidden files and directories - if let Some(name) = path.file_name() { - if name.to_string_lossy().starts_with('.') { - continue; - } + let is_dir = path.is_dir(); + if is_excluded_entry(config, &exclude_set, &path, is_dir) { + continue; } - if path.is_dir() { - let dir_name = path - .file_name() - .map(|n| n.to_string_lossy().to_string()) - .unwrap_or_default(); - - // Skip hardcoded exclude directories - if config.exclude_dirs.iter().any(|e| e == &dir_name) { - continue; - } - - // Skip directories matching user-configured exclude globs - let path_str = path.to_string_lossy(); - if exclude_set.is_match(path_str.as_ref()) - || exclude_set.is_match(dir_name.as_str()) - { - tracing::info!("Skipping {:?}: matched exclude pattern", path); - continue; - } - + if is_dir { let (t, p, s, child_by_lang, child_errors) = self .index_directory(graph, &path, config, depth + 1, counter.clone()) .await; @@ -532,12 +660,6 @@ impl Indexer { *parser_errors.entry(lang).or_insert(0) += count; } } else if path.is_file() { - // Skip files matching exclude globs - let path_str = path.to_string_lossy(); - if exclude_set.is_match(path_str.as_ref()) { - continue; - } - // Skip files that exceed the configurable size limit if let Ok(metadata) = std::fs::metadata(&path) { if metadata.len() > config.max_file_size_bytes { @@ -662,6 +784,102 @@ mod tests { ); } + fn filter_config(exclude_dirs: &[&str], max_depth: u32) -> IndexConfig { + IndexConfig { + exclude_dirs: exclude_dirs.iter().map(|d| d.to_string()).collect(), + exclude_patterns: vec![], + max_file_size_bytes: 1024 * 1024, + max_depth, + max_files: 1000, + } + } + + #[test] + fn workspace_filter_excludes_what_the_indexer_excludes() { + let tmp = tempfile::tempdir().unwrap(); + let root = tmp.path().to_path_buf(); + std::fs::create_dir_all(root.join("src/cache")).unwrap(); + let filter = + WorkspaceFilter::new(std::slice::from_ref(&root), &filter_config(&["cache"], 20)); + + assert!(filter.admits(&root.join("src/lib.rs"))); + // An excluded directory excludes everything beneath it, however deep - + // a watcher event names only the file, so every ancestor is checked. + assert!(!filter.admits(&root.join("src/cache/gen.rs"))); + assert!(!filter.admits(&root.join("cache/deep/er/gen.rs"))); + // Hidden entries, at any level, as the indexer's walk skips them. + assert!(!filter.admits(&root.join(".hidden.rs"))); + assert!(!filter.admits(&root.join("src/.git/x.rs"))); + // Outside every workspace root. + assert!(!filter.admits(Path::new("/somewhere/else/lib.rs"))); + } + + #[test] + fn workspace_filter_reads_codegraphignore_per_root() { + let tmp = tempfile::tempdir().unwrap(); + let root = tmp.path().to_path_buf(); + std::fs::write(root.join(".codegraphignore"), "**/generated/**\n").unwrap(); + let filter = WorkspaceFilter::new(std::slice::from_ref(&root), &filter_config(&[], 20)); + + assert!(!filter.admits(&root.join("src/generated/out.rs"))); + assert!(filter.admits(&root.join("src/handwritten.rs"))); + } + + #[test] + fn workspace_filter_enforces_max_depth() { + let tmp = tempfile::tempdir().unwrap(); + let root = tmp.path().to_path_buf(); + let filter = WorkspaceFilter::new(std::slice::from_ref(&root), &filter_config(&[], 1)); + + assert!(filter.admits(&root.join("a/file.rs"))); + assert!(!filter.admits(&root.join("a/b/file.rs"))); + } + + /// A workspace reached through a symlink is reported by the OS at its + /// resolved location. An explicit symlink is used rather than relying on + /// macOS's `/var` -> `/private/var`, which Linux CI would not exercise. + #[cfg(unix)] + #[test] + fn workspace_filter_matches_events_at_the_resolved_path() { + let tmp = tempfile::tempdir().unwrap(); + let real = tmp.path().join("real"); + std::fs::create_dir_all(real.join("src/cache")).unwrap(); + let link = tmp.path().join("link"); + std::os::unix::fs::symlink(&real, &link).unwrap(); + + let filter = + WorkspaceFilter::new(std::slice::from_ref(&link), &filter_config(&["cache"], 20)); + let resolved = real.canonicalize().unwrap(); + + // Admitted and excluded by the same rule whichever spelling arrives. + assert!(filter.admits(&resolved.join("src/lib.rs"))); + assert!(!filter.admits(&resolved.join("src/cache/gen.rs"))); + + // Translated back to the spelling the indexer stored, so lookups by + // path find the file's existing nodes instead of duplicating them. + assert_eq!( + filter.indexed_form(&resolved.join("src/lib.rs")), + link.join("src/lib.rs") + ); + assert_eq!( + filter.indexed_form(&link.join("src/lib.rs")), + link.join("src/lib.rs") + ); + } + + #[test] + fn workspace_filter_prefers_the_most_specific_root() { + let tmp = tempfile::tempdir().unwrap(); + let outer = tmp.path().to_path_buf(); + let inner = outer.join("inner"); + std::fs::create_dir_all(&inner).unwrap(); + // Only the inner root excludes `vendor`. + std::fs::write(inner.join(".codegraphignore"), "**/vendor/**\n").unwrap(); + let filter = WorkspaceFilter::new(&[outer.clone(), inner.clone()], &filter_config(&[], 20)); + + assert!(!filter.admits(&inner.join("vendor/x.rs"))); + } + #[test] fn extend_from_codegraphignore_appends_patterns() { let tmp = tempfile::tempdir().expect("tmp"); diff --git a/crates/codegraph-server/src/main.rs b/crates/codegraph-server/src/main.rs index ce49297..0f25c0b 100644 --- a/crates/codegraph-server/src/main.rs +++ b/crates/codegraph-server/src/main.rs @@ -125,6 +125,39 @@ struct Args { socket: Option, } +/// The settings an auto-spawned engine is started with. +/// +/// Only the model name used to be passed, so an engine started on a client's +/// behalf ignored its `--exclude`, `--max-files` and `--graph-only` and indexed +/// and embedded as if none had been given (issue #23). These are the +/// engine-level settings: they decide what is indexed and whether a model is +/// loaded, for every client of that engine. +/// +/// `--profile` is deliberately not forwarded. It filters the tools one client +/// is shown, and a shared engine would apply the first client's profile to +/// every client that connects after it. +/// +/// `--full-body-embedding` is not forwarded either: on this release it cannot +/// be set to false (a bare `default_value` on a bool parses as an always-true +/// flag), so the engine's default is already the only value a client can ask +/// for. +fn engine_args(args: &Args) -> Vec { + let mut out: Vec = vec![ + "--embedding-model".into(), + args.embedding_model.clone().into(), + "--max-files".into(), + args.max_files.to_string().into(), + ]; + for dir in &args.exclude { + out.push("--exclude".into()); + out.push(dir.into()); + } + if args.graph_only { + out.push("--graph-only".into()); + } + out +} + /// Default engine socket path (`~/.codegraph/cg-engine.sock`). fn default_socket_path() -> PathBuf { let home = std::env::var_os("HOME") @@ -336,7 +369,7 @@ async fn run() { std::env::current_dir().expect("Failed to get current directory") }); if let Err(e) = - codegraph_server::mcp::engine::connect(&sock, workspace, &args.embedding_model).await + codegraph_server::mcp::engine::connect(&sock, workspace, &engine_args(&args)).await { eprintln!("connect failed: {e}"); std::process::exit(1); @@ -422,6 +455,7 @@ async fn run() { exclude_dirs: args.exclude.clone(), max_files: args.max_files, full_body_embedding: args.full_body_embedding, + graph_only: args.graph_only, seeds: args.workspace.clone(), }; if let Err(e) = codegraph_server::mcp::engine::serve(cfg).await { @@ -598,3 +632,53 @@ mod crash_breadcrumb_tests { let _ = std::fs::remove_dir_all(&tmp); } } + +#[cfg(test)] +mod engine_args_tests { + use super::*; + + fn forwarded(argv: &[&str]) -> Vec { + let args = + Args::try_parse_from(std::iter::once("codegraph-server").chain(argv.iter().copied())) + .expect("parses"); + engine_args(&args) + .into_iter() + .map(|a| a.into_string().unwrap()) + .collect() + } + + #[test] + fn engine_level_resource_settings_reach_an_auto_spawned_engine() { + let got = forwarded(&[ + "--connect", + "--graph-only", + "--max-files", + "7", + "--exclude", + "cache", + "--exclude", + "generated", + "--embedding-model", + "static", + ]); + let pairs: Vec<&[String]> = got.chunks(2).collect(); + assert!(got.contains(&"--graph-only".to_string())); + assert!(pairs.iter().any(|p| p == &["--max-files", "7"])); + assert!(pairs.iter().any(|p| p == &["--embedding-model", "static"])); + let excludes: Vec<&String> = got + .windows(2) + .filter(|w| w[0] == "--exclude") + .map(|w| &w[1]) + .collect(); + assert_eq!(excludes, ["cache", "generated"]); + } + + #[test] + fn per_client_settings_stay_with_the_client() { + // A shared engine would apply one client's profile to every client. + let got = forwarded(&["--connect", "--profile", "core"]); + assert!(!got.iter().any(|a| a == "--profile" || a == "core")); + // And graph-only is not forced on when the client did not ask for it. + assert!(!got.iter().any(|a| a == "--graph-only")); + } +} diff --git a/crates/codegraph-server/src/mcp/engine.rs b/crates/codegraph-server/src/mcp/engine.rs index db5b523..561d0ee 100644 --- a/crates/codegraph-server/src/mcp/engine.rs +++ b/crates/codegraph-server/src/mcp/engine.rs @@ -29,6 +29,11 @@ pub struct EngineConfig { pub exclude_dirs: Vec, pub max_files: usize, pub full_body_embedding: bool, + /// Build graphs only and never load the embedding model. An auto-spawned + /// engine used to drop this along with every other resource setting but the + /// model name, so a client that asked for graph-only still caused a model + /// load in the engine it started (issue #23). + pub graph_only: bool, /// Workspaces to pre-load at startup (optional; others load on attach). pub seeds: Vec, } @@ -75,7 +80,10 @@ mod imp { use codegraph_memory::VectorEngine; - type Registry = Arc>>>; + /// One slot per workspace. The slot is created under the lock; the load + /// runs inside it, outside the lock, and every concurrent attach for the + /// same workspace awaits that one load. + type Registry = Arc>>>>>; struct Engine { cfg: EngineConfig, @@ -90,34 +98,42 @@ mod imp { } /// Load (or fetch from the registry) the backend for `workspace`, reusing the - /// shared model. Builds outside the registry lock so a slow first index of - /// one workspace doesn't block attaches to others. + /// shared model. + /// + /// This used to check the registry, release the lock, build, then insert - + /// so two attaches to a cold workspace both built a full backend, each + /// indexing the workspace and starting its own file watcher. Worse, the + /// loser kept the registry's entry but returned its own backend, so that + /// connection was served for its whole life by a second, unregistered copy + /// (issue #23). A per-workspace `OnceCell` makes the second attach wait for + /// the first load instead, while a slow first index of one workspace still + /// does not block attaches to others. async fn get_or_load(engine: &Arc, workspace: PathBuf) -> Arc { let ws = workspace .canonicalize() .unwrap_or_else(|_| workspace.clone()); - if let Some(s) = engine.registry.lock().await.get(&ws).cloned() { - return s; - } + let slot = { + let mut registry = engine.registry.lock().await; + Arc::clone(registry.entry(ws.clone()).or_default()) + }; + Arc::clone(slot.get_or_init(|| load_workspace(engine, ws)).await) + } + async fn load_workspace(engine: &Arc, ws: PathBuf) -> Arc { tracing::info!("Engine: loading workspace {}", ws.display()); let mut server = McpServer::new( - vec![ws.clone()], + vec![ws], engine.cfg.exclude_dirs.clone(), engine.cfg.max_files, engine.cfg.embedding_model.clone(), engine.cfg.full_body_embedding, - ); + ) + .with_graph_only(engine.cfg.graph_only); if let Some(shared) = &engine.shared_engine { server.set_shared_engine(Arc::clone(shared)).await; } server.ensure_indexed().await; - let server = Arc::new(server); - - let mut reg = engine.registry.lock().await; - // Another connection may have loaded it while we built — prefer theirs. - reg.entry(ws).or_insert_with(|| Arc::clone(&server)); - server + Arc::new(server) } /// First line of a connection: an attach frame selects the workspace. @@ -199,6 +215,47 @@ mod imp { } } + enum EngineLock { + /// This process is the engine for the socket. Released by the OS when + /// the process exits, crashes included. + Held(std::fs::File), + /// Another live process already is. + Contended, + /// The lock file could not be opened or locked. + Unavailable, + } + + /// Take the exclusive lock, kept beside the socket, that marks this process + /// as the engine for `socket_path`. + fn acquire_engine_lock(socket_path: &std::path::Path) -> EngineLock { + let mut lock_path = socket_path.as_os_str().to_owned(); + lock_path.push(".lock"); + let lock_path = PathBuf::from(lock_path); + if let Some(parent) = lock_path.parent() { + let _ = std::fs::create_dir_all(parent); + } + let file = match std::fs::OpenOptions::new() + .create(true) + .truncate(false) + .write(true) + .open(&lock_path) + { + Ok(f) => f, + Err(e) => { + tracing::warn!("Engine: cannot open {}: {e}", lock_path.display()); + return EngineLock::Unavailable; + } + }; + match file.try_lock() { + Ok(()) => EngineLock::Held(file), + Err(std::fs::TryLockError::WouldBlock) => EngineLock::Contended, + Err(std::fs::TryLockError::Error(e)) => { + tracing::warn!("Engine: cannot lock {}: {e}", lock_path.display()); + EngineLock::Unavailable + } + } + } + /// True if a live engine is already listening on the socket (probe by /// connecting). Distinguishes a stale socket file from a running engine so we /// neither steal a live socket nor refuse to start over a dead one. @@ -207,7 +264,25 @@ mod imp { } pub async fn serve(cfg: EngineConfig) -> Result<(), String> { - // Don't start a second engine over a live one (handles auto-spawn races). + // Held for the life of the engine. The socket probe alone did not stop + // two auto-spawned engines: both probed before either had bound, both + // loaded a model and their seed workspaces - which can take minutes - + // and the second then removed the first one's live socket file to bind + // its own, leaving the first running, holding a model, and unreachable + // until its idle timeout (issue #23). Taking an exclusive lock first + // makes the loser exit before it loads anything. + let _engine_lock = match acquire_engine_lock(&cfg.socket_path) { + EngineLock::Held(lock) => Some(lock), + EngineLock::Contended => { + tracing::info!("Engine: another instance owns this socket — exiting"); + return Ok(()); + } + // Could not lock at all: fall back to the socket probe below + // rather than refuse to start. + EngineLock::Unavailable => None, + }; + + // Still checked: an engine started by an older build holds no lock. if engine_is_live(&cfg.socket_path).await { tracing::info!("Engine: another instance is already live — exiting"); return Ok(()); @@ -216,7 +291,10 @@ mod imp { // One model for the whole engine. Gate on free memory the same way the // per-workspace path does, so a constrained box runs graph-only instead // of OOM-crashing on the model load. - let shared_engine = { + let shared_engine = if cfg.graph_only { + tracing::info!("Engine: graph-only — not loading an embedding model"); + None + } else { let mut sys = sysinfo::System::new(); sys.refresh_memory(); let avail = sys.available_memory(); @@ -310,7 +388,7 @@ mod imp { /// Spawn a detached engine for `socket_path` using this binary, so the engine /// outlives the shim. Best-effort; the caller retries the connect. - fn spawn_engine(socket_path: &std::path::Path, embedding_model: &str) { + fn spawn_engine(socket_path: &std::path::Path, engine_args: &[std::ffi::OsString]) { let exe = match std::env::current_exe() { Ok(e) => e, Err(e) => { @@ -322,8 +400,7 @@ mod imp { cmd.arg("--serve") .arg("--socket") .arg(socket_path) - .arg("--embedding-model") - .arg(embedding_model) + .args(engine_args) .stdin(std::process::Stdio::null()) .stdout(std::process::Stdio::null()) .stderr(std::process::Stdio::null()); @@ -336,10 +413,13 @@ mod imp { } } + /// `engine_args` are the engine-level settings an auto-spawned engine should + /// start with. They only take effect when this call is the one that starts + /// the engine: one already running keeps the settings it was started with. pub async fn connect( socket_path: &std::path::Path, workspace: PathBuf, - embedding_model: &str, + engine_args: &[std::ffi::OsString], ) -> Result<(), String> { use tokio::io::{AsyncReadExt, AsyncWriteExt}; @@ -347,7 +427,7 @@ mod imp { let stream = match UnixStream::connect(socket_path).await { Ok(s) => s, Err(_) => { - spawn_engine(socket_path, embedding_model); + spawn_engine(socket_path, engine_args); let mut connected = None; for _ in 0..60 { tokio::time::sleep(Duration::from_millis(500)).await; @@ -404,6 +484,33 @@ mod imp { } Ok(()) } + + #[cfg(test)] + mod lock_tests { + use super::*; + + /// flock is per open file description, not per process, so a second + /// open in the same process contends exactly as a second engine would. + #[test] + fn only_one_engine_may_own_a_socket() { + let dir = tempfile::tempdir().unwrap(); + let socket = dir.path().join("cg-engine.sock"); + + let first = acquire_engine_lock(&socket); + assert!( + matches!(first, EngineLock::Held(_)), + "first engine takes the lock" + ); + assert!( + matches!(acquire_engine_lock(&socket), EngineLock::Contended), + "a second engine must back off before loading anything" + ); + + // Released on drop, as on process exit: the socket is free again. + drop(first); + assert!(matches!(acquire_engine_lock(&socket), EngineLock::Held(_))); + } + } } #[cfg(unix)] @@ -418,7 +525,7 @@ pub async fn serve(_cfg: EngineConfig) -> Result<(), String> { pub async fn connect( _socket_path: &std::path::Path, _workspace: PathBuf, - _embedding_model: &str, + _engine_args: &[std::ffi::OsString], ) -> Result<(), String> { Err("the socket engine is not yet supported on this platform".to_string()) } diff --git a/crates/codegraph-server/src/mcp/file_watcher.rs b/crates/codegraph-server/src/mcp/file_watcher.rs index e45203f..dd1e286 100644 --- a/crates/codegraph-server/src/mcp/file_watcher.rs +++ b/crates/codegraph-server/src/mcp/file_watcher.rs @@ -9,11 +9,12 @@ //! resolves cross-file imports, and rebuilds search indexes. use crate::ai_query::QueryEngine; +use crate::indexer::{IndexConfig, WorkspaceFilter}; use crate::parser_registry::ParserRegistry; use crate::watcher::GraphUpdater; -use codegraph::CodeGraph; +use codegraph::{CodeGraph, NodeId}; use notify::{Config, Event, EventKind, RecommendedWatcher, RecursiveMode, Watcher}; -use std::collections::HashSet; +use std::collections::{HashMap, HashSet}; use std::path::{Path, PathBuf}; use std::sync::Arc; use std::time::{Duration, Instant}; @@ -22,20 +23,6 @@ use tokio::sync::{mpsc, RwLock}; /// Debounce interval — wait 2 seconds after last change before processing. const DEBOUNCE_MS: u64 = 2000; -/// Directories to skip when watching -const SKIP_DIRS: &[&str] = &[ - "node_modules", - "target", - "__pycache__", - ".git", - "dist", - "build", - "out", - "vendor", - "coverage", - "logs", -]; - /// Watches workspace directories for file changes and auto-indexes them. pub struct McpFileWatcher { _watcher: RecommendedWatcher, @@ -47,6 +34,9 @@ struct WatcherCtx { parsers: Arc, query_engine: Arc, supported_extensions: Vec, + /// The indexer's exclusion rule, so a file kept out of the initial index + /// cannot be let back in by a later write to it. + filter: WorkspaceFilter, } impl McpFileWatcher { @@ -58,6 +48,7 @@ impl McpFileWatcher { parsers: Arc, query_engine: Arc, directories: &[PathBuf], + config: &IndexConfig, ) -> Result { let (tx, mut rx) = mpsc::channel::(100); @@ -90,6 +81,7 @@ impl McpFileWatcher { parsers, query_engine, supported_extensions, + filter: WorkspaceFilter::new(directories, config), }; tokio::spawn(async move { @@ -104,9 +96,10 @@ impl McpFileWatcher { match event { Some(event) => { for path in &event.paths { - if !is_watchable(path, &ctx.supported_extensions) { + if !is_watchable(path, &ctx.supported_extensions, &ctx.filter) { continue; } + let path = &ctx.filter.indexed_form(path); match event.kind { EventKind::Create(_) | EventKind::Modify(_) => { deleted.remove(path); @@ -149,27 +142,11 @@ impl McpFileWatcher { } /// Check if a path is a supported source file worth watching. -fn is_watchable(path: &Path, supported_extensions: &[String]) -> bool { - // Skip directories - if path.is_dir() { +fn is_watchable(path: &Path, supported_extensions: &[String], filter: &WorkspaceFilter) -> bool { + if path.is_dir() || !filter.admits(path) { return false; } - // Skip paths in excluded directories - let path_str = path.to_string_lossy(); - for skip in SKIP_DIRS { - if path_str.contains(&format!("/{skip}/")) || path_str.contains(&format!("\\{skip}\\")) { - return false; - } - } - - // Skip hidden files - if let Some(name) = path.file_name().and_then(|n| n.to_str()) { - if name.starts_with('.') { - return false; - } - } - // Check extension if let Some(ext) = path.extension().and_then(|e| e.to_str()) { supported_extensions.iter().any(|se| se == ext) @@ -178,163 +155,138 @@ fn is_watchable(path: &Path, supported_extensions: &[String]) -> bool { } } +/// Every node in the graph, grouped by the file it came from. +/// +/// Built once per batch. The previous code ran a full-graph +/// `query().property("path", ..)` scan for each deleted file, each vanished +/// file, each changed file's dependents, each changed file and each dependent, +/// so a burst of generated files multiplied whole-graph scans by the size of +/// the burst (issue #23). One pass answers all of them. +fn nodes_by_path(graph: &CodeGraph) -> HashMap> { + let mut map: HashMap> = HashMap::new(); + for (id, node) in graph.iter_nodes() { + if let Some(path) = node.properties.get_string("path") { + map.entry(path.to_string()).or_default().push(id); + } + } + map +} + +fn path_key(path: &Path) -> String { + path.to_string_lossy().to_string() +} + /// Process accumulated file changes: re-index changed files, remove deleted files. async fn process_changes(ctx: &WatcherCtx, changed: &[PathBuf], removed: &[PathBuf]) { - let total = changed.len() + removed.len(); tracing::info!( "[file-watcher] Processing {} changes ({} modified, {} deleted)", - total, + changed.len() + removed.len(), changed.len(), removed.len() ); - // Handle deleted files - let mut had_deletes = false; - if !removed.is_empty() { - for path in removed { - let path_str = path.to_string_lossy().to_string(); - // Remove vectors before deleting nodes (needs node IDs still in graph) - ctx.query_engine.remove_file_vectors(&path_str).await; - // Remove nodes and connected edges - let mut graph = ctx.graph.write().await; - if let Ok(old_nodes) = graph.query().property("path", path_str.as_str()).execute() { - let count = old_nodes.len(); - for old_id in old_nodes { - let _ = graph.delete_node(old_id); - } - if count > 0 { - had_deletes = true; - tracing::info!( - "[file-watcher] Removed {} nodes for deleted {:?}", - count, - path - ); - } - } - } - } + let by_path = { + let graph = ctx.graph.read().await; + nodes_by_path(&graph) + }; - // Separate changed files into actual changes vs files that were deleted - // (macOS FSEvents sometimes reports deletes as modifications) - let mut actual_changed = Vec::new(); - for path in changed { - if path.exists() { - actual_changed.push(path.clone()); - } else { - // File was reported as modified but doesn't exist — treat as delete - let path_str = path.to_string_lossy().to_string(); - ctx.query_engine.remove_file_vectors(&path_str).await; - let mut graph = ctx.graph.write().await; - if let Ok(old_nodes) = graph.query().property("path", path_str.as_str()).execute() { - let count = old_nodes.len(); - for old_id in old_nodes { - let _ = graph.delete_node(old_id); - } - if count > 0 { - had_deletes = true; - tracing::info!( - "[file-watcher] Removed {} nodes for vanished {:?}", - count, - path - ); - } - } + // macOS FSEvents sometimes reports a delete as a modification, so a + // "changed" path that no longer exists is a delete. + let (changed, vanished): (Vec, Vec) = + changed.iter().cloned().partition(|p| p.exists()); + + let mut had_deletes = false; + for path in removed.iter().chain(vanished.iter()) { + // A file the index never held has nothing to remove. Skipping it here + // keeps a burst of deletes of excluded or generated files from taking + // the write lock at all. + let Some(ids) = by_path.get(&path_key(path)) else { + continue; + }; + let mut graph = ctx.graph.write().await; + for id in ids { + let _ = graph.delete_node(*id); } + had_deletes = true; + tracing::info!( + "[file-watcher] Removed {} nodes for deleted {:?}", + ids.len(), + path + ); } - let changed = &actual_changed; - // Handle changed/new files + // Dependents are found before any changed file is re-parsed, while its old + // nodes - and so their incoming edges - are still in the graph. + let mut dependents: HashSet = HashSet::new(); if !changed.is_empty() { - // Find dependents before deleting old nodes - let mut dependents: HashSet = HashSet::new(); - { - let graph = ctx.graph.read().await; - for path in changed { - let path_str = path.to_string_lossy().to_string(); - if let Ok(file_nodes) = graph.query().property("path", path_str.as_str()).execute() - { - for node_id in &file_nodes { - if let Ok(neighbors) = - graph.get_neighbors(*node_id, codegraph::Direction::Incoming) - { - for neighbor_id in neighbors { - if let Ok(neighbor) = graph.get_node(neighbor_id) { - if let Some(dep_path) = neighbor.properties.get_string("path") { - let dep = PathBuf::from(dep_path); - if !changed.contains(&dep) && dep.exists() { - dependents.insert(dep); - } - } - } + let graph = ctx.graph.read().await; + for path in &changed { + for node_id in by_path.get(&path_key(path)).into_iter().flatten() { + let Ok(neighbors) = graph.get_neighbors(*node_id, codegraph::Direction::Incoming) + else { + continue; + }; + for neighbor_id in neighbors { + if let Ok(neighbor) = graph.get_node(neighbor_id) { + if let Some(dep_path) = neighbor.properties.get_string("path") { + let dep = PathBuf::from(dep_path); + if !changed.contains(&dep) && dep.exists() { + dependents.insert(dep); } } } } } } + } - // Delete old nodes + re-parse changed files - let mut indexed = 0; - for path in changed { - { - let mut graph = ctx.graph.write().await; - let path_str = path.to_string_lossy().to_string(); - if let Ok(old_nodes) = graph.query().property("path", path_str.as_str()).execute() { - for old_id in old_nodes { - let _ = graph.delete_node(old_id); - } - } - } - { - let mut graph = ctx.graph.write().await; - if ctx.parsers.parse_file(path, &mut graph).is_ok() { - indexed += 1; - } - } + // One write lock per file covering both the removal and the re-parse. The + // old code took them separately, leaving a window where a concurrent + // search saw the file's symbols missing entirely. + let mut indexed = 0; + for path in changed.iter().chain(dependents.iter()) { + let mut graph = ctx.graph.write().await; + for id in by_path.get(&path_key(path)).into_iter().flatten() { + let _ = graph.delete_node(*id); } - - // Re-parse dependents - if !dependents.is_empty() { - tracing::info!("[file-watcher] Re-indexing {} dependents", dependents.len()); - for dep in &dependents { - { - let mut graph = ctx.graph.write().await; - let path_str = dep.to_string_lossy().to_string(); - if let Ok(old_nodes) = - graph.query().property("path", path_str.as_str()).execute() - { - for old_id in old_nodes { - let _ = graph.delete_node(old_id); - } - } - } - { - let mut graph = ctx.graph.write().await; - if ctx.parsers.parse_file(dep, &mut graph).is_ok() { - indexed += 1; - } - } - } + if ctx.parsers.parse_file(path, &mut graph).is_ok() { + indexed += 1; } + } + if !dependents.is_empty() { + tracing::info!("[file-watcher] Re-indexed {} dependents", dependents.len()); + } - if indexed > 0 || had_deletes { - // Resolve cross-file imports - { - let mut graph = ctx.graph.write().await; - GraphUpdater::resolve_cross_file_imports(&mut graph); - } - // Rebuild search indexes - ctx.query_engine.build_indexes().await; - // Incrementally re-embed changed files - for path in changed.iter().chain(dependents.iter()) { - let path_str = path.to_string_lossy().to_string(); - ctx.query_engine.update_file_vectors(&path_str).await; - } - - tracing::info!( - "[file-watcher] Indexed {} files (incl. dependents), indexes rebuilt", - indexed - ); - } + // Deletions change the graph as much as edits do, so they rebuild the + // indexes too. This used to sit inside the changed-files branch, so a batch + // of pure deletions left the text index holding entries for nodes that no + // longer existed until some later edit rebuilt it. Search did not show it - + // results are resolved against the graph, where the nodes were already + // gone - so this is index hygiene, not a visible correctness fix. + if indexed == 0 && !had_deletes { + return; } + { + let mut graph = ctx.graph.write().await; + GraphUpdater::resolve_cross_file_imports(&mut graph); + } + ctx.query_engine.build_indexes().await; + + // Deleting a file and re-parsing one both leave vectors behind for node IDs + // that no longer exist. One prune after all of it covers both; the old + // per-file call ran before the nodes were deleted, so it could not see the + // file it was meant for, and never ran at all after a re-parse. + ctx.query_engine.prune_orphan_vectors().await; + + let reembed: Vec = changed + .iter() + .chain(dependents.iter()) + .map(|p| path_key(p)) + .collect(); + ctx.query_engine.update_files_vectors(&reembed).await; + + tracing::info!( + "[file-watcher] Indexed {} files (incl. dependents), indexes rebuilt", + indexed + ); } diff --git a/crates/codegraph-server/src/mcp/server.rs b/crates/codegraph-server/src/mcp/server.rs index 1ecf815..5e4bc79 100644 --- a/crates/codegraph-server/src/mcp/server.rs +++ b/crates/codegraph-server/src/mcp/server.rs @@ -936,10 +936,17 @@ impl McpBackend { pub async fn index_workspace(&self) -> (usize, usize) { let config = self.index_config(); - // Initialize memory manager for each workspace folder - for folder in &self.workspace_folders { - if let Err(e) = self.memory_manager.initialize(folder).await { - tracing::warn!("Failed to initialize memory manager: {:?}", e); + // Initialising the memory manager loads the embedding model, which is + // exactly what --graph-only promises not to do - the comment on the + // graph-only branch below already said so, and it was not true, because + // this ran first (issue #23). Memory tools are unavailable in graph-only + // mode as a result, which is the same state the RAM gate already leaves + // them in on a low-memory machine. + if !self.graph_only { + for folder in &self.workspace_folders { + if let Err(e) = self.memory_manager.initialize(folder).await { + tracing::warn!("Failed to initialize memory manager: {:?}", e); + } } } @@ -1325,9 +1332,12 @@ impl McpServer { "version": crate::metadata::VERSION, })); - for folder in &self.backend.workspace_folders { - if let Err(e) = self.backend.memory_manager.initialize(folder).await { - tracing::warn!("Failed to initialize memory manager: {:?}", e); + // Same rule as index_workspace: no model load under --graph-only. + if !self.backend.graph_only { + for folder in &self.backend.workspace_folders { + if let Err(e) = self.backend.memory_manager.initialize(folder).await { + tracing::warn!("Failed to initialize memory manager: {:?}", e); + } } } // Build text/caller/callee indexes from the loaded graph (cheap — no @@ -1372,6 +1382,9 @@ impl McpServer { Arc::clone(&self.backend.parsers), Arc::clone(&self.backend.query_engine), &self.backend.workspace_folders, + // The config the initial index was built with, so the watcher + // excludes exactly what the index excluded. + &self.backend.index_config(), ) { Ok(watcher) => { self._file_watcher = Some(watcher); From a1b3c603ba38ecc3beea4e265b8280d49ce2e11f Mon Sep 17 00:00:00 2001 From: Andrey Vasilevsky Date: Wed, 30 Sep 2026 18:32:01 -0700 Subject: [PATCH 2/3] no-mistakes(review): Locate deleted-directory events in symlinked workspaces --- crates/codegraph-server/src/indexer.rs | 36 ++++++++++++++++++++++---- 1 file changed, 31 insertions(+), 5 deletions(-) diff --git a/crates/codegraph-server/src/indexer.rs b/crates/codegraph-server/src/indexer.rs index 41fbb9b..f098e07 100644 --- a/crates/codegraph-server/src/indexer.rs +++ b/crates/codegraph-server/src/indexer.rs @@ -129,11 +129,13 @@ impl WorkspaceFilter { /// opened as `/var/folders/...` or through any symlinked directory is /// reported by FSEvents at its resolved location, `/private/var/folders/...`, /// so a plain prefix test rejected every event and the watcher admitted - /// nothing at all. Both sides are compared resolved and as given. + /// nothing at all. The event path is compared, as reported, against both + /// spellings of the root; that needs no filesystem access, so it still + /// works when the file's directory was deleted along with it. /// - /// The event path is resolved through its parent directory, which still - /// exists when the file itself was just deleted. The most specific root - /// wins when workspace folders nest. + /// A path reported in unresolved form is additionally resolved through its + /// parent directory. The most specific root wins when workspace folders + /// nest. fn locate(&self, path: &Path) -> Option<(PathBuf, PathBuf, &IndexConfig, &globset::GlobSet)> { let resolved = path .parent() @@ -143,7 +145,11 @@ impl WorkspaceFilter { let mut best: Option<(usize, PathBuf, PathBuf, &IndexConfig, &globset::GlobSet)> = None; for (given, canonical, config, exclude_set) in &self.roots { - for (candidate, root) in [(Some(path), given), (resolved.as_deref(), canonical)] { + for (candidate, root) in [ + (Some(path), given), + (Some(path), canonical), + (resolved.as_deref(), canonical), + ] { let Some(candidate) = candidate else { continue }; let Ok(relative) = candidate.strip_prefix(root) else { continue; @@ -867,6 +873,26 @@ mod tests { ); } + /// Deleting a directory removes the parent an event path would be resolved + /// through, so a resolved-form path must be located without touching disk. + #[cfg(unix)] + #[test] + fn workspace_filter_locates_files_in_a_deleted_directory() { + let tmp = tempfile::tempdir().unwrap(); + let real = tmp.path().join("real"); + std::fs::create_dir_all(real.join("src/old")).unwrap(); + std::fs::write(real.join("src/old/a.rs"), "fn a() {}\n").unwrap(); + let link = tmp.path().join("link"); + std::os::unix::fs::symlink(&real, &link).unwrap(); + + let filter = WorkspaceFilter::new(std::slice::from_ref(&link), &filter_config(&[], 20)); + let deleted = real.canonicalize().unwrap().join("src/old/a.rs"); + std::fs::remove_dir_all(real.join("src/old")).unwrap(); + + assert!(filter.admits(&deleted)); + assert_eq!(filter.indexed_form(&deleted), link.join("src/old/a.rs")); + } + #[test] fn workspace_filter_prefers_the_most_specific_root() { let tmp = tempfile::tempdir().unwrap(); From d25203b9fd74c07cfdf9139092d26de3294cc547 Mon Sep 17 00:00:00 2001 From: Andrey Vasilevsky Date: Wed, 30 Sep 2026 18:42:54 -0700 Subject: [PATCH 3/3] no-mistakes(document): Document watcher ignore rules and graph-only memory tools --- README.md | 2 +- docs/tool-calling-guide.md | 3 ++- 2 files changed, 3 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 343a3fb..bbcb70b 100644 --- a/README.md +++ b/README.md @@ -97,7 +97,7 @@ one tool and exits without the MCP stdio handshake — ideal for scripting. | `--full-body-embedding` | `true` | Embed full function body (~50 lines) for better semantic search and duplicate detection | | `--max-files ` | 5000 | Maximum files to index | | `--profile ` | `all` | Filter the exposed MCP tool surface to a named subset (see below) | -| `--graph-only` | off | Skip embedding generation — build the graph and serve structural tools only. No ONNX model load, 10-50× faster indexing. Semantic search unavailable. For CI / one-shot graph queries. | +| `--graph-only` | off | Skip embedding generation — build the graph and serve structural tools only. No ONNX model load, 10-50× faster indexing. Semantic search and memory tools unavailable. For CI / one-shot graph queries. | | `--run-tool ` | — | One-shot mode: index, run a single tool, print its result, exit. No MCP handshake. Pair with `--tool-args ''`. | #### `--embedding-model static` — model2vec fast indexing diff --git a/docs/tool-calling-guide.md b/docs/tool-calling-guide.md index b125b95..8b77f01 100644 --- a/docs/tool-calling-guide.md +++ b/docs/tool-calling-guide.md @@ -589,7 +589,7 @@ Add an entire directory alongside existing data. ### Workspace exclusion: `.codegraphignore` + default skip list -`reindex_workspace` and `index_directory` honor a per-folder `.codegraphignore` file (gitignore-like syntax — one pattern per line, `#` comments, blank lines ignored; no `!` negation in v1) plus a built-in skip list for binary archives, compiled artifacts, OS metadata, and bulky non-source: +Initial indexing, `reindex_workspace`, `index_directory` and the file watcher honor a per-folder `.codegraphignore` file (gitignore-like syntax — one pattern per line, `#` comments, blank lines ignored; no `!` negation in v1) plus a built-in skip list for binary archives, compiled artifacts, OS metadata, and bulky non-source: - Archives: `**/*.tar.gz`, `**/*.tar.bz2`, `**/*.tar.xz`, `**/*.tgz`, `**/*.tbz2`, `**/*.zip`, `**/*.7z`, `**/*.rar`, `**/*.deb`, `**/*.rpm`, `**/*.pkg`, `**/*.dmg`, `**/*.iso`, `**/*.img` - Binaries: `**/*.exe`, `**/*.dll`, `**/*.so`, `**/*.dylib`, `**/*.bin`, `**/*.o`, `**/*.a`, `**/*.lib`, `**/*.obj`, `**/*.pdb`, `**/*.pyc`, `**/*.class`, `**/*.jar` @@ -599,6 +599,7 @@ Add an entire directory alongside existing data. - Misc: `**/*.sqlite`, `**/*.db`, `**/*.lock` Prevents fastembed/ONNX runaway on workspaces containing proof bundles, cloned upstream targets, prebuilt binaries, or triage doc folders. Per-folder so each workspace can have its own rules. +The file watcher reads `.codegraphignore` once at startup, so edits to it (like changes to `--exclude`) take effect on the next start. Example `.codegraphignore` for a bounty workspace: