diff --git a/README.de.md b/README.de.md index 53e026e..4f3af6f 100644 --- a/README.de.md +++ b/README.de.md @@ -9,6 +9,9 @@ Ein gemeinsamer Skill, eine CLI, gemeinsame Indizes. tgrep sucht Text und Code; Qwen übersetzt Fragen in Suchbegriffe und beantwortet sie anhand gefundener Ausschnitte. Die Hauptmodelle von Claude und Codex werden dabei nicht ersetzt. +Gemessene Performance: Der [lokale Benchmarkbericht](docs/benchmarks/README.md) +zeigt reproduzierbare CLI-Zeiten, Indexkosten und die Grenzen der synthetischen Tests. + ## Installation Voraussetzungen: macOS oder Linux (ARM64/x86_64), Git, Python **3.10+** mit venv. diff --git a/README.md b/README.md index 5ce34f5..a8f0555 100644 --- a/README.md +++ b/README.md @@ -9,6 +9,9 @@ One shared skill, one CLI, shared indexes. tgrep searches text and code; Qwen turns questions into search terms and answers them using matching excerpts. It does not replace the main models used by Claude and Codex. +Measured performance: see the [local benchmark report](docs/benchmarks/README.md) +for reproducible CLI timings, index costs, and the limits of the synthetic tests. + ## Installation Requirements: macOS or Linux (ARM64/x86_64), Git, Python **3.10+** with venv. diff --git a/docs/benchmarks/README.md b/docs/benchmarks/README.md new file mode 100644 index 0000000..8ad6bf7 --- /dev/null +++ b/docs/benchmarks/README.md @@ -0,0 +1,112 @@ +# Local performance measurements + +These measurements compare the complete skill CLI using a running tgrep index +with the same CLI using ripgrep through `--fresh`. They include Python startup, +backend selection, subprocess execution, JSON handling, and output capture. +They are not direct tgrep-versus-ripgrep binary timings. + +## Environment and method + +- Apple M4, 10 logical CPUs, 16 GiB RAM; macOS 26.5.2, Python 3.10.5. +- tgrep 1.0.5 and ripgrep 15.2.0 from the installer's pinned releases. +- Deterministic synthetic Python files in temporary directories, on local storage. + Files contain generated functions and known search markers. Their vocabulary is + deliberately simple and repetitive, which favors a compact trigram index. +- Two warmup searches per backend and query, then nine measured searches each. + Each pair runs sequentially in a seeded randomized backend order. +- Warm filesystem caches: neither OS caches nor storage caches are flushed. + The index build also reads the corpus. This measures repeated interactive search, + not cold disk performance. The machine is not a dedicated benchmark host. +- Identical root policy and literal/regex flags; maximum 1,000 returned matches. + Every query is checked against its known match count and exact sorted + `(path, line, text)` equality across both backends and every repetition. + Truncation, warnings, or fallback from tgrep abort the benchmark. +- No Ollama or Qwen requests. These numbers measure retrieval, not AI answer latency. + +Raw timings, corpus hashes, tool versions, source commit, and benchmark script hash +are recorded in [the JSON report](macos-arm64-2026-09-09.json). + +## Results + +Median wall time in milliseconds. Speedup is fresh / indexed; below 1× means +indexed search was slower. Every row passed exact match parity. + +| Files | Query | Indexed ms | Fresh ms | Speedup | +|---:|---|---:|---:|---:| +| 1,000 | Rare literal (1 match) | 44.0 | 49.7 | 1.13× | +| 1,000 | Distributed literal (1% of files) | 44.4 | 49.0 | 1.10× | +| 1,000 | Absent long literal | 44.2 | 50.0 | 1.13× | +| 1,000 | Selective regex (1 match) | 44.0 | 49.9 | 1.13× | +| 1,000 | Absent two-character literal | 45.0 | 50.1 | 1.11× | +| 20,000 | Rare literal (1 match) | 45.5 | 250.5 | 5.51× | +| 20,000 | Distributed literal (1% of files) | 53.2 | 262.7 | 4.94× | +| 20,000 | Absent long literal | 46.0 | 254.3 | 5.53× | +| 20,000 | Selective regex (1 match) | 45.9 | 236.6 | 5.15× | +| 20,000 | Absent two-character literal | 62.8 | 239.2 | 3.81× | +| 100,000 | Rare literal (1 match) | 51.9 | 1262.8 | 24.32× | +| 100,000 | Distributed literal (1% of files) | 89.7 | 1284.6 | 14.32× | +| 100,000 | Absent long literal | 51.5 | 1253.9 | 24.33× | +| 100,000 | Selective regex (1 match) | 51.8 | 1252.2 | 24.17× | +| 100,000 | Absent two-character literal | 3637.9 | 1262.0 | 0.35× | + +| Files | Source MiB | Build + start (s) | Index directory MiB | Server RSS MiB | +|---:|---:|---:|---:|---:| +| 1,000 | 3.0 | 0.136 | 2.4 | 24.0 | +| 20,000 | 63.2 | 0.685 | 47.4 | 108.9 | +| 100,000 | 323.2 | 4.245 | 237.8 | 504.2 | + +For the rare-literal query, estimated build break-even is 24 queries at 1,000 +files and 4 queries at both larger sizes. The two-character query at 100,000 files +has no break-even in this run: indexed search was about 2.88× slower than fresh. +For such queries, `--fresh` is worth comparing rather than assuming the index wins. + +The report contains 270 timed CLI invocations and 60 warmup invocations. +A preliminary run showed noticeably different timings, including better indexed +performance on the two-character query. The complete final run is reported here, +not a selection of the best samples across runs. Run-to-run variability has not +been quantified; the results do not establish its cause. + +## Interpretation and limits + +The median is the primary statistic. The JSON also records every sample and a +nearest-rank p95; with only nine measured samples, that p95 equals the maximum. +It is not a reliable estimate of production tail latency. + +Index build plus server startup is timed once per corpus, excluding generation +of the test files. Index directory size is logical file bytes, including metadata +and logs, not allocated disk blocks. Server RSS is sampled after the queries; +it is not peak memory, total system memory, or a process memory limit. + +The break-even estimate divides build-and-start time by the median per-query time +saving. It assumes repetition of that query, an unchanged corpus, and no further +index maintenance. It does not account for background CPU, power, storage costs, +or a changing working tree. + +Do not present these numbers as Chromium/gecko-dev results or a universal speedup. +Real repositories have richer trigram vocabularies, varying file sizes, ignores, +binary content, and update activity. This suite also does not cover broad queries +whose output exceeds the limit, filesystem watcher latency, cold server restarts, +or answer quality. A real-repository benchmark and a separate Qwen latency and +answer-quality evaluation are still needed for those claims. + +## Reproduce + +Install the project first, then pass the installed binaries explicitly. Their +runtime directory is recorded in `~/.local/share/tgrep-ai-skill/install.json`. +Run from the repository checkout: + +~~~sh +python3 scripts/benchmark.py \ + --tgrep /path/to/runtime/bin/tgrep \ + --rg /path/to/runtime/bin/rg \ + --sizes 1000 20000 100000 \ + --rounds 9 --warmups 2 \ + --output benchmark-results.json +~~~ + +Python 3.10+ and local loopback/process access are required. No extra Python +packages or model downloads are needed. The largest corpus contains roughly +323 MiB of source text, plus its index and temporary build files. Each corpus +is removed before the next starts. User configuration and existing indexes are +not changed. If a server cannot be stopped, its directory is preserved and +reported for diagnosis. diff --git a/docs/benchmarks/macos-arm64-2026-09-09.json b/docs/benchmarks/macos-arm64-2026-09-09.json new file mode 100644 index 0000000..d94379a --- /dev/null +++ b/docs/benchmarks/macos-arm64-2026-09-09.json @@ -0,0 +1,694 @@ +{ + "schema": 1, + "utc": "2026-09-09T13:52:19Z", + "platform": "macOS-26.5.2-arm64-arm-64bit", + "machine": "arm64", + "python": "3.10.5", + "logical_cpus": 10, + "cpu_model": "Apple M4", + "memory_bytes": 17179869184, + "source_commit": "da308252e0d79cfce222e4432f3615e4a317479f", + "benchmark_sha256": "384f45ab3350e8c426332e333a1b457cb6372c4aeb381d417d86ee737d7f8a24", + "tgrep_version": "tgrep 1.0.5", + "rg_version": "ripgrep 15.2.0 (rev e89fff89ac)\n\nfeatures:+pcre2\nsimd(compile):+NEON\nsimd(runtime):+NEON\n\nPCRE2 10.45 is available (JIT is available)", + "rounds": 9, + "warmups": 2, + "method": "Full CLI wall time; paired randomized order; warm OS caches; synthetic Python; identical filters; exact uncapped match parity; no LLM; RSS is a point sample, not a peak; index build is one run", + "corpora": [ + { + "files": 1000, + "bytes": 3131016, + "content_sha256": "1b1c0ec839759d3091b62daa13aef7374bfe6802a579988a7e80430932c720d2", + "build_and_start_seconds": 0.13603562500793487, + "index_directory_bytes": 2508252, + "server_rss_kib_after_queries": 24560, + "queries": [ + { + "query": "rare_literal", + "pattern": "unique_benchmark_needle", + "regex": false, + "matches": 1, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 43.61937500652857, + 44.31179098901339, + 43.95524997380562, + 44.13837500032969, + 43.9454999868758, + 44.36187501414679, + 44.33654199237935, + 42.993459006538615, + 43.78887498751283 + ], + "median_ms": 43.95524997380562, + "p95_ms": 44.36187501414679, + "min_ms": 42.993459006538615, + "max_ms": 44.36187501414679 + }, + "fresh": { + "samples_ms": [ + 49.94804199668579, + 51.16362500120886, + 47.71041599451564, + 50.432207994163036, + 49.47512500802986, + 49.03329099761322, + 49.73608400905505, + 50.42800001683645, + 48.66450000554323 + ], + "median_ms": 49.73608400905505, + "p95_ms": 51.16362500120886, + "min_ms": 47.71041599451564, + "max_ms": 51.16362500120886 + }, + "speedup": 1.131516349894368, + "build_break_even_queries": 24 + }, + { + "query": "distributed_literal", + "pattern": "batch_benchmark_marker", + "regex": false, + "matches": 10, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 44.348625000566244, + 44.457541982410476, + 44.44141700514592, + 44.35670800739899, + 44.05324999243021, + 45.027125015622005, + 45.141874987166375, + 44.143374980194494, + 44.272375002037734 + ], + "median_ms": 44.35670800739899, + "p95_ms": 45.141874987166375, + "min_ms": 44.05324999243021, + "max_ms": 45.141874987166375 + }, + "fresh": { + "samples_ms": [ + 50.79174999264069, + 51.26354200183414, + 47.867999994196, + 48.478584008989856, + 49.43908299901523, + 48.94704098114744, + 49.00087497662753, + 48.12095800298266, + 49.00637498940341 + ], + "median_ms": 49.00087497662753, + "p95_ms": 51.26354200183414, + "min_ms": 47.867999994196, + "max_ms": 51.26354200183414 + }, + "speedup": 1.104700442793317, + "build_break_even_queries": 30 + }, + { + "query": "absent_literal", + "pattern": "absent_benchmark_xyz987", + "regex": false, + "matches": 0, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 44.79016698314808, + 43.83229202358052, + 44.86550000729039, + 44.89724998711608, + 43.590999994194135, + 44.236082991119474, + 43.968166020931676, + 44.319958018604666, + 43.86754200095311 + ], + "median_ms": 44.236082991119474, + "p95_ms": 44.89724998711608, + "min_ms": 43.590999994194135, + "max_ms": 44.89724998711608 + }, + "fresh": { + "samples_ms": [ + 51.03270799736492, + 50.571750005474314, + 50.18066699267365, + 49.4932500005234, + 50.028375000692904, + 48.63591701723635, + 48.55274999863468, + 48.702333006076515, + 50.2279170032125 + ], + "median_ms": 50.028375000692904, + "p95_ms": 51.03270799736492, + "min_ms": 48.55274999863468, + "max_ms": 51.03270799736492 + }, + "speedup": 1.130940436356814, + "build_break_even_queries": 24 + }, + { + "query": "selective_regex", + "pattern": "unique_benchmark_[a-z]+", + "regex": true, + "matches": 1, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 44.6396250044927, + 43.618665979010984, + 44.53583399299532, + 44.004458002746105, + 44.15137498290278, + 44.25504201208241, + 43.58483399846591, + 43.54341700673103, + 44.02850000769831 + ], + "median_ms": 44.02850000769831, + "p95_ms": 44.6396250044927, + "min_ms": 43.54341700673103, + "max_ms": 44.6396250044927 + }, + "fresh": { + "samples_ms": [ + 49.173207982676104, + 50.38229198544286, + 48.61900000832975, + 48.50145798991434, + 49.86116700456478, + 50.28699999093078, + 50.19258300308138, + 50.690208008745685, + 48.744708008598536 + ], + "median_ms": 49.86116700456478, + "p95_ms": 50.690208008745685, + "min_ms": 48.50145798991434, + "max_ms": 50.690208008745685 + }, + "speedup": 1.1324748059971752, + "build_break_even_queries": 24 + }, + { + "query": "short_absent_literal", + "pattern": "ZQ", + "regex": false, + "matches": 0, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 44.28195900982246, + 45.14904198003933, + 44.985749991610646, + 44.815750006819144, + 44.532666011946276, + 47.29845799738541, + 45.33254200941883, + 45.52095799590461, + 44.8750409996137 + ], + "median_ms": 44.985749991610646, + "p95_ms": 47.29845799738541, + "min_ms": 44.28195900982246, + "max_ms": 47.29845799738541 + }, + "fresh": { + "samples_ms": [ + 49.157625006046146, + 49.074084003223106, + 50.866250006947666, + 49.54208401613869, + 50.290584011236206, + 48.762082995381206, + 50.253374996827915, + 50.214750022860244, + 50.08729099063203 + ], + "median_ms": 50.08729099063203, + "p95_ms": 50.866250006947666, + "min_ms": 48.762082995381206, + "max_ms": 50.866250006947666 + }, + "speedup": 1.113403488881985, + "build_break_even_queries": 27 + } + ] + }, + { + "files": 20000, + "bytes": 66311766, + "content_sha256": "24656adaa6852d218976d9e03a5ccad22e7873ee1898a44a505ae57f77a0711b", + "build_and_start_seconds": 0.6846851670125034, + "index_directory_bytes": 49737999, + "server_rss_kib_after_queries": 111536, + "queries": [ + { + "query": "rare_literal", + "pattern": "unique_benchmark_needle", + "regex": false, + "matches": 1, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 45.78887499519624, + 45.71445900364779, + 44.09775001113303, + 45.378749986412004, + 44.348958996124566, + 45.47354200622067, + 46.32125000352971, + 44.827875011833385, + 46.41525002080016 + ], + "median_ms": 45.47354200622067, + "p95_ms": 46.41525002080016, + "min_ms": 44.09775001113303, + "max_ms": 46.41525002080016 + }, + "fresh": { + "samples_ms": [ + 259.78191700414754, + 236.13649999606423, + 245.58208300732076, + 249.6871660114266, + 250.00729100429453, + 252.05895799444988, + 250.4555000050459, + 252.94183299411088, + 275.92891702079214 + ], + "median_ms": 250.4555000050459, + "p95_ms": 275.92891702079214, + "min_ms": 236.13649999606423, + "max_ms": 275.92891702079214 + }, + "speedup": 5.5077191913219385, + "build_break_even_queries": 4 + }, + { + "query": "distributed_literal", + "pattern": "batch_benchmark_marker", + "regex": false, + "matches": 200, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 53.4849590039812, + 52.302207972388715, + 53.51008297293447, + 53.13770900829695, + 53.515125007834285, + 53.90387499937788, + 53.103750018635765, + 53.02570800995454, + 53.203499992378056 + ], + "median_ms": 53.203499992378056, + "p95_ms": 53.90387499937788, + "min_ms": 52.302207972388715, + "max_ms": 53.90387499937788 + }, + "fresh": { + "samples_ms": [ + 250.48620800953358, + 262.7261670131702, + 263.6788750241976, + 247.8203329956159, + 277.5383749976754, + 240.2544590004254, + 249.16170799406245, + 270.4827080015093, + 271.8161670200061 + ], + "median_ms": 262.7261670131702, + "p95_ms": 277.5383749976754, + "min_ms": 240.2544590004254, + "max_ms": 277.5383749976754 + }, + "speedup": 4.938136909241091, + "build_break_even_queries": 4 + }, + { + "query": "absent_literal", + "pattern": "absent_benchmark_xyz987", + "regex": false, + "matches": 0, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 44.159957993542776, + 47.39529200014658, + 47.84866701811552, + 45.57337501319125, + 46.56354198232293, + 45.963167009176686, + 46.02929198881611, + 44.34749999199994, + 45.48666600021534 + ], + "median_ms": 45.963167009176686, + "p95_ms": 47.84866701811552, + "min_ms": 44.159957993542776, + "max_ms": 47.84866701811552 + }, + "fresh": { + "samples_ms": [ + 256.36966701131314, + 227.4162499816157, + 242.26879101479426, + 245.46620799810626, + 268.4037090220954, + 264.9954580119811, + 254.34612500248477, + 257.285250001587, + 253.62666699220426 + ], + "median_ms": 254.34612500248477, + "p95_ms": 268.4037090220954, + "min_ms": 227.4162499816157, + "max_ms": 268.4037090220954 + }, + "speedup": 5.533694511339999, + "build_break_even_queries": 4 + }, + { + "query": "selective_regex", + "pattern": "unique_benchmark_[a-z]+", + "regex": true, + "matches": 1, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 46.00400000344962, + 47.031791997142136, + 60.57037500431761, + 45.73762498330325, + 45.81824998604134, + 46.00641698925756, + 45.28516699792817, + 45.94125002040528, + 44.34270798810758 + ], + "median_ms": 45.94125002040528, + "p95_ms": 60.57037500431761, + "min_ms": 44.34270798810758, + "max_ms": 60.57037500431761 + }, + "fresh": { + "samples_ms": [ + 236.63554099039175, + 234.45345801883377, + 231.19745799340308, + 244.03829101356678, + 266.61883300403133, + 250.95816698740236, + 230.55508398101665, + 222.56295799161308, + 237.75845798081718 + ], + "median_ms": 236.63554099039175, + "p95_ms": 266.61883300403133, + "min_ms": 222.56295799161308, + "max_ms": 266.61883300403133 + }, + "speedup": 5.150829393742827, + "build_break_even_queries": 4 + }, + { + "query": "short_absent_literal", + "pattern": "ZQ", + "regex": false, + "matches": 0, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 62.238290993263945, + 61.93695799447596, + 61.6332910140045, + 60.99708398687653, + 62.79262498719618, + 87.44116697926074, + 67.36795799224637, + 67.13825001497753, + 67.49579100869596 + ], + "median_ms": 62.79262498719618, + "p95_ms": 87.44116697926074, + "min_ms": 60.99708398687653, + "max_ms": 87.44116697926074 + }, + "fresh": { + "samples_ms": [ + 223.0282919772435, + 235.64808300579898, + 229.89424999104813, + 256.0791669820901, + 249.86866599647328, + 239.1713339893613, + 246.4875830046367, + 237.04770900076255, + 247.23533299402334 + ], + "median_ms": 239.1713339893613, + "p95_ms": 256.0791669820901, + "min_ms": 223.0282919772435, + "max_ms": 256.0791669820901 + }, + "speedup": 3.8089080371800015, + "build_break_even_queries": 4 + } + ] + }, + { + "files": 100000, + "bytes": 338891766, + "content_sha256": "5fbbe8ac70e440c600f5d0e436221ac91b59c632142806fa781bd5723acaaed3", + "build_and_start_seconds": 4.24470595800085, + "index_directory_bytes": 249299483, + "server_rss_kib_after_queries": 516288, + "queries": [ + { + "query": "rare_literal", + "pattern": "unique_benchmark_needle", + "regex": false, + "matches": 1, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 52.986708004027605, + 52.21820899168961, + 47.65658298856579, + 51.93000001600012, + 48.175208998145536, + 51.420874980976805, + 52.188791014486924, + 49.957500013988465, + 52.81466600717977 + ], + "median_ms": 51.93000001600012, + "p95_ms": 52.986708004027605, + "min_ms": 47.65658298856579, + "max_ms": 52.986708004027605 + }, + "fresh": { + "samples_ms": [ + 1267.4464579904452, + 1250.1287919876631, + 1269.8567090264987, + 1244.8364160081837, + 1226.5618329984136, + 1262.7933750045486, + 1264.028707984835, + 1278.853291994892, + 1256.525833014166 + ], + "median_ms": 1262.7933750045486, + "p95_ms": 1278.853291994892, + "min_ms": 1226.5618329984136, + "max_ms": 1278.853291994892 + }, + "speedup": 24.31722269623474, + "build_break_even_queries": 4 + }, + { + "query": "distributed_literal", + "pattern": "batch_benchmark_marker", + "regex": false, + "matches": 1000, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 89.71654099877924, + 85.02966701053083, + 88.51554201100953, + 91.63329200237058, + 90.64820801722817, + 91.06499998597428, + 86.064000002807, + 89.87129098386504, + 89.1526660125237 + ], + "median_ms": 89.71654099877924, + "p95_ms": 91.63329200237058, + "min_ms": 85.02966701053083, + "max_ms": 91.63329200237058 + }, + "fresh": { + "samples_ms": [ + 1281.3319999841042, + 1284.5738749892917, + 1278.7110829958692, + 1278.9893750159536, + 1310.8776670123916, + 1328.5628340090625, + 1291.7908749950584, + 1282.07966600894, + 1302.044209005544 + ], + "median_ms": 1284.5738749892917, + "p95_ms": 1328.5628340090625, + "min_ms": 1278.7110829958692, + "max_ms": 1328.5628340090625 + }, + "speedup": 14.318138669732829, + "build_break_even_queries": 4 + }, + { + "query": "absent_literal", + "pattern": "absent_benchmark_xyz987", + "regex": false, + "matches": 0, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 47.1354580076877, + 51.449541992042214, + 54.270499997073784, + 51.953833986772224, + 51.542542001698166, + 53.11279100715183, + 51.19249998824671, + 46.72704200493172, + 52.20345899579115 + ], + "median_ms": 51.542542001698166, + "p95_ms": 54.270499997073784, + "min_ms": 46.72704200493172, + "max_ms": 54.270499997073784 + }, + "fresh": { + "samples_ms": [ + 1248.3060410013422, + 1253.853500005789, + 1248.6065829871222, + 1257.0107500068843, + 1265.3977500158362, + 1232.0264170120936, + 1260.5766249762382, + 1245.083749992773, + 1350.622707977891 + ], + "median_ms": 1253.853500005789, + "p95_ms": 1350.622707977891, + "min_ms": 1232.0264170120936, + "max_ms": 1350.622707977891 + }, + "speedup": 24.326574734410237, + "build_break_even_queries": 4 + }, + { + "query": "selective_regex", + "pattern": "unique_benchmark_[a-z]+", + "regex": true, + "matches": 1, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 51.999042014358565, + 52.205084008164704, + 50.98245799308643, + 50.01950002042577, + 52.03399999300018, + 52.279625000664964, + 46.93654100992717, + 51.81987499236129, + 47.0645840105135 + ], + "median_ms": 51.81987499236129, + "p95_ms": 52.279625000664964, + "min_ms": 46.93654100992717, + "max_ms": 52.279625000664964 + }, + "fresh": { + "samples_ms": [ + 1224.8768330027815, + 1263.8722909905482, + 1279.6648340008687, + 1251.7542080022395, + 1244.3842499924358, + 1273.0346659955103, + 1252.243124996312, + 1250.611333991401, + 1252.7479999989737 + ], + "median_ms": 1252.243124996312, + "p95_ms": 1279.6648340008687, + "min_ms": 1224.8768330027815, + "max_ms": 1279.6648340008687 + }, + "speedup": 24.165305786262582, + "build_break_even_queries": 4 + }, + { + "query": "short_absent_literal", + "pattern": "ZQ", + "regex": false, + "matches": 0, + "exact_match_parity": true, + "indexed": { + "samples_ms": [ + 3813.451374997385, + 3136.63400002406, + 3656.280249997508, + 3087.3178749752697, + 3637.896333995741, + 3078.818250010954, + 3662.4065420182887, + 3077.5334170029964, + 3688.4330000029877 + ], + "median_ms": 3637.896333995741, + "p95_ms": 3813.451374997385, + "min_ms": 3077.5334170029964, + "max_ms": 3813.451374997385 + }, + "fresh": { + "samples_ms": [ + 1278.5090419929475, + 1295.0507079949602, + 1287.7029590017628, + 1185.7254579954315, + 1262.0302079885732, + 1154.1498750157189, + 1374.8570830211975, + 1228.1200000143144, + 1250.1640419941396 + ], + "median_ms": 1262.0302079885732, + "p95_ms": 1374.8570830211975, + "min_ms": 1154.1498750157189, + "max_ms": 1374.8570830211975 + }, + "speedup": 0.346912086580104, + "build_break_even_queries": null + } + ] + } + ] +} diff --git a/scripts/benchmark.py b/scripts/benchmark.py new file mode 100644 index 0000000..5fc48a8 --- /dev/null +++ b/scripts/benchmark.py @@ -0,0 +1,253 @@ +#!/usr/bin/env python3 +"""Measure full CLI indexed/fresh latency on deterministic synthetic corpora.""" + +import argparse +import hashlib +import json +import math +import os +from pathlib import Path +import platform +import random +import shutil +import statistics +import subprocess +import sys +import tempfile +import time +from typing import Any + +SOURCE = Path(__file__).resolve().parents[1] +sys.path.insert(0, str(SOURCE / "src")) +from local_search import config, engine # noqa: E402 + +QUERIES = [ + ("rare_literal", "unique_benchmark_needle", False), + ("distributed_literal", "batch_benchmark_marker", False), + ("absent_literal", "absent_benchmark_xyz987", False), + ("selective_regex", "unique_benchmark_[a-z]+", True), + ("short_absent_literal", "ZQ", False), +] + + +def run(command: list[str], env: dict[str, str] | None = None) -> str: + return subprocess.run( + command, capture_output=True, text=True, env=env, timeout=120, check=True + ).stdout.strip() + + +def corpus(path: Path, count: int) -> dict[str, Any]: + """Create roughly 4 KiB per file, with known sparse and distributed matches.""" + path.mkdir() + digest = hashlib.sha256() + total = 0 + for number in range(count): + folder = path / f"package_{number // 100:05d}" + folder.mkdir(exist_ok=True) + lines = [f"# synthetic module {number}\n"] + if number == count // 2: + lines.append("# unique_benchmark_needle\n") + if number % 100 == 0: + lines.append("# batch_benchmark_marker\n") + for item in range(64): + lines.append( + f"def function_{number}_{item}(value): " + f"return value + {number * 64 + item}\n" + ) + data = "".join(lines).encode() + (folder / f"module_{number:06d}.py").write_bytes(data) + digest.update(data) + total += len(data) + return {"files": count, "bytes": total, "content_sha256": digest.hexdigest()} + + +def measure( + env: dict[str, str], pattern: str, regex: bool, fresh: bool, +) -> tuple[float, list[tuple[str, int, str]]]: + command = [ + sys.executable, "-m", "local_search.cli", "search", pattern, + "--root", "benchmark", "--limit", "1000", + ] + if regex: + command.append("--regex") + if fresh: + command.append("--fresh") + started = time.perf_counter() + result = subprocess.run( + command, env=env, capture_output=True, text=True, timeout=120 + ) + elapsed = (time.perf_counter() - started) * 1000 + if result.returncode not in (0, 1): + raise RuntimeError(f"Search failed: {result.stderr} {result.stdout}") + response = json.loads(result.stdout) + if response.get("truncated"): + raise RuntimeError("Truncated output cannot establish exact match parity") + reports = response["reports"] + expected = "rg" if fresh else "tgrep" + for report in reports: + if report["backend"] != expected or report["warnings"]: + raise RuntimeError(f"Unexpected backend or warning: {report}") + matches = sorted( + (item["path"], item["line"], item["text"]) + for item in response["matches"] + ) + return elapsed, matches + + +def summarize(samples: list[float]) -> dict[str, Any]: + ordered = sorted(samples) + return { + "samples_ms": samples, + "median_ms": statistics.median(samples), + "p95_ms": ordered[math.ceil(len(ordered) * 0.95) - 1], + "min_ms": min(samples), + "max_ms": max(samples), + } + + +def benchmark( + count: int, rounds: int, warmups: int, binaries: dict[str, str], +) -> dict[str, Any]: + base = Path(tempfile.mkdtemp(prefix="local-search-benchmark-")) + safe_to_remove = True + try: + metadata = corpus(base / "corpus", count) + env = { + **os.environ, + "XDG_CONFIG_HOME": str(base / "config"), + "XDG_DATA_HOME": str(base / "data"), + "PYTHONPATH": str(SOURCE / "src"), + **binaries, + } + previous = os.environ.copy() + os.environ.update(env) + root = None + try: + root = config.add_root("benchmark", base / "corpus") + started = time.perf_counter() + server = engine.start(root, rebuild=True) + deadline = time.monotonic() + 120 + while ( + server.get("indexing", True) + or server.get("reconcile_running") + or server.get("reconcile_pending") + ): + if time.monotonic() > deadline: + raise RuntimeError("Index did not become ready") + time.sleep(0.1) + server = engine.rpc_status(config.index_dir(root)) or {} + build_seconds = time.perf_counter() - started + if server["num_files"] != count: + raise RuntimeError("Unexpected indexed file count") + results = [] + rng = random.Random(20260909) + for name, pattern, regex in QUERIES: + timings = {False: [], True: []} + reference = None + for iteration in range(warmups + rounds): + order = [False, True] + rng.shuffle(order) + for fresh in order: + elapsed, matches = measure(env, pattern, regex, fresh) + if reference is None: + reference = matches + if reference != matches: + raise RuntimeError(f"Match mismatch for {name}") + if iteration >= warmups: + timings[fresh].append(elapsed) + expected_count = ( + 1 if name in ("rare_literal", "selective_regex") + else (count + 99) // 100 if name == "distributed_literal" + else 0 + ) + if len(reference) != expected_count: + raise RuntimeError(f"Unexpected match count for {name}") + indexed = summarize(timings[False]) + fresh = summarize(timings[True]) + saving = (fresh["median_ms"] - indexed["median_ms"]) / 1000 + results.append({ + "query": name, "pattern": pattern, "regex": regex, + "matches": len(reference), "exact_match_parity": True, + "indexed": indexed, "fresh": fresh, + "speedup": fresh["median_ms"] / indexed["median_ms"], + "build_break_even_queries": ( + math.ceil(build_seconds / saving) if saving > 0 else None + ), + }) + print( + f"{count} files / {name}: " + f"{indexed['median_ms']:.1f} vs {fresh['median_ms']:.1f} ms", + file=sys.stderr, flush=True, + ) + directory = config.index_dir(root) + index_bytes = sum(p.stat().st_size for p in directory.rglob("*") + if p.is_file()) + rss_kib = int(run(["ps", "-p", str(server["pid"]), "-o", "rss="])) + return { + **metadata, "build_and_start_seconds": build_seconds, + "index_directory_bytes": index_bytes, + "server_rss_kib_after_queries": rss_kib, "queries": results, + } + finally: + try: + if root is not None: + safe_to_remove = False + try: + engine.stop(root) + except BaseException: + print(f"Stop failed; preserving diagnostics: {base}", + file=sys.stderr) + raise + safe_to_remove = True + finally: + os.environ.clear() + os.environ.update(previous) + finally: + if safe_to_remove: + shutil.rmtree(base) + + +def main() -> None: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--tgrep", required=True, type=Path) + parser.add_argument("--rg", required=True, type=Path) + parser.add_argument("--sizes", nargs="+", type=int, default=[1000, 20000]) + parser.add_argument("--rounds", type=int, default=9) + parser.add_argument("--warmups", type=int, default=2) + parser.add_argument("--output", required=True, type=Path) + args = parser.parse_args() + if (args.rounds < 3 or args.warmups < 1 + or any(n < 1 or n > 100000 for n in args.sizes)): + parser.error("Use rounds >= 3, warmups >= 1 and sizes in 1..100000") + binaries = { + "LOCAL_SEARCH_TGREP": str(args.tgrep.resolve(strict=True)), + "LOCAL_SEARCH_RG": str(args.rg.resolve(strict=True)), + } + report = { + "schema": 1, "utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()), + "platform": platform.platform(), "machine": platform.machine(), + "python": platform.python_version(), "logical_cpus": os.cpu_count(), + "cpu_model": run(["sysctl", "-n", "machdep.cpu.brand_string"]) + if sys.platform == "darwin" else platform.processor(), + "memory_bytes": int(run(["sysctl", "-n", "hw.memsize"])) + if sys.platform == "darwin" else None, + "source_commit": run(["git", "-C", str(SOURCE), "rev-parse", "HEAD"]), + "benchmark_sha256": hashlib.sha256(Path(__file__).read_bytes()).hexdigest(), + "tgrep_version": run([binaries["LOCAL_SEARCH_TGREP"], "--version"]), + "rg_version": run([binaries["LOCAL_SEARCH_RG"], "--version"]), + "rounds": args.rounds, "warmups": args.warmups, + "method": "Full CLI wall time; paired randomized order; warm OS caches; " + "synthetic Python; identical filters; exact uncapped match parity; " + "no LLM; RSS is a point sample, not a peak; index build is one run", + "corpora": [], + } + for count in args.sizes: + report["corpora"].append( + benchmark(count, args.rounds, args.warmups, binaries) + ) + args.output.parent.mkdir(parents=True, exist_ok=True) + args.output.write_text(json.dumps(report, indent=2) + "\n") + + +if __name__ == "__main__": + main()