Every number on this page came from a script in bench/ in the Hush
repository, run against the real, live product and committed alongside its output. Nothing here
is estimated, rounded up, or written from memory. Where an axis has not been measured yet, this
page says so instead of guessing. Full methodology and the code itself:
bench/ on GitHub.
Hush's own live CPU and memory cost, sampled at idle from the running
engine's own instrumentation (_self_cost()), the
same number the in-app "Hush itself" line and the self-watchdog use.
| Metric | Min | Median | Mean | Max |
|---|---|---|---|---|
| CPU (% of a core, summed across Hush's whole process tree) | 0.0% | 2.0% | 2.58% | 6.4% |
| Resident memory, summed (MB) | 102.8 | 104.5 | 106.59 | 118.5 |
30 samples, one every 5 seconds, over a 152.4-second idle window (no synthetic
load; ambient state of the machine during the run). Process count stayed at 6–7 across the
window. Measured 2026-07-24 on a MacBook Air · Apple M5, Hush 0.95.1, engine pid 1170.
Raw samples: self_overhead.json.
Reproduce: python3 bench/self_overhead.py.
This is an adversarial self-test, not a representative accuracy figure. We ran the
attribution engine against Hush's OWN processes and a few system edge cases, deliberately the
hardest possible inputs: a Python engine launched under Xcode's bundled Python3 framework, an MCP server
behind Claude Desktop's disclaimer wrapper binary, and a bundle
living under /System/. Ground truth comes from an independent
ps read that shares no code with the engine being tested. The
point of a stress test is to break the heuristic and fix it: the first run scored 2 of 5, and each miss
traced to a real source-line bug. Two were fixed the same day (self-attribution: 2 of 5 → 4 of 5);
the third is a defensible system classification, documented below. A representative-corpus benchmark on
everyday user apps is pending, like the foreground-exemption test in section 3: that will be
the headline accuracy number, and this page will publish it when it is real.
| Case | Ground truth | Predicted (after fix) | Result |
|---|---|---|---|
| Claude Code CLI session | The claude binary itself (stream-json protocol) | kind=claude_session | Match |
| macOS system daemon | /usr/libexec/logd | kind=system, label "logd" | Match |
| Hush's own engine process | cpu_lens.py --serve, run via the Xcode-bundled Python3 framework | label "Hush Performance" (was "Xcode") | Match |
| Hush's own bundled MCP server | Hush/mcp-venv/bin/python .../mcp/server.py, launched behind Claude Desktop's disclaimer wrapper | label "Hush Performance" (was "Claude") | Match |
A running .app under /System/ (loginwindow) | A literal .app bundle | kind=system, not app | By design |
What the stress test found, and the fix:
disclaimer wrapper, so the launcher's
.app path (Xcode, Claude) came first in the command line and won.
_app_name() now prefers the app that OWNS the running SCRIPT
(.py/.js/… inside an .app
bundle) over the launcher whose path merely appears first. Both now resolve to "Hush Performance." The fix
fires only when a bundled script is present, so ordinary apps and browser helpers are unchanged (verified
against a Chromium helper, a native app, a plain binary, and the simulator)./System/-rooted bundle labeled "system":
loginwindow lives under /System/
and is a system-owned process; classifying it "system" rather than "app" is defensible, so it is left as-is and
reported here as a defensible classification rather than counted as a fix.Measured 2026-07-24, Hush 0.95.1 (re-measured after the fix). A case not present on a given run
is reported as "not present," never silently dropped or invented. Full per-case detail including raw command
lines:
attribution_accuracy.json.
Reproduce: python3 bench/attribution_accuracy.py.
The claim under test: the app you are actively using is never eased, even while Hush is easing a background CPU storm.
Proving this claim needs an actual sustained multi-core CPU
storm to test against: there is nothing to exempt the foreground app from if the machine is
idle. Running that storm on the machine used for this benchmark pass would compete with real work
in real time, so it was not run here. The full test (drive a background storm, hold a
different app in the foreground, then check Hush's own action log for zero easing actions against
the foreground app and at least one against the storm) is written and documented in
bench/foreground_exemption.py, ready to run on a
dedicated load rig. This section will carry real numbers once that run happens; no number is
published in its place.
bench/ at the root of the Hush repository, alongside a full
README. Reproduce any result yourself:
git clone https://github.com/belstone/hush-performance cd hush-performance python3 bench/self_overhead.py python3 bench/attribution_accuracy.pyBoth scripts are read-only against a running Hush engine: one polls its already-running local API, the other takes one independent process-list snapshot. Neither starts, stops, nor mutates anything.
This page is not a security or performance guarantee. It reports what was measured, when, and how, so the measurement can be checked and repeated.