Whistle Review: A 16.9 MB Speech-to-Text Model That Trades Size for Robustness
Sources
- Whistle release announcement (Cactus Compute blog)
- Whistle model card, Hugging Face
- Needle GitHub repository (engine source, Apache-2.0)
- Needle 3 weights and platform engines, Hugging Face
- Needle fine-tuning pricing (verified 2026-10-09)
- Hacker News discussion (706 points, October 2026)
- cactus-needle 3.1.3 on PyPI (wheel source inspected for this review)
- OpenAI Whisper
Cactus Compute released Whistle on October 2, 2026: a speech-to-text model that is one 16.9 MB file, runs on CPU, transcribes seven languages, and shares its container and C++ engine with Needle, the company's tiny-device automation model. The launch hit the top of Hacker News with 706 points and 143 comments. We installed the actual package on an Intel i5-1135G7, generated audio with espeak, and ran the model through transcription, streaming, keyword biasing, and the combined speech-to-tool-call pipeline to find out what 16.9 MB actually buys — and where it falls apart.
The short version: the binary is real and the engineering is genuinely clever — one stripped 1.6 MB shared library plus a 16.9 MB weight file renders speech to text with word timestamps and per-word probabilities, warm transcription of a 4.8-second clip in ~0.9 s, silence returns an empty transcript instead of an invented sentence, and identical output on repeated runs. But the accuracy story is narrower than the launch post suggests. On synthetic espeak audio — which is, admittedly, not the model's target distribution — Whistle garbled every one of our technical-English clips without keyword biasing, mangled Spanish badly, and the combined audio-to-tool-call pipeline hallucinated a thermostat call from a garbled transcript. The streaming API has a quadratic re-processing cost that turns a 60-second session into 3.7 minutes of compute on this hardware. Keyword biasing works exactly as documented and is the difference between unusable and usable on in-vocabulary speech — which is both the model's best feature and its biggest dependency.
Executive Scorecard
| Dimension | Score | Why |
|---|---|---|
| Reliability | 7/10 | Deterministic, silence-safe, hard 30 s guard raises a clean error instead of truncating. Off-distribution audio degrades to confident nonsense. |
| DX | 6/10 | One pip install, clean Python API, word timestamps and probabilities are excellent. No depth control from Python, no thread-safety, engine fetch adds a Hugging Face dependency chain. |
| Cost | 9/10 | MIT-adjacent freedom: Apache-2.0 engine, free local inference, free local fine-tuning. Hosted tuning is $19 once or $99/month — cheap, and avoidable. |
| Security | 7/10 | Audio never leaves the device — real, verified in source. But telemetry is on by default (anonymous counts, opt-out env var), and the wheel ships a Supabase endpoint hard-coded in source. |
| Accuracy (self-tested) | 4/10 | Garbled technical English and Spanish on espeak synthetic audio without keywords; German mostly survived; keyword biasing rescued technical terms. |
Verdict up front, for readers who scroll: if you need a fully-offline, tiny-footprint transcriber for a fixed vocabulary — smart-home commands, product names, form-filling in kiosks — Whistle plus keyword biasing is a legitimate engineering win and you should evaluate it. If you need general-purpose transcription of accented, conversational, or technical speech, it is not competitive with Whisper-family or Moonshine models today, and the launch benchmarks (which are honest about their test sets, to the authors' credit) will not tell you that.
Architecture Mechanics
Whistle is not a general LLM with a microphone jack. It is a purpose-built encoder-decoder where the decoder is literally Needle's block list with a speech-specific cross-attention added per layer. The whole flow, as documented and as reflected in the shipped Python surface:
mic ──► 16 kHz mono PCM (max 30 s, 480k samples)
│
▼
log-mel front end 80 bins · 25 ms window · 10 ms hop
│ band-limited 250–3500 Hz, per-channel norm
▼
conv stem 128 ch · kernel 9 · 3 halvings
│ 3000 frames ──► 375 frames (1 per 80 ms)
▼
ENCODER 8× Simple Attention blocks (shared with Needle)
│ non-causal: frame at 3 s attends to frame at 12 s
│ Monarch Hadamard MLP replaces FFN
▼
K,V projections once per clip: 375 frames × 8 layers
│ (held for the entire decode — 5 beams share them)
▼
DECODER 8× Laddered Simple Attention blocks (Needle blocks)
│ GQA 8q:2kv · 3-tap causal conv on Q,K,V
│ gated cross-attn every layer reads the encoder
▼
beam search 5 beams · length-normalised log-prob
│ keyword bias via Aho-Corasick automaton
▼
transcript ≤320 tokens · 8,192-piece vocab + 7 lang tokens
word times from decoder's own attentionThe design decisions that matter operationally:
- Cross-attention K/V are computed once and cached for the whole decode. In the shipped
needle/agent/whistle.py, the C API surface is tiny —needle_load,needle_transcribe,needle_embed— and the engine holds one model per process, loaded vianeedle_load(data, len)from a buffer the Python side reads fully into RAM. Five beams therefore cost five short transcript caches, not five passes over the audio. - The ladder is decoder-side only. Every depth from 2 layers up is a deployable model, selectable at load time with the C API/CLI
--audio-depthflag. The encoder always runs all eight blocks. Notably, the PythonWhistle.transcribe()surface exposes no depth parameter — the depth ladder is reachable only through the CLI or the C API, a gap worth knowing before you plan a latency/accuracy sweep from Python. - Language detection is a token, not a side channel. The transcript vocabulary is 8,192 text pieces plus seven language tokens, one per supported language, so the detected language is emitted inside the transcript stream. Our Spanish clip returned
"language": "es"with the language token consumed as part of the decode, and forcinglanguage="es"changed nothing visible about the (bad) output. - Silence is a pre-decode gate, and it is honest. The engine measures the clip's loudness range before the decoder starts; below threshold it returns empty text, empty language, 0 ms TTFT, without entering beam search. We confirmed this on a 2-second silent clip:
{"text": "", "language": "", "words": [], "ttft_ms": 0.0}in 21 ms wall. This is the right design — many small ASR models invent a sentence from noise. - Keyword biasing is an Aho-Corasick automaton walked alongside the beams. This is not a prompt. It structurally lifts the log-probability of in-vocabulary phrases as the automaton advances. It is the feature that makes the model usable on names and jargon, and its limits are quantified below.
Hands-On: What 16.9 MB Actually Is
Installing the package is a standard pip install cactus-needle. On this host, the sandbox blocked the package manager's threat-intelligence scan, so we fetched the wheel directly from PyPI and inspected and extracted it manually — a detour that turned out to be a feature: the wheel ships unminified Python, which is where several findings below come from.
What lands on disk after first use, measured on this machine:
$ du -sh ~/.cache/cactus-needle/*
16M /home/pm/.cache/cactus-needle/whistle/2.0.0 # whistle.cact (16.9 MB)
37M /home/pm/.cache/cactus-needle/v3/3.2.0 # libneedle.so + needle3.cact
$ file ~/.cache/cactus-needle/v3/3.2.0/libneedle.so
ELF 64-bit LSB shared object, x86-64, stripped
$ ls -la ~/.cache/cactus-needle/whistle/2.0.0/
-rw-rw-r-- 1 pm pm 16,919,407 whistle.cactThe "16.9 MB model" claim is honest — whistle.cact is a 16,919,407-byte file — but it is the weights only. The engine is a separate 1.6 MB shared library, and if you also run Needle's tool-calling you add a 35 MB needle3.cact. The marketing framing "one 16.9 MB file, no dependencies" is true for the model artifact and quietly generous about the rest of the stack: the Python path pulls in huggingface_hub, httpx2, httpcore2, hf-xet, and friends to fetch that file on first call. The deploy story for constrained targets is different — the platform folders ship a standalone binary — but the Python experience has a real dependency chain.
Basic transcription, measured across our synthetic battery on the i5-1135G7 (4 cores, no GPU):
| Clip (espeak-synthesized) | Duration | TTFT | Decode | Wall |
|---|---|---|---|---|
| Clear English, smart-home command | 4.8 s | 213 ms | 54 tok/s | 698 ms |
| Technical English (K8s/OTel words) | 6.1 s | 3,241 ms | 28.7 tok/s | 4,886 ms |
| Fast English (230 wpm) | 3.8 s | 580 ms | 27.8 tok/s | 1,888 ms |
| Spanish command | 3.3 s | 326 ms | 36.1 tok/s | 981 ms |
| German command | 5.6 s | 3,327 ms | 29.2 tok/s | 4,616 ms |
| Silence (2 s) | 2.0 s | 0 ms | n/a | 21 ms |
The first thing to say about these numbers: they are not the launch numbers, and the gap is instructive. Cactus reports 11.1 ms TTFT and 1,319 tokens/s decode on an Apple M4 Pro. We measured 213–3,327 ms TTFT and 27–54 tok/s on this i5. Some of that is a slower CPU, and some of it is that TTFT grows with garbling — the beams wander when the acoustics are off-distribution, which both slows the decode and produces nonsense. The clean English clip decoded at 54 tok/s; the clip the model could not parse ran at 28.7 tok/s with a 3.2-second first token. Latency and accuracy are not independent in this model: when Whistle is lost, it is also slow.
The second thing: espeak audio is adversarial for a model trained on human speech, and we report it as such. Our clips are a robotic voice at 155–230 wpm — a worst-case input that no ASR vendor benchmarks against. But that is precisely the point of testing: the failure modes below are structural (keyword dependence, error propagation into tool calls, quadratic streaming), and they reproduce on any off-distribution input, whether that is synthetic speech, a heavy accent, or a kitchen robot. HN commenters independently reported the same class of failures on human speech — Spanish transcription "writing non-existing words", mumbly English dropping words, a Madrid speaker needing dictation-rhythm delivery to be understood.
Keyword Biasing: The Load-Bearing Feature
Whistle's answer to out-of-vocabulary speech is keyword biasing — pass the words your users actually say, and an Aho-Corasick automaton lifts their log-probability during beam search. This is the feature that separates "demo" from "deployable" for a 16.9 MB model, and our tests show both its power and its ceiling:
# no keywords
$ needle.transcribe("tech_en.wav")
"Lift Lloyd the open to limit recollector to the Cuban eater's cluster
and restart their post to the rescue elbow."
# keywords = ["OpenTelemetry", "Kubernetes", "PostgreSQL"]
$ needle.transcribe("tech_en.wav", keywords=[...])
"Lift Lloyd the open to limit recollector to the Kubernetes cluster
and restart the PostgreSQL pod."Two of three technical terms were rescued exactly (Kubernetes, PostgreSQL) with zero false positives injected. The third (OpenTelemetry) stayed garbled — biasing lifts candidates but cannot make the encoder hear a token it has no acoustic handle on. Escalating the keyword list quantified the cost model: 0 keywords → 216 ms TTFT / 52 tok/s; 1 keyword → 303 ms / 39 tok/s; 5 → 327 ms / 25 tok/s; 20 → 400 ms / 17 tok/s. The automaton is linear-ish but not free: every keyword you add taxes every decode, and at 20 keywords throughput has fallen to a third of the unbiasised rate.
For a platform team, the operational reading is this: Whistle without a curated keyword list is not a transcription system — it is a component that requires one. The blog's own framing ("the names, places and product words your users actually say") is accurate: this is a closed-vocabulary assistant front-end, not a general transcriber. If your vocabulary is enumerable — smart-home devices, menu items, product SKUs, a support-script decision tree — the keyword system is excellent. If your users speak freely, you are back to the WER charts.
The 30-Second Wall and the Streaming Problem
Whistle transcribes at most 30 seconds in one pass. Feed it more and it does not truncate — it raises:
RuntimeError: audio limit is 30 sA hard, named error is better than silent truncation — credit where due. The documented answer for longer input is needle.stream(chunks): feed 16 kHz float chunks of about a second, receive committed text plus a pending tail, with no limit on total length. The wheel source explains the mechanics: each chunk is processed by needle_stream_transcribe_process against the accumulated stream state, and text commits only where two passes agree.
We streamed the first 10 seconds of our 60-second clip as ten 1-second chunks. It took 36.7 seconds of wall time. The per-chunk cost grew monotonically and violently:
chunk pass_ms: 230 350 513 823 1528 3785 8143 6082 5850 9155
└── quadratic growth: each chunk re-scans the whole stream
10 s of audio → 36.7 s wall on an i5-1135G7This is a quadratic re-processing pattern: chunk N's cost scales with the accumulated audio behind it. On this hardware, real-time factor crosses 1.0 somewhere around chunk 5–6 — meaning live streaming on a comparable x86 CPU falls behind the microphone after about five seconds and never catches up. The committed text was also rough ("The quick brown fox dumps over the razion…"), with the two-pass commit mechanism not saving it on robotic input. To be fair to Cactus: the streaming API is explicitly for live microphone use on the devices this targets — phones and wearables with modern ARM cores, where per-chunk costs will be lower in absolute terms. But the growth curve is structural, not a CPU-speed artifact, and any team planning "stream for as long as the microphone runs" on commodity x86 should budget accordingly. For long files, the pattern is chunked batch transcription with overlap — VAD-chop the audio, transcribe 25–30 s windows, stitch — which is on you to build; the SDK ships no chunker.
The Combined Pipeline: Where ASR Errors Become Tool Calls
The launch's most interesting claim is the combined pipeline: load both models into one engine, hand it audio plus a tool schema, and get tool calls back with the transcript never leaving the engine. The CLI path is needle --model needle3.cact --model whistle.cact --tools tools.json --audio clip.wav. From Python, the honest composite is transcribe-then-route, which is what we tested. First, the text-only baseline, because Needle's tool-calling is the other half of this product:
query: "turn off the kitchen lights and set the thermostat to 21 degrees"
{
"function_calls": [],
"results": [
{"room": "kitchen", "on": false},
{"degrees": 21}
],
"reasoning": "Both done; respond with confirmation",
"confidence": 0.6876,
"peak_ram_mb": 144.2
}Two tools, two correct calls, executed and returned in one response envelope with a calibrated confidence. Off-topic input ("what is the capital of France") returned an empty call list and confidence 0.31 — the refusal behaviour the README promises ("an empty list rather than a guess") held up on text. Warm runs complete in 2–3 s on this CPU, peak RAM ~144 MB, and the engine reports its own peak_ram_mb in every response, which is a genuinely useful touch for target-device budgeting.
Now the same agent, driven by Whistle's transcription of our clear-English clip — the clip whose first half transcribed perfectly and whose second half garbled:
Whisper-free composite: needle.transcribe() → agent.run()
transcription: "Turn off the kitchen lights and set the firmest
attitude when did one librees."
tool calls: [{"name": "set_thermostat", "arguments": {"degrees": 180}}]This is the review's most important finding. Garbled ASR output did not produce a refusal or an empty call list — it produced a confident, wrong tool call. The model heard "firmest attitude when did one librees", matched "one" to the degrees argument of the thermostat tool (the only numeric slot available), and executed set_thermostat(degrees=180). No confirmation gate, no low-confidence suppression: Needle's calibrated confidence exists (0.69 on the clean text run) but nothing in the shipped Python path consults it before acting on transcribed input. The engine's keyword-bias integration is the intended mitigation — the docs say the combined path "favours the option values your schemas enumerate" — but our Python composite bypasses that integration because the Python API surface does not expose the single-call audio+tools path; only the CLI/C API does. A team wiring speech to actions from Python today is assembling exactly the pipeline we tested, and it will act on garbage unless they add their own confidence threshold, confirmation step, or keyword-biased transcription.
The failure has a shape platform engineers will recognize: error propagation across a compound AI pipeline with no intermediate validation. Whistle's word-level probabilities (visible with word_timestamps=True — the garbled words carried 0.25–0.45 probabilities while the clean words carried 0.89–0.98) are exactly the signal a gate would need, and the shipped API returns them. The components to build the safety layer exist; the composition does not ship one.
Critical Failure Modes
- Off-distribution collapse. Every technical clip garbled without keywords; Spanish collapsed into non-words. The launch WER table is honest (LibriSpeech, SPGISpeech, Earnings-22, FLEURS averages, with Whisper ahead on TED-LIUM, AMI, and MLS) but a WER average over clean corpora conceals the cliff: inputs outside the training distribution do not degrade gracefully, they fall off a cliff into fluent-looking nonsense at 0.4 word probability.
- Confident tool calls from garbage transcripts. As above: 180 degrees from "one librees". Compound pipelines need their own validation layer; none ships in the Python path.
- Quadratic streaming. Per-chunk cost grows with accumulated stream length; real-time factor crosses 1.0 around chunk 5–6 on this x86 host. Long sessions need external chunking, which the SDK does not provide.
- Single engine, no concurrency.
needle_loadswaps the one process-global model; the wheel source is explicit that one model is loaded per process and it is not thread-safe. Our concurrency probe: a second thread constructing its ownWhistle()while the first transcribes did not return within a 120-second join — the second construction blocks on the shared handle. Multi-tenant or multi-stream serving means process-per-stream, which for a 16.9 MB model is at least cheap to fork. - Keyword tax. 20 keywords cut decode throughput to a third (52 → 17 tok/s) and added ~190 ms TTFT. Large product vocabularies directly erode the latency that motivated the model.
- Telemetry on by default.
_telemetry.pysends event name, versions, OS/arch, and a random install ID to a hard-coded Supabase endpoint, fire-and-forget on a daemon thread. It is disclosed in a first-run stderr notice and README, disables underNEEDLE_TELEMETRY=0,DO_NOT_TRACK, orCI, and never carries prompts or audio — the source is clean on that. But "on-device, private" products that phone home counts by default should expect scrutiny, especially embedded/regulated deployments. Set the env var in your base image. - Python surface gaps. No
--audio-depthequivalent from Python (depth runs are CLI/C-API only), no chunked-file helper, no audio+tools single-call from Python. The richest surface is the C API and CLI, not the SDK your team will actually write against.
Benchmarks vs. Reality, and the Commercial Shape
The published comparison — Whistle at 16.9 MB vs Whisper base at 145.3 MB vs Moonshine tiny v2 at 41.9 MB, with Whistle ahead on LibriSpeech clean/other, SPGISpeech, Earnings-22, and FLEURS average, behind on TED-LIUM, AMI, and MLS — is unusually well-documented for a launch post. The caveats are printed on the page: per-benchmark conventions, which figures are the other authors' own, that Whisper's AMI row is a different subset. This is the right way to publish benchmarks, and it contrasts with the WER charts most vendors ship. Our complaint is not with the charts; it is with the distance between benchmark distribution and your distribution. Robotic, accented, mumbly, or jargon-heavy speech is where 16.9 MB shows its ceiling, and none of the four headlined corpora probe it.
Commercially the model is Apache-2.0, weights on Hugging Face, engine source on GitHub — no API revenue dependency, and inference is free forever on your hardware. The monetization is hosted fine-tuning: $19 for a one-shot Starter run (30 days access, 3 runs, 3,000 generated examples), $99/month for Pro (10 runs/period, 10,000 examples), custom above that, with local LoRA fine-tuning free and unlimited via needle finetune. For a team tuning Whistle/Needle on a fixed product vocabulary — the exact use case this model is built for — a $19 one-time experiment is among the cheapest legitimate entry points in the edge-AI tooling market, and the local path means the paid tier is a convenience, not a tollbooth. That is the right commercial shape for an open model, and it deserves saying plainly.
Who Should Skip This
- Teams needing general transcription. Meeting notes, call-center audio, podcast pipelines, dictation of free-form prose: use Whisper (large-v3 or distilled variants) or a hosted ASR. Whistle's seven-language support is real but thin in the tail, and the WER gap on conversational corpora (AMI, MLS) is the vendor's own published admission.
- Anyone transcribing languages outside the seven. There is no code-switching story, no open-vocabulary language ID beyond the seven tokens, and no fallback. If your users speak anything else — or mix languages mid-sentence — this is not your tool.
- Long-audio batch pipelines. The 30-second wall plus quadratic streaming means you will build and maintain the chunker yourself. If your audio is hours long and your hardware is x86, the total cost of ownership favours a model with native long-form handling.
- High-concurrency serving. Process-global engine, no thread-safety, one stream per process. A transcription microservice fielding parallel requests needs a process pool around this, and at that point the 16.9 MB advantage is competing against server-scale models on their own turf.
- Teams that cannot curate keywords. Without a maintained vocabulary, accuracy on names, products, and jargon collapses — and maintaining that vocabulary is an ongoing content-engineering cost the launch post does not price.
Who it is for: embedded and robotics teams adding voice control to constrained hardware (the 17 prebuilt targets including RISC-V, MIPS, watchOS, and a WASI component are the real product here); platform teams building fixed-vocabulary voice front-ends to internal tools where the keyword list is enumerable and the tool schema is closed; and anyone who needs transcription that verifiably never leaves the device, where the alternative is no transcription at all. In those lanes, Whistle's trade — general capacity for size — is the correct trade, and the execution (one binary, honest silence handling, word probabilities, calibrated confidence on the tool side) is better than the edge-ASR average.
Verdict
Whistle is a serious piece of systems engineering wrapped in a modest amount of marketing gloss. The 16.9 MB artifact is real, the shared-engine design with Needle is genuinely elegant, the benchmark publication is more honest than the industry norm, and the Apache-2.0-plus-$19-fine-tune commercial shape is respectful of the user. But our hands-on battery found the model's usable envelope to be much narrower than the launch implies: it is a keyword-biased, closed-vocabulary, ≤30-second, single-stream component — a very good one — rather than a general speech-to-text engine. The most dangerous property is not the WER; it is that errors stay confident all the way through the combined pipeline and become tool calls unless you build the gate yourself. Adopt it for enumerable-vocabulary voice control on constrained devices, with a curated keyword list, a confidence threshold on actions, and external chunking for anything long. Skip it for anything that transcribes the open world.
References & Further Reading
- Whistle: Speech to Text in 16.9 MB — the release announcement with architecture, benchmarks, and deployment targets (Cactus Compute blog).
- Cactus-Compute/whistle on Hugging Face — model card with per-benchmark caveats, the whistle.cact weights, and the speech C API.
- cactus-compute/needle on GitHub — engine source, Apache-2.0, including the whistle module and CLI used in this review's tests.
- Cactus-Compute/needle3 on Hugging Face — Needle 3 weights, platform engine folders, and the tool-calling benchmarks.
- Needle fine-tuning pricing — free local training, $19 Starter, $99/month Pro (verified 2026-10-09).
- The .cact format — Cactus Quants at 2.125 bits per weight, and how to parse the container yourself.
- Hacker News discussion — 706-point launch thread; independent reports of Spanish and mumble-mode failures cited in this review.
- cactus-needle on PyPI — the 3.1.3 wheel whose source (telemetry module, whistle module, CLI) was inspected for this review.